# Aiven docs > Official technical documentation for Aiven, the AI-ready Open Source Data Platform. Comprehensive guides for deploying, scaling, and monitoring managed services including Apache Kafka, PostgreSQL, MySQL, OpenSearch, Redis, ClickHouse, and Flink. Covers platform features like Bring Your Own Cloud (BYOC), VPC peering, and authentication. Includes reference documentation for Aiven Developer Tools: CLI, API, Terraform Provider, and Kubernetes Operator. ## docs Aiven docs - Your AI-ready Open Source Data Platform. One unified platform to stream, store, and serve data on any cloud. - [Aiven docs](/index.md): Aiven docs - Your AI-ready Open Source Data Platform. One unified platform to stream, store, and serve data on any cloud. ### ai-features Connect AI coding assistants and agents to Aiven services with MCP servers and Skills, and find AI capabilities built into Aiven services. - [AI tools on Aiven](/ai-features.md): Connect AI coding assistants and agents to Aiven services with MCP servers and Skills, and find AI capabilities built into Aiven services. ### get-started Aiven provides managed open source services for streaming and databases across major cloud providers. - [Get started with Aiven](/get-started.md): Aiven provides managed open source services for streaming and databases across major cloud providers. ### integrations #### cloudlogging You can send your service logs to Google Cloud Logging to store, search, analyze, monitor, and alert on log data from your Aiven services. - [Google Cloud Logging](/integrations/cloudlogging.md): You can send your service logs to Google Cloud Logging to store, search, analyze, monitor, and alert on log data from your Aiven services. #### cloudwatch Amazon CloudWatch (AWS) is an AWS - [Amazon CloudWatch and Aiven](/integrations/cloudwatch.md): Amazon CloudWatch (AWS) is an AWS ##### cloudwatch-logs-cli Send logs from your Aiven service to the AWS CloudWatch using the Aiven client. - [Send logs to AWS CloudWatch from Aiven client](/integrations/cloudwatch/cloudwatch-logs-cli.md): Send logs from your Aiven service to the AWS CloudWatch using the Aiven client. ##### cloudwatch-logs-console Send your Aiven service logs to the AWS CloudWatch using the Aiven Console. - [Send logs to AWS CloudWatch from Aiven Console](/integrations/cloudwatch/cloudwatch-logs-console.md): Send your Aiven service logs to the AWS CloudWatch using the Aiven Console. ##### cloudwatch-metrics Aiven enables you to send your service metrics to your Amazon (AWS) CloudWatch. - [Send metrics to Amazon CloudWatch](/integrations/cloudwatch/cloudwatch-metrics.md): Aiven enables you to send your service metrics to your Amazon (AWS) CloudWatch. #### datadog Datadog is a monitoring platform, allowing - [Datadog and Aiven](/integrations/datadog.md): Datadog is a monitoring platform, allowing ##### add-custom-tags-to-datadog When using the Datadog integration in the Aiven Console, Aiven automatically includes a set of standard tags in all data sent to Datadog. - [Add custom tags Datadog integration](/integrations/datadog/add-custom-tags-to-datadog.md): When using the Datadog integration in the Aiven Console, Aiven automatically includes a set of standard tags in all data sent to Datadog. ##### datadog-logs Use the Aiven Rsyslog integration to send logs from your Aiven services to your external Datadog account. - [Send logs to Datadog](/integrations/datadog/datadog-logs.md): Use the Aiven Rsyslog integration to send logs from your Aiven services to your external Datadog account. ##### datadog-metrics Send metrics from your Aiven service to your external Datadog account. - [Send metrics to Datadog](/integrations/datadog/datadog-metrics.md): Send metrics from your Aiven service to your external Datadog account. #### prometheus-system-metrics Learn how to check what metrics are available for monitoring your service using Prometheus, and find out which of the available metrics are particularly worth monitoring and why. - [Prometheus system metrics](/integrations/prometheus-system-metrics.md): Learn how to check what metrics are available for monitoring your service using Prometheus, and find out which of the available metrics are particularly worth monitoring and why. #### rsyslog In addition to using Aiven for OpenSearch® to store the logs from your Aiven services, you can also integrate with an external monitoring system that supports the rsyslog protocol. - [Remote syslog integration](/integrations/rsyslog.md): In addition to using Aiven for OpenSearch® to store the logs from your Aiven services, you can also integrate with an external monitoring system that supports the rsyslog protocol. ##### loggly Aiven supports integrating logs with a number of external monitoring - [Log integration with Loggly](/integrations/rsyslog/loggly.md): Aiven supports integrating logs with a number of external monitoring ##### logtail Logtail is a logging service. You can use the Aiven Remote syslog integration to send your logs to Logtail. - [Send Aiven logs to Logtail](/integrations/rsyslog/logtail.md): Logtail is a logging service. You can use the Aiven Remote syslog integration to send your logs to Logtail. #### send-logs-to-elasticsearch You can store logs from one of your Aiven services in an external Elasticsearch service. - [Send logs to Elasticsearch®](/integrations/send-logs-to-elasticsearch.md): You can store logs from one of your Aiven services in an external Elasticsearch service. ### marketplace-setup Create an AWS, Azure, or Google Cloud marketplace subscription to use as a payment method for Aiven services. - [Set up marketplace subscriptions](/marketplace-setup.md): Create an AWS, Azure, or Google Cloud marketplace subscription to use as a payment method for Aiven services. ### platform #### concepts ##### aiven-node-firewall-configuration Aiven nodes are built using Linux. Firewall configuration is managed using native Linux kernel-level iptables rules that limit connectivity to nodes. - [Firewall configuration for service nodes](/platform/concepts/aiven-node-firewall-configuration.md): Aiven nodes are built using Linux. Firewall configuration is managed using native Linux kernel-level iptables rules that limit connectivity to nodes. ##### application-users An application user is a type of user that provides programmatic access to the Aiven platform and services through the Aiven API, CLI, Aiven Terraform Provider, and Aiven Kubernetes Operator. They're intended for non-human users that need to access Aiven. - [Application users](/platform/concepts/application-users.md): An application user is a type of user that provides programmatic access to the Aiven platform and services through the Aiven API, CLI, Aiven Terraform Provider, and Aiven Kubernetes Operator. They're intended for non-human users that need to access Aiven. ##### authentication-tokens There are 3 types of tokens used to access the Aiven platform: session tokens, personal tokens, and application tokens. - [Tokens](/platform/concepts/authentication-tokens.md): There are 3 types of tokens used to access the Aiven platform: session tokens, personal tokens, and application tokens. ##### availability-zones Availability zones (AZs) are physically isolated locations (data centers) where cloud services operate. - [Availability zones](/platform/concepts/availability-zones.md): Availability zones (AZs) are physically isolated locations (data centers) where cloud services operate. ##### backup-to-another-region Copy service backups to a secondary region for disaster recovery, and manage, fork from, or delete that secondary backup. - [Backup to another region for Aiven services Limited availability](/platform/concepts/backup-to-another-region.md): Copy service backups to a secondary region for disaster recovery, and manage, fork from, or delete that secondary backup. ##### billing-and-payment Billing on the Aiven Platform is managed through billing groups. - [Billing and payment](/platform/concepts/billing-and-payment.md): Billing on the Aiven Platform is managed through billing groups. ##### byoc Bring your own cloud (BYOC) allows you to use your own cloud infrastructure instead of relying on the Aiven-managed infrastructure. - [Bring your own cloud (BYOC)](/platform/concepts/byoc.md): Bring your own cloud (BYOC) allows you to use your own cloud infrastructure instead of relying on the Aiven-managed infrastructure. ##### byoc-enhanced-compliance Enhanced compliance clouds are - [Enhanced compliance BYOC clouds](/platform/concepts/byoc-enhanced-compliance.md): Enhanced compliance clouds are ##### carbon-footprint The carbon footprint page shows estimates of the greenhouse gas emissions associated with service usage in your organization to help you monitor and reduce your emissions. - [Track your carbon footprint](/platform/concepts/carbon-footprint.md): The carbon footprint page shows estimates of the greenhouse gas emissions associated with service usage in your organization to help you monitor and reduce your emissions. ##### cloud-security Learn about Aiven's access control, encryption, network security, data privacy and operator access. - [Cloud security](/platform/concepts/cloud-security.md): Learn about Aiven's access control, encryption, network security, data privacy and operator access. ##### disaster-recovery-test-scenarios Aiven provides disaster recovery testing services to help you plan for data center or service outages. - [Disaster recovery testing](/platform/concepts/disaster-recovery-test-scenarios.md): Aiven provides disaster recovery testing services to help you plan for data center or service outages. ##### discovered-organizations The discovered organizations feature lets organization admin identify other Aiven organizations that managed users have joined using the same email address associated with one of your verified domains. - [Discovered organizations](/platform/concepts/discovered-organizations.md): The discovered organizations feature lets organization admin identify other Aiven organizations that managed users have joined using the same email address associated with one of your verified domains. ##### enhanced-compliance-env Aiven collects, manages, and operates on sensitive data that is protected by privacy and compliance rules and regulations. Aiven meets the needs of its customers by providing specialized enhanced compliance environments (ECE) that comply with many of the most common compliance requirements. - [Enhanced compliance environments (ECE)](/platform/concepts/enhanced-compliance-env.md): Aiven collects, manages, and operates on sensitive data that is protected by privacy and compliance rules and regulations. Aiven meets the needs of its customers by providing specialized enhanced compliance environments (ECE) that comply with many of the most common compliance requirements. ##### maintenance-window Aiven applies maintenance updates automatically during a maintenance window that you choose for each service. - [Service maintenance, updates and upgrades](/platform/concepts/maintenance-window.md): Aiven applies maintenance updates automatically during a maintenance window that you choose for each service. ##### managed-users The managed users feature lets you centrally manage your organization's users and helps you to secure your organization in Aiven. - [Managed users](/platform/concepts/managed-users.md): The managed users feature lets you centrally manage your organization's users and helps you to secure your organization in Aiven. ##### orgs-units-projects The Aiven Platform uses organizations, organizational units, and projects to efficiently and securely organize your services and manage access. - [Organizations, units, and projects](/platform/concepts/orgs-units-projects.md): The Aiven Platform uses organizations, organizational units, and projects to efficiently and securely organize your services and manage access. ##### out-of-memory-conditions Understand how the Linux out-of-memory killer works, why it can affect a service, and how to avoid triggering it. - [Out of memory conditions](/platform/concepts/out-of-memory-conditions.md): Understand how the Linux out-of-memory killer works, why it can affect a service, and how to avoid triggering it. ##### permissions To give users access to projects and services in your organizations, you grant them permissions and roles: - [Roles and permissions](/platform/concepts/permissions.md): To give users access to projects and services in your organizations, you grant them permissions and roles: ##### rename-services Change the name of an Aiven service by forking it under a new name and deleting the original service. - [Rename a service](/platform/concepts/rename-services.md): Change the name of an Aiven service by forking it under a new name and deleting the original service. ##### service_backups Learn how Aiven backs up your services automatically, where backups are stored, and how to change your backup schedule. - [Service backups](/platform/concepts/service_backups.md): Learn how Aiven backs up your services automatically, where backups are stored, and how to change your backup schedule. ##### service-and-feature-releases New services and features follow a release lifecycle that promotes quality by giving customers the chance to test them and provide feedback. - [Service and feature releases](/platform/concepts/service-and-feature-releases.md): New services and features follow a release lifecycle that promotes quality by giving customers the chance to test them and provide feedback. ##### service-forking Fork an Aiven service to create an independent copy for testing, debugging, or development without affecting the original service. - [Fork a service](/platform/concepts/service-forking.md): Fork an Aiven service to create an independent copy for testing, debugging, or development without affecting the original service. ##### service-integration Service integrations provide additional functionality and features by connecting different Aiven services together. - [Service integrations overview](/platform/concepts/service-integration.md): Service integrations provide additional functionality and features by connecting different Aiven services together. ##### service-memory-limits Understand how Aiven limits the memory available to a service, and how that limit relates to the physical memory of the underlying node. - [Service memory limits](/platform/concepts/service-memory-limits.md): Understand how Aiven limits the memory available to a service, and how that limit relates to the physical memory of the underlying node. ##### service-power-cycle Power off an Aiven service to release resources and save credits, power it back on when you need it, or delete it permanently. - [Power on/off a service](/platform/concepts/service-power-cycle.md): Power off an Aiven service to release resources and save credits, power it back on when you need it, or delete it permanently. ##### service-pricing All Aiven services are billed based on actual usage so you only pay for the resources you use. - [Service pricing](/platform/concepts/service-pricing.md): All Aiven services are billed based on actual usage so you only pay for the resources you use. ##### static-ips Aiven services are normally addressed by their hostname, but static IP addresses are also available for an additional charge. - [Static IP addresses](/platform/concepts/static-ips.md): Aiven services are normally addressed by their hostname, but static IP addresses are also available for an additional charge. ##### tax-information Aiven services are provided by Aiven Ltd, a private limited company incorporated in Finland, but operating in various countries. None of Aiven's marketed prices include value-added or other taxes. - [Tax information for Aiven services](/platform/concepts/tax-information.md): Aiven services are provided by Aiven Ltd, a private limited company incorporated in Finland, but operating in various countries. None of Aiven's marketed prices include value-added or other taxes. ##### tls-ssl-certificates All traffic to Aiven services is always protected by TLS. It ensures that third parties can't eavesdrop or modify the data while in transit between Aiven services and the clients accessing them. - [TLS/SSL certificates](/platform/concepts/tls-ssl-certificates.md): All traffic to Aiven services is always protected by TLS. It ensures that third parties can't eavesdrop or modify the data while in transit between Aiven services and the clients accessing them. ##### user-access-management There are several types of users in the Aiven Platform: - [User and access management](/platform/concepts/user-access-management.md): There are several types of users in the Aiven Platform: ##### vpcs Virtual private clouds (VPCs) supported on the Aiven Platform provide enhanced security, flexibility, and network control, allowing efficient traffic, resource, and access management. - [Virtual private clouds (VPCs) in Aiven](/platform/concepts/vpcs.md): Virtual private clouds (VPCs) supported on the Aiven Platform provide enhanced security, flexibility, and network control, allowing efficient traffic, resource, and access management. #### howto ##### add-authentication-method You can authenticate directly with your email and password, or use single sign-on through providers like GitHub, Google, and Microsoft. - [Add authentication methods](/platform/howto/add-authentication-method.md): You can authenticate directly with your email and password, or use single sign-on through providers like GitHub, Google, and Microsoft. ##### add-storage-space Scale the disk storage of an Aiven service up or down without disrupting the running service. - [Scale disk storage](/platform/howto/add-storage-space.md): Scale the disk storage of an Aiven service up or down without disrupting the running service. ##### attach-vpc-aws-tgw AWS Transit Gateway (TGW) enables transitive routing from on-premises networks through VPN and from other VPC. - [Attach VPCs to AWS Transit Gateway](/platform/howto/attach-vpc-aws-tgw.md): AWS Transit Gateway (TGW) enables transitive routing from on-premises networks through VPN and from other VPC. ##### bring-your-own-key Register, list, update, or delete your customer managed keys (CMKs), associate CMKs with services, and view CMK usage across services in Aiven projects using the Aiven Provider for Terraform, Aiven API, or the Aiven CLI. - [Bring your own key (BYOK)](/platform/howto/bring-your-own-key.md): Register, list, update, or delete your customer managed keys (CMKs), associate CMKs with services, and view CMK usage across services in Aiven projects using the Aiven Provider for Terraform, Aiven API, or the Aiven CLI. ##### byoc ###### add-customer-info-custom-cloud Update the list of customer contacts for your custom cloud. - [Manage customer contacts for a custom cloud](/platform/howto/byoc/add-customer-info-custom-cloud.md): Update the list of customer contacts for your custom cloud. ###### assign-project-custom-cloud Select your organizations, units, or projects that can access and use your custom cloud. - [Enable a custom cloud in Aiven organizations, units, or projects](/platform/howto/byoc/assign-project-custom-cloud.md): Select your organizations, units, or projects that can access and use your custom cloud. ###### aws-privatelink-byoc Enable and manage AWS PrivateLink for your Aiven services deployed in your own cloud using bring your own cloud (BYOC). - [Use AWS PrivateLink with BYOC services Limited availability](/platform/howto/byoc/aws-privatelink-byoc.md): Enable and manage AWS PrivateLink for your Aiven services deployed in your own cloud using bring your own cloud (BYOC). ###### create-cloud - [Create an AWS-integrated custom cloud](/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md): Create a custom cloud for BYOC in your Aiven organization to better address your specific business needs or project requirements. - [Create a Microsoft Azure-integrated custom cloud](/platform/howto/byoc/create-cloud/create-azure-custom-cloud.md): Create a custom cloud for BYOC in your Aiven organization to better address your specific business needs or project requirements. - [Create a custom cloud](/platform/howto/byoc/create-cloud/create-custom-cloud.md): To create custom clouds in Aiven using self-service, select your cloud provider to integrate with. - [Create a Google-integrated custom cloud](/platform/howto/byoc/create-cloud/create-google-custom-cloud.md): Create a custom cloud for BYOC in your Aiven organization to better address your specific business needs or project requirements. ###### delete-custom-cloud Delete a custom cloud so that it's no longer available in your Aiven organization, units, or projects. - [Delete a custom cloud](/platform/howto/byoc/delete-custom-cloud.md): Delete a custom cloud so that it's no longer available in your Aiven organization, units, or projects. ###### download-infrastructure-template Download a Terraform template and a variables file that define the infrastructure of your - [Download an infrastructure template and a variables file](/platform/howto/byoc/download-infrastructure-template.md): Download a Terraform template and a variables file that define the infrastructure of your ###### enable-byoc Enabling the bring your own cloud (BYOC) feature allows you to create custom clouds in your Aiven organization. - [Enable bring your own cloud (BYOC)](/platform/howto/byoc/enable-byoc.md): Enabling the bring your own cloud (BYOC) feature allows you to create custom clouds in your Aiven organization. ###### manage-byoc-service Create a service in your custom cloud or migrate an existing service to your custom cloud. - [Manage services hosted in custom clouds](/platform/howto/byoc/manage-byoc-service.md): Create a service in your custom cloud or migrate an existing service to your custom cloud. ###### networking-security Aiven combines a multi-cloud strategy with a cloud-agnostic approach to make the bring your own cloud (BYOC) experience not only versatile and cost-efficient but also secure. - [Bring your own cloud networking and security](/platform/howto/byoc/networking-security.md): Aiven combines a multi-cloud strategy with a cloud-agnostic approach to make the bring your own cloud (BYOC) experience not only versatile and cost-efficient but also secure. ###### rename-custom-cloud Change the name of your custom cloud. - [Rename a custom cloud](/platform/howto/byoc/rename-custom-cloud.md): Change the name of your custom cloud. ###### store-data AWS BYOC environments use the tiered storage capability for data allocation. Cold data in your AWS custom cloud is stored in your AWS cloud account. - [Use tiered storage for AWS BYOC services Limited availability](/platform/howto/byoc/store-data.md): AWS BYOC environments use the tiered storage capability for data allocation. Cold data in your AWS custom cloud is stored in your AWS cloud account. ###### tag-custom-cloud-resources Tagging allows resource categorization, which simplifies governance, cost allocation, and system performance review. Custom cloud tags propagate to resources on the Aiven platform and in your own cloud infrastructure. - [Tag custom cloud resources](/platform/howto/byoc/tag-custom-cloud-resources.md): Tagging allows resource categorization, which simplifies governance, cost allocation, and system performance review. Custom cloud tags propagate to resources on the Aiven platform and in your own cloud infrastructure. ###### view-custom-cloud-status Find out whether your custom cloud is ready to use by viewing its status. - [View the status of a custom cloud](/platform/howto/byoc/view-custom-cloud-status.md): Find out whether your custom cloud is ready to use by viewing its status. ##### change-your-email-address You can't change your login email address directly. Instead, you can create a user with the new email address and remove the user with the old email address. - [Change your email address](/platform/howto/change-your-email-address.md): You can't change your login email address directly. Instead, you can create a user with the new email address and remove the user with the old email address. ##### configure-project-base-port The base port number for a project determines the Aiven service ports. The base port is randomly assigned, but you can also configure the port number for a project. - [Configure project base port](/platform/howto/configure-project-base-port.md): The base port number for a project determines the Aiven service ports. The base port is randomly assigned, but you can also configure the port number for a project. ##### controlled-upgrade Link Aiven services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. - [Control maintenance updates with upgrade pipelines Limited availability](/platform/howto/controlled-upgrade.md): Link Aiven services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. ##### create_authentication_token Create personal token in the Aiven Console to use with the Aiven CLI or API. - [Create personal tokens](/platform/howto/create_authentication_token.md): Create personal token in the Aiven Console to use with the Aiven CLI or API. ##### create_new_service Create an Aiven service from the Aiven Console. - [Create a service](/platform/howto/create_new_service.md): Create an Aiven service from the Aiven Console. ##### create_new_service_user Create additional users to access an Aiven service. - [Create service users](/platform/howto/create_new_service_user.md): Create additional users to access an Aiven service. ##### create-service-integration Create service integrations between different Aiven services and move telemetry data using these integrations. - [Create service integrations](/platform/howto/create-service-integration.md): Create service integrations between different Aiven services and move telemetry data using these integrations. ##### credits Use credits to pay for your Aiven services by adding them to a billing group. All projects assigned to the billing group automatically use its credits for the costs of the projects and the services within the projects. - [Use credits](/platform/howto/credits.md): Use credits to pay for your Aiven services by adding them to a billing group. All projects assigned to the billing group automatically use its credits for the costs of the projects and the services within the projects. ##### delete-user You can delete your personal user account as long as you are not a managed user of an organization. - [Delete user account](/platform/howto/delete-user.md): You can delete your personal user account as long as you are not a managed user of an organization. ##### disk-autoscaler Automatically increase the disk storage of an Aiven service when it's running out of space, instead of resizing it manually. - [Scale disk storage automatically](/platform/howto/disk-autoscaler.md): Automatically increase the disk storage of an Aiven service when it's running out of space, instead of resizing it manually. ##### download-invoices You can download monthly invoices from the Aiven Console in PDF or CSV format. - [Download invoices](/platform/howto/download-invoices.md): You can download monthly invoices from the Aiven Console in PDF or CSV format. ##### edit-user-profile You can edit your name, job title, location, or other personal details in the Aiven Console. - [Edit your user profile](/platform/howto/edit-user-profile.md): You can edit your name, job title, location, or other personal details in the Aiven Console. ##### feature-preview Before an official release, some features are available to our customers for testing. These feature previews let you try out upcoming enhancements and give our product teams feedback to help improve them. - [Feature previews](/platform/howto/feature-preview.md): Before an official release, some features are available to our customers for testing. These feature previews let you try out upcoming enhancements and give our product teams feedback to help improve them. ##### google-cloud-functions You can access Aiven service by creating a Serverless VPC access connector and Google Cloud Function. - [Access Aiven services from Google Cloud Functions via VPC peering](/platform/howto/google-cloud-functions.md): You can access Aiven service by creating a Serverless VPC access connector and Google Cloud Function. ##### integrations ###### access-jmx-metrics-jolokia Jolokia is one of the external metrics integrations supported on the Aiven platform, along with Datadog metrics and Prometheus metrics. - [Access JMX metrics via Jolokia](/platform/howto/integrations/access-jmx-metrics-jolokia.md): Jolokia is one of the external metrics integrations supported on the Aiven platform, along with Datadog metrics and Prometheus metrics. ###### datadog-increase-metrics-limit Monitoring services and applications are essential to know whether programs work as expected. To get started with monitoring, see Aiven and Datadog integration. - [Increase metrics limit setting for Datadog](/platform/howto/integrations/datadog-increase-metrics-limit.md): Monitoring services and applications are essential to know whether programs work as expected. To get started with monitoring, see Aiven and Datadog integration. ###### prometheus-metrics Discover Prometheus as a tool for monitoring your Aiven services. Check why use it and how it works. Learn how to enable and configure Prometheus on your project. - [Use Prometheus with Aiven](/platform/howto/integrations/prometheus-metrics.md): Discover Prometheus as a tool for monitoring your Aiven services. Check why use it and how it works. Learn how to enable and configure Prometheus on your project. ##### list-authentication Users can authenticate to the Aiven platform using a password, single sign-on (SSO), or tokens. The available authentication methods depend on the organization's authentication policy. - [Authentication methods](/platform/howto/list-authentication.md): Users can authenticate to the Aiven platform using a password, single sign-on (SSO), or tokens. The available authentication methods depend on the organization's authentication policy. ##### list-identity-providers Set up single sign-on (SSO) access to Aiven through a Security Assertion Markup Language (SAML) compliant identity provider (IdP). This lets you centrally manage your users in your IdP while giving them a seamless login experience. - [SAML identity providers and verified domains](/platform/howto/list-identity-providers.md): Set up single sign-on (SSO) access to Aiven through a Security Assertion Markup Language (SAML) compliant identity provider (IdP). This lets you centrally manage your users in your IdP while giving them a seamless login experience. ##### list-manage-vpc Create and manage a virtual private cloud (VPC) for your Aiven project or organization. - [Manage virtual private clouds (VPCs) in Aiven](/platform/howto/list-manage-vpc.md): Create and manage a virtual private cloud (VPC) for your Aiven project or organization. ##### list-marketplace-payments You can change the payment method for a billing group to a marketplace subscription. - [Use marketplace subscriptions to pay for Aiven services](/platform/howto/list-marketplace-payments.md): You can change the payment method for a billing group to a marketplace subscription. ##### list-monitoring Use metrics, logs, alerts, and dashboards to monitor the health of your services and integrations. - [Metrics, logs, and alerts](/platform/howto/list-monitoring.md): Use metrics, logs, alerts, and dashboards to monitor the health of your services and integrations. ##### list-organization-vpc-peering Establish a peering connection between your Aiven organization VPC and a cloud platform of - [Set up an organization VPC peering Limited availability](/platform/howto/list-organization-vpc-peering.md): Establish a peering connection between your Aiven organization VPC and a cloud platform of ##### list-project-vpc-peering Establish a peering connection between your Aiven project VPC and a cloud platform of your - [Set up a project VPC peering](/platform/howto/list-project-vpc-peering.md): Establish a peering connection between your Aiven project VPC and a cloud platform of your ##### list-service Overview - [Service management](/platform/howto/list-service.md): Overview ##### list-user-profile Browse through instructions for common Aiven platform tasks related to - [User profiles](/platform/howto/list-user-profile.md): Browse through instructions for common Aiven platform tasks related to ##### list-vpc-peering The VPC peering capability supported on the Aiven Platform improves network connectivity and security. It simplifies architecture, helps reduce network latency, and enhances resource sharing while maintaining isolation and control. - [Virtual private cloud (VPC) peering in Aiven](/platform/howto/list-vpc-peering.md): The VPC peering capability supported on the Aiven Platform improves network connectivity and security. It simplifies architecture, helps reduce network latency, and enhances resource sharing while maintaining isolation and control. ##### manage-application-users Application users give non-human users programmatic access to Aiven. You grant them access to organization resources using roles and permissions. - [Manage application users](/platform/howto/manage-application-users.md): Application users give non-human users programmatic access to Aiven. You grant them access to organization resources using roles and permissions. ##### manage-bank-transfers Aiven offers invoice billing and bank transfer payments for customers who have at least 1,000 USD in monthly recurring revenue. - [Pay with bank transfers](/platform/howto/manage-bank-transfers.md): Aiven offers invoice billing and bank transfer payments for customers who have at least 1,000 USD in monthly recurring revenue. ##### manage-billing-addresses Create addresses in the billing section of the Aiven Platform and use them as billing and shipping addresses in your billing groups. - [Manage billing and shipping addresses](/platform/howto/manage-billing-addresses.md): Create addresses in the billing section of the Aiven Platform and use them as billing and shipping addresses in your billing groups. ##### manage-domains Adding a verified domain in Aiven adds an extra layer of security to managing your organization's users. When you verify a domain, your organization users automatically become - [Manage domains](/platform/howto/manage-domains.md): Adding a verified domain in Aiven adds an extra layer of security to managing your organization's users. When you verify a domain, your organization users automatically become ##### manage-groups Create groups of users in your organization to make it easier to manage access to your organization's resources. - [Manage groups of users](/platform/howto/manage-groups.md): Create groups of users in your organization to make it easier to manage access to your organization's resources. ##### manage-org-users Adding users to your organization lets you give them access to specific projects and services within that organization. - [Manage users in an organization](/platform/howto/manage-org-users.md): Adding users to your organization lets you give them access to specific projects and services within that organization. ##### manage-org-vpc-peering-aws Set up a peering connection between your Aiven organization VPC and an AWS VPC. - [Manage an organization VPC peering with AWS Limited availability](/platform/howto/manage-org-vpc-peering-aws.md): Set up a peering connection between your Aiven organization VPC and an AWS VPC. ##### manage-org-vpc-peering-azure Set up a peering connection between your Aiven organization VPC and a Microsoft Azure virtual network. - [Manage an organization VPC peering with Microsoft Azure](/platform/howto/manage-org-vpc-peering-azure.md): Set up a peering connection between your Aiven organization VPC and a Microsoft Azure virtual network. ##### manage-org-vpc-peering-google Set up a peering connection between your Aiven organization VPC and a Google Cloud VPC. - [Manage organization VPC peering with Google Cloud Limited availability](/platform/howto/manage-org-vpc-peering-google.md): Set up a peering connection between your Aiven organization VPC and a Google Cloud VPC. ##### manage-organization-vpc Set up or delete an organization-wide VPC on the Aiven Platform. - [Manage organization virtual private clouds (VPCs) in Aiven Limited availability](/platform/howto/manage-organization-vpc.md): Set up or delete an organization-wide VPC on the Aiven Platform. ##### manage-organizations Learn how to manage your organizations via the Aiven Console. - [Manage organizations](/platform/howto/manage-organizations.md): Learn how to manage your organizations via the Aiven Console. ##### manage-payment-card Add credit cards to your organization and use them across different billing groups to pay for your Aiven services. - [Manage credit cards](/platform/howto/manage-payment-card.md): Add credit cards to your organization and use them across different billing groups to pay for your Aiven services. ##### manage-permissions You can grant organization users, application users, and groups access at the organization, organizational unit, and project level through roles and permissions. - [Manage permissions](/platform/howto/manage-permissions.md): You can grant organization users, application users, and groups access at the organization, organizational unit, and project level through roles and permissions. ##### manage-project Projects help you - [Manage projects](/platform/howto/manage-project.md): Projects help you ##### manage-project-vpc Set up or delete a project-wide VPC in your Aiven organization. - [Manage project virtual private clouds (VPCs) in Aiven](/platform/howto/manage-project-vpc.md): Set up or delete a project-wide VPC in your Aiven organization. ##### migrate-services-cloud-region When creating an Aiven service, you are not tied to a cloud provider or region. You can migrate your services later to better match your needs. - [Migrate service to another cloud or region](/platform/howto/migrate-services-cloud-region.md): When creating an Aiven service, you are not tied to a cloud provider or region. You can migrate your services later to better match your needs. ##### organization-event-logs Aiven consolidates all event logs for an organization into centralized event logs. - [Event logs](/platform/howto/organization-event-logs.md): Aiven consolidates all event logs for an organization into centralized event logs. ##### prepare-for-high-load Prepare an Aiven service for higher than usual traffic to avoid service outages. - [Prepare services for high load](/platform/howto/prepare-for-high-load.md): Prepare an Aiven service for higher than usual traffic to avoid service outages. ##### private-ip-resolution When an Aiven service is placed in a VPC (Virtual Private Cloud), the - [Handle resolution errors of private IP addresses](/platform/howto/private-ip-resolution.md): When an Aiven service is placed in a VPC (Virtual Private Cloud), the ##### public-access-in-vpc To enable public access for a service running within a virtual private cloud (VPC): - [Enable public access in VPCs](/platform/howto/public-access-in-vpc.md): To enable public access for a service running within a virtual private cloud (VPC): ##### reactivate-suspended-project If you have bills past due and didn't set up a payment method, your account may be suspended. - [Reactivate suspended accounts](/platform/howto/reactivate-suspended-project.md): If you have bills past due and didn't set up a payment method, your account may be suspended. ##### restore_progress_updates Track the restore progress of individual nodes in an Aiven service during node replacement, forking, or maintenance, using the Aiven API. - [Track service restore progress using the API](/platform/howto/restore_progress_updates.md): Track the restore progress of individual nodes in an Aiven service during node replacement, forking, or maintenance, using the Aiven API. ##### restrict-access Restrict access to your Aiven-managed service to a single IP, an address block, or any combination of both. - [Restrict network access to services](/platform/howto/restrict-access.md): Restrict access to your Aiven-managed service to a single IP, an address block, or any combination of both. ##### saml ###### add-auth0-idp Use Auth0 to give your organization users single sign-on (SSO) access to Aiven. - [Add Auth0 as an identity provider](/platform/howto/saml/add-auth0-idp.md): Use Auth0 to give your organization users single sign-on (SSO) access to Aiven. ###### add-azure-idp Use Microsoft Azure Active Directory (AD) to give your organization users single sign-on (SSO) access to Aiven. - [Add Microsoft Azure Active Directory as an identity provider](/platform/howto/saml/add-azure-idp.md): Use Microsoft Azure Active Directory (AD) to give your organization users single sign-on (SSO) access to Aiven. ###### add-fusionauth-idp Use FusionAuth to give your organization users single sign-on (SSO) access to Aiven. - [Add FusionAuth as an identity provider](/platform/howto/saml/add-fusionauth-idp.md): Use FusionAuth to give your organization users single sign-on (SSO) access to Aiven. ###### add-google-idp Use Google to give your organization users single sign-on (SSO) access to Aiven. - [Add Google as an identity provider](/platform/howto/saml/add-google-idp.md): Use Google to give your organization users single sign-on (SSO) access to Aiven. ###### add-identity-providers You can give your organization users access to Aiven through identity providers (IdPs) that support SAML. - [Add SAML identity providers](/platform/howto/saml/add-identity-providers.md): You can give your organization users access to Aiven through identity providers (IdPs) that support SAML. ###### add-jumpcloud-idp Use JumpCloud to give your organization users single sign-on (SSO) access to Aiven. - [Add JumpCloud as an identity provider](/platform/howto/saml/add-jumpcloud-idp.md): Use JumpCloud to give your organization users single sign-on (SSO) access to Aiven. ###### add-okta-idp Use Okta to give your organization users single sign-on (SSO) access to Aiven using SAML. Aiven also supports user provisioning for Okta with SCIM. - [Add Okta as an identity provider](/platform/howto/saml/add-okta-idp.md): Use Okta to give your organization users single sign-on (SSO) access to Aiven using SAML. Aiven also supports user provisioning for Okta with SCIM. ###### add-onelogin-idp Use OneLogin to give your organization users single sign-on (SSO) access to Aiven. - [Add OneLogin as an identity provider](/platform/howto/saml/add-onelogin-idp.md): Use OneLogin to give your organization users single sign-on (SSO) access to Aiven. ###### rotate-scim-token You can manually rotate the SCIM token for an identity provider to maintain the security of your user provisioning setup. - [Rotate SCIM tokens](/platform/howto/saml/rotate-scim-token.md): You can manually rotate the SCIM token for an identity provider to maintain the security of your user provisioning setup. ##### scale-services Change the plan of an Aiven service to scale it up or down and optimize costs. - [Change a service plan](/platform/howto/scale-services.md): Change the plan of an Aiven service to scale it up or down and optimize costs. ##### search-services On the page in Aiven Console, you can search for services by keywords and narrow down the results using filters. - [Search for services](/platform/howto/search-services.md): On the page in Aiven Console, you can search for services by keywords and narrow down the results using filters. ##### set-authentication-policies The authentication policy for your organization specifies the ways that users in your organization can access the organization on the Aiven Platform. - [Set authentication policies for organization users](/platform/howto/set-authentication-policies.md): The authentication policy for your organization specifies the ways that users in your organization can access the organization on the Aiven Platform. ##### support All customers using paid services have access to the Basic support tier. Aiven also offers paid support tiers with faster response times, phone support, and other services. Custom service level agreements are available for the Premium support tier. - [Support](/platform/howto/support.md): All customers using paid services have access to the Basic support tier. Aiven also offers paid support tiers with faster response times, phone support, and other services. Custom service level agreements are available for the Premium support tier. ##### tag-resources Add key-value tags to an Aiven service to organize services and track ownership, cost allocation, and governance. - [Use resource tags](/platform/howto/tag-resources.md): Add key-value tags to an Aiven service to organize services and track ownership, cost allocation, and governance. ##### technical-emails To stay up to date with the latest information about services and projects, you can set service and project contacts to receive email notifications. - [Manage project and service notifications](/platform/howto/technical-emails.md): To stay up to date with the latest information about services and projects, you can set service and project contacts to receive email notifications. ##### unsafe-passwords The Aiven Platform checks your email and password combination against a database of exposed credentials every time you log in and change your password. - [Change unsafe passwords](/platform/howto/unsafe-passwords.md): The Aiven Platform checks your email and password combination against a database of exposed credentials every time you log in and change your password. ##### use-aws-privatelinks AWS PrivateLink brings Aiven services to the selected virtual private cloud (VPC) in your AWS account. - [Use AWS PrivateLink with Aiven services](/platform/howto/use-aws-privatelinks.md): AWS PrivateLink brings Aiven services to the selected virtual private cloud (VPC) in your AWS account. ##### use-azure-privatelink Azure Private Link lets you bring your Aiven services into your virtual network (VNet) over a private endpoint. The endpoint creates a network interface into one of the VNet subnets, and receives a private IP address from its IP range. The private endpoint is routed to your Aiven service. - [Use Azure Private Link with Aiven services Early availability](/platform/howto/use-azure-privatelink.md): Azure Private Link lets you bring your Aiven services into your virtual network (VNet) over a private endpoint. The endpoint creates a network interface into one of the VNet subnets, and receives a private IP address from its IP range. The private endpoint is routed to your Aiven service. ##### use-billing-groups Costs associated with services and features in an Aiven project are charged to the payment method assigned to its billing group. - [Manage billing groups](/platform/howto/use-billing-groups.md): Costs associated with services and features in an Aiven project are charged to the payment method assigned to its billing group. ##### use-google-private-service-connect Enable Google Private Service Connect and use it with your Aiven-managed services. - [Use Google Private Service Connect with Aiven services Early availability](/platform/howto/use-google-private-service-connect.md): Enable Google Private Service Connect and use it with your Aiven-managed services. ##### user-2fa Two-factor authentication in Aiven provides an extra level of security by requiring a second authentication code in addition to the user password. - [Manage two-factor authentication](/platform/howto/user-2fa.md): Two-factor authentication in Aiven provides an extra level of security by requiring a second authentication code in addition to the user password. ##### view-organization-logs Monitor activity in your Aiven organization with the organization and organizational unit event logs. - [View event logs for organizations and organizational units](/platform/howto/view-organization-logs.md): Monitor activity in your Aiven organization with the organization and organizational unit event logs. ##### view-project-logs Monitor activity in your Aiven projects with the project event log. - [View project logs](/platform/howto/view-project-logs.md): Monitor activity in your Aiven projects with the project event log. ##### vnet-peering-azure Set up a peering connection between your Aiven project VPC and a Microsoft Azure virtual network. - [Manage a project VPC peering with Microsoft Azure](/platform/howto/vnet-peering-azure.md): Set up a peering connection between your Aiven project VPC and a Microsoft Azure virtual network. ##### vpc-peering-aws Set up a peering connection between your Aiven project VPC and an AWS VPC. - [Manage a project VPC peering with AWS](/platform/howto/vpc-peering-aws.md): Set up a peering connection between your Aiven project VPC and an AWS VPC. ##### vpc-peering-gcp Set up a peering connection between your Aiven project VPC and a Google Cloud VPC. - [Manage a project VPC peering with Google Cloud](/platform/howto/vpc-peering-gcp.md): Set up a peering connection between your Aiven project VPC and a Google Cloud VPC. ##### vpc-peering-upcloud Set up a peering connection between your Aiven project VPC and an UpCloud SDN network. - [Manage a project VPC peering with UpCloud](/platform/howto/vpc-peering-upcloud.md): Set up a peering connection between your Aiven project VPC and an UpCloud SDN network. ##### vpc-service-management Manage your Aiven services in a VPC, including setup, migration, and accessing resources securely within your project VPC. - [Manage a service in a VPC](/platform/howto/vpc-service-management.md): Manage your Aiven services in a VPC, including setup, migration, and accessing resources securely within your project VPC. #### reference ##### change-password You can change your password for the Aiven Platform in your account information. - [Change your password](/platform/reference/change-password.md): You can change your password for the Aiven Platform in your account information. ##### end-of-life Learn about the upcoming end of life (EOL) for select Aiven services, including timelines, actions after end of life, recommended migration options, and next steps. - [End of life for Aiven services](/platform/reference/end-of-life.md): Learn about the upcoming end of life (EOL) for select Aiven services, including timelines, actions after end of life, recommended migration options, and next steps. ##### eol-for-major-versions Learn about version lifecycle policies, end of life (EOL) schedules, upgrade procedures, and best practices for Aiven services and tools, including both multi-versioned services and single-versioned services. - [Aiven service and tool version lifecycle](/platform/reference/eol-for-major-versions.md): Learn about version lifecycle policies, end of life (EOL) schedules, upgrade procedures, and best practices for Aiven services and tools, including both multi-versioned services and single-versioned services. ##### get-resource-IDs Resource IDs like organization ID or user ID can be useful for working with the developer tools. You can get the IDs for resources in the Aiven Console. - [Get resource IDs](/platform/reference/get-resource-IDs.md): Resource IDs like organization ID or user ID can be useful for working with the developer tools. You can get the IDs for resources in the Aiven Console. ##### list_of_clouds This is a reference list of the default cloud regions available per provider on the Aiven Platform. - [Available cloud regions](/platform/reference/list_of_clouds.md): This is a reference list of the default cloud regions available per provider on the Aiven Platform. ##### password-policy Aiven is committed to keeping your data secure. Creating a strong - [Password policy](/platform/reference/password-policy.md): Aiven is committed to keeping your data secure. Creating a strong ##### referrals Invite someone to sign up to Aiven using your referral link and both of you get credits to spend when they start the Aiven trial. - [Refer Aiven and earn credits](/platform/reference/referrals.md): Invite someone to sign up to Aiven using your referral link and both of you get credits to spend when they start the Aiven trial. ##### service-ip-address When a new Aiven service is created, it automatically gets a hostname and one or more public IP addresses. - [Default service IP address and hostname](/platform/reference/service-ip-address.md): When a new Aiven service is created, it automatically gets a hostname and one or more public IP addresses. ### products #### clickhouse Aiven for ClickHouse® is a fully managed distributed columnar database based on open source ClickHouse - a fast, resource effective solution tailored for data warehouse and generation of real-time analytical data reports using advanced SQL queries. - [Aiven for ClickHouse®](/products/clickhouse.md): Aiven for ClickHouse® is a fully managed distributed columnar database based on open source ClickHouse - a fast, resource effective solution tailored for data warehouse and generation of real-time analytical data reports using advanced SQL queries. ##### concepts ###### choose-order-by-key In Aiven for ClickHouse®, the ORDER BY, PRIMARY KEY, and PARTITION BY for a MergeTree table work together to control how data is sorted, indexed, and grouped on disk. - [Choose ORDER BY and partition keys for MergeTree tables](/products/clickhouse/concepts/choose-order-by-key.md): In Aiven for ClickHouse®, the ORDER BY, PRIMARY KEY, and PARTITION BY for a MergeTree table work together to control how data is sorted, indexed, and grouped on disk. ###### clickhouse-tiered-storage The tiered storage feature introduces a method of organizing and storing data in two tiers for improved efficiency and cost optimization. The data is automatically moved to an appropriate tier based on your database's disk usage. - [Tiered storage in Aiven for ClickHouse®](/products/clickhouse/concepts/clickhouse-tiered-storage.md): The tiered storage feature introduces a method of organizing and storing data in two tiers for improved efficiency and cost optimization. The data is automatically moved to an appropriate tier based on your database's disk usage. ###### columnar-databases ClickHouse® is a columnar databases that handles data with specific benefits. - [ClickHouse® as a columnar database](/products/clickhouse/concepts/columnar-databases.md): ClickHouse® is a columnar databases that handles data with specific benefits. ###### data-integration-overview Aiven for ClickHouse® supports different types of integration allowing you to efficiently connect with other services or data sources and access the data to be processed. - [Aiven for ClickHouse® service integrations](/products/clickhouse/concepts/data-integration-overview.md): Aiven for ClickHouse® supports different types of integration allowing you to efficiently connect with other services or data sources and access the data to be processed. ###### databases-tables-views Databases, tables, and views organize data in Aiven for ClickHouse®. A database - [Databases, tables, and views in Aiven for ClickHouse®](/products/clickhouse/concepts/databases-tables-views.md): Databases, tables, and views organize data in Aiven for ClickHouse®. A database ###### disaster-recovery Aiven for ClickHouse® prevents and mitigates emergencies or crises with multiple disaster recovery methods to keep your data safe and sound. - [Disaster recovery in Aiven for ClickHouse®](/products/clickhouse/concepts/disaster-recovery.md): Aiven for ClickHouse® prevents and mitigates emergencies or crises with multiple disaster recovery methods to keep your data safe and sound. ###### federated-queries Discover federated queries and their capabilities in Aiven for ClickHouse® and how they simplify and speed up migrating into Aiven from external data sources. - [Querying external data in Aiven for ClickHouse®](/products/clickhouse/concepts/federated-queries.md): Discover federated queries and their capabilities in Aiven for ClickHouse® and how they simplify and speed up migrating into Aiven from external data sources. ###### indexing ClickHouse® processes data differently from other database management systems. ClickHouse uses sparse and skipping indexes and a vector computation engine. - [Indexing and data processing in ClickHouse®](/products/clickhouse/concepts/indexing.md): ClickHouse® processes data differently from other database management systems. ClickHouse uses sparse and skipping indexes and a vector computation engine. ###### olap Online analytical processing (OLAP) is an approach to producing - [Online analytical processing](/products/clickhouse/concepts/olap.md): Online analytical processing (OLAP) is an approach to producing ###### query-kafka-topic-data Query Kafka topic data in Aiven for ClickHouse® by connecting an Aiven for Apache Kafka® topic to a ClickHouse table. - [Query Kafka topic data in Aiven for ClickHouse®](/products/clickhouse/concepts/query-kafka-topic-data.md): Query Kafka topic data in Aiven for ClickHouse® by connecting an Aiven for Apache Kafka® topic to a ClickHouse table. ###### service-architecture Aiven for ClickHouse® is implemented as a multi-master cluster where data replication is - [Aiven for ClickHouse® service architecture](/products/clickhouse/concepts/service-architecture.md): Aiven for ClickHouse® is implemented as a multi-master cluster where data replication is ###### service-management Manage the security, configuration, and lifecycle of your Aiven for ClickHouse® - [Service management in Aiven for ClickHouse®](/products/clickhouse/concepts/service-management.md): Manage the security, configuration, and lifecycle of your Aiven for ClickHouse® ###### strings Aiven for ClickHouse® uses ClickHouse® databases, which can store diverse types of data, such as strings, decimals, booleans, or arrays. - [String data type in Aiven for ClickHouse®](/products/clickhouse/concepts/strings.md): Aiven for ClickHouse® uses ClickHouse® databases, which can store diverse types of data, such as strings, decimals, booleans, or arrays. ##### get-started Start using Aiven for ClickHouse® by creating and configuring a service, connecting to it, and loading sample data. - [Get started with Aiven for ClickHouse®](/products/clickhouse/get-started.md): Start using Aiven for ClickHouse® by creating and configuring a service, connecting to it, and loading sample data. ##### howto ###### change-cloud-region Move your Aiven for ClickHouse® service to a different cloud provider or region. - [Change the cloud or region for your Aiven for ClickHouse® service](/products/clickhouse/howto/change-cloud-region.md): Move your Aiven for ClickHouse® service to a different cloud provider or region. ###### change-service-plan Change the service plan for your Aiven for ClickHouse® service to scale resources up or down and optimize costs. - [Change the plan for your Aiven for ClickHouse® service](/products/clickhouse/howto/change-service-plan.md): Change the service plan for your Aiven for ClickHouse® service to scale resources up or down and optimize costs. ###### check-data-tiered-storage Monitor how your data is distributed between the two layers of your tiered storage: Network-attached block storage and object storage. - [Check data distribution between storage devices in Aiven for ClickHouse®'s tiered storage](/products/clickhouse/howto/check-data-tiered-storage.md): Monitor how your data is distributed between the two layers of your tiered storage: Network-attached block storage and object storage. ###### clickhouse-query-cache Aiven for ClickHouse® provides a query cache mechanism that helps improve query performance - [Use query cache in Aiven for ClickHouse®](/products/clickhouse/howto/clickhouse-query-cache.md): Aiven for ClickHouse® provides a query cache mechanism that helps improve query performance ###### configure-backup Set the time when - [Schedule Aiven for ClickHouse® backups](/products/clickhouse/howto/configure-backup.md): Set the time when ###### configure-tiered-storage Control how your data is distributed between storage devices in the tiered storage of an Aiven for ClickHouse® service. Configure tables so that ClickHouse automatically writes your data to network-attached block storage or object storage as needed. - [Configure data retention thresholds in Aiven for ClickHouse®'s tiered storage](/products/clickhouse/howto/configure-tiered-storage.md): Control how your data is distributed between storage devices in the tiered storage of an Aiven for ClickHouse® service. Configure tables so that ClickHouse automatically writes your data to network-attached block storage or object storage as needed. ###### connect-to-grafana You can visualise your ClickHouse® data using Grafana® and Aiven can help you connect the two services. - [Visualize ClickHouse® data with Grafana®](/products/clickhouse/howto/connect-to-grafana.md): You can visualise your ClickHouse® data using Grafana® and Aiven can help you connect the two services. ###### connect-with-clickhouse-cli It's recommended to connect to a ClickHouse® cluster with the ClickHouse® client. - [Connect to Aiven for ClickHouse® with clickhouse-client](/products/clickhouse/howto/connect-with-clickhouse-cli.md): It's recommended to connect to a ClickHouse® cluster with the ClickHouse® client. ###### connect-with-go To connect to your Aiven for ClickHouse® service with Go, you can use - [Connect to Aiven for ClickHouse® with Go](/products/clickhouse/howto/connect-with-go.md): To connect to your Aiven for ClickHouse® service with Go, you can use ###### connect-with-java Learn how to connect to your Aiven for ClickHouse® service with Java - [Connect to Aiven for ClickHouse® with Java](/products/clickhouse/howto/connect-with-java.md): Learn how to connect to your Aiven for ClickHouse® service with Java ###### connect-with-jdbc You can use [ClickHouse JDBC - [Connect Aiven for ClickHouse® to external databases via JDBC](/products/clickhouse/howto/connect-with-jdbc.md): You can use [ClickHouse JDBC ###### connect-with-nodejs Learn how to connect to your Aiven for ClickHouse® service with Node.js - [Connect to Aiven for ClickHouse® with Node.js](/products/clickhouse/howto/connect-with-nodejs.md): Learn how to connect to your Aiven for ClickHouse® service with Node.js ###### connect-with-php Learn how to connect to your Aiven for ClickHouse® service with PHP using the PHP ClickHouse client and the HTTPS port. - [Connect to Aiven for ClickHouse® with PHP](/products/clickhouse/howto/connect-with-php.md): Learn how to connect to your Aiven for ClickHouse® service with PHP using the PHP ClickHouse client and the HTTPS port. ###### connect-with-python To connect to your Aiven for ClickHouse® service with Python, you can - [Connect to Aiven for ClickHouse® with Python](/products/clickhouse/howto/connect-with-python.md): To connect to your Aiven for ClickHouse® service with Python, you can ###### controlled-upgrade-pipelines Link Aiven for ClickHouse® services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. - [Controlled upgrade pipelines for your Aiven for ClickHouse® service Limited availability](/products/clickhouse/howto/controlled-upgrade-pipelines.md): Link Aiven for ClickHouse® services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. ###### copy-data-across-instances You can copy data from one ClickHouse® server to another using the remoteSecure() function. - [Copy data between Aiven for ClickHouse® services](/products/clickhouse/howto/copy-data-across-instances.md): You can copy data from one ClickHouse® server to another using the remoteSecure() function. ###### create-dictionary Create dictionaries in Aiven for ClickHouse® to accelerate queries for better efficiency and performance. - [Create dictionaries in Aiven for ClickHouse®](/products/clickhouse/howto/create-dictionary.md): Create dictionaries in Aiven for ClickHouse® to accelerate queries for better efficiency and performance. ###### data-service-integration Connect your Aiven for ClickHouse® service with another Aiven-managed service or external data source to make your data available in the Aiven for ClickHouse service. - [Set up Aiven for ClickHouse® data source integrations](/products/clickhouse/howto/data-service-integration.md): Connect your Aiven for ClickHouse® service with another Aiven-managed service or external data source to make your data available in the Aiven for ClickHouse service. ###### disk-autoscaler Automatically increase the disk storage of your Aiven for ClickHouse® service when it's running out of space, instead of resizing it manually. - [Scale disk storage automatically for your Aiven for ClickHouse® service](/products/clickhouse/howto/disk-autoscaler.md): Automatically increase the disk storage of your Aiven for ClickHouse® service when it's running out of space, instead of resizing it manually. ###### enable-tiered-storage Enable the tiered storage feature on a table in your Aiven for ClickHouse® service. - [Enable tiered storage in Aiven for ClickHouse®](/products/clickhouse/howto/enable-tiered-storage.md): Enable the tiered storage feature on a table in your Aiven for ClickHouse® service. ###### fetch-query-statistics Usually, query statistics in ClickHouse can be obtained using the system.query_log table, which stores statistics of each executed query, including memory usage and duration. - [Fetch query statistics for Aiven for ClickHouse®](/products/clickhouse/howto/fetch-query-statistics.md): Usually, query statistics in ClickHouse can be obtained using the system.query_log table, which stores statistics of each executed query, including memory usage and duration. ###### fork-service Fork your Aiven for ClickHouse® service to create an independent copy for testing, - [Fork your Aiven for ClickHouse® service](/products/clickhouse/howto/fork-service.md): Fork your Aiven for ClickHouse® service to create an independent copy for testing, ###### integrate-kafka Integrate Aiven for ClickHouse® with either Aiven for Apache Kafka® service located in the same project, or an external Apache Kafka endpoint. - [Connect Apache Kafka® to Aiven for ClickHouse®](/products/clickhouse/howto/integrate-kafka.md): Integrate Aiven for ClickHouse® with either Aiven for Apache Kafka® service located in the same project, or an external Apache Kafka endpoint. ###### integrate-postgresql You can integrate Aiven for ClickHouse® with either Aiven for PostgreSQL service located in the same project, or an external PostgreSQL endpoint. - [Connect PostgreSQL® to Aiven for ClickHouse®](/products/clickhouse/howto/integrate-postgresql.md): You can integrate Aiven for ClickHouse® with either Aiven for PostgreSQL service located in the same project, or an external PostgreSQL endpoint. ###### integration-databases Create and manage integration databases in Aiven for ClickHouse® to query data from integrated services: - [Create and manage Aiven for ClickHouse® integration databases](/products/clickhouse/howto/integration-databases.md): Create and manage integration databases in Aiven for ClickHouse® to query data from integrated services: ###### list-connect-to-service Connect to the Aiven for ClickHouse® service using various programming languages or tools. - [Connect to Aiven for ClickHouse®](/products/clickhouse/howto/list-connect-to-service.md): Connect to the Aiven for ClickHouse® service using various programming languages or tools. ###### local-cache-tiered-storage Aiven for ClickHouse®'s tiered storage features local on-disk cache for remote files for improved query performance and reduced latency. - [Manage local cache for remote files in Aiven for ClickHouse®'s tiered storage](/products/clickhouse/howto/local-cache-tiered-storage.md): Aiven for ClickHouse®'s tiered storage features local on-disk cache for remote files for improved query performance and reduced latency. ###### maintenance-updates Manage maintenance updates and set the maintenance window for your Aiven for ClickHouse® service. - [Maintenance and updates for your Aiven for ClickHouse® service](/products/clickhouse/howto/maintenance-updates.md): Manage maintenance updates and set the maintenance window for your Aiven for ClickHouse® service. ###### manage-clickhouse-versions Aiven for ClickHouse® supports multiple ClickHouse versions. You can choose a version when you create a service and upgrade to a newer supported version later. - [Manage versions in Aiven for ClickHouse®](/products/clickhouse/howto/manage-clickhouse-versions.md): Aiven for ClickHouse® supports multiple ClickHouse versions. You can choose a version when you create a service and upgrade to a newer supported version later. ###### manage-databases-tables Create and work with databases and tables in Aiven for ClickHouse®. - [Manage Aiven for ClickHouse® databases and tables](/products/clickhouse/howto/manage-databases-tables.md): Create and work with databases and tables in Aiven for ClickHouse®. ###### manage-users-roles Create Aiven for ClickHouse® users and roles and grant them specific privileges to control or restrict access to your service. - [Manage Aiven for ClickHouse® users and roles](/products/clickhouse/howto/manage-users-roles.md): Create Aiven for ClickHouse® users and roles and grant them specific privileges to control or restrict access to your service. ###### materialized-views Use materialized views to persist data from the Kafka® table engine. - [Create materialized views in ClickHouse®](/products/clickhouse/howto/materialized-views.md): Use materialized views to persist data from the Kafka® table engine. ###### monitor-performance Push Aiven for ClickHouse® metrics to Aiven for Metrics or Aiven for PostgreSQL®, and integrate with Aiven for Grafana® to monitor your metrics on Grafana dashboards. - [Monitor Aiven for ClickHouse® metrics with Aiven for Grafana®](/products/clickhouse/howto/monitor-performance.md): Push Aiven for ClickHouse® metrics to Aiven for Metrics or Aiven for PostgreSQL®, and integrate with Aiven for Grafana® to monitor your metrics on Grafana dashboards. ###### power-cycle-service Power off your Aiven for ClickHouse® service to release resources and save credits, power - [Power on/off and delete your Aiven for ClickHouse® service](/products/clickhouse/howto/power-cycle-service.md): Power off your Aiven for ClickHouse® service to release resources and save credits, power ###### query-databases Run a query against an Aiven for ClickHouse® database using a tool of your choice. - [Query Aiven for ClickHouse® databases](/products/clickhouse/howto/query-databases.md): Run a query against an Aiven for ClickHouse® database using a tool of your choice. ###### rename-service Change the name of your Aiven for ClickHouse® service by forking it under a new name and - [Rename your Aiven for ClickHouse® service](/products/clickhouse/howto/rename-service.md): Change the name of your Aiven for ClickHouse® service by forking it under a new name and ###### restore-backup Restore an Aiven for ClickHouse® service from a - [Restore an Aiven for ClickHouse® backup](/products/clickhouse/howto/restore-backup.md): Restore an Aiven for ClickHouse® service from a ###### run-federated-queries With federated queries in Aiven for ClickHouse®, you can read and pull data from an external S3-compatible object storage or any web resource accessible over HTTP. - [Read and pull data from S3 object storages and web resources over HTTP](/products/clickhouse/howto/run-federated-queries.md): With federated queries in Aiven for ClickHouse®, you can read and pull data from an external S3-compatible object storage or any web resource accessible over HTTP. ###### scale-disk-storage Scale the disk storage of your Aiven for ClickHouse® service up or down without disrupting the running service. - [Scale disk storage for your Aiven for ClickHouse® service](/products/clickhouse/howto/scale-disk-storage.md): Scale the disk storage of your Aiven for ClickHouse® service up or down without disrupting the running service. ###### secure-service You can secure your Aiven for ClickHouse® service in a few different ways, for example by restricting network access, using Virtual Private Cloud (VPC), and enabling service termination protection. - [Secure a managed ClickHouse® service](/products/clickhouse/howto/secure-service.md): You can secure your Aiven for ClickHouse® service in a few different ways, for example by restricting network access, using Virtual Private Cloud (VPC), and enabling service termination protection. ###### set-up-kafka-topic-querying Send data from an Aiven for Apache Kafka® topic to Aiven for ClickHouse® and query it with SQL. - [Set up Kafka topic querying in Aiven for ClickHouse®](/products/clickhouse/howto/set-up-kafka-topic-querying.md): Send data from an Aiven for Apache Kafka® topic to Aiven for ClickHouse® and query it with SQL. ###### sql-user-defined-functions Use SQL user defined functions (UDFs) in Aiven for ClickHouse® to speed up your queries and optimize your application performance. - [Use SQL user defined functions in Aiven for ClickHouse®](/products/clickhouse/howto/sql-user-defined-functions.md): Use SQL user defined functions (UDFs) in Aiven for ClickHouse® to speed up your queries and optimize your application performance. ###### tag-service Add key-value tags to your Aiven for ClickHouse® service to organize services and track - [Tag your Aiven for ClickHouse® service](/products/clickhouse/howto/tag-service.md): Add key-value tags to your Aiven for ClickHouse® service to organize services and track ###### transfer-data-tiered-storage Moving data from network-attached block storage to object storage allows you to size down your block storage by selecting a service plan with less capacity. You can move the data back to network-attached block storage anytime. - [Transfer data between storage devices in Aiven for ClickHouse®'s tiered storage](/products/clickhouse/howto/transfer-data-tiered-storage.md): Moving data from network-attached block storage to object storage allows you to size down your block storage by selecting a service plan with less capacity. You can move the data back to network-attached block storage anytime. ###### use-shards-with-distributed-table If your Aiven for ClickHouse® service uses multiple shards, the data is replicated only between nodes of the same shard. - [Enable reading and writing data across shards in Aiven for ClickHouse®](/products/clickhouse/howto/use-shards-with-distributed-table.md): If your Aiven for ClickHouse® service uses multiple shards, the data is replicated only between nodes of the same shard. ###### vector-similarity-index-cache Tune the vector similarity index cache in Aiven for ClickHouse® to improve vector search performance for Hierarchical Navigable Small World (HNSW) indexes. - [Tune the vector similarity index cache in Aiven for ClickHouse®](/products/clickhouse/howto/vector-similarity-index-cache.md): Tune the vector similarity index cache in Aiven for ClickHouse® to improve vector search performance for Hierarchical Navigable Small World (HNSW) indexes. ##### reference ###### 25-8-default-settings Aiven for ClickHouse® uses a managed configuration that differs from upstream ClickHouse defaults. - [Aiven for ClickHouse® 25.8 default settings](/products/clickhouse/reference/25-8-default-settings.md): Aiven for ClickHouse® uses a managed configuration that differs from upstream ClickHouse defaults. ###### 26-3-default-settings Aiven for ClickHouse® uses a managed configuration that differs from upstream ClickHouse defaults. These differences help keep services reliable, secure, and predictable in Aiven-managed environments. - [Aiven for ClickHouse® 26.3 default settings](/products/clickhouse/reference/26-3-default-settings.md): Aiven for ClickHouse® uses a managed configuration that differs from upstream ClickHouse defaults. These differences help keep services reliable, secure, and predictable in Aiven-managed environments. ###### advanced-params See the configuration options available for - [Advanced parameters for Aiven for ClickHouse®](/products/clickhouse/reference/advanced-params.md): See the configuration options available for ###### clickhouse-metrics-datadog Learn what metrics are available via Datadog for Aiven for ClickHouse® - [Aiven for ClickHouse® metrics available via Datadog](/products/clickhouse/reference/clickhouse-metrics-datadog.md): Learn what metrics are available via Datadog for Aiven for ClickHouse® ###### clickhouse-metrics-prometheus List of all metrics available via Prometheus for Aiven for ClickHouse® services. - [Aiven for ClickHouse® metrics available via Prometheus](/products/clickhouse/reference/clickhouse-metrics-prometheus.md): List of all metrics available via Prometheus for Aiven for ClickHouse® services. ###### clickhouse-system-tables Aiven for ClickHouse® supports multiple types of system tables, which store metadata and system-level information. Querying system tables allows you to check the configuration, performance, and state of your database. - [System tables in Aiven for ClickHouse®](/products/clickhouse/reference/clickhouse-system-tables.md): Aiven for ClickHouse® supports multiple types of system tables, which store metadata and system-level information. Querying system tables allows you to check the configuration, performance, and state of your database. ###### limitations By respecting the Aiven for ClickHouse® restrictions and quotas, you can improve the security and productivity of your service workloads. - [Aiven for ClickHouse® limits and limitations](/products/clickhouse/reference/limitations.md): By respecting the Aiven for ClickHouse® restrictions and quotas, you can improve the security and productivity of your service workloads. ###### metrics-list Browse the Aiven for ClickHouse® service metrics shown in the Grafana® monitoring dashboard. - [Aiven for ClickHouse® monitoring dashboard metrics shown in Grafana®](/products/clickhouse/reference/metrics-list.md): Browse the Aiven for ClickHouse® service metrics shown in the Grafana® monitoring dashboard. ###### s3-supported-file-formats The S3 table function - [S3 table function file formats in Aiven for ClickHouse®](/products/clickhouse/reference/s3-supported-file-formats.md): The S3 table function ###### supported-database-engines A database engine controls how a database manages tables and metadata. - [Supported database engines in Aiven for ClickHouse®](/products/clickhouse/reference/supported-database-engines.md): A database engine controls how a database manages tables and metadata. ###### supported-input-output-formats When connecting Aiven for ClickHouse® to Aiven for Apache Kafka® using Aiven integrations, data exchange is possible with the following formats only: - [Formats for Aiven for ClickHouse® - Aiven for Apache Kafka® data exchange](/products/clickhouse/reference/supported-input-output-formats.md): When connecting Aiven for ClickHouse® to Aiven for Apache Kafka® using Aiven integrations, data exchange is possible with the following formats only: ###### supported-interfaces-drivers Find out what technologies and tools you can use to interact with Aiven for ClickHouse®. - [Interfaces and drivers supported in Aiven for ClickHouse®](/products/clickhouse/reference/supported-interfaces-drivers.md): Find out what technologies and tools you can use to interact with Aiven for ClickHouse®. ###### supported-table-engines Table engines define how data is stored and which queries a table supports. - [Supported table engines in Aiven for ClickHouse®](/products/clickhouse/reference/supported-table-engines.md): Table engines define how data is stored and which queries a table supports. ###### supported-table-functions [Table - [Table functions supported in Aiven for ClickHouse®](/products/clickhouse/reference/supported-table-functions.md): [Table ###### upgrade-to-26-3 Aiven for ClickHouse® 26.3 is a long-term support (LTS) release. - [Upgrade to Aiven for ClickHouse® 26.3](/products/clickhouse/reference/upgrade-to-26-3.md): Aiven for ClickHouse® 26.3 is a long-term support (LTS) release. ###### version-lifecycle Learn how Aiven manages Aiven for ClickHouse® version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. - [Aiven for ClickHouse® version lifecycle](/products/clickhouse/reference/version-lifecycle.md): Learn how Aiven manages Aiven for ClickHouse® version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. ###### version-support-policy Aiven for ClickHouse® follows the upstream ClickHouse long-term support, or LTS, release model. - [Aiven for ClickHouse® version support policy](/products/clickhouse/reference/version-support-policy.md): Aiven for ClickHouse® follows the upstream ClickHouse long-term support, or LTS, release model. #### datahub Aiven for DataHub is a cost-effective data catalog that integrates seamlessly with Aiven services. - [Aiven for DataHub](/products/datahub.md): Aiven for DataHub is a cost-effective data catalog that integrates seamlessly with Aiven services. ##### change-cloud You can change the cloud provider or region of an Aiven for DataHub service at any time. - [Change cloud for Aiven for DataHub](/products/datahub/change-cloud.md): You can change the cloud provider or region of an Aiven for DataHub service at any time. ##### configure-slack-notifications Get activity notifications for your DataHub service in a Slack channel, including new datasets, ownership changes, tags, and glossary updates. - [Configure Slack notifications for DataHub activity](/products/datahub/configure-slack-notifications.md): Get activity notifications for your DataHub service in a Slack channel, including new datasets, ownership changes, tags, and glossary updates. ##### configure-teams-notifications Get activity notifications for your DataHub service in a Microsoft Teams channel, including new datasets, ownership changes, tags, and glossary updates. - [Configure Teams notifications for DataHub activity Limited availability](/products/datahub/configure-teams-notifications.md): Get activity notifications for your DataHub service in a Microsoft Teams channel, including new datasets, ownership changes, tags, and glossary updates. ##### connect-datahub-to-services Add connectors to your DataHub service to ingest data. - [Connect DataHub to services](/products/datahub/connect-datahub-to-services.md): Add connectors to your DataHub service to ingest data. ##### datahub-lineage Data lineage is a map of how each of your data assets moves across your systems. Lineage can help you: - [View data lineage in DataHub](/products/datahub/datahub-lineage.md): Data lineage is a map of how each of your data assets moves across your systems. Lineage can help you: ##### datahub-mcp-server Make your data ecosystem visible to AI agents with the DataHub MCP server, enabling natural language search, end-to-end lineage tracking, and context-aware SQL generation. - [Use the DataHub MCP server](/products/datahub/datahub-mcp-server.md): Make your data ecosystem visible to AI agents with the DataHub MCP server, enabling natural language search, end-to-end lineage tracking, and context-aware SQL generation. ##### delete-service When you delete an Aiven for DataHub service, all service data and configuration are permanently deleted. - [Delete an Aiven for DataHub service](/products/datahub/delete-service.md): When you delete an Aiven for DataHub service, all service data and configuration are permanently deleted. ##### enable-oidc-auth-datahub Use OpenID Connect (OIDC) to configure single sign-on (SSO) to your DataHub service with your identity provider. - [Enable OIDC authentication for Aiven for DataHub](/products/datahub/enable-oidc-auth-datahub.md): Use OpenID Connect (OIDC) to configure single sign-on (SSO) to your DataHub service with your identity provider. ##### enable-prometheus-metrics Enable Prometheus metrics for your Aiven for DataHub service to monitor its performance and health. - [Enable Prometheus metrics for Aiven for DataHub](/products/datahub/enable-prometheus-metrics.md): Enable Prometheus metrics for your Aiven for DataHub service to monitor its performance and health. ##### fork-datahub-service Fork an Aiven for DataHub service to create a complete copy of it from its latest backups. This restores both its PostgreSQL metadata database and OpenSearch search index. - [Fork DataHub services](/products/datahub/fork-datahub-service.md): Fork an Aiven for DataHub service to create a complete copy of it from its latest backups. This restores both its PostgreSQL metadata database and OpenSearch search index. ##### get-started Start using DataHub by creating and configuring your first service. - [Get started with Aiven for DataHub](/products/datahub/get-started.md): Start using DataHub by creating and configuring your first service. ##### maintenance-updates Manage maintenance updates and set the maintenance window for your - [Maintenance updates for Aiven for DataHub](/products/datahub/maintenance-updates.md): Manage maintenance updates and set the maintenance window for your ##### manage-datahub-users Invite users to your DataHub service, giving them access to read or edit metadata. - [Manage DataHub users](/products/datahub/manage-datahub-users.md): Invite users to your DataHub service, giving them access to read or edit metadata. ##### permissions The following roles and permissions are required for specific Aiven for DataHub features. - [Permissions for Aiven for DataHub features](/products/datahub/permissions.md): The following roles and permissions are required for specific Aiven for DataHub features. ##### power-off-service You can power off an Aiven for DataHub service at any time to stop all processes and reduce costs. - [Power an Aiven for DataHub service off or on](/products/datahub/power-off-service.md): You can power off an Aiven for DataHub service at any time to stop all processes and reduce costs. ##### restore-datahub-indices Rebuild your OpenSearch indices for search and graph data if your search results or relationship graphs differ from the data in your metadata database. - [Reindex Aiven for DataHub search and graph indices Limited availability](/products/datahub/restore-datahub-indices.md): Rebuild your OpenSearch indices for search and graph data if your search results or relationship graphs differ from the data in your metadata database. ##### rotate-secrets Rotate your Aiven for DataHub authentication and token secrets to maintain security and prevent unauthorized access. - [Rotate Aiven for DataHub secrets](/products/datahub/rotate-secrets.md): Rotate your Aiven for DataHub authentication and token secrets to maintain security and prevent unauthorized access. ##### scale-datahub-service Scale your Aiven for DataHub service and its underlying resources to optimize costs and improve performance. - [Scale Aiven for DataHub services](/products/datahub/scale-datahub-service.md): Scale your Aiven for DataHub service and its underlying resources to optimize costs and improve performance. ##### tag-services Use tags to add metadata to Aiven services to categorize them or run custom logic on them. - [Tag Aiven for DataHub services](/products/datahub/tag-services.md): Use tags to add metadata to Aiven services to categorize them or run custom logic on them. ##### upgrade-datahub-version Upgrade your Aiven for DataHub service to a new version. - [Upgrade Aiven for DataHub Limited availability](/products/datahub/upgrade-datahub-version.md): Upgrade your Aiven for DataHub service to a new version. #### dragonfly Aiven for Dragonfly is an advanced, high-scale, and Aiven for Caching compatible in-memory database service that can be deployed in your preferred cloud environment. - [Aiven for Dragonfly®](/products/dragonfly.md): Aiven for Dragonfly is an advanced, high-scale, and Aiven for Caching compatible in-memory database service that can be deployed in your preferred cloud environment. ##### concepts ###### ha-dragonfly Aiven for Dragonfly® offers different plans with varying levels of high availability. The available features depend on the selected plan. - [High availability in Aiven for Dragonfly®](/products/dragonfly/concepts/ha-dragonfly.md): Aiven for Dragonfly® offers different plans with varying levels of high availability. The available features depend on the selected plan. ##### get-started Get started with Aiven for Dragonfly by creating your service, integrating it with other services, and connecting to it with your preferred programming language. - [Get started with Aiven for Dragonfly®](/products/dragonfly/get-started.md): Get started with Aiven for Dragonfly by creating your service, integrating it with other services, and connecting to it with your preferred programming language. ##### howto ###### compatibility-redisjson Learn how to optimize your experience with RedisJSON in Aiven for Dragonfly® with the v2 JSONPath syntax using the $ root node. - [RedisJSON v2 syntax compatibility](/products/dragonfly/howto/compatibility-redisjson.md): Learn how to optimize your experience with RedisJSON in Aiven for Dragonfly® with the v2 JSONPath syntax using the $ root node. ###### connect-go This example demonstrates how to connect to Dragonfly® using Go, using - [Connect to Aiven for Dragonfly® with Go](/products/dragonfly/howto/connect-go.md): This example demonstrates how to connect to Dragonfly® using Go, using ###### connect-node This example demonstrates how to connect to Dragonfly® from NodeJS using - [Connect to Aiven for Dragonfly® with NodeJS](/products/dragonfly/howto/connect-node.md): This example demonstrates how to connect to Dragonfly® from NodeJS using ###### connect-python This example demonstrates how to connect to Dragonfly® using Python, - [Connect to Aiven for Dragonfly® with Python](/products/dragonfly/howto/connect-python.md): This example demonstrates how to connect to Dragonfly® using Python, ###### connect-redis-cli This example demonstrates how to connect to Dragonfly® using - [Connect to Aiven for Dragonfly® with redis-cli](/products/dragonfly/howto/connect-redis-cli.md): This example demonstrates how to connect to Dragonfly® using ###### eviction-policy-df Aiven for Dragonfly® optimizes cache memory management with a low-overhead data eviction policy. - [Data eviction policy in Aiven for Dragonfly](/products/dragonfly/howto/eviction-policy-df.md): Aiven for Dragonfly® optimizes cache memory management with a low-overhead data eviction policy. ###### list-code-samples Connect to the Aiven for Dragonfly® service using various programming languages or tools. - [Connect to Aiven for Dragonfly®](/products/dragonfly/howto/list-code-samples.md): Connect to the Aiven for Dragonfly® service using various programming languages or tools. ###### migrate-aiven-caching-df-console Migrate your Aiven for Caching or Aiven for Valkey databases to Aiven for Dragonfly using the Aiven Console migration tool. - [Migrate from Aiven for Caching or Aiven for Valkey™ to Aiven for Dragonfly](/products/dragonfly/howto/migrate-aiven-caching-df-console.md): Migrate your Aiven for Caching or Aiven for Valkey databases to Aiven for Dragonfly using the Aiven Console migration tool. ###### migrate-ext-redis-df-console Migrate external Redis® or Valkey databases to Aiven for Dragonfly® using the Aiven Console migration tool. - [Migrate from external Redis®* or Valkey to Aiven for Dragonfly](/products/dragonfly/howto/migrate-ext-redis-df-console.md): Migrate external Redis® or Valkey databases to Aiven for Dragonfly® using the Aiven Console migration tool. ##### reference ###### advanced-params See the configuration options available for Aiven for Dragonfly®: - [Advanced parameters for Aiven for Dragonfly®](/products/dragonfly/reference/advanced-params.md): See the configuration options available for Aiven for Dragonfly®: ###### version-lifecycle Learn how Aiven manages the Aiven for Dragonfly® service version, end of life (EOL) dates, and what happens to your service after the version reaches EOL. - [Aiven for Dragonfly® version lifecycle](/products/dragonfly/reference/version-lifecycle.md): Learn how Aiven manages the Aiven for Dragonfly® service version, end of life (EOL) dates, and what happens to your service after the version reaches EOL. #### flink Aiven for Apache Flink® is a fully managed service that leverages the power of the open-source Apache Flink framework to provide distributed, stateful stream processing capabilities, allowing users to perform real-time computation with SQL efficiently. - [Aiven for Apache Flink®](/products/flink.md): Aiven for Apache Flink® is a fully managed service that leverages the power of the open-source Apache Flink framework to provide distributed, stateful stream processing capabilities, allowing users to perform real-time computation with SQL efficiently. ##### concepts ###### checkpoints Checkpoints in Aiven for Apache Flink® are a key feature for ensuring resiliency and fault tolerance in stateful functions. - [Checkpoints](/products/flink/concepts/checkpoints.md): Checkpoints in Aiven for Apache Flink® are a key feature for ensuring resiliency and fault tolerance in stateful functions. ###### custom-jars Aiven for Apache Flink enables you to upload, deploy, and manage your own Java code as custom JARs within a JAR application. - [Custom JARs in Aiven for Apache Flink® Limited availability](/products/flink/concepts/custom-jars.md): Aiven for Apache Flink enables you to upload, deploy, and manage your own Java code as custom JARs within a JAR application. ###### event-processing-time Event time refers to when events occur, and processing time is when a system observes or processes these events. Understanding the difference between these two is essential for data processing and streaming. It affects data handling, analysis, and storage. - [Event and processing times](/products/flink/concepts/event-processing-time.md): Event time refers to when events occur, and processing time is when a system observes or processes these events. Understanding the difference between these two is essential for data processing and streaming. It affects data handling, analysis, and storage. ###### flink-applications An Aiven for Apache Flink® Application is an abstraction layer that simplifies building data processing pipelines in Apache Flink. - [Aiven for Apache Flink® applications](/products/flink/concepts/flink-applications.md): An Aiven for Apache Flink® Application is an abstraction layer that simplifies building data processing pipelines in Apache Flink. ###### flink-architecture At a high level, Flink has a runtime architecture consisting of two types of processes: a JobManager and one or more TaskManager. - [Aiven for Apache Flink® architecture](/products/flink/concepts/flink-architecture.md): At a high level, Flink has a runtime architecture consisting of two types of processes: a JobManager and one or more TaskManager. ###### kafka-connector-requirements Explore the necessary settings for standard and upsert Kafka connectors in Aiven for Apache Flink®. - [Settings for Apache Kafka® connectors](/products/flink/concepts/kafka-connector-requirements.md): Explore the necessary settings for standard and upsert Kafka connectors in Aiven for Apache Flink®. ###### kafka-connectors In addition to integration with Apache Kafka® through a standard connector, Aiven for Apache Flink® also supports the use of upsert connectors, which allows you to create changelog-type data streams. - [Standard and upsert connectors for Apache Kafka®](/products/flink/concepts/kafka-connectors.md): In addition to integration with Apache Kafka® through a standard connector, Aiven for Apache Flink® also supports the use of upsert connectors, which allows you to create changelog-type data streams. ###### savepoints Savepoints in Aiven for Apache Flink® are snapshots of the current state of your Flink application. - [Savepoints](/products/flink/concepts/savepoints.md): Savepoints in Aiven for Apache Flink® are snapshots of the current state of your Flink application. ###### supported-syntax-sql-editor The built-in Table SQL editor in the [Aiven - [Built-in SQL editor](/products/flink/concepts/supported-syntax-sql-editor.md): The built-in Table SQL editor in the [Aiven ###### tables With Aiven for Apache Flink®, you can create and manage data pipelines using Flink tables. - [Tables in Aiven for Apache Flink®](/products/flink/concepts/tables.md): With Aiven for Apache Flink®, you can create and manage data pipelines using Flink tables. ###### watermarks Apache Flink® uses watermarks to synchronize and process events in data streams accurately. These watermarks are timestamps embedded in the data stream that track the progression of event time. - [Watermarks](/products/flink/concepts/watermarks.md): Apache Flink® uses watermarks to synchronize and process events in data streams accurately. These watermarks are timestamps embedded in the data stream that track the progression of event time. ###### windows Apache Flink® uses the concept of windows to manage the continuous flow of data in streams by segmenting it into manageable chunks. This approach is essential due to the continuous and unbounded nature of data streams, where waiting for all data to arrive is impractical. - [Windows](/products/flink/concepts/windows.md): Apache Flink® uses the concept of windows to manage the continuous flow of data in streams by segmenting it into manageable chunks. This approach is essential due to the continuous and unbounded nature of data streams, where waiting for all data to arrive is impractical. ##### get-started Begin your experience with Aiven for Apache Flink® by setting up a service, configuring data integrations, and building streaming applications. - [Get started with Aiven for Apache Flink®](/products/flink/get-started.md): Begin your experience with Aiven for Apache Flink® by setting up a service, configuring data integrations, and building streaming applications. ##### howto ###### connect-bigquery Connect Aiven for Apache Flink® with Google BigQuery as a sink using the Aiven client or the Aiven Console. - [Integrate Aiven for Apache Flink® with Google BigQuery](/products/flink/howto/connect-bigquery.md): Connect Aiven for Apache Flink® with Google BigQuery as a sink using the Aiven client or the Aiven Console. ###### connect-kafka To build data pipelines, Apache Flink® requires you to map source and target data structures as Flink tables within an application. - [Create an Apache Kafka®-based Apache Flink® table](/products/flink/howto/connect-kafka.md): To build data pipelines, Apache Flink® requires you to map source and target data structures as Flink tables within an application. ###### connect-opensearch To build data pipelines, Apache Flink® requires you to map source and target data structures as Flink tables within an application. - [Create an OpenSearch®-based Apache Flink® table](/products/flink/howto/connect-opensearch.md): To build data pipelines, Apache Flink® requires you to map source and target data structures as Flink tables within an application. ###### connect-pg To build data pipelines, Apache Flink® requires source and target data - [Create a PostgreSQL®-based Apache Flink® table](/products/flink/howto/connect-pg.md): To build data pipelines, Apache Flink® requires source and target data ###### create-flink-applications Aiven for Flink applications in Aiven for Apache Flink® servers as a container that includes everything connected to a Flink job, including source and sink connections and data processing logic. - [Use Aiven for Apache Flink® applications](/products/flink/howto/create-flink-applications.md): Aiven for Flink applications in Aiven for Apache Flink® servers as a container that includes everything connected to a Flink job, including source and sink connections and data processing logic. ###### create-integration With Aiven for Apache Flink®, you can create streaming data pipelines to connect various services. - [Create Apache Flink® data service integrations](/products/flink/howto/create-integration.md): With Aiven for Apache Flink®, you can create streaming data pipelines to connect various services. ###### create-jar-application Aiven for Apache Flink® enables you to upload and deploy custom code as a JAR file, enhancing your Apache Flink applications with advanced data processing capabilities. - [Create a JAR application Limited availability](/products/flink/howto/create-jar-application.md): Aiven for Apache Flink® enables you to upload and deploy custom code as a JAR file, enhancing your Apache Flink applications with advanced data processing capabilities. ###### create-sql-application Build data processing pipelines in Aiven for Apache Flink® by creating SQL applications using Apache Flink SQL. Set up source and sink tables, define processing logic, and manage your deployments. - [Create an SQL application](/products/flink/howto/create-sql-application.md): Build data processing pipelines in Aiven for Apache Flink® by creating SQL applications using Apache Flink SQL. Set up source and sink tables, define processing logic, and manage your deployments. ###### datagen-connector The DataGen source table is a built-in connector of the Apache Flink system that generates random data periodically, matching the specified data type of the source table. - [Create a DataGen-based Apache Flink® table](/products/flink/howto/datagen-connector.md): The DataGen source table is a built-in connector of the Apache Flink system that generates random data periodically, matching the specified data type of the source table. ###### ext-kafka-flink-integration Integrating external/self-hosted Apache Kafka® with Aiven for Apache Flink® allows users to leverage the power of both technologies to build scalable and robust real-time streaming applications. - [Integrate Aiven for Apache Flink® with Apache Kafka®](/products/flink/howto/ext-kafka-flink-integration.md): Integrating external/self-hosted Apache Kafka® with Aiven for Apache Flink® allows users to leverage the power of both technologies to build scalable and robust real-time streaming applications. ###### flink-confluent-avro Confluent Avro is a serialization format that requires integrating with a schema registry. - [Create Confluent Avro-based Apache Flink® table](/products/flink/howto/flink-confluent-avro.md): Confluent Avro is a serialization format that requires integrating with a schema registry. ###### manage-credentials-jars Learn how to use the AVNCREDENTIALSDIR environment variable to securely manage credentials for custom JARs in Aiven for Apache Flink®. - [Credential management for JAR applications](/products/flink/howto/manage-credentials-jars.md): Learn how to use the AVNCREDENTIALSDIR environment variable to securely manage credentials for custom JARs in Aiven for Apache Flink®. ###### manage-flink-applications This section provides information on managing your Aiven for Apache Flink® applications. - [Manage Aiven for Apache Flink® applications](/products/flink/howto/manage-flink-applications.md): This section provides information on managing your Aiven for Apache Flink® applications. ###### manage-flink-tables Aiven for Apache Flink® allows you to map source and target data structures as Flink tables and use transformation statements to reshape, filter or aggregate data. - [Manage tables in Aiven for Apache Flink® applications](/products/flink/howto/manage-flink-tables.md): Aiven for Apache Flink® allows you to map source and target data structures as Flink tables and use transformation statements to reshape, filter or aggregate data. ###### pg-cdc-connector Change Data Capture (CDC) is a technique that enables the tracking and capturing of changes made to data within a PostgreSQL® database. - [Create a PostgreSQL® CDC connector-based Apache Flink®](/products/flink/howto/pg-cdc-connector.md): Change Data Capture (CDC) is a technique that enables the tracking and capturing of changes made to data within a PostgreSQL® database. ###### restart-strategy-jar-applications Learn how Aiven for Apache Flink® applications uses restart strategies to recover from job failures, ensuring high availability and fault tolerance for your distributed applications. - [Restart strategy in SQL and JAR applications](/products/flink/howto/restart-strategy-jar-applications.md): Learn how Aiven for Apache Flink® applications uses restart strategies to recover from job failures, ensuring high availability and fault tolerance for your distributed applications. ###### slack-connector With Aiven's Slack Connector for Apache Flink®, you can create sink - [Create a Slack-based Apache Flink® table](/products/flink/howto/slack-connector.md): With Aiven's Slack Connector for Apache Flink®, you can create sink ###### timestamps_opensearch Frequently results in Apache Flink® data pipelines include one or more timestamps, either contained in the source events or generated by window aggregations. - [Define OpenSearch® timestamp data in SQL pipeline](/products/flink/howto/timestamps_opensearch.md): Frequently results in Apache Flink® data pipelines include one or more timestamps, either contained in the source events or generated by window aggregations. ###### upgrade-flink-version Upgrading to the latest version of Aiven for Apache Flink® allows you to benefit from improved features, enhanced performance, and better security. - [Upgrade Aiven for Apache Flink](/products/flink/howto/upgrade-flink-version.md): Upgrading to the latest version of Aiven for Apache Flink® allows you to benefit from improved features, enhanced performance, and better security. ##### reference ###### advanced-params See the configuration options available for Aiven for Apache Flink®: - [Advanced parameters for Aiven for Apache Flink®](/products/flink/reference/advanced-params.md): See the configuration options available for Aiven for Apache Flink®: ###### flink-limitations Because Aiven for Apache Flink is a fully managed service, there are differences between Aiven for Apache Flink and an Apache Flink service you run yourself. - [Aiven for Apache Flink® limitation](/products/flink/reference/flink-limitations.md): Because Aiven for Apache Flink is a fully managed service, there are differences between Aiven for Apache Flink and an Apache Flink service you run yourself. #### grafana Aiven for Grafana® is a fully managed analytics and monitoring solution, deployable in the cloud of your choice, which can bring unlimited scalability and high availability to your monitoring environment and other time series applications. - [Aiven for Grafana®](/products/grafana.md): Aiven for Grafana® is a fully managed analytics and monitoring solution, deployable in the cloud of your choice, which can bring unlimited scalability and high availability to your monitoring environment and other time series applications. ##### backups-migration Protect and recover your Aiven for Grafana® service data with cross-region backups and - [Backups and migration in Aiven for Grafana®](/products/grafana/backups-migration.md): Protect and recover your Aiven for Grafana® service data with cross-region backups and ##### concepts ###### service-memory Understand the memory limits and out-of-memory conditions that apply to your Aiven for Grafana® service. - [Memory and out-of-memory conditions in Aiven for Grafana®](/products/grafana/concepts/service-memory.md): Understand the memory limits and out-of-memory conditions that apply to your Aiven for Grafana® service. ##### get-started To start using Aiven for Grafana, the first step is to create a service. You can do this in the Aiven Console or with the Aiven CLI. - [Get started with Aiven for Grafana®](/products/grafana/get-started.md): To start using Aiven for Grafana, the first step is to create a service. You can do this in the Aiven Console or with the Aiven CLI. ##### howto ###### backup-to-another-region Copy your Aiven for Grafana® service backups to a secondary region for disaster recovery. - [Back up your Aiven for Grafana® service to another region](/products/grafana/howto/backup-to-another-region.md): Copy your Aiven for Grafana® service backups to a secondary region for disaster recovery. ###### change-cloud-region Move your Aiven for Grafana® service to a different cloud provider or region. - [Change the cloud or region for your Aiven for Grafana® service](/products/grafana/howto/change-cloud-region.md): Move your Aiven for Grafana® service to a different cloud provider or region. ###### change-service-plan Change the service plan for your Aiven for Grafana® service to scale resources up or down and optimize costs. - [Change the plan for your Aiven for Grafana® service](/products/grafana/howto/change-service-plan.md): Change the service plan for your Aiven for Grafana® service to scale resources up or down and optimize costs. ###### controlled-upgrade-pipelines Link Aiven for Grafana® services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. - [Controlled upgrade pipelines for your Aiven for Grafana® service Limited availability](/products/grafana/howto/controlled-upgrade-pipelines.md): Link Aiven for Grafana® services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. ###### dashboard-previews Grafana's dashboard previews provide a visual overview of your dashboards, displaying each configured dashboard as a graphical thumbnail. - [Dashboard preview for Aiven for Grafana®](/products/grafana/howto/dashboard-previews.md): Grafana's dashboard previews provide a visual overview of your dashboards, displaying each configured dashboard as a graphical thumbnail. ###### disk-autoscaler Automatically increase the disk storage of your Aiven for Grafana® service when it's running out of space, instead of resizing it manually. - [Scale disk storage automatically for your Aiven for Grafana® service](/products/grafana/howto/disk-autoscaler.md): Automatically increase the disk storage of your Aiven for Grafana® service when it's running out of space, instead of resizing it manually. ###### fork-service Fork your Aiven for Grafana® service to create an independent copy for testing, - [Fork your Aiven for Grafana® service](/products/grafana/howto/fork-service.md): Fork your Aiven for Grafana® service to create an independent copy for testing, ###### maintenance-updates Manage maintenance updates and set the maintenance window for your Aiven for Grafana® service. - [Maintenance and updates for your Aiven for Grafana® service](/products/grafana/howto/maintenance-updates.md): Manage maintenance updates and set the maintenance window for your Aiven for Grafana® service. ###### oauth-configuration Grafana version 9.5.5 introduced significant changes to the OAuth email lookup behavior to enhance security. However, some users may need to revert to the previous behavior as seen in Grafana 9.5.3. - [Aiven for Grafana® OAuth configuration and security considerations](/products/grafana/howto/oauth-configuration.md): Grafana version 9.5.5 introduced significant changes to the OAuth email lookup behavior to enhance security. However, some users may need to revert to the previous behavior as seen in Grafana 9.5.3. ###### power-cycle-service Power off your Aiven for Grafana® service to release resources and save credits, power it - [Power on/off and delete your Aiven for Grafana® service](/products/grafana/howto/power-cycle-service.md): Power off your Aiven for Grafana® service to release resources and save credits, power it ###### prepare-for-high-load Prepare your Aiven for Grafana® service for higher than usual traffic to avoid outages and keep performance stable. - [Prepare your Aiven for Grafana® service for high load](/products/grafana/howto/prepare-for-high-load.md): Prepare your Aiven for Grafana® service for higher than usual traffic to avoid outages and keep performance stable. ###### rename-service Change the name of your Aiven for Grafana® service by forking it under a new name and - [Rename your Aiven for Grafana® service](/products/grafana/howto/rename-service.md): Change the name of your Aiven for Grafana® service by forking it under a new name and ###### replace-expression-string Sometimes, it is useful to replace all occurrences of a string in Grafana metric expressions. - [Replace strings in Grafana® dashboard metric expressions](/products/grafana/howto/replace-expression-string.md): Sometimes, it is useful to replace all occurrences of a string in Grafana metric expressions. ###### rotating-grafana-service-credentials For improved security, it is recommended to periodically update your - [Update Aiven for Grafana® service credentials](/products/grafana/howto/rotating-grafana-service-credentials.md): For improved security, it is recommended to periodically update your ###### scale-disk-storage Scale the disk storage of your Aiven for Grafana® service up or down without disrupting the running service. - [Scale disk storage for your Aiven for Grafana® service](/products/grafana/howto/scale-disk-storage.md): Scale the disk storage of your Aiven for Grafana® service up or down without disrupting the running service. ###### send-emails Use the Aiven API or the Aiven client to configure the Simple Mail Transfer Protocol (SMTP) server settings and send the following emails from Aiven for Grafana: invite emails, reset password emails, and alert messages. - [Send emails from Aiven for Grafana®](/products/grafana/howto/send-emails.md): Use the Aiven API or the Aiven client to configure the Simple Mail Transfer Protocol (SMTP) server settings and send the following emails from Aiven for Grafana: invite emails, reset password emails, and alert messages. ###### tag-service Add key-value tags to your Aiven for Grafana® service to organize services and track - [Tag your Aiven for Grafana® service](/products/grafana/howto/tag-service.md): Add key-value tags to your Aiven for Grafana® service to organize services and track ###### track-restore-progress Track the restore progress of individual nodes in your Aiven for Grafana® service during node replacement, forking, or maintenance, using the Aiven API. - [Track restore progress for your Aiven for Grafana® service](/products/grafana/howto/track-restore-progress.md): Track the restore progress of individual nodes in your Aiven for Grafana® service during node replacement, forking, or maintenance, using the Aiven API. ##### maintenance-lifecycle Manage maintenance updates and the maintenance window for your Aiven for Grafana® service. - [Maintenance and lifecycle in Aiven for Grafana®](/products/grafana/maintenance-lifecycle.md): Manage maintenance updates and the maintenance window for your Aiven for Grafana® service. ##### reference ###### advanced-params See the configuration options available for - [Advanced parameters for Aiven for Grafana®](/products/grafana/reference/advanced-params.md): See the configuration options available for ###### plugins Aiven for Grafana includes several pre-installed plugins, updated regularly to ensure optimal performance. - [Plugins for Aiven for Grafana®](/products/grafana/reference/plugins.md): Aiven for Grafana includes several pre-installed plugins, updated regularly to ensure optimal performance. ###### version-lifecycle Learn how Aiven manages the Aiven for Grafana® service version, end of life (EOL) dates, and what happens to your service after the version reaches EOL. - [Aiven for Grafana® version lifecycle](/products/grafana/reference/version-lifecycle.md): Learn how Aiven manages the Aiven for Grafana® service version, end of life (EOL) dates, and what happens to your service after the version reaches EOL. ##### scaling-performance Scale resources and prepare your Aiven for Grafana® service for changing load. - [Scaling and performance in Aiven for Grafana®](/products/grafana/scaling-performance.md): Scale resources and prepare your Aiven for Grafana® service for changing load. #### kafka Aiven for Apache Kafka® is a fully managed Apache Kafka service for building event-driven applications, data pipelines, and stream processing systems. - [Aiven for Apache Kafka®](/products/kafka.md): Aiven for Apache Kafka® is a fully managed Apache Kafka service for building event-driven applications, data pipelines, and stream processing systems. ##### classic-kafka-overview Classic Kafka is an Aiven for Apache Kafka® service type that uses fixed plans with local broker storage. - [Classic Kafka overview](/products/kafka/classic-kafka-overview.md): Classic Kafka is an Aiven for Apache Kafka® service type that uses fixed plans with local broker storage. ##### concepts ###### acl Access Control Lists (ACLs) in Aiven for Apache Kafka® manage access to topics, consumer groups, clusters, and the Schema Registry. - [Access Control Lists in Aiven for Apache Kafka®](/products/kafka/concepts/acl.md): Access Control Lists (ACLs) in Aiven for Apache Kafka® manage access to topics, consumer groups, clusters, and the Schema Registry. ###### audit-logging Audit logging for Aiven for Apache Kafka® records Kafka client activity on your service. - [Audit logging for Aiven for Apache Kafka®](/products/kafka/concepts/audit-logging.md): Audit logging for Aiven for Apache Kafka® records Kafka client activity on your service. ###### auth-types Use modern encryption protocols to protect data in transit with Apache Kafka®. Aiven for Apache Kafka® provides multiple options to secure your data. - [Authentication types](/products/kafka/concepts/auth-types.md): Use modern encryption protocols to protect data in transit with Apache Kafka®. Aiven for Apache Kafka® provides multiple options to secure your data. ###### configuration-backup Aiven for Apache Kafka® includes configuration backups that automatically back up key service configurations at no additional cost. - [Configuration backups for Aiven for Apache Kafka®](/products/kafka/concepts/configuration-backup.md): Aiven for Apache Kafka® includes configuration backups that automatically back up key service configurations at no additional cost. ###### consumer-lag-predictor The consumer lag predictor for Aiven for Apache Kafka estimates the delay between the time a message is produced and when it's eventually consumed by a consumer group. - [Consumer lag predictor for Aiven for Apache Kafka® Limited availability](/products/kafka/concepts/consumer-lag-predictor.md): The consumer lag predictor for Aiven for Apache Kafka estimates the delay between the time a message is produced and when it's eventually consumed by a consumer group. ###### follower-fetching Follower fetching in Aiven for Apache Kafka allows consumers to retrieve data from the nearest replica instead of always fetching from the partition leader. - [Follower fetching in Aiven for Apache Kafka®](/products/kafka/concepts/follower-fetching.md): Follower fetching in Aiven for Apache Kafka allows consumers to retrieve data from the nearest replica instead of always fetching from the partition leader. ###### governance-overview Governance in Aiven for Apache Kafka® provides a structured and secure way to manage your Aiven for Apache Kafka clusters, ensuring security, compliance, and efficiency. - [Aiven for Apache Kafka® governance Limited availability](/products/kafka/concepts/governance-overview.md): Governance in Aiven for Apache Kafka® provides a structured and secure way to manage your Aiven for Apache Kafka clusters, ensuring security, compliance, and efficiency. ###### horizontal-vertical-scaling Aiven for Apache Kafka® has a number of predefined plans that specify the number of brokers and the capacity of individual brokers. - [Scaling options in Apache Kafka®](/products/kafka/concepts/horizontal-vertical-scaling.md): Aiven for Apache Kafka® has a number of predefined plans that specify the number of brokers and the capacity of individual brokers. ###### kafka-pricing You pay for the Kafka service plan, retained data, and traffic through your topics. - [Pricing for Aiven for Apache Kafka®](/products/kafka/concepts/kafka-pricing.md): You pay for the Kafka service plan, retained data, and traffic through your topics. ###### kafka-quotas Quotas ensure fair resource allocation, stability, and efficiency in your Kafka cluster. - [Quotas in Aiven for Apache Kafka®](/products/kafka/concepts/kafka-quotas.md): Quotas ensure fair resource allocation, stability, and efficiency in your Kafka cluster. ###### kafka-rest-api Produce and consume Apache Kafka messages over HTTP with the Karapace REST proxy. - [Apache Kafka® REST API](/products/kafka/concepts/kafka-rest-api.md): Produce and consume Apache Kafka messages over HTTP with the Karapace REST proxy. ###### kafka-tiered-storage Tiered storage in Aiven for Apache Kafka® helps you manage data more effectively by using two different storage types: local disk and remote cloud storage solutions like AWS S3, Google Cloud Storage, and Azure Blob Storage. - [Tiered storage in Aiven for Apache Kafka® overview](/products/kafka/concepts/kafka-tiered-storage.md): Tiered storage in Aiven for Apache Kafka® helps you manage data more effectively by using two different storage types: local disk and remote cloud storage solutions like AWS S3, Google Cloud Storage, and Azure Blob Storage. ###### kraft-mode Starting with Apache Kafka 3.9, Aiven for Apache Kafka uses KRaft (Kafka Raft) to manage metadata and controllers, replacing ZooKeeper. - [KRaft in Aiven for Apache Kafka®](/products/kafka/concepts/kraft-mode.md): Starting with Apache Kafka 3.9, Aiven for Apache Kafka uses KRaft (Kafka Raft) to manage metadata and controllers, replacing ZooKeeper. ###### list-kafka-tiered-storage Discover how tiered storage works in Aiven for Apache Kafka®, explore - [Tiered storage in Aiven for Apache Kafka® Early availability](/products/kafka/concepts/list-kafka-tiered-storage.md): Discover how tiered storage works in Aiven for Apache Kafka®, explore ###### log-compaction One way to reduce the disk space requirements in Apache Kafka® is to use compacted topics. - [Compacted topics](/products/kafka/concepts/log-compaction.md): One way to reduce the disk space requirements in Apache Kafka® is to use compacted topics. ###### monitor-consumer-group With Aiven for Apache Kafka® dashboards and telemetry, you can monitor the performance and system resources of your Aiven for Apache Kafka service. - [Monitoring consumer groups in Aiven for Apache Kafka®](/products/kafka/concepts/monitor-consumer-group.md): With Aiven for Apache Kafka® dashboards and telemetry, you can monitor the performance and system resources of your Aiven for Apache Kafka service. ###### partition-segments Apache Kafka® divides topics partition data into segment files (with .log suffix) stored on the file system. Each segment file is named using the offset of the first message (a.k.a. base offset) contained. - [Partition segments](/products/kafka/concepts/partition-segments.md): Apache Kafka® divides topics partition data into segment files (with .log suffix) stored on the file system. Each segment file is named using the offset of the first message (a.k.a. base offset) contained. ###### tiered-storage-guarantees With Aiven for Apache Kafka®'s tiered storage, there are two primary types of data retention guarantees: total retention and local retention. - [Guarantees](/products/kafka/concepts/tiered-storage-guarantees.md): With Aiven for Apache Kafka®'s tiered storage, there are two primary types of data retention guarantees: total retention and local retention. ###### tiered-storage-how-it-works Aiven for Apache Kafka® tiered storage optimizes data management across two distinct storage tiers: - [How tiered storage works in Aiven for Apache Kafka®](/products/kafka/concepts/tiered-storage-how-it-works.md): Aiven for Apache Kafka® tiered storage optimizes data management across two distinct storage tiers: ###### tiered-storage-limitations The main trade-off of tiered storage is the higher latency when accessing and reading data from remote storage compared to local disk storage. - [Trade-offs and limitations](/products/kafka/concepts/tiered-storage-limitations.md): The main trade-off of tiered storage is the higher latency when accessing and reading data from remote storage compared to local disk storage. ###### topic-catalog-overview The Aiven for Apache Kafka® topic catalog provides a centralized interface within the Aiven console to view and manage Apache Kafka topics across different projects and services. - [Aiven for Apache Kafka® topic catalog Early availability](/products/kafka/concepts/topic-catalog-overview.md): The Aiven for Apache Kafka® topic catalog provides a centralized interface within the Aiven console to view and manage Apache Kafka topics across different projects and services. ###### upgrade-procedure Aiven for Apache Kafka® provides an automated upgrade process during the following operations: - [Apache Kafka® upgrade procedure](/products/kafka/concepts/upgrade-procedure.md): Aiven for Apache Kafka® provides an automated upgrade process during the following operations: ##### dev-tier ###### create-dev-tier-kafka-service Create an Aiven for Apache Kafka® Developer tier service in the Aiven Console or with Skills. - [Create an Aiven for Apache Kafka® Developer tier service](/products/kafka/dev-tier/create-dev-tier-kafka-service.md): Create an Aiven for Apache Kafka® Developer tier service in the Aiven Console or with Skills. ###### kafka-dev-tier Aiven for Apache Kafka® Developer tier is a paid service tier for Classic Kafka that sits between the Free tier and Professional tier. - [Aiven for Apache Kafka® Developer tier](/products/kafka/dev-tier/kafka-dev-tier.md): Aiven for Apache Kafka® Developer tier is a paid service tier for Classic Kafka that sits between the Free tier and Professional tier. ##### diskless ###### concepts - [Batching and delivery in diskless topics](/products/kafka/diskless/concepts/batching-and-delivery.md): Diskless topics use a batching-based delivery model designed for cloud-native environments. - [Diskless topics for Apache Kafka®](/products/kafka/diskless/concepts/diskless-topic-overview.md): Diskless topics are a feature of Standard Kafka services for Aiven for Apache Kafka®. - [Diskless topics architecture](/products/kafka/diskless/concepts/diskless-topics-architecture.md): Diskless topics extend the Apache Kafka® storage model by replacing local disk storage with cloud object storage. - [Diskless topic limitations and behavior](/products/kafka/diskless/concepts/limitations.md): Diskless topics are compatible with Kafka APIs and clients, with some limitations: - [Partitions and objects in Diskless Topics](/products/kafka/diskless/concepts/partitions-and-objects.md): Diskless topics use the standard Kafka partitioning model but store data in cloud object storage instead of broker-local disks. - [Compare diskless and classic Apache Kafka® topics](/products/kafka/diskless/concepts/topics-vs-classic.md): Diskless topics are Apache Kafka®-compatible topics that store data in cloud object storage instead of broker-managed local disks. ###### howto - [Create diskless topics automatically using regular expressions](/products/kafka/diskless/howto/create-diskless-topics-automatically.md): Configure your Aiven for Apache Kafka® service to automatically create new topics as diskless topics when their names match configured regular expressions. This lets clients and connectors use diskless topics without setting diskless.enable=true in each create-topic request. ##### free-tier ###### create-free-tier-kafka-service You can create a free tier Aiven for Apache Kafka® service to learn Kafka, test producers and consumers, or run small proof-of-concept workloads. - [Create a free tier Aiven for Apache Kafka® service](/products/kafka/free-tier/create-free-tier-kafka-service.md): You can create a free tier Aiven for Apache Kafka® service to learn Kafka, test producers and consumers, or run small proof-of-concept workloads. ###### kafka-free-tier Get started with Apache Kafka® at no cost. The Aiven for Apache Kafka® free tier is a fully managed service for learning, prototyping, and evaluation. No credit card is required. - [Aiven for Apache Kafka® free tier](/products/kafka/free-tier/kafka-free-tier.md): Get started with Apache Kafka® at no cost. The Aiven for Apache Kafka® free tier is a fully managed service for learning, prototyping, and evaluation. No credit card is required. ##### get-started ###### create-kafka-service Create an Aiven for Apache Kafka® Professional tier service on Aiven Cloud. - [Create an Aiven for Apache Kafka® Professional tier service](/products/kafka/get-started/create-kafka-service.md): Create an Aiven for Apache Kafka® Professional tier service on Aiven Cloud. ###### create-kafka-service-byoc Create an Aiven for Apache Kafka® service in your own cloud account with Bring Your Own Cloud (BYOC). - [Create an Apache Kafka® service with bring your own cloud (BYOC)](/products/kafka/get-started/create-kafka-service-byoc.md): Create an Aiven for Apache Kafka® service in your own cloud account with Bring Your Own Cloud (BYOC). ###### get-started-kafka Create a managed Apache Kafka® service on Aiven. - [Get started with Aiven for Apache Kafka®](/products/kafka/get-started/get-started-kafka.md): Create a managed Apache Kafka® service on Aiven. ###### professional-tier The Professional tier is for production and workload-heavy Aiven for Apache Kafka® services. - [Aiven for Apache Kafka® Professional tier](/products/kafka/get-started/professional-tier.md): The Professional tier is for production and workload-heavy Aiven for Apache Kafka® services. ##### howto ###### add-manage-service-users Create and manage service users in Aiven for Apache Kafka® to enable secure access and interaction with your service. - [Manage service users in Aiven for Apache Kafka®](/products/kafka/howto/add-manage-service-users.md): Create and manage service users in Aiven for Apache Kafka® to enable secure access and interaction with your service. ###### add-missing-producer-consumer-metrics When you enable the - [Add client-side Apache Kafka® producer and consumer Datadog metrics](/products/kafka/howto/add-missing-producer-consumer-metrics.md): When you enable the ###### approvals The Approvals page allows you to manage requests for Aiven for Apache Kafka® resources owned by your group. - [Manage approvals Limited availability](/products/kafka/howto/approvals.md): The Approvals page allows you to manage requests for Aiven for Apache Kafka® resources owned by your group. ###### avoid-out-of-memory-error When a node in an Aiven for Apache Kafka® or Aiven for Apache Kafka® Connect cluster runs low on memory, the Java virtual machine (JVM) running the service may not be able to allocate the memory, and will raise a java.lang.OutOfMemoryError exception. - [Avoid OutOfMemoryError errors in Aiven for Apache Kafka®](/products/kafka/howto/avoid-out-of-memory-error.md): When a node in an Aiven for Apache Kafka® or Aiven for Apache Kafka® Connect cluster runs low on memory, the Java virtual machine (JVM) running the service may not be able to allocate the memory, and will raise a java.lang.OutOfMemoryError exception. ###### best-practices Follow these best practices to optimize the performance and reliability of your Aiven for Apache Kafka® service. - [Optimize Apache Kafka® performance](/products/kafka/howto/best-practices.md): Follow these best practices to optimize the performance and reliability of your Aiven for Apache Kafka® service. ###### change-retention-period To avoid running out of disk space, by default, Apache Kafka® drops the oldest messages from the beginning of each log after their retention period expires. - [Change data retention period](/products/kafka/howto/change-retention-period.md): To avoid running out of disk space, by default, Apache Kafka® drops the oldest messages from the beginning of each log after their retention period expires. ###### change-service-plan Change the service plan for your Aiven for Apache Kafka® service to scale resources up or down and optimize costs. - [Change the plan for your Aiven for Apache Kafka® service](/products/kafka/howto/change-service-plan.md): Change the service plan for your Aiven for Apache Kafka® service to scale resources up or down and optimize costs. ###### change-standard-kafka-plan Change the service plan for your Standard Kafka service to scale resources up or down and optimize costs. - [Change the plan for your Standard Kafka service](/products/kafka/howto/change-standard-kafka-plan.md): Change the service plan for your Standard Kafka service to scale resources up or down and optimize costs. ###### claim-topic To take ownership of a topic that your group does not currently own, you can submit a claim request. Once you send the request, the current owner can approve or decline it. - [Claim topic ownership Limited availability](/products/kafka/howto/claim-topic.md): To take ownership of a topic that your group does not currently own, you can submit a claim request. Once you send the request, the current owner can approve or decline it. ###### configure-audit-logging Turn audit logging on for your Aiven for Apache Kafka® service, change what it records, and manage audit log volume. - [Configure audit logging for Aiven for Apache Kafka®](/products/kafka/howto/configure-audit-logging.md): Turn audit logging on for your Aiven for Apache Kafka® service, change what it records, and manage audit log volume. ###### configure-custom-domain Configure a custom domain to replace the default Aiven service hostname for Kafka REST API, Schema Registry, and Kafka Connect. - [Configure a custom domain for Kafka REST API, Schema Registry, and Kafka Connect](/products/kafka/howto/configure-custom-domain.md): Configure a custom domain to replace the default Aiven service hostname for Kafka REST API, Schema Registry, and Kafka Connect. ###### configure-log-cleaner The log cleaner serves the purpose of preserving only the latest value associated with a specific message key in a partition for compacted topics. - [Configure the log cleaner for topic compaction](/products/kafka/howto/configure-log-cleaner.md): The log cleaner serves the purpose of preserving only the latest value associated with a specific message key in a partition for compacted topics. ###### configure-preferred-zones Configure preferred availability zones for Aiven for Apache Kafka®, Aiven for Apache Kafka® Connect, and Aiven for Apache Kafka® MirrorMaker 2. - [Configure preferred availability zones](/products/kafka/howto/configure-preferred-zones.md): Configure preferred availability zones for Aiven for Apache Kafka®, Aiven for Apache Kafka® Connect, and Aiven for Apache Kafka® MirrorMaker 2. ###### configure-topic-tiered-storage Aiven for Apache Kafka® allows you to configure tiered storage and set retention policies for individual topics. - [Enable and configure tiered storage for topics](/products/kafka/howto/configure-topic-tiered-storage.md): Aiven for Apache Kafka® allows you to configure tiered storage and set retention policies for individual topics. ###### configure-with-kafka-cli Aiven for Apache Kafka® services are fully manageable and customizable via the Aiven CLI. - [Manage configurations with Apache Kafka® CLI tools](/products/kafka/howto/configure-with-kafka-cli.md): Aiven for Apache Kafka® services are fully manageable and customizable via the Aiven CLI. ###### connect-with-command-line Use Quick connect to set up Apache Kafka® command-line tools for Aiven for - [Connect to Aiven for Apache Kafka® with command-line tools](/products/kafka/howto/connect-with-command-line.md): Use Quick connect to set up Apache Kafka® command-line tools for Aiven for ###### connect-with-cpp Use Quick connect to set up a C++ client for Aiven for Apache Kafka®. - [Connect to Aiven for Apache Kafka® with C++](/products/kafka/howto/connect-with-cpp.md): Use Quick connect to set up a C++ client for Aiven for Apache Kafka®. ###### connect-with-csharp Use Quick connect to set up a C# client for Aiven for Apache Kafka®. - [Connect to Aiven for Apache Kafka® with C#](/products/kafka/howto/connect-with-csharp.md): Use Quick connect to set up a C# client for Aiven for Apache Kafka®. ###### connect-with-go Use Quick connect to set up a Go client for Aiven for Apache Kafka®. - [Connect to Aiven for Apache Kafka® with Go](/products/kafka/howto/connect-with-go.md): Use Quick connect to set up a Go client for Aiven for Apache Kafka®. ###### connect-with-java Use Quick connect to set up a Java client for Aiven for Apache Kafka®. - [Connect to Aiven for Apache Kafka® with Java](/products/kafka/howto/connect-with-java.md): Use Quick connect to set up a Java client for Aiven for Apache Kafka®. ###### connect-with-kafka-rest Use Quick connect to produce and consume messages with the Apache Kafka® - [Connect to Aiven for Apache Kafka® with Kafka REST](/products/kafka/howto/connect-with-kafka-rest.md): Use Quick connect to produce and consume messages with the Apache Kafka® ###### connect-with-klaw Use Quick connect to set up Klaw for Aiven - [Connect to Aiven for Apache Kafka® with Klaw](/products/kafka/howto/connect-with-klaw.md): Use Quick connect to set up Klaw for Aiven ###### connect-with-nodejs Use Quick connect to set up a Node.js client for Aiven for Apache Kafka®. - [Connect to Aiven for Apache Kafka® with Node.js](/products/kafka/howto/connect-with-nodejs.md): Use Quick connect to set up a Node.js client for Aiven for Apache Kafka®. ###### connect-with-php Use Quick connect to set up a PHP client for Aiven for Apache Kafka®. - [Connect to Aiven for Apache Kafka® with PHP](/products/kafka/howto/connect-with-php.md): Use Quick connect to set up a PHP client for Aiven for Apache Kafka®. ###### connect-with-python Use Quick connect to set up a Python client for Aiven for Apache Kafka®. - [Connect to Aiven for Apache Kafka® with Python](/products/kafka/howto/connect-with-python.md): Use Quick connect to set up a Python client for Aiven for Apache Kafka®. ###### controlled-upgrade-pipelines Link Aiven for Apache Kafka® services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. - [Controlled upgrade pipelines for your Aiven for Apache Kafka® service Limited availability](/products/kafka/howto/controlled-upgrade-pipelines.md): Link Aiven for Apache Kafka® services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. ###### create-topic Create topics in your Aiven for Apache Kafka® service to organize message streams between producers and consumers. - [Create Apache Kafka® topics](/products/kafka/howto/create-topic.md): Create topics in your Aiven for Apache Kafka® service to organize message streams between producers and consumers. ###### create-topics-automatically When you send a message to a topic that does not exist, Apache Kafka® can - [Create Apache Kafka® topics automatically](/products/kafka/howto/create-topics-automatically.md): When you send a message to a topic that does not exist, Apache Kafka® can ###### datadog-customised-metrics When you configure a Datadog service integration for Aiven for Apache Kafka®, Aiven sends Kafka metrics to Datadog. - [Apache Kafka® metrics sent to Datadog](/products/kafka/howto/datadog-customised-metrics.md): When you configure a Datadog service integration for Aiven for Apache Kafka®, Aiven sends Kafka metrics to Datadog. ###### disk-autoscaler Automatically increase the disk storage of your Aiven for Apache Kafka® service when it's running out of space, instead of resizing it manually. - [Scale disk storage automatically for your Aiven for Apache Kafka® service](/products/kafka/howto/disk-autoscaler.md): Automatically increase the disk storage of your Aiven for Apache Kafka® service when it's running out of space, instead of resizing it manually. ###### enable-follower-fetching Enabling follower fetching in Aiven for Apache Kafka® allows your consumers to fetch data from the nearest replica instead of the leader, optimizing data fetching and enhancing performance. - [Enable follower fetching in Aiven for Apache Kafka®](/products/kafka/howto/enable-follower-fetching.md): Enabling follower fetching in Aiven for Apache Kafka® allows your consumers to fetch data from the nearest replica instead of the leader, optimizing data fetching and enhancing performance. ###### enable-governance Enable governance in Aiven for Apache Kafka® to create a secure and compliant framework to manage your Aiven for Apache Kafka services efficiently. - [Enable governance for Aiven for Apache Kafka® Limited availability](/products/kafka/howto/enable-governance.md): Enable governance in Aiven for Apache Kafka® to create a secure and compliant framework to manage your Aiven for Apache Kafka services efficiently. ###### enable-kafka-tiered-storage Tiered storage significantly improves the storage efficiency of your Aiven for Apache Kafka® service. - [Enable tiered storage for Aiven for Apache Kafka®](/products/kafka/howto/enable-kafka-tiered-storage.md): Tiered storage significantly improves the storage efficiency of your Aiven for Apache Kafka® service. ###### enable-oidc Aiven for Apache Kafka® supports OAuth 2.0/OIDC authentication for Kafka clients. - [Enable OAuth 2.0/OIDC authentication for Apache Kafka®](/products/kafka/howto/enable-oidc.md): Aiven for Apache Kafka® supports OAuth 2.0/OIDC authentication for Kafka clients. ###### enabled-consumer-lag-predictor The consumer lag predictor in Aiven for Apache Kafka® provides visibility into the time between message production and consumption, allowing for improved cluster performance and scalability. - [Enable the consumer lag predictor for Aiven for Apache Kafka® Limited availability](/products/kafka/howto/enabled-consumer-lag-predictor.md): The consumer lag predictor in Aiven for Apache Kafka® provides visibility into the time between message production and consumption, allowing for improved cluster performance and scalability. ###### flink-with-aiven-for-kafka Apache Flink® is an open-source platform for handling distributed streaming and batch data. It enhances Apache Kafka's® event streaming abilities by offering advanced features for consuming, transforming, aggregating, and enriching data. - [Use Apache Flink® with Aiven for Apache Kafka®](/products/kafka/howto/flink-with-aiven-for-kafka.md): Apache Flink® is an open-source platform for handling distributed streaming and batch data. It enhances Apache Kafka's® event streaming abilities by offering advanced features for consuming, transforming, aggregating, and enriching data. ###### generate-avro-java-classes Generate Java classes from Avro schema files (.avsc) to use with Apache Kafka® producers and consumers. Use the avro-tools JAR to create Java classes that match your schema structure. - [Generate Java classes from Avro schemas](/products/kafka/howto/generate-avro-java-classes.md): Generate Java classes from Avro schema files (.avsc) to use with Apache Kafka® producers and consumers. Use the avro-tools JAR to create Java classes that match your schema structure. ###### generate-json-java-classes Generate Java classes from JSON Schema (.json) files for use in Apache Kafka® applications. Use the jsonschema2pojo CLI tool to generate Java classes that match your schema structure. - [Generate Java classes from JSON Schema](/products/kafka/howto/generate-json-java-classes.md): Generate Java classes from JSON Schema (.json) files for use in Apache Kafka® applications. Use the jsonschema2pojo CLI tool to generate Java classes that match your schema structure. ###### generate-protobuf-java-classes Generate Java classes from Protocol Buffers (.proto) schema files for use in Apache Kafka® producers and consumers. Use the protoc compiler to generate Java classes that match your schema structure. - [Generate Java classes from Protobuf schemas](/products/kafka/howto/generate-protobuf-java-classes.md): Generate Java classes from Protocol Buffers (.proto) schema files for use in Apache Kafka® producers and consumers. Use the protoc compiler to generate Java classes that match your schema structure. ###### generate-sample-data Use the sample data generator to simulate streaming events and observe how data flows through topics and schemas in your Aiven for Apache Kafka® service. - [Stream sample data from the Aiven Console](/products/kafka/howto/generate-sample-data.md): Use the sample data generator to simulate streaming events and observe how data flows through topics and schemas in your Aiven for Apache Kafka® service. ###### generate-sample-data-manually Use a Docker-based producer to generate sample data in Aiven for Apache Kafka®. It creates a customizable stream of messages for testing and development. - [Generate sample data with Docker](/products/kafka/howto/generate-sample-data-manually.md): Use a Docker-based producer to generate sample data in Aiven for Apache Kafka®. It creates a customizable stream of messages for testing and development. ###### get-topic-partition-details Learn how to get partition details of an Apache Kafka® topic. - [Get partition details of an Apache Kafka® topic](/products/kafka/howto/get-topic-partition-details.md): Learn how to get partition details of an Apache Kafka® topic. ###### governance Governance in Aiven for Apache Kafka® helps manage your Aiven for Apache Kafka clusters securely and efficiently through structured policies, roles, and processes. - [Governance in Aiven for Apache Kafka® Limited availability](/products/kafka/howto/governance.md): Governance in Aiven for Apache Kafka® helps manage your Aiven for Apache Kafka clusters securely and efficiently through structured policies, roles, and processes. ###### group-requests The Group requests page allows you to view and track requests you and other members of your group made for Aiven for Apache Kafka® resources. - [Manage group requests Limited availability](/products/kafka/howto/group-requests.md): The Group requests page allows you to view and track requests you and other members of your group made for Aiven for Apache Kafka® resources. ###### integrate-external-kafka-cluster You can integrate an external Apache Kafka® cluster with your Aiven for Apache Kafka® service. - [Integrate an external Apache Kafka® cluster in Aiven](/products/kafka/howto/integrate-external-kafka-cluster.md): You can integrate an external Apache Kafka® cluster with your Aiven for Apache Kafka® service. ###### integrate-service-logs-into-kafka-topic You can send logs from your Aiven services into a specified Apache Kafka® topic. - [Integration of logs into Apache Kafka® topic](/products/kafka/howto/integrate-service-logs-into-kafka-topic.md): You can send logs from your Aiven services into a specified Apache Kafka® topic. ###### ipv6-client-connectivity Aiven for Apache Kafka® supports dual-stack IPv4 and IPv6 connectivity. - [Enable IPv6 connectivity for Aiven for Apache Kafka® Early availability](/products/kafka/howto/ipv6-client-connectivity.md): Aiven for Apache Kafka® supports dual-stack IPv4 and IPv6 connectivity. ###### kafbat-ui Kafbat UI is a popular Open-Source web GUI for Apache Kafka® management that allows you to monitor and manage Apache Kafka® clusters. - [Use Kafbat UI with Aiven for Apache Kafka®](/products/kafka/howto/kafbat-ui.md): Kafbat UI is a popular Open-Source web GUI for Apache Kafka® management that allows you to monitor and manage Apache Kafka® clusters. ###### kafdrop Kafdrop is a web UI for Apache Kafka® to monitor clusters, view topics and consumer groups, and integrate with the Schema Registry. - [Use Kafdrop Web UI with Aiven for Apache Kafka®](/products/kafka/howto/kafdrop.md): Kafdrop is a web UI for Apache Kafka® to monitor clusters, view topics and consumer groups, and integrate with the Schema Registry. ###### kafka-conduktor Conduktor is a friendly user interface for Apache Kafka and it works with Aiven and offers built-in support for setting up the connection. - [Connect to Apache Kafka® with Conduktor](/products/kafka/howto/kafka-conduktor.md): Conduktor is a friendly user interface for Apache Kafka and it works with Aiven and offers built-in support for setting up the connection. ###### kafka-custom-serde-encrypt With the Aiven platform, there are several deployment models available to meet your security and compliance needs: - [Encrypt client-side with a custom serializer and deserializer](/products/kafka/howto/kafka-custom-serde-encrypt.md): With the Aiven platform, there are several deployment models available to meet your security and compliance needs: ###### kafka-klaw Klaw is an open-source, web-based data governance toolkit for managing Apache Kafka® topics, ACLs, schemas, and connectors. - [Connect Aiven for Apache Kafka® with Klaw](/products/kafka/howto/kafka-klaw.md): Klaw is an open-source, web-based data governance toolkit for managing Apache Kafka® topics, ACLs, schemas, and connectors. ###### kafka-oauth2-aws-iam Use AWS IAM Outbound Identity Federation to authenticate Apache Kafka® clients with - [Set up Kafka OAuth 2.0/OIDC authentication with AWS IAM using Outbound Identity Federation](/products/kafka/howto/kafka-oauth2-aws-iam.md): Use AWS IAM Outbound Identity Federation to authenticate Apache Kafka® clients with ###### kafka-prometheus-privatelink You can integrate Prometheus with your Aiven for Apache Kafka® service using Privatelink for secure monitoring. - [Configure Prometheus for Aiven for Apache Kafka® using Privatelink](/products/kafka/howto/kafka-prometheus-privatelink.md): You can integrate Prometheus with your Aiven for Apache Kafka® service using Privatelink for secure monitoring. ###### kafka-quix Connect your Aiven for Apache Kafka® service with Quix to consume the data and process it in real-time, and produce it back to Kafka via Quix Cloud. - [Connect Aiven for Apache Kafka® with Quix](/products/kafka/howto/kafka-quix.md): Connect your Aiven for Apache Kafka® service with Quix to consume the data and process it in real-time, and produce it back to Kafka via Quix Cloud. ###### kafka-sasl-auth Aiven for Apache Kafka® supports multiple authentication methods, including Simple Authentication and Security Layer (SASL) over SSL. - [Enable and configure SASL authentication for Apache Kafka®](/products/kafka/howto/kafka-sasl-auth.md): Aiven for Apache Kafka® supports multiple authentication methods, including Simple Authentication and Security Layer (SASL) over SSL. ###### kafka-streams-with-aiven-for-kafka Apache Kafka® Streams is a client-side library for building real-time applications where input and output data are stored in Kafka clusters. - [Use Apache Kafka® Streams with Aiven for Apache Kafka®](/products/kafka/howto/kafka-streams-with-aiven-for-kafka.md): Apache Kafka® Streams is a client-side library for building real-time applications where input and output data are stored in Kafka clusters. ###### kafka-tiered-storage-get-started Aiven for Apache Kafka tiered storage optimizes resources by storing recent, frequently accessed data on faster local disks and moving less active data to more economical, slower storage. - [Get started with tiered storage](/products/kafka/howto/kafka-tiered-storage-get-started.md): Aiven for Apache Kafka tiered storage optimizes resources by storing recent, frequently accessed data on faster local disks and moving less active data to more economical, slower storage. ###### kafka-tools-config-file The open source Apache Kafka® code includes a series of tools under the bin directory that can be useful to manage and interact with an Aiven for Apache Kafka® service. - [Configure properties for Apache Kafka® toolbox](/products/kafka/howto/kafka-tools-config-file.md): The open source Apache Kafka® code includes a series of tools under the bin directory that can be useful to manage and interact with an Aiven for Apache Kafka® service. ###### kcat The kcat tool (formerly known as - [Use kcat with Aiven for Apache Kafka®](/products/kafka/howto/kcat.md): The kcat tool (formerly known as ###### keystore-truststore Aiven for Apache Kafka® utilises TLS (SSL) to secure the traffic between its services and client applications. - [Configure Java SSL keystore and truststore to access Apache Kafka®](/products/kafka/howto/keystore-truststore.md): Aiven for Apache Kafka® utilises TLS (SSL) to secure the traffic between its services and client applications. ###### kpow Kpow by Factor House is an enterprise solution for Kafka management and monitoring. - [Use Kpow with Aiven for Apache Kafka®](/products/kafka/howto/kpow.md): Kpow by Factor House is an enterprise solution for Kafka management and monitoring. ###### ksql-docker Aiven provides a managed Apache Kafka® solution together with a number of auxiliary services like Apache Kafka Connect, Kafka REST and Schema Registry via Karapace. - [Use ksqlDB with Aiven for Apache Kafka®](/products/kafka/howto/ksql-docker.md): Aiven provides a managed Apache Kafka® solution together with a number of auxiliary services like Apache Kafka Connect, Kafka REST and Schema Registry via Karapace. ###### list-code-samples - [Connect to service](/products/kafka/howto/list-code-samples.md) ###### maintenance-updates Manage maintenance updates and set the maintenance window for your Aiven for Apache - [Maintenance and updates for your Aiven for Apache Kafka® service](/products/kafka/howto/maintenance-updates.md): Manage maintenance updates and set the maintenance window for your Aiven for Apache ###### manage-acls Access control lists (ACLs) in Aiven for Apache Kafka® define permissions for topics, schemas, consumer groups, and transactional IDs. - [Manage access control lists in Aiven for Apache Kafka®](/products/kafka/howto/manage-acls.md): Access control lists (ACLs) in Aiven for Apache Kafka® define permissions for topics, schemas, consumer groups, and transactional IDs. ###### manage-quotas Manage quotas in your Aiven for Apache Kafka® service to control network throughput and CPU usage per client. - [Manage quotas in Aiven for Apache Kafka®](/products/kafka/howto/manage-quotas.md): Manage quotas in your Aiven for Apache Kafka® service to control network throughput and CPU usage per client. ###### manage-resource-requests Streamline the creation and ownership requests of Apache Kafka resources to enhance governance and ensure efficient management within Aiven for Apache Kafka. - [Manage Aiven for Apache Kafka® resource requests Limited availability](/products/kafka/howto/manage-resource-requests.md): Streamline the creation and ownership requests of Apache Kafka resources to enhance governance and ensure efficient management within Aiven for Apache Kafka. ###### manage-topics-details Explore advanced topic management features in the Apache Kafka topic catalog. Navigate through various options and configurations for your Apache Kafka topics. - [Manage Apache Kafka® topics in detail Early availability](/products/kafka/howto/manage-topics-details.md): Explore advanced topic management features in the Apache Kafka topic catalog. Navigate through various options and configurations for your Apache Kafka topics. ###### monitor-logs-acl-failure Aiven for Apache Kafka® uses access control lists (ACL) and user - [Monitor and alert logs for denied ACL](/products/kafka/howto/monitor-logs-acl-failure.md): Aiven for Apache Kafka® uses access control lists (ACL) and user ###### optimizing-resource-usage Aiven for Apache Kafka® service plans with CPUs of 2 or less are optimized for lightweight operations, making them suitable for applications that handle fewer messages per second and do not require high throughput. - [Optimizing resource usage for Aiven for Apache Kafka®](/products/kafka/howto/optimizing-resource-usage.md): Aiven for Apache Kafka® service plans with CPUs of 2 or less are optimized for lightweight operations, making them suitable for applications that handle fewer messages per second and do not require high throughput. ###### power-cycle-service Power off your Aiven for Apache Kafka® service to release resources and save credits, power it back on when you need it, or delete it permanently. - [Power on/off and delete your Aiven for Apache Kafka® service](/products/kafka/howto/power-cycle-service.md): Power off your Aiven for Apache Kafka® service to release resources and save credits, power it back on when you need it, or delete it permanently. ###### prevent-full-disks Ensure your Aiven for Apache Kafka® services run smoothly by preventing low disk space. The Aiven platform actively monitors disk usage, triggering notifications when it exceeds 90%. - [Prevent full disks](/products/kafka/howto/prevent-full-disks.md): Ensure your Aiven for Apache Kafka® services run smoothly by preventing low disk space. The Aiven platform actively monitors disk usage, triggering notifications when it exceeds 90%. ###### provectus-kafka-ui Provectus® UI for Apache Kafka® is a popular Open-Source web GUI for Apache Kafka® management that allows you to monitor and manage Apache Kafka® clusters. - [Use Provectus® UI for Apache Kafka® with Aiven for Apache Kafka®](/products/kafka/howto/provectus-kafka-ui.md): Provectus® UI for Apache Kafka® is a popular Open-Source web GUI for Apache Kafka® management that allows you to monitor and manage Apache Kafka® clusters. ###### renew-ssl-certs Aiven for Apache Kafka® automatically generates a new SSL certificate for service users about three months before the existing certificate's expiration date. This new certificate includes a renewed private key. - [Renew and acknowledge service user SSL certificates](/products/kafka/howto/renew-ssl-certs.md): Aiven for Apache Kafka® automatically generates a new SSL certificate for service users about three months before the existing certificate's expiration date. This new certificate includes a renewed private key. ###### request-access-topic Request access to an Apache Kafka topic in Aiven for Apache Kafka Governance to produce or consume messages using access control lists (ACLs). - [Request access to an Apache Kafka topic](/products/kafka/howto/request-access-topic.md): Request access to an Apache Kafka topic in Aiven for Apache Kafka Governance to produce or consume messages using access control lists (ACLs). ###### rotate-credentials Rotate credentials for an Apache Kafka® subscription to replace outdated credentials and maintain secure, approved access. - [Rotate credentials for an Apache Kafka® subscription](/products/kafka/howto/rotate-credentials.md): Rotate credentials for an Apache Kafka® subscription to replace outdated credentials and maintain secure, approved access. ###### scale-disk-storage Use dynamic disk sizing (DDS) to add or remove disk storage on a Classic Kafka service. - [Scale disk storage for your Classic Kafka service](/products/kafka/howto/scale-disk-storage.md): Use dynamic disk sizing (DDS) to add or remove disk storage on a Classic Kafka service. ###### set-kafka-parameters Every Aiven for Apache Kafka® service comes with a set of configuration - [Manage Apache Kafka® parameters](/products/kafka/howto/set-kafka-parameters.md): Every Aiven for Apache Kafka® service comes with a set of configuration ###### set-up-kafka-with-skills Use Skills to automate Kafka workflows to create and configure an Aiven for Apache Kafka® service from the command line. - [Set up Aiven for Apache Kafka® using Skills](/products/kafka/howto/set-up-kafka-with-skills.md): Use Skills to automate Kafka workflows to create and configure an Aiven for Apache Kafka® service from the command line. ###### switch-topic-to-diskless Switch an existing - [Switch a classic topic to a diskless topic Early availability](/products/kafka/howto/switch-topic-to-diskless.md): Switch an existing ###### tag-service Add key-value tags to your Aiven for Apache Kafka® service to organize services and track ownership, cost allocation, and governance. - [Tag your Aiven for Apache Kafka® service](/products/kafka/howto/tag-service.md): Add key-value tags to your Aiven for Apache Kafka® service to organize services and track ownership, cost allocation, and governance. ###### terraform-governance-approvals Aiven for Apache Kafka® Governance lets you manage approval workflows for Apache Kafka topic changes using Terraform and GitHub Actions. - [Manage approvals in Aiven for Apache Kafka® Governance using Terraform & GitHub](/products/kafka/howto/terraform-governance-approvals.md): Aiven for Apache Kafka® Governance lets you manage approval workflows for Apache Kafka topic changes using Terraform and GitHub Actions. ###### use-schema-registry-in-java Configure a Java producer and consumer to use Karapace schema registry with Aiven for Apache Kafka. - [Use schema registry with Java producers and consumers](/products/kafka/howto/use-schema-registry-in-java.md): Configure a Java producer and consumer to use Karapace schema registry with Aiven for Apache Kafka. ###### view-kafka-storage-in-console Use the storage overview page to review storage usage, billing, and retention settings for your Aiven for Apache Kafka® service. - [Storage usage and settings in Aiven Console](/products/kafka/howto/view-kafka-storage-in-console.md): Use the storage overview page to review storage usage, billing, and retention settings for your Aiven for Apache Kafka® service. ###### view-kafka-topic-catalog The Aiven for Apache Kafka® topic catalog offers a user-friendly interface to manage your Apache Kafka topics within your Aiven for Apache Kafka services. - [Manage Apache Kafka® topics with topic catalog Early availability](/products/kafka/howto/view-kafka-topic-catalog.md): The Aiven for Apache Kafka® topic catalog offers a user-friendly interface to manage your Apache Kafka topics within your Aiven for Apache Kafka services. ###### viewing-resetting-offset The open source Apache Kafka® code includes a kafka-consumer-groups.sh utility enabling you to view and manipulate the state of consumer groups. - [View and reset consumer group offsets](/products/kafka/howto/viewing-resetting-offset.md): The open source Apache Kafka® code includes a kafka-consumer-groups.sh utility enabling you to view and manipulate the state of consumer groups. ##### kafka-connect Aiven for Apache Kafka® Connect is a fully managed distributed Apache Kafka® integration component, deployable in the cloud of your choice. - [Aiven for Apache Kafka® Connect](/products/kafka/kafka-connect.md): Aiven for Apache Kafka® Connect is a fully managed distributed Apache Kafka® integration component, deployable in the cloud of your choice. ###### concepts - [Troubleshoot connector list unavailable in Apache Kafka® Connect](/products/kafka/kafka-connect/concepts/connect-plugin-list-not-available.md): When you try to view connectors in Aiven for Apache Kafka® Connect, you might see the message connector list not currently available. - [JDBC source connector modes](/products/kafka/kafka-connect/concepts/jdbc-source-modes.md): JDBC source connector extracts data from a relational database, such as PostgreSQL® or MySQL, and pushes it to Apache Kafka® where can be transformed and read by multiple consumers. - [Available Apache Kafka® Connect connectors](/products/kafka/kafka-connect/concepts/list-of-connector-plugins.md): Discover a variety of connectors available for use with any Aiven for Apache Kafka® service with Apache Kafka® Connect enabled. ###### get-started Get started with Aiven for Apache Kafka® Connect and integrate it with an Aiven for Apache Kafka® service. - [Get started with Aiven for Apache Kafka® Connect](/products/kafka/kafka-connect/get-started.md): Get started with Aiven for Apache Kafka® Connect and integrate it with an Aiven for Apache Kafka® service. ###### howto - [Create an AMQP source connector for Aiven for Apache Kafka® Early availability](/products/kafka/kafka-connect/howto/amqp-source-connector.md): The AMQP source connector retrieves messages from an AMQP-compatible queue and writes them to Apache Kafka® topics. - [Configure the Iceberg sink connector with AWS Glue catalog](/products/kafka/kafka-connect/howto/aws-glue-catalog.md): The AWS Glue catalog directly manages Iceberg metadata within AWS Glue. It supports automatic table creation and schema evolution. - [Configure the Iceberg sink connector with AWS Glue REST catalog](/products/kafka/kafka-connect/howto/aws-glue-rest-catalog.md): The AWS Glue REST catalog stores metadata using the Iceberg REST API. It integrates Apache Kafka with AWS Glue using REST-based communication. - [Create an Azure Blob Storage sink connector for Aiven for Apache Kafka®](/products/kafka/kafka-connect/howto/azure-blob-sink.md): The Azure Blob Storage sink connector moves data from Apache Kafka® topics to Azure Blob Storage containers for long-term storage, such as archiving or creating backups. - [Create an Azure Blob Storage source connector for Aiven for Apache Kafka®](/products/kafka/kafka-connect/howto/azure-blob-source.md): Use the Azure Blob source connector to stream data from Blob Storage into Apache Kafka® for real-time processing, analytics, or recovery. - [Get the best from Apache Kafka® Connect](/products/kafka/kafka-connect/howto/best-practices.md): We recommend to follow these best practices to ensure that your Apache - [Bring your own Apache Kafka® Connect cluster](/products/kafka/kafka-connect/howto/bring-your-own-kafka-connect-cluster.md): Aiven provides Apache Kafka® Connect as a managed service in combination with the Aiven for Apache Kafka® managed service. However, there are circumstances where you may want to roll your own Kafka Connect cluster. - [Create a Stream Reactor sink connector from Apache Kafka® to Apache Cassandra®](/products/kafka/kafka-connect/howto/cassandra-streamreactor-sink.md): The Apache Cassandra® Stream Reactor sink connector enables you to move data from an Aiven for Apache Kafka® cluster to a Apache Cassandra® database. - [Create a Stream Reactor source connector from Apache Cassandra® to Apache Kafka®](/products/kafka/kafka-connect/howto/cassandra-streamreactor-source.md): The Apache Cassandra® Stream Reactor source connector enables you to move data from an Apache Cassandra® database to an Aiven for Apache Kafka® cluster. - [Create a ClickHouse sink connector for Aiven for Apache Kafka®](/products/kafka/kafka-connect/howto/clickhouse-sink-connector.md): The ClickHouse sink connector delivers data from Apache Kafka® topics to a ClickHouse database for efficient querying and analysis. - [Configure AWS Secrets Manager](/products/kafka/kafka-connect/howto/configure-aws-secrets-manager.md): Configure and use AWS Secrets Manager as a secret provider in Aiven for Apache Kafka® Connect services. - [Configure Azure Key Vault](/products/kafka/kafka-connect/howto/configure-azure-key-vault.md): Configure and use Azure Key Vault as a secret provider in Aiven for Apache Kafka® Connect services. - [Configure the ENV secret provider](/products/kafka/kafka-connect/howto/configure-env-secret-provider.md): Configure and use the ENV secret provider in Aiven for Apache Kafka® Connect services. - [Configure HashiCorp Vault](/products/kafka/kafka-connect/howto/configure-hashicorp-vault.md): Configure and use HashiCorp Vault as a secret provider in Aiven for Apache Kafka® Connect services. - [Aiven for Apache Kafka® Connect secret providers](/products/kafka/kafka-connect/howto/configure-secret-providers.md): Configure and use secret providers in Apache Kafka Connect services on Aiven for Apache Kafka. - [Create a sink connector from Apache Kafka® to Couchbase](/products/kafka/kafka-connect/howto/couchbase-sink.md): The Couchbase sink connector pushes Apache Kafka® data to the NoSQL database. - [Create a source connector from Couchbase to Apache Kafka®](/products/kafka/kafka-connect/howto/couchbase-source.md): The Couchbase source connector pushes data - [Create a Debezium source connector from MongoDB to Apache Kafka®](/products/kafka/kafka-connect/howto/debezium-source-connector-mongodb.md): Track and write MongoDB database changes to an Apache Kafka® topic in a standard format with the Debezium source connector, enabling transformation and access by multiple consumers using a MongoDB replica set or sharded cluster. - [Create a Debezium source connector from MySQL to Apache Kafka®](/products/kafka/kafka-connect/howto/debezium-source-connector-mysql.md): The MySQL Debezium source connector extracts the changes committed to the database binary log (binlog), and writes them to an Apache Kafka® topic in a standard format where they can be transformed and read by multiple consumers. - [Create an Oracle Debezium source connector for Aiven for Apache Kafka® Early availability](/products/kafka/kafka-connect/howto/debezium-source-connector-oracle.md): The Oracle Debezium source connector streams change data from an Oracle database to Apache Kafka® topics. - [Create a Debezium source connector from PostgreSQL® to Apache Kafka®](/products/kafka/kafka-connect/howto/debezium-source-connector-pg.md): The Debezium source connector extracts the changes committed to the transaction log in a relational database, such as PostgreSQL®, and writes them to an Apache Kafka® topic in a standard format where they can be transformed and read by multiple consumers. - [Handle PostgreSQL® node replacements when using Debezium for change data capture](/products/kafka/kafka-connect/howto/debezium-source-connector-pg-node-replacement.md): When you run a Debezium source connector for PostgreSQL® with an Aiven for PostgreSQL® service, some database operations can interrupt change data capture (CDC). - [Create a Debezium source connector from SQL Server to Apache Kafka® with CDC](/products/kafka/kafka-connect/howto/debezium-source-connector-sql-server.md): The SQL Server Debezium source connector uses the change data capture (CDC) feature to extract database changes from designated tables and write them to Apache Kafka® topic in a standard format for multiple consumers to read and transform. - [Create a sink connector from Apache Kafka® to Elasticsearch](/products/kafka/kafka-connect/howto/elasticsearch-sink.md): The Elasticsearch sink connector enables you to move data from an Aiven for Apache Kafka® cluster to an Elasticsearch instance for further processing and analysis. - [Enable automatic restart for Apache Kafka® Connect connectors](/products/kafka/kafka-connect/howto/enable-automatic-restart.md): Automatic restart can help recover a connector task after a rare transient failure, such as an out-of-memory error caused by a sudden data surge. - [Enable Apache Kafka® Connect on Aiven for Apache Kafka®](/products/kafka/kafka-connect/howto/enable-connect.md): For a low-cost way to get started with Aiven for Apache Kafka® Connect, you can run Kafka Connect on the same nodes as your Apache Kafka cluster, sharing the resources. The Kafka service must be running on a business or premium plan. - [Create a sink connector from Apache Kafka® to Google BigQuery](/products/kafka/kafka-connect/howto/gcp-bigquery-sink.md): Set up the BigQuery sink connector to move data from Aiven for Apache Kafka® into BigQuery tables for analysis and storage. - [Create a sink connector from Apache Kafka® to Google Pub/Sub Lite](/products/kafka/kafka-connect/howto/gcp-pubsub-lite-sink.md): The [Google Pub/Sub Lite sink - [Create a Google Pub/Sub Lite source connector to Apache Kafka®](/products/kafka/kafka-connect/howto/gcp-pubsub-lite-source.md): The [Google Pub/Sub Lite source - [Create a sink connector from Apache Kafka® to Google Pub/Sub](/products/kafka/kafka-connect/howto/gcp-pubsub-sink.md): The [Google Pub/Sub sink - [Create a Google Pub/Sub source connector to Apache Kafka®](/products/kafka/kafka-connect/howto/gcp-pubsub-source.md): The Google Pub/Sub source connector enables you to push from a Google Pub/Sub subscription to an Aiven for Apache Kafka® topic. - [Create a Google Cloud Storage sink connector for Apache Kafka®](/products/kafka/kafka-connect/howto/gcs-sink.md): The Google Cloud Storage (GCS) sink connector moves data from Aiven for Apache Kafka® topics to a Google Cloud Storage bucket for long-term storage. - [Create a sink connector from Apache Kafka® via HTTP](/products/kafka/kafka-connect/howto/http-sink.md): The HTTP sink connector enables you to move data from an Aiven for Apache Kafka® cluster to a remote server via HTTP. - [Create an IBM MQ sink connector in Aiven for Apache Kafka®](/products/kafka/kafka-connect/howto/ibm-mq-sink-connector.md): The IBM MQ sink connector allows you to route messages from Apache Kafka® topics to IBM MQ queues. - [Create an Iceberg sink connector for Aiven for Apache Kafka®](/products/kafka/kafka-connect/howto/iceberg-sink-connector.md): Use the Iceberg sink connector to write real-time Apache Kafka® data to Iceberg tables for analytics and long-term storage. - [Create a sink connector from Apache Kafka® to InfluxDB®](/products/kafka/kafka-connect/howto/influx-sink.md): The InfluxDB® Stream Reactor sink connector enables you to move data from an Aiven for Apache Kafka® cluster to an InfluxDB® instance. - [Configure the Iceberg sink connector with a PostgreSQL JDBC catalog](/products/kafka/kafka-connect/howto/jdbc-catalog-postgres.md): The JDBC catalog stores Iceberg table metadata in a PostgreSQL® database. Use it to - [Create a JDBC sink connector from Apache Kafka® to another database](/products/kafka/kafka-connect/howto/jdbc-sink.md): The JDBC (Java Database Connectivity) sink connector enables you to move data from an Aiven for Apache Kafka® cluster to any relational database offering JDBC drivers like PostgreSQL® or MySQL. - [Create a JDBC source connector from MySQL to Apache Kafka®](/products/kafka/kafka-connect/howto/jdbc-source-connector-mysql.md): The JDBC source connector pushes data from a relational database, such as MySQL, to Apache Kafka® where can be transformed and read by multiple consumers. - [Create a JDBC source connector from PostgreSQL® to Apache Kafka®](/products/kafka/kafka-connect/howto/jdbc-source-connector-pg.md): The JDBC source connector pushes data from a relational database, such as PostgreSQL®, to Apache Kafka® where can be transformed and read by multiple consumers. - [Create a JDBC source connector from SQL Server to Apache Kafka®](/products/kafka/kafka-connect/howto/jdbc-source-connector-sql-server.md): The JDBC source connector pushes data from a relational database, such - [Integrate Aiven for Apache Kafka® Connect with PostgreSQL using Debezium with mutual TLS](/products/kafka/kafka-connect/howto/kafka-connect-debezium-tls-pg.md): Integrate Aiven for Apache Kafka® Connect with PostgreSQL using Debezium with mutual TLS (mTLS) to enhance security with mutual authentication. - [Manage connector versions in Aiven for Apache Kafka Connect®](/products/kafka/kafka-connect/howto/manage-connector-versions.md): Aiven for Apache Kafka Connect® lets you control which connector version is used in your service. - [Manage Apache Kafka® Connect logging level](/products/kafka/kafka-connect/howto/manage-logging-level.md): During the operation of an Aiven for Apache Kafka® Connect cluster, you may encounter errors from one or more running connectors. Sometimes the stack trace printed in the logs is useful in determining the root cause of an issue, while other times the information provided just isn't enough to work with. - [Create a MongoDB source connector for Aiven for Apache Kafka®](/products/kafka/kafka-connect/howto/mongodb-poll-source-connector.md): Use the MongoDB source connector to stream data from MongoDB collections into Apache Kafka® topics for processing and analytics. - [Create a Lenses.io MongoDB sink connector for Aiven for Apache Kafka®](/products/kafka/kafka-connect/howto/mongodb-sink-lenses.md): Use the MongoDB sink connector by Lenses.io to write data from Apache Kafka® topics into a MongoDB database. - [Create a MongoDB sink connector for Aiven for Apache Kafka®](/products/kafka/kafka-connect/howto/mongodb-sink-mongo.md): Use the MongoDB sink connector to move data from an Aiven for Apache Kafka® service to a MongoDB database. - [Create an MQTT sink connector](/products/kafka/kafka-connect/howto/mqtt-sink-connector.md): The MQTT sink connector copies messages from an Apache Kafka® topic to an MQTT queue. - [Create a source connector from MQTT to Apache Kafka®](/products/kafka/kafka-connect/howto/mqtt-source-connector.md): The Stream Reactor MQTT source connector transfers messages from an MQTT topic to an Aiven for Apache Kafka® topic, where they can be processed and consumed by multiple applications. - [Create a sink connector from Apache Kafka® to OpenSearch®](/products/kafka/kafka-connect/howto/opensearch-sink.md): The OpenSearch sink connector writes data from Aiven for Apache Kafka® to OpenSearch®. - [Create a Stream Reactor sink connector from Apache Kafka® to Redis®*](/products/kafka/kafka-connect/howto/redis-streamreactor-sink.md): The Redis®\ Stream Reactor sink connector enables you to move data from an Aiven for Apache Kafka® cluster to a Redis®\ database. - [Request a new connector](/products/kafka/kafka-connect/howto/request-new-connector.md): If you know about new and interesting Apache Kafka® connectors you'd - [Use AWS IAM assume role credentials provider](/products/kafka/kafka-connect/howto/s3-iam-assume-role.md): The Aiven for Apache Kafka® S3 sink connector moves data from an Aiven for Apache Kafka cluster to Amazon S3 for long-term storage. - [Amazon S3 sink connector](/products/kafka/kafka-connect/howto/s3-sink.md): Use the Amazon S3 sink connectors to move data from Apache Kafka® topics to Amazon S3 for storage and analytics. - [S3 sink connector by Aiven naming and data formats](/products/kafka/kafka-connect/howto/s3-sink-additional-parameters.md): The Apache Kafka Connect® S3 sink connector by Aiven moves data from an Aiven for Apache Kafka® cluster to Amazon S3. You can configure object naming and output data formats. - [S3 sink connector by Confluent naming and data formats](/products/kafka/kafka-connect/howto/s3-sink-additional-parameters-confluent.md): The Apache Kafka Connect® S3 sink connector moves data from Aiven for Apache Kafka® to Amazon S3 for long-term storage. - [Create an Amazon S3 sink connector for Aiven from Apache Kafka®](/products/kafka/kafka-connect/howto/s3-sink-connector-aiven.md): The Amazon S3 sink connector sends data from Aiven for Apache Kafka® to Amazon S3 for long-term storage. - [Create an Amazon S3 sink connector by Confluent from Apache Kafka®](/products/kafka/kafka-connect/howto/s3-sink-connector-confluent.md): The Amazon S3 sink connector by Confluent moves data from Apache Kafka® topics to Amazon S3 buckets for long-term storage. - [Prepare AWS for Amazon S3 sink](/products/kafka/kafka-connect/howto/s3-sink-prepare.md): Set up AWS to allow the S3 sink connector to write data from Apache Kafka® to Amazon S3. - [Create an Amazon S3 source connector for Aiven for Apache Kafka®](/products/kafka/kafka-connect/howto/s3-source-connector.md): The Amazon S3 source connector allows you to ingest data from S3 buckets into Apache Kafka® topics for real-time processing and analytics. - [Create a Salesforce sink connector for Aiven for Apache Kafka® Early availability](/products/kafka/kafka-connect/howto/salesforce-sink-connector.md): The Salesforce sink connector writes records from Apache Kafka® topics to Salesforce objects, such as Account or Contact. - [Create a Salesforce source connector for Aiven for Apache Kafka® Early availability](/products/kafka/kafka-connect/howto/salesforce-source-connector.md): The Salesforce source connector retrieves data from Salesforce and writes it to Apache Kafka® topics. - [Configure the Iceberg sink connector with Snowflake Open Catalog](/products/kafka/kafka-connect/howto/snowflake-open-catalog.md): Snowflake Open Catalog is a managed Apache Polaris™ service that supports the Apache Iceberg™ REST catalog API and uses Amazon S3 for storage. - [Create and configure a Snowflake sink connector for Apache Kafka®](/products/kafka/kafka-connect/howto/snowflake-sink.md): The Apache Kafka Connect® Snowflake sink connector moves data from Aiven for Apache Kafka® topics to a Snowflake database. It requires configuration in both Snowflake and Aiven for Apache Kafka. - [Create a sink connector from Apache Kafka® to Splunk](/products/kafka/kafka-connect/howto/splunk-sink.md): The Splunk sink connector enables you to move ###### reference - [Advanced parameters for Apache Kafka® Connect](/products/kafka/kafka-connect/reference/advanced-params.md): See the configuration options available for - [Aiven for Apache Kafka® Connect metrics available via Prometheus](/products/kafka/kafka-connect/reference/connect-metrics-prometheus.md): Discover metrics offered by Prometheus for the Aiven for Apache Kafka® Connect service. ##### kafka-mirrormaker Aiven for Apache Kafka® MirrorMaker 2 is a fully managed distributed Apache Kafka® data replication utility, deployable in the cloud of your choice. Apache Kafka® MirrorMaker 2 lets you sync your topic data across Apache Kafka® cluster deployed anywhere in the world. - [Aiven for Apache Kafka® MirrorMaker 2](/products/kafka/kafka-mirrormaker.md): Aiven for Apache Kafka® MirrorMaker 2 is a fully managed distributed Apache Kafka® data replication utility, deployable in the cloud of your choice. Apache Kafka® MirrorMaker 2 lets you sync your topic data across Apache Kafka® cluster deployed anywhere in the world. ###### concepts - [Configuration and tuning for Aiven for Apache Kafka® MirrorMaker 2](/products/kafka/kafka-mirrormaker/concepts/configuration-layers.md): Learn where Aiven for Apache Kafka® MirrorMaker 2 settings are configured across service, replication-flow, and integration layers, which parameters affect performance, and what restarts when you change them. - [Disaster recovery and migration](/products/kafka/kafka-mirrormaker/concepts/disaster-recovery-migration.md): MirrorMaker 2 is the standard replication tool packaged with the Apache Kafka® and can be run as a managed service on the Aiven platform. - [MirrorMaker 2 active-active setup](/products/kafka/kafka-mirrormaker/concepts/disaster-recovery/active-active-setup.md): An active-active setup of MM2 allows data to be replicated between two clusters simultaneously. - [Active-passive setup](/products/kafka/kafka-mirrormaker/concepts/disaster-recovery/active-passive-setup.md): In this setup, there are two Apache Kafka® clusters, the primary and secondary clusters. The primary cluster contains the topic topic. - [Configure permissions for MirrorMaker 2 with external Kafka clusters](/products/kafka/kafka-mirrormaker/concepts/permissions-internal-topics.md): By default, Apache Kafka® MirrorMaker 2 creates internal topics to store metadata, - [Include or exclude topics in a replication flow](/products/kafka/kafka-mirrormaker/concepts/replication-flow-topics-regex.md): When you define a replication flow, specify which topics in the source Apache Kafka® cluster to include or exclude from the cross-cluster replica. ###### get-started Create an Apache Kafka® MirrorMaker 2 service and connect it to your Aiven for Apache Kafka® service. - [Get started with Apache Kafka® MirrorMaker 2](/products/kafka/kafka-mirrormaker/get-started.md): Create an Apache Kafka® MirrorMaker 2 service and connect it to your Aiven for Apache Kafka® service. ###### howto - [Configure Apache Kafka® MirrorMaker 2 metrics sent to Datadog](/products/kafka/kafka-mirrormaker/howto/datadog-customised-metrics.md): When creating a Datadog service integration, customize which metrics are sent to the Datadog endpoint using the Aiven CLI. - [Exactly-once delivery in Aiven for Apache Kafka MirrorMaker 2](/products/kafka/kafka-mirrormaker/howto/exactly-once-delivery.md): Exactly-once delivery in Aiven for Apache Kafka MirrorMaker 2 replicates each message exactly once between clusters, preventing duplicates or data loss. - [Offset sync status analysis for Apache Kafka® MirrorMaker 2](/products/kafka/kafka-mirrormaker/howto/log-analysis-offset-sync-tool.md): Analyze offset synchronization status between source and target clusters with Aiven’s offset sync inspection tool for Aiven for Apache Kafka MirrorMaker 2. - [Configure rack awareness in Aiven for Apache Kafka® MirrorMaker 2](/products/kafka/kafka-mirrormaker/howto/mm2-rack-awareness.md): Configure rack awareness in Aiven for Apache Kafka® MirrorMaker 2 to reduce cross-availability zone (AZ) network traffic by directing MirrorMaker to read from local follower replicas instead of remote partition leaders. - [Monitor replication execution](/products/kafka/kafka-mirrormaker/howto/monitor-replication-execution.md): Apache Kafka® MirrorMaker 2 uses Kafka® Connect for monitoring and state management, helping you track replication flows and address issues. - [Remove topic prefix when replicating with Apache Kafka® MirrorMaker 2](/products/kafka/kafka-mirrormaker/howto/remove-mirrormaker-prefix.md): When you use Apache Kafka® MirrorMaker 2 to replicate topics across Apache - [Set up an Apache Kafka® MirrorMaker 2 replication flow](/products/kafka/kafka-mirrormaker/howto/setup-replication-flow.md): Apache Kafka® MirrorMaker 2 replication flows sync topics from a source - [Update integration configurations for Aiven for Apache Kafka® MirrorMaker 2](/products/kafka/kafka-mirrormaker/howto/update-integration-configurations.md): Update the producer and consumer settings on a MirrorMaker 2 service integration to control how MirrorMaker 2 communicates with the source and target Kafka clusters. ###### reference - [Advanced parameters for Aiven for Apache Kafka® MirrorMaker 2](/products/kafka/kafka-mirrormaker/reference/advanced-params.md): See the configuration options available for - [Known issues](/products/kafka/kafka-mirrormaker/reference/known-issues.md): MirrorMaker 2 may translate offsets incorrectly if LAG shows negative on the target cluster - [Terminology for Aiven for Apache Kafka® MirrorMaker 2](/products/kafka/kafka-mirrormaker/reference/terminology.md): - Cluster alias: The name alias defined in MirrorMaker 2 for a ###### troubleshooting - [Why topics or partitions not replicated](/products/kafka/kafka-mirrormaker/troubleshooting/topic-not-replicated.md): Apache Kafka® MirrorMaker 2 provides reliable message replication across Kafka clusters using configurations, states, and offsets. ##### karapace Karapace is an Aiven-built open-source Schema Registry and REST Proxy for Aiven for Apache Kafka®. - [Karapace](/products/kafka/karapace.md): Karapace is an Aiven-built open-source Schema Registry and REST Proxy for Aiven for Apache Kafka®. ###### concepts - [Schema registry ACL definitions](/products/kafka/karapace/concepts/acl-definition.md): Learn the username, operation, and resource fields used in Karapace Schema Registry ACLs. - [Schema references in Karapace](/products/kafka/karapace/concepts/schema-references.md): Schema references let you register a schema that depends on other schemas already stored in the Schema Registry. - [Karapace schema registry authorization](/products/kafka/karapace/concepts/schema-registry-authorization.md): The schema registry authorization feature when enabled in ###### howto - [Enable Apache Kafka® REST proxy authorization](/products/kafka/karapace/howto/enable-kafka-rest-proxy-authorization.md): Enable REST proxy authorization so Karapace enforces Apache Kafka ACLs on REST API requests. - [Enable schema registry and REST proxy](/products/kafka/karapace/howto/enable-karapace.md): You can enable the Karapace schema registry and REST proxy independently on Aiven for Apache Kafka®. - [Enable OAuth 2.0/OIDC support for Apache Kafka® REST proxy](/products/kafka/karapace/howto/enable-oauth-oidc-kafka-rest-proxy.md): Secure your Apache Kafka® resources by integrating OAuth 2.0/OpenID Connect (OIDC) with the Karapace REST proxy and enabling REST proxy authorization. - [Enable OAuth 2.0/OIDC authentication for Aiven for Apache Kafka® Schema Registry](/products/kafka/karapace/howto/enable-oauth-oidc-schema-registry.md): Authenticate Karapace Schema Registry requests with OAuth 2.0/OIDC bearer tokens and optionally enforce role-based authorization. - [Shut down Karapace on invalid schema records](/products/kafka/karapace/howto/enable-schema-reader-strict-mode.md): Configure Karapace Schema Registry to shut down when it detects invalid records in the _schemas topic. - [Enable Karapace schema registry authorization](/products/kafka/karapace/howto/enable-schema-registry-authorization.md): Most Aiven for Apache Kafka® services will automatically have schema registry authorization enabled, and the functionality cannot be disabled or enabled once a service has been created. - [Manage Karapace schema registry authorization](/products/kafka/karapace/howto/manage-schema-registry-authorization.md): Karapace schema registry authorization allows you to authenticate the user, to control access to individual Karapace schema registry REST API endpoints, and to filter the content the endpoints return. - [Register schemas with references in Karapace](/products/kafka/karapace/howto/register-schemas-with-references.md): Use the Karapace Schema Registry API and curl to register Avro and Protobuf schemas that use schema references. - [Set the Karapace version](/products/kafka/karapace/howto/set-karapace-version.md): You can set which Karapace version your Aiven for Apache Kafka® service runs. ##### reference ###### advanced-params View the configuration options for Aiven for Apache Kafka® Classic. - [Advanced parameters for Aiven for Apache Kafka® Classic](/products/kafka/reference/advanced-params.md): View the configuration options for Aiven for Apache Kafka® Classic. ###### advanced-params-dev-tier View the configuration options for Aiven for Apache Kafka® Developer tier. - [Advanced parameters for Aiven for Apache Kafka® Developer tier](/products/kafka/reference/advanced-params-dev-tier.md): View the configuration options for Aiven for Apache Kafka® Developer tier. ###### advanced-params-free-tier View the configuration options for Aiven for Apache Kafka® Free tier. - [Advanced parameters for Aiven for Apache Kafka® Free tier](/products/kafka/reference/advanced-params-free-tier.md): View the configuration options for Aiven for Apache Kafka® Free tier. ###### advanced-params-index Advanced parameters let you customize your Aiven for Apache Kafka® service - [Advanced parameters for Aiven for Apache Kafka®](/products/kafka/reference/advanced-params-index.md): Advanced parameters let you customize your Aiven for Apache Kafka® service ###### advanced-params-standard View the configuration options for Standard Kafka. These configurations apply to - [Advanced parameters for Standard Kafka](/products/kafka/reference/advanced-params-standard.md): View the configuration options for Standard Kafka. These configurations apply to ###### advanced-params-standard-topic View the topic-level configuration options for Standard Kafka. These configurations - [Advanced topic parameters for Standard Kafka](/products/kafka/reference/advanced-params-standard-topic.md): View the topic-level configuration options for Standard Kafka. These configurations ###### kafka-metrics-prometheus Explore common metrics available via Prometheus for your Aiven for Apache Kafka® service. - [Aiven for Apache Kafka® metrics available via Prometheus](/products/kafka/reference/kafka-metrics-prometheus.md): Explore common metrics available via Prometheus for your Aiven for Apache Kafka® service. ###### version-lifecycle Learn how Aiven manages Aiven for Apache Kafka® version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. - [Aiven for Apache Kafka® version lifecycle](/products/kafka/reference/version-lifecycle.md): Learn how Aiven manages Aiven for Apache Kafka® version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. ##### standard-kafka-overview Standard Kafka is an Aiven for Apache Kafka® service type that stores topic data in object storage through diskless topics. - [Standard Kafka overview](/products/kafka/standard-kafka-overview.md): Standard Kafka is an Aiven for Apache Kafka® service type that stores topic data in object storage through diskless topics. ##### terminology A comprehensive glossary of essential Apache Kafka® terms and their meaning. - [Apache Kafka® terminology](/products/kafka/terminology.md): A comprehensive glossary of essential Apache Kafka® terms and their meaning. ##### troubleshooting ###### non-leader-for-partition Aiven continuously monitors services to ensure they are healthy. If problems arise, nodes can be recycled: new nodes are created to substitute old, malfunctioning ones. - [NOT_LEADER_FOR_PARTITION errors](/products/kafka/troubleshooting/non-leader-for-partition.md): Aiven continuously monitors services to ensure they are healthy. If problems arise, nodes can be recycled: new nodes are created to substitute old, malfunctioning ones. ###### troubleshoot-consumer-disconnections Apache Kafka® consumers sometimes experience disconnections from a cluster node. - [Troubleshoot Apache Kafka® consumer connections](/products/kafka/troubleshooting/troubleshoot-consumer-disconnections.md): Apache Kafka® consumers sometimes experience disconnections from a cluster node. #### metrics Aiven for Metrics, powered by Thanos, simplifies the management and analysis of large volumes of metrics data. The service is scalable, reliable, and efficient, suitable for organizations of all sizes. - [Aiven for Metrics](/products/metrics.md): Aiven for Metrics, powered by Thanos, simplifies the management and analysis of large volumes of metrics data. The service is scalable, reliable, and efficient, suitable for organizations of all sizes. ##### concepts ###### retention-rules Retention rules in Aiven for Metrics define how long your metrics data is stored. - [Retention rules in Aiven for Metrics](/products/metrics/concepts/retention-rules.md): Retention rules in Aiven for Metrics define how long your metrics data is stored. ###### service-memory Understand the memory limits and out-of-memory conditions that apply to your Aiven for Metrics service. - [Memory and out-of-memory conditions in Aiven for Metrics](/products/metrics/concepts/service-memory.md): Understand the memory limits and out-of-memory conditions that apply to your Aiven for Metrics service. ###### storage-resource-scaling Aiven for Metrics optimizes storage and compute resources using tiered storage, disk storage, memory, and compute power to balance cost and performance. - [Optimize storage and resources](/products/metrics/concepts/storage-resource-scaling.md): Aiven for Metrics optimizes storage and compute resources using tiered storage, disk storage, memory, and compute power to balance cost and performance. ##### get-started Get started with Aiven for Metrics by creating your service using the Aiven Console or Aiven CLI. - [Get started with Aiven for Metrics](/products/metrics/get-started.md): Get started with Aiven for Metrics by creating your service using the Aiven Console or Aiven CLI. ##### howto ###### change-cloud-region Move your Aiven for Metrics service to a different cloud provider or region. - [Change the cloud or region for your Aiven for Metrics service](/products/metrics/howto/change-cloud-region.md): Move your Aiven for Metrics service to a different cloud provider or region. ###### change-service-plan Change the service plan for your Aiven for Metrics service to scale resources up or down and optimize costs. - [Change the plan for your Aiven for Metrics service](/products/metrics/howto/change-service-plan.md): Change the service plan for your Aiven for Metrics service to scale resources up or down and optimize costs. ###### controlled-upgrade-pipelines Link Aiven for Metrics services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. - [Controlled upgrade pipelines for your Aiven for Metrics service Limited availability](/products/metrics/howto/controlled-upgrade-pipelines.md): Link Aiven for Metrics services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. ###### maintenance-updates Manage maintenance updates and set the maintenance window for your Aiven for Metrics service. - [Maintenance and updates for your Aiven for Metrics service](/products/metrics/howto/maintenance-updates.md): Manage maintenance updates and set the maintenance window for your Aiven for Metrics service. ###### power-cycle-service Power off your Aiven for Metrics service to release resources and save credits, power it - [Power on/off and delete your Aiven for Metrics service](/products/metrics/howto/power-cycle-service.md): Power off your Aiven for Metrics service to release resources and save credits, power it ###### prepare-for-high-load Prepare your Aiven for Metrics service for higher than usual traffic to avoid outages and keep performance stable. - [Prepare your Aiven for Metrics service for high load](/products/metrics/howto/prepare-for-high-load.md): Prepare your Aiven for Metrics service for higher than usual traffic to avoid outages and keep performance stable. ###### storage-usage Get a comprehensive view of your storage usage with Aiven for Metrics. The Storage page in the Aiven console provides valuable data to help you effectively manage and understand your object storage consumption and associated costs. - [Manage storage](/products/metrics/howto/storage-usage.md): Get a comprehensive view of your storage usage with Aiven for Metrics. The Storage page in the Aiven console provides valuable data to help you effectively manage and understand your object storage consumption and associated costs. ###### tag-service Add key-value tags to your Aiven for Metrics service to organize services and track - [Tag your Aiven for Metrics service](/products/metrics/howto/tag-service.md): Add key-value tags to your Aiven for Metrics service to organize services and track ###### track-restore-progress Track the restore progress of individual nodes in your Aiven for Metrics service during node replacement, forking, or maintenance, using the Aiven API. - [Track restore progress for your Aiven for Metrics service](/products/metrics/howto/track-restore-progress.md): Track the restore progress of individual nodes in your Aiven for Metrics service during node replacement, forking, or maintenance, using the Aiven API. ##### maintenance-lifecycle Manage maintenance updates, maintenance windows, and node restore progress for your - [Maintenance and lifecycle in Aiven for Metrics](/products/metrics/maintenance-lifecycle.md): Manage maintenance updates, maintenance windows, and node restore progress for your ##### scaling-performance Scale resources and prepare your Aiven for Metrics service for changing load. - [Scaling and performance in Aiven for Metrics](/products/metrics/scaling-performance.md): Scale resources and prepare your Aiven for Metrics service for changing load. #### mysql Aiven for MySQL® is a fully managed relational database service, deployable in the cloud of your choice. - [Aiven for MySQL®](/products/mysql.md): Aiven for MySQL® is a fully managed relational database service, deployable in the cloud of your choice. ##### backups-migration Back up, restore, and migrate your Aiven for MySQL® service data. - [Backups and migration in Aiven for MySQL®](/products/mysql/backups-migration.md): Back up, restore, and migrate your Aiven for MySQL® service data. ##### concepts ###### high-availability Aiven for MySQL® is available on a variety of plans, offering different levels of service availability. The selected plan defines the features available. - [High availability of Aiven for MySQL®](/products/mysql/concepts/high-availability.md): Aiven for MySQL® is available on a variety of plans, offering different levels of service availability. The selected plan defines the features available. ###### max-number-of-connections Calculate the total number of simultaneous connections available to all users combined on your Aiven for MySQL® service, and learn why the per-user connection limit is a separate setting. - [MySQL max_connections](/products/mysql/concepts/max-number-of-connections.md): Calculate the total number of simultaneous connections available to all users combined on your Aiven for MySQL® service, and learn why the per-user connection limit is a separate setting. ###### mysql-backups Aiven for MySQL databases are automatically backed-up, with full backups daily, and binary logs recorded continuously. - [Understand MySQL backups in Aiven](/products/mysql/concepts/mysql-backups.md): Aiven for MySQL databases are automatically backed-up, with full backups daily, and binary logs recorded continuously. ###### mysql-free-tier Use Aiven for MySQL® for free. You don't need a credit card to sign up - [Aiven for MySQL® free tier](/products/mysql/concepts/mysql-free-tier.md): Use Aiven for MySQL® for free. You don't need a credit card to sign up ###### mysql-memory-usage MySQL memory utilization can appear high, even if the service is relatively idle. - [Understand MySQL memory usage](/products/mysql/concepts/mysql-memory-usage.md): MySQL memory utilization can appear high, even if the service is relatively idle. ###### mysql-replication Replication in Aiven for MySQL® is always based on replicating logical changes. This means that the replication protocol may contain an actual statement that the target server should apply or it may have an entry saying update row with these old attributes to have these new attributes. - [Understand MySQL replication in Aiven](/products/mysql/concepts/mysql-replication.md): Replication in Aiven for MySQL® is always based on replicating logical changes. This means that the replication protocol may contain an actual statement that the target server should apply or it may have an entry saying update row with these old attributes to have these new attributes. ###### mysql-tuning-and-concurrency Determining how much memory is available for queries, and tuning concurrency accordingly, requires calculation of service memory, query analysis, and monitoring. - [MySQL tuning for concurrency](/products/mysql/concepts/mysql-tuning-and-concurrency.md): Determining how much memory is available for queries, and tuning concurrency accordingly, requires calculation of service memory, query analysis, and monitoring. ##### get-started Start using Aiven for MySQL® by creating a service, connecting to it, and loading sample data. - [Get started with Aiven for MySQL®](/products/mysql/get-started.md): Start using Aiven for MySQL® by creating a service, connecting to it, and loading sample data. ##### howto ###### ai-insights Use Aiven's artificial intelligence capabilities to identify slow queries and get optimization suggestions. - [AI database optimizer for Aiven for MySQL®](/products/mysql/howto/ai-insights.md): Use Aiven's artificial intelligence capabilities to identify slow queries and get optimization suggestions. ###### backup-to-another-region Copy your Aiven for MySQL® service backups to a secondary region for disaster recovery. - [Back up your Aiven for MySQL® service to another region](/products/mysql/howto/backup-to-another-region.md): Copy your Aiven for MySQL® service backups to a secondary region for disaster recovery. ###### change-cloud-region Move your Aiven for MySQL® service to a different cloud provider or region. - [Change the cloud or region for your Aiven for MySQL® service](/products/mysql/howto/change-cloud-region.md): Move your Aiven for MySQL® service to a different cloud provider or region. ###### change-service-plan Change the service plan for your Aiven for MySQL® service to scale resources up or down and optimize costs. - [Change the plan for your Aiven for MySQL® service](/products/mysql/howto/change-service-plan.md): Change the service plan for your Aiven for MySQL® service to scale resources up or down and optimize costs. ###### connect-from-cli Connect to your Aiven for MySQL® service via the command line with the following tools: - [Connect to Aiven for MySQL® from the command line](/products/mysql/howto/connect-from-cli.md): Connect to your Aiven for MySQL® service via the command line with the following tools: ###### connect-from-mysql-workbench You can use a graphical client like MySQL Workbench to connect to Aiven for MySQL® services. - [Connect to Aiven for MySQL® with MySQL Workbench](/products/mysql/howto/connect-from-mysql-workbench.md): You can use a graphical client like MySQL Workbench to connect to Aiven for MySQL® services. ###### connect-using-mysqlx-with-python Enabling the MySQLx protocol support allows you to use your MySQL - [Connect to Aiven for MySQL® using MySQLx with Python](/products/mysql/howto/connect-using-mysqlx-with-python.md): Enabling the MySQLx protocol support allows you to use your MySQL ###### connect-with-datagrip Use DataGrip to connect to your Aiven for MySQL® service. - [Connect to Aiven for MySQL® with DataGrip](/products/mysql/howto/connect-with-datagrip.md): Use DataGrip to connect to your Aiven for MySQL® service. ###### connect-with-dbeaver Use DBeaver to connect to your Aiven for MySQL® service. - [Connect to Aiven for MySQL® with DBeaver](/products/mysql/howto/connect-with-dbeaver.md): Use DBeaver to connect to your Aiven for MySQL® service. ###### connect-with-java This example connects your Java application to an Aiven for MySQL® service. - [Connect to Aiven for MySQL® with Java](/products/mysql/howto/connect-with-java.md): This example connects your Java application to an Aiven for MySQL® service. ###### connect-with-php This example connects to an Aiven for MySQL® service from PHP, making use of the built-in PDO module. - [Connect to Aiven for MySQL® with PHP](/products/mysql/howto/connect-with-php.md): This example connects to an Aiven for MySQL® service from PHP, making use of the built-in PDO module. ###### connect-with-python This example connects your Python application to an Aiven for MySQL® - [Connect to Aiven for MySQL® with Python](/products/mysql/howto/connect-with-python.md): This example connects your Python application to an Aiven for MySQL® ###### create-database Once you've created your Aiven for MySQL® service, you can add additional databases, for security purposes or to isolate your data per application. - [Create Aiven for MySQL® databases](/products/mysql/howto/create-database.md): Once you've created your Aiven for MySQL® service, you can add additional databases, for security purposes or to isolate your data per application. ###### create-missing-primary-keys Learn strategies to create missing primary keys in your Aiven for MySQL® service. They are important for MySQL replication process. - [Create missing primary keys](/products/mysql/howto/create-missing-primary-keys.md): Learn strategies to create missing primary keys in your Aiven for MySQL® service. They are important for MySQL replication process. ###### create-remote-replica Learn how to create an Aiven for MySQL® read replica to provide a read-only instance of your managed MySQL service in another geographically autonomous region. - [Create Aiven for MySQL® read replicas](/products/mysql/howto/create-remote-replica.md): Learn how to create an Aiven for MySQL® read replica to provide a read-only instance of your managed MySQL service in another geographically autonomous region. ###### create-tables-without-primary-keys If your Aiven for MySQL® service was created after 2020-06-03, by default it does not allow creating new tables without primary keys. - [Create new tables without primary keys](/products/mysql/howto/create-tables-without-primary-keys.md): If your Aiven for MySQL® service was created after 2020-06-03, by default it does not allow creating new tables without primary keys. ###### disable-foreign-key-checks All Aiven for MySQL® services have foreign key checks enabled by default helping in keeping referential integrity across tables. However, you might want to disable it for a particular session. - [Disable foreign key checks](/products/mysql/howto/disable-foreign-key-checks.md): All Aiven for MySQL® services have foreign key checks enabled by default helping in keeping referential integrity across tables. However, you might want to disable it for a particular session. ###### disk-autoscaler Automatically increase the disk storage of your Aiven for MySQL® service when it's running out of space, instead of resizing it manually. - [Scale disk storage automatically for your Aiven for MySQL® service](/products/mysql/howto/disk-autoscaler.md): Automatically increase the disk storage of your Aiven for MySQL® service when it's running out of space, instead of resizing it manually. ###### do-check-service-migration Learn how to find potential errors before starting your database migration process. This can be done by using either the Aiven CLI or the Aiven REST API. - [Perform pre-migration checks](/products/mysql/howto/do-check-service-migration.md): Learn how to find potential errors before starting your database migration process. This can be done by using either the Aiven CLI or the Aiven REST API. ###### enable-slow-queries You can identify inefficient or time-consuming queries by enabling slow query log in your Aiven for MySQL® service. - [Enable slow query logging](/products/mysql/howto/enable-slow-queries.md): You can identify inefficient or time-consuming queries by enabling slow query log in your Aiven for MySQL® service. ###### fork-service Fork your Aiven for MySQL® service to create an independent copy for testing, - [Fork your Aiven for MySQL® service](/products/mysql/howto/fork-service.md): Fork your Aiven for MySQL® service to create an independent copy for testing, ###### identify-disk-usage-issues Aiven for MySQL® is configured to use innodbfileper_table=ON, which - [Identify disk usage issues](/products/mysql/howto/identify-disk-usage-issues.md): Aiven for MySQL® is configured to use innodbfileper_table=ON, which ###### list-code-samples Connect to the Aiven for MySQL® service using various programming languages or tools. - [Connect to Aiven for MySQL®](/products/mysql/howto/list-code-samples.md): Connect to the Aiven for MySQL® service using various programming languages or tools. ###### maintenance-updates Manage maintenance updates and set the maintenance window for your Aiven for MySQL® service. - [Maintenance and updates for your Aiven for MySQL® service](/products/mysql/howto/maintenance-updates.md): Manage maintenance updates and set the maintenance window for your Aiven for MySQL® service. ###### manage-mysql-version Aiven for MySQL® supports multiple versions of MySQL running concurrently in the platform. Choose a version that best fits your needs and upgrade your service when ready. - [Manage Aiven for MySQL® versions](/products/mysql/howto/manage-mysql-version.md): Aiven for MySQL® supports multiple versions of MySQL running concurrently in the platform. Choose a version that best fits your needs and upgrade your service when ready. ###### manage-service-users Create and manage service users in your Aiven for MySQL® service to control access to - [Manage Aiven for MySQL® service users](/products/mysql/howto/manage-service-users.md): Create and manage service users in your Aiven for MySQL® service to control access to ###### migrate-database-mysqldump Copy your Aiven for MySQL® data to a file, back it up to another Aiven for MySQL database, and restore it using mysqldump/restore or mydumper/myloader. - [Backup and restore Aiven for MySQL® with mysqldump or mydumper](/products/mysql/howto/migrate-database-mysqldump.md): Copy your Aiven for MySQL® data to a file, back it up to another Aiven for MySQL database, and restore it using mysqldump/restore or mydumper/myloader. ###### migrate-db-to-aiven-via-console Use the Aiven Console to migrate MySQL® databases to managed MySQL clusters in your Aiven organization. - [Migrate to Aiven for MySQL® via console](/products/mysql/howto/migrate-db-to-aiven-via-console.md): Use the Aiven Console to migrate MySQL® databases to managed MySQL clusters in your Aiven organization. ###### migrate-from-external-mysql Migrate your external MySQL database to an Aiven-hosted one using either a one-time dump-and-restore or continuous data synchronization through MySQL's built-in replication. - [Migrate to Aiven for MySQL® via CLI](/products/mysql/howto/migrate-from-external-mysql.md): Migrate your external MySQL database to an Aiven-hosted one using either a one-time dump-and-restore or continuous data synchronization through MySQL's built-in replication. ###### mysql-long-running-queries Aiven does not terminate any customer queries even if they run - [Detect and terminate long-running queries in Aiven for MySQL®](/products/mysql/howto/mysql-long-running-queries.md): Aiven does not terminate any customer queries even if they run ###### power-cycle-service Power off your Aiven for MySQL® service to release resources and save credits, power it - [Power on/off and delete your Aiven for MySQL® service](/products/mysql/howto/power-cycle-service.md): Power off your Aiven for MySQL® service to release resources and save credits, power it ###### prepare-for-high-load Prepare your Aiven for MySQL® service for higher than usual traffic to avoid outages and keep performance stable. - [Prepare your Aiven for MySQL® service for high load](/products/mysql/howto/prepare-for-high-load.md): Prepare your Aiven for MySQL® service for higher than usual traffic to avoid outages and keep performance stable. ###### prevent-disk-full Learn how Aiven prevents running out of disk space from happening and how you can make more space available on your disk when needed. Running out of disk space makes the service start malfunctioning and prevents backups from being properly created. - [Prevent running out of disk space](/products/mysql/howto/prevent-disk-full.md): Learn how Aiven prevents running out of disk space from happening and how you can make more space available on your disk when needed. Running out of disk space makes the service start malfunctioning and prevents backups from being properly created. ###### reclaim-disk-space You can configure InnoDB to release disk space back to the operating system by running the OPTIMIZE TABLE command. - [Reclaim disk space](/products/mysql/howto/reclaim-disk-space.md): You can configure InnoDB to release disk space back to the operating system by running the OPTIMIZE TABLE command. ###### rename-service Change the name of your Aiven for MySQL® service by forking it under a new name and - [Rename your Aiven for MySQL® service](/products/mysql/howto/rename-service.md): Change the name of your Aiven for MySQL® service by forking it under a new name and ###### scale-disk-storage Scale the disk storage of your Aiven for MySQL® service up or down without disrupting the running service. - [Scale disk storage for your Aiven for MySQL® service](/products/mysql/howto/scale-disk-storage.md): Scale the disk storage of your Aiven for MySQL® service up or down without disrupting the running service. ###### tag-service Add key-value tags to your Aiven for MySQL® service to organize services and track - [Tag your Aiven for MySQL® service](/products/mysql/howto/tag-service.md): Add key-value tags to your Aiven for MySQL® service to organize services and track ###### track-restore-progress Track the restore progress of individual nodes in your Aiven for MySQL® service during node replacement, forking, or maintenance, using the Aiven API. - [Track restore progress for your Aiven for MySQL® service](/products/mysql/howto/track-restore-progress.md): Track the restore progress of individual nodes in your Aiven for MySQL® service during node replacement, forking, or maintenance, using the Aiven API. ###### use-incremental-backups Streamline your Aiven for MySQL® backups using incremental backups. - [Use Aiven for MySQL® incremental backups Early availability](/products/mysql/howto/use-incremental-backups.md): Streamline your Aiven for MySQL® backups using incremental backups. ##### maintenance-lifecycle Keep your Aiven for MySQL® service current by upgrading versions and applying maintenance - [Maintenance and lifecycle in Aiven for MySQL®](/products/mysql/maintenance-lifecycle.md): Keep your Aiven for MySQL® service current by upgrading versions and applying maintenance ##### reference ###### advanced-params See the configuration options available for Aiven for MySQL: - [Advanced parameters for Aiven for MySQL](/products/mysql/reference/advanced-params.md): See the configuration options available for Aiven for MySQL: ###### resource-capability - [Resource capability of Aiven for MySQL® plans](/products/mysql/reference/resource-capability.md) ###### version-lifecycle Learn how Aiven manages Aiven for MySQL® version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. - [Aiven for MySQL® version lifecycle](/products/mysql/reference/version-lifecycle.md): Learn how Aiven manages Aiven for MySQL® version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. ##### scaling-performance Tune performance and manage disk usage to keep your Aiven for MySQL® service running - [Scaling and performance in Aiven for MySQL®](/products/mysql/scaling-performance.md): Tune performance and manage disk usage to keep your Aiven for MySQL® service running #### opensearch Aiven for OpenSearch® is a fully managed distributed search and analytics suite, deployable in the cloud of your choice. - [Aiven for OpenSearch®](/products/opensearch.md): Aiven for OpenSearch® is a fully managed distributed search and analytics suite, deployable in the cloud of your choice. ##### concepts ###### access_control Access control is a crucial security measure that allows you to control who can access your data and resources. By setting up access control rules, you can restrict access to sensitive data and prevent unauthorized changes or deletions. - [Access control in Aiven for OpenSearch®](/products/opensearch/concepts/access_control.md): Access control is a crucial security measure that allows you to control who can access your data and resources. By setting up access control rules, you can restrict access to sensitive data and prevent unauthorized changes or deletions. ###### aggregations Alongside the search functionality, OpenSearch® offers a powerful - [Analyze data with Aiven for OpenSearch® aggregations](/products/opensearch/concepts/aggregations.md): Alongside the search functionality, OpenSearch® offers a powerful ###### cross-cluster-replication-opensearch Cross-cluster replication (CCR) in Aiven for OpenSearch lets you replicate indices, including their data, mappings, and metadata, from one service to another across different regions and cloud providers. - [OpenSearch® cross-cluster replication Limited availability](/products/opensearch/concepts/cross-cluster-replication-opensearch.md): Cross-cluster replication (CCR) in Aiven for OpenSearch lets you replicate indices, including their data, mappings, and metadata, from one service to another across different regions and cloud providers. ###### dedicated-node-roles Aiven for OpenSearch® supports dedicated node roles, enabling workload isolation across specialized node groups for optimized performance and scaling. - [Dedicated node roles in Aiven for OpenSearch®](/products/opensearch/concepts/dedicated-node-roles.md): Aiven for OpenSearch® supports dedicated node roles, enabling workload isolation across specialized node groups for optimized performance and scaling. ###### high-availability-for-opensearch Aiven for OpenSearch® is available on a variety of plans, offering - [High availability in Aiven for OpenSearch®](/products/opensearch/concepts/high-availability-for-opensearch.md): Aiven for OpenSearch® is available on a variety of plans, offering ###### hot-warm-tiering Hot/warm data tiering lets you store recent, frequently queried data on fast nodes and older data on cheaper nodes—without splitting your cluster or changing how you search. - [Hot/warm data tiering in Aiven for OpenSearch® Limited availability](/products/opensearch/concepts/hot-warm-tiering.md): Hot/warm data tiering lets you store recent, frequently queried data on fast nodes and older data on cheaper nodes—without splitting your cluster or changing how you search. ###### index-replication The replication factor in Aiven for OpenSearch® determines the number of copies (replicas) of each index shard. Replicas protect against data loss, and OpenSearch can also route search queries to replicas, so more replicas can spread search load across more nodes. - [Replication factors in Aiven for OpenSearch®](/products/opensearch/concepts/index-replication.md): The replication factor in Aiven for OpenSearch® determines the number of copies (replicas) of each index shard. Replicas protect against data loss, and OpenSearch can also route search queries to replicas, so more replicas can spread search load across more nodes. ###### indices Learn how documents, indices, shards, and replicas relate to each other in Aiven for OpenSearch®, and how mapping, aliases, and index lifecycle management keep your indices efficient. - [Manage indices in Aiven for OpenSearch®](/products/opensearch/concepts/indices.md): Learn how documents, indices, shards, and replicas relate to each other in Aiven for OpenSearch®, and how mapping, aliases, and index lifecycle management keep your indices efficient. ###### opensearch-free-tier Get started with Aiven for OpenSearch® at no cost. The Aiven for OpenSearch free tier is a fully managed service for learning, prototyping, and evaluation. No credit card is required. - [Aiven for OpenSearch® free tier](/products/opensearch/concepts/opensearch-free-tier.md): Get started with Aiven for OpenSearch® at no cost. The Aiven for OpenSearch free tier is a fully managed service for learning, prototyping, and evaluation. No credit card is required. ###### opensearch-security-considerations Before enabling OpenSearch Security management for Aiven for OpenSearch service, understand its impact on your current system and adapt your infrastructure accordingly. - [Key considerations and system adaptation for OpenSearch® Security management](/products/opensearch/concepts/opensearch-security-considerations.md): Before enabling OpenSearch Security management for Aiven for OpenSearch service, understand its impact on your current system and adapt your infrastructure accordingly. ###### opensearch-vs-elasticsearch OpenSearch® is the open-source version of the Elasticsearch project, which has a restrictive license. Third parties cannot offer Elasticsearch as a service. - [OpenSearch® vs Elasticsearch](/products/opensearch/concepts/opensearch-vs-elasticsearch.md): OpenSearch® is the open-source version of the Elasticsearch project, which has a restrictive license. Third parties cannot offer Elasticsearch as a service. ###### os-security OpenSearch Security is a powerful feature that enhances the security of - [OpenSearch Security for Aiven for OpenSearch®](/products/opensearch/concepts/os-security.md): OpenSearch Security is a powerful feature that enhances the security of ###### service-memory Understand the memory limits and out-of-memory conditions that apply to your Aiven for OpenSearch® service. - [Memory and out-of-memory conditions in Aiven for OpenSearch®](/products/opensearch/concepts/service-memory.md): Understand the memory limits and out-of-memory conditions that apply to your Aiven for OpenSearch® service. ###### shards-number A key component of using OpenSearch® is determining the optimal number of shards for your index. - [Optimal number of shards](/products/opensearch/concepts/shards-number.md): A key component of using OpenSearch® is determining the optimal number of shards for your index. ###### when-create-index Determine wether to create an index per customer/project/entity or to look for alternatives. - [When to create an index](/products/opensearch/concepts/when-create-index.md): Determine wether to create an index per customer/project/entity or to look for alternatives. ##### dashboards OpenSearch® Dashboards is both a visualisation tool for data in the cluster and a user interface for OpenSearch plugins. - [OpenSearch® Dashboards](/products/opensearch/dashboards.md): OpenSearch® Dashboards is both a visualisation tool for data in the cluster and a user interface for OpenSearch plugins. ###### get-started To start using Aiven for OpenSearch® Dashboards, create Aiven for OpenSearch® service first and OpenSearch Dashboards service will be added alongside it. - [Get started with Aiven for OpenSearch® Dashboards](/products/opensearch/dashboards/get-started.md): To start using Aiven for OpenSearch® Dashboards, create Aiven for OpenSearch® service first and OpenSearch Dashboards service will be added alongside it. ###### howto - [Run queries from OpenSearch Dashboards Dev Tools](/products/opensearch/dashboards/howto/dev-tools-usage-example.md): Similarly to how you can work with the OpenSearch® service using cURL you can run the queries directly from OpenSearch Dashboards Dev Tools. The console contains both a request editor and a command output window. - [Create alerts with OpenSearch® Dashboards](/products/opensearch/dashboards/howto/opensearch-alerting-dashboard.md): Set up alerts in OpenSearch® Dashboards to send notifications when your data meets specific conditions. ##### get-started Learn how to use Aiven for OpenSearch®, create a service, secure access, manage indices, and explore your data. - [Get started with Aiven for OpenSearch®](/products/opensearch/get-started.md): Learn how to use Aiven for OpenSearch®, create a service, secure access, manage indices, and explore your data. ##### howto ###### audit-logs Aiven for OpenSearch® enables audit logging functionality via the OpenSearch Security dashboard, which allows OpenSearch Security administrators to track system events, security-related events, and user activity. - [Enable and manage OpenSearch® Audit logs](/products/opensearch/howto/audit-logs.md): Aiven for OpenSearch® enables audit logging functionality via the OpenSearch Security dashboard, which allows OpenSearch Security administrators to track system events, security-related events, and user activity. ###### backup-to-another-region Copy your Aiven for OpenSearch® service backups to a secondary region for disaster recovery. - [Back up your Aiven for OpenSearch® service to another region](/products/opensearch/howto/backup-to-another-region.md): Copy your Aiven for OpenSearch® service backups to a secondary region for disaster recovery. ###### change-cloud-region Move your Aiven for OpenSearch® service to a different cloud provider or region. - [Change the cloud or region for your Aiven for OpenSearch® service](/products/opensearch/howto/change-cloud-region.md): Move your Aiven for OpenSearch® service to a different cloud provider or region. ###### change-service-plan Change the service plan for your Aiven for OpenSearch® service to scale resources up or down and optimize costs. - [Change the plan for your Aiven for OpenSearch® service](/products/opensearch/howto/change-service-plan.md): Change the service plan for your Aiven for OpenSearch® service to scale resources up or down and optimize costs. ###### connect-with-nodejs The most convenient way to work with the cluster when using NodeJS is to rely on OpenSearch® JavaScript client. - [Connect to Aiven for OpenSearch® with NodeJS](/products/opensearch/howto/connect-with-nodejs.md): The most convenient way to work with the cluster when using NodeJS is to rely on OpenSearch® JavaScript client. ###### connect-with-python You can interact with your cluster with the help of the Python OpenSearch® client. - [Connect to Aiven for OpenSearch® with Python](/products/opensearch/howto/connect-with-python.md): You can interact with your cluster with the help of the Python OpenSearch® client. ###### control_access_to_content Manage users and permissions in Aiven for OpenSearch by creating Access Control Lists (ACLs) in the Aiven Console. - [Manage users and access control in Aiven for OpenSearch®](/products/opensearch/howto/control_access_to_content.md): Manage users and permissions in Aiven for OpenSearch by creating Access Control Lists (ACLs) in the Aiven Console. ###### controlled-upgrade-pipelines Link Aiven for OpenSearch® services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. - [Controlled upgrade pipelines for your Aiven for OpenSearch® service Limited availability](/products/opensearch/howto/controlled-upgrade-pipelines.md): Link Aiven for OpenSearch® services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. ###### create-free-tier-opensearch You can create a free tier Aiven for OpenSearch® service to learn OpenSearch, test indexing and queries, or run small proof-of-concept workloads. - [Create a free tier Aiven for OpenSearch® service](/products/opensearch/howto/create-free-tier-opensearch.md): You can create a free tier Aiven for OpenSearch® service to learn OpenSearch, test indexing and queries, or run small proof-of-concept workloads. ###### custom-dictionary-files Custom dictionary files are user-defined files that enhance query analysis and improve search relevance in OpenSearch. By adding domain-specific vocabulary and rules, these files refine search results to be more accurate and relevant. - [Custom dictionary files](/products/opensearch/howto/custom-dictionary-files.md): Custom dictionary files are user-defined files that enhance query analysis and improve search relevance in OpenSearch. By adding domain-specific vocabulary and rules, these files refine search results to be more accurate and relevant. ###### custom-repositories Use the Aiven Console or API for configuring custom repositories in Aiven for OpenSearch to store snapshots in your cloud storage. - [Manage Aiven for OpenSearch® custom repositories in the Aiven Console or API](/products/opensearch/howto/custom-repositories.md): Use the Aiven Console or API for configuring custom repositories in Aiven for OpenSearch to store snapshots in your cloud storage. ###### datadog-metrics Send Aiven for OpenSearch® metrics to Datadog, and choose which metric categories the integration collects. - [Aiven for OpenSearch® metrics sent to Datadog](/products/opensearch/howto/datadog-metrics.md): Send Aiven for OpenSearch® metrics to Datadog, and choose which metric categories the integration collects. ###### disk-autoscaler Automatically increase the disk storage of your Aiven for OpenSearch® service when it's running out of space, instead of resizing it manually. - [Scale disk storage automatically for your Aiven for OpenSearch® service](/products/opensearch/howto/disk-autoscaler.md): Automatically increase the disk storage of your Aiven for OpenSearch® service when it's running out of space, instead of resizing it manually. ###### enable-opensearch-security OpenSearch Security provides a range of security features, including fine-grained access controls, SAML authentication, and audit logging to monitor activity within your Aiven for OpenSearch® service. - [Enable OpenSearch Security management for Aiven for OpenSearch®](/products/opensearch/howto/enable-opensearch-security.md): OpenSearch Security provides a range of security features, including fine-grained access controls, SAML authentication, and audit logging to monitor activity within your Aiven for OpenSearch® service. ###### enable-slow-query-log Identify inefficient or time-consuming queries by enabling slow query logging in your Aiven for OpenSearch® service. - [Enable slow query logging](/products/opensearch/howto/enable-slow-query-log.md): Identify inefficient or time-consuming queries by enabling slow query logging in your Aiven for OpenSearch® service. ###### fork-service Fork your Aiven for OpenSearch® service to create an independent copy for testing, - [Fork your Aiven for OpenSearch® service](/products/opensearch/howto/fork-service.md): Fork your Aiven for OpenSearch® service to create an independent copy for testing, ###### handle-low-disk-space Free up disk space in Aiven for OpenSearch® and recover from the flood stage watermark when a node runs low on disk. - [Handle low disk space in Aiven for OpenSearch®](/products/opensearch/howto/handle-low-disk-space.md): Free up disk space in Aiven for OpenSearch® and recover from the flood stage watermark when a node runs low on disk. ###### hot-warm-tiering Set up Index State Management (ISM) policies to automate index lifecycle management across hot and warm data nodes in Aiven for OpenSearch®. - [Manage hot/warm data tiering in Aiven for OpenSearch® Limited availability](/products/opensearch/howto/hot-warm-tiering.md): Set up Index State Management (ISM) policies to automate index lifecycle management across hot and warm data nodes in Aiven for OpenSearch®. ###### import-opensearch-data-elasticsearch-dump-to-aiven Backup your OpenSearch® data into Aiven for Opensearch. - [Copy data from OpenSearch to Aiven for OpenSearch® using elasticsearch-dump](/products/opensearch/howto/import-opensearch-data-elasticsearch-dump-to-aiven.md): Backup your OpenSearch® data into Aiven for Opensearch. ###### import-opensearch-data-elasticsearch-dump-to-aws Backup your OpenSearch® data into an AWS S3 bucket. - [Copy data from Aiven for OpenSearch® to AWS S3 using elasticsearch-dump](/products/opensearch/howto/import-opensearch-data-elasticsearch-dump-to-aws.md): Backup your OpenSearch® data into an AWS S3 bucket. ###### integrate-with-grafana You can monitor and set up alerts for the data in your Aiven for OpenSearch® service with Grafana®. - [Integrate with Grafana®](/products/opensearch/howto/integrate-with-grafana.md): You can monitor and set up alerts for the data in your Aiven for OpenSearch® service with Grafana®. ###### jwt-authentication Configure JSON Web Token (JWT) authentication to enable secure, stateless authentication for Aiven for OpenSearch®. - [Enable JSON Web Token authentication on Aiven for OpenSearch®](/products/opensearch/howto/jwt-authentication.md): Configure JSON Web Token (JWT) authentication to enable secure, stateless authentication for Aiven for OpenSearch®. ###### list-connect-to-service Connect to the Aiven for OpenSearch® service using various programming languages or tools. - [Connect to Aiven for OpenSearch®](/products/opensearch/howto/list-connect-to-service.md): Connect to the Aiven for OpenSearch® service using various programming languages or tools. ###### list-opensearch-security Using OpenSearch Security can significantly strengthen the security of - [OpenSearch® Security management in Aiven for OpenSearch®](/products/opensearch/howto/list-opensearch-security.md): Using OpenSearch Security can significantly strengthen the security of ###### list-search-service Learn how to write and execute search queries and aggregate data using OpenSearch clients in two widely used programming languages: Python and NodeJS. - [Search and aggregations with Aiven for OpenSearch®](/products/opensearch/howto/list-search-service.md): Learn how to write and execute search queries and aggregate data using OpenSearch clients in two widely used programming languages: Python and NodeJS. ###### maintenance-updates Manage maintenance updates and set the maintenance window for your Aiven for OpenSearch® service. - [Maintenance and updates for your Aiven for OpenSearch® service](/products/opensearch/howto/maintenance-updates.md): Manage maintenance updates and set the maintenance window for your Aiven for OpenSearch® service. ###### manage-custom-repo - [Manage Aiven for OpenSearch® custom repositories in OpenSearch® API](/products/opensearch/howto/manage-custom-repo/custom-repositories-os-api.md): Use the OpenSearch® API for configuring custom repositories in Aiven for OpenSearch to store snapshots in your cloud storage. - [Manage Aiven for OpenSearch® custom repositories](/products/opensearch/howto/manage-custom-repo/list-manage-custom-repo.md): Set up custom repositories in Aiven for OpenSearch® in the Aiven Console, the Aiven API, ###### manage-snapshots Create, list, retrieve, or delete snapshots in your Aiven for OpenSearch custom repositories. - [Create and manage snapshots in Aiven for OpenSearch®](/products/opensearch/howto/manage-snapshots.md): Create, list, retrieve, or delete snapshots in your Aiven for OpenSearch custom repositories. ###### migrate-external-snapshots-aiven-opensearch Migrate an existing OpenSearch or Elasticsearch® snapshot to Aiven for OpenSearch® with minimal downtime and data integrity. - [Migrate external OpenSearch or Elasticsearch snapshots to Aiven](/products/opensearch/howto/migrate-external-snapshots-aiven-opensearch.md): Migrate an existing OpenSearch or Elasticsearch® snapshot to Aiven for OpenSearch® with minimal downtime and data integrity. ###### migrate-ism-policies Reapply Index State Management (ISM) policies to Aiven for OpenSearch® using a script. - [Reapply ISM policies after snapshot restore](/products/opensearch/howto/migrate-ism-policies.md): Reapply Index State Management (ISM) policies to Aiven for OpenSearch® using a script. ###### migrate-knn-nmslib-engine Identify indices that use the deprecated nmslib k-NN engine, and reindex them to - [Migrate Aiven for OpenSearch® k-NN indices off the nmslib engine](/products/opensearch/howto/migrate-knn-nmslib-engine.md): Identify indices that use the deprecated nmslib k-NN engine, and reindex them to ###### migrate-opendistro-security-config-aiven Migrate your security configuration from an OpenDistro service to Aiven for OpenSearch® using a migration script. - [Migrate OpenDistro security configuration to Aiven for OpenSearch](/products/opensearch/howto/migrate-opendistro-security-config-aiven.md): Migrate your security configuration from an OpenDistro service to Aiven for OpenSearch® using a migration script. ###### migrating_elasticsearch_data_to_aiven To migrate Elasticsearch data to Aiven for OpenSearch®, reindex from a remote Elasticsearch cluster. - [Migrate Elasticsearch data to Aiven for OpenSearch®](/products/opensearch/howto/migrating_elasticsearch_data_to_aiven.md): To migrate Elasticsearch data to Aiven for OpenSearch®, reindex from a remote Elasticsearch cluster. ###### oidc-authentication OpenID Connect (OIDC) is an authentication protocol that builds on top of the OAuth 2.0 protocol. - [Enable OpenID Connect authentication on Aiven for OpenSearch®](/products/opensearch/howto/oidc-authentication.md): OpenID Connect (OIDC) is an authentication protocol that builds on top of the OAuth 2.0 protocol. ###### opensearch-aggregations-and-nodejs Learn how to aggregate data using OpenSearch and its NodeJS client. - [Use Aggregations with OpenSearch® and NodeJS](/products/opensearch/howto/opensearch-aggregations-and-nodejs.md): Learn how to aggregate data using OpenSearch and its NodeJS client. ###### opensearch-alerting-api OpenSearch® alerting feature sends notifications when data from one or more indices meets certain conditions that can be customized. - [Create alerts with OpenSearch® API](/products/opensearch/howto/opensearch-alerting-api.md): OpenSearch® alerting feature sends notifications when data from one or more indices meets certain conditions that can be customized. ###### opensearch-and-nodejs Learn how the OpenSearch® JavaScript client gives a clear and useful interface to communicate with an OpenSearch cluster and run search queries. - [Write search queries with OpenSearch® and NodeJS](/products/opensearch/howto/opensearch-and-nodejs.md): Learn how the OpenSearch® JavaScript client gives a clear and useful interface to communicate with an OpenSearch cluster and run search queries. ###### opensearch-dashboard-multi_tenancy Aiven for OpenSearch® provides support for multi-tenancy through OpenSearch Security Dashboard. - [Set up OpenSearch® Dashboard multi-tenancy](/products/opensearch/howto/opensearch-dashboard-multi_tenancy.md): Aiven for OpenSearch® provides support for multi-tenancy through OpenSearch Security Dashboard. ###### opensearch-log-integration Aiven provides a log integration that sends logs from any Aiven service, such as Aiven for Apache Kafka®, Aiven for PostgreSQL®, and Aiven for Grafana®, to Aiven for OpenSearch®, so you can search, analyze, and monitor all your logs in one place. - [Manage OpenSearch® log integration](/products/opensearch/howto/opensearch-log-integration.md): Aiven provides a log integration that sends logs from any Aiven service, such as Aiven for Apache Kafka®, Aiven for PostgreSQL®, and Aiven for Grafana®, to Aiven for OpenSearch®, so you can search, analyze, and monitor all your logs in one place. ###### opensearch-search-and-python Learn how to write and run search queries on your OpenSearch cluster using a Python OpenSearch client. - [Write search queries with OpenSearch® and Python](/products/opensearch/howto/opensearch-search-and-python.md): Learn how to write and run search queries on your OpenSearch cluster using a Python OpenSearch client. ###### opensearch-with-curl Connect to your Aiven for OpenSearch® service with cURL. - [Connect to Aiven for OpenSearch® with cURL](/products/opensearch/howto/opensearch-with-curl.md): Connect to your Aiven for OpenSearch® service with cURL. ###### os-metrics Monitor and optimize your Aiven for OpenSearch service with metrics available via Prometheus. - [Aiven for OpenSearch® metrics available via Prometheus](/products/opensearch/howto/os-metrics.md): Monitor and optimize your Aiven for OpenSearch service with metrics available via Prometheus. ###### os-version-upgrade Aiven for OpenSearch® allows you to choose the version that best fits your needs and upgrade when ready. - [Upgrade Aiven for OpenSearch®](/products/opensearch/howto/os-version-upgrade.md): Aiven for OpenSearch® allows you to choose the version that best fits your needs and upgrade when ready. ###### power-cycle-service Power off your Aiven for OpenSearch® service to release resources and save credits, power - [Power on/off and delete your Aiven for OpenSearch® service](/products/opensearch/howto/power-cycle-service.md): Power off your Aiven for OpenSearch® service to release resources and save credits, power ###### prepare-for-high-load Prepare your Aiven for OpenSearch® service for higher than usual traffic to avoid outages and keep performance stable. - [Prepare your Aiven for OpenSearch® service for high load](/products/opensearch/howto/prepare-for-high-load.md): Prepare your Aiven for OpenSearch® service for higher than usual traffic to avoid outages and keep performance stable. ###### reindex-opensearch When upgrading Aiven for OpenSearch® to a newer version, reindex indices created with an earlier version to ensure compatibility with the target version. - [Reindex Aiven for OpenSearch® data on a newer version](/products/opensearch/howto/reindex-opensearch.md): When upgrading Aiven for OpenSearch® to a newer version, reindex indices created with an earlier version to ensure compatibility with the target version. ###### rename-service Change the name of your Aiven for OpenSearch® service by forking it under a new name and - [Rename your Aiven for OpenSearch® service](/products/opensearch/howto/rename-service.md): Change the name of your Aiven for OpenSearch® service by forking it under a new name and ###### resolve-shards-too-large Resolve the large shard size alert in Aiven for OpenSearch® by deleting old data, splitting an index, or reindexing it with more shards. - [Manage large shards in Aiven for OpenSearch®](/products/opensearch/howto/resolve-shards-too-large.md): Resolve the large shard size alert in Aiven for OpenSearch® by deleting old data, splitting an index, or reindexing it with more shards. ###### restore_opensearch_backup Depending on your service plan, you can restore OpenSearch® backups from - [Restore an OpenSearch® backup](/products/opensearch/howto/restore_opensearch_backup.md): Depending on your service plan, you can restore OpenSearch® backups from ###### saml-sso-authentication SAML (Security Assertion Markup Language) is a standard protocol for exchanging authentication and authorization data between an identity provider (IdP) and a Service Provider (SP). - [Enable SAML authentication on Aiven for OpenSearch®](/products/opensearch/howto/saml-sso-authentication.md): SAML (Security Assertion Markup Language) is a standard protocol for exchanging authentication and authorization data between an identity provider (IdP) and a Service Provider (SP). ###### sample-dataset Databases are more fun with data, so to get you started on your OpenSearch® journey we picked this open data set of recipes as a great example you can try out yourself. - [Sample dataset](/products/opensearch/howto/sample-dataset.md): Databases are more fun with data, so to get you started on your OpenSearch® journey we picked this open data set of recipes as a great example you can try out yourself. ###### scale-disk-storage Scale the disk storage of your Aiven for OpenSearch® service up or down without disrupting the running service. - [Scale disk storage for your Aiven for OpenSearch® service](/products/opensearch/howto/scale-disk-storage.md): Scale the disk storage of your Aiven for OpenSearch® service up or down without disrupting the running service. ###### set_index_retention_patterns Learn to set index retention patterns and manage maximum indices in your Aiven for OpenSearch® instance. - [Index retention patterns](/products/opensearch/howto/set_index_retention_patterns.md): Learn to set index retention patterns and manage maximum indices in your Aiven for OpenSearch® instance. ###### setup-cross-cluster-replication-opensearch Set up cross-cluster replication (CCR) for your Aiven for OpenSearch service to synchronize data across regions and cloud providers efficiently. - [Set up cross-cluster replication for Aiven for OpenSearch® Limited availability](/products/opensearch/howto/setup-cross-cluster-replication-opensearch.md): Set up cross-cluster replication (CCR) for your Aiven for OpenSearch service to synchronize data across regions and cloud providers efficiently. ###### snapshot-credentials Use custom_keystores in Aiven for OpenSearch® to store object storage credentials in Amazon S3, Google Cloud Storage, or Azure. - [Store and manage snapshot repository credentials in Aiven for OpenSearch®](/products/opensearch/howto/snapshot-credentials.md): Use custom_keystores in Aiven for OpenSearch® to store object storage credentials in Amazon S3, Google Cloud Storage, or Azure. ###### tag-service Add key-value tags to your Aiven for OpenSearch® service to organize services and track - [Tag your Aiven for OpenSearch® service](/products/opensearch/howto/tag-service.md): Add key-value tags to your Aiven for OpenSearch® service to organize services and track ###### track-restore-progress Track the restore progress of individual nodes in your Aiven for OpenSearch® service during node replacement, forking, or maintenance, using the Aiven API. - [Track restore progress for your Aiven for OpenSearch® service](/products/opensearch/howto/track-restore-progress.md): Track the restore progress of individual nodes in your Aiven for OpenSearch® service during node replacement, forking, or maintenance, using the Aiven API. ###### upgrade-clients-to-opensearch Elasticsearch has introduced breaking changes into their client libraries as early as 7.13.\*, meaning newer Elasticsearch clients won't work with OpenSearch®. - [Upgrade Elasticsearch clients to OpenSearch®](/products/opensearch/howto/upgrade-clients-to-opensearch.md): Elasticsearch has introduced breaking changes into their client libraries as early as 7.13.\*, meaning newer Elasticsearch clients won't work with OpenSearch®. ##### maintenance-lifecycle Keep your Aiven for OpenSearch® service current by upgrading versions and applying - [Maintenance and lifecycle in Aiven for OpenSearch®](/products/opensearch/maintenance-lifecycle.md): Keep your Aiven for OpenSearch® service current by upgrading versions and applying ##### reference ###### advanced-params See the configuration options available for Aiven for OpenSearch®: - [Advanced parameters for Aiven for OpenSearch®](/products/opensearch/reference/advanced-params.md): See the configuration options available for Aiven for OpenSearch®: ###### list-of-plugins-for-each-version Plugin availability and versions in Aiven for OpenSearch® vary by OpenSearch major version. Each plugin version corresponds to the OpenSearch core version. - [Plugin versions per OpenSearch release](/products/opensearch/reference/list-of-plugins-for-each-version.md): Plugin availability and versions in Aiven for OpenSearch® vary by OpenSearch major version. Each plugin version corresponds to the OpenSearch core version. ###### low-space-watermarks OpenSearch® relies on three indicators to identify and respond to low disk space. - [Low disk space watermarks](/products/opensearch/reference/low-space-watermarks.md): OpenSearch® relies on three indicators to identify and respond to low disk space. ###### opensearch-limitations Aiven for OpenSearch® has configuration, API, and feature restrictions that differ from upstream OpenSearch to maintain service stability and security. - [Aiven for OpenSearch® limits and limitations](/products/opensearch/reference/opensearch-limitations.md): Aiven for OpenSearch® has configuration, API, and feature restrictions that differ from upstream OpenSearch to maintain service stability and security. ###### plugins Aiven for OpenSearch® includes a standard set of plugins. In addition to the plugins that were previously available in Aiven for Elasticsearch, Aiven for OpenSearch also includes plugins that are designed and developed specifically for OpenSearch. - [Plugins available with Aiven for OpenSearch®](/products/opensearch/reference/plugins.md): Aiven for OpenSearch® includes a standard set of plugins. In addition to the plugins that were previously available in Aiven for Elasticsearch, Aiven for OpenSearch also includes plugins that are designed and developed specifically for OpenSearch. ###### version-lifecycle Learn how Aiven manages Aiven for OpenSearch® version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. - [Aiven for OpenSearch® version lifecycle](/products/opensearch/reference/version-lifecycle.md): Learn how Aiven manages Aiven for OpenSearch® version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. ##### scaling-performance Scale resources and tune performance to keep your Aiven for OpenSearch® service running - [Scaling and performance in Aiven for OpenSearch®](/products/opensearch/scaling-performance.md): Scale resources and tune performance to keep your Aiven for OpenSearch® service running ##### troubleshooting ###### troubleshooting-opensearch-dashboards OpenSearch® Dashboards version must match your OpenSearch cluster version. - [OpenSearch® Dashboards incompatible version issues](/products/opensearch/troubleshooting/troubleshooting-opensearch-dashboards.md): OpenSearch® Dashboards version must match your OpenSearch cluster version. #### postgresql Aiven for PostgreSQL® is a fully managed and hosted relational database service. It's a high-performance data warehouse that offers maximum flexibility and functionality with a variety of advanced extensions out of the box. - [Aiven for PostgreSQL®](/products/postgresql.md): Aiven for PostgreSQL® is a fully managed and hosted relational database service. It's a high-performance data warehouse that offers maximum flexibility and functionality with a variety of advanced extensions out of the box. ##### backups-migration Back up, restore, and migrate your Aiven for PostgreSQL® service data. - [Backups and migration in Aiven for PostgreSQL®](/products/postgresql/backups-migration.md): Back up, restore, and migrate your Aiven for PostgreSQL® service data. ##### concepts ###### aiven-db-migrate The aiven-db-migrate tool is the recommended approach for migrating your PostgreSQL® to Aiven. - [Prepare for migrating PostgreSQL® to Aiven using aiven-db-migrate](/products/postgresql/concepts/aiven-db-migrate.md): The aiven-db-migrate tool is the recommended approach for migrating your PostgreSQL® to Aiven. ###### dba-tasks-pg Aiven doesn't allow superuser access to Aiven for PostgreSQL® services. However, most DBA-type actions are still available through other methods. - [Perform DBA-type tasks in Aiven for PostgreSQL®](/products/postgresql/concepts/dba-tasks-pg.md): Aiven doesn't allow superuser access to Aiven for PostgreSQL® services. However, most DBA-type actions are still available through other methods. ###### high-availability Aiven for PostgreSQL® is available on a variety of plans, offering different levels of service availability. The selected plan defines the features available. - [High availability of Aiven for PostgreSQL®](/products/postgresql/concepts/high-availability.md): Aiven for PostgreSQL® is available on a variety of plans, offering different levels of service availability. The selected plan defines the features available. ###### pg-audit-logging The path to optimal data security, compliance, incident management, and system performance starts with collecting robust audit logs. - [Aiven for PostgreSQL® audit logging](/products/postgresql/concepts/pg-audit-logging.md): The path to optimal data security, compliance, incident management, and system performance starts with collecting robust audit logs. ###### pg-backups Aiven for PostgreSQL® databases are automatically backed up, with full backups made daily, and write-ahead logs (WAL) copied at 5 minute intervals, or for every new file generated. - [Aiven for PostgreSQL® backups](/products/postgresql/concepts/pg-backups.md): Aiven for PostgreSQL® databases are automatically backed up, with full backups made daily, and write-ahead logs (WAL) copied at 5 minute intervals, or for every new file generated. ###### pg-connection-pooling Connection pooling in Aiven for PostgreSQL® services allows you to maintain very large numbers of connections to a database while minimizing the consumption of server resources. - [Aiven for PostgreSQL® connection pooling with PgBouncer](/products/postgresql/concepts/pg-connection-pooling.md): Connection pooling in Aiven for PostgreSQL® services allows you to maintain very large numbers of connections to a database while minimizing the consumption of server resources. ###### pg-disk-usage When you create your first Aiven for PostgreSQL® service, you may see the disk usage gradually increasing even though you are not yet inserting or updating much data. - [About PostgreSQL® disk usage](/products/postgresql/concepts/pg-disk-usage.md): When you create your first Aiven for PostgreSQL® service, you may see the disk usage gradually increasing even though you are not yet inserting or updating much data. ###### pg-free-tier Use Aiven for PostgreSQL® for free. You don't need a credit card to sign up - [Aiven for PostgreSQL® free tier](/products/postgresql/concepts/pg-free-tier.md): Use Aiven for PostgreSQL® for free. You don't need a credit card to sign up ###### pg-shared-buffers Use shared buffers to share memory over multiple sessions. Discover how to inspect the database cache performance and the query cache performance and learn how to put data into cache manually. - [Aiven for PostgreSQL® shared buffers](/products/postgresql/concepts/pg-shared-buffers.md): Use shared buffers to share memory over multiple sessions. Discover how to inspect the database cache performance and the query cache performance and learn how to put data into cache manually. ###### pgvector In machine learning (ML) models, all data items in a particular data set are mapped into one unified n-dimensional vector space, no matter how big the input data set is. - [pgvector for AI-powered search in Aiven for PostgreSQL®](/products/postgresql/concepts/pgvector.md): In machine learning (ML) models, all data items in a particular data set are mapped into one unified n-dimensional vector space, no matter how big the input data set is. ###### timescaledb TimescaleDB is an open-source database designed to make your existing relational database scalable for time series data. - [Use TimescaleDB with Aiven for PostgreSQL®](/products/postgresql/concepts/timescaledb.md): TimescaleDB is an open-source database designed to make your existing relational database scalable for time series data. ###### upgrade-failover Aiven for PostgreSQL® Business and Premium plans include standby read-replica servers. If the primary server fails, a standby replica server is automatically promoted as new primary server. - [Aiven for PostgreSQL® upgrade and failover procedures](/products/postgresql/concepts/upgrade-failover.md): Aiven for PostgreSQL® Business and Premium plans include standby read-replica servers. If the primary server fails, a standby replica server is automatically promoted as new primary server. ##### crdr Related pages - [Cross-region disaster recovery in Aiven for PostgreSQL®](/products/postgresql/crdr.md): Related pages ###### crdr-overview The cross-region disaster recovery (CRDR) feature ensures your business continuity by recovering your workloads to a remote region in the event of a region-wide - [Cross-region disaster recovery in Aiven for PostgreSQL® Limited availability](/products/postgresql/crdr/crdr-overview.md): The cross-region disaster recovery (CRDR) feature ensures your business continuity by recovering your workloads to a remote region in the event of a region-wide ###### enable-crdr Enable cross-region disaster recovery (CRDR) in Aiven for PostgreSQL® by creating a recovery service, which takes over from a primary service in case of a region outage. - [Set up cross-region disaster recovery in Aiven for PostgreSQL® Limited availability](/products/postgresql/crdr/enable-crdr.md): Enable cross-region disaster recovery (CRDR) in Aiven for PostgreSQL® by creating a recovery service, which takes over from a primary service in case of a region outage. ###### failover - [Perform Aiven for PostgreSQL® failover to the recovery region Limited availability](/products/postgresql/crdr/failover/crdr-failover-to-recovery.md): Move your workload to another region for disaster recovery or testing purposes. - [Perform Aiven for PostgreSQL® failback to the primary region Limited availability](/products/postgresql/crdr/failover/crdr-revert-to-primary.md): Shift your workloads back to the primary region, where your service was hosted originally before failing over to the recovery region. Restore your CRDR setup. - [Failover & failback](/products/postgresql/crdr/failover/list-failover.md): Perform a ###### switchover - [Perform Aiven for PostgreSQL® switchback to the primary region Limited availability](/products/postgresql/crdr/switchover/crdr-switchback.md): Shift your workloads back to the primary region, where your service was hosted originally before switching over to the recovery region. - [Perform Aiven for PostgreSQL® switchover to the recovery region Limited availability](/products/postgresql/crdr/switchover/crdr-switchover.md): Perform a planned promotion of your recovery service while the primary service is healthy. - [Switchover & switchback](/products/postgresql/crdr/switchover/list-switchover.md): Perform a ##### database-management Create and manage databases, extensions, and routine maintenance tasks in your - [Database management in Aiven for PostgreSQL®](/products/postgresql/database-management.md): Create and manage databases, extensions, and routine maintenance tasks in your ##### get-started Start using Aiven for PostgreSQL® by creating a service, connecting to it, and loading sample data. - [Get started with Aiven for PostgreSQL®](/products/postgresql/get-started.md): Start using Aiven for PostgreSQL® by creating a service, connecting to it, and loading sample data. ##### howto ###### ai-insights Use Aiven's artificial intelligence capabilities to identify slow queries and get optimization suggestions. - [AI database optimizer for Aiven for PostgreSQL®](/products/postgresql/howto/ai-insights.md): Use Aiven's artificial intelligence capabilities to identify slow queries and get optimization suggestions. ###### analyze-with-google-data-studio Google Looker Studio (previously Data Studio) allows you to create reports and visualisations of the data in your Aiven for PostgreSQL® database, and combine these with data from many other data sources. - [Report and analyze with Google Looker Studio](/products/postgresql/howto/analyze-with-google-data-studio.md): Google Looker Studio (previously Data Studio) allows you to create reports and visualisations of the data in your Aiven for PostgreSQL® database, and combine these with data from many other data sources. ###### backup-to-another-region Copy your Aiven for PostgreSQL® service backups to a secondary region for disaster recovery. - [Back up your Aiven for PostgreSQL® service to another region](/products/postgresql/howto/backup-to-another-region.md): Copy your Aiven for PostgreSQL® service backups to a secondary region for disaster recovery. ###### change-service-plan Change the service plan for your Aiven for PostgreSQL® service to scale resources up or down and optimize costs. - [Change the plan for your Aiven for PostgreSQL® service](/products/postgresql/howto/change-service-plan.md): Change the service plan for your Aiven for PostgreSQL® service to scale resources up or down and optimize costs. ###### check-avoid-transaction-id-wraparound The PostgreSQL® transaction control mechanism assigns a transaction ID - [Check and avoid transaction ID wraparound](/products/postgresql/howto/check-avoid-transaction-id-wraparound.md): The PostgreSQL® transaction control mechanism assigns a transaction ID ###### claim-public-schema-ownership When an Aiven for PostgreSQL® instance is created, the public schema - [Claim public schema ownership](/products/postgresql/howto/claim-public-schema-ownership.md): When an Aiven for PostgreSQL® instance is created, the public schema ###### connect-datagrip Use DataGrip to connect to your Aiven for - [Connect to Aiven for PostgreSQL® with DataGrip](/products/postgresql/howto/connect-datagrip.md): Use DataGrip to connect to your Aiven for ###### connect-dbeaver Use DBeaver to connect to your Aiven for PostgreSQL® service. - [Connect to Aiven for PostgreSQL® with DBeaver](/products/postgresql/howto/connect-dbeaver.md): Use DBeaver to connect to your Aiven for PostgreSQL® service. ###### connect-go This example connects to PostgreSQL® service from Go, making use of the - [Connect to Aiven for PostgreSQL® with Go](/products/postgresql/howto/connect-go.md): This example connects to PostgreSQL® service from Go, making use of the ###### connect-java This example connects to PostgreSQL® service from Java, making use of JDBC Driver. - [Connect to Aiven for PostgreSQL® with Java](/products/postgresql/howto/connect-java.md): This example connects to PostgreSQL® service from Java, making use of JDBC Driver. ###### connect-libredb-studio Use LibreDB Studio to connect to your Aiven for PostgreSQL® - [Connect to Aiven for PostgreSQL® with LibreDB Studio](/products/postgresql/howto/connect-libredb-studio.md): Use LibreDB Studio to connect to your Aiven for PostgreSQL® ###### connect-node This example connects to PostgreSQL® service from NodeJS, making use of the pg package. - [Connect to Aiven for PostgreSQL® with NodeJS](/products/postgresql/howto/connect-node.md): This example connects to PostgreSQL® service from NodeJS, making use of the pg package. ###### connect-pgadmin pgAdmin is one of the most popular PostgreSQL® clients. Use it to manage and query your database. - [Connect to Aiven for PostgreSQL® with pgAdmin](/products/postgresql/howto/connect-pgadmin.md): pgAdmin is one of the most popular PostgreSQL® clients. Use it to manage and query your database. ###### connect-php This example connects to PostgreSQL® service from PHP, making use of the - [Connect to Aiven for PostgreSQL® with PHP](/products/postgresql/howto/connect-php.md): This example connects to PostgreSQL® service from PHP, making use of the ###### connect-psql psql is a command line tool for PostgreSQL®, useful to manage and - [Connect to Aiven for PostgreSQL® with psql](/products/postgresql/howto/connect-psql.md): psql is a command line tool for PostgreSQL®, useful to manage and ###### connect-python This example connects to a PostgreSQL® service from Python, making use - [Connect to Aiven for PostgreSQL® with Python](/products/postgresql/howto/connect-python.md): This example connects to a PostgreSQL® service from Python, making use ###### connect-rivery Configure PostgreSQL® as a connection in Rivery, a fully managed solution for data ingestion, transformation, orchestration, reverse ETL and more. - [Connect to Aiven for PostgreSQL® with Rivery](/products/postgresql/howto/connect-rivery.md): Configure PostgreSQL® as a connection in Rivery, a fully managed solution for data ingestion, transformation, orchestration, reverse ETL and more. ###### connect-skyvia Skyvia is an universal cloud data platform. This - [Connect to Aiven for PostgreSQL® with Skyvia](/products/postgresql/howto/connect-skyvia.md): Skyvia is an universal cloud data platform. This ###### connect-zapier Zapier is an automation platform that connects - [Connect to Aiven for PostgreSQL® with Zapier](/products/postgresql/howto/connect-zapier.md): Zapier is an automation platform that connects ###### controlled-upgrade-pipelines Link Aiven for PostgreSQL® services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. - [Controlled upgrade pipelines for your Aiven for PostgreSQL® service Limited availability](/products/postgresql/howto/controlled-upgrade-pipelines.md): Link Aiven for PostgreSQL® services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. ###### create-database Once you've created your Aiven for PostgreSQL® service, you can add additional databases, whether for security purposes or to isolate your data per application. - [Create PostgreSQL® databases](/products/postgresql/howto/create-database.md): Once you've created your Aiven for PostgreSQL® service, you can add additional databases, whether for security purposes or to isolate your data per application. ###### create-manual-backups Aiven provides fully automated backup management for PostgreSQL. - [Create manual PostgreSQL® backups with pg_dump](/products/postgresql/howto/create-manual-backups.md): Aiven provides fully automated backup management for PostgreSQL. ###### create-read-replica Use Aiven for PostgreSQL® read-only replicas to reduce the load on the primary server and optimize query response times across different geographical locations. - [Create and use Aiven for PostgreSQL® read-only replicas](/products/postgresql/howto/create-read-replica.md): Use Aiven for PostgreSQL® read-only replicas to reduce the load on the primary server and optimize query response times across different geographical locations. ###### data-api Data API turns your Aiven for PostgreSQL® database into a backend by exposing its tables as secure REST endpoints, without backend code. - [Data API for Aiven for PostgreSQL® Limited availability](/products/postgresql/howto/data-api.md): Data API turns your Aiven for PostgreSQL® database into a backend by exposing its tables as secure REST endpoints, without backend code. - [Configure authentication for Aiven for PostgreSQL® Data API Limited availability](/products/postgresql/howto/data-api/authentication.md): Authenticate Data API requests with JWTs from your own identity provider, and authorize with PostgreSQL roles. - [Enable Aiven for PostgreSQL® Data API Limited availability](/products/postgresql/howto/data-api/get-started.md): Expose an Aiven for PostgreSQL database as REST endpoints. - [Manage Aiven for PostgreSQL® Data API Limited availability](/products/postgresql/howto/data-api/manage.md): Check status, expose more databases, and remove the Data API. - [Call the Aiven for PostgreSQL® Data API endpoints Limited availability](/products/postgresql/howto/data-api/use-endpoints.md): Find your API URL and call your database over HTTPS with bearer token authentication. ###### datasource-integration There are two types of datasource integrations you can use with Aiven for PostgreSQL®: - [Connect two PostgreSQL® services via datasource integration](/products/postgresql/howto/datasource-integration.md): There are two types of datasource integrations you can use with Aiven for PostgreSQL®: ###### disk-autoscaler Automatically increase the disk storage of your Aiven for PostgreSQL® service when it's running out of space, instead of resizing it manually. - [Scale disk storage automatically for your Aiven for PostgreSQL® service](/products/postgresql/howto/disk-autoscaler.md): Automatically increase the disk storage of your Aiven for PostgreSQL® service when it's running out of space, instead of resizing it manually. ###### enable-jit PostgreSQL® 11 introduces a new component in the execution engine, a Just-in-Time (JIT) expression compiler. - [Enable JIT in PostgreSQL®](/products/postgresql/howto/enable-jit.md): PostgreSQL® 11 introduces a new component in the execution engine, a Just-in-Time (JIT) expression compiler. ###### fork-service Fork your Aiven for PostgreSQL® service to create an independent copy for testing, - [Fork your Aiven for PostgreSQL® service](/products/postgresql/howto/fork-service.md): Fork your Aiven for PostgreSQL® service to create an independent copy for testing, ###### identify-pg-slow-queries Use the PostgreSQL® pgstatstatements extension to find slow queries. - [Identify PostgreSQL® slow queries with pg_stat_statements](/products/postgresql/howto/identify-pg-slow-queries.md): Use the PostgreSQL® pgstatstatements extension to find slow queries. ###### list-code-samples Connect to the Aiven for PostgreSQL® service using various programming languages or tools. All connections to PostgreSQL are encrypted and protected with TLS. - [Connect to Aiven for PostgreSQL® services](/products/postgresql/howto/list-code-samples.md): Connect to the Aiven for PostgreSQL® service using various programming languages or tools. All connections to PostgreSQL are encrypted and protected with TLS. ###### list-pgaudit - [pgaudit logging](/products/postgresql/howto/list-pgaudit.md) ###### logical-replication-aws-aurora If you have not enabled logical replication on Aurora already, the - [Enable logical replication on Amazon Aurora PostgreSQL®](/products/postgresql/howto/logical-replication-aws-aurora.md): If you have not enabled logical replication on Aurora already, the ###### logical-replication-aws-rds If you have not enabled logical replication on RDS already, the - [Enable logical replication on Amazon RDS PostgreSQL®](/products/postgresql/howto/logical-replication-aws-rds.md): If you have not enabled logical replication on RDS already, the ###### logical-replication-gcp-cloudsql If you have not enabled logical replication on Google Cloud SQL PostgreSQL® already, set the cloudsql.logical_decoding parameter to On: - [Enable logical replication on Google Cloud SQL](/products/postgresql/howto/logical-replication-gcp-cloudsql.md): If you have not enabled logical replication on Google Cloud SQL PostgreSQL® already, set the cloudsql.logical_decoding parameter to On: ###### maintenance-updates Manage maintenance updates and set the maintenance window for your Aiven for PostgreSQL® service. - [Maintenance and updates for your Aiven for PostgreSQL® service](/products/postgresql/howto/maintenance-updates.md): Manage maintenance updates and set the maintenance window for your Aiven for PostgreSQL® service. ###### manage-extensions Install, update, and remove PostgreSQL® extensions on Aiven for PostgreSQL using SQL commands. - [Manage Aiven for PostgreSQL® extensions](/products/postgresql/howto/manage-extensions.md): Install, update, and remove PostgreSQL® extensions on Aiven for PostgreSQL using SQL commands. ###### manage-pool Connection pooling lets you maintain very large numbers of connections to a database while minimizing the consumption of server resources. - [Manage connection pooling](/products/postgresql/howto/manage-pool.md): Connection pooling lets you maintain very large numbers of connections to a database while minimizing the consumption of server resources. ###### manage-service-users Create and manage service users in your Aiven for PostgreSQL® service to control access - [Manage Aiven for PostgreSQL® service users](/products/postgresql/howto/manage-service-users.md): Create and manage service users in your Aiven for PostgreSQL® service to control access ###### migrate-aiven-db-migrate The aiven-db-migrate tool is an open source project available on GitHub, and it is the preferred way to perform PostgreSQL® database migration. - [Migrate PostgreSQL® databases to Aiven using aiven-db-migrate](/products/postgresql/howto/migrate-aiven-db-migrate.md): The aiven-db-migrate tool is an open source project available on GitHub, and it is the preferred way to perform PostgreSQL® database migration. ###### migrate-cloud-region Any Aiven service can be relocated to a different cloud vendor or region. This is also valid for PostgreSQL® where the migration happens without downtime. - [Migrate to a different cloud provider or region](/products/postgresql/howto/migrate-cloud-region.md): Any Aiven service can be relocated to a different cloud vendor or region. This is also valid for PostgreSQL® where the migration happens without downtime. ###### migrate-db-to-aiven-via-console Migrate PostgreSQL databases to the Aiven platform using the Aiven Console. - [Migrate PostgreSQL® databases to Aiven using the Aiven Console](/products/postgresql/howto/migrate-db-to-aiven-via-console.md): Migrate PostgreSQL databases to the Aiven platform using the Aiven Console. ###### migrate-pg-dump-restore Aiven for PostgreSQL® supports the same tools as a regular PostgreSQL database, so you can migrate using the standard pgdump and pgrestore tools. - [Migrate PostgreSQL® databases to Aiven using pg_dump and pg_restore](/products/postgresql/howto/migrate-pg-dump-restore.md): Aiven for PostgreSQL® supports the same tools as a regular PostgreSQL database, so you can migrate using the standard pgdump and pgrestore tools. ###### migrate-using-bucardo The preferred approach to migrating a database to Aiven for PostgreSQL® is to use Aiven's open source migration tool (About aiven-db-migrate). - [Migrate PostgreSQL® databases to Aiven using Bucardo](/products/postgresql/howto/migrate-using-bucardo.md): The preferred approach to migrating a database to Aiven for PostgreSQL® is to use Aiven's open source migration tool (About aiven-db-migrate). ###### monitor-database-with-datadog Database Monitoring with Datadog enables you to capture key metrics on the Datadog platform for any Aiven for PostgreSQL® service with Datadog Metrics integration. - [Monitor a database with Datadog](/products/postgresql/howto/monitor-database-with-datadog.md): Database Monitoring with Datadog enables you to capture key metrics on the Datadog platform for any Aiven for PostgreSQL® service with Datadog Metrics integration. ###### monitor-pgbouncer-with-datadog Integrate PgBouncer with Datadog to track connection pool metrics and monitor application traffic on the Datadog platform. - [Monitor PgBouncer with Datadog for Aiven for PostgreSQL®](/products/postgresql/howto/monitor-pgbouncer-with-datadog.md): Integrate PgBouncer with Datadog to track connection pool metrics and monitor application traffic on the Datadog platform. ###### monitor-relation-function-metrics-datadog Configure the Datadog Metrics integration to collect per-table and per-index statistics, and per-function call statistics for PL/pgSQL functions, for your Aiven for PostgreSQL® service. - [Collect relation and function metrics with Datadog for Aiven for PostgreSQL®](/products/postgresql/howto/monitor-relation-function-metrics-datadog.md): Configure the Datadog Metrics integration to collect per-table and per-index statistics, and per-function call statistics for PL/pgSQL functions, for your Aiven for PostgreSQL® service. ###### monitor-with-pgwatch2 pgwatch2 is an open source monitoring solution for PostgreSQL®, created by CYBERTEC and can be used to monitor instances of Aiven for PostgreSQL collecting key PostgreSQL metrics and also gathering data from a wide range of PostgreSQL extensions. - [Monitor PostgreSQL® metrics with pgwatch2](/products/postgresql/howto/monitor-with-pgwatch2.md): pgwatch2 is an open source monitoring solution for PostgreSQL®, created by CYBERTEC and can be used to monitor instances of Aiven for PostgreSQL collecting key PostgreSQL metrics and also gathering data from a wide range of PostgreSQL extensions. ###### optimize-pg-slow-queries Aiven for PostgreSQL allows you to identify slow queries using the pgstatstatements view. - [Optimize Aiven for PostgreSQL® slow queries](/products/postgresql/howto/optimize-pg-slow-queries.md): Aiven for PostgreSQL allows you to identify slow queries using the pgstatstatements view. ###### pagila Use a sample database that you can import in your Aiven for PostgreSQL® service. - [Sample dataset for PostgreSQL®: Pagila](/products/postgresql/howto/pagila.md): Use a sample database that you can import in your Aiven for PostgreSQL® service. ###### pg-controlled-switchover Control when primary node switchover happens during Aiven for PostgreSQL® maintenance. - [Controlled maintenance updates in Aiven for PostgreSQL® Limited availability](/products/postgresql/howto/pg-controlled-switchover.md): Control when primary node switchover happens during Aiven for PostgreSQL® maintenance. ###### pg-long-running-queries Aiven does not terminate any customer queries even if they run - [Detect and terminate long-running queries in Aiven for PostgreSQL®](/products/postgresql/howto/pg-long-running-queries.md): Aiven does not terminate any customer queries even if they run ###### pg-object-size PostgreSQL® offers different commands and functions to get disk space usage for a database, a table, or an index. - [Check the size of a database, a table or an index](/products/postgresql/howto/pg-object-size.md): PostgreSQL® offers different commands and functions to get disk space usage for a database, a table, or an index. ###### pg-reads-failover-to-primary Enable automatic failover for your Aiven for PostgreSQL® read workloads to ensure uninterrupted access when standby nodes are unavailable. - [Aiven for PostgreSQL® reads failover to the primary](/products/postgresql/howto/pg-reads-failover-to-primary.md): Enable automatic failover for your Aiven for PostgreSQL® read workloads to ensure uninterrupted access when standby nodes are unavailable. ###### pg-studio Aiven PG Studio lets you write, understand, and run SQL queries in the Aiven Console using natural language. It combines an SQL editor with an AI assistant that uses your database schema to generate and explain queries. - [PG Studio for Aiven for PostgreSQL® Early availability](/products/postgresql/howto/pg-studio.md): Aiven PG Studio lets you write, understand, and run SQL queries in the Aiven Console using natural language. It combines an SQL editor with an AI assistant that uses your database schema to generate and explain queries. - [Get started with PG Studio](/products/postgresql/howto/pg-studio/get-started.md): Open PG Studio and run your first queries. - [Manage queries in PG Studio](/products/postgresql/howto/pg-studio/manage-queries.md): Save and organize your queries for reuse. - [Security and connections in PG Studio](/products/postgresql/howto/pg-studio/security-connections.md): Understand how PG Studio connects and protects your data. - [Use AI Assistant in PG Studio](/products/postgresql/howto/pg-studio/use-ai-assistant.md): Generate and explain SQL queries with natural language. - [Write and run queries in PG Studio](/products/postgresql/howto/pg-studio/write-run-queries.md): Use the SQL editor to write, edit, and execute queries. ###### pgbouncer-stats PgBouncer is used at Aiven as a connection pooler to lower the performance impact of opening new connections to Aiven for PostgreSQL®. - [Access PgBouncer statistics for Aiven for PostgreSQL®](/products/postgresql/howto/pgbouncer-stats.md): PgBouncer is used at Aiven as a connection pooler to lower the performance impact of opening new connections to Aiven for PostgreSQL®. ###### power-cycle-service Power off your Aiven for PostgreSQL® service to release resources and save credits, power it back on when you need it, or delete it permanently. - [Power on/off and delete your Aiven for PostgreSQL® service](/products/postgresql/howto/power-cycle-service.md): Power off your Aiven for PostgreSQL® service to release resources and save credits, power it back on when you need it, or delete it permanently. ###### prepare-for-high-load Prepare your Aiven for PostgreSQL® service for higher than usual traffic to avoid outages and keep performance stable. - [Prepare your Aiven for PostgreSQL® service for high load](/products/postgresql/howto/prepare-for-high-load.md): Prepare your Aiven for PostgreSQL® service for higher than usual traffic to avoid outages and keep performance stable. ###### prevent-full-disk If your Aiven for PostgreSQL® service runs out of disk space, the service will start malfunctioning, which will preclude proper backup creation. - [Prevent PostgreSQL® full disk issues](/products/postgresql/howto/prevent-full-disk.md): If your Aiven for PostgreSQL® service runs out of disk space, the service will start malfunctioning, which will preclude proper backup creation. ###### readonly-user You can restrict access to Aiven for PostgreSQL® databases and tables by - [Restrict access to databases or tables in Aiven for PostgreSQL®](/products/postgresql/howto/readonly-user.md): You can restrict access to Aiven for PostgreSQL® databases and tables by ###### rename-service Change the name of your Aiven for PostgreSQL® service by forking it under a new name and - [Rename your Aiven for PostgreSQL® service](/products/postgresql/howto/rename-service.md): Change the name of your Aiven for PostgreSQL® service by forking it under a new name and ###### repair-pg-index PostgreSQL® indexes can become corrupted due to a variety of reasons including software bugs, hardware failures or unexpected duplicated data. REINDEX allows you to rebuild the index in such situations. - [Identify and repair issues with PostgreSQL® indexes with REINDEX](/products/postgresql/howto/repair-pg-index.md): PostgreSQL® indexes can become corrupted due to a variety of reasons including software bugs, hardware failures or unexpected duplicated data. REINDEX allows you to rebuild the index in such situations. ###### report-metrics-grafana As well as offering PostgreSQL-as-a-service, the Aiven platform gives - [Monitor PostgreSQL® metrics with Grafana®](/products/postgresql/howto/report-metrics-grafana.md): As well as offering PostgreSQL-as-a-service, the Aiven platform gives ###### restore-backup Aiven for PostgreSQL® databases are automatically backed up and can be restored from a backup at any point in time within the backup retention period, which varies by plan. - [Restore PostgreSQL® from a backup](/products/postgresql/howto/restore-backup.md): Aiven for PostgreSQL® databases are automatically backed up and can be restored from a backup at any point in time within the backup retention period, which varies by plan. ###### run-aiven-db-migrate-python The aiven-db-migrate tool is an open source project available on - [Migrate between PostgreSQL® instances using aiven-db-migrate in Python](/products/postgresql/howto/run-aiven-db-migrate-python.md): The aiven-db-migrate tool is an open source project available on ###### scale-disk-storage Scale the disk storage of your Aiven for PostgreSQL® service up or down without disrupting the running service. - [Scale disk storage for your Aiven for PostgreSQL® service](/products/postgresql/howto/scale-disk-storage.md): Scale the disk storage of your Aiven for PostgreSQL® service up or down without disrupting the running service. ###### setup-logical-replication Aiven for PostgreSQL® represents an ideal managed solution for a variety of use cases; remote production systems can be completely migrated to Aiven using different methods including using Aiven-db-migrate or the standard dump and restore method. - [Set up logical replication to Aiven for PostgreSQL®](/products/postgresql/howto/setup-logical-replication.md): Aiven for PostgreSQL® represents an ideal managed solution for a variety of use cases; remote production systems can be completely migrated to Aiven using different methods including using Aiven-db-migrate or the standard dump and restore method. ###### tag-service Add key-value tags to your Aiven for PostgreSQL® service to organize services and track ownership, cost allocation, and governance. - [Tag your Aiven for PostgreSQL® service](/products/postgresql/howto/tag-service.md): Add key-value tags to your Aiven for PostgreSQL® service to organize services and track ownership, cost allocation, and governance. ###### track-restore-progress Track the restore progress of individual nodes in your Aiven for PostgreSQL® service during node replacement, forking, or maintenance, using the Aiven API. - [Track restore progress for your Aiven for PostgreSQL® service](/products/postgresql/howto/track-restore-progress.md): Track the restore progress of individual nodes in your Aiven for PostgreSQL® service during node replacement, forking, or maintenance, using the Aiven API. ###### upgrade PostgreSQL® in-place upgrades allows to upgrade an instances to a new major version without needing to fork and redirect the traffic. - [Perform a PostgreSQL® major version upgrade](/products/postgresql/howto/upgrade.md): PostgreSQL® in-place upgrades allows to upgrade an instances to a new major version without needing to fork and redirect the traffic. ###### upgrade-postgis-topology-columns Troubleshoot issues that block PostGIS® extension upgrades on Aiven for PostgreSQL® services, and complete the upgrade safely. - [Troubleshoot PostGIS® upgrade issues](/products/postgresql/howto/upgrade-postgis-topology-columns.md): Troubleshoot issues that block PostGIS® extension upgrades on Aiven for PostgreSQL® services, and complete the upgrade safely. ###### use-dblink-extension dblink is a PostgreSQL® extension that allows you to connect to other PostgreSQL databases and to run arbitrary queries. - [Use the PostgreSQL® dblink extension](/products/postgresql/howto/use-dblink-extension.md): dblink is a PostgreSQL® extension that allows you to connect to other PostgreSQL databases and to run arbitrary queries. ###### use-pg-audit-logging Enable and configure the Aiven for PostgreSQL® audit logging feature on your service. Access and visualize your logs to monitor activities on your databases. - [Collect audit logs in Aiven for PostgreSQL®](/products/postgresql/howto/use-pg-audit-logging.md): Enable and configure the Aiven for PostgreSQL® audit logging feature on your service. Access and visualize your logs to monitor activities on your databases. ###### use-pg-cron-extension The pg_cron extension is a cron-based job scheduler for PostgreSQL (10 or higher) that runs inside the database. - [Use the PostgreSQL® pg_cron extension](/products/postgresql/howto/use-pg-cron-extension.md): The pg_cron extension is a cron-based job scheduler for PostgreSQL (10 or higher) that runs inside the database. ###### use-pg-repack-extension pgrepack is a PostgreSQL® extension that allows you to efficiently reorganize tables to remove any excess bloat the tables have accumulated. - [Use the PostgreSQL® pg_repack extension](/products/postgresql/howto/use-pg-repack-extension.md): pgrepack is a PostgreSQL® extension that allows you to efficiently reorganize tables to remove any excess bloat the tables have accumulated. ###### use-pgvector The pgvector extension allows you to perform the vector similarity search and use embedding techniques directly in Aiven for PostgreSQL. - [Enable and use pgvector on Aiven for PostgreSQL®](/products/postgresql/howto/use-pgvector.md): The pgvector extension allows you to perform the vector similarity search and use embedding techniques directly in Aiven for PostgreSQL. ###### visualize-grafana PostgreSQL® can hold a wide variety of types of data, and creating visualisations helps gather insights on top of raw figures. Aiven can set up the Grafana® and the integration between the two services for you. - [Visualize PostgreSQL® data with Grafana®](/products/postgresql/howto/visualize-grafana.md): PostgreSQL® can hold a wide variety of types of data, and creating visualisations helps gather insights on top of raw figures. Aiven can set up the Grafana® and the integration between the two services for you. ##### maintenance-lifecycle Keep your Aiven for PostgreSQL® service current by upgrading versions and applying - [Maintenance and lifecycle in Aiven for PostgreSQL®](/products/postgresql/maintenance-lifecycle.md): Keep your Aiven for PostgreSQL® service current by upgrading versions and applying ##### reference ###### advanced-params On creating a PostgreSQL® database, you can customize it using a series - [Advanced parameters for Aiven for PostgreSQL®](/products/postgresql/reference/advanced-params.md): On creating a PostgreSQL® database, you can customize it using a series ###### idle-connections PostgreSQL® keep-alive connection parameters are useful to manage Idle - [Keep-alive connections parameters](/products/postgresql/reference/idle-connections.md): PostgreSQL® keep-alive connection parameters are useful to manage Idle ###### list-of-extensions PostgreSQL® extensions allow you to extend the functionality by adding more capabilities to your Aiven for PostgreSQL. - [Extensions on Aiven for PostgreSQL®](/products/postgresql/reference/list-of-extensions.md): PostgreSQL® extensions allow you to extend the functionality by adding more capabilities to your Aiven for PostgreSQL. ###### list-of-extensions-for-each-version Extension availability and versions in Aiven for PostgreSQL® vary by PostgreSQL major version, with each extension having one default version per PostgreSQL release. - [Extension versions per PostgreSQL release](/products/postgresql/reference/list-of-extensions-for-each-version.md): Extension availability and versions in Aiven for PostgreSQL® vary by PostgreSQL major version, with each extension having one default version per PostgreSQL release. ###### log-formats-supported Aiven for PostgreSQL® supports setting different log formats which are compatible with popular log analysis tools like pgbadger or pganalyze. - [Supported log formats](/products/postgresql/reference/log-formats-supported.md): Aiven for PostgreSQL® supports setting different log formats which are compatible with popular log analysis tools like pgbadger or pganalyze. ###### pg-connection-limits Find the default max_connections value for each Aiven for PostgreSQL® plan, and learn - [Connection limits per plan for Aiven for PostgreSQL®](/products/postgresql/reference/pg-connection-limits.md): Find the default max_connections value for each Aiven for PostgreSQL® plan, and learn ###### pg-metrics The metrics/dashboard integration in the Aiven console enables you to push PostgreSQL® metrics to an external endpoint like Datadog or to create an integration and a prebuilt dashboard in Aiven for Grafana®. - [PostgreSQL® metrics exposed in Grafana®](/products/postgresql/reference/pg-metrics.md): The metrics/dashboard integration in the Aiven console enables you to push PostgreSQL® metrics to an external endpoint like Datadog or to create an integration and a prebuilt dashboard in Aiven for Grafana®. ###### resource-capability Aiven also uses [industry-standard - [Resource capability of Aiven for PostgreSQL® plans](/products/postgresql/reference/resource-capability.md): Aiven also uses [industry-standard ###### terminology - Primary node: The PostgreSQL® primary node is the main server - [Terminology for PostgreSQL®](/products/postgresql/reference/terminology.md): - Primary node: The PostgreSQL® primary node is the main server ###### use-of-deprecated-tls-versions TLS versions TLSv1 and TLSv1.1 are considered insecure, and are no longer supported in Aiven for PostgreSQL® deployments. - [Use of deprecated TLS versions](/products/postgresql/reference/use-of-deprecated-tls-versions.md): TLS versions TLSv1 and TLSv1.1 are considered insecure, and are no longer supported in Aiven for PostgreSQL® deployments. ###### version-lifecycle Learn how Aiven manages Aiven for PostgreSQL® version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. - [Aiven for PostgreSQL® version lifecycle](/products/postgresql/reference/version-lifecycle.md): Learn how Aiven manages Aiven for PostgreSQL® version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. ##### scaling-performance Tune performance and manage disk usage to keep your Aiven for PostgreSQL® service - [Scaling and performance in Aiven for PostgreSQL®](/products/postgresql/scaling-performance.md): Tune performance and manage disk usage to keep your Aiven for PostgreSQL® service ##### troubleshooting ###### pg-password-encryption-upgrade Verify that your Aiven for PostgreSQL® connections use scram-sha-256 password encryption. - [Verify the Aiven for PostgreSQL® password encryption method](/products/postgresql/troubleshooting/pg-password-encryption-upgrade.md): Verify that your Aiven for PostgreSQL® connections use scram-sha-256 password encryption. ###### troubleshooting-connection-pooling Discover the PgBouncer connection pooler and learn how to cope with some specific connection pooling issues. - [Troubleshoot connection pooling issues in Aiven for PostgreSQL®](/products/postgresql/troubleshooting/troubleshooting-connection-pooling.md): Discover the PgBouncer connection pooler and learn how to cope with some specific connection pooling issues. ###### troubleshooting-fatal-out-of-shared-mem Identify and resolve the out of shared memory issue caused by stuck sessions. - [Troubleshoot out-of-shared-memory errors](/products/postgresql/troubleshooting/troubleshooting-fatal-out-of-shared-mem.md): Identify and resolve the out of shared memory issue caused by stuck sessions. #### runtime Aiven Runtime lets you deploy and run containerized applications directly within your existing Aiven project infrastructure. - [Aiven Runtime overview](/products/runtime.md): Aiven Runtime lets you deploy and run containerized applications directly within your existing Aiven project infrastructure. ##### change-cloud You can change the cloud provider or region of an Aiven Runtime application. - [Change cloud for Aiven Runtime](/products/runtime/change-cloud.md): You can change the cloud provider or region of an Aiven Runtime application. ##### connect-github-account Connect your GitHub account to deploy applications from your GitHub repositories. - [Connect or configure a GitHub account](/products/runtime/connect-github-account.md): Connect your GitHub account to deploy applications from your GitHub repositories. ##### connect-services-to-apps Connect your deployed application to Aiven services. - [Connect services to Aiven Runtime](/products/runtime/connect-services-to-apps.md): Connect your deployed application to Aiven services. ##### custom-domain Connect a custom domain to an Aiven Runtime application using Cloudflare. Cloudflare receives traffic for your custom domain at its edge, and a Cloudflare Worker forwards each request to the Aiven-generated application hostname. - [Connect a custom domain to an Aiven Runtime](/products/runtime/custom-domain.md): Connect a custom domain to an Aiven Runtime application using Cloudflare. Cloudflare receives traffic for your custom domain at its edge, and a Cloudflare Worker forwards each request to the Aiven-generated application hostname. ##### deploy-apps Build and deploy applications using Aiven Runtime from source code in a GitHub repository. - [Deploy an application](/products/runtime/deploy-apps.md): Build and deploy applications using Aiven Runtime from source code in a GitHub repository. ##### deployment-information You can change your application's branch at any time. - [Change branch](/products/runtime/deployment-information.md): You can change your application's branch at any time. ##### manifest-files ###### compose-files Aiven Runtime scans your repository for Compose files, such as Docker Compose files, to detect applications, identify supported data services, and create integrations. - [Create Compose files for Aiven Runtime](/products/runtime/manifest-files/compose-files.md): Aiven Runtime scans your repository for Compose files, such as Docker Compose files, to detect applications, identify supported data services, and create integrations. ###### containerfiles Aiven Runtime automatically detects and analyzes Containerfiles and Dockerfiles in your repository to configure applications. - [Create Containerfiles and Dockerfiles for Aiven Runtime](/products/runtime/manifest-files/containerfiles.md): Aiven Runtime automatically detects and analyzes Containerfiles and Dockerfiles in your repository to configure applications. ###### manifests Aiven Runtime uses container manifests to understand how to build and deploy your applications. You can define applications using two types of container manifests that work together to create complete solutions. - [Manifest files for Aiven Runtime](/products/runtime/manifest-files/manifests.md): Aiven Runtime uses container manifests to understand how to build and deploy your applications. You can define applications using two types of container manifests that work together to create complete solutions. ##### ports To make your application available on public networks, you can configure it to listen on ports for HTTP/S traffic. - [Manage ports for Aiven Runtime](/products/runtime/ports.md): To make your application available on public networks, you can configure it to listen on ports for HTTP/S traffic. ##### power-off-apps You can power an Aiven Runtime application on or off at any time. - [Power off Aiven Runtime applications](/products/runtime/power-off-apps.md): You can power an Aiven Runtime application on or off at any time. ##### scale-apps Adjust the plan of your applications at any time to scale them and optimize costs. - [Change application plan for Aiven Runtime](/products/runtime/scale-apps.md): Adjust the plan of your applications at any time to scale them and optimize costs. ##### secrets-and-variables Environment variables and secrets let you configure your application at runtime instead of embedding settings and sensitive information into your code. - [Manage secrets and environment variables for Aiven Runtime](/products/runtime/secrets-and-variables.md): Environment variables and secrets let you configure your application at runtime instead of embedding settings and sensitive information into your code. #### services Deploy fully managed and scalable open source data technologies as individual services and advanced data pipelines in minutes. - [Services](/products/services.md): Deploy fully managed and scalable open source data technologies as individual services and advanced data pipelines in minutes. #### valkey Aiven for Valkey™ is a fully managed in-memory NoSQL database service that offers high performance, scalability, and security. Deployable in the cloud of your choice, it helps you store and access data efficiently. - [Aiven for Valkey™](/products/valkey.md): Aiven for Valkey™ is a fully managed in-memory NoSQL database service that offers high performance, scalability, and security. Deployable in the cloud of your choice, it helps you store and access data efficiently. ##### backups-migration Back up and migrate your Aiven for Valkey™ service data. - [Backups and migration in Aiven for Valkey™](/products/valkey/backups-migration.md): Back up and migrate your Aiven for Valkey™ service data. ##### concepts ###### high-availability Explore high availability with Aiven for Valkey™ across multiple plans. Gain insights into service continuity and understand the approach to handling failures. - [High availability in Aiven for Valkey™](/products/valkey/concepts/high-availability.md): Explore high availability with Aiven for Valkey™ across multiple plans. Gain insights into service continuity and understand the approach to handling failures. ###### lua-scripts Learn how to leverage the built-in support for Lua scripting in Aiven for Valkey™. - [Lua scripts with Aiven for Valkey™](/products/valkey/concepts/lua-scripts.md): Learn how to leverage the built-in support for Lua scripting in Aiven for Valkey™. ###### memory-usage Learn how Aiven for Valkey™ addresses the challenges of high memory usage and high change rate. Discover how it implements robust memory management and persistence strategies. - [Memory management and persistence in Aiven for Valkey™](/products/valkey/concepts/memory-usage.md): Learn how Aiven for Valkey™ addresses the challenges of high memory usage and high change rate. Discover how it implements robust memory management and persistence strategies. ###### read-replica Aiven for Valkey™ read replica replicates data from a primary service to a replica service across different DNS zones, clouds, or regions, enhancing data availability and supporting disaster recovery. - [Aiven for Valkey™ read replica Early availability](/products/valkey/concepts/read-replica.md): Aiven for Valkey™ read replica replicates data from a primary service to a replica service across different DNS zones, clouds, or regions, enhancing data availability and supporting disaster recovery. ###### valkey-cluster Aiven for Valkey™ clustering provides a managed, scalable solution for distributed in-memory data storage with built-in high availability and automatic failover capabilities. - [Aiven for Valkey™ clustering Limited availability](/products/valkey/concepts/valkey-cluster.md): Aiven for Valkey™ clustering provides a managed, scalable solution for distributed in-memory data storage with built-in high availability and automatic failover capabilities. ###### valkey-free-tier Use Aiven for Valkey™ for free. You don't need a credit card to sign up - [Aiven for Valkey™ free tier](/products/valkey/concepts/valkey-free-tier.md): Use Aiven for Valkey™ for free. You don't need a credit card to sign up ##### get-started Begin your journey with Aiven for Valkey™, the versatile in-memory data store offering high-performance capabilities for caching, message queues, and efficient data storage solutions. - [Get started with Aiven for Valkey™](/products/valkey/get-started.md): Begin your journey with Aiven for Valkey™, the versatile in-memory data store offering high-performance capabilities for caching, message queues, and efficient data storage solutions. ##### howto ###### backup-to-another-region Copy your Aiven for Valkey™ service backups to a secondary region for disaster recovery. - [Back up your Aiven for Valkey™ service to another region](/products/valkey/howto/backup-to-another-region.md): Copy your Aiven for Valkey™ service backups to a secondary region for disaster recovery. ###### benchmark-performance Aiven for Valkey™ uses memtier_benchmark, a command-line tool by Redis, for load generation and performance evaluation of NoSQL key-value databases. - [Benchmark Aiven for Valkey™ performance](/products/valkey/howto/benchmark-performance.md): Aiven for Valkey™ uses memtier_benchmark, a command-line tool by Redis, for load generation and performance evaluation of NoSQL key-value databases. ###### change-cloud-region Move your Aiven for Valkey™ service to a different cloud provider or region. - [Change the cloud or region for your Aiven for Valkey™ service](/products/valkey/howto/change-cloud-region.md): Move your Aiven for Valkey™ service to a different cloud provider or region. ###### change-service-plan Change the service plan for your Aiven for Valkey™ service to scale resources up or down and optimize costs. - [Change the plan for your Aiven for Valkey™ service](/products/valkey/howto/change-service-plan.md): Change the service plan for your Aiven for Valkey™ service to scale resources up or down and optimize costs. ###### configure-acl-permissions Aiven for Valkey™ uses access control lists (ACLs) to manage the usage of commands and keys based on specific username and password combinations. - [Configure ACL permissions in Aiven for Valkey™](/products/valkey/howto/configure-acl-permissions.md): Aiven for Valkey™ uses access control lists (ACLs) to manage the usage of commands and keys based on specific username and password combinations. ###### configure-backups Learn how backups work for your Aiven for Valkey™ service and set the time when automatic - [Aiven for Valkey™ service backups](/products/valkey/howto/configure-backups.md): Learn how backups work for your Aiven for Valkey™ service and set the time when automatic ###### connect-go Establish a connection to the Aiven for Valkey™ service using Go. This example demonstrates how to connect to Aiven for Valkey from Go using the go-valkey/valkey library, designed to interact with the Valkey protocol. - [Connect to Aiven for Valkey™ with Go](/products/valkey/howto/connect-go.md): Establish a connection to the Aiven for Valkey™ service using Go. This example demonstrates how to connect to Aiven for Valkey from Go using the go-valkey/valkey library, designed to interact with the Valkey protocol. ###### connect-java Establish a connection to your Aiven for Valkey™ service using Java and the jedis library. - [Connect to Aiven for Valkey™ with Java](/products/valkey/howto/connect-java.md): Establish a connection to your Aiven for Valkey™ service using Java and the jedis library. ###### connect-node Connect to the Aiven for Valkey™ service using NodeJS with the ioredis library. - [Connect to Aiven for Valkey™ with NodeJS](/products/valkey/howto/connect-node.md): Connect to the Aiven for Valkey™ service using NodeJS with the ioredis library. ###### connect-php Connect to the Aiven for Valkey™ database using PHP, making use of the predis library. - [Connect to Aiven for Valkey™ with PHP](/products/valkey/howto/connect-php.md): Connect to the Aiven for Valkey™ database using PHP, making use of the predis library. ###### connect-python Connect to the Aiven for Valkey™ service using Python with the valkey-py library. valkey-py is a Python interface specifically designed for the Valkey key-value store. - [Connect to Aiven for Valkey™ with Python](/products/valkey/howto/connect-python.md): Connect to the Aiven for Valkey™ service using Python with the valkey-py library. valkey-py is a Python interface specifically designed for the Valkey key-value store. ###### connect-services Connect to the Aiven for Valkey™ service using various programming languages or tools. - [Connect to Aiven for Valkey™](/products/valkey/howto/connect-services.md): Connect to the Aiven for Valkey™ service using various programming languages or tools. ###### connect-valkey-cli Learn how to establish a connection to an Aiven for Valkey™ service using the valkey-cli. - [Connect to Aiven for Valkey™ with valkey-cli](/products/valkey/howto/connect-valkey-cli.md): Learn how to establish a connection to an Aiven for Valkey™ service using the valkey-cli. ###### controlled-upgrade-pipelines Link Aiven for Valkey™ services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. - [Controlled upgrade pipelines for your Aiven for Valkey™ service Limited availability](/products/valkey/howto/controlled-upgrade-pipelines.md): Link Aiven for Valkey™ services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. ###### create-valkey-read-replica Aiven for Valkey™ read replica enables data replication from a primary to a replica service, improving performance and increasing redundancy for high availability and disaster recovery. - [Create read replica in Aiven for Valkey™ Early availability](/products/valkey/howto/create-valkey-read-replica.md): Aiven for Valkey™ read replica enables data replication from a primary to a replica service, improving performance and increasing redundancy for high availability and disaster recovery. ###### estimate-max-number-of-connections The number of simultaneous connections for Aiven for Valkey™ depends on the total available memory on the server. - [Estimate the maximum number of connections for Aiven for Valkey™](/products/valkey/howto/estimate-max-number-of-connections.md): The number of simultaneous connections for Aiven for Valkey™ depends on the total available memory on the server. ###### fork-service Fork your Aiven for Valkey™ service to create an independent copy for testing, - [Fork your Aiven for Valkey™ service](/products/valkey/howto/fork-service.md): Fork your Aiven for Valkey™ service to create an independent copy for testing, ###### maintenance-updates Manage maintenance updates and set the maintenance window for your Aiven for Valkey™ service. - [Maintenance and updates for your Aiven for Valkey™ service](/products/valkey/howto/maintenance-updates.md): Manage maintenance updates and set the maintenance window for your Aiven for Valkey™ service. ###### manage-service-users Create and manage service users in your Aiven for Valkey™ service to control access to - [Manage Aiven for Valkey™ service users](/products/valkey/howto/manage-service-users.md): Create and manage service users in your Aiven for Valkey™ service to control access to ###### manage-ssl-connectivity Manage SSL connectivity for your Aiven for Valkey™ service by enabling secure connections and configuring stunnel for clients without SSL support. - [Manage SSL connectivity in Aiven for Valkey™](/products/valkey/howto/manage-ssl-connectivity.md): Manage SSL connectivity for your Aiven for Valkey™ service by enabling secure connections and configuring stunnel for clients without SSL support. ###### migrate-caching-valkey-to-aiven-for-valkey Migrate your Valkey™ databases to Aiven for Valkey™ using the Aiven Console migration tool. - [Migrate Valkey™ databases to Aiven for Valkey™](/products/valkey/howto/migrate-caching-valkey-to-aiven-for-valkey.md): Migrate your Valkey™ databases to Aiven for Valkey™ using the Aiven Console migration tool. ###### migrate-dragonfly-to-valkey Migrate your data from Aiven for Dragonfly® to Aiven for Valkey™ using a - [Migrate from Aiven for Dragonfly® to Aiven for Valkey™](/products/valkey/howto/migrate-dragonfly-to-valkey.md): Migrate your data from Aiven for Dragonfly® to Aiven for Valkey™ using a ###### migrate-redis-aiven-cli Move your data from a source, standalone Redis®* data store to an Aiven-managed Valkey™ service. The migration process first attempts to use the replication method, and if it fails, it switches to scan. - [Migrate from Redis®* to Aiven for Valkey™ using the CLI](/products/valkey/howto/migrate-redis-aiven-cli.md): Move your data from a source, standalone Redis®* data store to an Aiven-managed Valkey™ service. The migration process first attempts to use the replication method, and if it fails, it switches to scan. ###### migrate-redis-aiven-via-console Migrate your Redis®* databases, whether on-premise or cloud-hosted, to Aiven for Valkey™, using Aiven Console's guided wizard. - [Migrate from Redis®* to Aiven for Valkey™ using Aiven Console](/products/valkey/howto/migrate-redis-aiven-via-console.md): Migrate your Redis®* databases, whether on-premise or cloud-hosted, to Aiven for Valkey™, using Aiven Console's guided wizard. ###### power-cycle-service Power off your Aiven for Valkey™ service to release resources and save credits, power it - [Power on/off and delete your Aiven for Valkey™ service](/products/valkey/howto/power-cycle-service.md): Power off your Aiven for Valkey™ service to release resources and save credits, power it ###### prepare-for-high-load Prepare your Aiven for Valkey™ service for higher than usual traffic to avoid outages and keep performance stable. - [Prepare your Aiven for Valkey™ service for high load](/products/valkey/howto/prepare-for-high-load.md): Prepare your Aiven for Valkey™ service for higher than usual traffic to avoid outages and keep performance stable. ###### rename-service Change the name of your Aiven for Valkey™ service by forking it under a new name and - [Rename your Aiven for Valkey™ service](/products/valkey/howto/rename-service.md): Change the name of your Aiven for Valkey™ service by forking it under a new name and ###### tag-service Add key-value tags to your Aiven for Valkey™ service to organize services and track - [Tag your Aiven for Valkey™ service](/products/valkey/howto/tag-service.md): Add key-value tags to your Aiven for Valkey™ service to organize services and track ###### track-restore-progress Track the restore progress of individual nodes in your Aiven for Valkey™ service during node replacement, forking, or maintenance, using the Aiven API. - [Track restore progress for your Aiven for Valkey™ service](/products/valkey/howto/track-restore-progress.md): Track the restore progress of individual nodes in your Aiven for Valkey™ service during node replacement, forking, or maintenance, using the Aiven API. ###### valkey-version-upgrade Aiven for Valkey™ supports multiple versions of Valkey running concurrently in the platform. Choose a version that best fits your needs and upgrade your service when ready. - [Manage Aiven for Valkey™ versions](/products/valkey/howto/valkey-version-upgrade.md): Aiven for Valkey™ supports multiple versions of Valkey running concurrently in the platform. Choose a version that best fits your needs and upgrade your service when ready. ##### maintenance-lifecycle Keep your Aiven for Valkey™ service current by upgrading versions and applying maintenance - [Maintenance and lifecycle in Aiven for Valkey™](/products/valkey/maintenance-lifecycle.md): Keep your Aiven for Valkey™ service current by upgrading versions and applying maintenance ##### reference ###### advanced-params See the configuration options available for Aiven for Valkey™: - [Advanced parameters for Aiven for Valkey™](/products/valkey/reference/advanced-params.md): See the configuration options available for Aiven for Valkey™: ###### restricted-commands For optimal performance, stability, and security, Aiven for Valkey™ disables specific commands. - [Restricted commands in Aiven for Valkey™](/products/valkey/reference/restricted-commands.md): For optimal performance, stability, and security, Aiven for Valkey™ disables specific commands. ###### valkey-metrics-in-prometheus Monitor and optimize your Aiven for Valkey™ service with metrics available via Prometheus. - [Aiven for Valkey™ metrics available via Prometheus](/products/valkey/reference/valkey-metrics-in-prometheus.md): Monitor and optimize your Aiven for Valkey™ service with metrics available via Prometheus. ###### valkey-modules Aiven for Valkey™ includes pre-enabled modules that extend core Valkey functionality with additional data types and commands. - [Supported Valkey™ modules](/products/valkey/reference/valkey-modules.md): Aiven for Valkey™ includes pre-enabled modules that extend core Valkey functionality with additional data types and commands. ###### version-lifecycle Learn how Aiven manages Aiven for Valkey™ version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. - [Aiven for Valkey™ version lifecycle](/products/valkey/reference/version-lifecycle.md): Learn how Aiven manages Aiven for Valkey™ version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. ##### scaling-performance Scale your Aiven for Valkey™ service vertically or horizontally, and tune performance and - [Scaling and performance in Aiven for Valkey™](/products/valkey/scaling-performance.md): Scale your Aiven for Valkey™ service vertically or horizontally, and tune performance and ##### troubleshooting ###### troubleshoot-connection-issues Learn troubleshooting techniques for your Aiven for Valkey™ service and resolve common connection issues. - [Troubleshoot Aiven for Valkey™ connection issues](/products/valkey/troubleshooting/troubleshoot-connection-issues.md): Learn troubleshooting techniques for your Aiven for Valkey™ service and resolve common connection issues. ###### warning-overcommit_memory When starting an Aiven for Valkey™ service in the Aiven Console, - [Handle the overcommit memory warning](/products/valkey/troubleshooting/warning-overcommit_memory.md): When starting an Aiven for Valkey™ service in the Aiven Console, ### tools You can interact with the Aiven platform with various interfaces and tools that best suit your workflow. - [Aiven dev tools](/tools.md): You can interact with the Aiven platform with various interfaces and tools that best suit your workflow. #### agents Create and run AI agents on the Aiven Platform with built-in tools and MCP integrations. - [Managed Agents Limited availability](/tools/agents.md): Create and run AI agents on the Aiven Platform with built-in tools and MCP integrations. ##### chat-with-agent Interact with an agent on demand. - [Chat with an agent Limited availability](/tools/agents/chat-with-agent.md): Interact with an agent on demand. ##### create-agent Create an agent by describing a task or configuring it manually. - [Create an agent Limited availability](/tools/agents/create-agent.md): Create an agent by describing a task or configuring it manually. ##### manage-agent View and update an agent's configuration, integrations, and schedules. - [Manage an agent Limited availability](/tools/agents/manage-agent.md): View and update an agent's configuration, integrations, and schedules. ##### manage-integrations Configure the tools and integrations an agent can use, including access to Aiven resources. - [Manage integrations Limited availability](/tools/agents/manage-integrations.md): Configure the tools and integrations an agent can use, including access to Aiven resources. ##### schedule-agent Create scheduled tasks so an agent runs automatically. - [Schedule an agent Limited availability](/tools/agents/schedule-agent.md): Create scheduled tasks so an agent runs automatically. #### aiven-console In the Aiven Console you can create and manage Aiven services, update your user profile, manage settings across organizations and projects, set up billing groups, view invoices, and more. - [Aiven Console overview](/tools/aiven-console.md): In the Aiven Console you can create and manage Aiven services, update your user profile, manage settings across organizations and projects, set up billing groups, view invoices, and more. ##### howto ###### create-orgs-and-units Organizations and organizational units help you group projects and apply common settings like authentication and access. - [Create organizations and organizational units](/tools/aiven-console/howto/create-orgs-and-units.md): Organizations and organizational units help you group projects and apply common settings like authentication and access. #### api Use the Aiven API to programmatically access and automate tasks in the Aiven platform. - [Aiven API](/tools/api.md): Use the Aiven API to programmatically access and automate tasks in the Aiven platform. ##### secret-redaction Service user passwords, secret service user_config fields, and integration endpoint secrets are redacted in API responses by default. - [Secret redaction in Aiven API](/tools/api/secret-redaction.md): Service user passwords, secret service user_config fields, and integration endpoint secrets are redacted in API responses by default. #### cli The Aiven command line interface (CLI) lets you use the Aiven platform and services in a scriptable way through the API. - [Aiven CLI](/tools/cli.md): The Aiven command line interface (CLI) lets you use the Aiven platform and services in a scriptable way through the API. ##### byoc Set up and manage your custom clouds using the Aiven client and avn byoc commands. - [avn byoc](/tools/cli/byoc.md): Set up and manage your custom clouds using the Aiven client and avn byoc commands. ##### cloud The avn cloud command allows you to list the clouds available in a given project. - [avn cloud](/tools/cli/cloud.md): The avn cloud command allows you to list the clouds available in a given project. ##### credits Full list of commands for avn credits. - [avn credits](/tools/cli/credits.md): Full list of commands for avn credits. ##### events The avn events command is an audit log of things that have happened in - [avn events](/tools/cli/events.md): The avn events command is an audit log of things that have happened in ##### mirrormaker Full list of commands for avn mirrormaker. - [avn mirrormaker](/tools/cli/mirrormaker.md): Full list of commands for avn mirrormaker. ##### service-cli Full list of commands for avn service. - [avn service](/tools/cli/service-cli.md): Full list of commands for avn service. ##### service ###### acl Full list of commands for avn service acl. - [avn service acl](/tools/cli/service/acl.md): Full list of commands for avn service acl. ###### connection-info Full list of commands for avn service connection-info. - [avn service connection-info](/tools/cli/service/connection-info.md): Full list of commands for avn service connection-info. ###### connection-pool Full list of commands for - [avn service connection-pool](/tools/cli/service/connection-pool.md): Full list of commands for ###### connector Full list of commands for avn service connector. - [avn service connector](/tools/cli/service/connector.md): Full list of commands for avn service connector. ###### database Full list of commands for avn service database. - [avn service database](/tools/cli/service/database.md): Full list of commands for avn service database. ###### es-acl Full list of commands for avn service es-acl. - [avn service es-acl](/tools/cli/service/es-acl.md): Full list of commands for avn service es-acl. ###### flink Full list of commands for avn service flink. - [avn service flink](/tools/cli/service/flink.md): Full list of commands for avn service flink. ###### integration A full list of commands for avn service integration. - [avn service integration](/tools/cli/service/integration.md): A full list of commands for avn service integration. ###### kafka-acl Full list of commands for avn service kafka-acl. - [avn service kafka-acl](/tools/cli/service/kafka-acl.md): Full list of commands for avn service kafka-acl. ###### privatelink Full list of commands for - [avn service privatelink](/tools/cli/service/privatelink.md): Full list of commands for ###### quota Full list of commands for avn service quota. - [avn service quota](/tools/cli/service/quota.md): Full list of commands for avn service quota. ###### schema-registry-acl Full list of commands for avn service schema-registry-acl. - [avn service schema-registry-acl](/tools/cli/service/schema-registry-acl.md): Full list of commands for avn service schema-registry-acl. ###### service-index Full list of commands for avn service index. - [avn service index](/tools/cli/service/service-index.md): Full list of commands for avn service index. ###### tags Full list of commands for avn service tags. - [avn service tags](/tools/cli/service/tags.md): Full list of commands for avn service tags. ###### topic Full list of commands for avn service topic. - [avn service topic](/tools/cli/service/topic.md): Full list of commands for avn service topic. ###### user Full list of commands for avn service user. - [avn service user](/tools/cli/service/user.md): Full list of commands for avn service user. ##### user Manage users and personal tokens with the avn user commands. - [avn user](/tools/cli/user.md): Manage users and personal tokens with the avn user commands. ##### vpc The list of commands for project VPCs (avn vpc) and organization VPCs (avn organization vpc) - [avn vpc](/tools/cli/vpc.md): The list of commands for project VPCs (avn vpc) and organization VPCs (avn organization vpc) #### doc-diff-llms Set up automated monitoring of the Aiven documentation to track changes and get notifications when content is updated. - [Monitor Aiven documentation changes with GitHub Actions](/tools/doc-diff-llms.md): Set up automated monitoring of the Aiven documentation to track changes and get notifications when content is updated. #### kubernetes Manage Aiven infrastructure with Aiven Operator for Kubernetes® by using Custom Resource Definitions (CRD). - [Aiven Operator for Kubernetes®](/tools/kubernetes.md): Manage Aiven infrastructure with Aiven Operator for Kubernetes® by using Custom Resource Definitions (CRD). #### mcp-server Create and manage Aiven services from AI assistants. - [Aiven MCP](/tools/mcp-server.md): Create and manage Aiven services from AI assistants. #### query-optimizer Use Aiven's AI-powered SQL query optimizer for PostgreSQL® and MySQL® to get query optimization recommendations for an ad-hoc query. - [Standalone SQL query optimizer Early availability](/tools/query-optimizer.md): Use Aiven's AI-powered SQL query optimizer for PostgreSQL® and MySQL® to get query optimization recommendations for an ad-hoc query. #### terraform Use the Aiven Provider for Terraform to provision and manage your Aiven infrastructure. - [Aiven Provider for Terraform](/tools/terraform.md): Use the Aiven Provider for Terraform to provision and manage your Aiven infrastructure. ##### howto ###### use-opentofu OpenTofu is an open source infrastructure-as-code tool that you can use to configure your Aiven infrastructure. - [Use OpenTofu with Aiven Provider for Terraform](/tools/terraform/howto/use-opentofu.md): OpenTofu is an open source infrastructure-as-code tool that you can use to configure your Aiven infrastructure. --- # Full Documentation Content ### [Get started](/docs/get-started.md) [Your first steps to set up your account, for free.](/docs/get-started.md) ### [Managed services](/docs/products/services.md) [Discover our managed services and how to set them up.](/docs/products/services.md) ### [Runtime applications](/docs/products/runtime.md) [Run stateless apps within your existing Aiven project infrastructure.](/docs/products/runtime.md) ### [Aiven dev tools](/docs/tools.md) [Manage your Aiven infrastructure with the Aiven API, Terraform Provider, Kubernetes Operator, or CLI.](/docs/tools.md) ### [Integrations](/docs/platform/concepts/service-integration.md) [Explore the integrations offered by Aiven to connect your services with other systems and tools. Unlock new possibilities and improve interoperability.](/docs/platform/concepts/service-integration.md) ### [API documentation](/docs/tools/api.md) [Interact programmatically with the Aiven platform. Automate your workflows, integrate with your existing tools, and extend the functionality.](/docs/tools/api.md) --- # AI tools on Aiven Aiven provides tools that connect your AI agents to Aiven services, metadata, and data. It also builds AI capabilities directly into its managed services, from vector search to AI assistants. Use the following sections to find the right starting point for your use case. ## Choose your path[​](#choose-your-path "Direct link to Choose your path") [Managed Agents](/docs/tools/agents.md) [Create and run agents on the Aiven Platform.](/docs/tools/agents.md) [Connect AI agents and tools to Aiven](#connect-ai-agents-and-tools-to-aiven) [Connect AI assistants and agents to Aiven services, metadata, and data.](#connect-ai-agents-and-tools-to-aiven) [AI built into Aiven services](#ai-built-into-aiven-services) [Use vector search, query optimization, and SQL generation built into supported Aiven services.](#ai-built-into-aiven-services) ## Run agents on Aiven[​](#run-agents-on-aiven "Direct link to Run agents on Aiven") Create agents on the Aiven Platform, connect the tools they need, and run them on demand or on a schedule. | Tool | What you can do | Get started | | ---------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | --------------------------------------- | | Managed Agents [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) | Create agents, connect MCP integrations, and run agents on demand or on a schedule | [Managed Agents](/docs/tools/agents.md) | ## Connect AI agents and tools to Aiven[​](#connect-ai-agents-and-tools-to-aiven "Direct link to Connect AI agents and tools to Aiven") Connect AI assistants and agents to Aiven to manage services, explore metadata, and work with data using MCP servers and Skills. For example, connect an agent to Aiven for Apache Kafka® instead of running manual commands. | Tool | What you can do | Get started | | -------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------- | | Aiven MCP server | Create and manage services, plans, metrics, logs, and configuration from Cursor, Claude Code, and other MCP clients | [Set up Aiven MCP](/docs/tools/mcp-server.md) | | DataHub MCP server [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) | Give AI agents natural language search, lineage tracking, and context-aware SQL generation over your data ecosystem | [Use the DataHub MCP server](/docs/products/datahub/datahub-mcp-server.md) | | Kafka Skills | Create and configure a Kafka service, topics, ACLs, and Schema Registry from the command line | [Set up using Skills](/docs/products/kafka/howto/set-up-kafka-with-skills.md) | ## AI built into Aiven services[​](#ai-built-into-aiven-services "Direct link to AI built into Aiven services") These aren't separate products. They're capabilities built into the services you're already using. Vector search[Postgres (pgvector)](/docs/products/postgresql/concepts/pgvector.md)[ClickHouse](/docs/products/clickhouse/howto/vector-similarity-index-cache.md)[OpenSearch](/docs/products/opensearch/reference/plugins.md)[Valkey](/docs/products/valkey/reference/valkey-modules.md#valkey-search) Query optimization[Postgres](/docs/products/postgresql/howto/ai-insights.md)[MySQL](/docs/products/mysql/howto/ai-insights.md)[Standalone optimizer](/docs/tools/query-optimizer.md) SQL generation[PG Studio AI assistant](/docs/products/postgresql/howto/pg-studio/use-ai-assistant.md) Related pages * [Get started with Aiven](/docs/get-started.md) * [Aiven dev tools](/docs/tools.md) --- # Get started with Aiven Aiven provides managed open source services for streaming and databases across major cloud providers. The Aiven Platform streamlines your operations by centralizing your cloud infrastructure, security, and observability in one unified control plane. You can access the platform through the Aiven Console, Aiven API, Aiven Provider for Terraform, Aiven CLI, and Aiven Operator for Kubernetes®. ## Try the Aiven Platform for free[​](#try-the-aiven-platform-for-free "Direct link to Try the Aiven Platform for free") Sign up for a free trial to explore the Aiven Platform and try Aiven's managed services. Aiven services are also available on the AWS, Azure, and Google Cloud marketplaces. [Sign up for free](https://console.aiven.io/signup) [Marketplace signup](/docs/marketplace-setup.md) ## First steps[​](#first-steps "Direct link to First steps") Set up your organization and create your first service. ### Step 1: Set up billing[​](#step-1-set-up-billing "Direct link to Step 1: Set up billing") [Billing overview](/docs/platform/concepts/billing-and-payment.md) [Learn how billing and payments work.](/docs/platform/concepts/billing-and-payment.md) [Create a payment method](/docs/platform/howto/manage-payment-card.md) [Add a credit card to pay for your Aiven services.](/docs/platform/howto/manage-payment-card.md) [Configure your billing group](/docs/platform/howto/use-billing-groups.md) [Add your payment method and other details to the default billing group.](/docs/platform/howto/use-billing-groups.md) ### Step 2: Set up your organization[​](#step-2-set-up-your-organization "Direct link to Step 2: Set up your organization") [Organizations overview](/docs/platform/concepts/orgs-units-projects.md) [Learn about organizing your resources with organizations, units, and projects.](/docs/platform/concepts/orgs-units-projects.md) [Optional: Create organizational units](/docs/tools/aiven-console/howto/create-orgs-and-units.md) [Use units to group your Aiven projects.](/docs/tools/aiven-console/howto/create-orgs-and-units.md) [Create projects](/docs/platform/howto/manage-project.md) [Create projects in your organization or units to hold your Aiven services.](/docs/platform/howto/manage-project.md) [Organization setup with Terraform](https://github.com/aiven/terraform-provider-aiven/blob/main/examples/organization/README.md) [Follow an example to set up your organization using the Aiven Provider for Terraform.](https://github.com/aiven/terraform-provider-aiven/blob/main/examples/organization/README.md) ### Step 3: Manage organization users[​](#step-3-manage-organization-users "Direct link to Step 3: Manage organization users") Start collaborating by adding users to your organization, creating groups, and assigning them to projects. You can add users manually to your organization or create managed users through your identity provider (IdP). #### Add users to your organization[​](#add-users-to-your-organization "Direct link to Add users to your organization") * Add users manually * Create managed users [Invite users to your organization](/docs/platform/howto/manage-org-users.md) [Email your team invites to join your organization on Aiven.](/docs/platform/howto/manage-org-users.md) Make your organization users managed users by verifying a domain and configuring an identity provider. Aiven also supports automatic [user provisioning with Okta](/docs/platform/howto/saml/add-okta-idp.md) through System for Cross-domain Identity Management (SCIM). [Managed users](/docs/platform/concepts/managed-users.md) [Understand the benefits of managed users.](/docs/platform/concepts/managed-users.md) [Add a domain](/docs/platform/howto/manage-domains.md) [Add a verified domain to your organization using a DNS TXT record or HTML file.](/docs/platform/howto/manage-domains.md) [Add an identity provider](/docs/platform/howto/saml/add-identity-providers.md) [Let your users access Aiven through your preferred IdP.](/docs/platform/howto/saml/add-identity-providers.md) #### Set up user groups[​](#set-up-user-groups "Direct link to Set up user groups") Add users to groups to streamline access management to your Aiven projects and services. [Create groups](/docs/platform/howto/manage-groups.md) [Create and add users to groups.](/docs/platform/howto/manage-groups.md) [Give groups access to projects](/docs/platform/howto/manage-permissions.md) [Grant roles and permissions to a group of users to access a project and its services.](/docs/platform/howto/manage-permissions.md) [Create and assign groups with Terraform](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/get-started) [Follow an example to create a user group and give it access to a project.](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/get-started) ### Step 4: Secure your organization[​](#step-4-secure-your-organization "Direct link to Step 4: Secure your organization") [Configure an authentication policy](/docs/platform/howto/set-authentication-policies.md) [Determine how your organization users log in and use tokens.](/docs/platform/howto/set-authentication-policies.md) [Permissions and roles](/docs/platform/concepts/permissions.md) [Learn how access is controlled at the organization, project, and service level.](/docs/platform/concepts/permissions.md) [Manage organization admin](/docs/platform/concepts/permissions.md) [Control who can manage the organization, its billing, and all projects.](/docs/platform/concepts/permissions.md) [Application users](/docs/platform/concepts/application-users.md) [Learn how application users provide more secure programmatic access to the Aiven Platform.](/docs/platform/concepts/application-users.md) [Create application users](/docs/platform/howto/manage-application-users.md) [Use application users to access the Aiven API, Terraform Provider, CLI, and Kubernetes Operator.](/docs/platform/howto/manage-application-users.md) [Create a virtual private cloud](/docs/platform/howto/manage-project-vpc.md) [Connect private networks with each other without going through the public internet.](/docs/platform/howto/manage-project-vpc.md) ### Step 5: Create your first service[​](#step-5-create-your-first-service "Direct link to Step 5: Create your first service") Start deploying services in your project to stream, store, or analyze your data. #### Create a service in the Aiven Console[​](#create-a-service-in-the-aiven-console "Direct link to Create a service in the Aiven Console") [View all services](/docs/products/services.md) [Choose a service to learn more about it.](/docs/products/services.md) #### Create a service using the dev tools[​](#create-a-service-using-the-dev-tools "Direct link to Create a service using the dev tools") The get started guide for each [service](https://aiven.io/docs/products/services) includes examples for creating a service using the [Aiven Provider for Terraform](/docs/tools/terraform.md). You can try out more service and integration examples using code samples for the Aiven Terraform Provider or [Aiven Operator for Kubernetes®](/docs/tools/kubernetes.md). [Aiven Provider for Terraform examples](https://registry.terraform.io/providers/aiven/aiven/latest/docs/guides/examples) [Aiven Operator for Kubernetes® examples](https://aiven.github.io/aiven-operator/resources/project.html) Create a service using the Aiven CLI or API. [Aiven CLI](/docs/tools/cli/service-cli.md#avn-cli-service-create) [Aiven API](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) Create and manage services using AI assistants. [Aiven MCP](/docs/tools/mcp-server.md) [Create and manage services using AI assistants like Claude and Cursor.](/docs/tools/mcp-server.md) ### Step 6: Earn credits by referring others to Aiven[​](#step-6-earn-credits-by-referring-others-to-aiven "Direct link to Step 6: Earn credits by referring others to Aiven") Invite colleagues and friends to sign up to Aiven. [Both of you earn credits](/docs/platform/reference/referrals.md) when they sign up and start using services. ## Next steps[​](#next-steps "Direct link to Next steps") [Explore Aiven Console](/docs/tools/aiven-console.md) [Read about cloud security](/docs/platform/concepts/cloud-security.md) [Integrate your services](/docs/platform/concepts/service-integration.md) --- # Google Cloud Logging You can send your service logs to Google Cloud Logging to store, search, analyze, monitor, and alert on log data from your Aiven services. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * You have a Google Project ID and Log ID. More information about Google Cloud projects is available in the [Google Cloud documentation](https://cloud.google.com/resource-manager/docs/creating-managing-projects). * You have Google Cloud service account credentials in JSON format created within the same Google Cloud project where logs will be sent. See Google Cloud's documentation for [instructions on how to create and get service account credentials](https://developers.google.com/workspace/guides/create-credentials). * The service account has permission to create log entries. See the Google Cloud documentation for information on [access control with IAM](https://cloud.google.com/logging/docs/access-control). ## Set up Cloud Logging integration in Aiven Console[​](#set-up-cloud-logging-integration-in-aiven-console "Direct link to Set up Cloud Logging integration in Aiven Console") ### Step 1. Create the integration endpoint[​](#step-1-create-the-integration-endpoint "Direct link to Step 1. Create the integration endpoint") 1. Go to **Integration endpoints**. 2. Select **Google Cloud Logging**. 3. Click **Add new endpoint**. 4. Enter a name. 5. Enter the **GCP Project ID** and **Log ID** from Google Cloud. 6. Enter the **Google Service Account Credentials** in JSON format. 7. Click **Create**. warning Cross-project service account credentials do not work with **Google Service Account Credentials** for this integration. Ensure that the service account is created within the same Google Cloud project where logs are sent. ### Step 2. Add the integration endpoint to your service[​](#step-2-add-the-integration-endpoint-to-your-service "Direct link to Step 2. Add the integration endpoint to your service") 1. Go to the service to add the logs integration to. 2. On the sidebar, click **Integrations**. 3. Select **Google Cloud Logging**. 4. Choose the endpoint that you created. 5. Click **Enable**. ## Set up Cloud Logging integration using the CLI[​](#set-up-cloud-logging-integration-using-the-cli "Direct link to Set up Cloud Logging integration using the CLI") ### Step 1. Create the integration endpoint[​](#step-1-create-the-integration-endpoint-1 "Direct link to Step 1. Create the integration endpoint") ``` avn service integration-endpoint-create --project your-project-name \ -d "Google Cloud Logging" -t external_google_cloud_logging \ -c project_id=your-gcp-project-id \ -c log_id=my-aiven-service-logs \ -c service_account_credentials='{"type": "service_account"...} ``` ### Step 2. Add the integration endpoint to your service[​](#step-2-add-the-integration-endpoint-to-your-service-1 "Direct link to Step 2. Add the integration endpoint to your service") 1. Get the endpoint identifier: ``` avn service integration-endpoint-list --project your-project-name ``` 2. Use the `endpoint_id` to attach the service to the endpoint: ``` avn service integration-create --project your-project-name \ -t external_google_cloud_logging -s your-service \ -D ``` --- # Amazon CloudWatch and Aiven [Amazon CloudWatch (AWS)](https://aws.amazon.com/cloudwatch/) is an AWS monitoring service that helps to observe your applications and infrastructure resources. Aiven provides integrations that enable you to include Aiven service data into Amazon CloudWatch metrics and logs. ## CloudWatch metrics[​](#cloudwatch-metrics "Direct link to CloudWatch metrics") You can send the metrics from any or all of your Aiven services to CloudWatch. You can specify the namespace where your metrics should be send or one will be created for you. Find out [how to send your Aiven service metrics to AWS CloudWatch](/docs/integrations/cloudwatch/cloudwatch-metrics.md). ## CloudWatch logs[​](#cloudwatch-logs "Direct link to CloudWatch logs") You can send any of your Aiven service logs to your Amazon CloudWatch logs service. See [Send logs to AWS CloudWatch from Aiven Console](https://aiven.io/docs/integrations/cloudwatch/cloudwatch-logs-console) and [Send logs to AWS CloudWatch from Aiven client](https://aiven.io/docs/integrations/cloudwatch/cloudwatch-logs-console). --- # Send logs to AWS CloudWatch from Aiven client Send logs from your Aiven service to the AWS CloudWatch using the [Aiven client](/docs/tools/cli.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") This is what you'll need to send your logs from the AWS CloudWatch using the [Aiven client](/docs/tools/cli.md). * Aiven client installed. * An Aiven account with a service running. * An AWS account, and which region it is in. * An AWS Access Key and Secret Key. Generate the credentials by visiting **IAM dashboard** then click **Users**, open the **Security credentials** tab, and choose **Create access key**. Click **Download** and keep the file. important Your AWS credentials should have appropriate access rights. According to the official AWS documentation, the access rights required for the credentials are: * `logs:DescribeLogStreams` which lists the log streams for the specified log group endpoint. * `logs:CreateLogGroup` which creates a log group with the specified name endpoint. * `logs:CreateLogStream` which creates a log stream for the specified log group. * `logs:PutLogEvents` which uploads a batch of log events to the specified log stream. Find more information about [CloudWatch API](https://docs.aws.amazon.com/AmazonCloudWatchLogs/latest/APIReference/API_Operations). ## Configure the integration[​](#configure-the-integration "Direct link to Configure the integration") 1. Open the Aiven client, and log in: ``` avn user login --token ``` See also [avn user](/docs/tools/cli/user.md). 2. Collect the following information for the creation of the endpoint between your Aiven account and AWS CloudWatch. These are the placeholders you will need to replace in the code sample: | Variable | Description | | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | | `PROJECT` | Aiven project where your endpoint will be saved to. | | `LOG_GROUP_NAME` | Used to group your log streams on AWS CloudWatch. It is an optional field. If the value is not provided, it'll be generated for you. | | `AWS_REGION` | The AWS region of your account. | | `AWS_ACCESS_KEY_ID` | Your AWS access key ID. | | `AWS_SECRET_ACCESS_KEY` | Your AWS secret access key. | | `ENDPOINT_NAME` | Reference name for this log integration when linking it to other Aiven services. | 3. Create the endpoint between your Aiven account and AWS CloudWatch. ``` avn service integration-endpoint-create --project PROJECT \ -d ENDPOINT_NAME -t external_aws_cloudwatch_logs \ -c log_group_name=LOG_GROUP_NAME \ -c access_key=AWS_ACCESS_KEY\ -c secret_key=AWS_SECRET_ACCESS_KEY \ -c region=AWS_REGION ``` 4. Collect the `ENDPOINT_ID` value. You should be able to see information about your endpoint by running: ``` avn service integration-endpoint-list --project PROJECT ``` Output example ``` ENDPOINT_ID ENDPOINT_NAME ENDPOINT_TYPE ==================================== =================== =============================== 50020216-61dc-60ca-b72b-000d3cd726cb ENDPOINT_NAME external_aws_cloudwatch_logs ``` The output will provide you with the `ENDPOINT_ID` to identify your endpoint, your customized endpoint name and the endpoint type. ## Send logs from an Aiven service to AWS CloudWatch[​](#send-logs-from-an-aiven-service-to-aws-cloudwatch "Direct link to Send logs from an Aiven service to AWS CloudWatch") 1. Collect the following information for sending the service logs of an Aiven service to your CloudWatch: | Variable | Description | | -------------------- | -------------------------------------------------------------------------------- | | `PROJECT` | The Aiven project where your endpoint is saved. | | `ENDPOINT_ID` | Reference name for this log integration when linking it to other Aiven services. | | `AIVEN_SERVICE_NAME` | The Aiven service name that you want send the logs from. | 2. Send logs from the Aiven service to AWS CloudWatch by running: ``` avn service integration-create --project PROJECT\ -t external_aws_cloudwatch_logs -s AIVEN_SERVICE_NAME \ -D ENDPOINT_ID ``` --- # Send logs to AWS CloudWatch from Aiven Console Send your Aiven service logs to the AWS CloudWatch using the [Aiven Console](https://console.aiven.io). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An AWS account, and which region it is in. * An Aiven account with a service running. * An AWS Access Key and Secret Key. Generate the credentials by visiting **IAM dashboard** then click **Users**, open the **Security credentials** tab, and choose **Create access key**. Click **Download** and keep them for a later instruction. ## Configure the integration[​](#configure-the-integration "Direct link to Configure the integration") Start by configuring the link between the Aiven service and the AWS CloudWatch. This setup only needs to be done once. 1. Select **Integration endpoints** in the [Aiven Console](https://console.aiven.io/), then choose **AWS CloudWatch Logs**. 2. Select **Add new endpoint** or **Create new**. 3. Configure the settings for the new endpoint: * **Endpoint name** is how you will refer to this logs integration when linking it to other Aiven services. * Your AWS credentials: **Access Key** and **Secret Key**. * Your AWS account **Region**. * **Log Group Name** where your logs streams can be grouped in a group on AWS CloudWatch. If this field is not provided, it will be generated for you. 4. Select **Create** to save this endpoint. ## Send logs from an Aiven service to AWS CloudWatch[​](#send-logs-from-an-aiven-service-to-aws-cloudwatch "Direct link to Send logs from an Aiven service to AWS CloudWatch") 1. In your service, select **Integrations** and choose the **Amazon CloudWatch Logs** option. 2. Pick the endpoint by the **Endpoint name** you created earlier and choose **Enable**. 3. Visit your AWS account and look under **CloudWatch** and explore the **Logs** section to see the data flowing within a few minutes. Related pages Learn more about [Amazon CloudWatch and Aiven](/docs/integrations/cloudwatch.md). --- # Send metrics to Amazon CloudWatch Aiven enables you to send your service metrics to your [Amazon (AWS) CloudWatch](https://aws.amazon.com/cloudwatch/). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An AWS account, and which region it is in. * An Aiven account with a service running. * An AWS Access Key and Secret Key. tip To generate your AWS credentials: 1. Open your AWS console under the **IAM dashboard**. 2. Click **Users** and open the **Security credentials** tab. 3. Choose **Create access key**. Click **Download** and keep the file. ## Configure the integration[​](#configure-the-integration "Direct link to Configure the integration") Your first step is to create the endpoint to be used between the Aiven service and the AWS CloudWatch. This setup only needs to be done once. 1. In your project, click **Integration endpoints**. 2. Click **AWS CloudWatch Metrics** and **Add a new endpoint** or **Create new**. 3. Configure the settings for the new endpoint: * **Endpoint name** is how you will refer to the AWS CloudWatch metrics integration when linking it to an Aiven service. * **CloudWatch Namespace** where your metrics can be organized in different spaces. * Your AWS credentials: **Access Key** and **Secret Key**. * Your AWS account **Region**. 4. To save this endpoint, click **Create**. ## Send metrics from an Aiven service to AWS CloudWatch[​](#send-metrics-from-an-aiven-service-to-aws-cloudwatch "Direct link to Send metrics from an Aiven service to AWS CloudWatch") For each of the services whose metrics should be sent to your AWS CloudWatch: 1. From your service, click **Integrations** and choose the **Amazon CloudWatch Metrics** option. 2. Choose the endpoint by the **Endpoint name** you created earlier and choose **Continue**. 3. Customize which metrics to send to the CloudWatch. To do this, select a metric group or individual metric field. 4. Go to your AWS account and check the **CloudWatch** service. You can go to the **Metrics** section to see your Aiven service metrics data. It may take a few minutes until the data arrives. Related pages Learn more about [Amazon CloudWatch and Aiven](/docs/integrations/cloudwatch.md). --- # Datadog and Aiven [Datadog](https://www.datadoghq.com/) is a monitoring platform, allowing you to keep an eye on all aspects of your cloud estate. Aiven has integrations that make it easy to include an Aiven service in your Datadog dashboards. ## Datadog for metrics[​](#datadog-for-metrics "Direct link to Datadog for metrics") You can send the metrics from any or all of your Aiven services to Datadog. The integration also supports adding tags to the data, either for all metrics, or on a per-service basis. Find out [how to send your Aiven service metrics to Datadog](/docs/integrations/datadog/datadog-metrics.md). tip If you're using Aiven for Apache Kafka® you can also [customise the metrics sent to Datadog](/docs/products/kafka/howto/datadog-customised-metrics.md). tip If you're using Aiven for PostgreSQL® you can also [collect relation and function metrics with Datadog](/docs/products/postgresql/howto/monitor-relation-function-metrics-datadog.md). note Datadog integration is not available for new Startup-2 plans in Aiven for Apache Kafka®. Existing customers who already use Startup-2 with Datadog integration can continue to create Startup-2 services with Datadog integration and use their existing services without upgrading to a higher plan. Aiven recommends Business-4 or higher for Aiven for Apache Kafka® services with Datadog integration to avoid resource pressure on Startup-2 plans. If you are an existing customer and cannot create a Startup-2 service with Datadog integration in a new project, contact [Aiven Support](/docs/platform/howto/support.md). ## Datadog for logs[​](#datadog-for-logs "Direct link to Datadog for logs") The RSyslog integration can be used with any Aiven service to send the service logs to Datadog. We have a handy guide to show you [how to ship logs to Datadog from your Aiven service](/docs/integrations/datadog/datadog-logs.md) Related pages * [Send metrics to Datadog](/docs/integrations/datadog/datadog-metrics.md) * [Database monitoring with Datadog](/docs/products/postgresql/howto/monitor-database-with-datadog.md) * [Collect relation and function metrics with Datadog](/docs/products/postgresql/howto/monitor-relation-function-metrics-datadog.md) * [Ship logs to Datadog](/docs/integrations/datadog/datadog-logs.md) --- # Add custom tags Datadog integration When using the Datadog integration in the [Aiven Console](https://console.aiven.io/), Aiven automatically includes a set of standard tags in all data sent to Datadog. These tags consist of: `aiven-cloud:`, `aiven-service-type:`, `aiven-service:` and `aiven-project:`. In addition to the standard tags, you have the flexibility to include your own custom tags, which will then be appended to the data sent to Datadog. You have the option to configure these tags at both the endpoint configuration level and on a per-service integration level. ## Configure tags for the Datadog endpoint[​](#h_0e3d855c3f "Direct link to Configure tags for the Datadog endpoint") When configuring tags at the service integration level, it's important to note that these tags apply exclusively to the specific integration or connection being configured. Any tags configured at the endpoint level will be included in addition to these tags. To add tags to the endpoint: 1. Log in to [Aiven Console](https://console.aiven.io/), and select **Integration endpoints**. 2. Select **Datadog** from the list of available integration endpoints. 3. Select the **Edit endpoint** icon next to the endpoint name to which should get tagged. 4. Enter the desired tags in the provided field. You can add multiple tags by selecting the **Add** icon and optionally include descriptions for each tag. 5. Select **Save changes**. ## Configure tags for a service[​](#h_e11242c546 "Direct link to Configure tags for a service") When configuring tags at the service integration level, the tags are exclusively applied to that specific integration (connection). Additionally, any tags configured at the endpoint level will be appended to these tags. To add tags to the service integration: 1. Log in to [Aiven Console](https://console.aiven.io/), and select your service. 2. On the **Overview** page of your service, go to the **Service integrations** section and select **Manage integrations**. 3. Next to the Datadog integration listed at the top on the Integrations screen, select **Edit** from the drop-down menu (ellipsis). 4. Enter the desired tags in the provided field. You can add multiple tags by selecting the **Add** icon and optionally include descriptions for each tag. 5. Select **Save configuration** to apply the changes. --- # Send logs to Datadog Use the Aiven Rsyslog integration to send logs from your Aiven services to your external Datadog account. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A Datadog [API key](https://docs.datadoghq.com/account_management/api-app-keys/) ## Add a Syslog integration endpoint[​](#add-a-syslog-integration-endpoint "Direct link to Add a Syslog integration endpoint") This integration uses TCP endpoints to send logs to Datadog. You can use the same integration endpoint for multiple services. 1. In the project, click **Integration endpoints**. 2. Select **Syslog** > **Create new** or **Add new endpoint**. 3. Enter an **Endpoint name**. 4. Configure the **Server**: * For the US region enter `intake.logs.datadoghq.com`. * For the EU region use `tcp-intake.logs.datadoghq.eu`. 5. Configure the **Port**: * For the US region enter `10516`. * For the EU region enter `443`. 6. Enable **TLS**. 7. Set the **Format** to `custom`. 8. To configure the **Log Template**, enter: ``` DATADOG_API_KEY <%pri%>1 %timestamp:::date-rfc3339% %HOSTNAME%.AIVEN_PROJECT_NAME %app-name% - - - %msg% ``` Where: * `DATADOG_API_KEY` is your Datadog API key. * `AIVEN_PROJECT_NAME` is the name of the project your service is in. note Datadog correlates metrics and logs by hostname. The integration appends the project name to the hostname to disambiguate between services with the same name in different projects. However, without the project name no log data is lost. Don't edit the values surrounded by `%`, such as `%msg%`, as these are used in constructing the log line. For example: ``` 01234567890123456789abcdefabcdef <%pri%>1 %timestamp:::date-rfc3339% %HOSTNAME%.example-project %app-name% - - - %msg% ``` 9. Click **Create**. ## Send logs from an Aiven service to Datadog[​](#send-logs-from-an-aiven-service-to-datadog "Direct link to Send logs from an Aiven service to Datadog") 1. In the service, click **Integrations**. 2. In the **Endpoint integrations** select **Rsyslog**. 3. Select the integration endpoint you created and click **Enable**. Related pages * Read more about Datadog log collection and integrations in the [Datahub documentation](https://docs.datadoghq.com/logs/log_collection). * Learn more about [Datadog and Aiven](/docs/integrations/datadog.md). * [Monitor PgBouncer with Datadog](/docs/products/postgresql/howto/monitor-pgbouncer-with-datadog.md). * Enable [database monitoring with Datadog](/docs/products/postgresql/howto/monitor-database-with-datadog.md). --- # Send metrics to Datadog Send metrics from your Aiven service to your external Datadog account. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A Datadog account * A Datadog [API key](https://docs.datadoghq.com/account_management/api-app-keys/) * A running Aiven service ## Add a Datadog integration endpoint[​](#add-a-datadog-integration-endpoint "Direct link to Add a Datadog integration endpoint") You can use the Datadog integration endpoint for multiple services. 1. In the project, click **Integration endpoints**. 2. Select **Datadog** > **Create new** or **Add new endpoint**. 3. Enter a name for the endpoint and your Datadog **API key**. 4. Select the Datadog **Site** that you use. The site is shown in your [Datadog website URL](https://docs.datadoghq.com/getting_started/site/). 5. Optional: Add [custom tags](/docs/integrations/datadog/add-custom-tags-to-datadog.md) to send with the metrics to Datadog. 6. Click **Add endpoint**. ## Add a Datadog metrics integration to an Aiven service[​](#add-a-datadog-metrics-integration-to-an-aiven-service "Direct link to Add a Datadog metrics integration to an Aiven service") note Datadog integration is not available for new Startup-2 plans in Aiven for Apache Kafka®. Existing customers who already use Startup-2 with Datadog integration can continue to create Startup-2 services with Datadog integration and use their existing services without upgrading to a higher plan. Aiven recommends Business-4 or higher for Aiven for Apache Kafka® services with Datadog integration to avoid resource pressure on Startup-2 plans. If you are an existing customer and cannot create a Startup-2 service with Datadog integration in a new project, contact [Aiven Support](/docs/platform/howto/support.md). 1. In the service, click **Integrations**. 2. In the **Endpoint integrations** select **Datadog Metrics**. 3. Select your Datadog integration endpoint and click **Enable**. tip For Aiven for Apache Kafka® services you can also [customize the metrics sent to Datadog](/docs/products/kafka/howto/datadog-customised-metrics.md). You can see the metrics data on your Datadog dashboard. Related pages * Learn more about [Datadog and Aiven](/docs/integrations/datadog.md). * [Monitor PgBouncer with Datadog](/docs/products/postgresql/howto/monitor-pgbouncer-with-datadog.md). * Enable [database monitoring with Datadog](/docs/products/postgresql/howto/monitor-database-with-datadog.md). * [Collect relation and function metrics with Datadog](/docs/products/postgresql/howto/monitor-relation-function-metrics-datadog.md). --- # Prometheus system metrics Learn how to check what metrics are available for monitoring your service using Prometheus, and find out which of the available metrics are particularly worth monitoring and why. ## About the Prometheus integration[​](#about-the-prometheus-integration "Direct link to About the Prometheus integration") The Prometheus integration allows you to monitor your Aiven services and understand the resource usage. Using this integration, you can also track some non-service-specific metrics that may be worth monitoring. To start using Prometheus for monitoring the metrics, [configure the Prometheus integration and set up the Prometheus server](/docs/platform/howto/integrations/prometheus-metrics.md). ## Get a list of available service metrics[​](#get-a-list-of-available-service-metrics "Direct link to Get a list of available service metrics") To discover the metrics available for your services, make an HTTP `GET` request to your Prometheus service endpoint. 1. Once your Prometheus integration is configured, collect the following Prometheus service details from [Aiven Console](https://console.aiven.io/) > the **Overview** page of your service > the **Connection information** section > the **Prometheus** tab: * Prometheus URL * Username * Password 2. Make a request to get a snapshot of your metrics, replacing the placeholders in the following code with the values for your service: ``` curl -k --user USERNAME:PASSWORD PROMETHEUS_URL/metrics ``` The resulting output is a full list of the metrics available for your service. ## Metrics[​](#metrics "Direct link to Metrics") ### CPU usage[​](#cpu-usage "Direct link to CPU usage") CPU usage metrics are helpful in determining if the CPU is constantly being maxed out. For a high-level view of the CPU usage for a single CPU service, you can use the following: ``` 100 - cpu_usage_idle{cpu="cpu-total"} ``` note A process with a `nice` value larger than `0` is categorized as `cpu_usage_nice`, which is not included in `cpu_usage_user`. tip It can be useful to monitor `cpu_usage_iowait{cpu="cpu-total"}`. Its high value indicates that the service node is working on something I/O intensive. For example, if `cpu_usage_iowait{cpu="cpu-total"}` equals `40`, the CPU is idle waiting for disk or network I/O operations for 40% of time. Some important CPU-related metrics you can collect and monitor are generated from the [Telegraf plugin](https://github.com/influxdata/telegraf/tree/master/plugins/inputs/cpu). They are as follows: | Metrics | Description | | ---------------------- | --------------------------------------------------------------------------------------- | | `cpu_usage_idle` | Percentage of time the CPU is idle | | `cpu_usage_system` | Percentage of time the Kernel code is consuming the CPU | | `cpu_usage_user` | Percentage of time the CPU is in the user-space program with a `nice` value <= `0` | | `cpu_usage_nice` | Percentage of time the CPU is in the user-space program with a `nice` value > `0` | | `cpu_usage_iowait` | Percentage of time that the CPU is idle when the system has pending disk I/O operations | | `cpu_usage_steal` | Percentage of time waiting for the hypervisor to give CPU cycles to the VM | | `cpu_usage_irq` | Percentage of time the system is handling interrupts | | `cpu_usage_softirq` | Percentage of time the system is handling software interrupts | | `cpu_usage_guest` | Percentage of time the CPU is running for a guest OS | | `cpu_usage_guest_nice` | Percentage of time the CPU is running for a guest OS with a low priority | ### Disk usage[​](#disk-usage "Direct link to Disk usage") Monitoring the disk usage ensures that applications or processes don't fail due to an insufficient disk storage. tip Consider monitoring `disk_used_percent` and `disk_free`. The following table lists some important disk usage metrics you can collect and monitor: | Metrics | Description | | ------------------- | --------------------------------------------------------------------------------------------------------------------- | | `disk_free` | Free space on the service disk | | `disk_used` | Used space on the disk, for example, `1.0e+9` (8,000,000,000 bytes) | | `disk_total` | Total space on the disk (free and used) | | `disk_used_percent` | Percentage of the disk space used equal to `disk_used / disk_total * 100`, for example, `80` (80% service disk usage) | | `disk_inodes_free` | Number of index nodes available on the service disk | | `disk_inodes_used` | Number of index nodes used on the service disk | | `disk_inodes_total` | Total number of index nodes on the service disk | ### Memory usage[​](#memory-usage "Direct link to Memory usage") Metrics for monitoring the memory consumption are essential to ensure the performance of your service. tip Consider monitoring `mem_available` (in bytes) or `mem_available_percent`, as this is the estimated amount of memory available for application without swapping. ### Network usage[​](#network-usage "Direct link to Network usage") Monitoring the network provides visibility of your network and an understanding of the network utilization and traffic, allowing you to act immediately in case of network issues. tip It may be worth monitoring the number of established TCP sessions available in the `netstat_tcp_established` metric. --- # Remote syslog integration In addition to using Aiven for OpenSearch® to store the logs from your Aiven services, you can also integrate with an external monitoring system that supports the rsyslog protocol. ## Creating rsyslog integration[​](#creating-rsyslog-integration "Direct link to Creating rsyslog integration") ### Add rsyslog integration endpoint[​](#add-rsyslog-integration-endpoint "Direct link to Add rsyslog integration endpoint") Add the remote syslog you want to send the log to into the project that contains the service you want to integrate. This can be configured from the **Integration endpoints** page in the Aiven Console. ![\"Create new Syslog endpoint\" dialog](/docs/assets/images/remote-syslog-endpoint-157bc0b094dda2ac89b172da043e34d8.png) Another option is to use the [Aiven Client](https://github.com/aiven/aiven-client) . ``` avn service integration-endpoint-create --project your-project \ -d example-syslog -t rsyslog \ -c server=logs.example.com -c port=514 \ -c format=rfc5424 -c tls=true ``` When defining the remote syslog server the following parameters can be applied using the `-c` switch. Required: * `server` - DNS name or IPv4 address of the server * `port` - port to connect to * `format` - message format used by the server, this can be either `rfc3164` (the old BSD style message format), `rfc5424` (current syslog message format) or `custom` * `tls` - use TLS (as the messages are not filtered and may contain sensitive information, it is highly recommended to set this to true if the remote server supports it) Conditional (required if `format` == `custom`): * `logline` - syslog log line template for a custom format, supporting limited rsyslog style templating (using `%tag%` ). Supported tags are: `HOSTNAME`, `app-name`, `msg`, `msgid`, `pri`, `procid`, `structured-data`, `timestamp` and `timestamp:::date-rfc3339`. Optional: * `sd` - content of the structured data block of `rfc5424` message * `ca` - (PEM format) Certificate Authority to use for verifying the servers certificate (typically not needed unless the server's certificate is issued by an internal CA or it uses a self-signed certificate) * `key` - (PEM format) client key if the server requires client authentication * `cert` - (PEM format) client cert to use * `max_message_size` - Rsyslog maximum message size; default value: 8192 ### Add rsyslog integration to service[​](#add_rsyslog_integration "Direct link to Add rsyslog integration to service") This can be configured in the [Aiven Console](https://console.aiven.io/) by navigating to the **Overview** page of the target service > the **Service integrations** section and selecting **Manage integrations**. You should be able to select your previously configured Rsyslog service integration by selecting **Enable** in the modal window. Alternately, with the Aiven Client, first you need the id of the endpoint previously created. ``` avn service integration-endpoint-list --project your-project ENDPOINT_ID ENDPOINT_NAME ENDPOINT_TYPE ==================================== ============== ============= 618fb764-5832-4636-ba26-0d9857222cfd example-syslog rsyslog ``` Then you can link the service to the endpoint. ``` avn service integration-create --project your-project \ -t rsyslog -s your-service \ -D 618fb764-5832-4636-ba26-0d9857222cfd ``` ## Example configurations[​](#example-configurations "Direct link to Example configurations") Rsyslog is a standard integration so you can use it with any external system. We have collected some examples of how to integrate with popular third party platforms to get you started. note All integrations can be configured using the Aiven Console or the Aiven CLI though the examples are easier to copy and paste in the CLI form. ### Coralogix[​](#rsyslog_coralogix "Direct link to Coralogix") For a [Coralogix](https://coralogix.com/) integration, use a custom `logline` format with your key and company ID. The Syslog Endpoint to use for `server` depends on your account: * If it ends with `.com`, use `syslogserver.coralogix.com`. * If it ends with `.us`, use `syslogserver.coralogix.us`. * If it ends with `.in`, use `syslogserver.app.coralogix.in`. See the Coralogix [Rsyslog](https://coralogix.com/docs/) documentation for more information. ``` avn service integration-endpoint-create --project your-project \ -d coralogix -t rsyslog \ -c server=syslogserver.coralogix.us -c port=5142 \ -c tls=false -c format=custom \ -c logline="{\"fields\": {\"private_key\":\"YOUR_CORALOGIX_KEY\",\"company_id\":\"YOUR_COMPANY_ID\",\"app_name\":\"%app-name%\",\"subsystem_name\":\"programname\"},\"message\": {\"message\":\"%msg%\",\"program_name\":\"%programname%\",\"pri_text\":\"%pri%\",\"hostname\":\"%HOSTNAME%\"}}" ``` note `tls` needs to be set to `false`. ### Loggly®[​](#rsyslog_loggly "Direct link to Loggly®") For [Loggly](https://www.loggly.com/) integration, use a custom `logline` format with your token. ``` avn service integration-endpoint-create --project your-project \ -d loggly -t rsyslog \ -c server=logs-01.loggly.com -c port=6514 \ -c tls=true -c format=custom \ -c logline='<%pri%>%protocol-version% %timestamp:::date-rfc3339% %HOSTNAME% %app-name% %procid% %msgid% TOKEN tag="RsyslogTLS"] %msg%' ``` ### Mezmo (LogDNA)[​](#rsyslog_mezmo "Direct link to Mezmo (LogDNA)") For [Mezmo](https://www.mezmo.com/) syslog integration, use a custom `logline` format with your key. ``` avn service integration-endpoint-create --project your-project \ -d logdna -t rsyslog \ -c server=syslog-a.logdna.com -c port=6514 \ -c tls=true -c format=custom \ -c logline='<%pri%>%protocol-version% %timestamp:::date-rfc3339% %HOSTNAME% %app-name% %procid% %msgid% [logdna@48950 key="YOUR_KEY_GOES_HERE"] %msg%' ``` ### New Relic[​](#rsyslog_new_relic "Direct link to New Relic") For [New Relic](https://newrelic.com/) Syslog integration, use a custom `logline` format with your license key. This is so you can prepend your [New Relic License Key](https://docs.newrelic.com/docs/apis/intro-apis/new-relic-api-keys/#license-key) and ensure the format matches the [built-in Grok pattern](https://docs.newrelic.com/docs/logs/ui-data/built-log-parsing-rules/#syslog-rfc5424). The value to use for `server` depends on the account location: * `newrelic.syslog.eu.nr-data.net` for an EU region account (the US endpoint will not work for an EU account) * `newrelic.syslog.nr-data.net` for other regions For more information, see [Use TCP endpoint to forward logs to New Relic](https://docs.newrelic.com/docs/logs/log-api/use-tcp-endpoint-forward-logs-new-relic/) ``` avn service integration-endpoint-create --project your-project \ -d newrelic -t rsyslog \ -c server=newrelic.syslog.nr-data.net -c port=6514 \ -c tls=true -c format=custom \ -c logline='YOUR_LICENSE_KEY <%pri%>%protocol-version% %timestamp:::date-rfc3339% %hostname% %app-name% %procid% %msgid% %structured-data% %msg%' ``` ### Papertrail[​](#rsyslog_papertrail "Direct link to Papertrail") As [Papertrail](https://www.papertrail.com/) identifies the client based on the server and port you only need to copy the appropriate values from the "Log Destinations" page and use those as the values for `server` and `port` respectively. You **do not need** the ca-bundle as the Papertrail servers use certificates signed by a known CA. You also need to set the format to `rfc3164`. ``` avn service integration-endpoint-create --project your-project \ -d papertrail -t rsyslog \ -c server=logsN.papertrailapp.com -c port=XXXXX \ -c tls=true -c format=rfc3164 ``` ### Sumo Logic®[​](#rsyslog_sumo_logic "Direct link to Sumo Logic®") For [Sumo Logic](https://www.sumologic.com/), use a custom `logline` format with your collector token, use the server and port of the collector, and replace `YOUR_DEPLOYMENT` with one of `au`, `ca`, `de`, `eu`, `fed`, `in`, `jp`, `us1` or `us2`. See [Cloud Syslog Source](https://help.sumologic.com/03Send-Data/Sources/02Sources-for-Hosted-Collectors/Cloud-Syslog-Source) for more information. ``` avn service integration-endpoint-create --project your-project \ -d sumologic -t rsyslog \ -c server=syslog.collection.YOUR_DEPLOYMENT.sumologic.com -c port=6514 \ -c tls=true -c format=custom \ -c logline='<%pri%>%protocol-version% %timestamp:::date-rfc3339% %HOSTNAME% %app-name% %procid% %msgid% YOUR_TOKEN %msg%' ``` *The Loggly trademark is the exclusive property of SolarWinds Worldwide, LLC or its affiliates, is registered with the U.S. Patent and Trademark Office, and may be registered or pending registration in other countries. All other SolarWinds trademarks, service marks, and logos may be common law marks or are registered or pending registration.* --- # Log integration with Loggly Aiven supports integrating logs with a number of external monitoring systems that support rsyslog protocol, including [Loggly](https://www.loggly.com/). To integrate your service with Loggly, a new endpoint needs to be added into the project that contains the service to integrate. This can be done using through Aiven console or command line using [Aiven CLI](/docs/tools/cli.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before creating the integration, you will need to generate a Loggly **customer token**. From the Loggly dashboard: * Go to **Source Setup** tab * Select **Customer Token** tab underneath it. * Generate a new custom token (or use a previously generated one) and copy its value. ## Define the Loggly integration endpoint[​](#define-the-loggly-integration-endpoint "Direct link to Define the Loggly integration endpoint") To create a Loggly integration using the [Aiven Console](https://console.aiven.io): * Select the Project where the integration needs to be defined * Select **Integration endpoints** * Go to **Syslog** configuration * Define a new endpoint with the following parameters * **Endpoint name** - the name for the endpoint (for example `Loggly`) * **Server** - the Loggly hostname `logs-01.loggly.com` * **Port** - the Loggly port `514` * **TLS** - disabled (see below how to enable TLS with avn client) * **Format** - `rfc5424` * **Structured Data** - `TOKEN@NNNNN TAG="your-tag"` replacing * `TOKEN` needs to be replaced with your Loggly **customer token** retrieved in the prerequisite stage * `NNNNN` is Loggly Private Enterprise Number (PEN) which is `41058` (check [Loggly documentation](https://documentation.solarwinds.com/en/success_center/loggly/content/admin/streaming-syslog-without-using-files.htm) for up to date information) * `your-tag` with any arbitrary tag value wrapped in double quotes tip You can automate the creation of a Loggly integration endpoint using the [Aiven CLI dedicated command](/docs/tools/cli/service/integration.md#avn_service_integration_endpoint_create). ## Enable the Loggly integration[​](#enable-the-loggly-integration "Direct link to Enable the Loggly integration") To enable the Loggly integration for a particular Aiven service: * Open the service details in the [Aiven Console](https://console.aiven.io) * Browse to the **Service Integrations** option * Click **Manage Integrations** which will bring up a list of available integrations for your service * Select **Rsyslog** from the provided list * Click **Use integration** * Select the Loggly endpoint that you created in previous step * Click **Enable** to enable service integration. After enabling this service integration, it will be shown as active in the [Aiven Console](https://console.aiven.io), and the logs will be now integrated with Loggly. note It may take a few moments to setup the new log, and you can track the status on the **Overview** page of your service > the **Service integrations** section. Your logs should now be visible on Loggly **Search** tab. Enter the tag name your previously specified, for example `tag:your-tag`, and it will populate the dashboard with the log events from the Aiven service. tip You can automate the creation of the Loggly integration using the [Aiven CLI dedicated command](/docs/tools/cli/service/integration.md#avn_service_integration_create). --- # Send Aiven logs to Logtail [Logtail](https://betterstack.com/logs) is a logging service. You can use the Aiven [Remote syslog integration](/docs/integrations/rsyslog.md) to send your logs to Logtail. 1. Set up an Rsyslog source on Logtail. Choose **Connect source**, give your source a **Name**, and select "Rsyslog" as the **Platform**. 2. Copy the **Source token** of your new source to configure the Aiven side of the integration. 3. Create the service integration on Aiven. Choose **Integration endpoints** in the web console, click **Syslog** and choose **Add new endpoint**. 4. Configure the new endpoint: * Set an **Endpoint name** for this integration * **Server**: `in.logtail.com` * **Port**: `6514` * **Format**: `custom` * Now replace `YOUR_LOGTAIL_SOURCE_TOKEN` in the log template below with the token you copied in step 2, and paste into the **Log template** field: ``` <%pri%>%protocol-version% %timestamp:::date-rfc3339% %HOSTNAME% %app-name% %procid% %msgid% [logtail@11993 source_token="YOUR_LOGTAIL_SOURCE_TOKEN"] %msg% ``` 5. Add your new logs integration to any of your Aiven services (more information [in the Rsyslog article](/docs/integrations/rsyslog.md#add_rsyslog_integration)) 6. Check the **Live tail** page on Logtail to see the logs coming in. ## Create the Logtail service integration endpoint with Aiven client[​](#create-the-logtail-service-integration-endpoint-with-aiven-client "Direct link to Create the Logtail service integration endpoint with Aiven client") To use the CLI, use the following command to create the service integration endpoint. Replace the placeholder with your token: ``` avn service integration-endpoint-create --project your-project \ -d logtail -t rsyslog \ -c server=in.logtail.com -c port=6514 \ -c tls=true -c format=custom \ -c logline='<%pri%>%protocol-version% %timestamp:::date-rfc3339% %HOSTNAME% %app-name% %procid% %msgid% [logtail@11993 source_token="TOKEN-FROM-LOGTAIL"] %msg%' ``` This replaces steps 3 and 4 above. Related pages * [Log integration with Loggly](/docs/integrations/rsyslog/loggly.md) --- # Send logs to Elasticsearch® You can store logs from one of your Aiven services in an external Elasticsearch service. Collect these values for the connection: | Variable | Description | | ------------------------ | ---------------------------------------------------------------- | | `ELASTICSEARCH_USER` | User name to access the Elasticsearch service. | | `ELASTICSEARCH_PASSWORD` | Password to access the Elasticsearch service. | | `ELASTICSEARCH_HOST` | HTTPS service host of your external Elasticsearch service. | | `ELASTICSEARCH_PORT` | Port to use for the connection. | | `CA_CERTIFICATE` | CA certificate in PEM structure, if necessary. | | `CONNECTION_NAME` | Name of this external connection to be used with Aiven services. | ## Create external Elasticsearch integration[​](#create-external-elasticsearch-integration "Direct link to Create external Elasticsearch integration") Start by setting up an external service integration for Elasticsearch. 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. In the project, click **Integration endpoints**. 3. Select **External Elasticsearch** from the list. 4. Select **Add new endpoint**. 5. Set a preferred endpoint name, we'll call it `CONNECTION_NAME` later. 6. In the connection URL field set the connection string in a format `https://ELASTICSEARCH_USER:ELASTICSEARCH_PASSWORD@ELASTICSEARCH_HOST:ELASTICSEARCH_PORT`, using your own values for those parameters. 7. Set desired index prefix, that doesn't overlap with any of already existing indexes in your Elasticsearch service. 8. Optional: Add the body of your CA certificate in PEM format. 9. Click **Create**. ## Send service logs to Elasticsearch[​](#send-service-logs-to-elasticsearch "Direct link to Send service logs to Elasticsearch") 1. Click **Services** and open a a service. 2. On the sidebar, click **Integrations**. 3. Select **Elasticsearch Logs** from the list. 4. Select the **Endpoint name** and click **Enable**. note Logs are split per day with index name consisting of your desired index prefix and a date in a format year-month-day, for example `logs-2022-08-30`. note You can also set up the integration using Aiven CLI and the following commands: * [avn service integration-endpoint-create](/docs/tools/cli/service/integration.md#avn_service_integration_endpoint_create) * [avn service integration-endpoint-list](/docs/tools/cli/service/integration.md#avn_service_integration_endpoint_list) * [avn service integration-create](/docs/tools/cli/service/integration.md#avn_service_integration_create) Related pages [Set up a log integration with an Aiven for OpenSearch® service](/docs/products/opensearch/howto/opensearch-log-integration.md) --- # Set up marketplace subscriptions Create an AWS, Azure, or Google Cloud marketplace subscription to use as a payment method for Aiven services. To set up your marketplace subscription as a payment method, first subscribe to Aiven on the cloud provider's marketplace, then link your subscription to your Aiven organization. If you already have Aiven services, you can also [change the billing to a marketplace subscription](/docs/platform/howto/list-marketplace-payments.md). * AWS Marketplace * Azure Marketplace * Google Cloud Marketplace 1. Go to [Aiven Platform on AWS Marketplace](https://aws.amazon.com/marketplace/pp/prodview-fx7pxfq5uaxha). 2. Click **View purchase options**. 3. Click **Subscribe**. note You won't be charged. This only sets up a billing subscription between AWS and Aiven. You are charged after creating Aiven services. 1. Click **Set up your account**. This takes you to the Aiven Console to complete the process. 2. On the AWS signup page at Aiven, register or log in. 3. Choose or create an Aiven organization to link the AWS subscription to. If you have any issues linking Aiven to your AWS subscription, in the AWS web console find the Aiven subscription and click **Set up your account**. note The URL to log into your AWS subscription is . 1. Go to [Aiven Platform on the Azure Marketplace](https://azuremarketplace.microsoft.com/en-us/marketplace/apps/aivenltd1590663507662.aiven_managed_database_services?tab=Overview). 2. Click **Get it now**. 3. Select **Recurring billing**. 4. Enter the other details to [complete the purchase](https://learn.microsoft.com/en-us/marketplace/purchase-software-appsource). note You won't be charged. This only sets up a billing subscription between Azure and Aiven. You will be charged after creating Aiven services. 5. To complete the setup in Aiven Console, click **Configure account now**. 6. On the [Azure signup page at Aiven](https://console.azure.aiven.io/login), log in using your Azure console email address. If you have any issues linking Aiven to your Azure subscription, in the Azure web console find the Aiven subscription and click **Open SaaS Account on publisher's site** to complete the subscription process. note The URL to log into your Azure subscription is . 1. Go to [Aiven Platform on the Google Cloud Marketplace](https://console.cloud.google.com/marketplace/product/aiven-public/aiven). 2. Click **Subscribe**. 3. Select your billing account and click **Subscribe**. note You won't be charged. This only sets up a billing subscription between Google Cloud and Aiven. You will be charged after creating Aiven services. 1. Click **Go to product page**. 2. To complete the setup on Aiven Console, click **Manage on provider**. 3. Log in or register for Aiven. note The URL to log into your Google Cloud subscription is [console.gcp.aiven.io](https://console.gcp.aiven.io). Services you create through your subscription are billed per hour and metering is sent from Aiven to Google. You can view this in the **Billing** section in Google Cloud. The subscription is shown as **Use of Aiven** and the following labels show your usage per service: * `container_name`: The name of the Aiven project * `resource_name`: The name of Aiven service --- # Firewall configuration for service nodes Aiven nodes are built using Linux. Firewall configuration is managed using native Linux kernel-level iptables rules that limit connectivity to nodes. The iptables configuration is generated dynamically at runtime depending on service type, deployment parameters, and user preferences. Rules are updated when required, for example, when deploying multi-node clusters of services. Intra-node connections are limited to point-to-point connections to specific IP addresses. All traffic to ports that are not required for the service to function is rejected instead of dropped to avoid timeouts. Service ports that you can connect to depend on the service type and deployment type. The configuration can also affect the ports that are available: * Is the service in a public network, [dedicated VPC](/docs/platform/howto/manage-project-vpc.md), virtual cloud account, or a [Bring Your Own Cloud (BYOC)](/docs/platform/concepts/byoc.md) setup? * Have you configured IP ranges in `user_config.ip_filter`? * Have you [enabled public Internet access for services in a VPC](/docs/platform/howto/public-access-in-vpc.md)? ## Commonly opened ports[​](#commonly-opened-ports "Direct link to Commonly opened ports") Aiven services commonly assign the following ports for services when deployed without any special configuration: | Port | Description | | ----------------------------- | ------------------------------------------------------------------------------------------------------ | | 22 | Aiven management plane traffic over SSH | | 80 (proxy, not open on nodes) | Redirect HTTP web traffic to HTTPS | | 443 | Web user interface traffic- Kafka® Connect
- Flink®
- Grafana®
- OpenSearch® Dashboards | | 30287 | Aiven platform management port | | 500, 4500 (UDP) | IPsec (IKE, IPsec NAT-T) | ## Service ports[​](#service-ports "Direct link to Service ports") Aiven service ports are assigned randomly as offsets of a base port number. The base port number [is set per project](/docs/platform/howto/configure-project-base-port.md). That means that a PostgreSQL® service and a MySQL® service in the same project will have closely resembling or even overlapping port numbers. These ports are in the 10000 to 30000 range. If a base port number is not defined, the service is assigned a random port number. This is defined during runtime when the service is started. ## Cloud management[​](#cloud-management "Direct link to Cloud management") Local access to the metadata address is allowed via 169.254.169.254/32. This includes ports 123 and 52 for services like NTP and local DNS. Azure health checks, DHCP, and DNS are allowed from IP 168.63.129.16/32 using ports 67 and 53. This is an Azure-specific management address. ## Enhanced compliance environments[​](#enhanced-compliance-environments "Direct link to Enhanced compliance environments") In [Enhanced Compliance Environments (ECE)](/docs/platform/concepts/enhanced-compliance-env.md), there is additional filtering at VPC level and a SOCKS5 proxy. ECE environments have more variable configurations because we provide more flexibility for configuring these to meet your requirements. Typically, ECE nodes are accessible only over VPC connections and are not exposed to the internet. This results in layered firewalls with cloud-provider SDN firewalls and individual node-specific iptables rules. ## BYOC environments[​](#byoc-environments "Direct link to BYOC environments") With the BYOC deployment model, you deploy Aiven services under your own cloud accounts. This gives you greater control over deployment configuration, but the VM-level firewall configurations are set at deployment time according to Aiven base configurations. You can apply additional firewalls using your cloud service provider's configuration options. --- # Application users An application user is a type of user that provides programmatic access to the Aiven platform and services through the [Aiven API](/docs/tools/api.md), [CLI](/docs/tools.md), [Aiven Terraform Provider](/docs/tools/terraform.md), and [Aiven Kubernetes Operator](/docs/tools/kubernetes.md). They're intended for non-human users that need to access Aiven. You must be an [organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) to access this feature. ## Application user permissions[​](#application-user-permissions "Direct link to Application user permissions") You [create and manage application users](/docs/platform/howto/manage-application-users.md) at the organization level and you [give them access to projects and services](/docs/platform/howto/manage-permissions.md) in the same way as organization users. You cannot make application users super admin, but you can grant them the organization admin role, giving them full access to your organization, its organizational units, projects, services, billing, and other settings. Unlike organization users, application users can't log in to the Aiven Console and the authentication policies don't apply to them. ## Security best practices[​](#security-best-practices "Direct link to Security best practices") Because application users can have the same level of access to projects and services it's important to secure these accounts and their tokens to avoid abuse. The following are some suggested best practices for using Aiven application users. ### Create dedicated application users for each application[​](#create-dedicated-application-users-for-each-application "Direct link to Create dedicated application users for each application") Try to create a different application user for each tool or application. For example, if you have an application that needs to connect to services in one of your projects and you're using Aiven Terraform Provider in the same project, create two application users. Use the description field for each user to clearly indicate what it's used for. This helps you manage the lifecycle of the users and ensure the access permissions are correct for each use case. ### Restrict access to trusted networks[​](#restrict-access-to-trusted-networks "Direct link to Restrict access to trusted networks") Specify allowed IP address ranges for each token. This prevents tokens from being used outside of your trusted networks, reducing the risk of breaches. You can also specify these ranges in your organization's [authentication policy](/docs/platform/howto/set-authentication-policies.md), limiting all access to the Aiven Platform to these IP addresses, including through application tokens. ### Keep tokens secure and rotate them regularly[​](#keep-tokens-secure-and-rotate-them-regularly "Direct link to Keep tokens secure and rotate them regularly") Make sure tokens are securely stored and only accessible by people who need them. Tokens should also be routinely [revoked](/docs/platform/howto/manage-application-users.md#revoke-a-token-for-an-application-user) and replaced with new tokens. ### Delete unused users and tokens[​](#delete-unused-users-and-tokens "Direct link to Delete unused users and tokens") Regularly audit your list of application users to delete unused users. You can view a list of your organization's application users and the last time they were used in **Admin** > **Application users**. Click **Actions** > **View profile** to see a user's tokens. You can [delete unused users](/docs/platform/howto/manage-application-users.md#delete-an-application-user) and [revoke specific tokens](/docs/platform/howto/manage-application-users.md#revoke-a-token-for-an-application-user). --- # Tokens There are 3 types of tokens used to access the Aiven platform: session tokens, personal tokens, and application tokens. Session tokens are created when you log in or make an API call. These tokens are revoked when you log out of the Aiven Console or the CLI. You can [create personal tokens](/docs/platform/howto/create_authentication_token.md) to access resources instead of using your password. Application tokens are linked to [application users](/docs/platform/concepts/application-users.md). Application users and tokens are a more secure option for non-human users like external applications. You can create multiple personal or application tokens for different use cases. To keep your personal and application tokens secure: * Set a session duration to limit the impact of exposure * Refrain from letting users share tokens * Rotate your tokens regularly * Restrict usage to trusted networks by specifying an allowed IP address range * Use application users for non-human users and follow [security best practices](/docs/platform/concepts/application-users.md) for their tokens * Control access to your organization's resources with the [authentication policy](/docs/platform/howto/set-authentication-policies.md) --- # Availability zones Availability zones (AZs) are physically isolated locations (data centers) where cloud services operate. There are multiple AZs within a region, each with independent power, cooling, and network infrastructure. The choice of AZs is usually affected by the latency/proximity to customers, compliance, SLA, redundancy/data security requirements, and cost. All AZs in a region are interconnected for an easy resource replication and application partitioning. ## Cross-availability-zone data distribution[​](#cross-zone-data-distro "Direct link to Cross-availability-zone data distribution") Services can be replicated across multiple availability zones (AZs), which simplifies handling failures, decreases network latency, and enhances resource protection. This deployment model provides redundancy and failover in case an AZ goes down. Deploying services across multiple AZs enables smooth traffic transfer between AZs. It improves resiliency and reliability of workloads. ## Aiven services across availability zones[​](#aiven-services-across-availability-zones "Direct link to Aiven services across availability zones") For Aiven services, nodes automatically spread across multiple AZs. All Aiven's multi-node service plans are automatically spread among AZs of a region as long as the underlying cloud provider supports it. The virtual machines (VMs) are distributed evenly across zones to provide the best possible service availability in cases an entire AZ (which may include one or more datacenters) goes down. Cloud providers with AZs available to support your Aiven services are the following: * Amazon Web Services * Google Cloud Platform * Microsoft Azure * OVHcloud * UpCloud ## Supported availability zones[​](#supported-availability-zones "Direct link to Supported availability zones") To learn what availability zones per cloud provider and region are supported for Aiven-managed services, check the **Cloud** column in [List of available cloud regions](/docs/platform/reference/list_of_clouds.md). Example With UpCloud, the only location where Aiven can automatically balance replicas of services is `upcloud-fi-hel`. For `upcloud-fi-hel`, UpCloud provides two datacenters (`fi-hel1` and `fi-hel2`). With a two-node plan, for example, it will result in one of the servers in `fi-hel1` and the other in `fi-hel2`. note With OVHcloud, AZ support depends on OVHcloud's own regional infrastructure, so it isn't available in every region. Only some OVHcloud regions provide multiple datacenters for AZ distribution, and OVHcloud decides which regions have this support. ## Smart availability zones for Apache Kafka®[​](#smart-availability-zones-for-apache-kafka "Direct link to Smart availability zones for Apache Kafka®") On top of spreading service's nodes across the availability zones of a cloud region, Aiven automatically balances replicas of your Apache Kafka® partitions into different AZs. Since Aiven automatically rebalances the data in your Apache Kafka® cluster, your data remains fully available when a node or a whole AZ is lost. Related pages * [List of available cloud regions](/docs/platform/reference/list_of_clouds.md) * [PostgreSQL® backups](/docs/products/postgresql/concepts/pg-backups.md) * [High availability](/docs/products/postgresql/concepts/high-availability.md) * [Create and use read-only replicas](/docs/products/postgresql/howto/create-read-replica.md) * [Migrate service to another cloud or region](/docs/platform/howto/migrate-services-cloud-region.md) * [Aiven for Apache Kafka® MirrorMaker 2](/docs/products/kafka/kafka-mirrormaker.md) * [MySQL backups](/docs/products/mysql/concepts/mysql-backups.md) --- # Backup to another region for Aiven services [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Copy service backups to a secondary region for disaster recovery, and manage, fork from, or delete that secondary backup. In addition to the primary service backup, you can have a secondary backup in an alternative location. important This feature is in [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-). Contact your account team to enable it. Backup to another region (BTAR) is a disaster recovery feature that allows backup files to be copied from the service's primary backup region to an additional (secondary) region. BTAR can bolster data resilience and helps improve data protection against disasters in the primary backup region. When the primary region is down, BTAR allows forking the service from an additional copy of the backup residing in a secondary region. ## Limitations[​](#limitations "Direct link to Limitations") * The cloud provider for your additional backup region must match the cloud provider for your service and the primary backup. * Secondary backup can only be restored in the region where it was stored. For a service that has the backup to another region (BTAR) feature enabled, you can check the service backup status, change the backup region, monitor the replication lag, fork and restore using the cross-region backup, or migrate to another cloud or region. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * At least one Aiven service with BTAR enabled * Access to the [Aiven Console](https://console.aiven.io/) * [Aiven API](/docs/tools/api.md) * [Aiven CLI](/docs/tools/cli.md) ## Change a backup region[​](#change-a-backup-region "Direct link to Change a backup region") 1. Log in to the [Aiven Console](https://console.aiven.io/) and go to your project. 2. On the **Services** page, select an Aiven service on which you'd like to enable BTAR. 3. On your service page, **Backups**. 4. On the **Backups** page, click **Actions** > **Edit secondary backup location**. 5. In the **Edit secondary backup location** window, use the **Backup location** menu to select a region for your additional backup. Confirm your choice by selecting **Save**. important You can change the backup region once in 24 hours. ## Monitor a service with BTAR[​](#monitor-a-service-with-btar "Direct link to Monitor a service with BTAR") There are a few things you may want to check for your Aiven service in the context of BTAR: * What is the status of a secondary backup? * Does your service have a backup in another region? * What is the target region of the secondary backup? * What is the replication lag between data availability in the primary region and the secondary region? ### Check BTAR status[​](#check-btar-status "Direct link to Check BTAR status") To see the availability, the status, and the target region of a secondary (BTAR) backup in the [Aiven Console](https://console.aiven.io/), go to your service page > **Backups** > **Secondary backup location**. ### Determine replication lag[​](#determine-replication-lag "Direct link to Determine replication lag") Determine the target region and the replication lag for a secondary (BTAR) backup of your service, call the [ServiceBackupToAnotherRegionReport](https://api.aiven.io/doc/#tag/Service/operation/ServiceBackupToAnotherRegionReport) endpoint. Configure the call as follows: 1. Enter `YOUR-PROJECT-NAME` and `YOUR-SERVICE-NAME` into the URL. 2. Specify `DESIRED-TIME-PERIOD` depending on the time period you need the metrics for: select one of the following values for the `period` key: `hour`, `day`, `week`, `month`, or `year`. ``` curl --request POST \ --url https://api.aiven.io/v1/project/YOUR-PROJECT-NAME/service/YOUR-SERVICE-NAME/backup_to_another_region/report \ --header 'Authorization: Bearer YOUR-BEARER-TOKEN' \ --header 'content-type: application/json' \ --data '{"period":"DESIRED-TIME-PERIOD"}' ``` As output, you get metrics including replication lags at specific points in time. ## Fork and restore a service with BTAR[​](#fork-and-restore "Direct link to Fork and restore a service with BTAR") You can use the [Aiven Console](/docs/tools/aiven-console.md), [API](/docs/tools/api.md), or [CLI](/docs/tools/cli.md) to recover your service from a backup in another region. To restore your service using BTAR, create a fork of the original service in the region where the secondary backup resides. note When you **fork & restore** from the secondary backup, your new fork service is created in the cloud and region where the secondary backup is located. The fork service gets the same plan that the primary service uses. Backups of the fork service are located in the region where this new service is hosted. * Aiven Console * Aiven CLI * Aiven API 1. Open the [Aiven Console](https://console.aiven.io/) and go to your service homepage. 2. Click **Backups**. 3. On the **Backups** page, select **Fork & restore**. 4. In the **New database fork** window: 1. Set **Backup location** to either **Primary location** or **Secondary location**. 2. Set **Backup version** to one of the following: * **Latest transaction** * **Point in time**: Set it up to no earlier than the time of taking the oldest replicated base backup. 3. Specify a name for the new fork service. 4. Select **Create fork**. Run the [avn service create](/docs/tools/cli/service-cli.md#avn-cli-service-create) command with the `--service-to-fork-from` option and the `--recovery-target-time`option. Set `--recovery-target-time` to no earlier than the time of taking the oldest replicated base backup. ``` avn service create FORK_SERVICE_NAME \ --plan SERVICE_PLAN \ --project PROJECT_NAME \ --service-type SERVICE_TYPE \ --cloud SECONDARY_BACKUP_REGION \ --recovery-target-time "YYYY-MM-DDTHH:MM:SS+00:00" \ --service-to-fork-from PRIMARY_SERVICE_NAME ``` Replace the following with meaningful data: * `FORK_SERVICE_NAME` * `SERVICE_PLAN` * `PROJECT_NAME` * `SERVICE_TYPE` * `SECONDARY_BACKUP_REGION` * `PRIMARY_SERVICE_NAME` * `YYYY-MM-DDTHH:MM:SS+00:00` Use the [ServiceCreate](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) API to create a fork service. When constructing the API request, add the `user_config` object to the request body and nest the `service_to_fork_from` field and the `recovery_target_time` field inside. Set `recovery_target_time` to no earlier than the time of taking the oldest replicated base backup. ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "cloud": "SECONDARY_BACKUP_REGION", "plan": "SERVICE_PLAN", "service_name": "FORK_SERVICE_NAME", "service_type": "SERVICE_TYPE", "user_config": { "service_to_fork_from": "PRIMARY_SERVICE_NAME", "recovery_target_time": "YYYY-MM-DDTHH:MM:SS+00:00" } }' ``` Replace the following with meaningful data: * `FORK_SERVICE_NAME` * `SERVICE_PLAN` * `PROJECT_NAME` * `SERVICE_TYPE` * `SECONDARY_BACKUP_REGION` * `PRIMARY_SERVICE_NAME` * `YYYY-MM-DDTHH:MM:SS+00:00` ## Migrate a service with BTAR[​](#migrate-a-service-with-btar "Direct link to Migrate a service with BTAR") You can migrate a service with BTAR the same way you [migrate a service with a regular backup](/docs/platform/howto/migrate-services-cloud-region.md). note When you migrate your service, locations of service backups, both primary and secondary ones, do not change. ## Delete a cross-region backup[​](#delete-a-cross-region-backup "Direct link to Delete a cross-region backup") Delete an additional service backup created in a region different from your primary backup region. You can delete a cross-region backup using the [Aiven Console](/docs/tools/aiven-console.md), [API](/docs/tools/api.md), or [CLI](/docs/tools/cli.md). When you delete the additional cross-region backup, you still have the default backup located in the primary, service-hosting region. * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. From the **Services** view, select an Aiven service on which you'd like to disable BTAR. 3. On your service's page, click **Backups**. 4. On the **Backups** page, click **Actions** > **Secondary backup location**. 5. In the **Edit secondary backup location** window, select **Disable**. Your additional service backup is no longer visible on your service's **Backups** page in the **Secondary backup location** column. To remove secondary backups for your service, use the [avn service update](/docs/tools/cli/service-cli.md) command to remove all target region names from the `additional_backup_regions` array. ``` avn service update your-sevice-name \ -c additional_backup_regions=\[\] ``` To remove secondary backups for your service, update the service configuration. Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint to remove all target regions names from the `additional_backup_regions` array. ``` curl --request PUT \ --url https://api.aiven.io/v1/project/YOUR_PROJECT_NAME/service/YOUR_SERVICE_NAME \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "additional_backup_regions": [] } }' ``` Related pages * [Service backups](/docs/platform/concepts/service_backups.md) * [Track service restore progress using the API](/docs/platform/howto/restore_progress_updates.md) --- # Billing and payment Billing on the Aiven Platform is managed through billing groups. To pay for services, you assign every project to a [billing group](#billing-groups). The [costs for all services](/docs/platform/concepts/service-pricing.md) in a project are charged to the [payment method](#payment-methods) of the billing group assigned to that project. ## Billing groups[​](#billing-groups "Direct link to Billing groups") [Billing groups](/docs/platform/howto/use-billing-groups.md) store payment details in one place such as: * payment method * [billing and shipping addresses](/docs/platform/howto/manage-billing-addresses.md) * billing contact emails This lets you use the same payment information across different projects within your organization, including those in other organizational units. You receive a consolidated invoice for all projects assigned to a billing group. Use billing groups to combine costs based on categories like an organization's departments or IT environments. You create and manage billing groups at the organization level, and you [assign billing groups to projects](/docs/platform/howto/use-billing-groups.md#assign-a-billing-group-to-a-project) in the project settings. You cannot use a billing group for projects that are in other organizations. ## Payment methods[​](#payment-methods "Direct link to Payment methods") The default payment method for new customers is [credit card](/docs/platform/howto/manage-payment-card.md). All costs accrued over a calendar month are charged to the billing group's card on the first day of the following month. You can also make payments using your AWS, Google Cloud, or Azure [marketplace subscriptions](/docs/platform/howto/list-marketplace-payments.md). Alternatively, you can [request to pay by bank transfer](/docs/platform/howto/manage-bank-transfers.md). When you redeem Aiven [credits](/docs/platform/howto/credits.md), they're assigned to a billing group as a payment method. Credits are automatically used to cover charges of any project assigned to that billing group. Related pages * Create [billing groups](/docs/platform/howto/use-billing-groups.md) for your organization. * Use the [invoice API](https://api.aiven.io/doc/#tag/BillingGroup) to export cost information to business intelligence tools. --- # Bring your own cloud (BYOC) Bring your own cloud (BYOC) allows you to use your own cloud infrastructure instead of relying on the Aiven-managed infrastructure. Aiven services are usually deployed on Aiven-managed infrastructure, using Aiven-managed security protocols, and backed by Aiven-managed storage and backups. This provides a straightforward and safe approach to deploying Aiven services. However, you might need a different configuration if your business, project, or organization has specific requirements. With BYOC, your Aiven organization gets connected with your cloud provider account by creating custom clouds in your Aiven organization. ## How it works[​](#how-it-works "Direct link to How it works") A custom cloud is a secure environment within your cloud provider account to run Aiven-managed data services. By enabling BYOC, creating custom clouds, and setting up Aiven services within the custom clouds, you can manage your infrastructure on the Aiven platform while keeping your data in your own cloud. ![How BYOC works](/docs/assets/images/byoc-how-it-works-18dcb2b27dac34240f7f5e005233eaac.png) 1. [Enable BYOC](/docs/platform/howto/byoc/enable-byoc.md) in your Aiven organization by setting up a call with the Aiven sales team to share your use case and its requirements. 2. [Create a custom cloud](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in the Aiven Console or CLI by providing cloud setup details essential to generate your custom cloud infrastructure template. 3. **Integrate your cloud account with Aiven** by applying the infrastructure template for [AWS](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md#deploy-the-template), [Google Cloud](/docs/platform/howto/byoc/create-cloud/create-google-custom-cloud.md#deploy-the-template), or [Microsoft Azure](/docs/platform/howto/byoc/create-cloud/create-azure-custom-cloud.md). 4. [Deploy services](/docs/platform/howto/byoc/manage-byoc-service.md) by creating new Aiven-managed services in the custom cloud or migrating existing Aiven-managed services to the custom cloud. 5. **View Aiven-managed assets in your cloud account**: You can preview Aiven-managed services and infrastructure in your cloud account. tip Use an AI assistant connected to [Aiven MCP](/docs/tools/mcp-server.md) to create, update, and view details for services running in your custom clouds from clients such as Cursor and Claude Code. ## Why use BYOC[​](#why-use-byoc "Direct link to Why use BYOC") Consider using BYOC and custom clouds if you have specific business needs or project requirements, such as: * **Compliance**: Aiven offers managed environments for several standard compliance regulations, such as HIPAA, PCI DSS, and GDPR. However, if you have strict regulatory requirements or special compliance requirements, BYOC may be the best option for you. * **Network auditing**: If you require the visibility of all traffic within any VPC you operate in or need frequent auditing capabilities, BYOC is potentially a good fit. BYOC gives you the ability to audit network metadata but not the actual contents. * **Fine-grained network control**: BYOC only requires specific network access for Aiven (for example, service management or troubleshooting) to deploy and manage open source data services, otherwise allowing you to customize your network to meet any internal requirements or requirements of your customers. * **Cost optimization**: Depending on your cloud provider, with BYOC you can use cost savings plans, committed use discounts, or other strategies to save on compute and storage infrastructure costs related to Aiven services. ## Who is eligible for BYOC[​](#who-is-eligible-for-byoc "Direct link to Who is eligible for BYOC") The BYOC setup is a bespoke service offered on a case-by-case basis, and not all cloud providers support it yet. You're eligible for BYOC if: * Your cloud providers are: * AWS, GCP, or Microsoft Azure: [BYOC self-service](/docs/platform/howto/byoc/enable-byoc.md) * OCI: BYOC on request ([limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-)). [Contact Aiven](https://aiven.io/contact) for access. * You have a commitment deal with Aiven. * You have the [Advanced or Premium support tier](/docs/platform/howto/support.md). note See [Aiven support tiers](https://aiven.io/support-services) and [Aiven responsibility matrix](https://aiven.io/responsibility-matrix) for BYOC. Contact your account team to learn more or upgrade your support tier. ## When to use the regular Aiven deployment[​](#when-to-use-the-regular-aiven-deployment "Direct link to When to use the regular Aiven deployment") BYOC deployments are not automated, and they add additional complexity to communicating to the Aiven control plane, service management, key management, and security. In most cases, you can meet your regulatory and business requirements by utilizing a regular Aiven deployment or [Enhanced Compliance Environment](/docs/platform/concepts/enhanced-compliance-env.md). tip If you would like to understand BYOC better or are unsure which deployment model is the best fit for you, contact your account team. ## BYOC pricing and billing[​](#byoc-pricing-and-billing "Direct link to BYOC pricing and billing") Unlike Aiven's standard all-inclusive pricing, the BYOC setup has custom pricing depending on the nature of your requirements. If you enter this arrangement, you are responsible for all cloud infrastructure and network traffic charges. You receive two separate monthly invoices, one from Aiven for their managed services and another from the cloud service provider for the cloud infrastructure costs. This enables you to use any cloud commit you may have and potentially leverage enterprise discounts in certain cases. note For a cost estimate and analysis, contact your account team. ## BYOC architecture[​](#byoc-architecture "Direct link to BYOC architecture") * AWS BYOC private * AWS BYOC public * AWS BYOC ECE HIPAA * AWS BYOC ECE PCI DSS * GCP BYOC private * GCP BYOC public * Azure BYOC private * Azure BYOC public ![BYOC AWS private architecture](/docs/assets/images/byoc-aws-private-c02fe720751537117db5a9a0a607f3e1.png) In the AWS private deployment model, a Virtual Private Cloud (**BYOC VPC**) for your Aiven services is created within a particular cloud region in your remote cloud account. Aiven accesses this VPC from a static IP address and routes traffic through a proxy for additional security. To accomplish this, Aiven utilizes a bastion host (**Bastion node**) logically separated from the Aiven services you deploy. The service VMs reside in a privately addressed subnet (**Private subnet**) and are accessed by the Aiven management plane via the bastion. They are not accessible through the internet. note Although the bastion host and the service nodes reside in the VPC under your management (**BYOC VPC**), they are not accessible (for example, via SSH) to anyone outside Aiven. The bastion and workload nodes require outbound access to the internet to work properly (supporting HA signaling to the Aiven management node and RPM download from Aiven repositories). [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Object storage in your AWS cloud account is where your service's [backups](/docs/platform/concepts/byoc.md#byoc-service-backups) and [cold data](/docs/platform/howto/byoc/store-data.md) are stored using two S3 buckets. To enable this optional feature for your new BYOC clouds, [contact Aiven](https://aiven.io/contact). ![BYOC AWS public architecture](/docs/assets/images/byoc-aws-public-51ec70fa7fdee0d32bc8714c026b5fa9.png) In the AWS public deployment model, a Virtual Private Cloud (**BYOC VPC**) for your Aiven services is created within a particular cloud region in your remote cloud account. Aiven accesses this VPC through an internet gateway. Service VMs reside in a publicly addressed subnet (**Public subnet**), and Aiven services can be accessed through the public internet: the Aiven control plane connects to the nodes using the public address, and the Aiven management plane can access the service VMs directly. To restrict access to your service, you can use the [IP filter](/docs/platform/howto/restrict-access.md). [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Object storage in your AWS cloud account is where your service's [backups](/docs/platform/concepts/byoc.md#byoc-service-backups) and [cold data](/docs/platform/howto/byoc/store-data.md) are stored using two S3 buckets. To enable this optional feature for your new BYOC clouds, [contact Aiven](https://aiven.io/contact). note The service endpoint has two hostnames. The public hostname is derived from the private hostname by adding a `public-` prefix to it. The Service URI shown by the Aiven Console displays only the private hostname. ![BYOC AWS HIPAA architecture](/docs/assets/images/aws-byoc-ece-hipaa-d13d873fc8294ddb91f3b8e3a32c6e24.png) The AWS `hipaa` deployment model is for healthcare workloads that handle protected health information (PHI) under HIPAA. It builds on the AWS private model and shares its network topology: a dedicated **BYOC VPC** with workload nodes in a private subnet and a **Bastion node** that proxies traffic for the Aiven management plane. On top of that topology, it enforces the following controls: * Workload nodes have no public IPs, and services are not reachable from the public internet. * Workload subnets have no internet egress. The bastion host proxies outbound traffic, and Aiven creates no NAT gateways by default. * Amazon S3 access is restricted to buckets in your own AWS account through a VPC gateway endpoint, where Aiven stores service [backups](/docs/platform/concepts/byoc.md#byoc-service-backups) and [cold data](/docs/platform/howto/byoc/store-data.md). * Aiven tags all resources with the `hipaa` compliance model for governance and audit traceability. For the full list of controls, requirements, and limitations, see [Enhanced compliance BYOC clouds](/docs/platform/concepts/byoc-enhanced-compliance.md). ![BYOC AWS PCI DSS architecture](/docs/assets/images/aws-byoc-ece-pci-dss-c602dbf0fac2572f22ee0c725ec21944.png) The AWS `pci_dss` deployment model is for payment workloads that require cardholder data environment (CDE) isolation under PCI DSS. It builds on the AWS private model and shares its network topology: a dedicated **BYOC VPC** with workload nodes in a private subnet and a **Bastion node** that proxies traffic for the Aiven management plane. On top of that topology, it enforces the following controls: * Workload nodes have no public IPs, and services are not reachable from the public internet. * Workload subnets have no internet egress. The bastion host proxies outbound traffic, and Aiven creates no NAT gateways by default. * Amazon S3 access is restricted to buckets in your own AWS account through a VPC gateway endpoint, where Aiven stores service [backups](/docs/platform/concepts/byoc.md#byoc-service-backups) and [cold data](/docs/platform/howto/byoc/store-data.md). * Aiven tags all resources with the `pci_dss` compliance model for governance and audit traceability. For the full list of controls, requirements, and limitations, see [Enhanced compliance BYOC clouds](/docs/platform/concepts/byoc-enhanced-compliance.md). ![BYOC Google Cloud private architecture](/docs/assets/images/byoc-gcp-private-1b8a8ec8e4d2c8c1846a13c9f5851036.png) In the Google Cloud private deployment model, a Virtual Private Cloud (**BYOC VPC**) for your Aiven services is created within a particular cloud region in your remote cloud account. Within the **BYOC VPC**, there are: * **Public subnet** for the bastion node * **Private subnet** for the workload nodes (your Aiven services) Aiven accesses the **BYOC VPC** from a static IP address and routes traffic through a proxy for additional security. To accomplish this, Aiven utilizes a bastion host (**Bastion note**) logically separated from the Aiven services you deploy. The service VMs reside in a privately addressed subnet (**Private subnet**) and are accessed by the Aiven management plane via the bastion. They are not accessible through the internet. note Although the bastion host and the service nodes reside in the VPC under your management (**BYOC VPC**), they are not accessible (for example, via SSH) to anyone outside Aiven. The bastion and workload nodes require outbound access to the internet to work properly (supporting HA signaling to the Aiven management node and RPM download from Aiven repositories). ![BYOC Google Cloud public architecture](/docs/assets/images/byoc-gcp-public-356815a22fc7cfad23b5b73fca7c7d35.png) In the Google Cloud public deployment model, a Virtual Private Cloud (**Workload VPC**) for your Aiven services is created within a particular cloud region in your remote cloud account. Aiven accesses this VPC through an internet gateway. Service VMs reside in a publicly addressed subnet (**Public subnet**), and Aiven services can be accessed through the public internet: the Aiven control plane connects to the nodes using the public address, and the Aiven management plane can access the service VMs directly. To restrict access to your service, you can use the [IP filter](/docs/platform/howto/restrict-access.md). note The service endpoint has two hostnames. The public hostname is derived from the private hostname by adding a `public-` prefix to it. The Service URI shown by the Aiven Console displays only the private hostname. ![BYOC Azure private architecture](/docs/assets/images/byoc-azure-private-414195076d542b32cd0687576bfc2414.png) In the Azure standard deployment model, two separate Virtual Networks are created in your Azure subscription within a particular cloud region: * **Bastion VNet**: Contains a **Bastion subnet** with the bastion host. * **Workload VNet**: Contains a **Private subnet** with the workload nodes (your Aiven services). The two VNets are connected via **VNet peering**. Aiven accesses the **Bastion VNet** from static IP addresses and routes traffic through a proxy for additional security. To accomplish this, Aiven uses a bastion host (**Bastion node**) in the Bastion subnet, logically separated from the Aiven services you deploy. The service VMs reside in the **Private subnet** and are accessed by the Aiven management plane via the bastion. They are not accessible from the internet. **Network Security Groups (NSGs)** control inbound and outbound traffic on each VNet. Each VNet has its own **NAT gateway** for outbound internet access. **Azure Blob Storage** accounts (Premium LRS and Standard LRS) are provisioned in your Azure subscription for service backups. note Although the bastion host and the service nodes reside in the VNets under your management, they are not accessible (for example, via SSH) to anyone outside Aiven. The bastion and workload nodes require outbound access to the internet to work properly (supporting HA signaling to the Aiven management node and RPM download from Aiven repositories). ![BYOC Azure public architecture](/docs/assets/images/byoc-azure-public-148d74f01763033635b2092e170f5c4f.png) In the Azure public deployment model, a Virtual Network (**Workload VNet**) for your Aiven services is created within a particular cloud region in your Azure subscription. Service VMs reside in a publicly addressed subnet (**Public subnet**) and are assigned public IP addresses. Aiven services can be accessed through the public internet: The Aiven control plane connects to the nodes using the public address, and the Aiven management plane can access the service VMs directly. To restrict access to your service, you can use the [IP filter](/docs/platform/howto/restrict-access.md). **Network Security Groups (NSGs)** allow all public inbound TCP and UDP traffic to workload nodes. You are responsible for ensuring this configuration complies with your organization's policies and regulations. **Azure Blob Storage** accounts (Premium LRS and Standard LRS) are provisioned in your Azure subscription for service backups. note The service endpoint has two hostnames. The public hostname is derived from the private hostname by adding a `public-` prefix to it. The Service URI shown by the Aiven Console displays only the private hostname. Firewall rules are enforced on the subnet level. You can integrate your services using standard VPC peering techniques. All Aiven communication is encrypted. tip For AWS custom clouds, you can also choose a compliance deployment model (`pci_dss` or `hipaa`) to run services under PCI DSS or HIPAA requirements. See [Enhanced compliance BYOC clouds](/docs/platform/concepts/byoc-enhanced-compliance.md). ## BYOC service backups[​](#byoc-service-backups "Direct link to BYOC service backups") Depending on the BYOC service, Aiven takes regular service backups to enable forking, point in time recovery (PITR), and disaster recovery. important * All backups are encrypted using Aiven-managed keys. * You are responsible for managing object storage configuration. ## Dev tools for BYOC[​](#dev-tools-for-byoc "Direct link to Dev tools for BYOC") With BYOC, you can use any standard Aiven method (for example, `avn` [CLI client](/docs/tools/cli.md) or [Aiven Terraform Provider](/docs/tools/terraform.md)) to manage your services and generally have the same user experience as with the regular Aiven deployment model. Related pages * [Bring your own cloud networking and security](/docs/platform/howto/byoc/networking-security.md) * [Enable bring your own cloud (BYOC)](/docs/platform/howto/byoc/enable-byoc.md) * [Create a custom cloud in Aiven](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) --- # Enhanced compliance BYOC clouds Enhanced compliance clouds are [bring your own cloud (BYOC)](/docs/platform/concepts/byoc.md) custom clouds that you create in your own AWS account to run Aiven services under specific compliance requirements. important To enable this feature, contact your account team. After the **compliance deployment models** are enabled for your organization, you select a compliance deployment model when you [create an AWS custom cloud](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md). note Enhanced compliance BYOC clouds are different from [Aiven-managed enhanced compliance environments (ECE)](/docs/platform/concepts/enhanced-compliance-env.md), which run on Aiven-managed infrastructure. Enhanced compliance BYOC clouds run in your own AWS account through BYOC. ## Compliance deployment models[​](#compliance-deployment-models "Direct link to Compliance deployment models") When you create an AWS custom cloud, you choose a deployment model. In addition to the private (`standard`) and public (`standard_public`) models, two compliance models are available: * `hipaa`: For healthcare workloads that handle protected health information (PHI) under the Health Insurance Portability and Accountability Act (HIPAA). * `pci_dss`: For payment workloads that require cardholder data environment (CDE) isolation under the Payment Card Industry Data Security Standard (PCI DSS). Both models apply the same controls. They remain separate so that you can record which standard applies and so that their requirements can evolve independently. ## How enhanced compliance clouds differ[​](#how-enhanced-compliance-clouds-differ "Direct link to How enhanced compliance clouds differ") An enhanced compliance cloud builds on the private (`standard`) BYOC deployment model. It runs in a dedicated VPC with no shared infrastructure and adds the following controls: * **No public service access**: Workload nodes have no public IPs, and services are not reachable from the public internet. You access them over [VPC peering](/docs/platform/howto/vpc-peering-aws.md) or [AWS PrivateLink](/docs/platform/howto/byoc/aws-privatelink-byoc.md). Indirect access paths, such as public query APIs, are also blocked. * **Bastion-proxied egress**: Workload subnets have no internet egress. The bastion host proxies outbound traffic to the Aiven management plane and package repositories. Aiven creates no NAT gateways by default. If your compliance requirements allow it, you can permit limited egress to specific address ranges, but it is never required. * **Customer-owned object storage**: Amazon S3 access is restricted to buckets in your own AWS account through a VPC gateway endpoint. Aiven stores service [backups](/docs/platform/concepts/byoc.md#byoc-service-backups) and [cold data](/docs/platform/howto/byoc/store-data.md) there, so your data stays in your account and Aiven-owned buckets are not used. * **Resource tagging for governance**: Aiven tags all resources with their compliance model (`pci_dss` or `hipaa`) for governance and audit traceability. * **No forking or cross-cloud migration**: You cannot fork or migrate a service running in an enhanced compliance cloud to another cloud. This keeps your data in your AWS account and prevents a migrated service from depending on backup storage left behind in its original cloud. ## Requirements and limitations[​](#requirements-and-limitations "Direct link to Requirements and limitations") * Compliance deployment models must be enabled for your organization before you can use them. Contact your account team to enable them. * Compliance deployment models are available for **Amazon Web Services (AWS)** only. * Object storage in your AWS account is required. Aiven uses it for service backups and cold data, and you cannot turn it off for these models. * Plan your network connectivity before you create the cloud, because services are reachable only over VPC peering or AWS PrivateLink. * Meet the [BYOC eligibility requirements](/docs/platform/concepts/byoc.md#who-is-eligible-for-byoc): a commitment deal with Aiven and the [Advanced or Premium support tier](/docs/platform/howto/support.md). tip If you are unsure which deployment model fits your compliance needs, contact your account team. Related pages * [About bring your own cloud](/docs/platform/concepts/byoc.md) * [Create an AWS-integrated custom cloud](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md) * [Bring your own cloud networking and security](/docs/platform/howto/byoc/networking-security.md) * [Use AWS PrivateLink with BYOC](/docs/platform/howto/byoc/aws-privatelink-byoc.md) * [Store data in your BYOC object storage](/docs/platform/howto/byoc/store-data.md) * [Aiven-managed enhanced compliance environments (ECE)](/docs/platform/concepts/enhanced-compliance-env.md) --- # Track your carbon footprint The carbon footprint page shows estimates of the greenhouse gas emissions associated with service usage in your organization to help you monitor and reduce your emissions. The estimates are powered by [OxygenIT](https://www.oxygenit.io), a European climate-tech company that measures and optimizes IT and cloud emissions. OxygenIT works with hyperscalers, cloud providers, and enterprises to improve transparency and efficiency. Emissions are calculated for services in AWS, Microsoft Azure, and Google Cloud. Services in other cloud providers are not included in the estimates. Emissions are measured in metric tons of carbon dioxide-equivalent, which includes multiple greenhouse gasses. OxygenIT uses region-specific electricity data from each data center's energy mix and hardware efficiency metrics. The estimates follow recognized methodologies from frameworks like the Cloud Carbon Footprint and the Green Software Foundation. The estimates are not direct measurements of actual emissions. Use these estimates for internal monitoring to track trends and identify optimization opportunities, not for external assurance or regulatory reporting. ## Data privacy[​](#data-privacy "Direct link to Data privacy") All computations are done within Aiven’s environment using anonymized usage data such as CPU, storage, network, and memory for each of your active services. No data is transferred to OxygenIT. Aiven processes all data in compliance with GDPR and contractual data protection terms between Aiven and OxygenIT. ## View carbon emissions data[​](#view-carbon-emissions-data "Direct link to View carbon emissions data") You have the [`organization:sustainability:read` permission](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) to access this feature. 1. In your organization, click **Admin**. 2. Click **Carbon footprint**. In the **Carbon emissions explorer**, you can filter the data by: * **Time range**: View emissions for a predefined range or choose a custom date range. * **Billing groups**: View emissions for all services in the projects assigned to a specific billing group. * **Projects**: View emissions for all services in a specific project. ## Reduce your carbon footprint[​](#reduce-your-carbon-footprint "Direct link to Reduce your carbon footprint") Emissions estimates are based on the actual use of your Aiven services and the energy mix in each cloud region. Compute time, storage, network traffic, and data center efficiency affect energy consumption. Changes in workloads or cloud regions impact your footprint. Electricity from high-emission sources like coal and extended service use increase your carbon footprint. To reduce your carbon footprint, you can: * Optimize resource usage by scaling down unused services. * Choose greener regions powered by renewable energy. * Use tiered storage and more aggressive caching. * Use disk autoscaler to increase storage only when needed. --- # Cloud security Learn about Aiven's access control, encryption, network security, data privacy and operator access. ## Cloud provider accounts[​](#cloud-provider-accounts "Direct link to Cloud provider accounts") Aiven services are hosted under cloud provider accounts controlled by Aiven. These accounts are managed only by Aiven operations personnel, and customers cannot directly access the Aiven cloud provider account resources. ## Virtual machines[​](#virtual-machines "Direct link to Virtual machines") Each Aiven service consists of one or more virtual machines, which are automatically launched to the target cloud region chosen by the customer. In cloud regions that have multiple Availability Zones (or a similar mechanism), virtual machines are distributed evenly across the zones to provide the best possible service in cases when an entire availability zone (which can include one or more datacenters) becomes unavailable. Service-providing virtual machines (VMs) are dedicated to a single customer. This means there is no multi-tenancy on a VM basis, and the customer data never leaves the machine, except when uploaded to the offsite backup location. Virtual machines are not reused and are terminated and wiped during service upgrade or termination. ## Data encryption[​](#data-encryption "Direct link to Data encryption") Aiven at-rest data encryption covers both active service instances as well as service backups in cloud object storage. Service instances and the underlying VMs use full volume encryption using LUKS with a randomly generated ephemeral key per each instance and each volume. The key is never re-used and will be trashed at the destruction of the instance, so there's a natural key rotation with roll-forward upgrades. We use the LUKS2 default mode aes-xts-plain64:sha256 with a 512-bit key. Backups are encrypted with a randomly generated key per file. These keys are in turn encrypted with RSA key-encryption key-pair and stored in the header section of each backup segment. The file encryption is performed with AES-256 in CTR mode with HMAC-SHA256 for integrity protection. The RSA key-pair is randomly generated for each service. The key lengths are 256-bit for block encryption, 512-bit for the integrity protection and 3072-bits for the RSA key. Aiven-encrypted backup files are stored in the object storage in the same region where the service virtual machines are located. For cloud providers that do not provide object storage services like DigitalOcean or UpCloud, the nearest region from another cloud provider that provides an object storage is used. ## Networking security[​](#networking-security "Direct link to Networking security") Aiven enforces TLS encrypted connections for customer access to services by default, with two exceptions: * Aiven for MySQL® doesn't enforce TLS at all, so it accepts unencrypted connections unless you require TLS for your own database users with `ALTER USER ... REQUIRE SSL`. * Aiven for Valkey™ lets you [turn off TLS](/docs/products/valkey/howto/manage-ssl-connectivity.md#allow-plain-text-connections). Communication between virtual machines within Aiven is secured with either TLS or IPsec. There are no unencrypted plaintext connections. Virtual machine network interfaces are protected by a dynamically configured iptables-based firewall that only allows connections from specific addresses both from the internal network (other VMs in the same service) or external public network (customer client connections). The allowed source IP addresses for establishing connections are user controlled on a per-service basis. ## Networking with VPC peering[​](#networking-with-vpc-peering "Direct link to Networking with VPC peering") When using VPC peering, no public internet based access is provided to the services. Service addresses are published in public DNS, but they can only be connected to from the customer's peered VPC using private network addresses. The services providing virtual machines are still contained under Aiven cloud provider accounts. Aiven VPCs are created per project and not shared between customers, organizations, organizational units, or projects. When services are deployed in a VPC, the DNS names resolve to their Private IP. Some services allow both public and private access to be enabled while inside a VPC and see [Enable public access in a VPC](/docs/platform/howto/public-access-in-vpc.md). Connections to Aiven services are always established from your VPC to the peered Aiven VPC. So, you can use VPC firewall rules to prevent any ingress connections from the Aiven side to your network. ## Operator access[​](#operator-access "Direct link to Operator access") Normally all the resources required for providing an Aiven service are automatically created, maintained and terminated by the Aiven infrastructure and there is no manual Aiven operator intervention required. However, the Aiven Operations Team has the capability to securely login to the service Virtual Machines for troubleshooting purposes. These accesses are audit logged. No customer access to the virtual machine level is provided. ## Customer data privacy[​](#customer-data-privacy "Direct link to Customer data privacy") Customer data privacy is of utmost importance to Aiven, and is covered by internal security and customer privacy policies as well as strict EU regulations. Aiven has partnered with GitHub in scanning for service token leaks and will inform customers of tokens being made public through an email notification. For more details, check [Improving security: Aiven and GitHub's secret scanning partnership](https://aiven.io/blog/aiven-and-github-secret-scanning-partnership). Aiven operators never access customer data, unless explicitly requested to do so by the customer in order to troubleshoot a technical issue. Aiven operations team has mandatory recurring training regarding the applicable policies. ## Periodic security evaluation[​](#periodic-security-evaluation "Direct link to Periodic security evaluation") Aiven services are periodically assessed, and penetration tested for any security issues by an independent professional cyber security vendor. ## Software bill of materials (SBOM)[​](#software-bill-of-materials-sbom "Direct link to Software bill of materials (SBOM)") The SBOM is a list of all packages that are being used by Aiven in the services we provide. The list details packages installed in the VM operating system, as well as packages installed by the service itself. SBOM reports are being widely adopted and may eventually be required for compliance or security assessments. We provide these reports as a file download via the [Aiven CLI](/docs/tools/cli.md), in CSV or SPDX format. To get the SBOM report download link for a project, run: ``` avn project generate-sbom --project PROJECT_NAME --output csv ``` SBOM reports are only available when all services within the project have the latest maintenance patches applied. ## Time synchronization[​](#time-synchronization "Direct link to Time synchronization") All Aiven backend and customer services are configured to use trusted NTP (Network Time Protocol) servers of the respective cloud provider where each service is deployed. Related pages * [Manage SSL connectivity in Aiven for Valkey™](/docs/products/valkey/howto/manage-ssl-connectivity.md) * [Advanced parameters for Aiven for Valkey™](/docs/products/valkey/reference/advanced-params.md) * [Manage Aiven for MySQL® service users](/docs/products/mysql/howto/manage-service-users.md) * [Roles and permissions](/docs/platform/concepts/permissions.md) --- # Disaster recovery testing Aiven provides disaster recovery testing services to help you plan for data center or service outages. An Aiven specialist simulates an issue with the virtual machines for a test service in your environment. For example, they can test failovers when a primary instance of your Aiven service fails and recovery times when there's a complete outage. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * The [Premium support tier](/docs/platform/howto/support.md) * At least 7 working days' notice * A service created specifically for disaster recovery testing. This cannot be a service that's used in production. ## Request a disaster recovery test[​](#request-a-disaster-recovery-test "Direct link to Request a disaster recovery test") To request a test, contact your account team with the following information: * A date and time, including time zone, for the test that's within your account team's normal business hours * The virtual machine and availability zone to target --- # Discovered organizations The discovered organizations feature lets organization admin identify other Aiven organizations that [managed users](/docs/platform/concepts/managed-users.md) have joined using the same email address associated with one of your verified domains. This visibility helps you monitor external memberships and maintain organizational security. It provides insight into the organizations managed users are part of as well as any organizations they independently created before becoming managed users. The discovered organizations page lists the organizations and the managed users that joined them. Managed users who are super admin of another organization are also identified, so you can contact them to remove or modify the organization to comply with your internal security policies. ## View discovered organizations[​](#view-discovered-organizations "Direct link to View discovered organizations") 1. Click **Admin**. 2. Click **Discovered organizations**. --- # Enhanced compliance environments (ECE) Aiven collects, manages, and operates on sensitive data that is protected by privacy and compliance rules and regulations. Aiven meets the needs of its customers by providing specialized enhanced compliance environments (ECE) that comply with many of the most common compliance requirements. The vendors that assist with this collection, management, and operation are subject to the same rules and regulations. An enhanced compliance environment runs on Aiven managed infrastructure with the additional compliance requirement that no ECE VPC is shared and the managed environment is logically separated from the standard Aiven deployment environment. This decreases the blast radius of the environment to prevent inadvertent data sharing. Users of an ECE **must** encrypt all regulated data before reaching an Aiven service. As part of the increased compliance of the environment, enhanced logging is enabled for - `stderr`, `stdout`, and `stdin`. ## Eligibility[​](#eligibility "Direct link to Eligibility") Enhanced compliance environments, although similar to standard environments, involve added setup and maintenance complexity. The following are requirements to utilize an ECE: * You use Amazon Web Services (AWS), Google Cloud, or Microsoft Azure (excluding Azure Germany). * You have a commitment deal with Aiven. * You have the [Advanced or Premium support tier](/docs/platform/howto/support.md). ## Cost of ECE[​](#cost-of-ece "Direct link to Cost of ECE") The cost of an enhanced compliance environment, beyond the initial eligibility requirements, is exactly the same as a standard Aiven deployment. All service costs are the same and all egress networking charges are still covered in an ECE. Customers can still take advantage of any applicable marketplaces and spend down their commitment accordingly. ## Similarities to a standard environment[​](#similarities-to-a-standard-environment "Direct link to Similarities to a standard environment") In many ways, an ECE is the same as a standard Aiven deployment. Aiven's tooling, such as [Aiven CLI](/docs/tools/cli.md) and [Aiven Terraform Provider](/docs/tools/terraform.md), interact with ECEs seamlessly, you will still be able to take advantage of Aiven's service integrations, and you can access the environment through VPC peering or Privatelink (on AWS or Azure). However, there are some key differences from standard environments as well: * No internet access. Internet access cannot even be allow-listed as would be the case in standard private VPCs * VPCs cannot be automatically provisioned and must be manually configured by our networking team. * VPC peering or Privatelink connections must be manually approved on the Aiven end With these differences in mind, Aiven requires the following to provision an Enhanced Compliance Environment: * A CIDR block for all region/environment combinations. For example, if you have a development, QA and production environment and operate in 3 regions in each of those, we will need 9 CIDR blocks. The necessary peering information to enable the peer from our end. This differs between clouds: | Cloud name | Required peering information | | ---------- | ---------------------------------------------------------- | | AWS | - AWS account ID
- VPC ID | | GCP | - GCP Project ID
- VPC Network Name | | Azure | - Azure Tenant ID
- Azure App ID
- Azure VNet ID | ## Compliance[​](#compliance "Direct link to Compliance") Although not exhaustive, Aiven's ECE is designed to support compliance standards like the Health Insurance Portability and Accountability Act (HIPAA) and the Payment Card Industry Data Security Standard (PCI DSS). Support for these standards is not available in the standard environments. If you have compliance requirements beyond these standards, contact the [sales team](https://aiven.io/contact) so we can better understand your specific needs. Additionally, we offer an alternative deployment option that runs in your own AWS account. See [Bring Your Own Cloud (BYOC)](/docs/platform/concepts/byoc.md) and [Enhanced compliance BYOC clouds](/docs/platform/concepts/byoc-enhanced-compliance.md) for PCI DSS or HIPAA workloads. ## Migrate to an ECE[​](#migrate-to-an-ece "Direct link to Migrate to an ECE") Migrations to Aiven are a standard procedure, but migrating to an ECE can add complexity. To migrate a new service to an ECE, contact [the sales team](https://aiven.io/contact) to request help. To migrate an existing Aiven service to an ECE, the Aiven network team creates a VPN tunnel between your non-compliant VPC and your ECE VPC to enable a one-click migration. The migration is then performed in an automated and zero-downtime fashion. Once the migration is complete, the VPN tunnel is removed. --- # Service maintenance, updates and upgrades Aiven applies maintenance updates automatically during a maintenance window that you choose for each service. ## Maintenance window[​](#maintenance-window "Direct link to Maintenance window") The maintenance window is the time period when Aiven can automatically apply maintenance updates to a service. When an update becomes available, Aiven schedules it for the next available maintenance window for each service. The update runs in the first window after it becomes available, and can begin any time after the start time. For example, if a service has a maintenance window of Monday 12:00 UTC, and an update becomes available on Tuesday, the update will be applied on the following Monday. During maintenance, Aiven might restart or replace service nodes. This can cause brief connection interruptions, but services are designed to minimize downtime. Aiven performs maintenance in a rolling-forward style, creating new nodes alongside existing ones and retiring the old nodes after the upgrade completes. Major service upgrades are triggered manually. A manually triggered upgrade starts immediately, regardless of the maintenance window. important You cannot control the order in which services are updated. Each service updates according to its own configured maintenance window, and there is no guaranteed way to control the update sequence. Manual updates and maintenance window adjustments only help for non-critical updates. ## Set the maintenance window[​](#set-the-maintenance-window "Direct link to Set the maintenance window") To set the maintenance window for your service: * Console * Terraform 1. In the Aiven Console, open your service. 2. In the **Maintenance** section, click **Actions** > **Change maintenance window**. 3. Set the day and time. 4. Click **Save changes**. Use the `maintenance_window_dow` and `maintenance_window_time` attributes in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Control maintenance updates with upgrade pipelines](/docs/platform/howto/controlled-upgrade.md) * [Static IP addresses](/docs/platform/concepts/static-ips.md) --- # Managed users The managed users feature lets you centrally manage your organization's users and helps you to secure your organization in Aiven. With managed users, you can: * Control how users log in with [authentication policies](/docs/platform/howto/set-authentication-policies.md), not just how they access the organization * Have visibility of all users in your domain even if they weren't added to the Aiven organization * Set their state, including deactivating and deleting user accounts Managed users are also restricted from making changes to their profiles and creating new organizations. [Organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) can create organizations. ## Make organizations user managed users[​](#make-organizations-user-managed-users "Direct link to Make organizations user managed users") To make your users managed users, [verify a domain](/docs/platform/howto/manage-domains.md). Users in an organization with a verified domain automatically become managed users. ## View all managed users in an organization[​](#view-all-managed-users-in-an-organization "Direct link to View all managed users in an organization") 1. Click **Admin** > **Users**. ## Deactivate a managed user[​](#deactivate-a-managed-user "Direct link to Deactivate a managed user") 1. Click **Admin**. 2. Select **Users**. 3. Find the user and click **Actions** > **Deactivate**. You can follow the same process to reactivate the user. --- # Organizations, units, and projects The Aiven Platform uses organizations, organizational units, and projects to efficiently and securely organize your services and manage access. There are three levels in this hierarchy: * **Organization**: Contains all your projects and services. It's recommended to have one Aiven organization. * **Organizational units**: Added to the organization, units give you greater flexibility to organize your infrastructure based on your specific use cases. For example, you can split production and testing workloads into different organizational units. * **Projects**: Created in the organization or organizational units to group your services together. ## Organizations[​](#organizations "Direct link to Organizations") When you sign up to Aiven, an organization is created for you. You can use your organization to create a hierarchical structure that fits your needs. ![Hierarchy showing two organizational units, each with two projects, nested within one organization.](/docs/assets/images/organizations-hierarchy-52524611c731c8b24097d24a17344071.png) Organizations also let you centrally manage settings like: * [Billing information](/docs/platform/concepts/billing-and-payment.md): Managed only at the organization level, you can use billing groups across all projects in the organization and its units. You can't share billing information between organizations. * [Users](/docs/platform/concepts/user-access-management.md) and [groups](/docs/platform/howto/manage-groups.md): Managed at the organization level. * [Permissions and roles](/docs/platform/concepts/permissions.md) Grant users and groups access to resources at the organization, unit, and project level. * [Domains](/docs/platform/howto/manage-domains.md) and [identity providers (IdPs)](/docs/platform/howto/saml/add-identity-providers.md): Only available at the organization level, verified domains and IdPs provide greater control and security for your organization's resources. * [Authentication policy](/docs/platform/howto/set-authentication-policies.md): Managed at the organization level, letting you control how users access your organization on the Aiven Platform. * [Support tiers](/docs/platform/howto/support.md): Specific to a single organization and applied to all units, projects, and services within that organization. They cannot be shared between organizations. * Access control lists (ACLs): Available on the organization, organizational unit, and project level. * ACLs for service plans are inherited, meaning all projects within an organization or organizational unit have the same service plan. ## Organizational units[​](#organizational-units "Direct link to Organizational units") Organizational units are collections of projects. Customers often use these to group projects based on things like: * Departments in their company like finance, marketing, and engineering. * Environments such as development, testing, and production. You can create as many units as you need in your organization, but you cannot nest units within other units. You can grant access to units and their resources by assigning [roles and permissions](/docs/platform/concepts/permissions.md) at the organizational unit level. You cannot configure things like verified domains or billing at the unit level. ## Projects[​](#projects "Direct link to Projects") Projects are collections of services. You can [create projects](/docs/platform/howto/manage-project.md) in an organization or in organizational units. Projects help you group your services based on your organization's structure or processes. You can grant access to projects and their resources using project-level [roles and permissions](/docs/platform/concepts/permissions.md). They also let you apply uniform network security settings across all services within the project. The following are some examples of how customers organize their services: * Single project: One project containing services that are distinguished by their names. For example, services have names based on the type of environment: `demo_pg_project.postgres-prod` and `demo_pg_project.postgres-staging`. * Environment-based projects: Each project represents a deployment environment, for example: `dev`, `qa`, and `production`. This can make it easier to apply uniform user permissions, such as developer access to production infrastructure. * Project-based projects: Each project contains all the services for an internal project, with naming that highlights the relevant environment. For example: `customer-success-prod` and `business-analytics-test`. ## Best practices for organizations[​](#best-practices-for-organizations "Direct link to Best practices for organizations") ### Small organizations[​](#small-organizations "Direct link to Small organizations") For smaller organizations that have a limited number of projects and services, it's recommended to consolidate all your projects within one organization. This makes it easier for your teams to navigate between projects and services. Good naming conventions also help with finding projects and services. For example, you can include the environment type like `dev` or `prod` at the beginning of project names. Use project-level permissions to grant access to only those users who need access to those services. ### Medium-sized organizations[​](#medium-sized-organizations "Direct link to Medium-sized organizations") For more complex cases, take advantage of the organizational units to group related projects. You can, for example, group projects into units that correspond to your internal departments. Alternatively, you can group them by categories like testing, staging, and production environments. Create user groups and assign permissions to the groups at the unit or project level. ### Large organizations[​](#large-organizations "Direct link to Large organizations") Keep all projects in organizational units instead of the organization. Use clear naming conventions for the units and projects. Add all users to groups that represent similar roles and, therefore, similar access needs. Assign permissions and roles at the project level where possible and the unit level where necessary. Restrict the number of organization admin and users with organization-level roles and permissions. Use the granular billing permissions to give your finance team access to invoices without the ability to make changes to projects or services. Add your domain to your organization and configure other security settings like single sign-on and the authentication policy for your organization. For complex infrastructure, consider using the [Aiven Provider for Terraform](/docs/tools/terraform.md) to manage your organization and its resources. --- # Out of memory conditions Understand how the Linux out-of-memory killer works, why it can affect a service, and how to avoid triggering it. Many processes request more memory from the kernel than they will ever use or need. In these cases, the kernel overallocates memory. This allows it to satisfy multiple processes requesting more memory than is available, which is not used or is freed by the time any other process actually needs it. However, if enough processes start using all their allocated memory simultaneously there may not be enough physical memory available and an `Out Of Memory` (`OOM`) condition occurs. warning This situation is critical and must be resolved immediately. The solution that the Linux kernel employs is to invoke the `Out of Memory Killer` (or `OOM Killer`). This reviews all running processes and kills one or more of them to free up system memory and keep the system running. The `OOM Killer` selects process to kill based on an `oom_score`; a calculation that balances how much memory the process is using with how long the process has been running. Processes that have been running for a long time are less likely to be killed. Subprocesses are summed with parent processes in terms of memory usage, so a process which forks many subprocesses, but itself does not use a lot of memory, may still be killed. In most instances, the hosted data service, or a child process, will have the highest memory footprint and be a prime candidate for termination when the OOM Killer inspects the running processes. Aiven's cloud data platform leverages kernel namespaces or containers to isolate processes from each other. Isolation has several benefits, including: * A smaller footprint for security‑related concerns * A smaller blast radius for failure * Greater control of system resources Left unchecked, the `OOM Killer` may opt to kill the primary service. This is undesirable as unclean termination of the primary service can lead to data loss, inconsistency, or corrupted backups. Further, if Aiven's management platform detects that the primary service is unavailable for , the service will be marked as down and a failover will occur. To mitigate this scenario, namespaces are used, some with additional memory limits, in combination with an `oom_score_adjust` on the primary process, to coax the `OOM Killer` into selection of less critical processes. This will still result in a service restart, but in a more controlled process, where the database is shut down, rather than killed; exposure to data loss is limited and recovery is faster when the service restarts, often avoiding failover. warning Out of Memory conditions can still lead to unexpected behavior, including data unavailable or data loss conditions. ## Avoid running low on memory[​](#avoid-running-low-on-memory "Direct link to Avoid running low on memory") The OOM killer only runs when the system is critically low on memory. To prevent it from running, either reduce your memory usage or increase the available memory. For most databases, the service memory footprint can often be reduced by: * Reducing concurrency or implementing connection pooling * Tuning queries to limit result sets * Tuning indexes for query load * Dropping unused objects from storage In cases where the working set no longer fits into memory, consider scaling your service. Related pages * [Service memory limits](/docs/platform/concepts/service-memory-limits.md) * [Change a service plan](/docs/platform/howto/scale-services.md) * [Prepare services for high load](/docs/platform/howto/prepare-for-high-load.md) --- # Roles and permissions To give users access to projects and services in your organizations, you grant them permissions and roles: * **Permissions**: Actions that a principal can perform on a resource or group of resources. * **Roles**: Sets of permissions that you can assign to a principal. Principals are [organization users](/docs/platform/howto/manage-org-users.md), [application users](/docs/platform/concepts/application-users.md), and [groups](/docs/platform/howto/manage-groups.md). You can grant access to principals at the organization, organizational unit, and project level. To give users access to a specific service, create service users. Roles and permissions are cumulative. This means that a user's effective access is the combination of all roles and permissions granted to them at every level. This includes roles and permissions granted directly to the user and those granted to the groups they are a member of. For example, if you grant a user the `project:services:write` permission at the organization level, they have write access to all services in all projects in the organization. If you also assign the user the `read_only` role on a specific project, they still have write access to the services in that project. The less permissive role does not negate the more permissive permission. ## Organization roles and permissions[​](#organization-roles-and-permissions "Direct link to Organization roles and permissions") Roles and permissions at the organization level apply to the organization and all units, projects, and services within it. ### Organization roles[​](#organization-roles "Direct link to Organization roles") | Console name | API name | Allowed actions | | ------------------- | ------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Organization member | None | This is the default role for all organization users. **You cannot grant this role to users.**

All non-managed organization users can: - Edit their profiles.
- Create organizations.
- Leave organizations.
- Add [allowed authentication methods](/docs/platform/howto/set-authentication-policies.md).
- Generate and revoke personal tokens, if allowed by the [authentication policy](/docs/platform/howto/set-authentication-policies.md).
- Enable and disable feature previews.
- View organization users using the Aiven CLI and Aiven API.
- View organization user groups and their membership using the Aiven CLI and Aiven API.
[Managed users](/docs/platform/concepts/managed-users.md) have more restrictions. | | Super admin | None | - Completely unrestricted access to all organization resources and settings, including: all units and projects, billing information, the authentication policy, [other super admin](/docs/platform/howto/manage-permissions.md#make-users-super-admin), organization users, application users, groups, domains, and identity providers.
- Rename the organization.
- Delete the organization. | | Admin | `role:organization:admin` | - Full access to the organization.
- View and change billing information.
- Change the authentication policy.
- Create and delete organizational units and projects.
- Move projects within an organization and to other organizations.
- Invite, deactivate, and remove organization users.
- Create, edit, and delete groups.
- Create and delete application users and their tokens.
- Add and remove domains.
- Add, enable, disable, and remove identity providers. Cannot delete an organization or manage its super admin.

Users who are granted this role on the **unit level** can:- Create and manage projects within the unit. - Grant users and groups permission to the unit. | ### Organization permissions[​](#organization-permissions "Direct link to Organization permissions") | Console name | API name | Allowed actions | | ------------------------------ | ---------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Manage application users | `organization:app_users:write` | - Create, edit, and delete application users.
- View all application users.
- Generate tokens for application users that are not super admin and have not been granted any permissions.
- Revoke application tokens.
- List all application tokens. | | View organization event logs | `organization:event_logs:read` | - View the [event logs](/docs/platform/howto/view-organization-logs.md) for the organization and its organizational units. | | View billing | `organization:billing:read` | - View all billing groups, billing addresses, and payment methods.
- View and download invoices. | | Manage billing | `organization:billing:write` | - Create, edit, and delete billing groups.
- Add, edit, and delete payment methods.
- Add, edit, and delete addresses.
- View and download invoices. | | Manage domains | `organization:domains:write` | - Add, edit, and remove domains.
- View all organization domains. | | View user groups | `organization:groups:read` | View all groups and their members in the Aiven Console. All organization members can view user groups using the Aiven CLI and Aiven API. | | Manage user groups | `organization:groups:write` | - Create and delete groups.
- Rename groups and update group descriptions.
- Add organization and application users to groups that have not been granted any permissions.
- Remove organization and application users from groups. | | View organization networking | `organization:networking:read` | - View all organization VPCs. | | Manage organization networking | `organization:networking:write` | - Add, edit, and remove organization VPCs.
- Create and manage VPC peering connections. | | Manage projects | `organization:projects:write` | - Create and delete projects.
- Assign projects to billing groups.
- Add and remove project tags. **Cannot otherwise access or move the project or its services.** | | View carbon footprint | `organization:sustainability:read` | View the emissions data for the organization. | | View organization users | `organization:users:read` | View organization users and their group memberships in the Aiven Console. All organization members can view users and their group memberships using the Aiven CLI and Aiven API. | | Manage organization users | `organization:users:write` | - Invite new users to the organization.
- View all invited users.
- Remove user invites.
- Deactivate, edit and delete [managed users](/docs/platform/concepts/managed-users.md), including organization admin.
- Remove non-managed users from the organization, including organization admin.
- Reset passwords for managed users.
- View all authentication methods for an organization user.
- Revoke tokens for managed users.
- View all tokens generated by managed users. | ## Project roles and permissions[​](#project-roles-and-permissions "Direct link to Project roles and permissions") You can grant the following roles and permissions to principals. Roles and permissions granted at the project level apply to the project and all services within it. Project roles and permissions granted at the unit level apply to all projects and services within the unit. These permissions apply to the [project API endpoint](https://api.aiven.io/doc/#tag/Organizations/operation/OrganizationProjectsCreate) `/v1/organization/{organization_id}/projects`. ### Project roles[​](#project-roles "Direct link to Project roles") | Console name | API name | Permissions | | ------------------- | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Admin | `admin` | - Full access to the project except billing settings.
- Full access to all of the services in the project.

**This role will be replaced by the `role:project:admin` role.** | | Developer | `developer` | - View project event log.
- View project tags.
- View all services in the project.
- View project permissions.
- View service users.
- View project VPCs.
- Create databases.
- View service connection information.
- View service logs.
- View integration endpoints.
- Get the project's software bill of materials download link.
- Create and change service database users.
- View static IP addresses.
- Remove Aiven for OpenSearch® indexes.
- Create and change Aiven for Apache Kafka® topics.
- Create and change Aiven for PostgreSQL® connection pools. | | Operator | `operator` | - Add, edit, and delete project tags.
- View project tags.
- View project permissions.
- Create, edit, and delete services and their configuration.
- Add and remove dynamic disk sizing and tiered storage.
- Power on and off services.
- Create a fork of a service.
- Enable and disable termination protection.
- Add and remove service contacts.
- Add, edit, and delete service tags.
- View service tags.
- Change clouds and regions.
- Change deployment model of a service.
- Perform service maintenance updates.
- Create, edit, and delete project VPCs and peering connections.
- Update IP allowlists.
- Change the network configuration options.
- View all project VPCs.
- List all peering connections.
- View project event log.
- Get the project's software bill of materials report download link.
- Create, edit, and delete integration endpoints.
- Enable and disable service integrations.
- View integration endpoints.
- View all service integrations for the project, including integrations with services in other projects.
- View service users.
- Manage service users.
- View service user credentials.
- View the list of service backups.
- Configure backup settings.
- View service logs.
- Create, edit, delete, associate and dissociate static IP addresses.

**This role will be replaced by the `role:project:manager` role.** | | Read only | `read_only` | - View project event log.
- View project tags.
- View project permissions.
- View all services and their configuration.
Cannot view Kafka Connect connector configurations. Viewing connector configurations requires the `service:data:write` permission because they can contain secrets in plain text.
- View integration endpoints.
- View static IP addresses. | | Project admin | `role:project:admin` | - Manage project permissions.
- Manage project tags.
- View project audit logs and event logs.
- Manage project integration endpoints and service integrations.
- Manage all services in the project and their configuration.
- Manage project VPCs and peering connections.
- Associate and remove static IP addresses.
- Perform service queries through the API and Console.
- Manage service-specific features such as Kafka topics and schemas, PostgreSQL connection pools, and OpenSearch indexes.
- View and manage service users and credentials.
- View service logs and metrics.
- Update project settings.
- Delete the project. | | Project manager | `role:project:manager` | - View project permissions.
- View project event log.
- Manage project VPCs and peering connections.
- Manage project integration endpoints and service integrations.
- View all services in the project.
- Create and delete services.
- Power on and off services.
- Create a fork of a service.
- Change service plans.
- Change clouds and regions.
- Change deployment model of a service.
- Update IP allowlists.
- Change the network configuration options.
- Add, edit, and delete service tags.
- Enable and disable termination protection.
- Configure backup settings.
- Add and remove service contacts.
- Add and remove dynamic disk sizing and tiered storage.
- Manage service-specific features such as Kafka topics and schemas, PostgreSQL connection pools, and OpenSearch indexes.
- View and manage service users and credentials.
- View service logs.
- View service metrics. | | Project read access | `role:project:read` | - View project event logs.
- View project permissions.
- View integration endpoints.
- View service integrations for the project.
- View all services and their configuration.
- View all project VPCs and peering connections.
- View project tags.
- View logs for all services in the project.
- View metrics for all services in the project.
- View service users in all services in the project.
- View service tags for all services in the project. Service configuration secrets and service user passwords are redacted. | | Maintain services | `role:services:maintenance` | - Perform service maintenance updates.
- Change maintenance windows.
- Upgrade service versions. | | Recover services | `role:services:recover` | - View all details for services in a project.
- Add and remove dynamic disk sizing and tiered storage.
- Change service plans.
- Create a fork of a service.
- Promote read replicas. | ### Project permissions[​](#project-permissions "Direct link to Project permissions") | Console name | API name | Allowed actions | | ------------------------------------- | ----------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | \[Deprecated] View project audit logs | `project:audit_logs:read` | **This permissions is deprecated and will be sunset on 30 October 2026. Use `project:event_logs:read` instead.**

- [View the logs](/docs/platform/howto/view-project-logs.md) for the project.
- View all services in the project. | | View project event logs | `project:event_logs:read` | - [View the logs](/docs/platform/howto/view-project-logs.md) for the project.
- View all services in the project. | | View project integrations | `project:integrations:read` | - View all integration endpoints for the project.
- View all service integrations for the project, including integrations with services in other projects. | | Manage project integrations | `project:integrations:write` | - Add and remove integration endpoints.
- Enable and disable service integrations.
- Create services to integrate an existing service with.
- Read and write integration secrets. | | View project networking | `project:networking:read` | - View all project VPCs.
- List all peering connections. | | Manage project networking | `project:networking:write` | - Create, edit, and delete project VPCs and peering connections.
- View all project VPCs and peering connections. | | View project permissions | `project:permissions:read` | - View all users granted permissions to a project. | | View services | `project:services:read` | - View all details for services in a project, except the service logs and metrics.
- View project settings. | | Manage services | `project:services:write` | - Create and delete services.
- Power on and off services.
- Add and remove dynamic disk sizing and tiered storage.
- Change service plans.
- Change deployment model of a service.
- Change clouds and regions.
- Update IP allowlists.
- Change the network configuration options.
- Add, edit, and delete service tags.
- Enable and disable termination protection.
- Configure backup settings.
- Add and remove service contacts.
- Create a fork of a service. | | Manage service configuration | `service:configuration:write` | - Change clouds and regions.
- Change deployment model of a service.
- Update IP allowlists.
- Change the network configuration options.
- Add and remove service tags.
- Enable and disable termination protection.
- Configure backup settings.
- Add and remove service contacts. | | Access data | `service:data:write` | - Perform service queries through the API and Console.
- View query statistics and current queries.
- Manage service-specific features like Kafka Topics and Schemas, PostgreSQL connection pools, and OpenSearch indexes. | | View service logs | `service:logs:read` | - View logs for all services in the project. **Service logs may contain sensitive information.** | | View service metrics | `service:metrics:read` | - View metrics for all services in the project. | | View configuration secrets | `service:secrets:read` | - Read service configuration secrets such as keys.
- View service users. | | Manage service users | `service:users:write` | - Create and delete service users.
- View service users.
- View, update, and reset connection information for services.
- View service user credentials.
- Manage service user credentials.
- View all services in a project. | Related pages * [Manage permissions](/docs/platform/howto/manage-permissions.md) --- # Rename a service Change the name of an Aiven service by forking it under a new name and deleting the original service. You cannot rename a service after creation. Instead, you can create a fork with the new name and delete the original service. ## Rename a service[​](#rename-a-service "Direct link to Rename a service") 1. Stop writing to the service. 2. Fork the service. 3. Add any integrations or SSO configurations that weren't copied. 4. Connect your clients to the new service. 5. Test the forked service. 6. Delete the original service. Related pages * [Fork a service](/docs/platform/concepts/service-forking.md) * [Power on/off a service](/docs/platform/concepts/service-power-cycle.md) --- # Service backups Learn how Aiven backs up your services automatically, where backups are stored, and how to change your backup schedule. All Aiven services, except for Aiven for Apache Kafka®, have automatic encrypted backups. Backups are stored in the object storage of the cloud region where the service is created, for example, S3 for AWS or Google Cloud Storage for Google Cloud. note If you change a service's cloud provider or an availability zone, its backups are not migrated from their original location. Whenever a service is powered on from a powered-off state, the latest available backup is automatically restored. Backups are automatically deleted 30 days after the service's deletion date. ## Access to backups[​](#access-to-backups "Direct link to Access to backups") Backups are encrypted and not available for download, but you can create your own backups with the appropriate tooling: * [PostgreSQL®](https://www.postgresql.org/docs/current/app-pgdump.html): `pgdump` * [MySQL®](https://dev.mysql.com/doc/refman/8.4/en/mysqldump.html): `mysqldump` * [OpenSearch®](https://github.com/elasticsearch-dump/elasticsearch-dump): `elasticdump` * [Valkey™](https://valkey.io/topics/cli/): `valkey-cli` ## Edit the backup schedule[​](#edit-the-backup-schedule "Direct link to Edit the backup schedule") To edit the backup schedule for your service: * Console * Aiven API * Aiven CLI * Terraform 1. In your service, **Backups**. 2. Click **Actions** > **Configure backup settings**. 3. Click **Add configuration options**. 4. Add `backup_hour` and `backup_minute`, and set their values. 5. Click **Save configuration**. Call the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint, and add the following properties to the `user_config` object: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "backup_hour": BACKUP_HOUR, "backup_minute": BACKUP_MINUTE } }' ``` Replace the following: * `SERVICE_NAME`: the name of your service. * `PROJECT_NAME`: the name of your project. * `BACKUP_HOUR`: the hour when the service backup starts. Accepted values are integers between `0` and `23`. * `BACKUP_MINUTE`: the minute when the service backup starts. Accepted values are integers between `0` and `59`. Run the [avn service update](/docs/tools/cli/service-cli.md#avn-cli-service-update) command, and add the following properties to the `user_config` object: ``` avn service update SERVICE_NAME \ --project PROJECT_NAME \ --user-config '{ "backup_hour": BACKUP_HOUR, "backup_minute": BACKUP_MINUTE }' ``` Replace the following: * `SERVICE_NAME`: the name of your service. * `PROJECT_NAME`: the name of your project. * `BACKUP_HOUR`: the hour when the service backup starts. Accepted values are integers between `0` and `23`. * `BACKUP_MINUTE`: the minute when the service backup starts. Accepted values are integers between `0` and `59`. Use the `backup_hour` and `backup_minute` attributes in [your service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs) to set the start time for backups. If a backup was recently made, it can take another backup cycle before the new backup time takes effect. Related pages * [Track service restore progress using the API](/docs/platform/howto/restore_progress_updates.md) * [Backup to another region](/docs/platform/concepts/backup-to-another-region.md) --- # Service and feature releases New services and features follow a release lifecycle that promotes quality by giving customers the chance to test them and provide feedback. The release stages and labels for new services and features vary across the Aiven Console and other interfaces like the Aiven API and Aiven Terraform Provider. ## Aiven Console[​](#aiven-console "Direct link to Aiven Console") The release stages for features in the Aiven Console are: 1. Limited availability 2. Early availability 3. General availability ### Limited availability [Limited availability](/docs/platform/concepts/service-and-feature-releases.md)[​](#limited-availability- "Direct link to limited-availability-") The limited availability stage is an initial release of a new feature or service that you can try out by invitation only. To try a service or feature in this stage, contact the [sales team](https://aiven.io/contact). ### Early availability [Early availability](/docs/platform/concepts/service-and-feature-releases.md)[​](#early-availability- "Direct link to early-availability-") Features and services in the early availability stage are released to all users for testing and feedback. There are some restrictions on the functionality and [service level agreement](https://aiven.io/sla). They are intended for use in non-production environments, but you can test them with production-like workloads to gauge their behavior under heavy load. Support tickets for these services and features are considered [low severity](https://aiven.io/support-services). You can enable most early availability features yourself on the [feature preview page](/docs/platform/howto/feature-preview.md) in the user profile. note Using an early availability means that you consent to Aiven contacting you to request feedback that helps shape the products. ### General availability[​](#general-availability "Direct link to General availability") New services and features become generally available once they are ready for production use at scale. The duration of the early and limited availability stages depends on customer adoption and the time it takes to address any issues discovered during testing. ### Billing for limited and early availability services[​](#billing-for-limited-and-early-availability-services "Direct link to Billing for limited and early availability services") Aiven provides credits for customers to try out limited and early availability services. After the credit code expires or after you have used the credits, you are charged for the service at the usual rate. ## Aiven API[​](#aiven-api "Direct link to Aiven API") Some new API endpoints are marked experimental, but the feature is stable and functional. The API specification for experimental endpoints can be adjusted based on feedback or evolving requirements. Any such changes are communicated to ensure a smooth transition. ## Aiven Provider for Terraform[​](#aiven-provider-for-terraform "Direct link to Aiven Provider for Terraform") Resources and data sources can be beta or deprecated. ### Beta[​](#beta "Direct link to Beta") Resources and data sources that are in development have the beta label. You can use these by setting the [beta environment variable](https://registry.terraform.io/providers/aiven/aiven/latest/docs#environment-variables). ### Deprecated[​](#deprecated "Direct link to Deprecated") Deprecated resources and data sources have warnings in the documentation and when you plan or apply changes. Warnings include details on which resources to use as replacements and the [migration guide](https://registry.terraform.io/providers/aiven/aiven/latest/docs/guides/update-deprecated-resources) provides instructions on moving to the new resources. ## Feature roadmap[​](#feature-roadmap "Direct link to Feature roadmap") You can see Aiven's public roadmap, track the progress of specific features, submit your own suggestions, and vote on other suggestions on [Aiven Ideas](https://ideas.aiven.io/). --- # Fork a service Fork an Aiven service to create an independent copy for testing, debugging, or development without affecting the original service. Fork an Aiven service to create a complete copy of it from its latest backup. Forked services are independent and don't share resources with or increase the load on the original service. Common use cases for forking include: * Creating a snapshot to analyze an issue. * Creating a development copy of your production environment. * Testing upgrades before applying them to production services. * Creating an instance in a different cloud provider, region, or with a different plan. * Renaming a service. During the forking process, the fork might initially have only one node while backups are being taken. The other nodes appear after the backup process is complete. When you fork a service, its configuration and data is copied to the new service from the latest backup. Support for forking from a specific point in time depends on the service type. ## Fork a service[​](#fork-a-service "Direct link to Fork a service") * Console * CLI * API * Terraform 1. In your service, in the **Backups** section, click **Backup management**. 2. Click **Fork & restore**. 3. Choose the backup to fork from. 4. Enter a name, and select the cloud and plan. 5. Click **Create fork**. Use the [create service command](/docs/tools/cli/service-cli.md#avn-cli-service-create) with: * `--service-to-fork-from`: the name of the service to use as the source. * `--project-to-fork-from`: to fork a service in a different project, set this to the project name the source service is in. Use the [`ServiceCreate` endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) and in the `user_config` property set: * `service_to_fork_from`: the name of the source service. * `project_to_fork_from`: to fork a service in a different project, set this to the name of the project the source service is in. Use the `service_to_fork_from` attribute in the user config of your service resource. To fork a service in a different project, set the `project_to_fork_from` attribute. More information on the service resources and their configuration options is available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Service backups](/docs/platform/concepts/service_backups.md) * [Rename a service](/docs/platform/concepts/rename-services.md) * [Change a service plan](/docs/platform/howto/scale-services.md) --- # Service integrations overview Service integrations provide additional functionality and features by connecting different Aiven services together. tip Use an AI assistant connected to [Aiven MCP](/docs/tools/mcp-server.md) to create, manage, and view details for service integrations from clients such as Cursor and Claude Code. ## About service integrations[​](#about-service-integrations "Direct link to About service integrations") Service integrations include: * Metrics integration, which allows sending advanced telemetry data to an [Aiven for PostgreSQL®](https://aiven.io/postgresql) database for metrics and visualizing them in [Aiven for Grafana®](https://aiven.io/grafana) * [Log integration](/docs/products/opensearch/howto/opensearch-log-integration.md), which allows sending logs from any Aiven service to [Aiven for OpenSearch®](https://aiven.io/opensearch). ## Benefits of integrated telemetry[​](#benefits-of-integrated-telemetry "Direct link to Benefits of integrated telemetry") Aiven automatically provides basic host-level resource metrics (CPU, memory, disk and network) for every service under the **Metrics** tab on the Aiven Console. The advanced telemetry feature brings much more detailed, service-specific metrics to the Aiven user, who can then drill down into the service behavior and identify issues and performance bottlenecks. ### Service integrations billing[​](#service-integrations-billing "Direct link to Service integrations billing") The advanced telemetry data and predefined dashboards do not cost anything extra, but they require you to have an Aiven for PostgreSQL® and an Aiven for Grafana® service, which will be billed hourly as regular Aiven services. ### Can I define my own metrics dashboards?[​](#can-i-define-my-own-metrics-dashboards "Direct link to Can I define my own metrics dashboards?") Yes. You can use the predefined dashboards as a starting point and save them under a different name or you can build one from scratch. Dashboards whose title starts with the word "Aiven" are automatically managed, so it is better to name the dashboards differently from that. ### Can I access the telemetry data directly in PostgreSQL?[​](#can-i-access-the-telemetry-data-directly-in-postgresql "Direct link to Can I access the telemetry data directly in PostgreSQL?") Yes. The PostgreSQL service is a normal PostgreSQL database service that can be accessed via any PostgreSQL client software or your own application or script. You can perform complex queries over the data and so on. ### Can I add alerts in Grafana for the telemetry data?[​](#can-i-add-alerts-in-grafana-for-the-telemetry-data "Direct link to Can I add alerts in Grafana for the telemetry data?") Yes. You can add alert thresholds for individual graphs and attach different alerting mechanisms to them for sending out alerts. See the [Grafana documentation](/docs/products/grafana.md) for more information. ## Use service integrations[​](#use-service-integrations "Direct link to Use service integrations") You can [create service integrations](/docs/platform/howto/create-service-integration.md) or start using existing integrations. --- # Service memory limits Understand how Aiven limits the memory available to a service, and how that limit relates to the physical memory of the underlying node. The practical memory limit will always be less than the service physical memory limit. **All services are subject to operating overhead:** * A small amount of memory is required by the operating system kernel to manage system resources, including networking functions and disk cache. * Aiven's cloud data platform requires memory to monitor availability, provide metrics, logging and manage backups. A server or node's **usable memory** can be calculated as: `usable memory = RAM - overhead` Where: * `overhead` is 350 MiB (≈ 0.34 GiB). Services may utilize optional components, service integrations, connection pooling, or plug-ins, which are not included in overhead calculations. If a service is overcommitted, the operating system, management layer, backups or availability monitoring, may fail status checks or operations due to resource contention. In severe instances, the node may fail completely with an out-of-memory condition. Related pages * [Out of memory conditions](/docs/platform/concepts/out-of-memory-conditions.md) * [Prepare services for high load](/docs/platform/howto/prepare-for-high-load.md) * [Change a service plan](/docs/platform/howto/scale-services.md) --- # Power on/off a service Power off an Aiven service to release resources and save credits, power it back on when you need it, or delete it permanently. ## Power off a service[​](#power-off-a-service "Direct link to Power off a service") When you power off a service: * All virtual machines are removed from the public cloud. * The service configuration is stored on the Aiven Platform. * If there are no backups, all service data is lost. * If the service has time-based or point in time recovery backups, the backups remain on the Aiven Platform. Services powered off for more than 180 days are automatically deleted. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power off service**. To power off a service, run: ``` avn service update SERVICE_NAME --power-off ``` ## Power on a service[​](#power-on-a-service "Direct link to Power on a service") When you power on a service: * New virtual machines are created on the service's public cloud. * The service starts with the stored configuration parameters. * The latest time-based backup is restored. * Maintenance updates are automatically applied. * If a point in time recovery backup is available, the database transaction logs are replayed to recover the service data to a specific point in time. The restoration takes from a few minutes to a few hours, depending on the network bandwidth, the disk IOPS allocated to the service, and the size of the backup. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power on service**. To power on a service, run: ``` avn service update SERVICE_NAME --power-on ``` To see when the service is running, run: ``` avn service wait SERVICE_NAME ``` note When you power on a service, your data is restored from the latest available backup. An automatic backup is also taken before the service is powered off. note Static IP addresses are not removed when a service is powered off or deleted. They continue to generate the usual costs. To avoid these costs, [remove the static IP addresses](/docs/platform/concepts/static-ips.md). ## Delete a service[​](#delete-a-service "Direct link to Delete a service") * Console * CLI 1. In your project, click **Services**. 2. Open the service to delete, and click **Actions** > **Delete service**. To delete a service, run: ``` avn service terminate SERVICE_NAME ``` Related pages * [Fork a service](/docs/platform/concepts/service-forking.md) * [Service backups](/docs/platform/concepts/service_backups.md) --- # Service pricing All Aiven services are billed based on actual usage so you only pay for the resources you use. Services are charged by the hour while they are powered on. The minimum hourly charge unit is one hour. For example, if you create an Aiven service and power it off after 40 minutes, you are charged for one hour of usage. After 20.5 hours, you are charged for 21 hours. Powering off a service stops the accumulation of new charges immediately. Costs for all services in a project are charged separately, but you can consolidate the charges for multiple projects by assigning them to a [billing group](/docs/platform/howto/use-billing-groups.md). ## Service plans[​](#service-plans "Direct link to Service plans") The cost of an Aiven service plan is all-inclusive and covers: * Virtual machine costs * Network costs * Backup costs * Setup and maintenance costs * Migration between clouds or plans * Migration to another region * [Credit card and processing fees payable by Aiven](#credit-card-fees) There are additional costs for some features such as PrivateLink and additional storage. Network traffic is not charged separately, but your application cloud service provider might charge you for the network traffic going to or from their services. Learn more about and compare the different plans on the [plans and pricing page](https://aiven.io/pricing?product=opensearch\&tab=plan-pricing). To ensure consistent performance and allow for seamless migration, all Aiven service plans have standardized resources regardless of the underlying cloud provider. note Since available virtual machine (VM) sizes differ between clouds, services can be provisioned on a VM with more resources (CPU, RAM, or disk storage) than advertised. These additional resources are not guaranteed. Aiven reserves the right to switch to a more appropriately sized VM if one becomes available from the cloud provider. ### Free tier[​](#free-tier "Direct link to Free tier") The Free tier is available for the following service types: * [Aiven for Apache Kafka®](/docs/products/kafka/free-tier/kafka-free-tier.md) * [Aiven for MySQL®](/docs/products/mysql/concepts/mysql-free-tier.md) * [Aiven for OpenSearch®](/docs/products/opensearch/concepts/opensearch-free-tier.md) * [Aiven for PostgreSQL®](/docs/products/postgresql/concepts/pg-free-tier.md) * [Aiven for Valkey™](/docs/products/valkey/concepts/valkey-free-tier.md) You don't need a credit card to create a free service and you can use them indefinitely free of charge. Free services do not have any time limitations. However, Aiven reserves the right to: * Power off free services with no initial usage within the first few hours after the service is running. You can power them back on at any time. * Power off free services with no continuative activity on the service. A notification is sent before the service is powered off. You can power them back on at any time. * Shut down services if Aiven believes they violate the [acceptable use policy](https://aiven.io/terms). * Change the cloud provider, region, or configuration at any time. You can run free services alongside a free trial without affecting your trial credits. Free services also continue running after your trial has expired. You can upgrade your free service to a paid plan at any time by adding a payment method to the project's billing group. ### Developer tier[​](#developer-tier "Direct link to Developer tier") Developer tier plans, pricing, and limits vary by product. Developer tier services are not automatically powered off when inactive. #### Aiven for Apache Kafka®[​](#aiven-for-apache-kafka "Direct link to Aiven for Apache Kafka®") Aiven for Apache Kafka® has a separate [Developer tier](/docs/products/kafka/dev-tier/kafka-dev-tier.md) paid plan for development and testing, with its own throughput, storage, retention, and pricing. #### Aiven for OpenSearch[​](#aiven-for-opensearch "Direct link to Aiven for OpenSearch") The Aiven for OpenSearch Developer tier provides you with always-on OpenSearch for prototypes and personal projects. The Aiven for OpenSearch Developer tier includes: * Single node * 1 CPU per virtual machine * 4 GB RAM * 30 GB storage * 2 shards, maximum 20 shards per node * Up to 50 concurrent connections * Daily backups * Aiven and third-party service integrations * [Basic tier support](/docs/platform/howto/support.md) Limitations of the Aiven for OpenSearch Developer tier are: * No choice of cloud provider or specific cloud region * Cannot create the service in a VPC * No static IPs * No dynamic disk scaling * No forking Aiven reserves the right to change the cloud provider, region, or configuration of these Developer tier services at any point in time. #### Aiven for PostgreSQL and Aiven for MySQL[​](#aiven-for-postgresql-and-aiven-for-mysql "Direct link to Aiven for PostgreSQL and Aiven for MySQL") The Developer tier is available for Aiven for PostgreSQL® and Aiven for MySQL® services, letting you scale up your service in a cost-effective way. The Developer tier includes: * A single node * 1 CPU per virtual machine * 1 GB RAM * up to 8 GB disk storage * Monitoring for metrics and logs * Backups * [Basic tier support](/docs/platform/howto/support.md) Limitations of the Developer tier are: * No choice of cloud provider or specific cloud region * Cannot create the service in a VPC * No static IPs * No integrations * No forking * For PostgreSQL: No connection pooling * For PostgreSQL: `max_connections` limit set to `20` Aiven reserves the right to change the cloud provider, region, or configuration of these Developer tier services at any point in time. #### Aiven for Valkey™[​](#aiven-for-valkey "Direct link to Aiven for Valkey™") The Aiven for Valkey Developer tier includes: * Single node * 2 CPUs per VM * 4 GB RAM per VM * Monitoring for metrics and logs * Backups * [Basic tier support](/docs/platform/howto/support.md) Limitations of the Aiven for Valkey Developer tier are: * No choice of cloud provider or specific cloud region * Cannot create services in VPCs * No static IPs * No integrations * No forking Aiven reserves the right to change the cloud provider, region, or configuration of Aiven for Valkey Developer tier services at any point in time. ### Custom plans[​](#custom-plans "Direct link to Custom plans") If the service plans don't fit your use cases, you can request a custom plan. Custom plans are most useful for special cases, like a very high throughput cluster. Custom plans are available for all Aiven service types. The starting price is $5,000 USD per month. You can adjust the following in custom plans: * Storage capacity * Backup frequency * Number of nodes * CPU/RAM configuration per node There can be limitations depending on the cloud provider, region, available instance types, and service type. #### Request a custom plan[​](#request-a-custom-plan "Direct link to Request a custom plan") To get a quote for a custom plan: Contact the [sales team](https://aiven.io/contact) with the following information: * **Cloud**: The cloud provider and region, for example: Google Cloud `us-east1`. * **Service type**: For example: Aiven for PostgreSQL®. * **Aiven project**: The name of the project to add the custom plan to. * **Configuration**: The storage, backup frequency, or other variables to change from the default plan. ### Credit card fees[​](#credit-card-fees "Direct link to Credit card fees") The prices listed on the website and in your invoices are inclusive of all credit card and processing fees that are payable by Aiven. Some credit card issuers add extra charges on top of the fees Aiven charges you. The most common fee is an international transaction fee. Some issuers charge this fee for transactions where the native country of the merchant, processor, bank, and card are different. Aiven is based in Finland and the processor is based in the United States. Such fees are not added by or visible to Aiven, so they cannot be included in the prices or waived. ## Free trials[​](#free-trials "Direct link to Free trials") Aiven offers a free trial for 30 days with credits you can use to explore the Aiven Platform. You don't need a credit card to sign up. You can use trial credits for any paid services and paid features like virtual private cloud peering. The trial starts when you create your Aiven user account. Trials include: * Up to 10 VMs * 1 [Virtual Private Cloud (VPC)](/docs/platform/howto/manage-project-vpc.md) * Up to 10 VPC peering connections There are some limitations: * You can only have one trial. * Trial credits are added to a single [billing group](/docs/platform/howto/use-billing-groups.md). You cannot transfer them to another billing group. * You cannot create new services if your remaining credits would be spent too quickly. * You cannot end a free trial yourself before the end of the trial period. If you didn't add a payment card to the billing group for your project, services in that project are automatically powered off when the trial ends or your credits run out. To keep your services running [add a payment card](/docs/platform/howto/manage-payment-card.md) and assign it to the billing group for the project with your services. Related pages * [Tax information for Aiven services](/docs/platform/concepts/tax-information.md) --- # Static IP addresses Aiven services are normally addressed by their hostname, but static IP addresses are also available for an additional charge. Static IP address are useful with the following: * Firewall rules for specific IP addresses * Tools such as proxies that use IP addresses rather than hostnames Static IP addresses on the Aiven platform are created in a specific cloud and belong to a specific project. During their lifecycle, the IP addresses can be in a number of states, and move between states by being created/deleted, by being associated with a service, or when the service is reconfigured to use the static IP addresses. ## Manage IP addresses[​](#manage-ip-addresses "Direct link to Manage IP addresses") To create, delete, associate or dissociate IP addresses, use the `avn static-ip` CLI: ``` avn static-ip associate|create|delete|dissociate|list ``` To reconfigure a service, set its `static_ips` configuration value to `true` to use static IP addresses, or `false` to stop using them. note * The `static_ip` configuration can only be enabled when enough static IP addresses have been created and associated with the service. * Static IP addresses are not removed when a service is powered off or deleted. They continue to generate the usual costs. To avoid these costs, [remove the static IP addresses](/docs/platform/concepts/static-ips.md). By default, an Aiven service will use public IP addresses allocated from the cloud provider's shared pool of addresses for the cloud region. As a result, Aiven service IP addresses cannot be predicted. This is a good approach for most use cases, but we do also offer static IP addresses, should you need them. This allows you to configure a firewall rule for your own services. note Each static IP address will incur a small charge. ## Calculate the number of IP addresses needed[​](#platform_howto_setup_static_ip "Direct link to Calculate the number of IP addresses needed") Work out how many IP addresses your service will need to guarantee that the service has enough static IP addresses available to handle upgrades and failures. The table below shows how to calculate the number of IP addresses needed, depending on the number of nodes your service has. | Number of nodes | IP addresses needed | Example | | --------------- | -------------------------- | --------------------------------------------------------- | | up to 6 nodes | double the number of nodes | Service with 3 nodes needs 6 static IP addresses (2 \* 3) | | 6+ nodes | number of nodes + 6 | Service with 9 nodes needs 15 static IP addresses (9 + 6) | ## Reserve static IP addresses[​](#reserve-static-ip-addresses "Direct link to Reserve static IP addresses") Create a static IP addresses using the [Aiven CLI](/docs/tools/cli.md), repeat as many times as needed to create enough IP addresses for your service. Specify the name of the cloud that the IP address should be created in, to match the service that will use it. ``` avn static-ip create --cloud azure-westeurope ``` The command returns some information about the newly created static IP address. ``` CLOUD_NAME IP_ADDRESS SERVICE_NAME STATE STATIC_IP_ADDRESS_ID ================ ========== ============ ======== ==================== azure-westeurope null null creating ip359373e5e56 ``` When the IP address has been provisioned, the state turns to `created`. The list of static IP addresses in the current project is available using the `static-ip list` command: ``` avn static-ip list ``` The output shows the state of each one, and its `ID`, which you will need in the next step. ``` CLOUD_NAME IP_ADDRESS SERVICE_NAME STATE STATIC_IP_ADDRESS_ID ================ ============= ============ ======= ==================== azure-westeurope 13.81.29.69 null created ip359373e5e56 azure-westeurope 13.93.221.175 null created ip358375b2765 ``` Once the status shows as `created` IP address can be associated with a service. ## Associate static IP addresses with a service[​](#associate-static-ip-addresses-with-a-service "Direct link to Associate static IP addresses with a service") Using the name of the service, and the ID of the static IP address, you can assign which service a static IP should be used by: ``` avn static-ip associate --service my-static-pg ip359373e5e56 avn static-ip associate --service my-static-pg ip358375b2765 ``` When you have the required number of addresses available, you can enable static IP addresses on your service. ## Configure a service to use static IP[​](#configure-a-service-to-use-static-ip "Direct link to Configure a service to use static IP") Enable static IP addresses for the service by setting the `static_ips` user configuration option: ``` avn service update -c static_ips=true my-static-pg ``` note This leads to a rolling forward replacement of service nodes, similar to applying a maintenance upgrade. The new nodes will use the static IP addresses associated with the service Once the nodes have been replaced, they will be using the static IP addresses that you associated with the service. You can check which ones are in use by running the `avn static-ip list` command again. The `available` state means that the IP is associated to the service, and `assigned` means that it is in active use. ## Remove static IP addresses[​](#platform_howto_remove_static_ip "Direct link to Remove static IP addresses") note To dissociate an IP from a service, the service must either have enough IP addresses still available, or the `static_ips` configuration setting for the service must be set to `false`. Static IP addresses are removed by first dissociating them from a service, while they are not in use. This returns them back to the `created` state to either be associated with another service, or deleted. ``` avn static-ip dissociate ip358375b2765 ``` To delete a static IP: ``` avn static-ip delete ip358375b2765 ``` Neither the `dissociate` or `delete` commands return any information on success. --- # Tax information for Aiven services Aiven services are provided by Aiven Ltd, a private limited company incorporated in Finland, but operating in various countries. None of Aiven's marketed prices include value-added or other taxes. ## Aiven's tax status[​](#aivens-tax-status "Direct link to Aiven's tax status") Aiven has the following tax status in the jurisdictions it operates in. ### European Union[​](#european-union "Direct link to European Union") Finnish law requires Aiven to charge a value-added tax for services provided within the European Union. The value-added tax percentage depends on the domicile of the customer. For business customers in EU countries other than Finland, Aiven can apply the reverse charge mechanism of 2006/112/EC article 196 when a valid VAT ID [is added to a billing group](#add-a-vat-id-to-a-billing-group) for a billing group. ### United States[​](#united-states "Direct link to United States") According to the tax treaty between Finland and the United States, no direct tax needs to be withheld from payments by American entities to Aiven. Although a W-8 form is not required to confirm this status, we're happy to provide a W-8BEN-E form describing Aiven's status. Contact to request one. Aiven charges American sales taxes in several states in the United-States. The applicability of these taxes depends on various factors, such as the taxability of Aiven services, tax thresholds, amongst others. ### Rest of the world[​](#rest-of-the-world "Direct link to Rest of the world") No European Uninion value-added tax is applied to services sold outside the European Union as described in the Value-Added Tax Act of Finland section 69h. Finland has tax treaties with most countries in the world affirming this status. Aiven reserves the right to charge any necessary taxes, similar to American sales taxes, depending on the local tax legislation of the customer. ## Add a VAT ID to a billing group[​](#add-a-vat-id-to-a-billing-group "Direct link to Add a VAT ID to a billing group") You must have the `organization:billing:write` [permission](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) to add a VAT ID to a billing group. 1. Click **Billing**. 2. Click **Billing groups**. 3. Find the billing group to update and click **Details**. 4. Click **Billing details**. 5. In the **Other details** section, click **Change**. 6. Enter your **VAT ID** and click **Save changes**. It can take up to an hour for the new tax status to be visible on the invoice estimates. --- # TLS/SSL certificates All traffic to Aiven services is always protected by TLS. It ensures that third parties can't eavesdrop or modify the data while in transit between Aiven services and the clients accessing them. Every Aiven project has its own private Certificate Authority (CA) which is used to sign certificates that are used internally by the Aiven services to communicate between different cluster nodes and to Aiven management systems. Some service types uses the Aiven project's CA for external connections. To access these services, download the CA certificate and configure it on your browser or client. For other services a browser-recognized CA is used, which is normally already marked as trusted in browsers and operating systems, so downloading the CA certificate is not normally required. note All the services in a project share the same Certificate Authority (CA). ## Certificate requirements[​](#certificate-requirements "Direct link to Certificate requirements") Most of our services use a browser-recognized CA certificate, but there are exceptions: * **Aiven for PostgreSQL®** requires the Aiven project CA certificate to connect when using `verify-ca` or `verify-full` as `sslmode`. The first mode requires the client to verify that the server certificate is actually emitted by the Aiven CA, while the second provides maximum security by performing HTTPS-like validation on the hostname as well. The default `sslmode=require` ensures TLS is used when connecting to the database, but does not verify the server certificate. For more information, see the [PostgreSQL documentation](https://www.postgresql.org/docs/current/ssl-tcp.html) * **Aiven for MySQL®** requires the Aiven project CA certificate to connect when using `VERIFY_CA` or `VERIFY_IDENTITY` as the SSL mode. `VERIFY_CA` requires the client to verify that the server certificate is signed by the Aiven CA, while `VERIFY_IDENTITY` also validates the hostname. For more information, see the [MySQL documentation](https://dev.mysql.com/doc/refman/8.4/en/using-encrypted-connections.html). * **Aiven for Apache Kafka®** supports different authentication methods: * **Client certificate**. The client authenticates with a client certificate and key. This method requires the Aiven project CA certificate, the client certificate, and the client key. * **SASL over SSL**. The client authenticates with a service username and password. Communication is encrypted with the project CA certificate by default. You can enable the `letsencrypt_sasl` setting to use a public CA instead of the project CA. For details, see [Enable and configure SASL authentication](/docs/products/kafka/howto/kafka-sasl-auth.md). * **Aiven for Valkey™** uses a browser-recognized (Let's Encrypt) certificate by default, so no CA certificate download is required. Services created before this certificate mode was enabled still use the Aiven project CA certificate. If the **Overview** page for your service offers a CA certificate to download, your service uses the project CA. There's no self-service option to use the project CA certificate for a service that uses a browser-recognized certificate. To request this, [open a support ticket](/docs/platform/howto/support.md). For details, see [Manage SSL connectivity in Aiven for Valkey™](/docs/products/valkey/howto/manage-ssl-connectivity.md). If your service uses the project CA certificate, it also goes through periodic [certificate rotation](#certificate-rotation) like other services that use this CA. You can download the project CA certificates from the **Overview** page of your service. For steps, see [Download the project CA certificates](/docs/platform/concepts/tls-ssl-certificates.md#download-ca-certificates). note Some older services use the Aiven project CA certificate. To switch to a browser-recognized certificate, [open a support ticket](/docs/platform/howto/support.md). ## Certificate rotation[​](#certificate-rotation "Direct link to Certificate rotation") To keep certificates secure, Aiven periodically rotates the project CA certificate, even though its listed expiration date can be many years away. A rotation can happen because the certificate is approaching expiration, or for other operational or security reasons. All services in a project share the same CA, so a rotation happens at the project level, but each service picks up the new certificate during its own maintenance window, using the same [maintenance process](/docs/platform/concepts/maintenance-window.md) as other updates. Because of this, services in the same project can start trusting the new certificate at different times. During a rotation, your service trusts both the current and the new CA certificate. This overlap is sometimes called a certificate bundle. This matters if your client verifies the server certificate against a specific CA. Examples include PostgreSQL's `verify-ca` or `verify-full` modes, MySQL's `VERIFY_CA` or `VERIFY_IDENTITY` modes, and an older Aiven for Valkey™ service that still uses the project CA certificate. In these cases, update your client to trust the new certificate in the bundle before the rotation completes. Otherwise, your client can't verify the server certificate and the connection fails. Aiven sends an email notification to your project and service contacts before a certificate rotation. Confirm that your [project and service contacts](/docs/platform/howto/technical-emails.md) are up to date so that you receive these notifications. ## Download CA certificates[​](#download-ca-certificates "Direct link to Download CA certificates") If your service needs a CA certificate, download one: 1. Open your service's **Overview** page. 2. In the **Connection information** section, find **CA Certificate** and click **Download**. You can also use the `avn service user-creds-download` [CLI](/docs/tools/cli/service/user.md#avn_service_user_creds_download): ``` avn service user-creds-download --username ``` Related pages * [Manage project and service notifications](/docs/platform/howto/technical-emails.md) * [Service maintenance, updates and upgrades](/docs/platform/concepts/maintenance-window.md) * [Manage SSL connectivity in Aiven for Valkey™](/docs/products/valkey/howto/manage-ssl-connectivity.md) * [Support](/docs/platform/howto/support.md) --- # User and access management There are several types of users in the Aiven Platform: * Super admin: Users with full access to the organization. Limit the number of these users in an organization for increased security. * [Organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) with full access to an organization's users and resources. * [Organization members](/docs/platform/howto/manage-org-users.md): All users in an organization. * [Managed users](/docs/platform/concepts/managed-users.md): Users that have more restrictions and are centrally managed. With a verified domain organization users automatically become managed users. * [Application users](/docs/platform/concepts/application-users.md): Non-human users intended for programmatic access to the platform. All users can be added to [groups](/docs/platform/howto/manage-groups.md) to help streamline the process of granting access to organization resources. You grant them access to resources with [roles and permissions](/docs/platform/concepts/permissions.md). --- # Virtual private clouds (VPCs) in Aiven Virtual private clouds (VPCs) supported on the Aiven Platform provide enhanced security, flexibility, and network control, allowing efficient traffic, resource, and access management. A VPC is a logically isolated section of a cloud provider's network, which makes it a private network within a public cloud. It's a secure customizable network environment that you define and control to deploy and manage resources. ### VPC characteristics[​](#vpc-characteristics "Direct link to VPC characteristics") * Isolation: Each VPC operates independently from other VPCs, ensuring secure separation. * Security: Control over network traffic and isolation. * Internet connectivity: Control whether the VPC connects to the internet via internet gateways or remains isolated. * Network control: Configure route tables, network gateways, and security settings. * Customizable IP range: You can define your own IP address range (CIDR block). * Subnets: Divide the VPC into smaller sub-networks (subnets) for organizing resources based on availability zones or functional groups. * Flexibility: Custom network architecture tailored to your application's needs. * Scalability: Expand or modify the network as demand grows. Custom domain restrictions in VPCs When you set a custom domain for a service deployed in a VPC, the service certificate is only created for the `public-*` hostname and the custom domain. Certificate verification will fail for the `private-*` hostname and the dynamic service name. To avoid certificate verification issues, ensure your applications connect using either the `public-*` hostname or the custom domain when accessing VPC-deployed services. ### VPC components[​](#vpc-components "Direct link to VPC components") * Subnets: Represent smaller public or private networks within the VPC. * [Peering connection](/docs/platform/howto/list-vpc-peering.md): Connect VPCs for intercommunication. * NAT (Network Address Translation) gateway: Allows outbound internet access for private subnets. * Internet gateway (IGW): Enables public traffic to access the internet. * Security groups: Represent firewall rules controlling inbound and outbound traffic for resources. * Route tables: Specify how traffic is directed within the VPC. * Network Access Control Lists (NACLs): Constitute an extra layer of security at the subnet level ### VPC use cases[​](#vpc-use-cases "Direct link to VPC use cases") * Data isolation: Keeping sensitive data within a private network * Hosting applications: Deploying scalable web and database applications * Multi-tier architecture: Separating application layers (web, app, database) within distinct subnets * Hybrid cloud architecture: Connecting on-premises networks to the cloud securely ## VPC types[​](#vpc-types "Direct link to VPC types") The Aiven Platform allows creating and using two types of VPCs, which differ in scope: project-wide VPCs and organization-wide VPCs. ### Project VPCs[​](#project-vpcs "Direct link to Project VPCs") A project VPC is a VPC that spans a single Aiven project within your Aiven organization. A project-wide VPC allows all resources in that project to interconnect and share a common VPC network, simplifying network management and promoting consistency across your Aiven project's services. Learn how to [create and manage projects VPCs in Aiven](/docs/platform/howto/manage-project-vpc.md). ### Organization VPCs [Limited availability](/docs/platform/concepts/service-and-feature-releases.md)[​](#organization-vpcs- "Direct link to organization-vpcs-") An organization VPC is a VPC that spans multiple Aiven projects within your Aiven organization. An organization-wide VPC allows different projects to share a centralized network infrastructure while maintaining isolation and control. Learn how to [create and manage organization VPCs in Aiven](/docs/platform/howto/manage-organization-vpc.md). Related pages For information on VPCs supported by particular cloud providers, see the following: * AWS: [How Amazon VPC works](https://docs.aws.amazon.com/vpc/latest/userguide/how-it-works.html) * Google Cloud: [VPC networks](https://cloud.google.com/vpc/docs/vpc) * Azure: [What is Azure Virtual Network?](https://learn.microsoft.com/en-us/azure/virtual-network/virtual-networks-overview) * UpCloud: * [How to configure SDN Private networks](https://upcloud.com/docs/guides/configure-sdn-private-networks/) * [How to configure SDN Private networks using the UpCloud API](https://upcloud.com/docs/guides/configure-sdn-private-networks-upcloud-api/) --- # Add authentication methods You can authenticate directly with your email and password, or use single sign-on through providers like GitHub, Google, and Microsoft. To add an authentication method for your user account in the [Aiven Console](https://console.aiven.io/): 1. Click the user information icon in the top right and select **Authentication**. 2. Click **Add authentication method**. 3. Select Aiven password or an external account provider. 4. Click **Add authentication method**. After authorizing access, the new method is shown in the list. --- # Scale disk storage Scale the disk storage of an Aiven service up or down without disrupting the running service. /eol-for-major-versions#aiven-for-flinkAdding or removing disk storage does not disrupt the running service. You pay only for extra storage instead of upgrading compute resources. You can add extra storage when you create a service or after it is running. When you add storage to a running service, the Aiven Platform provisions the extra disk and adds it to the running instances. For a clustered service such as Aiven for Apache Kafka®, Aiven divides extra storage equally between the nodes. For a shared service, each node receives the full extra capacity. ## Limitations[​](#limitations "Direct link to Limitations") * Disk added for extra storage is slower than the original disk until the next maintenance update. The slower disk can reduce performance for I/O-intensive workloads. * Maximum storage depends on the plan, service type, and cloud provider. It can be up to five times the plan's base storage size. * Cloud providers limit how many times you can increase storage between maintenance updates. If you reach the limit, run a maintenance update to optimize performance. * You cannot add storage during a maintenance update. * Dynamic disk sizing (DDS) is not supported on custom service plans. Pricing If you add storage when you create a service, **Additional disk storage** shows an estimated monthly cost. The **Service summary** lists plan storage plus additional storage. The estimated monthly price includes the additional storage cost. If you add storage to a running service, the Aiven Console shows the cost of the additional storage and related backups. The same costs appear on your invoices. ## Add or remove storage[​](#add-or-remove-storage "Direct link to Add or remove storage") ### Add storage when you create a service[​](#add-storage-when-you-create-a-service "Direct link to Add storage when you create a service") To add storage while you create a service: 1. In **Additional disk storage**, set the size with the slider or enter a value in GB. 2. Review the estimated monthly cost. 3. In the **Service summary**, click **Create service**. Change additional storage later on the running service, or enable automatic disk scaling with Aiven Autoscaler. ### Change storage on a running service[​](#change-storage-on-a-running-service "Direct link to Change storage on a running service") You cannot add or remove storage when service nodes are in the rebuilding state, for example during a maintenance update or a service upgrade. If you are removing disk storage: * Make sure the data in your service does not exceed the allocated storage. If it does, you cannot remove the additional storage. * Plan for the time it takes to rebuild the service. The time depends on the service. - Console - CLI - Terraform 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Actions** > **Manage additional storage**. 3. Change the disk storage. note * The price shown for the additional storage includes backup costs. * You can only remove storage that you previously added using this feature. To downgrade further, you can change your service plan. 4. Click **Save Changes**. Use [Aiven CLI](/docs/tools/cli.md) to add or remove additional storage using [`avn service update`](/docs/tools/cli/service-cli.md#avn-cli-service-update) with the `--disk-space-gib` flag to specify the total disk space to provide to your service. For example, if your service has a 80-GiB disk and you would like to add an extra 10-GiB disk, use: ``` avn service update --disk-space-gib 90 --project PROJECT_NAME SERVICE_NAME ``` note * When you perform a horizontal service upgrade or downgrade, remember to include all additional disks the service uses. For example, when switching from `Startup-4` to `Business-4` or from `Business-4` to `Startup-4`, include all the additional disks available for this service. * When you fork an existing service, include all additional disks the service uses. Use the `additional_disk_space` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). If you added storage, the additional storage is available immediately. If you removed additional storage, the service nodes go through a rolling restart. Depending on the service type and configuration, there might be a short downtime for services with no HA capabilities. note Storage optimization is performed at the next maintenance update after a change to the storage size. Due to cloud provider limitations, there is a limit on how many times storage can be increased between two maintenance updates. When this limit is reached, perform a maintenance update for performance optimization. Plan increases to avoid reaching this limit. Related pages * [Disk autoscaler](/docs/platform/howto/disk-autoscaler.md) * [Change a service plan](/docs/platform/howto/scale-services.md) --- # Attach VPCs to AWS Transit Gateway [AWS Transit Gateway (TGW)](https://aws.amazon.com/transit-gateway/) enables transitive routing from on-premises networks through VPN and from other VPC. By creating a Transit Gateway VPC attachment, services in an Aiven VPC can route traffic to all other networks attached, directly or indirectly, to the Transit Gateway. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * AWS Transit Gateway * Aiven VPC in the same region as your Transit Gateway * Tools * [Aiven Console](https://console.aiven.io/) * [AWS Management Console](https://console.aws.amazon.com) * [Aiven CLI](/docs/tools/cli.md) * [AWS CLI](https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html) ## Prepare for an attachment[​](#prepare-for-an-attachment "Direct link to Prepare for an attachment") 1. Find your AWS account ID and AWS Transit Gateway ID, for example, in the [AWS Management Console](https://console.aws.amazon.com). The AWS account ID and the AWS Transit Gateway ID will be referred to as `USER_ACCOUNT_ID` and `USER_TGW_ID`, respectively. 2. Share the TGW with the Aiven AWS account using one of the following: * [AWS Resource Access Manager](https://console.aws.amazon.com/ram/home) in [AWS Management Console](https://console.aws.amazon.com) * [AWS CLI](https://aws.amazon.com/cli/) command [`create-resource-share`](https://docs.aws.amazon.com/cli/latest/reference/ram/create-resource-share) important * Add the Transit Gateway as a shared resource. * Add Aiven AWS account ID `675999398324` as a principal. 3. Find your Aiven VPC ID using either the [Aiven Console](https://console.aiven.io/) or the [Aiven CLI](/docs/tools/cli/vpc.md#list-vpcs). Depending on the type of your Aiven VPC, select **Project VPC** or **Organization VPC**: * Project VPC * Organization VPC - In the [Aiven CLI](/docs/tools/cli.md), run the [avn vpc list](/docs/tools/cli/vpc.md) command. - In the [Aiven Console](https://console.aiven.io/): 1. Go to your organization, and open your project's page. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, select your project VPC. 4. On the **VPC details** page, go to the **Overview** section, and copy **ID**. * In the [Aiven CLI](/docs/tools/cli.md), run the [avn organization vpc list](/docs/tools/cli/vpc.md) command. * In the [Aiven Console](https://console.aiven.io/): 1. Go to your organization, and click **Admin** in the top navigation bar. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, select your organization VPC. 4. On the **VPC details** page, go to the **Overview** section, and copy **ID**. The Aiven VPC ID will be referred to as `VPC_ID`. 4. If your Aiven VPC is an [organization VPC](/docs/platform/concepts/vpcs.md#vpc-types), find your organization ID. Otherwise, skip this step. [Find your organization ID in the Aiven Console](/docs/platform/reference/get-resource-IDs.md#get-an-organization-id) or retrieve your organization ID from the output of the `avn organization list` command. The organization ID, if applicable, will be referred to as `ORG_ID`. 5. Determine the IP ranges to route from the VPC to the AWS Transit Gateway. note A Transit Gateway has a route table of its own and, by default, routes traffic to each attached network (directly to an attached VPC or indirectly through VPN attachments). Attached route tables of the VPC need to be updated to include the TGW as a target for any IP range (CIDR) to be routed using the VPC attachment. The IP ranges will be referred to as `USER_PEER_NETWORK_CIDR`. ## Request the attachment[​](#request-the-attachment "Direct link to Request the attachment") ### Send the connection request[​](#send-the-connection-request "Direct link to Send the connection request") To create the Transit Gateway - VPC attachment, make a request to the Aiven API for a peering connection. note In the request, you can use the `--user-peer-network-cidr` argument multiple times to define more than one peer network CIDR. You can also create the attachment without a CIDR, but you would need to add it later to make your attachment operational. Depending on the type of your Aiven VPC, select **Project VPC** or **Organization VPC**, and run the provided commands, replacing placeholders with meaningful values. * Project VPC * Organization VPC ``` avn vpc peering-connection create \ --project-vpc-id VPC_ID \ --peer-cloud-account USER_ACCOUNT_ID \ --peer-vpc USER_TGW_ID \ --user-peer-network-cidr USER_PEER_NETWORK_CIDR ``` ``` avn organization vpc peering-connection create \ --organization-id ORG_ID \ --organization-vpc-id VPC_ID \ --peer-cloud-account USER_ACCOUNT_ID \ --peer-vpc USER_TGW_ID \ --user-peer-network-cidr USER_PEER_NETWORK_CIDR ``` Once submitted, the request gets the `APPROVED` status, and the Aiven Platform starts building the connection by creating the AWS Transit Gateway VPC attachment, which can take a few minutes. Once the connection is built, the status changes to `PENDING_PEER`. ### Check the connection status[​](#check-the-connection-status "Direct link to Check the connection status") To check the status of your peering connection request, select either **Project VPC** or **Organization VPC**, depending on your Aiven VPC type, and run the provided commands, replacing placeholders with meaningful values. * Project VPC * Organization VPC ``` avn vpc peering-connection get \ --project-vpc-id VPC_ID \ --peer-cloud-account USER_ACCOUNT_ID \ --peer-vpc USER_TGW_ID -v ``` ``` avn organization vpc list \ --organization-id ORG_ID \ --organization-vpc-id VPC_ID ``` The status that allows you to move on to accepting the attachment is `PENDING_PEER`. ## Accept the attachment[​](#accept-the-attachment "Direct link to Accept the attachment") Once the status is `PENDING_PEER`, accept the attachment in your AWS account. 1. Log in to the [AWS Management Console](https://console.aws.amazon.com), and go to [Transit gateway attachments](https://console.aws.amazon.com/vpc/home#TransitGatewayAttachments). 2. In the **Transit gateway attachments** list, find and select your pending attachment. 3. Click **Actions** > **Accept request**. The Aiven Platform monitors the attachment until it is accepted. Once it detects that the request is accepted, the status changes to `ACTIVE`. This indicates that your attachment is operational, the VPC route table is updated to route `USER_PEER_NETWORK_CIDR` to the Transit Gateway, and service nodes in the VPC have open firewall access to those networks. Related pages [AWS VPC peering connections](https://docs.aws.amazon.com/vpc/latest/peering/what-is-vpc-peering) --- # Bring your own key (BYOK) Register, list, update, or delete your customer managed keys (CMKs), associate CMKs with services, and view CMK usage across services in Aiven projects using the [Aiven Provider for Terraform](/docs/tools/terraform.md), [Aiven API](/docs/tools/api.md), or the [Aiven CLI](/docs/tools/cli.md). important Bring your own key (BYOK) is a [BYOC](/docs/platform/concepts/byoc.md) enterprise feature, available for newly created BYOC services only. [Contact Aiven](https://aiven.io/contact) to request access. ## Encryption scope[​](#encryption-scope "Direct link to Encryption scope") BYOK encrypts the following using your CMKs: * **Backups**: All backups created by Aiven services are encrypted with your CMK. * **Service data at rest**: CMKs protect all data stored by the service. * **Data in transit between the service and backups**: Encryption occurs on the service node before data leaves the cluster, so backup transfers use your CMK. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * **Key management service** (KMS) that supports customer-managed keys in one of the supported cloud providers: * **Google Cloud KMS**: asymmetric RSA 2048 or RSA 4096 keys * **Oracle Cloud Infrastructure (OCI) Vault**: AES keys * **AWS KMS**: symmetric encryption keys (`ENCRYPT_DECRYPT`) * **Azure Key Vault**: RSA keys (software-protected or HSM-backed) * **Public internet access** to your KMS or key vault: Aiven connects to your key management service over the public internet to perform wrap and unwrap operations, so the key management service must be reachable from the public internet. Key vaults that block public internet access aren't currently supported. * [**Authentication token**](/docs/platform/howto/create_authentication_token.md) to use the [Aiven API](/docs/tools/api.md) * [**Aiven CLI**](/docs/tools/cli.md) installed and configured (for CLI instructions) * [**Aiven Provider for Terraform**](/docs/tools/terraform.md) installed and configured (for Terraform instructions) * More information on the `aiven_cmk` resource and its configuration options are available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/cmk). ## List CMK accessors[​](#list-cmk-accessors "Direct link to List CMK accessors") List customer managed key (CMK) accessors - principals that need to be granted access to perform encrypt/decrypt operations on your behalf. * API * CLI * Terraform #### API endpoint[​](#api-endpoint "Direct link to API endpoint") `GET /v1/project/PROJECT_ID/secrets/cmks/accessors` Reference: [CMKAccessorsList API](https://api.aiven.io/doc/#tag/Secrets/operation/CMKAccessorsList) #### Path parameters[​](#path-parameters "Direct link to Path parameters") | Parameter | Type | Required | Description | | ------------ | ------ | -------- | ------------------ | | `PROJECT_ID` | String | True | Project identifier | #### Sample request[​](#sample-request "Direct link to Sample request") ``` curl -X GET https://api.aiven.io/v1/project/PROJECT_ID/secrets/cmks/accessors \ -H "Authorization: Bearer AIVEN_API_TOKEN" ``` #### Sample response[​](#sample-response "Direct link to Sample response") A successful request returns a `200 OK` status code and a JSON object with the accessors for each provider, for example: ``` { "accessors": { "gcp": { "access_group": "access.example.12345678-1234-1234-1234-123456789abc@aiven.io" }, "oci": { "access_group": "ocid1.group.oc1..abcdABCD....", "access_tenant": "ocid1.tenancy.oc1..abcdABCD...." }, "aws": { "role_arn": "arn:aws:iam::123456789012:role/aiven-cmk-anchor" }, "azure": { "app_id": "12345678-1234-1234-1234-123456789abc" } } } ``` #### Command[​](#command "Direct link to Command") ``` avn project cmks accessors --project PROJECT_NAME ``` #### Parameters[​](#parameters "Direct link to Parameters") | Parameter | Type | Required | Description | | ----------- | ------ | -------- | ------------ | | `--project` | String | True | Project name | #### Sample request[​](#sample-request-1 "Direct link to Sample request") ``` avn project cmks accessors --project my-project ``` #### Sample output[​](#sample-output "Direct link to Sample output") The output is always in the JSON format: ``` { "gcp": { "access_group": "access.example.12345678-1234-1234-1234-123456789abc@aiven.io" }, "oci": { "access_group": "ocid1.group.oc1..abcdABCD....", "access_tenant": "ocid1.tenancy.oc1..abcdABCD...." }, "aws": { "role_arn": "arn:aws:iam::123456789012:role/aiven-cmk-anchor" }, "azure": { "app_id": "12345678-1234-1234-1234-123456789abc" } } ``` #### Data sources[​](#data-sources "Direct link to Data sources") Use the per-provider CMK accessor data source for the cloud provider hosting your KMS. Each data source takes the project name and returns the accessor values as read-only attributes: * [aiven\_cmk\_accessor\_gcp](https://registry.terraform.io/providers/aiven/aiven/latest/docs/data-sources/cmk_accessor_gcp): returns `access_group` for Google Cloud KMS. * [aiven\_cmk\_accessor\_oci](https://registry.terraform.io/providers/aiven/aiven/latest/docs/data-sources/cmk_accessor_oci): returns `access_group` and `access_tenant` for OCI Vault. * [aiven\_cmk\_accessor\_aws](https://registry.terraform.io/providers/aiven/aiven/latest/docs/data-sources/cmk_accessor_aws): returns `principal`, the IAM role ARN, for AWS KMS. * [aiven\_cmk\_accessor\_azure](https://registry.terraform.io/providers/aiven/aiven/latest/docs/data-sources/cmk_accessor_azure): returns `app_id` for Azure Key Vault. #### Sample configuration (Google Cloud)[​](#sample-configuration-google-cloud "Direct link to Sample configuration (Google Cloud)") ``` data "aiven_cmk_accessor_gcp" "example" { project = "my-project" } output "gcp_access_group" { value = data.aiven_cmk_accessor_gcp.example.access_group } ``` #### Sample configuration (OCI)[​](#sample-configuration-oci "Direct link to Sample configuration (OCI)") ``` data "aiven_cmk_accessor_oci" "example" { project = "my-project" } output "oci_access_group" { value = data.aiven_cmk_accessor_oci.example.access_group } output "oci_access_tenant" { value = data.aiven_cmk_accessor_oci.example.access_tenant } ``` #### Sample configuration (AWS)[​](#sample-configuration-aws "Direct link to Sample configuration (AWS)") ``` data "aiven_cmk_accessor_aws" "example" { project = "my-project" } output "aws_principal" { value = data.aiven_cmk_accessor_aws.example.principal } ``` #### Sample configuration (Azure)[​](#sample-configuration-azure "Direct link to Sample configuration (Azure)") ``` data "aiven_cmk_accessor_azure" "example" { project = "my-project" } output "azure_app_id" { value = data.aiven_cmk_accessor_azure.example.app_id } ``` note Use the accessor values returned by this operation when granting Aiven access to your key: * **Google Cloud KMS**: Grant the `access_group` email address the `roles/cloudkms.cryptoOperator` role on your key. * **OCI Vault**: Use `access_tenant` and `access_group` OCIDs to create cross-tenancy IAM policies. * **AWS KMS**: Use the `role_arn` as the trusted principal in your KMS key policy (see [AWS KMS setup](#aws-kms-setup)). * **Azure Key Vault**: Use the `app_id` to create a service principal in your Azure AD tenant (see [Azure Key Vault setup](#azure-key-vault-setup)). ## Set up customer-managed keys on your cloud provider[​](#set-up-customer-managed-keys-on-your-cloud-provider "Direct link to Set up customer-managed keys on your cloud provider") Before registering a CMK with Aiven, set up the key and grant Aiven access on your cloud provider. ### Google Cloud KMS setup[​](#google-cloud-kms-setup "Direct link to Google Cloud KMS setup") #### Create a key ring[​](#create-a-key-ring "Direct link to Create a key ring") ``` gcloud kms keyrings create \ --location \ --project ``` #### Create a CryptoKey[​](#create-a-cryptokey "Direct link to Create a CryptoKey") ``` gcloud kms keys create \ --location \ --keyring \ --purpose encryption \ --project ``` For HSM-backed keys: ``` gcloud kms keys create \ --location \ --keyring \ --purpose encryption \ --protection-level hsm \ --project ``` Record the key resource name: ``` projects//locations//keyRings//cryptoKeys/ ``` #### Grant Aiven access to your key[​](#grant-aiven-access-to-your-key "Direct link to Grant Aiven access to your key") 1. Get Aiven's access group email using [List CMK accessors](#list-cmk-accessors). 2. Grant the `Cloud KMS CryptoKey Encrypter/Decrypter` role to Aiven's group: ``` gcloud kms keys add-iam-policy-binding \ --location \ --keyring \ --project \ --member "group:@aiven.io" \ --role "roles/cloudkms.cryptoKeyEncrypterDecrypter" ``` ### Oracle Cloud Infrastructure (OCI) Vault setup[​](#oracle-cloud-infrastructure-oci-vault-setup "Direct link to Oracle Cloud Infrastructure (OCI) Vault setup") OCI key validation can fail with a generic error when the key region is not available for BYOK. If key validation fails after you confirm the key OCID and IAM policy, contact Aiven support. #### Create cross-tenancy IAM policies[​](#create-cross-tenancy-iam-policies "Direct link to Create cross-tenancy IAM policies") 1. Get Aiven's tenancy OCID and group OCID using [List CMK accessors](#list-cmk-accessors). Use OCI `access_tenant` for `` and OCI `access_group` for ``. 2. Create the policy in the root compartment of your tenancy. 3. Create cross-tenancy IAM policies in your tenancy to grant Aiven access to the key: ``` oci iam policy create \ --compartment-id \ --name aiven-cmk-access \ --statements '[ "define tenancy AT as ", "define group AG as ", "admit group AG of tenancy AT to use keys in tenancy" ]' ``` All three statements are required. Do not remove the `define group` statement. Optional: Restrict access to a specific key by adding a condition to the `admit` statement: ``` oci iam policy create \ --compartment-id \ --name aiven-cmk-access \ --statements '[ "define tenancy AT as ", "define group AG as ", "admit group AG of tenancy AT to use keys in tenancy where target.key.id = \"\"" ]' ``` #### Create a Vault[​](#create-a-vault "Direct link to Create a Vault") ``` oci kms management vault create \ --compartment-id \ --display-name \ --vault-type DEFAULT ``` For HSM-backed vaults: ``` oci kms management vault create \ --compartment-id \ --display-name \ --vault-type VIRTUAL_PRIVATE ``` Record the Vault's management endpoint and crypto endpoint. #### Create a Master Encryption Key[​](#create-a-master-encryption-key "Direct link to Create a Master Encryption Key") ``` oci kms management key create \ --compartment-id \ --display-name \ --endpoint \ --key-shape '{"algorithm": "AES", "length": 32}' ``` Record the key OCID: ``` ocid1.key.oc1.. ``` ### AWS KMS setup[​](#aws-kms-setup "Direct link to AWS KMS setup") Aiven authenticates to your AWS KMS key using cross-account IAM access. You grant access by adding Aiven's IAM role ARN as a trusted principal in your KMS key policy. No resources need to be created in Aiven's AWS account. Everything is controlled through your key policy. #### Step 1: Create a KMS key[​](#step-1-create-a-kms-key "Direct link to Step 1: Create a KMS key") Create a symmetric encryption key in the AWS region where your Aiven services will run: ``` aws kms create-key \ --description "Aiven CMK for data-at-rest encryption" \ --key-usage ENCRYPT_DECRYPT \ --origin AWS_KMS ``` Record the key ARN from the output: ``` arn:aws:kms:::key/ ``` #### Step 2: Create a key alias (optional but recommended)[​](#step-2-create-a-key-alias-optional-but-recommended "Direct link to Step 2: Create a key alias (optional but recommended)") ``` aws kms create-alias \ --alias-name alias/aiven-cmk \ --target-key-id ``` #### Step 3: Grant Aiven access via the key policy[​](#step-3-grant-aiven-access-via-the-key-policy "Direct link to Step 3: Grant Aiven access via the key policy") Get Aiven's IAM role ARN using [List CMK accessors](#list-cmk-accessors) (`role_arn`), then update your KMS key policy to allow Aiven to perform encrypt and decrypt operations. The key policy must include the following statement: ``` { "Sid": "Allow Aiven to use this key for CMK operations", "Effect": "Allow", "Principal": { "AWS": "" }, "Action": [ "kms:Encrypt", "kms:Decrypt" ], "Resource": "*" } ``` The full key policy must also retain the root account statement so that your IAM policies can still manage the key: ``` { "Version": "2012-10-17", "Statement": [ { "Sid": "Enable IAM policies for key management", "Effect": "Allow", "Principal": { "AWS": "arn:aws:iam:::root" }, "Action": "kms:*", "Resource": "*" }, { "Sid": "Allow Aiven to use this key for CMK operations", "Effect": "Allow", "Principal": { "AWS": "" }, "Action": [ "kms:Encrypt", "kms:Decrypt" ], "Resource": "*" } ] } ``` Apply the key policy: ``` aws kms put-key-policy \ --key-id \ --policy-name default \ --policy file://key-policy.json ``` note You can revoke Aiven's access at any time by removing the Aiven principal from the key policy, or by disabling or deleting the key. All KMS operations performed by Aiven are logged in your AWS CloudTrail, giving you a full audit trail. ### Azure Key Vault setup[​](#azure-key-vault-setup "Direct link to Azure Key Vault setup") Aiven authenticates to your Azure Key Vault using a multi-tenant application registered in Aiven's Azure AD tenant. You grant access by creating a service principal from Aiven's application in your tenant and assigning it the **Key Vault Crypto User** role on your key. #### Step 1: Create a service principal for Aiven's application[​](#step-1-create-a-service-principal-for-aivens-application "Direct link to Step 1: Create a service principal for Aiven's application") Get Aiven's `app_id` using [List CMK accessors](#list-cmk-accessors), then register Aiven's application as a service principal in your Azure AD tenant: ``` az ad sp create --id ``` This allows Aiven to authenticate into your Azure AD tenant using its multi-tenant application. #### Step 2: Create a Key Vault with Azure RBAC[​](#step-2-create-a-key-vault-with-azure-rbac "Direct link to Step 2: Create a Key Vault with Azure RBAC") If this is the first Key Vault in your subscription, register the `Microsoft.KeyVault` resource provider first: ``` az provider register --namespace Microsoft.KeyVault --wait ``` The Key Vault must use **Azure RBAC** for its authorization model (not the legacy access policy model): ``` az keyvault create \ --name \ --resource-group \ --location \ --enable-rbac-authorization true ``` To protect against accidental or malicious key deletion, enable purge protection by adding `--enable-purge-protection true` to the command. If you are using an existing Key Vault, ensure RBAC authorization is enabled: ``` az keyvault update \ --name \ --enable-rbac-authorization true ``` For a newly created Key Vault, assign yourself the **Key Vault Administrator** role so you can create keys: ``` az role assignment create \ --role "Key Vault Administrator" \ --assignee $(az ad signed-in-user show --query id -o tsv) \ --scope $(az keyvault show \ --name --resource-group --query id -o tsv) ``` #### Step 3: Create an RSA key[​](#step-3-create-an-rsa-key "Direct link to Step 3: Create an RSA key") Create an RSA key with the **wrapKey** and **unwrapKey** operations enabled. Software-protected key: ``` az keyvault key create \ --vault-name \ --name \ --kty RSA \ --size 2048 \ --ops wrapKey unwrapKey ``` Dedicated HSM-backed key (for higher security requirements): note HSM-backed (Hardware Security Module) keys require an Azure Managed HSM, which must be provisioned separately before creating keys. Provisioning takes a few minutes and incurs an hourly cost. See the [Azure Managed HSM quickstart](https://learn.microsoft.com/en-us/azure/key-vault/managed-hsm/quick-create-cli) for single-tenant (dedicated) HSM setup instructions. ``` az keyvault key create \ --hsm-name \ --name \ --kty RSA-HSM \ --size 2048 \ --ops wrapKey unwrapKey ``` Supported key sizes are `2048`, `3072`, and `4096`. Record the key URL, which becomes the resource identifier you register with Aiven: * Software-protected key: `https://.vault.azure.net/keys/` * HSM-backed key: `https://.vault.azure.net/keys/` If you include a specific version in the URL, Aiven uses that version exclusively. Omitting the version means Aiven always uses the latest active version. #### Step 4: Grant Aiven access to the key[​](#step-4-grant-aiven-access-to-the-key "Direct link to Step 4: Grant Aiven access to the key") Find the object ID of the Aiven service principal in your tenant: ``` az ad sp show --id --query id -o tsv ``` **Software-protected key**: assign the **Key Vault Crypto User** built-in role to the Aiven service principal. This role grants only key wrapping permissions, so Aiven cannot manage, rotate, or delete your key. Scope to the specific key (recommended): ``` az role assignment create \ --assignee \ --role "Key Vault Crypto User" \ --scope "/subscriptions//resourceGroups//providers/Microsoft.KeyVault/vaults//keys/" ``` Or scope to the entire vault: ``` az role assignment create \ --assignee \ --role "Key Vault Crypto User" \ --scope "/subscriptions//resourceGroups//providers/Microsoft.KeyVault/vaults/" ``` **HSM-backed key**: use the Managed HSM CLI command and assign the **Managed HSM Crypto User** role instead: Scope to the specific key (recommended): ``` az keyvault role assignment create \ --hsm-name \ --role "Managed HSM Crypto User" \ --assignee-object-id \ --scope /keys/ ``` Or scope to all keys in the HSM: ``` az keyvault role assignment create \ --hsm-name \ --role "Managed HSM Crypto User" \ --assignee-object-id \ --scope /keys ``` note You can revoke Aiven's access at any time by removing the role assignment or deleting the service principal from your tenant. This immediately prevents any further cryptographic operations by Aiven. ## Manage a project CMK[​](#manage-a-project-cmk "Direct link to Manage a project CMK") Use the Aiven Provider for Terraform, Aiven API, or Aiven CLI to manage customer managed keys (CMKs) for encrypting service data. For per-service CMK assignment and rotation, see [Manage service CMK associations](#manage-service-cmk-associations). ### Register CMK resource identifier[​](#register-cmk-resource-identifier "Direct link to Register CMK resource identifier") Register a customer managed key resource identifier for an Aiven project. * API * CLI * Terraform #### API endpoint[​](#api-endpoint-1 "Direct link to API endpoint") `POST /v1/project/PROJECT_ID/secrets/cmks` #### Path parameters[​](#path-parameters-1 "Direct link to Path parameters") | Parameter | Type | Required | Description | | ------------ | ------ | -------- | ------------------ | | `PROJECT_ID` | String | True | Project identifier | #### Request body parameters[​](#request-body-parameters "Direct link to Request body parameters") | Parameter | Type | Required | Description | | ------------- | ------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | | `provider` | String | True | Cloud provider hosting the KMS: `gcp`, `oci`, `aws`, or `azure` | | `resource` | String | True | CMK reference (key identifier of max 512 characters). Format depends on provider: GCP resource name, OCI OCID, AWS KMS key ARN, or Azure Key Vault key URL | | `default_cmk` | Boolean | False | Mark this key as default for new service creation | #### Sample request (Google Cloud)[​](#sample-request-google-cloud "Direct link to Sample request (Google Cloud)") ``` curl -X POST https://api.aiven.io/v1/project/PROJECT_ID/secrets/cmks \ -H "Content-Type: application/json" \ -H "Authorization: Bearer AIVEN_API_TOKEN" \ -d '{ "provider": "gcp", "resource": "projects/aiven-example/locations/us-central1/keyRings/example-keyring/cryptoKeys/example-key", "default_cmk": true }' ``` #### Sample response (Google Cloud)[​](#sample-response-google-cloud "Direct link to Sample response (Google Cloud)") A successful request returns a `201 CREATED` status code and a JSON object representing the newly registered CMK configuration, for example: ``` { "cmk": { "id": "12345678-1234-1234-1234-12345678abcd", "provider": "gcp", "default_cmk": true, "resource": "projects/aiven-example/locations/us-central1/keyRings/example-keyring/cryptoKeys/example-key", "status": "current", "created_at": "YYYY-MM-DDTHH:MM:SSZ", "updated_at": "YYYY-MM-DDTHH:MM:SSZ" } } ``` #### Sample request (AWS)[​](#sample-request-aws "Direct link to Sample request (AWS)") ``` curl -X POST https://api.aiven.io/v1/project/PROJECT_ID/secrets/cmks \ -H "Content-Type: application/json" \ -H "Authorization: Bearer AIVEN_API_TOKEN" \ -d '{ "provider": "aws", "resource": "arn:aws:kms:us-east-1:123456789012:key/mrk-1234abcd12ab34cd56ef1234567890ab", "default_cmk": true }' ``` #### Sample response (AWS)[​](#sample-response-aws "Direct link to Sample response (AWS)") ``` { "cmk": { "id": "12345678-1234-1234-1234-12345678abcd", "provider": "aws", "default_cmk": true, "resource": "arn:aws:kms:us-east-1:123456789012:key/mrk-1234abcd12ab34cd56ef1234567890ab", "status": "current", "created_at": "YYYY-MM-DDTHH:MM:SSZ", "updated_at": "YYYY-MM-DDTHH:MM:SSZ" } } ``` #### Sample request (Azure)[​](#sample-request-azure "Direct link to Sample request (Azure)") ``` curl -X POST https://api.aiven.io/v1/project/PROJECT_ID/secrets/cmks \ -H "Content-Type: application/json" \ -H "Authorization: Bearer AIVEN_API_TOKEN" \ -d '{ "provider": "azure", "resource": "https://my-vault.vault.azure.net/keys/my-cmk-key", "default_cmk": true }' ``` #### Sample response (Azure)[​](#sample-response-azure "Direct link to Sample response (Azure)") ``` { "cmk": { "id": "12345678-1234-1234-1234-12345678abcd", "provider": "azure", "default_cmk": true, "resource": "https://my-vault.vault.azure.net/keys/my-cmk-key", "status": "current", "created_at": "YYYY-MM-DDTHH:MM:SSZ", "updated_at": "YYYY-MM-DDTHH:MM:SSZ" } } ``` #### Response fields[​](#response-fields "Direct link to Response fields") | Parameter | Description | | ------------- | --------------------------------------------------------- | | `id` | Identifier of the specific key | | `provider` | Provider type: `gcp`, `oci`, `aws`, or `azure` | | `resource` | CMK reference | | `status` | One of `current`, `old`, or `deleted` | | `default_cmk` | Whether this CMK has been marked default for new services | | `created_at` | CMK creation timestamp | | `updated_at` | CMK update timestamp | #### Command[​](#command-1 "Direct link to Command") ``` avn project cmks create --project PROJECT_NAME --provider PROVIDER --resource RESOURCE ``` #### Parameters[​](#parameters-1 "Direct link to Parameters") | Parameter | Type | Required | Description | | --------------- | ------ | -------- | ------------------------------------------------------------------------------------------------------------------- | | `--project` | String | True | Project name | | `--provider` | String | True | Cloud provider hosting the KMS: `gcp`, `oci`, `aws`, or `azure` | | `--resource` | String | True | CMK reference. Format depends on provider: GCP resource name, OCI OCID, AWS KMS key ARN, or Azure Key Vault key URL | | `--default-cmk` | Flag | False | Mark this key as default for new service creation | | `--json` | Flag | False | Output in JSON format | #### Sample request (Google Cloud)[​](#sample-request-google-cloud-1 "Direct link to Sample request (Google Cloud)") ``` avn project cmks create \ --project my-project \ --provider gcp \ --resource "projects/aiven-example/locations/us-central1/keyRings/example-keyring/cryptoKeys/example-key" \ --default-cmk ``` #### Sample output (Google Cloud)[​](#sample-output-google-cloud "Direct link to Sample output (Google Cloud)") Table format: ``` property value ============ ================================================ id 12345678-1234-1234-1234-12345678abcd provider gcp default_cmk True resource projects/aiven-example/locations/us-central1/keyRings/example-keyring/cryptoKeys/example-key status current created_at YYYY-MM-DDTHH:MM:SSZ updated_at YYYY-MM-DDTHH:MM:SSZ ``` #### Sample request (AWS)[​](#sample-request-aws-1 "Direct link to Sample request (AWS)") ``` avn project cmks create \ --project my-project \ --provider aws \ --resource "arn:aws:kms:us-east-1:123456789012:key/mrk-1234abcd12ab34cd56ef1234567890ab" \ --default-cmk ``` #### Sample request (Azure)[​](#sample-request-azure-1 "Direct link to Sample request (Azure)") ``` avn project cmks create \ --project my-project \ --provider azure \ --resource "https://my-vault.vault.azure.net/keys/my-cmk-key" \ --default-cmk ``` For JSON output, use `--json` flag. #### Resource[​](#resource "Direct link to Resource") Use the [aiven\_cmk](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/cmk) resource to register a CMK with your Aiven project. #### Sample configuration[​](#sample-configuration "Direct link to Sample configuration") ``` Loading... ``` ### Update CMK[​](#update-cmk "Direct link to Update CMK") Update attributes or parameters on an existing customer managed key configuration. * API * CLI * Terraform #### API endpoint[​](#api-endpoint-2 "Direct link to API endpoint") `POST /v1/project/PROJECT_ID/secrets/cmks/CMK_ID` #### Path parameters[​](#path-parameters-2 "Direct link to Path parameters") | Parameter | Type | Required | Description | | ------------ | ------ | -------- | ------------------ | | `PROJECT_ID` | String | True | Project identifier | | `CMK_ID` | String | True | CMK identifier | #### Request body parameters[​](#request-body-parameters-1 "Direct link to Request body parameters") | Parameter | Type | Required | Description | | ------------- | ------- | -------- | ------------------------------------------------------- | | `default_cmk` | Boolean | False | Mark a specific key as default for new service creation | #### Sample request[​](#sample-request-2 "Direct link to Sample request") ``` curl -X POST https://api.aiven.io/v1/project/PROJECT_ID/secrets/cmks/CMK_ID \ -H "Content-Type: application/json" \ -H "Authorization: Bearer AIVEN_API_TOKEN" \ -d '{ "default_cmk": false }' ``` #### Sample response[​](#sample-response-1 "Direct link to Sample response") A successful request returns a `200 OK` status code and a JSON object representing the updated CMK configuration, for example: ``` { "cmk": { "id": "12345678-1234-1234-1234-12345678abcd", "provider": "gcp", "default_cmk": false, "resource": "projects/aiven-example/locations/us-central1/keyRings/example-keyring/cryptoKeys/example-key", "status": "current", "created_at": "YYYY-MM-DDTHH:MM:SSZ", "updated_at": "YYYY-MM-DDTHH:MM:SSZ" } } ``` #### Response fields[​](#response-fields-1 "Direct link to Response fields") | Parameter | Description | | ------------- | --------------------------------------------------------- | | `id` | Identifier of the specific key | | `provider` | Provider type: `gcp`, `oci`, `aws`, or `azure` | | `resource` | CMK reference | | `status` | One of `current`, `old`, or `deleted` | | `default_cmk` | Whether this CMK has been marked default for new services | | `created_at` | CMK creation timestamp | | `updated_at` | CMK update timestamp | #### Command[​](#command-2 "Direct link to Command") ``` avn project cmks update --project PROJECT_NAME --cmk-id CMK_ID --default-cmk true ``` #### Parameters[​](#parameters-2 "Direct link to Parameters") | Parameter | Type | Required | Description | | --------------- | ------ | -------- | -------------------------------------------------------------------------- | | `--project` | String | True | Project name | | `--cmk-id` | String | True | CMK identifier | | `--default-cmk` | String | False | Mark a specific key as default for new service creation: `true` or `false` | | `--json` | Flag | False | Output in JSON format | #### Sample request[​](#sample-request-3 "Direct link to Sample request") ``` avn project cmks update \ --project my-project \ --cmk-id 12345678-1234-1234-1234-12345678abcd \ --default-cmk false ``` #### Sample output[​](#sample-output-1 "Direct link to Sample output") Table format: ``` property value ============ ================================================ id 12345678-1234-1234-1234-12345678abcd provider gcp default_cmk False resource projects/aiven-example/locations/us-central1/keyRings/example-keyring/cryptoKeys/example-key status current created_at YYYY-MM-DDTHH:MM:SSZ updated_at YYYY-MM-DDTHH:MM:SSZ ``` For JSON output, use `--json` flag. #### Resource[​](#resource-1 "Direct link to Resource") Update the [aiven\_cmk](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/cmk) resource configuration to modify the CMK settings. note You can only update the `default_cmk` attribute. Changing the `project`, `cmk_provider`, or `resource` attributes forces the recreation of the resource. #### Sample configuration[​](#sample-configuration-1 "Direct link to Sample configuration") Update the `default_cmk` setting: ``` resource "aiven_cmk" "example" { project = "my-project" cmk_provider = "gcp" resource = "projects/aiven-example/locations/us-central1/keyRings/example-keyring/cryptoKeys/example-key" - default_cmk = true + default_cmk = false } ``` ### Get CMK details[​](#get-cmk-details "Direct link to Get CMK details") Get the details of a customer managed key configuration. * API * CLI * Terraform #### API endpoint[​](#api-endpoint-3 "Direct link to API endpoint") `GET /v1/project/PROJECT_ID/secrets/cmks/CMK_ID` #### Path parameters[​](#path-parameters-3 "Direct link to Path parameters") | Parameter | Type | Required | Description | | ------------ | ------ | -------- | ------------------ | | `PROJECT_ID` | String | True | Project identifier | | `CMK_ID` | String | True | CMK identifier | #### Sample request[​](#sample-request-4 "Direct link to Sample request") ``` curl -X GET https://api.aiven.io/v1/project/PROJECT_ID/secrets/cmks/CMK_ID \ -H "Authorization: Bearer AIVEN_API_TOKEN" ``` #### Sample response (OCI)[​](#sample-response-oci "Direct link to Sample response (OCI)") A successful request returns a `200 OK` status code and a JSON object representing the specified CMK configuration, for example: ``` { "cmk": { "id": "a1b2c3d4-e5f6-4789-a0b1-c2d3e4f5a6b7", "provider": "oci", "default_cmk": false, "resource": "ocid1.key.oc1.iad.xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx", "status": "current", "created_at": "YYYY-MM-DDTHH:MM:SSZ", "updated_at": "YYYY-MM-DDTHH:MM:SSZ" } } ``` #### Sample response (AWS)[​](#sample-response-aws-1 "Direct link to Sample response (AWS)") ``` { "cmk": { "id": "a1b2c3d4-e5f6-4789-a0b1-c2d3e4f5a6b7", "provider": "aws", "default_cmk": false, "resource": "arn:aws:kms:us-east-1:123456789012:key/mrk-1234abcd12ab34cd56ef1234567890ab", "status": "current", "created_at": "YYYY-MM-DDTHH:MM:SSZ", "updated_at": "YYYY-MM-DDTHH:MM:SSZ" } } ``` #### Sample response (Azure)[​](#sample-response-azure-1 "Direct link to Sample response (Azure)") ``` { "cmk": { "id": "a1b2c3d4-e5f6-4789-a0b1-c2d3e4f5a6b7", "provider": "azure", "default_cmk": false, "resource": "https://my-vault.vault.azure.net/keys/my-cmk-key", "status": "current", "created_at": "YYYY-MM-DDTHH:MM:SSZ", "updated_at": "YYYY-MM-DDTHH:MM:SSZ" } } ``` #### Command[​](#command-3 "Direct link to Command") ``` avn project cmks get --project PROJECT_NAME --cmk-id CMK_ID -v ``` For JSON output, use `--json` flag: ``` avn project cmks get --project PROJECT_NAME --cmk-id CMK_ID -v [--json] ``` #### Parameters[​](#parameters-3 "Direct link to Parameters") | Parameter | Type | Required | Description | | ----------- | ------ | -------- | --------------------- | | `--project` | String | True | Project name | | `--cmk-id` | String | True | CMK identifier | | `--json` | Flag | False | Output in JSON format | #### Sample request[​](#sample-request-5 "Direct link to Sample request") ``` avn project cmks get \ --project my-project \ --cmk-id a1b2c3d4-e5f6-4789-a0b1-c2d3e4f5a6b7 \ -v ``` #### Sample output (OCI)[​](#sample-output-oci "Direct link to Sample output (OCI)") Table format: ``` ID PROVIDER RESOURCE DEFAULT_CMK STATUS CREATED_AT UPDATED_AT ==================================== ======== ========================================================= =========== ======= =========================== =========================== 12345678-1234-1234-1234-12345678abcd oci ocid1.key.oc1.iad.xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx false current 2026-05-05T10:55:52.719109Z 2026-05-05T13:04:59.294447Z Associated services: SERVICE_NAME STATUS ============ ====== my-pg active ``` JSON format: ``` { "created_at": "2026-05-05T10:55:52.719109Z", "default_cmk": false, "id": "12345678-1234-1234-1234-12345678abcd", "provider": "oci", "resource": "ocid1.key.oc1.iad.xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx", "service_associations": [ { "service_name": "my-pg", "status": "active" } ], "status": "current", "updated_at": "2026-05-05T13:04:59.294447Z" } ``` #### Resource[​](#resource-2 "Direct link to Resource") Use [aiven\_cmk](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/cmk) outputs to get CMK details. #### Accessing CMK details from resource[​](#accessing-cmk-details-from-resource "Direct link to Accessing CMK details from resource") ``` output "cmk_id" { value = aiven_cmk.example.cmk_id } output "cmk_status" { value = aiven_cmk.example.status } output "cmk_created_at" { value = aiven_cmk.example.created_at } ``` #### Importing existing CMK[​](#importing-existing-cmk "Direct link to Importing existing CMK") Import an existing CMK to manage it with Terraform: ``` terraform import aiven_cmk.example PROJECT_NAME/CMK_ID ``` ### List CMKs[​](#list-cmks "Direct link to List CMKs") List all customer managed key configurations for a project. * API * CLI * Terraform #### API endpoint[​](#api-endpoint-4 "Direct link to API endpoint") `GET /v1/project/PROJECT_ID/secrets/cmks` #### Path parameters[​](#path-parameters-4 "Direct link to Path parameters") | Parameter | Type | Required | Description | | ------------ | ------ | -------- | ------------------ | | `PROJECT_ID` | String | True | Project identifier | #### Sample request[​](#sample-request-6 "Direct link to Sample request") ``` curl -X GET https://api.aiven.io/v1/project/PROJECT_ID/secrets/cmks \ -H "Authorization: Bearer AIVEN_API_TOKEN" ``` #### Sample response (OCI)[​](#sample-response-oci-1 "Direct link to Sample response (OCI)") A successful request returns a `200 OK` status code and a JSON object containing a list of all CMK configurations for a project, for example: ``` { "cmks": [ { "id": "a1b2c3d4-e5f6-4789-a0b1-c2d3e4f5a6b7", "provider": "oci", "default_cmk": false, "resource": "ocid1.key.oc1.iad.xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx", "status": "current", "created_at": "YYYY-MM-DDTHH:MM:SSZ", "updated_at": "YYYY-MM-DDTHH:MM:SSZ" }, { "id": "8c9d0e1f-2a3b-4c5d-6e7f-8a9b0c1d2e3f", "provider": "oci", "default_cmk": false, "resource": "ocid1.key.oc1.phx.yyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyy", "status": "old", "created_at": "YYYY-MM-DDTHH:MM:SSZ", "updated_at": "YYYY-MM-DDTHH:MM:SSZ" } ] } ``` #### Command[​](#command-4 "Direct link to Command") ``` avn project cmks list --project PROJECT_NAME ``` #### Parameters[​](#parameters-4 "Direct link to Parameters") | Parameter | Type | Required | Description | | ----------- | ------ | -------- | --------------------- | | `--project` | String | True | Project name | | `--json` | Flag | False | Output in JSON format | #### Sample request[​](#sample-request-7 "Direct link to Sample request") ``` avn project cmks list --project my-project ``` #### Sample output (OCI)[​](#sample-output-oci-1 "Direct link to Sample output (OCI)") Table format: ``` id status provider resource default_cmk created_at ==================================== ======= ======== ======================================================== =========== =================== a1b2c3d4-e5f6-4789-a0b1-c2d3e4f5a6b7 current oci ocid1.key.oc1.iad.xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx False YYYY-MM-DDTHH:MM:SSZ 8c9d0e1f-2a3b-4c5d-6e7f-8a9b0c1d2e3f old oci ocid1.key.oc1.phx.yyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyyy False YYYY-MM-DDTHH:MM:SSZ ``` For JSON output, use `--json` flag. Use Terraform state to list all CMKs managed in your [CMK resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/cmk): ``` terraform state list | grep aiven_cmk ``` View details of a specific CMK: ``` terraform state show aiven_cmk.cmk_oci_iad ``` note To list all CMKs in a project (including those not managed by Terraform), use the API or CLI. ### Remove CMK[​](#remove-cmk "Direct link to Remove CMK") Delete a customer managed key configuration. note You can delete a CMK only when it has no service associations in `active`, `activating`, or `deactivating` status. Move each linked service to another CMK or remove the CMK association before deletion. * API * CLI * Terraform #### API endpoint[​](#api-endpoint-5 "Direct link to API endpoint") `DELETE /v1/project/PROJECT_ID/secrets/cmks/CMK_ID` #### Path parameters[​](#path-parameters-5 "Direct link to Path parameters") | Parameter | Type | Required | Description | | ------------ | ------ | -------- | ------------------ | | `PROJECT_ID` | String | True | Project identifier | | `CMK_ID` | String | True | CMK identifier | #### Sample request[​](#sample-request-8 "Direct link to Sample request") ``` curl -X DELETE https://api.aiven.io/v1/project/PROJECT_ID/secrets/cmks/CMK_ID \ -H "Authorization: Bearer AIVEN_API_TOKEN" ``` #### Sample response (OCI)[​](#sample-response-oci-2 "Direct link to Sample response (OCI)") A successful request returns a `200 OK` status code and a JSON object representing the deleted CMK configuration, for example: ``` { "cmk": { "id": "a1b2c3d4-e5f6-4789-a0b1-c2d3e4f5a6b7", "provider": "oci", "default_cmk": false, "resource": "ocid1.key.oc1.iad.xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx", "status": "current", "created_at": "YYYY-MM-DDTHH:MM:SSZ", "updated_at": "YYYY-MM-DDTHH:MM:SSZ" } } ``` If the CMK is still associated with one or more services, the request returns `409 Conflict`. ``` { "message": "CMK cannot be deleted because it is still associated with one or more services" } ``` #### Command[​](#command-5 "Direct link to Command") ``` avn project cmks delete --project PROJECT_NAME --cmk-id CMK_ID ``` #### Parameters[​](#parameters-5 "Direct link to Parameters") | Parameter | Type | Required | Description | | ----------- | ------ | -------- | --------------------- | | `--project` | String | True | Project name | | `--cmk-id` | String | True | CMK identifier | | `--json` | Flag | False | Output in JSON format | #### Sample request[​](#sample-request-9 "Direct link to Sample request") ``` avn project cmks delete \ --project my-project \ --cmk-id a1b2c3d4-e5f6-4789-a0b1-c2d3e4f5a6b7 ``` #### Sample output (OCI)[​](#sample-output-oci-2 "Direct link to Sample output (OCI)") The command returns the details of the deleted CMK in the format you specified. For JSON output, use `--json` flag. #### Resource[​](#resource-3 "Direct link to Resource") warning Deleting a CMK fails when the CMK is still associated with one or more services. Move each linked service to another CMK or remove the CMK association first. Remove or comment out the [CMK resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/cmk) in your Terraform configuration: ``` # resource "aiven_cmk" "example" { # project = "my-project" # cmk_provider = "oci" # resource = "ocid1.key.oc1.iad.xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx" # default_cmk = false # } ``` Alternatively, use `terraform destroy` to target a specific resource: ``` terraform destroy -target=aiven_cmk.example ``` #### What happens[​](#what-happens "Direct link to What happens") When you remove a [CMK resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/cmk): * Terraform apply fails if the CMK is still associated with one or more services. * After all associations are removed, the CMK configuration is deleted from the Aiven project. * The resource is removed from Terraform state. ## Manage service CMK associations[​](#manage-service-cmk-associations "Direct link to Manage service CMK associations") Associate a specific customer managed key (CMK) with individual services during creation or update. This allows you to use different CMKs for different services, change CMKs for existing services, or remove CMK associations altogether. ### Associate a CMK when creating a service[​](#associate-a-cmk-when-creating-a-service "Direct link to Associate a CMK when creating a service") Create a service with a specific CMK by providing the CMK ID in the service creation request. Set `cloud` to control the cloud region for the new service. * API * CLI * Terraform #### API endpoint[​](#api-endpoint-6 "Direct link to API endpoint") `POST /v1/project/PROJECT_ID/service` #### Request body parameters[​](#request-body-parameters-2 "Direct link to Request body parameters") | Parameter | Type | Required | Description | | -------------------------------------------------------------------------- | ------ | -------- | --------------------------------------------------------------- | | `service_name` | String | True | Name of the service | | `service_type` | String | True | Type of service (for example, `pg`, `mysql`, `redis`) | | `plan` | String | True | Service plan | | `cloud` | String | False | Cloud region for the service, for example `google-europe-west3` | | `cmk_id` | String | False | Customer managed key (CMK) identifier. If omitted, the | | project's default CMK is used, or Aiven-managed keys if no default is set. | | | | #### Sample request[​](#sample-request-10 "Direct link to Sample request") ``` curl -X POST https://api.aiven.io/v1/project/PROJECT_ID/service \ -H "Content-Type: application/json" \ -H "Authorization: Bearer AIVEN_API_TOKEN" \ -d '{ "service_name": "my-pg-service", "service_type": "pg", "plan": "startup-4", "cloud": "google-europe-west3", "cmk_id": "12345678-1234-1234-1234-12345678abcd" }' ``` #### Sample response[​](#sample-response-2 "Direct link to Sample response") A successful request returns a `201 CREATED` status code and a JSON object representing the newly created service with the CMK association: ``` { "service": { "service_name": "my-pg-service", "service_type": "pg", "plan": "startup-4", "cloud_name": "google-europe-west3", "state": "REBUILDING", "cmk_id": "12345678-1234-1234-1234-12345678abcd" } } ``` #### Command[​](#command-6 "Direct link to Command") ``` avn service create \ --project PROJECT_NAME \ --service-type SERVICE_TYPE \ --plan PLAN_NAME \ --cloud CLOUD_NAME \ --cmk-id CMK_ID \ SERVICE_NAME ``` #### Parameters[​](#parameters-6 "Direct link to Parameters") | Parameter | Type | Required | Description | | ---------------- | ------ | -------- | ----------------------------------------------------- | | `--project` | String | True | Project name | | `--service-type` | String | True | Type of service (for example, `pg`, `mysql`, `redis`) | | `--plan` | String | True | Service plan | | `--cloud` | String | False | Cloud region for the service | | `--cmk-id` | String | False | Customer managed key (CMK) identifier | #### Sample request[​](#sample-request-11 "Direct link to Sample request") ``` avn service create \ --project my-project \ --service-type pg \ --plan startup-4 \ --cloud google-europe-west3 \ --cmk-id 12345678-1234-1234-1234-12345678abcd \ my-pg-service ``` #### Resource[​](#resource-4 "Direct link to Resource") Use the `cmk_id` parameter in the service resource to associate a CMK at creation time. #### Sample configuration[​](#sample-configuration-2 "Direct link to Sample configuration") ``` resource "aiven_cmk" "example_key" { project = "my-project" cmk_provider = "gcp" resource = "projects/aiven-example/locations/us-central1/keyRings/example-keyring/cryptoKeys/example-key" default_cmk = false } resource "aiven_pg" "example" { project = "my-project" service_name = "my-pg-service" plan = "startup-4" cloud_name = "google-europe-west3" cmk_id = aiven_cmk.example_key.cmk_id } ``` ### Change or remove the CMK for an existing service[​](#change-or-remove-the-cmk-for-an-existing-service "Direct link to Change or remove the CMK for an existing service") Update a service to use a different CMK or remove its CMK association. * API * CLI * Terraform #### API endpoint[​](#api-endpoint-7 "Direct link to API endpoint") `PUT /v1/project/PROJECT_ID/service/SERVICE_NAME` #### Path parameters[​](#path-parameters-6 "Direct link to Path parameters") | Parameter | Type | Required | Description | | -------------- | ------ | -------- | ------------------ | | `PROJECT_ID` | String | True | Project identifier | | `SERVICE_NAME` | String | True | Service name | #### Request body parameters[​](#request-body-parameters-3 "Direct link to Request body parameters") | Parameter | Type | Required | Description | | --------- | ------ | -------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cmk_id` | String | False | Customer managed key (CMK) identifier to use for this service. Pass an empty UUID (`00000000-0000-0000-0000-000000000000`) to remove the CMK association and use Aiven-managed keys instead. | #### Sample request (change CMK)[​](#sample-request-change-cmk "Direct link to Sample request (change CMK)") ``` curl -X PUT https://api.aiven.io/v1/project/PROJECT_ID/service/SERVICE_NAME \ -H "Content-Type: application/json" \ -H "Authorization: Bearer AIVEN_API_TOKEN" \ -d '{ "cmk_id": "87654321-4321-4321-4321-87654321dcba" }' ``` #### Sample request (remove CMK association)[​](#sample-request-remove-cmk-association "Direct link to Sample request (remove CMK association)") ``` curl -X PUT https://api.aiven.io/v1/project/PROJECT_ID/service/SERVICE_NAME \ -H "Content-Type: application/json" \ -H "Authorization: Bearer AIVEN_API_TOKEN" \ -d '{ "cmk_id": "00000000-0000-0000-0000-000000000000" }' ``` #### Sample response[​](#sample-response-3 "Direct link to Sample response") A successful request returns a `200 OK` status code and a JSON object representing the updated service: ``` { "service": { "service_name": "my-pg-service", "service_type": "pg", "cloud_name": "google-europe-west3", "state": "REBALANCING", "cmk_id": "87654321-4321-4321-4321-87654321dcba" } } ``` #### Command[​](#command-7 "Direct link to Command") ``` avn service update --project PROJECT_NAME --cmk-id CMK_ID SERVICE_NAME ``` #### Parameters[​](#parameters-7 "Direct link to Parameters") | Parameter | Type | Required | Description | | ----------- | ------ | -------- | ----------------------------------------------------------------------------------------------------------------- | | `--project` | String | True | Project name | | `--cmk-id` | String | False | Customer managed key (CMK) identifier. Pass `00000000-0000-0000-0000-000000000000` to remove the CMK association. | #### Sample request (change CMK)[​](#sample-request-change-cmk-1 "Direct link to Sample request (change CMK)") ``` avn service update \ --project my-project \ --cmk-id 87654321-4321-4321-4321-87654321dcba \ my-pg-service ``` #### Sample request (remove CMK association)[​](#sample-request-remove-cmk-association-1 "Direct link to Sample request (remove CMK association)") ``` avn service update \ --project my-project \ --cmk-id 00000000-0000-0000-0000-000000000000 \ my-pg-service ``` #### Resource[​](#resource-5 "Direct link to Resource") Update the `cmk_id` parameter in your service resource to change or remove CMK association. #### Sample configuration (change CMK)[​](#sample-configuration-change-cmk "Direct link to Sample configuration (change CMK)") ``` resource "aiven_cmk" "example_key_new" { project = "my-project" cmk_provider = "gcp" resource = "projects/aiven-example/locations/us-central1/keyRings/example-keyring/cryptoKeys/example-key-v2" default_cmk = false } resource "aiven_pg" "example" { project = "my-project" service_name = "my-pg-service" plan = "startup-4" cloud_name = "google-europe-west3" cmk_id = aiven_cmk.example_key_new.cmk_id } ``` #### Sample configuration (remove CMK)[​](#sample-configuration-remove-cmk "Direct link to Sample configuration (remove CMK)") ``` resource "aiven_pg" "example" { project = "my-project" service_name = "my-pg-service" plan = "startup-4" cloud_name = "google-europe-west3" cmk_id = "00000000-0000-0000-0000-000000000000" } ``` ### View CMK details in service information[​](#view-cmk-details-in-service-information "Direct link to View CMK details in service information") When you retrieve service information, the response now includes the `cmk_id` field showing which CMK is actively protecting that service's data. This allows you to verify encryption key usage and track which services are using which CMKs. * API * CLI #### API endpoint[​](#api-endpoint-8 "Direct link to API endpoint") `GET /v1/project/PROJECT_ID/service/SERVICE_NAME` #### Sample request[​](#sample-request-12 "Direct link to Sample request") ``` curl -X GET https://api.aiven.io/v1/project/PROJECT_ID/service/SERVICE_NAME \ -H "Authorization: Bearer AIVEN_API_TOKEN" ``` #### Sample response[​](#sample-response-4 "Direct link to Sample response") The service details now include the `cmk_id` field: ``` { "service": { "service_name": "my-pg-service", "service_type": "pg", "plan": "startup-4", "cloud_name": "google-europe-west3", "state": "RUNNING", "cmk_id": "12345678-1234-1234-1234-12345678abcd" } } ``` If the service is not using a CMK, the `cmk_id` field is `null`: ``` { "service": { "service_name": "my-mysql-service", "service_type": "mysql", "plan": "startup-4", "cloud_name": "google-europe-west3", "state": "RUNNING", "cmk_id": null } } ``` #### Command[​](#command-8 "Direct link to Command") ``` avn service get --project PROJECT_NAME SERVICE_NAME ``` #### Sample output[​](#sample-output-2 "Direct link to Sample output") The output is always in the JSON format: ``` { "service": { "service_name": "my-pg-service", "service_type": "pg", "plan": "startup-4", "cloud_name": "google-europe-west3", "state": "RUNNING", "cmk_id": "12345678-1234-1234-1234-12345678abcd" } } ``` If no CMK is associated with the service, the `cmk_id` value is empty or `null` in JSON output: ``` { "service": { "service_name": "my-mysql-service", "service_type": "mysql", "plan": "startup-4", "cloud_name": "google-europe-west3", "state": "RUNNING", "cmk_id": null } } ``` ### List services associated with a CMK[​](#list-services-associated-with-a-cmk "Direct link to List services associated with a CMK") Find all services in a project that are using a specific customer managed key. This is useful for auditing, capacity planning, or managing CMK usage across your infrastructure. * API * CLI * Terraform #### API endpoint[​](#api-endpoint-9 "Direct link to API endpoint") `GET /v1/project/PROJECT_ID/secrets/cmks/CMK_ID/service_associations` #### Path parameters[​](#path-parameters-7 "Direct link to Path parameters") | Parameter | Type | Required | Description | | ------------ | ------ | -------- | ------------------ | | `PROJECT_ID` | String | True | Project identifier | | `CMK_ID` | String | True | CMK identifier | #### Sample request[​](#sample-request-13 "Direct link to Sample request") ``` curl -X GET https://api.aiven.io/v1/project/PROJECT_ID/secrets/cmks/CMK_ID/service_associations \ -H "Authorization: Bearer AIVEN_API_TOKEN" ``` #### Sample response[​](#sample-response-5 "Direct link to Sample response") A successful request returns a `200 OK` status code and a JSON object containing a list of services associated with the CMK: ``` { "service_associations": [ { "service_name": "my-pg-service", "status": "active" }, { "service_name": "my-kafka-cluster", "status": "activating" } ] } ``` #### Response fields[​](#response-fields-2 "Direct link to Response fields") | Field | Description | | -------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | | `service_name` | Name of the service using this CMK | | `status` | Association status: `active` (service is actively using the CMK), or `activating` (service is in the process of transitioning to use this CMK) | The response includes services currently associated with the CMK, including services in transition (`activating`) and services already using the CMK (`active`). Services not associated with the specified CMK are not included. note If a CMK has no associated services, the `service_associations` array is empty. #### Command[​](#command-9 "Direct link to Command") The service-association lookup for a specific CMK is currently available through the API endpoint only. #### Resource[​](#resource-6 "Direct link to Resource") The Aiven Provider for Terraform does not currently provide a data source or resource for listing services associated with a specific CMK. Use the API endpoint for this operation. --- # Manage customer contacts for a custom cloud Update the list of customer contacts for your [custom cloud](/docs/platform/concepts/byoc.md). With the [BYOC feature enabled](/docs/platform/howto/byoc/enable-byoc.md), you can [create custom clouds](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organizations. While [creating a custom cloud](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md), you add at least the **Admin** contact so that the Aiven team can reach out to them if needed. You can change the provided contacts any time later by following [Update the contacts list](#update-the-contacts-list). important While you can add multiple different customer contacts for your custom cloud, **Admin** is a mandatory role that is always required as a primary support contact. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven Console * Aiven CLI - At least one [custom cloud created](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organization - Access to the [Aiven Console](https://console.aiven.io/) - [Organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role in your Aiven organization * At least one [custom cloud created](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organization * [Aiven CLI client](/docs/tools/cli.md) installed * [Organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role in your Aiven organization ## Update the contacts list[​](#update-the-contacts-list "Direct link to Update the contacts list") * Aiven Console * Aiven CLI 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your organization. 2. Click **Admin** in the top navigation, and click **Bring your own cloud** in the sidebar. 3. In the **Bring your own cloud** view, select a cloud. 4. On the selected cloud's page, click **Actions** > **Customer contact**. 5. In the **Customer contact** window, select a new contact's role from the menu, enter the email address, and click to add the provided contact's details. 6. When you're done adding all the contacts, select **Save changes**. Use the [avn byoc update](/docs/tools/cli/byoc.md#avn-byoc-update) command with the `--contact-email` option to edit the list of individuals from your organization to be contacted by the Aiven team if needed. ``` avn byoc update \ --organization-id "ORGANIZATION_IDENTIFIER" \ --byoc-id "CUSTOM_CLOUD_IDENTIFIER" \ --contact-email 'email="EMAIL_ADDRESS",real_name="John Doe",role="Admin"' ``` Each `--contact-email` value takes the `email`, `real_name`, and `role` fields in the format `email="EMAIL",real_name="NAME",role="ROLE"`. The `email` field is required, and all values must be quoted. To set multiple contacts, repeat the option: ``` avn byoc update \ --organization-id "ORGANIZATION_IDENTIFIER" \ --byoc-id "CUSTOM_CLOUD_IDENTIFIER" \ --contact-email 'email="admin@example.com",real_name="John Doe",role="Admin"' \ --contact-email 'email="ops@example.com",real_name="Jane Smith",role="Operator"' ``` note The `--contact-email` option replaces the entire list of customer contacts. List every contact, including the mandatory **Admin** contact, each time you run the command. Any contact you omit is removed. Related pages * [View the status of a custom cloud](/docs/platform/howto/byoc/view-custom-cloud-status.md) * [Assign a project to your custom cloud](/docs/platform/howto/byoc/assign-project-custom-cloud.md) * [Rename a custom cloud](/docs/platform/howto/byoc/rename-custom-cloud.md) * [Tag custom cloud resources](/docs/platform/howto/byoc/tag-custom-cloud-resources.md) --- # Enable a custom cloud in Aiven organizations, units, or projects Select your organizations, units, or projects that can access and use your [custom cloud](/docs/platform/concepts/byoc.md). With the [BYOC feature enabled](/docs/platform/howto/byoc/enable-byoc.md), you can [create custom clouds](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organization. As a part of the [initial custom cloud's setup](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md), you select in what projects you'll be able to use your new custom cloud to host Aiven services. You can update this setting any time later by following by following [Enable projects to use your custom cloud](#enable-projects-to-use-your-custom-cloud). You can make your custom cloud available by default in all existing and future projects, or you can pick specific projects and/or organizational units where you want your custom cloud to be available. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven Console * Aiven CLI - At least one [custom cloud created](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organization - Access to the [Aiven Console](https://console.aiven.io/) - [Organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role in your Aiven organization * At least one [custom cloud created](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organization * [Aiven CLI client](/docs/tools/cli.md) installed * [Organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role in your Aiven organization ## Enable projects to use your custom cloud[​](#enable-projects-to-use-your-custom-cloud "Direct link to Enable projects to use your custom cloud") * Aiven Console * Aiven CLI 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your organization. 2. Click **Admin** in the top navigation, and click **Bring your own cloud** in the sidebar. 3. In the **Bring your own cloud** view, select a cloud. 4. On the selected cloud's page, go to the **Available projects** tab, and modify the settings: * **Set availability** 1. Click **Set availability** to decide if your custom cloud is available in all projects in your organization or in selected projects/units only. 2. In the **Custom cloud's availability in your organization** window, choose between: * **By default for all projects** * **By selection** * **Assign organizational units**: select organizational units * **Assign projects**: select projects note By selecting an organizational unit, you make your custom cloud available from all the projects in this unit. 3. Click **Save**. * **Assign projects** 1. Click **Assign projects** to enable your custom cloud in specific organizational units and/or projects. 2. In the **Assign projects** window, use the available menus to select units and/or projects. 3. Click **Assign projects**. Use the [`avn byoc cloud permissions add`](/docs/tools/cli/byoc.md#avn-byoc-cloud-permissions-add) command to enable your custom cloud in organizations, projects, or units. ``` avn byoc cloud permissions add \ --organization-id "ORGANIZATION_IDENTIFIER" \ --byoc-id "CUSTOM_CLOUD_IDENTIFIER" \ --account "ACCOUNT_IDENTIFIER" ``` In the organizations, projects, or organizational units for which you enable your custom cloud, you can: * Create services in the custom cloud * Migrate existing services to your custom cloud if your service and networking configuration allows it. For more information, contact your account team. Related pages * [View the status of a custom cloud](/docs/platform/howto/byoc/view-custom-cloud-status.md) * [Add customer's contact information for your custom cloud](/docs/platform/howto/byoc/add-customer-info-custom-cloud.md) * [Rename a custom cloud](/docs/platform/howto/byoc/rename-custom-cloud.md) * [Tag custom cloud resources](/docs/platform/howto/byoc/tag-custom-cloud-resources.md) --- # Use AWS PrivateLink with BYOC services [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Enable and manage AWS PrivateLink for your Aiven services deployed in your own cloud using [bring your own cloud (BYOC)](/docs/platform/concepts/byoc.md). ## Limitations[​](#limitations "Direct link to Limitations") * AWS PrivateLink for Aiven BYOC is a feature with [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-). * You can enable AWS PrivateLink for your Aiven BYOC service only if it uses the [AWS BYOC private deployment](/docs/platform/concepts/byoc.md#byoc-architecture) model. * To use [cross-region connections](/docs/platform/howto/use-aws-privatelinks.md#allow-cross-region-connections) with a BYOC service, [set up the required permissions](#set-up-permissions) first. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Aiven CLI](/docs/tools/cli.md) * [AWS CLI](https://aws.amazon.com/cli/) * [Terraform](https://developer.hashicorp.com/terraform) * Optionally, access to: * [Aiven Console](https://console.aiven.io/) * [AWS Management Console](https://console.aws.amazon.com) ## Set up permissions[​](#set-up-permissions "Direct link to Set up permissions") Use your Terraform template to grant Aiven the permissions to set up and manage secure PrivateLink connections within your AWS environment for your BYOC service. To add the required permissions to your own AWS account: 1. [Download the latest version of your Terraform template](/docs/platform/howto/byoc/download-infrastructure-template.md). 2. [Apply the updated template in your own AWS account using Terraform](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md#deploy-the-template). ## Enable AWS PrivateLink[​](#enable-aws-privatelink "Direct link to Enable AWS PrivateLink") See [Use AWS PrivateLink with Aiven services](/docs/platform/howto/use-aws-privatelinks.md#enable-aws-privatelink). ## Allow cross-region connections[​](#allow-cross-region-connections "Direct link to Allow cross-region connections") See [Allow cross-region connections](/docs/platform/howto/use-aws-privatelinks.md#allow-cross-region-connections). Related pages * [Create a custom cloud (BYOC environment in Aiven)](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md#create-a-custom-cloud) * [Download an infrastructure template and a variables file](/docs/platform/howto/byoc/download-infrastructure-template.md) --- # Create an AWS-integrated custom cloud Create a [custom cloud](/docs/platform/concepts/byoc.md) for BYOC in your Aiven organization to better address your specific business needs or project requirements. To configure a custom cloud in your Aiven organization and prepare your AWS account so that Aiven can access it: 1. In the Aiven Console or with the Aiven CLI client, you specify new cloud details to generate a Terraform infrastructure-as-code template. 2. You download the generated template and deploy it in your AWS account to acquire IAM Role ARN (Amazon Resource Name). 3. You deploy your custom cloud resources supplying the acquired IAM Role ARN to the Aiven platform, which gives Aiven the permissions to securely access your AWS account, create resources, and manage them onward. 4. You select projects that can use your new custom clouds for creating services. 5. You add contact details for individuals from your organization that Aiven can reach out to in case of technical issues with the new cloud. ## Before you start[​](#before-you-start "Direct link to Before you start") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * You have [enabled the BYOC feature](/docs/platform/howto/byoc/enable-byoc.md). * You have an active account with your cloud provider. * Depending on the tool to use for creating a custom cloud: * Console: Access to the [Aiven Console](https://console.aiven.io/) or * CLI: * [Aiven CLI client](/docs/tools/cli.md) installed * Aiven organization ID from the output of the `avn organization list` command or from the [Aiven Console](https://console.aiven.io/) > **User information** > **Organizations**. * You have the [organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role in your Aiven organization. * You have Terraform installed. * You have required [IAM permissions](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md#iam-permissions). ### IAM permissions[​](#iam-permissions "Direct link to IAM permissions") You need cloud account credentials set up on your machine so that your user or role has required Terraform permissions [to integrate with your cloud provider](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md#create-a-custom-cloud). Show permissions required for creating resources for bastion and workload networks ``` { "Statement": [ { "Action": [ "iam:AttachRolePolicy", "iam:CreateRole", "iam:DeleteRole", "iam:DeleteRolePolicy", "iam:GetRole", "iam:GetRolePolicy", "iam:ListAttachedRolePolicies", "iam:ListInstanceProfilesForRole", "iam:ListRolePolicies", "iam:PutRolePolicy", "iam:TagRole", "iam:UpdateAssumeRolePolicy" ], "Effect": "Allow", "Resource": "arn:aws:iam::*:role/cce-*-iam-role" }, { "Action": [ "ec2:DescribeAddresses", "ec2:DescribeAddressesAttribute", "ec2:DescribeAvailabilityZones", "ec2:DescribeInternetGateways", "ec2:DescribeNatGateways", "ec2:DescribeNetworkInterfaces", "ec2:DescribePrefixLists", "ec2:DescribeRouteTables", "ec2:DescribeSecurityGroups", "ec2:DescribeSecurityGroupRules", "ec2:DescribeStaleSecurityGroups", "ec2:DescribeSubnets", "ec2:DescribeVpcs", "ec2:DescribeVpcEndpoints", "ec2:DescribeVpcAttribute", "ec2:DescribeTags" ], "Effect": "Allow", "Resource": [ "*" ], "Sid": "Describe" }, { "Action": [ "ec2:CreateTags" ], "Condition": { "StringEquals": { "ec2:CreateAction": [ "AllocateAddress", "CreateInternetGateway", "CreateNatGateway", "CreateRoute", "CreateRouteTable", "CreateSecurityGroup", "CreateSubnet", "CreateVpc", "CreateVpcEndpoint" ] } }, "Effect": "Allow", "Resource": [ "*" ], "Sid": "CreateTag" }, { "Action": [ "ec2:DeleteTags" ], "Effect": "Allow", "Resource": [ "arn:aws:ec2:*:*:elastic-ip/*", "arn:aws:ec2:*:*:internet-gateway/*", "arn:aws:ec2:*:*:natgateway/*", "arn:aws:ec2:*:*:route-table/*", "arn:aws:ec2:*:*:security-group/*", "arn:aws:ec2:*:*:security-group-rule/*", "arn:aws:ec2:*:*:subnet/*", "arn:aws:ec2:*:*:vpc/*" ], "Sid": "DeleteTag" }, { "Action": [ "ec2:AllocateAddress", "ec2:CreateInternetGateway", "ec2:CreateVpc" ], "Condition": { "StringLike": { "aws:RequestTag/Name": "cce-*" } }, "Effect": "Allow", "Resource": [ "*" ], "Sid": "Create" }, { "Action": [ "ec2:CreateNatGateway" ], "Condition": { "StringNotLike": { "ec2:ResourceTag/Name": "cce-*" } }, "Effect": "Deny", "Resource": [ "arn:aws:ec2:*:*:elastic-ip/*", "arn:aws:ec2:*:*:subnet/*" ], "Sid": "CreateNGWAllowCCESubnetOnly" }, { "Action": [ "ec2:CreateNatGateway" ], "Condition": { "StringNotLike": { "aws:RequestTag/Name": "cce-*" } }, "Effect": "Deny", "Resource": [ "arn:aws:ec2:*:*:natgateway/*" ], "Sid": "CreateNGWAllowCCEOnly" }, { "Action": [ "ec2:CreateNatGateway" ], "Effect": "Allow", "Resource": [ "arn:aws:ec2:*:*:elastic-ip/*", "arn:aws:ec2:*:*:natgateway/*", "arn:aws:ec2:*:*:subnet/*" ], "Sid": "CreateNGW" }, { "Action": [ "ec2:CreateRouteTable", "ec2:CreateSecurityGroup", "ec2:CreateSubnet" ], "Condition": { "StringNotLike": { "ec2:ResourceTag/Name": "cce-*" } }, "Effect": "Deny", "Resource": [ "arn:aws:ec2:*:*:vpc/*" ], "Sid": "CreateSubAllowCCEVPCOnly" }, { "Action": [ "ec2:CreateRouteTable" ], "Condition": { "StringNotLike": { "aws:RequestTag/Name": "cce-*" } }, "Effect": "Deny", "Resource": [ "arn:aws:ec2:*:*:route-table/*" ], "Sid": "CreateRTAllowCCEOnly" }, { "Action": [ "ec2:CreateRouteTable" ], "Effect": "Allow", "Resource": [ "arn:aws:ec2:*:*:route-table/*", "arn:aws:ec2:*:*:vpc/*" ], "Sid": "CreateRT" }, { "Action": [ "ec2:CreateSecurityGroup" ], "Condition": { "StringNotLike": { "aws:RequestTag/Name": "cce-*" } }, "Effect": "Deny", "Resource": [ "arn:aws:ec2:*:*:security-group/*" ], "Sid": "CreateSGsAllowCCEOnly" }, { "Action": [ "ec2:CreateSecurityGroup" ], "Effect": "Allow", "Resource": [ "arn:aws:ec2:*:*:security-group/*", "arn:aws:ec2:*:*:vpc/*" ], "Sid": "CreateSG" }, { "Action": [ "ec2:CreateSubnet" ], "Condition": { "StringNotLike": { "aws:RequestTag/Name": "cce-*" } }, "Effect": "Deny", "Resource": [ "arn:aws:ec2:*:*:subnet/*" ], "Sid": "CreateSubAllowCCEOnly" }, { "Action": [ "ec2:CreateSubnet" ], "Effect": "Allow", "Resource": [ "arn:aws:ec2:*:*:subnet/*", "arn:aws:ec2:*:*:vpc/*" ], "Sid": "CreateSubnets" }, { "Action": [ "ec2:CreateVpcEndpoint" ], "Effect": "Allow", "Resource": [ "*" ], "Sid": "CreateVpcEndpoint" }, { "Action": [ "ec2:AssociateAddress", "ec2:AssociateRouteTable", "ec2:AssociateSubnetCidrBlock", "ec2:AssociateVpcCidrBlock", "ec2:AssignPrivateNatGatewayAddress", "ec2:AttachInternetGateway", "ec2:AuthorizeSecurityGroupEgress", "ec2:AuthorizeSecurityGroupIngress", "ec2:CreateRoute", "ec2:ModifySecurityGroupRules", "ec2:ModifySubnetAttribute", "ec2:ModifyVpcAttribute", "ec2:ModifyVpcEndpoint", "ec2:ReplaceRoute", "ec2:ReplaceRouteTableAssociation", "ec2:UpdateSecurityGroupRuleDescriptionsEgress", "ec2:UpdateSecurityGroupRuleDescriptionsIngress" ], "Condition": { "StringLike": { "ec2:ResourceTag/Name": "cce-*" } }, "Effect": "Allow", "Resource": [ "*" ], "Sid": "Modify" }, { "Action": [ "ec2:DisassociateAddress" ], "Condition": { "StringNotLike": { "ec2:ResourceTag/Name": "cce-*" } }, "Effect": "Deny", "Resource": [ "arn:aws:ec2:*:*:elastic-ip/*" ], "Sid": "DisassociateEIPAllowCCEOnly" }, { "Action": [ "ec2:DisassociateAddress" ], "Effect": "Allow", "Resource": [ "arn:aws:ec2:*:*:*/*" ], "Sid": "DisassociateEIP" }, { "Action": [ "ec2:DetachInternetGateway", "ec2:DisassociateNatGatewayAddress", "ec2:DisassociateRouteTable", "ec2:DisassociateSubnetCidrBlock", "ec2:DisassociateVpcCidrBlock", "ec2:DeleteInternetGateway", "ec2:DeleteNatGateway", "ec2:DeleteNetworkInterface", "ec2:DeleteRoute", "ec2:DeleteRouteTable", "ec2:DeleteSecurityGroup", "ec2:DeleteSubnet", "ec2:DeleteVpc", "ec2:DeleteVpcEndpoints", "ec2:ReleaseAddress", "ec2:RevokeSecurityGroupEgress", "ec2:RevokeSecurityGroupIngress", "ec2:UnassignPrivateNatGatewayAddress" ], "Condition": { "StringLike": { "ec2:ResourceTag/Name": "cce-*" } }, "Effect": "Allow", "Resource": [ "*" ], "Sid": "Delete" }, { "Action": [ "s3:*" ], "Effect": "Allow", "Resource": [ "arn:aws:s3:::cce-*" ] } ], "Version": "2012-10-17" } ``` ## Create a custom cloud[​](#create-a-custom-cloud "Direct link to Create a custom cloud") Create a custom cloud either in the Aiven Console or with the Aiven CLI. * Aiven Console * Aiven CLI #### Launch the BYOC setup[​](#launch-the-byoc-setup "Direct link to Launch the BYOC setup") 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to an organization. 2. Click **Admin** in the top navigation, and click **Bring your own cloud** in the sidebar. 3. In the **Bring your own cloud** view, select **Create custom cloud**. #### Generate an infrastructure template[​](#generate-an-infrastructure-template "Direct link to Generate an infrastructure template") In this step, an IaC template is generated in the Terraform format. In [the next step](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md#deploy-the-template), you'll deploy this template in your AWS account to acquire Role ARN (Amazon Resource Name), which Aiven needs for accessing your AWS account. In the **Create custom cloud** wizard: 1. Specify cloud details: * Cloud provider * Region * Custom cloud name * [Infrastructure tags](/docs/platform/howto/byoc/tag-custom-cloud-resources.md) 2. Click **Next**. 3. Specify deployment and storage details: * [Deployment model](/docs/platform/concepts/byoc.md#byoc-architecture) Choose a deployment model: * **Private**: Routes traffic through a proxy using a bastion host logically separated from the Aiven services. Provides network isolation but does not add HIPAA or PCI DSS compliance controls. * **Public**: Lets the Aiven control plane connect to the service nodes over the public internet. Use it for public-facing workloads. * **HIPAA**: Builds on the private model for healthcare workloads that handle protected health information (PHI). * **PCI DSS**: Builds on the private model for payment workloads that require cardholder data environment (CDE) isolation. If **HIPAA** or **PCI DSS** options are not visible in the console, these compliance deployment models have not been enabled for your organization. Contact your account team to request access. They require object storage in your AWS account and restrict outbound traffic and public access. See [Enhanced compliance BYOC clouds](/docs/platform/concepts/byoc-enhanced-compliance.md). note The service endpoint has two hostnames. The public hostname is derived from the private hostname by adding a `public-` prefix to it. The Service URI shown by the Aiven Console displays only the private hostname. * CIDR for BYOC resources The **CIDR** block defines the IP address range of the VPC that Aiven creates in your own cloud account. Any Aiven service created in the custom cloud will be placed in the VPC and will get an IP address within this address range. In the **CIDR** field, specify an IP address range for the BYOC VPC using a CIDR block notation, for example: `10.0.0.0/16`, `172.31.0.0/16`, or `192.168.0.0/20`. Make sure that an IP address range you use meets the following requirements: * IP address range is within the private IP address ranges allowed in [RFC 1918](https://datatracker.ietf.org/doc/html/rfc1918). * CIDR block size is between `/16` (65536 IP addresses) and `/24` (256 IP addresses). * CIDR block is large enough to host the desired number of services after splitting it into per-availability-zone subnets. For example, the smallest `/24` CIDR block might be enough for a few services but can pose challenges during node replacements or maintenance upgrades if running low on available free IP addresses. * CIDR block of your BYOC VCP doesn't overlap with the CIDR blocks of VPCs you plan to peer your BYOC VPC with. You cannot change the BYOC VPC CIDR block after your custom cloud is created. * Object storage [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) By default, the following data is stored in the BYOC object storage in your own cloud account: * [Cold data managed by the service](/docs/platform/howto/byoc/store-data.md) * [Backups of the service](/docs/platform/concepts/byoc.md#byoc-service-backups) note * Data is stored in your BYOC object storage using one S3 bucket per custom cloud. * Permissions for S3 bucket management will be included in the Terraform infrastructure template to be generated upon completing this step. 4. Click **Generate template**. Your IaC Terraform template gets generated based on your inputs. You can view, copy, or download it. Now, you can use the template to [acquire Role ARN](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md#deploy-the-template). #### Deploy the template[​](#deploy-the-template "Direct link to Deploy the template") Role ARN is an [identifier of the role](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles.html) created when running the infrastructure template in your AWS account. Aiven uses Role ARN to [assume the role](https://docs.aws.amazon.com/STS/latest/APIReference/API_AssumeRole.html) and run operations such as creating VMs for service nodes in your BYOC account. Use the [generated Terraform template](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md#generate-an-infrastructure-template) to create your Role ARN by deploying the template in your AWS account. 1. Copy or download the template and the variables file from the **Create custom cloud** wizard. 2. Optionally, modify the template as needed. note To connect to a custom-cloud service from different security groups (other than the one dedicated for the custom cloud) or from IP address ranges, add specific ingress rules before you apply a Terraform infrastructure template in your AWS account in the process of creating a custom cloud resources. Before adding ingress rules, see the examples provided in the Terraform template you generated and downloaded from [Aiven Console](https://console.aiven.io/). 3. Set up Terraform to authenticate with AWS. Configure your AWS credentials using one of the following methods: * **Environment variables** (quick setup for testing): ``` export AWS_ACCESS_KEY_ID="your_access_key" export AWS_SECRET_ACCESS_KEY="your_secret_key" export AWS_DEFAULT_REGION="your_region" ``` * **AWS CLI profile** (recommended for local development): 1. Configure credentials using the AWS CLI: ``` aws configure --profile your-profile-name ``` 2. Reference the profile when running Terraform: ``` export AWS_PROFILE=your-profile-name ``` * **IAM roles** (recommended for production and CI/CD environments): If running on an EC2 instance, in AWS CloudShell, or in a CI/CD pipeline, use IAM roles attached to the compute resource instead of static credentials. tip For enhanced security, consider using [aws-vault](https://github.com/99designs/aws-vault) to store encrypted credentials or [AWS Single Sign-On (SSO)](https://docs.aws.amazon.com/cli/latest/userguide/cli-configure-sso.html) for centralized identity management. For more authentication options and configuration details, see the [AWS Provider authentication documentation](https://registry.terraform.io/providers/hashicorp/aws/latest/docs#authentication-and-configuration). 4. Deploy the infrastructure template using Terraform: ``` terraform init terraform plan -var-file=FILE_NAME.tfvars terraform apply -var-file=FILE_NAME.tfvars ``` Replace `FILE_NAME.tfvars` with the name of the variables file you downloaded. important The `-var-file` option is required to pass the configuration variables to Terraform. 5. Find the role identifier (Role ARN) in the Terraform output after running `terraform apply`. 6. Enter Role ARN into the **IAM role ARN** field in the **Create custom cloud** wizard. 7. Click **Next** to proceed or park your cloud setup and save your current configuration as a draft by selecting **Save draft**. You can resume creating your cloud later. #### Set up your custom cloud's availability[​](#set-up-your-custom-clouds-availability "Direct link to Set up your custom cloud's availability") Select projects where the new custom cloud will be available for hosting your services. These projects will support creating new services in the custom cloud or migrating your existing services to the custom cloud if your service and networking configuration allows it. For more information on migrating your existing services to the custom cloud, contact your account team. Your cloud can be available in: * All the projects in your organization * Selected organizational units * Specific projects only To set up your cloud's availability in the **Create custom cloud** wizard > the **Assign BYOC to projects** section, select one of the two following options: * **By default for all projects** to make your custom cloud available in all existing and future projects in the organization * **By selection** to pick specific projects or organizational units where you want your custom cloud to be available. note By selecting an organizational unit, you make your custom cloud available from all the projects in this unit. #### Add customer contacts[​](#add-customer-contacts "Direct link to Add customer contacts") Select at least one person whom Aiven can contact in case of any technical issues with your custom cloud. note **Admin** is a mandatory role, which is required as a primary support contact. In the **Create custom cloud** wizard > the **Customer contacts** section: 1. Select a contact person's role using the **Job title** menu, and provide their email address in the **Email** field. 2. Use **+ Add another contact** to add as many customer contacts as needed for your custom cloud. 3. Click **Save and validate**. The custom cloud process has been initiated. #### Complete the cloud setup[​](#complete-the-cloud-setup "Direct link to Complete the cloud setup") Select **Done** to close the **Create custom cloud** wizard. The deployment of your new custom cloud might take a few minutes. As soon as it's over, and your custom cloud is ready to use, you'll be able to see it on the list of your custom clouds in the **Bring your own cloud** view. note Your new custom cloud is ready to use only after its status changes to **Active**. 1. Generate an infrastructure template by running [avn byoc create](/docs/tools/cli/byoc.md#avn-byoc-create). ``` avn byoc create \ --organization-id "ORGANIZATION_ID" \ --deployment-model "DEPLOYMENT_MODEL_NAME" \ --cloud-provider "aws" \ --cloud-region "CLOUD_REGION_NAME" \ --reserved-cidr "CIDR_BLOCK" \ --display-name "CUSTOM_CLOUD_DISPLAY_NAME" ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization to connect with your own cloud account to create the custom cloud, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md#prerequisites). * `DEPLOYMENT_MODEL_NAME` with the type of [network architecture](/docs/platform/concepts/byoc.md#byoc-architecture) your custom cloud uses: * `standard_public` (public) model: The nodes have public IPs and can be configured to be publicly accessible for authenticated users. The Aiven control plane can connect to the service nodes via the public internet. The service endpoint has two hostnames: a private one and a public one derived from it by adding a `public-` prefix. The Service URI shown in the Aiven Console displays only the private hostname. * `standard` (private) model: The nodes reside in a VPC without public IP addresses and are by default not accessible from outside. Traffic is routed through a proxy for additional security utilizing a bastion host logically separated from the Aiven services. * `hipaa` or `pci_dss` (compliance) models: Build on the `standard` model to run services under HIPAA or PCI DSS requirements. Use `hipaa` for healthcare workloads that handle protected health information (PHI), or `pci_dss` for payment workloads that require cardholder data environment (CDE) isolation. These models must be enabled for your organization before you can use them. Contact your account team to request access. They require object storage in your AWS account and restrict outbound traffic and public access. See [Enhanced compliance BYOC clouds](/docs/platform/concepts/byoc-enhanced-compliance.md). * `CLOUD_REGION_NAME` with the name of an AWS cloud region where to create your custom cloud: 1. Pick a region from the **Cloud** column in the supported [AWS cloud regions](/docs/platform/reference/list_of_clouds.md#amazon-web-services) table. 2. Drop the `aws-` prefix from the selected region name, for example, `aws-eu-north-1` > `eu-north-1`. * `CIDR_BLOCK` with a CIDR block defining the IP address range of the VPC that Aiven creates in your own cloud account, for example: `10.0.0.0/16`, `172.31.0.0/16`, or `192.168.0.0/20`. * `CUSTOM_CLOUD_DISPLAY_NAME` with the name of your custom cloud, which you can set arbitrarily. Show sample output ``` { "custom_cloud_environment": { "cloud_provider": "aws", "cloud_region": "europe-north1", "contact_emails": [ { "email": "firstname.secondname@domain.com", "real_name": "Test User", "role": "Admin" } ], "custom_cloud_environment_id": "018b6442-c602-42bc-b63d-438026133f60", "deployment_model": "standard", "display_name": "My BYOC Cloud on AWS", "errors": [], "reserved_cidr": "10.0.0.0/16", "state": "draft", "tags": {}, "update_time": "2024-05-07T14:24:18Z" } } ``` 2. Deploy the IaC template. 1. Download the template and the variable file: * [avn byoc template terraform get-template](/docs/tools/cli/byoc.md#avn-byoc-template-terraform-get-template) ``` avn byoc template terraform get-template \ --organization-id "ORGANIZATION_ID" \ --byoc-id "CUSTOM_CLOUD_ID" >| "tf_dir/tf_file.tf" ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization to connect with your own cloud account to create the custom cloud, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md#prerequisites). * `CUSTOM_CLOUD_ID` with the identifier of your custom cloud, which you can extract from the output of the [avn byoc list](/docs/tools/cli/byoc.md#avn-byoc-list) command, for example `018b6442-c602-42bc-b63d-438026133f60`. * [avn byoc template terraform get-vars](/docs/tools/cli/byoc.md#avn-byoc-template-terraform-get-vars) ``` avn byoc template terraform get-vars \ --organization-id "ORGANIZATION_ID" \ --byoc-id "CUSTOM_CLOUD_ID" >| "tf_dir/tf_file.tfvars" ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization to connect with your own cloud account to create the custom cloud, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md#prerequisites). * `CUSTOM_CLOUD_ID` with the identifier of your custom cloud, which you can extract from the output of the [avn byoc list](/docs/tools/cli/byoc.md#avn-byoc-list) command, for example `018b6442-c602-42bc-b63d-438026133f60`. 2. Optionally, modify the template as needed. note To connect to a custom-cloud service from different security groups (other than the one dedicated for the custom cloud) or from IP address ranges, add specific ingress rules before you apply a Terraform infrastructure template in your AWS cloud account in the process of creating a custom cloud resources. Before adding ingress rules, see the examples provided in the Terraform template you generated and downloaded from the [Aiven Console](https://console.aiven.io/). 3. Set up Terraform to authenticate with AWS. Configure your AWS credentials using one of the following methods: * **Environment variables** (quick setup for testing): ``` export AWS_ACCESS_KEY_ID="your_access_key" export AWS_SECRET_ACCESS_KEY="your_secret_key" export AWS_DEFAULT_REGION="your_region" ``` * **AWS CLI profile** (recommended for local development): 1. Configure credentials using the AWS CLI: ``` aws configure --profile your-profile-name ``` 2. Reference the profile when running Terraform: ``` export AWS_PROFILE=your-profile-name ``` * **IAM roles** (recommended for production and CI/CD environments): If running on an EC2 instance, in AWS CloudShell, or in a CI/CD pipeline, use IAM roles attached to the compute resource instead of static credentials. tip For enhanced security, consider using [aws-vault](https://github.com/99designs/aws-vault) to store encrypted credentials or [AWS Single Sign-On (SSO)](https://docs.aws.amazon.com/cli/latest/userguide/cli-configure-sso.html) for centralized identity management. For more authentication options and configuration details, see the [AWS Provider authentication documentation](https://registry.terraform.io/providers/hashicorp/aws/latest/docs#authentication-and-configuration). 4. Deploy the infrastructure template using Terraform with the provided variables file: ``` terraform init terraform plan -var-file=FILE_NAME.tfvars terraform apply -var-file=FILE_NAME.tfvars ``` Replace `FILE_NAME.tfvars` with the name of the variables file you downloaded. important The `-var-file` option is required to pass the configuration variables to Terraform. 5. Find `aws-iam-role-arn` in the Terraform output after running `terraform apply`. 3. Provision resources by running [avn byoc provision](/docs/tools/cli/byoc.md#avn-byoc-provision) and passing the generated `aws-iam-role-arn` as an option. ``` avn byoc provision \ --organization-id "ORGANIZATION_ID" \ --byoc-id "CUSTOM_CLOUD_ID" \ --aws-iam-role-arn "GENERATED_ROLE_ARN" ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization to connect with your own cloud account to create the custom cloud, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md#prerequisites). * `CUSTOM_CLOUD_ID` with the identifier of your custom cloud, which you can extract from the output of the [avn byoc list](/docs/tools/cli/byoc.md#avn-byoc-list) command, for example `018b6442-c602-42bc-b63d-438026133f60`. * `GENERATED_ROLE_ARN` with the identifier of the role created when running the infrastructure template in your AWS cloud account. You can extract `GENERATED_ROLE_ARN` from the output of the `terraform apply` command or `terraform output` command. 4. Enable your custom cloud in organizations, projects, or units by running [avn byoc cloud permissions add](/docs/tools/cli/byoc.md#avn-byoc-cloud-permissions-add). ``` avn byoc cloud permissions add \ --organization-id "ORGANIZATION_ID" \ --byoc-id "CUSTOM_CLOUD_ID" \ --account "ACCOUNT_ID" ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization to connect with your own cloud account to create the custom cloud, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md#prerequisites). * `CUSTOM_CLOUD_ID` with the identifier of your custom cloud, which you can extract from the output of the [avn byoc list](/docs/tools/cli/byoc.md#avn-byoc-list) command, for example `018b6442-c602-42bc-b63d-438026133f60`. * `ACCOUNT_ID` with the identifier of your account (organizational unit) in Aiven, for example `a484338c34d7`. You can extract `ACCOUNT_ID` from the output of the `avn organization list` command. 5. Add customer contacts for the new cloud by running [avn byoc update](/docs/tools/cli/byoc.md#avn-byoc-update). ``` avn byoc update \ --organization-id "ORGANIZATION_ID" \ --byoc-id "CUSTOM_CLOUD_ID" \ ' { "contact_emails": [ { "email": "EMAIL_ADDRESS", "real_name": "John Doe", "role": "Admin" } ] } ' ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization to connect with your own cloud account to create the custom cloud, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md#prerequisites). * `CUSTOM_CLOUD_ID` with the identifier of your custom cloud, which you can extract from the output of the [avn byoc list](/docs/tools/cli/byoc.md#avn-byoc-list) command, for example `018b6442-c602-42bc-b63d-438026133f60`. Related pages * [Bring your own cloud networking and security](/docs/platform/howto/byoc/networking-security.md) * [View the status of a custom cloud](/docs/platform/howto/byoc/view-custom-cloud-status.md) * [Manage services hosted in custom clouds](/docs/platform/howto/byoc/manage-byoc-service.md) --- # Create a Microsoft Azure-integrated custom cloud Create a [custom cloud](/docs/platform/concepts/byoc.md) for BYOC in your Aiven organization to better address your specific business needs or project requirements. Azure supports two deployment models: * **Standard** (`standard`): Two separate VNets (Bastion and Workload) connected via VNet peering. Aiven routes management traffic through a bastion host proxy, and workload nodes are not accessible from the public internet. * **Standard public** (`standard_public`): A single Workload VNet with publicly addressed service VMs. Aiven connects to service nodes directly over the public internet. To configure a custom cloud in your Aiven organization and prepare your Azure subscription so that Aiven can access it: 1. In the Aiven Console or with the Aiven CLI client, you specify new cloud details to generate a Terraform infrastructure-as-code template. 2. You download the generated template and deploy it in your Azure subscription using the Azure CLI and Terraform. 3. You provision the custom cloud by supplying your Azure subscription ID and tenant ID to the Aiven platform, which gives Aiven the permissions to access your Azure subscription, create resources, and manage them onward. 4. You select Aiven projects that can use your new custom cloud for creating services. 5. You add contact details for individuals from your organization that Aiven can reach out to in case of technical issues with the new cloud. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * You have [enabled the BYOC feature](/docs/platform/howto/byoc/enable-byoc.md). * You have an active Azure subscription where the BYOC infrastructure will be deployed. * Your Azure identity (user or service principal) has the [required Azure permissions](#azure-permissions). * You have the [organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role in your Aiven organization. * Depending on the tool you use to create the custom cloud: * **Console**: Access to the [Aiven Console](https://console.aiven.io/), or * **CLI**: * [Aiven CLI client](/docs/tools/cli.md) installed * Aiven organization ID from the output of the `avn organization list` command or from the [Aiven Console](https://console.aiven.io/) > **User information** > **Organizations**. * [Azure CLI](https://learn.microsoft.com/en-us/cli/azure/install-azure-cli) (`az`) installed. * Terraform >= 1.0 installed. ## Azure permissions[​](#azure-permissions "Direct link to Azure permissions") To deploy the Aiven BYOC Terraform template, your Azure identity needs the following subscription permissions. Assign them before running `terraform apply`. ### Azure subscription permissions[​](#azure-subscription-permissions "Direct link to Azure subscription permissions") Assign one of the following to your Azure identity on the subscription: * **Owner** built-in role (simplest, but broad), or * A custom role with the minimum permissions defined below. Show minimum custom role permissions for the BYOC deployer ``` { "Actions": [ "Microsoft.Resources/subscriptions/resourceGroups/read", "Microsoft.Resources/subscriptions/resourceGroups/write", "Microsoft.Resources/subscriptions/resourceGroups/delete", "Microsoft.Network/virtualNetworks/read", "Microsoft.Network/virtualNetworks/write", "Microsoft.Network/virtualNetworks/delete", "Microsoft.Network/virtualNetworks/subnets/read", "Microsoft.Network/virtualNetworks/subnets/write", "Microsoft.Network/virtualNetworks/subnets/delete", "Microsoft.Network/virtualNetworks/subnets/join/action", "Microsoft.Network/virtualNetworks/peer/action", "Microsoft.Network/virtualNetworks/virtualNetworkPeerings/read", "Microsoft.Network/virtualNetworks/virtualNetworkPeerings/write", "Microsoft.Network/virtualNetworks/virtualNetworkPeerings/delete", "Microsoft.Network/networkSecurityGroups/read", "Microsoft.Network/networkSecurityGroups/write", "Microsoft.Network/networkSecurityGroups/delete", "Microsoft.Network/networkSecurityGroups/join/action", "Microsoft.Network/networkSecurityGroups/securityRules/read", "Microsoft.Network/networkSecurityGroups/securityRules/write", "Microsoft.Network/networkSecurityGroups/securityRules/delete", "Microsoft.Network/natGateways/read", "Microsoft.Network/natGateways/write", "Microsoft.Network/natGateways/delete", "Microsoft.Network/natGateways/join/action", "Microsoft.Network/publicIPAddresses/read", "Microsoft.Network/publicIPAddresses/write", "Microsoft.Network/publicIPAddresses/delete", "Microsoft.Network/publicIPAddresses/join/action", "Microsoft.Storage/storageAccounts/read", "Microsoft.Storage/storageAccounts/write", "Microsoft.Storage/storageAccounts/delete", "Microsoft.Storage/storageAccounts/listkeys/action", "Microsoft.Storage/storageAccounts/blobServices/read", "Microsoft.Storage/storageAccounts/blobServices/containers/read", "Microsoft.Storage/storageAccounts/blobServices/containers/write", "Microsoft.Storage/storageAccounts/blobServices/containers/delete", "Microsoft.Authorization/roleAssignments/read", "Microsoft.Authorization/roleAssignments/write", "Microsoft.Authorization/roleAssignments/delete", "Microsoft.Authorization/roleDefinitions/read", "Microsoft.Authorization/roleDefinitions/write", "Microsoft.Authorization/roleDefinitions/delete" ], "AssignableScopes": [ "/subscriptions/{subscriptionId}" ], "DataActions": [], "Description": "Minimum permissions for running the Aiven BYOC Azure Terraform template.", "Name": "Aiven BYOC Terraform Operator", "NotActions": [], "NotDataActions": [] } ``` ## Create a custom cloud[​](#create-a-custom-cloud "Direct link to Create a custom cloud") Create a custom cloud either in the Aiven Console or with the Aiven CLI. * Aiven Console * Aiven CLI #### Launch the BYOC setup[​](#launch-the-byoc-setup "Direct link to Launch the BYOC setup") 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to an organization. 2. Click **Admin** in the top navigation, and click **Bring your own cloud** in the sidebar. 3. In the **Bring your own cloud** view, select **Create custom cloud**. #### Generate an infrastructure template[​](#generate-an-infrastructure-template "Direct link to Generate an infrastructure template") In the **Create custom cloud** wizard: 1. Specify cloud details: * **Cloud provider**: Select **Microsoft Azure**. * **Deployment model**: Choose a model: * **Standard**: Two VNets (Bastion and Workload) connected via VNet peering. Workload nodes are not accessible from the public internet. * **Standard public**: A single Workload VNet with publicly addressed service VMs. The service endpoint has two hostnames: a private one and a public one derived from it by adding a `public-` prefix. The Service URI shown in the Aiven Console displays only the private hostname. * **Cloud region**: Select an Azure region, for example `westeurope`. * **CIDR**: Enter an IP address range for the virtual networks Aiven creates in your Azure subscription, for example `10.0.0.0/16`, `172.31.0.0/16`, or `192.168.0.0/20`. * **Display name**: Enter a name for your custom cloud. 2. Click **Next** and review the deployment settings. 3. Click **Next** to generate the template. #### Deploy the template[​](#deploy-the-template "Direct link to Deploy the template") 1. On the **Infrastructure template** page, download the **Infrastructure template** and the **Variables file**. important Do not modify the downloaded files. Changing any parameters, names, or configurations may result in provisioning failures or unexpected behavior. 2. Install the Aiven CCE enterprise application on your Entra tenant: ``` az login --tenant "AZURE_TENANT_ID" az ad sp create --id b40b60e2-10c8-4917-bc74-18a87950e767 ``` Replace `AZURE_TENANT_ID` with your Azure tenant ID. To look it up, run: `az account show --query tenantId -o tsv`. The app ID is also available in the variables file you downloaded as `aiven_cce_client_id`. note Run these commands **once per tenant**, regardless of how many custom clouds you create on the same tenant. If the Aiven CCE enterprise application is already installed on your tenant, skip this step. To remove the Aiven CCE enterprise application from your tenant after you have deleted all custom clouds on it, run: ``` az login --tenant "AZURE_TENANT_ID" az ad sp delete --id b40b60e2-10c8-4917-bc74-18a87950e767 ``` 3. Deploy the infrastructure template using Terraform: ``` terraform init terraform apply -var-file=FILE_NAME.tfvars ``` Replace `FILE_NAME.tfvars` with the name of the variables file you downloaded. 4. In the **Create custom cloud** wizard, enter the identifiers from the Terraform output: * **Subscription ID**: Run `terraform output -raw azure_subscription_id`. * **Tenant ID**: Run `terraform output -raw azure_tenant_id`. 5. Click **Next**. #### Assign to projects and add contacts[​](#assign-to-projects-and-add-contacts "Direct link to Assign to projects and add contacts") 1. Select the projects that can use your new custom cloud, then click **Next**. 2. Add contact details for team members Aiven can reach out to in case of technical issues with the new cloud: * **Email** * **Real name** * **Role** (for example, **Admin**) 3. Click **Create custom cloud**. When your custom cloud's [status is **Active**](/docs/platform/howto/byoc/view-custom-cloud-status.md), it's ready to use. 1. Generate an IaC template by running [avn byoc create](/docs/tools/cli/byoc.md#avn-byoc-create). ``` avn byoc create \ --organization-id "ORGANIZATION_ID" \ --deployment-model "DEPLOYMENT_MODEL" \ --cloud-provider "azure" \ --cloud-region "CLOUD_REGION_NAME" \ --reserved-cidr "CIDR_BLOCK" \ --display-name "CUSTOM_CLOUD_DISPLAY_NAME" ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization to connect with your own cloud account to create the custom cloud, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-azure-custom-cloud.md#prerequisites). * `DEPLOYMENT_MODEL` with the deployment model to use: * `standard`: Two VNets (Bastion and Workload) connected via VNet peering. Workload nodes are not accessible from the public internet. * `standard_public`: A single Workload VNet with publicly addressed service VMs. The service endpoint has two hostnames: a private one and a public one derived from it by adding a `public-` prefix. The Service URI shown in the Aiven Console displays only the private hostname. * `CLOUD_REGION_NAME` with the name of an Azure region where to create your custom cloud: 1. Pick a region from the **Cloud** column in the supported [Azure cloud regions](/docs/platform/reference/list_of_clouds.md#azure) table. 2. Drop the `azure-` prefix from the selected region name, for example, `azure-westeurope` > `westeurope`. * `CIDR_BLOCK` with a CIDR block defining the IP address range for the virtual networks that Aiven creates in your own cloud account, for example: `10.0.0.0/16`, `172.31.0.0/16`, or `192.168.0.0/20`. * `CUSTOM_CLOUD_DISPLAY_NAME` with the name of your custom cloud, which you can set arbitrarily. Show sample output ``` { "custom_cloud_environment": { "cloud_provider": "azure", "cloud_region": "westeurope", "contact_emails": [ { "email": "firstname.secondname@domain.com", "real_name": "Test User", "role": "Admin" } ], "custom_cloud_environment_id": "018b6442-c602-42bc-b63d-438026133f60", "deployment_model": "standard", "display_name": "My BYOC Cloud on Azure", "errors": [], "reserved_cidr": "10.0.0.0/16", "state": "draft", "tags": {}, "update_time": "2024-05-07T14:24:18Z" } } ``` 2. Deploy the IaC template. 1. Download the template and the variable file: * [avn byoc template terraform get-template](/docs/tools/cli/byoc.md#avn-byoc-template-terraform-get-template) ``` avn byoc template terraform get-template \ --organization-id "ORGANIZATION_ID" \ --byoc-id "CUSTOM_CLOUD_ID" >| "tf_dir/tf_file.tf" ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-azure-custom-cloud.md#prerequisites). * `CUSTOM_CLOUD_ID` with the identifier of your custom cloud, which you can extract from the output of the [avn byoc list](/docs/tools/cli/byoc.md#avn-byoc-list) command, for example `018b6442-c602-42bc-b63d-438026133f60`. * [avn byoc template terraform get-vars](/docs/tools/cli/byoc.md#avn-byoc-template-terraform-get-vars) ``` avn byoc template terraform get-vars \ --organization-id "ORGANIZATION_ID" \ --byoc-id "CUSTOM_CLOUD_ID" >| "tf_dir/tf_file.tfvars" ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-azure-custom-cloud.md#prerequisites). * `CUSTOM_CLOUD_ID` with the identifier of your custom cloud, which you can extract from the output of the [avn byoc list](/docs/tools/cli/byoc.md#avn-byoc-list) command, for example `018b6442-c602-42bc-b63d-438026133f60`. 2. Optionally, modify the template as needed. note To connect to a custom-cloud service from IP address ranges outside the custom cloud, add specific ingress rules before you apply the Terraform infrastructure template. Before adding ingress rules, see the examples provided in the downloaded Terraform template. 3. Authenticate with Azure using the Azure CLI: ``` az login ``` For more authentication options, see the [Azure CLI authentication documentation](https://learn.microsoft.com/en-us/cli/azure/authenticate-azure-cli). 4. Install the Aiven CCE enterprise application on your Entra tenant: ``` az login --tenant "AZURE_TENANT_ID" az ad sp create --id b40b60e2-10c8-4917-bc74-18a87950e767 ``` Replace `AZURE_TENANT_ID` with your Azure tenant ID. To look it up, run: `az account show --query tenantId -o tsv`. The app ID is also available in the variables file you downloaded as `aiven_cce_client_id`. note Run these commands **once per tenant**, regardless of how many custom clouds you create on the same tenant. If the Aiven CCE enterprise application is already installed on your tenant, skip this step. To remove the Aiven CCE enterprise application from your tenant after you have deleted all custom clouds on it with `terraform destroy`, run: ``` az login --tenant "AZURE_TENANT_ID" az ad sp delete --id b40b60e2-10c8-4917-bc74-18a87950e767 ``` 5. Deploy the infrastructure template using Terraform with the provided variables file: ``` terraform init terraform apply -var-file=FILE_NAME.tfvars ``` Replace `FILE_NAME.tfvars` with the name of the variables file you downloaded. important The `-var-file` option is required to pass the configuration variables to Terraform. The template creates the following resources in your Azure subscription: * **Role assignments** granting the Aiven CCE enterprise application operator and quota-reader access in your subscription * A **resource group** containing all BYOC resources * Two **custom role definitions** in your subscription: `{deployment_name}-aiven-operator` (least-privilege operator on the resource group) and `{deployment_name}-aiven-quota-reader` (read-only compute quota access on the subscription) * **Storage accounts** (Premium LRS and Standard LRS), with the Storage Account Key Operator Service Role assigned to the Aiven CCE enterprise application * For the `standard` deployment model: * Two **virtual networks** (Bastion VNet and Workload VNet) with subnets * **VNet peering** between the bastion and workload networks * **Network security groups (NSGs)** controlling ingress and egress * **NAT gateways** for outbound internet access from both networks * For the `standard_public` deployment model: * A **Workload VNet** with a single subnet for service VMs * A **network security group (NSG)** allowing all public inbound TCP and UDP traffic to workload nodes 6. Retrieve the values required for the provisioning step: ``` terraform output -raw azure_subscription_id terraform output -raw azure_tenant_id ``` 3. Provision resources by running [avn byoc provision](/docs/tools/cli/byoc.md#avn-byoc-provision) and passing your Azure subscription ID and tenant ID. ``` avn byoc provision \ --organization-id "ORGANIZATION_ID" \ --byoc-id "CUSTOM_CLOUD_ID" \ --azure-subscription-id "AZURE_SUBSCRIPTION_ID" \ --azure-tenant-id "AZURE_TENANT_ID" ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-azure-custom-cloud.md#prerequisites). * `CUSTOM_CLOUD_ID` with the identifier of your custom cloud, which you can extract from the output of the [avn byoc list](/docs/tools/cli/byoc.md#avn-byoc-list) command, for example `018b6442-c602-42bc-b63d-438026133f60`. * `AZURE_SUBSCRIPTION_ID` with your Azure subscription ID from the Terraform output: `terraform output -raw azure_subscription_id`. * `AZURE_TENANT_ID` with your Azure tenant ID from the Terraform output: `terraform output -raw azure_tenant_id`. 4. Enable your custom cloud in organizations, projects, or units by running [avn byoc cloud permissions add](/docs/tools/cli/byoc.md#avn-byoc-cloud-permissions-add). ``` avn byoc cloud permissions add \ --organization-id "ORGANIZATION_ID" \ --byoc-id "CUSTOM_CLOUD_ID" \ --account "ACCOUNT_ID" ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-azure-custom-cloud.md#prerequisites). * `CUSTOM_CLOUD_ID` with the identifier of your custom cloud, which you can extract from the output of the [avn byoc list](/docs/tools/cli/byoc.md#avn-byoc-list) command, for example `018b6442-c602-42bc-b63d-438026133f60`. * `ACCOUNT_ID` with the identifier of your account (organizational unit) in Aiven, for example `a484338c34d7`. You can extract `ACCOUNT_ID` from the output of the `avn organization list` command. 5. Add customer contacts for the new cloud by running [avn byoc update](/docs/tools/cli/byoc.md#avn-byoc-update). ``` avn byoc update \ --organization-id "ORGANIZATION_ID" \ --byoc-id "CUSTOM_CLOUD_ID" \ ' { "contact_emails": [ { "email": "EMAIL_ADDRESS", "real_name": "John Doe", "role": "Admin" } ] } ' ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-azure-custom-cloud.md#prerequisites). * `CUSTOM_CLOUD_ID` with the identifier of your custom cloud, which you can extract from the output of the [avn byoc list](/docs/tools/cli/byoc.md#avn-byoc-list) command, for example `018b6442-c602-42bc-b63d-438026133f60`. ## Limitations[​](#limitations "Direct link to Limitations") The following features are not supported for Azure custom clouds: * **Enhanced compliance (ECE)** deployment models (`pci_dss`, `hipaa`) * **Static IPs** * **VNet peering** from the Aiven Console: manage peering directly in your Azure subscription * **PrivateLink** Related pages * [Bring your own cloud networking and security](/docs/platform/howto/byoc/networking-security.md) * [View the status of a custom cloud](/docs/platform/howto/byoc/view-custom-cloud-status.md) * [Manage services hosted in custom clouds](/docs/platform/howto/byoc/manage-byoc-service.md) --- # Create a custom cloud To create custom clouds in Aiven using self-service, select your cloud provider to integrate with. note BYOC for Oracle Cloud Infrastructure (OCI) is in [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-) and not available as self-service. To use BYOC with OCI, contact [Aiven](https://aiven.io/contact). [Amazon Web Services](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md) [Create an AWS-integrated custom cloud.](/docs/platform/howto/byoc/create-cloud/create-aws-custom-cloud.md) [Google Cloud](/docs/platform/howto/byoc/create-cloud/create-google-custom-cloud.md) [Create a Google-integrated custom cloud.](/docs/platform/howto/byoc/create-cloud/create-google-custom-cloud.md) [Microsoft Azure](/docs/platform/howto/byoc/create-cloud/create-azure-custom-cloud.md) [Create an Azure-integrated custom cloud.](/docs/platform/howto/byoc/create-cloud/create-azure-custom-cloud.md) #### Limitations[​](#limitations "Direct link to Limitations") * You need at least the Advanced tier of Aiven support services to be eligible for activating BYOC. tip See [Aiven support tiers](https://aiven.io/support-services) and [Aiven responsibility matrix](https://aiven.io/responsibility-matrix) for BYOC. Contact your account team to learn more or upgrade your support tier. * Only [organization admins](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) can create custom clouds. Related pages * [About bring your own cloud](/docs/platform/concepts/byoc.md) * [Bring your own cloud networking and security](/docs/platform/howto/byoc/networking-security.md) --- # Create a Google-integrated custom cloud Create a [custom cloud](/docs/platform/concepts/byoc.md) for BYOC in your Aiven organization to better address your specific business needs or project requirements. To configure a custom cloud in your Aiven organization and prepare your Google Cloud account so that Aiven can access it: 1. In the Aiven Console or with the Aiven CLI client, you specify new cloud details to generate a Terraform infrastructure-as-code template. 2. You download the generated template and deploy it in your Google Cloud account to acquire a privilege-bearing service account, which Aiven needs for accessing your Google Cloud account only with permissions that are required. note Privilege-bearing service account is an [identifier](https://registry.terraform.io/providers/hashicorp/google/latest/docs/resources/google_service_account#id) of the [service account](https://cloud.google.com/iam/docs/service-account-types#user-managed) created when running the infrastructure template in your Google Cloud account. Aiven [impersonates this service account](https://cloud.google.com/iam/docs/create-short-lived-credentials-direct) and runs operations, such as creating VMs for service nodes, in your BYOC account. 3. You deploy your custom cloud resources supplying the generated privilege-bearing service account to the Aiven platform, which gives Aiven the permissions to securely access your Google Cloud account, create resources, and manage them onward. 4. You select Aiven projects that can use your new custom clouds for creating services. 5. You add contact details for individuals from your organization that Aiven can reach out to in case of technical issues with the new cloud. ## Before you start[​](#before-you-start "Direct link to Before you start") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * You have [enabled the BYOC feature](/docs/platform/howto/byoc/enable-byoc.md). * You have an active account with your cloud provider. * Depending on the tool to use for creating a custom cloud: * Console: Access to the [Aiven Console](https://console.aiven.io/) or * CLI: * [Aiven CLI client](/docs/tools/cli.md) installed * Aiven organization ID from the output of the `avn organization list` command or from the [Aiven Console](https://console.aiven.io/) > **User information** > **Organizations**. * You have the [organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role in your Aiven organization. * You have Terraform installed. * You have required [IAM permissions](/docs/platform/howto/byoc/create-cloud/create-google-custom-cloud.md#iam-permissions). * Your Google Cloud project is not under an organization policy that: * Enforces the `constraints/compute.requireShieldedVm` constraint. For more information, see [Shielded VMs](https://cloud.google.com/compute/shielded-vm/docs/shielded-vm). * Restricts the `constraints/compute.vmExternalIpAccess` constraint. Aiven requires the ability to create instances with external IP addresses. An organization policy that only allows external IPs for specific listed instances is not compatible with BYOC. ### IAM permissions[​](#iam-permissions "Direct link to IAM permissions") You need cloud account credentials set up on your machine so that your user or role has required Terraform permissions [to integrate with your cloud provider](/docs/platform/howto/byoc/create-cloud/create-google-custom-cloud.md#create-a-custom-cloud). Show permissions needed by your service account that will run the Terraform script in your Google project * `roles/iam.serviceAccountAdmin` (sets up impersonation to the privilege-bearing service account) * `roles/resourcemanager.projectIamAdmin` (provides permissions to the privilege-bearing service account to use your project) * `roles/compute.instanceAdmin.v1` (manages networks and instances) * `roles/compute.securityAdmin` (creates firewall rules) * Enable [Identity and Access Management (IAM) API](https://cloud.google.com/iam/docs/reference/rest) to create the privilege-bearing service account * Enable [Cloud Resource Manager (CRM) API](https://cloud.google.com/resource-manager/reference/rest) to set IAM policies to the privilege-bearing service account * Enable [Compute Engine API](https://console.cloud.google.com/marketplace/product/google/compute.googleapis.com). For more information on Google Cloud roles, see [IAM basic and predefined roles reference](https://cloud.google.com/iam/docs/understanding-roles) in the Google Cloud documentation. ## Create a custom cloud[​](#create-a-custom-cloud "Direct link to Create a custom cloud") Create a custom cloud either in the Aiven Console or with the Aiven CLI. * Aiven Console * Aiven CLI #### Launch the BYOC setup[​](#launch-the-byoc-setup "Direct link to Launch the BYOC setup") 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to an organization. 2. Click **Admin** in the top navigation, and click **Bring your own cloud** in the sidebar. 3. In the **Bring your own cloud** view, select **Create custom cloud**. #### Generate an infrastructure template[​](#generate-an-infrastructure-template "Direct link to Generate an infrastructure template") In this step, an IaC template is generated in the Terraform format. In [the next step](/docs/platform/howto/byoc/create-cloud/create-google-custom-cloud.md#deploy-the-template), you'll deploy this template in your Google Cloud account to acquire a privilege-bearing service account (SA), which Aiven needs for accessing your Google Cloud account. In the **Create custom cloud** wizard: 1. Specify cloud details: * Cloud provider * Region * Custom cloud name * [Infrastructure tags](/docs/platform/howto/byoc/tag-custom-cloud-resources.md) 2. Click **Next**. 3. Specify deployment and storage details: * [Deployment model](/docs/platform/concepts/byoc.md#byoc-architecture) Choose between: * Private model, which routes traffic through a proxy for additional security utilizing a bastion host logically separated from the Aiven services. * Public model, which allows the Aiven control plane to connect to the service nodes via the public internet. note The service endpoint has two hostnames. The public hostname is derived from the private hostname by adding a `public-` prefix to it. The Service URI shown by the Aiven Console displays only the private hostname. * CIDR for BYOC resources The **CIDR** block defines the IP address range of the VPC that Aiven creates in your own cloud account. Any Aiven service created in the custom cloud will be placed in the VPC and will get an IP address within this address range. In the **CIDR** field, specify an IP address range for the BYOC VPC using a CIDR block notation, for example: `10.0.0.0/16`, `172.31.0.0/16`, or `192.168.0.0/20`. Make sure that an IP address range you use meets the following requirements: * IP address range is within the private IP address ranges allowed in [RFC 1918](https://datatracker.ietf.org/doc/html/rfc1918). * CIDR block size is between `/16` (65536 IP addresses) and `/24` (256 IP addresses). * CIDR block is large enough to host the desired number of services after splitting it into per-availability-zone subnets. For example, the smallest `/24` CIDR block might be enough for a few services but can pose challenges during node replacements or maintenance upgrades if running low on available free IP addresses. * CIDR block of your BYOC VCP doesn't overlap with the CIDR blocks of VPCs you plan to peer your BYOC VPC with. You cannot change the BYOC VPC CIDR block after your custom cloud is created. 4. Click **Generate template**. Your infrastructure Terraform template gets generated based on your inputs. You can view, copy, or download it. Now, you can use the template to acquire a privilege-bearing service account. #### Deploy the template[​](#deploy-the-template "Direct link to Deploy the template") Use the [generated Terraform template](/docs/platform/howto/byoc/create-cloud/create-google-custom-cloud.md#generate-an-infrastructure-template) to create a privilege-bearing service account by deploying the template in your Google Cloud account. 1. Copy or download the template and the variables file from the **Create custom cloud** wizard. 2. Optionally, modify the template as needed. note To connect to a custom-cloud service from different security groups (other than the one dedicated for the custom cloud) or from IP address ranges, add specific ingress rules before you apply the Terraform infrastructure template in your Google Cloud account in the process of creating a custom cloud resources. Before adding ingress rules, see the examples provided in the Terraform template you generated and downloaded from [Aiven Console](https://console.aiven.io/). 3. Set up Terraform to authenticate with Google Cloud. Configure your Google Cloud credentials using one of the following methods: * Set the `GOOGLE_APPLICATION_CREDENTIALS` environment variable: ``` export GOOGLE_APPLICATION_CREDENTIALS="/path/to/your/service-account-key.json" ``` * Use `gcloud` CLI to set application default credentials: ``` gcloud auth application-default login ``` * Set the service account key in the Terraform provider configuration. For more information, see the [Google Provider authentication documentation](https://registry.terraform.io/providers/hashicorp/google/latest/docs/guides/getting_started#adding-credentials). 4. Deploy the infrastructure template using Terraform: ``` terraform init terraform plan -var-file=FILE_NAME.tfvars terraform apply -var-file=FILE_NAME.tfvars ``` Replace `FILE_NAME.tfvars` with the name of the variables file you downloaded. important The `-var-file` option is required to pass the configuration variables to Terraform. 5. Find the privilege-bearing service account in the Terraform output after running `terraform apply`. 6. Supply the privilege-bearing service account into the **Create custom cloud** wizard. 7. Click **Next** to proceed or park your cloud setup and save your current configuration as a draft by selecting **Save draft**. You can resume creating your cloud later. #### Set up your custom cloud's availability[​](#set-up-your-custom-clouds-availability "Direct link to Set up your custom cloud's availability") Select projects where the new custom cloud will be available for hosting your services. These projects will support creating new services in the custom cloud or migrating your existing services to the custom cloud if your service and networking configuration allows it. For more information on migrating your existing services to the custom cloud, contact your account team. Your cloud can be available in: * All the projects in your organization * Selected organizational units * Specific projects only To set up your cloud's availability in the **Create custom cloud** wizard > the **Assign BYOC to projects** section, select one of the two following options: * **By default for all projects** to make your custom cloud available in all existing and future projects in the organization * **By selection** to pick specific projects or organizational units where you want your custom cloud to be available. note By selecting an organizational unit, you make your custom cloud available from all the projects in this unit. #### Add customer contacts[​](#add-customer-contacts "Direct link to Add customer contacts") Select at least one person whom Aiven can contact in case of any technical issues with your custom cloud. note **Admin** is a mandatory role, which is required as a primary support contact. In the **Create custom cloud** wizard > the **Customer contacts** section: 1. Select a contact person's role using the **Job title** menu, and provide their email address in the **Email** field. 2. Use **+ Add another contact** to add as many customer contacts as needed for your custom cloud. 3. Click **Save and validate**. The custom cloud process has been initiated. #### Complete the cloud setup[​](#complete-the-cloud-setup "Direct link to Complete the cloud setup") Select **Done** to close the **Create custom cloud** wizard. The deployment of your new custom cloud might take a few minutes. As soon as it's over, and your custom cloud is ready to use, you'll be able to see it in the list of your custom clouds in the **Bring your own cloud** view. note Your new custom cloud is ready to use only after its status changes to **Active**. 1. Generate an IaC template by running [avn byoc create](/docs/tools/cli/byoc.md#avn-byoc-create). ``` avn byoc create \ --organization-id "ORGANIZATION_ID" \ --deployment-model "DEPLOYMENT_MODEL_NAME" \ --cloud-provider "google" \ --cloud-region "CLOUD_REGION_NAME" \ --reserved-cidr "CIDR_BLOCK" \ --display-name "CUSTOM_CLOUD_DISPLAY_NAME" ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization to connect with your own cloud account to create the custom cloud, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-google-custom-cloud.md#prerequisites). * `DEPLOYMENT_MODEL_NAME` with the type of [network architecture](/docs/platform/concepts/byoc.md#byoc-architecture) your custom cloud uses: * `standard_public` (public) model: The nodes have public IPs and can be configured to be publicly accessible for authenticated users. The Aiven control plane can connect to the service nodes via the public internet. The service endpoint has two hostnames: a private one and a public one derived from it by adding a `public-` prefix. The Service URI shown in the Aiven Console displays only the private hostname. * `standard` (private) model: The nodes reside in a VPC without public IP addresses and are by default not accessible from outside. Traffic is routed through a proxy for additional security utilizing a bastion host logically separated from the Aiven services. * `CLOUD_REGION_NAME` with the name of a Google region where to create your custom cloud, for example `europe-north1`. See all available options in [Google Cloud regions](/docs/platform/reference/list_of_clouds.md#google-cloud). * `CIDR_BLOCK` with a CIDR block defining the IP address range of the VPC that Aiven creates in your own cloud account, for example: `10.0.0.0/16`, `172.31.0.0/16`, or `192.168.0.0/20`. * `CUSTOM_CLOUD_DISPLAY_NAME` with the name of your custom cloud, which you can set arbitrarily. Show sample output ``` { "custom_cloud_environment": { "cloud_provider": "google", "cloud_region": "europe-north1", "contact_emails": [ { "email": "firstname.secondname@domain.com", "real_name": "Test User", "role": "Admin" } ], "custom_cloud_environment_id": "018b6442-c602-42bc-b63d-438026133f60", "deployment_model": "standard", "display_name": "My BYOC Cloud on Google", "errors": [], "reserved_cidr": "10.0.0.0/16", "state": "draft", "tags": {}, "update_time": "2024-05-07T14:24:18Z" } } ``` 2. Deploy the IaC template. 1. Download the template and the variable file: * [avn byoc template terraform get-template](/docs/tools/cli/byoc.md#avn-byoc-template-terraform-get-template) ``` avn byoc template terraform get-template \ --organization-id "ORGANIZATION_ID" \ --byoc-id "CUSTOM_CLOUD_ID" >| "tf_dir/tf_file.tf" ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization to connect with your own cloud account to create the custom cloud, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-google-custom-cloud.md#prerequisites). * `CUSTOM_CLOUD_ID` with the identifier of your custom cloud, which you can extract from the output of the [avn byoc list](/docs/tools/cli/byoc.md#avn-byoc-list) command, for example `018b6442-c602-42bc-b63d-438026133f60`. * [avn byoc template terraform get-vars](/docs/tools/cli/byoc.md#avn-byoc-template-terraform-get-vars) ``` avn byoc template terraform get-vars \ --organization-id "ORGANIZATION_ID" \ --byoc-id "CUSTOM_CLOUD_ID" >| "tf_dir/tf_file.tfvars" ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization to connect with your own cloud account to create the custom cloud, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-google-custom-cloud.md#prerequisites). * `CUSTOM_CLOUD_ID` with the identifier of your custom cloud, which you can extract from the output of the [avn byoc list](/docs/tools/cli/byoc.md#avn-byoc-list) command, for example `018b6442-c602-42bc-b63d-438026133f60`. 2. Optionally, modify the template as needed. note To connect to a custom-cloud service from different security groups (other than the one dedicated for the custom cloud) or from IP address ranges, add specific ingress firewall rules before you apply the Terraform infrastructure template in your Google Cloud account in the process of creating a custom cloud resources. Before adding ingress rules, see the examples provided in the Terraform template you generated and downloaded from the [Aiven Console](https://console.aiven.io/). 3. Set up Terraform to authenticate with Google Cloud. Configure your Google Cloud credentials using one of the following methods: * Set the `GOOGLE_APPLICATION_CREDENTIALS` environment variable: ``` export GOOGLE_APPLICATION_CREDENTIALS="/path/to/your/service-account-key.json" ``` * Use `gcloud` CLI to set application default credentials: ``` gcloud auth application-default login ``` * Set the service account key in the Terraform provider configuration. For more information, see the [Google Provider authentication documentation](https://registry.terraform.io/providers/hashicorp/google/latest/docs/guides/getting_started#adding-credentials). 4. Deploy the infrastructure template using Terraform with the provided variables file: ``` terraform init terraform plan -var-file=FILE_NAME.tfvars terraform apply -var-file=FILE_NAME.tfvars ``` Replace `FILE_NAME.tfvars` with the name of the variables file you downloaded. important The `-var-file` option is required to pass the configuration variables to Terraform. 5. Find `privilege_bearing_service_account_id` in the Terraform output after running `terraform apply`. 3. Provision resources by running [avn byoc provision](/docs/tools/cli/byoc.md#avn-byoc-provision) and passing the generated `google-privilege-bearing-service-account-id` as an option. ``` avn byoc provision \ --organization-id "ORGANIZATION_ID" \ --byoc-id "CUSTOM_CLOUD_ID" \ --google-privilege-bearing-service-account-id "GENERATED_SERVICE_ACCOUNT_ID" ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization to connect with your own cloud account to create the custom cloud, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-google-custom-cloud.md#prerequisites). * `CUSTOM_CLOUD_ID` with the identifier of your custom cloud, which you can extract from the output of the [avn byoc list](/docs/tools/cli/byoc.md#avn-byoc-list) command, for example `018b6442-c602-42bc-b63d-438026133f60`. * `GENERATED_SERVICE_ACCOUNT_ID` with the identifier of the service account created when running the infrastructure template in your Google Cloud account, for example `projects/your-project/serviceAccounts/cce-cce0123456789a@your-project.iam.gserviceaccount.com`. You can extract `GENERATED_SERVICE_ACCOUNT_ID` from the output of the `terraform apply` command or `terraform output` command. 4. Enable your custom cloud in organizations, projects, or units by running [avn byoc cloud permissions add](/docs/tools/cli/byoc.md#avn-byoc-cloud-permissions-add). ``` avn byoc cloud permissions add \ --organization-id "ORGANIZATION_ID" \ --byoc-id "CUSTOM_CLOUD_ID" \ --account "ACCOUNT_ID" ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization to connect with your own cloud account to create the custom cloud, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-google-custom-cloud.md#prerequisites). * `CUSTOM_CLOUD_ID` with the identifier of your custom cloud, which you can extract from the output of the [avn byoc list](/docs/tools/cli/byoc.md#avn-byoc-list) command, for example `018b6442-c602-42bc-b63d-438026133f60`. * `ACCOUNT_ID` with the identifier of your account (organizational unit) in Aiven, for example `a484338c34d7`. You can extract `ACCOUNT_ID` from the output of the `avn organization list` command. 5. Add customer contacts for the new cloud by running [avn byoc update](/docs/tools/cli/byoc.md#avn-byoc-update). ``` avn byoc update \ --organization-id "ORGANIZATION_ID" \ --byoc-id "CUSTOM_CLOUD_ID" \ ' { "contact_emails": [ { "email": "EMAIL_ADDRESS", "real_name": "John Doe", "role": "Admin" } ] } ' ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization to connect with your own cloud account to create the custom cloud, for example `org123a456b789`. Get your `ORGANIZATION_ID` [from the Aiven Console or CLI](/docs/platform/howto/byoc/create-cloud/create-google-custom-cloud.md#prerequisites). * `CUSTOM_CLOUD_ID` with the identifier of your custom cloud, which you can extract from the output of the [avn byoc list](/docs/tools/cli/byoc.md#avn-byoc-list) command, for example `018b6442-c602-42bc-b63d-438026133f60`. Related pages * [Bring your own cloud networking and security](/docs/platform/howto/byoc/networking-security.md) * [View the status of a custom cloud](/docs/platform/howto/byoc/view-custom-cloud-status.md) * [Manage services hosted in custom clouds](/docs/platform/howto/byoc/manage-byoc-service.md) --- # Delete a custom cloud Delete a [custom cloud](/docs/platform/concepts/byoc.md) so that it's no longer available in your Aiven organization, units, or projects. After deleting a custom cloud, the data co-hosted in this cloud will no longer be available from the Aiven platform. Before deleting your custom cloud, make sure there are no active services using this cloud. ### Impact on your Aiven resources[​](#impact-on-your-aiven-resources "Direct link to Impact on your Aiven resources") The deletion impacts mostly resources on the Aiven site, such as cloud configuration files. ### Impact on your remote cloud resources[​](#impact-on-your-remote-cloud-resources "Direct link to Impact on your remote cloud resources") A bastion service and the corresponding compute instance are deleted as a consequence of your custom cloud's removal. As for resources created when applying the Terraform template to create the custom cloud, they are not removed after deleting the custom cloud. Unless you've removed them earlier, you're advised to do that after deleting your cloud. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven Console * Aiven CLI - You have at least one [custom cloud created](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organization. - You have access to the [Aiven Console](https://console.aiven.io/). - You have the [organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role in your Aiven organization. * You have at least one [custom cloud created](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organization. * You have the [Aiven CLI client](/docs/tools/cli.md) installed. * You have the [organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role in your Aiven organization. ## Delete your cloud[​](#delete-your-cloud "Direct link to Delete your cloud") ### Delete BYOC-deployed services[​](#delete-byoc-deployed-services "Direct link to Delete BYOC-deployed services") Before deleting your custom cloud, delete all services hosted in it. important * You cannot delete your custom cloud until all services deployed in it are deleted. * Deleting BYOC-deployed services permanently removes all hosted data, service backups, and configurations. ### Delete the cloud from Aiven[​](#delete-the-cloud-from-aiven "Direct link to Delete the cloud from Aiven") * Aiven Console * Aiven CLI 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your organization. 2. Click **Admin** in the top navigation, and click **Bring your own cloud** in the sidebar. 3. In the **Bring your own cloud** view, select a cloud. 4. On the selected cloud's page, click **Actions** > **Delete**. 5. Click **Delete** in the **Warning** window. Use the [avn byoc delete](/docs/tools/cli/byoc.md#avn-byoc-delete) command to delete your custom cloud. ``` avn byoc delete \ --organization-id "ORGANIZATION_IDENTIFIER" \ --byoc-id "CUSTOM_CLOUD_IDENTIFIER" ``` ### Delete cloud resources from Terraform[​](#delete-cloud-resources-from-terraform "Direct link to Delete cloud resources from Terraform") Remove the resources created in your remote cloud account when applying the Terraform template to create the custom cloud. They are not removed automatically after deleting the cloud. To delete the resources, run: ``` terraform destroy -var-file=FILE_NAME.tfvars ``` important Since running `terraform destroy` requires specifying the variables file, add `-var-file=FILE_NAME.tfvars` as an option. Related pages * [View the status of a custom cloud](/docs/platform/howto/byoc/view-custom-cloud-status.md) * [Bring your own cloud networking and security](/docs/platform/howto/byoc/networking-security.md) --- # Download an infrastructure template and a variables file Download a Terraform template and a variables file that define the infrastructure of your [custom cloud](/docs/platform/concepts/byoc.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven Console * Aiven CLI - At least one [custom cloud created](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organization - Access to the [Aiven Console](https://console.aiven.io/) - [Organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role in your Aiven organization * At least one [custom cloud created](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organization * [Aiven CLI client](/docs/tools/cli.md) installed * [Organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role in your Aiven organization ## Download an infrastructure template[​](#download-an-infrastructure-template "Direct link to Download an infrastructure template") * Aiven Console * Aiven CLI 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your organization. 2. Click **Admin** in the top navigation, and click **Bring your own cloud** in the sidebar. 3. Select a cloud, and find the **Infrastructure template** on the **Overview**. 4. Click **Download**. [Run the `avn byoc template terraform get-template` command](/docs/tools/cli/byoc.md#avn-byoc-template-terraform-get-template) to download your infrastructure template. ``` avn byoc template terraform get-template \ --organization-id "ORGANIZATION_IDENTIFIER" \ --byoc-id "CUSTOM_CLOUD_IDENTIFIER" >| "tf_dir/tf_file.tf" ``` ## Download a variable file[​](#download-a-variable-file "Direct link to Download a variable file") * Aiven Console * Aiven CLI 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your organization. 2. Click **Admin** in the top navigation, and click **Bring your own cloud** in the sidebar. 3. Select a cloud, and find the **Variables file** on the **Overview**. 4. Click **Download**. [Run the `avn byoc template terraform get-vars` command](/docs/tools/cli/byoc.md#avn-byoc-template-terraform-get-vars) to download your variables file. ``` avn byoc template terraform get-vars \ --organization-id "ORGANIZATION_IDENTIFIER" \ --byoc-id "CUSTOM_CLOUD_IDENTIFIER" >| "tf_dir/tf_file.tfvars" ``` Related pages * [Bring your own cloud networking and security](/docs/platform/howto/byoc/networking-security.md) * [Tag custom cloud resources](/docs/platform/howto/byoc/tag-custom-cloud-resources.md) --- # Enable bring your own cloud (BYOC) Enabling [the bring your own cloud (BYOC) feature](/docs/platform/concepts/byoc.md) allows you to [create custom clouds](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organization. note BYOC supports Amazon Web Services (AWS), Google Cloud, and Microsoft Azure with self-service, and Oracle Cloud Infrastructure (OCI) in [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-) upon [request](https://aiven.io/contact). To enable [BYOC](/docs/platform/concepts/byoc.md), open the [Aiven Console](https://console.aiven.io/) and [set up a call with the Aiven sales team](/docs/platform/howto/byoc/enable-byoc.md#enable-byoc). note Enabling [the BYOC feature](/docs/platform/concepts/byoc.md) or creating custom clouds in your Aiven environment does not affect the configuration of your existing Aiven organizations, projects, or services. It only allows you to run Aiven services in your cloud provider account. important Before enabling BYOC, check [who is eligible for BYOC](/docs/platform/concepts/byoc.md#who-is-eligible-for-byoc) and review [feature limitations](/docs/platform/howto/byoc/enable-byoc.md#byoc-enable-limitations) and [prerequisites](/docs/platform/howto/byoc/enable-byoc.md#byoc-enable-prerequisites). ## Limitations[​](#byoc-enable-limitations "Direct link to Limitations") * You need at least the Advanced tier of [Aiven support services](https://aiven.io/support-services) to be eligible for activating BYOC. * Only [organization admins](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) can request enabling BYOC. ## Prerequisites[​](#byoc-enable-prerequisites "Direct link to Prerequisites") * [Organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role for your Aiven organization * Access to the [Aiven Console](https://console.aiven.io/) * Active account with your cloud provider ## Enable BYOC[​](#enable-byoc "Direct link to Enable BYOC") 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your organization. 2. Click **Admin** in the top navigation, and click **Bring your own cloud** in the sidebar. 3. In the **Bring your own cloud** view, click **Contact us**. 4. In the **Contact us** window, enter your email address and country. Select the cloud provider, add any other information you think might be relevant, and click **Confirm**. The scheduling assistant shows up so that you can schedule a short call with the Aiven sales team to proceed on your BYOC enablement request. 5. Using the scheduling assistant, select a date and time when to talk to our sales team to share your requirements and make sure BYOC suits your needs. Confirm the selected time, make sure you add the call to your calendar, and close the scheduling assistant. 6. Join the scheduled call with our sales team to follow up with us on enabling BYOC in your environment. If the call reveals BYOC addresses your needs and your environment is eligible for BYOC, the feature will be enabled for your Aiven organization. ## Next steps[​](#next-steps "Direct link to Next steps") With BYOC activated in your Aiven organization, you can [create and use custom clouds](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md). Related pages * [About bring your own cloud](/docs/platform/concepts/byoc.md) * [Bring your own cloud networking and security](/docs/platform/howto/byoc/networking-security.md) * [Create a custom cloud in Aiven](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) --- # Manage services hosted in custom clouds Create a service in your custom cloud or migrate an existing service to your custom cloud. ## Create a service in a custom cloud[​](#create-a-service-in-a-custom-cloud "Direct link to Create a service in a custom cloud") * Aiven Console * Aiven CLI To create a service in the [Aiven Console](https://console.aiven.io/) in your new custom cloud, follow the instructions in the create a service guide for the service type. When creating a service in the [Aiven Console](https://console.aiven.io/), at the **Select service region** step, select **Custom clouds** from the available regions. To create a service hosted in your new custom cloud, run [avn service create](/docs/tools/cli/service-cli.md#avn-cli-service-create) passing your new custom cloud name as an option: ``` avn service create \ --project "PROJECT_NAME" \ --service-type "TYPE_OF_BYOC_SERVICE" \ --plan "SERVICE_PLAN" \ --cloud "CUSTOM_CLOUD_NAME" \ "NEW_BYOC_SERVICE_NAME" ``` ## Migrate an existing service to a custom cloud[​](#migrate-an-existing-service-to-a-custom-cloud "Direct link to Migrate an existing service to a custom cloud") You can migrate a non-BYOC Aiven-managed service to your custom cloud. How you do that depends on the [deployment mode](/docs/platform/concepts/byoc.md#byoc-architecture) of your custom cloud: public or private. ### Migrate to public BYOC[​](#migrate-to-public-byoc "Direct link to Migrate to public BYOC") To migrate a service to a custom cloud in the public deployment model, [change a cloud provider and a cloud region](/docs/platform/howto/migrate-services-cloud-region.md) to point to your custom cloud. ### Migrate to private BYOC[​](#migrate-to-private-byoc "Direct link to Migrate to private BYOC") Migrating a service to a custom cloud in the private deployment model requires network reconfiguration. Services are never exposed to the internet, and correct private communication must be established. [Create a support ticket](/docs/platform/howto/support.md#create-a-support-ticket) if you need assistance migrating to a private-deployment–model custom cloud. Related pages * [Bring your own cloud networking and security](/docs/platform/howto/byoc/networking-security.md) * [View the status of a custom cloud](/docs/platform/howto/byoc/view-custom-cloud-status.md) --- # Bring your own cloud networking and security Aiven combines a multi-cloud strategy with a cloud-agnostic approach to make the [bring your own cloud (BYOC)](/docs/platform/concepts/byoc.md) experience not only versatile and cost-efficient but also secure. ## Bastion proxy networking[​](#bastion-proxy-networking "Direct link to Bastion proxy networking") The [bastion](https://en.wikipedia.org/wiki/Bastion_host) proxy service acts as a trusted relay between the Aiven Management plane and your BYOC environment. A typical BYOC implementation results in workloads that are not directly accessible from external networks. The bastion proxy is located in a virtual private cloud (VPC) and facilitates the communication between Aiven Management and monitoring systems and your private workloads. The bastion proxy service is also responsible for facilitating a communication channel that allows the Aiven Management plane to launch services within your private cloud. ### Traffic to the bastion[​](#traffic-to-the-bastion "Direct link to Traffic to the bastion") Aiven Management plane traffic to your BYOC environment originates from a set of redundant gateways with fixed IPs. Therefore, the bastion subnet needs to whitelist four static IPs on relevant ports. | Source | Destination | Description | | ---------------------- | -------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | Aiven Management plane | Bastion subnet | Aiven uses SSH from the Aiven Management plane to the bastion service node for troubleshooting, setup, or configuration tasks on the bastion. | | Aiven Management plane | Bastion subnet | Proxy service is used to tunnel the traffic from the Aiven Management plane to workload nodes. | | Aiven Management plane | Bastion subnet | Aiven Management plane accesses the bastion service node to retrieve status information and perform routine operations. | ### Traffic from the bastion[​](#traffic-from-the-bastion "Direct link to Traffic from the bastion") | Source | Destination | Description | | -------------- | --------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | | Bastion subnet | Aiven Management | Communication channel with the Aiven Management plane used to collect metrics and logs from the bastion and workload nodes | | Bastion subnet | Aiven Management | Required for the bastion and workload nodes to make calls to the Aiven Management plane | | Bastion subnet | Aiven software repositories | Used for downloading RPM packages for bastion nodes setup | | Bastion subnet | Workload subnet | Aiven uses SSH from the bastion subnet to the workload subnet for troubleshooting, setup, or configuration tasks on the workload nodes. | | Bastion subnet | Workload subnet | Aiven Management plane accesses the workload nodes via the bastion node to retrieve status information and perform routine operations. | | Bastion subnet | DNS & NTP | Destination dependant on cloud provider | note Aiven Management and CloudFront IP ranges occasionally change and are considered dynamic for firewall policies. To accommodate traffic to these destinations, an egress of `0.0.0.0/0` is generally required. ## Workload nodes networking[​](#workload-nodes-networking "Direct link to Workload nodes networking") Management traffic to and from workload nodes is sent through the bastion proxy and encrypted. Aiven Management and API calls are always encrypted. SSH access is permitted if Aiven support staff require direct access to the bastion or workload nodes. SSH connections from support staff are logged, monitored, and require validation by Aiven's security operations team. SSH connections originate from Aiven Management plane's static IP addresses. ### Traffic to the workload[​](#traffic-to-the-workload "Direct link to Traffic to the workload") On top of SSH and Aiven Management Plane connectivity, there are also ingress and egress rules from the Customer Networks that need to be defined to allow customer applications to connect to Aiven services. This depends on the cloud provider being used, and the customer's network architecture. | Source | Destination | Description | | ------------------------ | --------------- | --------------------------------------------------------------------------------------------------------------------------------------- | | Bastion subnet | Workload subnet | Aiven uses SSH from the bastion subnet to the workload subnet for troubleshooting, setup, or configuration tasks on the workload nodes. | | Bastion subnet | Workload subnet | Aiven Management plane accesses the workload nodes via the bastion node to retrieve status information and perform routine operations. | | Your remote applications | Workload subnet | Varies depending on services being used. | ### Traffic from the workload[​](#traffic-from-the-workload "Direct link to Traffic from the workload") The workload subnet nodes communicate with each other. The nature of these communications varies depending on the services, features, and plugins being used. | Source | Destination | Description | | --------------- | --------------------------- | -------------------------------------------------------------------------------------------------------------------------- | | Workload subnet | Aiven Management | Communication channel with the Aiven Management plane used to collect metrics and logs from the bastion and workload nodes | | Workload subnet | Aiven Management | Required for the bastion and workload nodes to make calls to the Aiven Management plane | | Workload subnet | Aiven software repositories | Used for downloading RPM packages for workload nodes setup | | Workload subnet | Workload subnet | Intra-subnet communication between Aiven nodes. Ports are dynamic and dependent on services, features, and plugins used. | | Workload subnet | Your remote applications | Varies depending on services being used. | | Workload subnet | DNS & NTP | Destination dependant on cloud provider | note Aiven Management and CloudFront IP ranges occasionally change and are considered dynamic for firewall policies. To accommodate traffic to these destinations, an egress of `0.0.0.0/0` is generally required. ## Security and compliance[​](#security-and-compliance "Direct link to Security and compliance") ### System security[​](#system-security "Direct link to System security") The bastion and workload nodes leverage hardened images customized to only include components required for the operation of Aiven services. All nodes use full-disk encryption specific to the customer and service, and are rotated during maintenance events. Encryption keys are stored securely in an Aiven proprietary KMS, and only specific and approved operators have access to this system. All access is logged and monitored. If required for troubleshooting or incident investigation, SSH connections to the nodes can only be performed by specific and approved operators, and can only originate from specific Aiven gateway IPs. All access is logged and monitored, and Aiven's security team follows up on access to ensure approval and validity. Aiven base system images are routinely updated and patched. During maintenance events, service nodes, including the bastion service, are replaced with updated images. ### Independent audits[​](#independent-audits "Direct link to Independent audits") The BYOC deployment model is subject to external audits and falls within scope of Aiven's routine obligations. For more information on Aiven security and compliance, see [Aiven Security](https://aiven.io/security-compliance). Related pages * [About bring your own cloud](/docs/platform/concepts/byoc.md) * [Enable bring your own cloud (BYOC)](/docs/platform/howto/byoc/enable-byoc.md) * [Create a custom cloud in Aiven](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) --- # Rename a custom cloud Change the name of your [custom cloud](/docs/platform/concepts/byoc.md). With the [BYOC feature enabled](/docs/platform/howto/byoc/enable-byoc.md), you can [create custom clouds](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organizations. While [creating a custom cloud](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md), you specify the custom cloud name. You can change this name any time later by following [Rename your cloud](#rename-your-cloud). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven Console * Aiven CLI - At least one [custom cloud created](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organization - Access to the [Aiven Console](https://console.aiven.io/) - [Organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role in your Aiven organization * At least one [custom cloud created](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organization * [Aiven CLI client](/docs/tools/cli.md) installed * [Organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role in your Aiven organization ## Rename your cloud[​](#rename-your-cloud "Direct link to Rename your cloud") * Aiven Console * Aiven CLI 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your organization. 2. Click **Admin** in the top navigation, and click **Bring your own cloud** in the sidebar. 3. In the **Bring your own cloud** view, select a cloud. 4. On the selected cloud's page, click **Actions** > **Rename**. 5. In the **Rename custom cloud** window, enter a new name, and click **Rename**. Use the [`avn byoc update`](/docs/tools/cli/byoc.md#avn-byoc-update) command to change the name of your custom cloud. ``` avn byoc update \ --organization-id "ORGANIZATION_IDENTIFIER" \ --byoc-id "CUSTOM_CLOUD_IDENTIFIER" \ --display-name "NAME_OF_CUSTOM_CLOUD" ``` Related pages * [View the status of a custom cloud](/docs/platform/howto/byoc/view-custom-cloud-status.md) * [Assign a project to your custom cloud](/docs/platform/howto/byoc/assign-project-custom-cloud.md) * [Add customer's contact information for your custom cloud](/docs/platform/howto/byoc/add-customer-info-custom-cloud.md) * [Tag custom cloud resources](/docs/platform/howto/byoc/tag-custom-cloud-resources.md) --- # Use tiered storage for AWS BYOC services [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) AWS BYOC environments use the tiered storage capability for data allocation. Cold data in your AWS custom cloud is stored in your AWS cloud account. note This is a [limited availability feature](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-). AWS [BYOC](/docs/platform/concepts/byoc.md) environments allow using tiered storage to store data. The tiered storage is a data allocation mechanism for improved efficiency and cost optimization of data management. When enabled, tiered storage allows moving data automatically between hot storage (for frequently accessed, critical, and often updated data) and cold storage (for rarely accessed, static, or archived data). Cold data of AWS BYOC-hosted services is stored in object storage in your AWS cloud account. One bucket is created per custom cloud. important * AWS [BYOC](/docs/platform/concepts/byoc.md) tiered storage is only supported for [Aiven for Apache Kafka](/docs/products/kafka/howto/kafka-tiered-storage-get-started.md) and [Aiven for ClickHouse](/docs/products/clickhouse/concepts/clickhouse-tiered-storage.md). * Tiered storage enabled on non-BYOC services is owned by Aiven and as such doesn't allow to store cold data in your own cloud account. * Non-BYOC services with Aiven-owned tiered storage cannot be migrated to BYOC. To use tiered storage in an AWS BYOC-hosted service, tiered storage needs to be enabled both [in your custom cloud](/docs/platform/howto/byoc/store-data.md#enable-tiered-storage-in-an-aws-custom-cloud) and [in the BYOC-hosted service](/docs/platform/howto/byoc/store-data.md#enable-tiered-storage-on-a-service). ## Enable tiered storage in an AWS custom cloud[​](#enable-tiered-storage-in-an-aws-custom-cloud "Direct link to Enable tiered storage in an AWS custom cloud") [Contact the Aiven support team](mailto:support@aiven.io) to request enabling tiered storage in your AWS custom cloud. ## Enable tiered storage on a service[​](#enable-tiered-storage-on-a-service "Direct link to Enable tiered storage on a service") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * At least one AWS [custom cloud](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) * At least one Aiven-managed service, either Aiven for Apache Kafka® or Aiven for ClickHouse®, hosted in a custom cloud note If your Aiven-managed service is not hosted in an AWS custom cloud, you can [migrate it](/docs/platform/howto/byoc/manage-byoc-service.md#migrate-an-existing-service-to-a-custom-cloud). ### Activate tiered storage[​](#activate-tiered-storage "Direct link to Activate tiered storage") * [Enable for Aiven for Apache Kafka](/docs/products/kafka/howto/enable-kafka-tiered-storage.md) * [Enable for Aiven for ClickHouse](/docs/products/clickhouse/howto/enable-tiered-storage.md) Related pages * [Bring your own cloud networking and security](/docs/platform/howto/byoc/networking-security.md) * [View the status of a custom cloud](/docs/platform/howto/byoc/view-custom-cloud-status.md) * [Manage services hosted in custom clouds](/docs/platform/howto/byoc/manage-byoc-service.md) --- # Tag custom cloud resources Tagging allows resource categorization, which simplifies governance, cost allocation, and system performance review. Custom cloud tags propagate to resources on the Aiven platform and in your own cloud infrastructure. important Since the tags propagate to your own cloud infrastructure, the contents and total number of the tags need to stay within limits imposed by your cloud provider. ## Types of tagging[​](#types-of-tagging "Direct link to Types of tagging") You can tag your custom cloud resources by: * [Service tagging](#service-tagging), affecting Aiven service nodes and VMs * [Infrastructure tagging](#infrastructure-tagging), affecting all taggable BYOC infrastructure components ## Tagging service nodes and VMs[​](#service-tagging "Direct link to Tagging service nodes and VMs") If you add a tag to a BYOC service, all the service nodes and VMs inherit this tag, and the tag propagates to your own cloud infrastructure. ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven Console * Aiven CLI - At least one [custom cloud created](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organization - Access to the [Aiven Console](https://console.aiven.io/) - [Organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role in your Aiven organization * At least one [custom cloud created](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organization * [Aiven CLI client](/docs/tools/cli.md) installed * [Organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) role in your Aiven organization ### Tag service nodes and VMs[​](#tag-service-nodes-and-vms "Direct link to Tag service nodes and VMs") * Aiven Console * Aiven CLI Create a resource tag for a BYOC service the same way you create it for a regular Aiven-managed service. Ensure you use the `byoc_resource_tag` prefix in the tag key. For example, to label all VMs running a particular BYOC service with the tag that has `my-cost-center` as a key and `12345` as a value, create a resource tag for this service with key `byoc_resource_tag:my-cost-center` and value `12345`. For your BYOC service, [create tags using the Aiven CLI the same way you create them for a regular Aiven-managed service](/docs/tools/cli/service-cli.md#avn-service-tags). Ensure you use the `byoc_resource_tag` prefix in the tag key. ``` avn service tags update SERVICE_NAME --add-tag byoc_resource_tag:business_unit=sales --add-tag byoc_resource_tag:env=smoke_test ``` For instructions on how to add, update, remove, or list service tags via Aiven CLI, see [the avn service tags documentation](/docs/tools/cli/service-cli.md#avn-service-tags), where you also find limits and limitations that apply to service tags and tagging. ## Tagging infrastructure components[​](#infrastructure-tagging "Direct link to Tagging infrastructure components") You can define a set of tags for each taggable infrastructure component created by the Terraform infrastructure template (for example, VPCs, subnets, or security groups). You can manage the tags using the [Aiven CLI client](/docs/tools/cli.md) or directly in the variable file used to run the Terraform infrastructure template. note Tagging GCP BYOC infrastructure uses [Google labels](https://cloud.google.com/resource-manager/docs/labels-overview), not [Google tags](https://cloud.google.com/resource-manager/docs/tags/tags-overview). ### Limitations[​](#limitations "Direct link to Limitations") * Tag keys are in lower case and can include ASCII alphanumeric printable English characters. Punctuation characters other than dashes (between words) are not allowed. * Do not use tag keys that start with `aiven`. * Do not change the tags that the Terraform template applies by default. ### Before you start[​](#before-you-start "Direct link to Before you start") * You have at least one [custom cloud created](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in your Aiven organization. * You have the [Aiven CLI client](/docs/tools/cli.md) installed. ### Manage infrastructure tags[​](#manage-infrastructure-tags "Direct link to Manage infrastructure tags") Use the [`avn byoc tags update`](/docs/tools/cli/byoc.md#avn-byoc-tags-update) command to add or update infrastructure tags for your custom cloud. Pass the tags as an option. ``` avn byoc tags update \ --organization-id "ORGANIZATION_IDENTIFIER" \ --byoc-id "CUSTOM_CLOUD_IDENTIFIER" \ --add-tag TAG_KEY_A=TAG_VALUE_A \ --add-tag TAG_KEY_B=TAG_VALUE_B ``` important Any change to infrastructure tags requires reapplying the Terraform template. Related pages * [Bring your own cloud networking and security](/docs/platform/howto/byoc/networking-security.md) * [View the status of a custom cloud](/docs/platform/howto/byoc/view-custom-cloud-status.md) * [Manage services hosted in custom clouds](/docs/platform/howto/byoc/manage-byoc-service.md) * [Download an infrastructure template and a variables file](/docs/platform/howto/byoc/download-infrastructure-template.md) --- # View the status of a custom cloud Find out whether your custom cloud is ready to use by viewing its status. 1. Log in to [Aiven Console](https://console.aiven.io/) as an administrator, and go to an organization. 2. From the top navigation bar, select **Admin**. 3. From the left sidebar, select **Bring your own cloud**. 4. In the **Bring your own cloud** view, identify your new cloud on the list of available clouds and check its status in the **Status** column. When your custom cloud's status is **Active**, its deployment has been completed. Your custom cloud is ready to use and you can see it on the list of your custom clouds in the **Bring your own cloud** view. Now you can [create new services in the custom cloud](/docs/platform/howto/byoc/manage-byoc-service.md#create-a-service-in-a-custom-cloud) or [migrate your existing services to the custom cloud](/docs/platform/howto/byoc/manage-byoc-service.md#migrate-an-existing-service-to-a-custom-cloud) if your service and networking configuration allows it. For more information on migrating your existing services to the custom cloud, contact your account team. If the status is **Creation failed**, provisioning validation encountered an unrecoverable error. Delete the custom cloud and re-create it after resolving the issue. Related pages * [Bring your own cloud networking and security](/docs/platform/howto/byoc/networking-security.md) * [Manage services hosted in custom clouds](/docs/platform/howto/byoc/manage-byoc-service.md) --- # Change your email address You can't change your login email address directly. Instead, you can create a user with the new email address and remove the user with the old email address. You must be an [organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) to do this. Aiven support cannot change your email address. 1. In the organization, click **Admin**. 2. Click **Users**. 3. Click **Invite users**. 4. Enter the new email address. 5. Click **Invite users**. 6. Remove the old user from all groups and [add the new user to those groups](/docs/platform/howto/manage-groups.md). 7. Optional: Update the new user's [permissions](/docs/platform/concepts/permissions.md). 8. After the user accepts the invite and you confirm that they have the correct access, [remove the old user](/docs/platform/howto/manage-org-users.md) from the organization. --- # Configure project base port The base port number for a project determines the Aiven service ports. The base port is randomly assigned, but you can also configure the port number for a project. To set the base port number for a project: * Console * Terraform 1. In the project, click **Settings**. 2. In the **Networking settings** section, enter a **Base port** number. 3. Click **Save changes**. Use the `base_port` attribute in [your `aiven_organization_project` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/organization_project). --- # Control maintenance updates with upgrade pipelines [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Link Aiven services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. Control when your Aiven managed services receive maintenance updates and test maintenance updates in development or staging environments before they reach production. important Controlled upgrade pipeline is a [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-) feature. [Contact Aiven](https://aiven.io/contact) to request access. Aiven performs automatic service maintenance for security fixes, minor software updates, and other platform changes. The controlled upgrade pipeline feature lets you link services of the same type in an ordered sequence to control when each service receives updates. After a maintenance update upgrades a service at the initial pipeline step, you validate that service version before the update proceeds to the service at the next pipeline step. Validating means approving the new version as safe to roll out to the next service. Validation can be manual or automatic after a configurable delay. ## Why use controlled upgrade pipelines[​](#why-use-controlled-upgrade-pipelines "Direct link to Why use controlled upgrade pipelines") Controlled upgrade pipelines prevent production incidents caused by automatic updates reaching production before teams can test the new version in a lower environment. They give you full oversight of the update process: * **Risk mitigation**: Prevents unexpected maintenance updates from breaking your production environment by ensuring they are tested in a non-production setting first. * **Stability**: Keeps destination services (such as production) on a known-good version until you, or the automatic timer, confirm the new version is safe. * **Process control**: Allows platform teams to standardize their deployment and maintenance lifecycle across environments. ## About controlled upgrade pipelines[​](#about-controlled-upgrade-pipelines "Direct link to About controlled upgrade pipelines") ### Upgrade steps[​](#upgrade-steps "Direct link to Upgrade steps") An upgrade step is a pair of services linked by an upgrade constraint: * **Source service**: The service that receives maintenance updates first * **Destination service**: The service that waits for validation before receiving updates Each destination service can have only one source service. A source service can have multiple destination services. ### Upgrade pipelines[​](#upgrade-pipelines "Direct link to Upgrade pipelines") An upgrade pipeline is a chain of upgrade steps that spans multiple environments. For example: * Single chain: development → staging → production * Multiple destinations: development → production-eu and development → production-na ## How validation works[​](#how-validation-works "Direct link to How validation works") When a maintenance update upgrades your source service: 1. The source service receives the update first. 2. Test the updated source service to verify it works as expected. 3. Validate the update manually using the API or CLI, or wait for automatic validation after the configured delay. The default delay is 7 days. 4. After validation, the destination service becomes eligible for the same maintenance update. 5. The destination service receives the update during its next maintenance window. If one source service has multiple destination services, one validation for the source service applies to all connected destination services. ### Validation and maintenance windows[​](#validation-and-maintenance-windows "Direct link to Validation and maintenance windows") Validation and the maintenance window control different things: * **Validation** controls *what* version the destination service upgrades to. * The maintenance window controls *when* the upgrade happens. After you validate an update, or automatic validation applies, the destination service receives the validated version during its next scheduled maintenance window. Validation does not trigger an immediate upgrade outside the maintenance window. Upgrade pipelines add a constraint on what is installed during a maintenance update; they do not change when maintenance runs. Nodes in the destination service maintain the validated version until a newer version is validated, either when you validate it manually or when automatic validation applies after the configured delay. When a node is recycled, it uses the same validated version, not the latest available version. When you create a step, the destination service keeps the newest version that is already validated at that moment. If the destination service is already applying maintenance during step creation, the in-progress target version becomes the initial validated version. warning A powered-off source service cannot receive maintenance updates, so you cannot validate it. If you power off services earlier in the chain, the destination service upgrades regardless. For example, in a development → staging → production chain, if both development and staging are powered off, production upgrades without testing and validation in the earlier environments. Keep services in the chain powered on to preserve the protection that upgrade pipelines provide. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * **Same service type**: You can only link services of the same type. For example, two Aiven for PostgreSQL services. * **Chain length**: The default maximum chain depth is 3 services, which is 2 steps. If you need a longer chain, [contact Aiven](https://aiven.io/contact). * **No cycles**: You cannot create circular dependencies between services. * **Emergency overrides**: Aiven can apply critical security or stability fixes to a destination service before explicit validation. * **Supported services**: This feature supports all Aiven service types except Aiven for Apache Flink® and Aiven for MySQL. * **Automatic maintenance updates only**: Pipelines apply to automatic maintenance updates, such as minor service version updates and node image updates. Major version upgrades, for example Aiven for PostgreSQL® 15 to 16, require manual action and are not promoted automatically through the pipeline. * **No permanent blocking**: You cannot prevent an update indefinitely. Automatic validation applies after the configured delay, up to the maximum delay. * **No validation rollback**: You cannot undo a validation after it is recorded. ## Use controlled upgrade pipelines[​](#use-controlled-upgrade-pipelines "Direct link to Use controlled upgrade pipelines") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") To use controlled upgrade pipelines, you need the following: * The feature enabled by Aiven ([Limited availability](/docs/platform/concepts/service-and-feature-releases.md)) * Dev tool of your choice: * [Aiven CLI](/docs/tools/cli.md) Install the latest version of the Aiven CLI to access the `upgrade-pipeline` commands. * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) * Set `PROVIDER_AIVEN_ENABLE_BETA=true` before running Terraform. * See the [resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) for full schema, import format, and lifecycle behavior. * [Aiven Operator for Kubernetes](/docs/tools/kubernetes.md) Install the operator and create an Aiven token secret named `aiven-token` that the operator uses to authenticate against the Aiven API. * Write access to the source and destination projects * At least two services of the same type (for example, two Aiven for PostgreSQL® services) * Services can be in different projects in the same organization ### Set up an upgrade pipeline[​](#set-up-an-upgrade-pipeline "Direct link to Set up an upgrade pipeline") Use the Aiven CLI or API to create upgrade steps between your services. note The `upgrade-pipeline` CLI commands require Aiven CLI version 4.x or later. Command names and parameters may change before general availability. #### Create an upgrade step[​](#create-an-upgrade-step "Direct link to Create an upgrade step") Create a step to link a source service and a destination service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ [--source-project SOURCE_PROJECT] SOURCE_SERVICE \ [--destination-project DESTINATION_PROJECT] DESTINATION_SERVICE \ [--auto-validation-delay-days DAYS] ``` **Options** * `--organization-id` is required. * `--source-project` and `--destination-project` are optional. If you omit either project option, Aiven CLI uses the current default project set with `avn project switch`. * `--auto-validation-delay-days` is optional. Defaults to 7 days if not specified. ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "SOURCE_PROJECT_NAME", "source_service_name": "SOURCE_SERVICE_NAME", "destination_project_name": "DESTINATION_PROJECT_NAME", "destination_service_name": "DESTINATION_SERVICE_NAME", "auto_validation_delay_days": 7 }' ``` Use the [`aiven_upgrade_step`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) resource: ``` resource "aiven_upgrade_step" "example" { organization_id = "ORGANIZATION_ID" source_project_name = "SOURCE_PROJECT_NAME" source_service_name = "SOURCE_SERVICE_NAME" destination_project_name = "DESTINATION_PROJECT_NAME" destination_service_name = "DESTINATION_SERVICE_NAME" auto_validation_delay_days = 7 } ``` Apply an `UpgradePipelineStep` manifest with `kubectl`: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: upgrade-step-sample spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: SOURCE_PROJECT_NAME sourceServiceName: SOURCE_SERVICE_NAME destinationProjectName: DESTINATION_PROJECT_NAME destinationServiceName: DESTINATION_SERVICE_NAME autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable after the resource is created. Parameters: * `source_project_name`: Name of the project containing the source service * `source_service_name`: Name of the source service * `destination_project_name`: Name of the project containing the destination service * `destination_service_name`: Name of the destination service * `auto_validation_delay_days`: Optional. Number of days before automatic validation. The value must be at least `1`. The default is 7 days. The maximum delay you can configure is 30 days. #### List upgrade steps[​](#list-upgrade-steps "Direct link to List upgrade steps") View all upgrade steps you have access to: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step list --organization-id ORGANIZATION_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) To list managed upgrade steps, use: ``` terraform state list 'aiven_upgrade_step.*' terraform state show 'aiven_upgrade_step.example' ``` List `UpgradePipelineStep` resources in the current namespace: ``` kubectl get upgradepipelinesteps ``` #### View a specific step[​](#view-a-specific-step "Direct link to View a specific step") Get details about a specific upgrade step, including the last validation: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step get \ --organization-id ORGANIZATION_ID \ STEP_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` terraform state show aiven_upgrade_step.example ``` Show the manifest and full status, including `id`, `conditions`, and `lastValidation`: ``` kubectl describe upgradepipelinestep RESOURCE_NAME kubectl get upgradepipelinestep RESOURCE_NAME -o yaml ``` The step details include `last_validation` values such as `validated_at`, `validated_by_user`, and `comment` when validation exists (available through the API). ### Validate an upgrade[​](#validate-an-upgrade "Direct link to Validate an upgrade") After testing your source service with the new update, validate the version to allow the destination service to receive the same update. #### Manual validation[​](#manual-validation "Direct link to Manual validation") Validate the current version of your source service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step validate-for-service \ --project SOURCE_PROJECT \ SERVICE_NAME \ [--comment "COMMENT"] ``` `--comment` is optional. Use it to record a note about the validation, for example `"Tested and verified in development"`. ``` curl -X POST https://api.aiven.io/v1/project/SOURCE_PROJECT/service/SOURCE_SERVICE/upgrade-validation \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "comment": "Tested and verified in development" }' ``` Terraform manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. The operator manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. #### Automatic validation[​](#automatic-validation "Direct link to Automatic validation") If you do not manually validate an update, the system automatically validates the source service version after the configured delay. Auto-validation starts from when the source service receives the update. ### Manage upgrade steps[​](#manage-upgrade-steps "Direct link to Manage upgrade steps") #### Update a step[​](#update-a-step "Direct link to Update a step") Modify the automatic validation delay for an existing step: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step update \ --organization-id ORGANIZATION_ID \ --auto-validation-delay-days 14 \ STEP_ID ``` ``` curl -X PATCH https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "auto_validation_delay_days": 14 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` resource "aiven_upgrade_step" "example" { # ...required fields... auto_validation_delay_days = 14 # Updated from 7 to 14 } ``` Apply the changes: ``` terraform plan terraform apply ``` Edit `autoValidationDelayDays` in your manifest and re-apply: ``` spec: autoValidationDelayDays: 14 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable. To change them, delete the resource and create a new one. #### Delete a step[​](#delete-a-step "Direct link to Delete a step") Remove an upgrade step to allow the destination service to receive updates independently: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step delete --organization-id ORGANIZATION_ID STEP_ID ``` Find `STEP_ID` from the upgrade step list command. ``` curl -X DELETE https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) Remove the resource from configuration and apply, or destroy it directly: ``` terraform apply terraform destroy -target=aiven_upgrade_step.example ``` Delete the `UpgradePipelineStep` resource: ``` kubectl delete upgradepipelinestep RESOURCE_NAME ``` Deleting a step removes all associated validations. ### Example: Three-environment pipeline[​](#example-three-environment-pipeline "Direct link to Example: Three-environment pipeline") Create a pipeline that promotes updates from development to staging to production: * CLI * API * Terraform * Kubernetes 1. Create a step from development to staging: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project dev-project \ --destination-project staging-project \ --auto-validation-delay-days 3 \ pg-dev pg-staging ``` 2. Create a step from staging to production: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project staging-project \ --destination-project prod-project \ --auto-validation-delay-days 7 \ pg-staging pg-prod ``` 1) Create a step from development to staging: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "dev-project", "source_service_name": "pg-dev", "destination_project_name": "staging-project", "destination_service_name": "pg-staging", "auto_validation_delay_days": 3 }' ``` 2) Create a step from staging to production: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "staging-project", "source_service_name": "pg-staging", "destination_project_name": "prod-project", "destination_service_name": "pg-prod", "auto_validation_delay_days": 7 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` # Step 1: Development → Staging resource "aiven_upgrade_step" "dev_to_staging" { organization_id = "ORGANIZATION_ID" source_project_name = "dev-project" source_service_name = "pg-dev" destination_project_name = "staging-project" destination_service_name = "pg-staging" auto_validation_delay_days = 3 } # Step 2: Staging → Production resource "aiven_upgrade_step" "staging_to_prod" { organization_id = "ORGANIZATION_ID" source_project_name = "staging-project" source_service_name = "pg-staging" destination_project_name = "prod-project" destination_service_name = "pg-prod" auto_validation_delay_days = 7 } ``` Apply the configuration: ``` export PROVIDER_AIVEN_ENABLE_BETA=true terraform init terraform plan terraform apply ``` Define both steps in a single manifest and apply it: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: dev-to-staging spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: dev-project sourceServiceName: pg-dev destinationProjectName: staging-project destinationServiceName: pg-staging autoValidationDelayDays: 3 --- apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: staging-to-prod spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: staging-project sourceServiceName: pg-staging destinationProjectName: prod-project destinationServiceName: pg-prod autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-pipeline.yaml ``` When a maintenance update arrives: 1. The development service receives the update. 2. After testing, validate the development version or wait 3 days for auto-validation. 3. The staging service receives the update during its next maintenance window. 4. After testing, validate the staging version or wait 7 days for auto-validation. 5. The production service receives the update during its next maintenance window. Related pages * [Service maintenance, updates and upgrades](/docs/platform/concepts/maintenance-window.md) * [Fork a service](/docs/platform/concepts/service-forking.md) --- # Create personal tokens Create personal token in the Aiven Console to use with the Aiven CLI or API. 1. Click **User information** and select **Tokens**. 2. Click **Generate token**. 3. Enter a description and set the session duration. 4. Click **Generate token**. 5. Click the **Copy** icon and save your token somewhere safe. important You cannot view the token after you close this window. 6. Click **Close**. Related pages * Use [authentication policies](/docs/platform/howto/set-authentication-policies.md) to control how organization users use personal tokens * [Create a token using the Aiven CLI](/docs/tools/cli/user.md) --- # Create a service Create an Aiven service from the Aiven Console. 1. In your project, click **Services**. 2. Click **Create service**. 3. Select the service type. 4. Enter a name for your service. important You cannot change the name after you create the service. You can [fork the service](/docs/platform/concepts/service-forking.md) with a new name instead. 5. Optional: Add tags. 6. Select the cloud provider, region, and plan. note Available plans and pricing for the same service can vary between cloud providers and regions. 7. Optional: Add disk storage. 8. In the **Service summary**, click **Create service**. The status of the service is **Rebuilding** during its creation. When the status is **Running**, you can start using the service. This typically takes couple of minutes and can vary between cloud providers and regions. Related pages * [Create service users](/docs/platform/howto/create_new_service_user.md) * [Services](/docs/products/services.md) --- # Create service users Create additional users to access an Aiven service. Service users only exist in the scope of the Aiven service. They are unique to the service and not shared with any other services. Every service has a default `avnadmin` user with full access to the service. 1. In your service, **Users**. 2. Click **Add service user** or **Create user**. 3. Enter a name for your service user. 4. Set up all the other configuration options. If a password is required, a random password is generated automatically. You can change it later. 5. Click **Add service user**. By default, there is a limit of 50 service users per service. To request a higher limit for a service, [create a support ticket](/docs/platform/howto/support.md). Related pages * [Create a service](/docs/platform/howto/create_new_service.md) * [Use resource tags](/docs/platform/howto/tag-resources.md) --- # Create service integrations Create [service integrations](/docs/platform/concepts/service-integration.md) between different Aiven services and move telemetry data using these integrations. tip For help with setting up a service integration or to request an integration type or endpoint not yet available, contact the [support team](mailto:support@aiven.io). The following example shows you how to create and integrate these services: * An Aiven for Apache Kafka® service that produces the telemetry data. * An Aiven for PostgreSQL® service where the telemetry data is stored and can be queried. * An Aiven for Grafana® service for the telemetry data. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Console * Terraform - Access to the [Aiven Console](https://console.aiven.io) * [Terraform installed](https://www.terraform.io/downloads) * A [personal token](/docs/platform/howto/create_authentication_token.md) ## Create the integrations[​](#create-the-integrations "Direct link to Create the integrations") * Console * Terraform 1. In the Aiven Console, create 3 services: Aiven for Apache Kafka®, Aiven for PostgreSQL®, and Aiven for Grafana®. note Integrations are not available on Hobbyist plans. 2. On the **Services** page, open the PostgreSQL service. 3. Click **Integrations**. 4. Under **Aiven services**, click **Grafana Metrics Dashboard**. 5. Select **Existing service** and choose the Grafana service you created. 6. Click **Enable**. 7. To send metrics from the Apache Kafka service, under **Aiven services** select **Receive Metrics**. 8. Select **Existing service** and choose the Apache Kafka service you created. 9. Click **Enable**. 1) Create a file named `provider.tf` and add the following: ``` Loading... ``` 2) For the services, create a file named `services.tf` and add the following: ``` Loading... ``` 3) To integrate the services, create a file named `integrations.tf` and add the following: ``` Loading... ``` 4) Create a file named `variables.tf` and add the following: ``` Loading... ``` 5) Create a file named `terraform.tfvars` and add values for your token and Aiven project. To apply your Terraform configuration: 1. Initialize Terraform by running: ``` terraform init ``` The output is similar to the following: ``` Initializing the backend... Initializing provider plugins... - Finding aiven/aiven versions matching ">= 4.0.0, < 5.0.0"... - Installing aiven/aiven v4.9.2... - Installed aiven/aiven v4.9.2 ... Terraform has been successfully initialized! ... ``` 2. To create an execution plan and preview the changes, run: ``` terraform plan ``` 3. To deploy your changes, run: ``` terraform apply --auto-approve ``` To view the Apache Kafka metrics data in Grafana: 1. In the [Aiven Console](https://console.aiven.io/), go to your Grafana service. 2. In the **Connection information** section, click the **Service URI** to open Grafana. 3. Log in using the **User** and **Password** provided in the **Connection information** section. 4. On the dashboards page in Grafana, click the dashboard name to view it. If you don't see a dashboard after logging in, search for `Aiven Kafka - - Resources` in the Grafana console. This is a predefined dashboard automatically maintained by Aiven. Data can take a minute to appear on the dashboard after enabling the integrations. Refresh the view by reloading the page. warning Any modifications you make to the predefined dashboard are automatically overwritten by the system during updates. You can create your own custom dashboards, or make a copy of this predefined dashboard to customize it. Don't use `Aiven` at the start of your dashboard names. --- # Use credits Use credits to pay for your Aiven services by adding them to a billing group. All projects assigned to the billing group automatically use its credits for the costs of the projects and the services within the projects. You must have the `organization:billing:write` [permission](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) to add credits to a billing group. To add credits to a billing group: 1. Click **Billing**. 2. Click **Billing groups**. 3. Find the billing group and click **Details**. 4. On the **Credits** tab, enter the credit code and click **Claim credits**. --- # Delete user account You can delete your personal user account as long as you are not a managed user of an organization. warning This action is irreversible. All data is permanently removed and your cannot be recovered. To delete your account: 1. Delete all services in all projects. 2. [Delete all projects](/docs/platform/howto/manage-project.md#delete-a-project). 3. [Delete all organizations](/docs/platform/howto/manage-organizations.md) that you manage and [leave all organizations](/docs/platform/howto/manage-organizations.md) you are part of. 4. Go to **User information** > **User profile**. 5. Click **Delete account**. note Billing groups with trial credits and the organizations they are assigned to cannot be deleted. You can delete them after the trial period ends. To delete the billing group, organization, or your account during the trial period [contact Aiven support](/docs/platform/howto/support.md). --- # Scale disk storage automatically Automatically increase the disk storage of an Aiven service when it's running out of space, instead of resizing it manually. Use the Aiven Autoscaler to automatically increase the storage capacity of a service disk when it's running out of space. Disk autoscaler only increases storage, it doesn't scale storage down. ## Why use disk autoscaling[​](#why-use-disk-autoscaling "Direct link to Why use disk autoscaling") * **Cost efficiency**: Start with a regular-sized disk and let Aiven scale it up only when needed, without the risk of running out of disk space. * **Resiliency**: Avoid a service becoming non-functional because it ran out of disk space, including during unexpected spikes in demand. ## How it works[​](#how-it-works "Direct link to How it works") 1. You create an autoscaler integration endpoint in your project, setting the maximum total disk size to allow. 2. You enable an autoscaler integration for your service using that endpoint. 3. Aiven monitors the disk space usage of your service. 4. When disk usage reaches the threshold for your service type, Aiven increases the available storage by at least 10%, using the current used space as a baseline. note The exact increase depends on the service type and cloud provider. Some providers enforce a minimum increase of 10 GB. Autoscale thresholds per service type The threshold that triggers disk autoscaling is a percentage of the available disk storage capacity: * Aiven for OpenSearch®: 75% of the available disk storage capacity * All other supported service types: 85% of the available disk storage capacity 5. The disk increase is recorded in the project event log, and you receive a notification about the added disk space. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * **Maximum storage**: The maximum storage that the autoscaler can allocate for your service is limited by both the maximum disk size set on the autoscaler endpoint and the maximum disk storage supported for your service plan. * **Timing**: Autoscaling takes a moment to complete. In the meantime, the service disk might fill up and the service might enter read-only mode until autoscaling finishes, unless the autoscaler's disk capacity limit is reached. * **Maintenance updates**: Autoscaling works only on fully running services and can't happen during a maintenance update. * **Manual changes**: Changing disk space manually can delay an autoscaling event. * **Terraform**: Don't manage disk space with the Aiven Terraform Provider on a service that uses the autoscaler, to avoid conflicts between the two. * **Performance**: Disk added through autoscaling is slower than the original disk until the next maintenance update applies. This might affect I/O-intensive workloads. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven organization, project, and service that's up and running * The operator role for the organization, project, and service * Dynamic disk sizing support on your service plan and cloud region * One of the following to manage the autoscaler: * [Aiven Console](https://console.aiven.io/) * [Aiven API](https://api.aiven.io/doc/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) ### Enable disk autoscaling[​](#enable-disk-autoscaling "Direct link to Enable disk autoscaling") To enable disk autoscaling, create an autoscaler integration endpoint, then enable an autoscaler integration on your service using that endpoint. * Console * API * CLI * Terraform Create an autoscaler endpoint: 1. Log in to [Aiven Console](https://console.aiven.io/), and go to your organization and project. 2. On the left sidebar, click **Integration endpoints**. 3. Click **Aiven Autoscaler** > **Add new endpoint**. 4. Set the endpoint name and the maximum total disk storage in GB, and click **Add endpoint**. Enable the autoscaler on a service: 1. On the left sidebar, click **Services**, and open your service. 2. On the left sidebar, click **Integrations**. 3. In **Endpoint integrations**, click **Aiven Autoscaler**. 4. Click the endpoint you created, and click **Enable**. 1) Call [ServiceIntegrationEndpointCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointCreate) to create an autoscaler integration endpoint on your project: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "endpoint_name": "ENDPOINT_NAME", "endpoint_type": "autoscaler", "user_config": { "autoscaling": [ { "type": "autoscale_disk", "cap_gb": 300 } ] } }' ``` 2) Call [ServiceIntegrationCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationCreate) to enable the autoscaler integration on your service, using the endpoint ID from the previous response: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "dest_endpoint_id": "ENDPOINT_ID", "integration_type": "autoscaler", "source_project": "PROJECT_NAME", "source_service": "SERVICE_NAME" }' ``` 1. Create an autoscaler integration endpoint using [avn service integration-endpoint-create](/docs/tools/cli.md): ``` avn service integration-endpoint-create \ --project PROJECT_NAME \ --endpoint-name ENDPOINT_NAME \ --endpoint-type autoscaler \ --user-config-json '{"autoscaling": [{"type": "autoscale_disk", "cap_gb": 300}]}' ``` 2. Find the ID of the new endpoint: ``` avn service integration-endpoint-list --project PROJECT_NAME ``` 3. Enable the autoscaler integration on your service, using the endpoint ID from the previous step: ``` avn service integration-create \ --dest-service SERVICE_NAME \ --integration-type autoscaler \ --source-endpoint-id ENDPOINT_ID ``` Use the [`aiven_service_integration_endpoint`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration_endpoint) and [`aiven_service_integration`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration) resources: ``` resource "aiven_service_integration_endpoint" "autoscaler_endpoint" { project = "PROJECT_NAME" endpoint_name = "ENDPOINT_NAME" endpoint_type = "autoscaler" autoscaler_user_config { autoscaling { type = "autoscale_disk" cap_gb = 300 } } } resource "aiven_service_integration" "autoscaler_integration" { project = "PROJECT_NAME" integration_type = "autoscaler" source_service_name = "SERVICE_NAME" destination_endpoint_id = aiven_service_integration_endpoint.autoscaler_endpoint.id } ``` See the [disk autoscaler guide](https://registry.terraform.io/providers/aiven/aiven/latest/docs/guides/disk-autoscaler) for more details. ### Change the maximum disk space for autoscaling[​](#change-the-maximum-disk-space-for-autoscaling "Direct link to Change the maximum disk space for autoscaling") After you enable disk autoscaling, you can update the maximum total disk size at any time. * Console * API * CLI * Terraform 1. On the left sidebar, click **Integration endpoints**. 2. Click **Aiven Autoscaler**. 3. Find your endpoint, click **Actions**, and click the option to edit it. 4. Set a new maximum disk storage value, and save your changes. Call [ServiceIntegrationEndpointUpdate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointUpdate) with the new `cap_gb` value: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint/ENDPOINT_ID \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "user_config": { "autoscaling": [ { "type": "autoscale_disk", "cap_gb": 500 } ] } }' ``` ``` avn service integration-endpoint-update ENDPOINT_ID \ --user-config-json '{"autoscaling": [{"type": "autoscale_disk", "cap_gb": 500}]}' ``` Update the `cap_gb` value in the `autoscaling` block of your [`aiven_service_integration_endpoint`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration_endpoint) resource, then apply the change. ### Turn off disk autoscaling[​](#turn-off-disk-autoscaling "Direct link to Turn off disk autoscaling") To turn off disk autoscaling, remove the autoscaler integration from your service. You can also delete the integration endpoint if you no longer need it. * Console * API * CLI Disconnect the service from the autoscaler: 1. On the left sidebar, click **Services**, and open your service. 2. On the left sidebar, click **Integrations**. 3. In **Endpoint integrations**, find **Aiven Autoscaler**, click **Actions**, and click the option to disconnect it. Delete the autoscaler endpoint, if you no longer need it: 1. On the left sidebar, click **Integration endpoints**. 2. Click **Aiven Autoscaler**. 3. Find your endpoint, and delete it. 1) Call [ServiceIntegrationDelete](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationDelete) to remove the autoscaler integration from your service: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration/INTEGRATION_ID \ --header "Authorization: Bearer TOKEN" ``` 2) Call [ServiceIntegrationEndpointDelete](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointDelete) to delete the autoscaler integration endpoint, if you no longer need it: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint/ENDPOINT_ID \ --header "Authorization: Bearer TOKEN" ``` 1. Find the ID of the integration to remove: ``` avn service integration-list SERVICE_NAME ``` 2. Remove the autoscaler integration from your service: ``` avn service integration-delete INTEGRATION_ID ``` 3. Find the ID of the integration endpoint to delete, if you no longer need it: ``` avn service integration-endpoint-list --project PROJECT_NAME ``` 4. Delete the autoscaler integration endpoint: ``` avn service integration-endpoint-delete ENDPOINT_ID ``` Related pages * [Scale disk storage manually](/docs/platform/howto/add-storage-space.md) * [Change a service plan](/docs/platform/howto/scale-services.md) --- # Download invoices You can download monthly invoices from the Aiven Console in PDF or CSV format. You must have the `organization:billing:read` [permission](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) to download invoices. 1. Click **Billing**. 2. Click **Invoices**. 3. Click **Actions** > **Download PDF** or **Download CSV**. --- # Edit your user profile You can edit your name, job title, location, or other personal details in the Aiven Console. 1. Click **User information** in the top right. 2. Select **User profile**. 3. Click **Edit**. 4. Edit your profile and click **Save changes**. --- # Feature previews Before an official release, some features are available to our customers for testing. These feature previews let you try out upcoming enhancements and give our product teams feedback to help improve them. ## Enable a feature preview[​](#enable-a-feature-preview "Direct link to Enable a feature preview") To try upcoming features before they are released: 1. Click the **User information** icon in the top right and select **Feature preview**. 2. On the **Feature preview** tab, click **Enable** for any of the features you want to test. After enabling a feature preview and testing it, you can provide feedback by clicking **Give feedback**. --- # Access Aiven services from Google Cloud Functions via VPC peering You can access Aiven service by creating a **Serverless VPC access connector** and **Google Cloud Function**. By default, **Google Cloud Functions** can only access the Internet and is not able to access your GCP VPC or Aiven VPC. For **Google Cloud Functions** to access VPC, **Serverless VPC access connector** is required. **Serverless VPC access connector** consists of two or more Google-managed VM that forward requests (and perform NAT) from Cloud Functions to your GCP VPC and Aiven VPC. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") You have: * Created a [VPC on the Aiven platform](/docs/platform/howto/manage-project-vpc.md). * Set up [VPC peering on GCP](/docs/platform/howto/manage-project-vpc.md). ## Create a Serverless VPC access connector[​](#create-a-serverless-vpc-access-connector "Direct link to Create a Serverless VPC access connector") 1. Open GCP console and go to **Navigation menu** > **Networking** > **VPC network** and select [Serverless VPC access](https://console.cloud.google.com/networking/connectors/list). 2. Click **Create connector**: * **Name**: The connector name of your choice. * **Region**: The region where to create the Cloud Function. * **Network**: Your GCP VPC, which is already peered to Aiven VPC already. * **Subnet**: Select **custom IP range** and enter a **/28** private subnet that is not in use. 3. If you have **allowed IP addresses** configured on your Aiven service, ensure the subnet of **serverless VPC access connector** is listed there ## Create a Cloud Function[​](#create-a-cloud-function "Direct link to Create a Cloud Function") 1. Open GCP console and under **Navigation menu**, **Serverless** section, select [Cloud Functions](https://console.cloud.google.com/functions/list). 2. Click **create function** * **Environment**: Your choice of environment. You can use the default value (2nd gen). * **Function name**: the name of your choice. * **Region**: The region of the **serverless VPC access connector**. * Click and expand the **runtime, build connections and security settings** section, select **Connections** tab, and select the **serverless VPC access connector** you have created. * Click **Next** 3. Select the runtime of your choice. warning Do not click **Test function**. 4. Click **Deploy** 5. Wait for GCP to deploy the cloud function. Once deployed, use the **Source** tab to edit the function if needed. warning Do not click **Test function**. 6. Click the **Testing** tab and test the command in Cloud Shell to ensure it can access VPC. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") If you cannot access your VPC or Aiven VPC from the Cloud Function, consider using the following example for troubleshooting purposes. ``` # Cound Function 2nd gen, Python 3.11 import functions_framework import socket CLOUD_FUNCTION_KEY = 'gcf-aiven-test-CHANGE_ME_FOR_SECURITY_REASON' @functions_framework.http def hello_http(request): request_json = request.get_json(silent=True) if request_json and "cloud_function_key" in request_json and request_json["cloud_function_key"] == CLOUD_FUNCTION_KEY: result = "" try: host = request_json['host'] port = request_json['port'] timeout = request_json.get('timeout', 10) s = socket.socket(socket.AF_INET, socket.SOCK_STREAM) s.settimeout(timeout) s.connect((host, port)) result = "OK" except Exception as e: result = repr(e) pass return 'Result: {}\n'.format(result) return "HTTP 401\n", 401 ``` The request body contains: * `CLOUD_FUNCTION_KEY` Change this to protect your Cloud Function endpoint, especially if it does not require authentication. * `host`: FQDN or IP address if your Aiven service or VM in your GCP VPC. * `port`: Destination TCP port number. **Example:** In the **Testing** tab in your **Cloud Function**: ``` { "cloud_function_key": "gcf-aiven-test-CHANGE_ME_FOR_SECURITY_REASON", "host": "fqdn-or-ip-to-your-aiven-service.a.aivencloud.com", "port": 12345 } ``` The request returns: * `OK` if it can establish TCP 3-way handshake. * `TimeoutError` if it cannot reach the specified port. For assistance, contact Aiven support. When you do, share your your Cloud Function endpoint and `CLOUD_FUNCTION_KEY`. --- # Access JMX metrics via Jolokia [Jolokia](https://jolokia.org/) is one of the external metrics integrations supported on the Aiven platform, along with [Datadog metrics](/docs/integrations/datadog/datadog-metrics.md) and [Prometheus metrics](/docs/platform/howto/integrations/prometheus-metrics.md). note JMX metrics via Jolokia are only supported for Aiven for Apache Kafka® services. ## Configure a Jolokia endpoint[​](#configure-a-jolokia-endpoint "Direct link to Configure a Jolokia endpoint") * Console * Aiven CLI To enable Jolokia integration, create a Jolokia endpoint: 1. In the [Aiven Console](https://console.aiven.io/), select your project. 2. Click **Integration endpoints** in the sidebar. 3. Click **Jolokia** in the list of integration types. 4. Click **Add new endpoint**. 5. Enter an endpoint name and click **Create**. A username and password are generated automatically. You can reuse the same Jolokia endpoint for multiple services in the same project. Use the [avn service integration-endpoint create](/docs/tools/cli/service/integration.md#avn_service_integration_endpoint_create) command to create a Jolokia endpoint: ``` avn service integration-endpoint create \ --project \ --endpoint-name \ --endpoint-type jolokia ``` ## Enable Jolokia integration[​](#enable-jolokia-integration "Direct link to Enable Jolokia integration") To enable Jolokia integration for an Aiven for Apache Kafka® service: 1. In the [Aiven Console](https://console.aiven.io/), select your Aiven for Apache Kafka service. 2. On the **Overview** page, click **Integrations**. 3. Under **Endpoint integrations**, click **Jolokia**. 4. Select the Jolokia endpoint you created and click **Enable**. This configures the endpoint on all service nodes and provides access to JMX metrics. Jolokia supports HTTP POST requests to retrieve service-specific metrics and bulk requests to collect metrics in batches. For details, see the [Jolokia protocol documentation](https://jolokia.org/reference/html/manual/jolokia_protocol.html). Several metrics are specific to individual Kafka® brokers. To get a complete view of the cluster, you might need to query each broker. The brokers share a DNS name. Use the `host` command (on Unix) or `nslookup` (on Windows) to list the associated IP addresses. ``` host kafka-67bd7c5-myproject.aivencloud.com kafka-67bd7c5-myproject.aivencloud.com has address 35.228.218.115 kafka-67bd7c5-myproject.aivencloud.com has address 35.228.234.106 kafka-67bd7c5-myproject.aivencloud.com has address 35.228.157.197 ``` ## Kafka topic-level metrics availability[​](#kafka-topic-level-metrics-availability "Direct link to Kafka topic-level metrics availability") Some Kafka topic-level JMX metrics might not be available if the topic has no recent traffic. For example: * Messages in per second: `kafka.server:type=BrokerTopicMetrics,name=MessagesInPerSec,topic=` * Replication bytes in per second: `kafka.server:type=BrokerTopicMetrics,name=ReplicationBytesInPerSec` Kafka stops exposing these metrics after several minutes (typically 10 to 15) without message production or consumption on the topic. When traffic resumes, the metrics typically reappear after a short delay. If you request one of these metrics while the topic is inactive, Jolokia returns a `404` response. This behavior is expected and does not indicate an issue with Kafka or Jolokia. For more information, see the [Kafka JMX documentation](https://kafka.apache.org/42/operations/monitoring/#security-considerations-for-remote-monitoring-using-jmx). ## Example cURL requests[​](#example-curl-requests "Direct link to Example cURL requests") Before sending a cURL request, [download the CA certificate](/docs/platform/concepts/tls-ssl-certificates.md#download-ca-certificates) for your project. The same certificate applies to all endpoints and services within the project. To read a specific metric, use port `6733`, which is the default port for Jolokia. Replace `joljkr2l:PWD` with the username and password generated during the Jolokia endpoint setup. View the credentials in the endpoint details on the **Integration endpoints** page. ``` curl --cacert ca.pem \ -X POST \ https://joljkr2l:PWD@HOST_IP:6733/jolokia/ \ -d \ '{"type":"read","mbean":"kafka.server:type=ReplicaManager,name=PartitionCount"}' ``` Jolokia supports searching beans using `search` command: ``` curl --cacert ca.pem \ -X POST \ https://joljkr2l:PWD@HOST_IP:6733/jolokia/ \ -d \ '{"type":"search","mbean":"kafka.server:*"}' ``` --- # Increase metrics limit setting for Datadog Monitoring services and applications are essential to know whether programs work as expected. To get started with monitoring, see [Aiven and Datadog integration](/docs/integrations/datadog.md). Sometimes, you cannot find the metrics you expected, or some values are missing on dashboards for large service clusters. You can overcome this limitation and get more metrics from the integration. ## Identify that metrics have been dropped[​](#identify-that-metrics-have-been-dropped "Direct link to Identify that metrics have been dropped") The following is an example log of a large Apache Kafka® service cluster where some metrics are missing and cannot be found in the Datadog dashboards after service integration. These metrics have been dropped by user Telegraf. ``` 2022-02-15T22:47:30.601220+0000 scoober-kafka-3c1132a3-82 user-telegraf: 2022-02-15T22:47:30Z W! [outputs.prometheus_client] Metric buffer overflow; 3378 metrics have been dropped 2022-02-15T22:47:30.625696+0000 scoober-kafka-3c1132a3-86 user-telegraf: 2022-02-15T22:47:30Z W! [outputs.prometheus_client] Metric buffer overflow; 1197 metrics have been dropped ``` ## Configure the maximum metric limit[​](#configure-the-maximum-metric-limit "Direct link to Configure the maximum metric limit") You can set the maximum metric limit for services in the Datadog integration configuration, where the Datadog agents gather metrics via `JMX` using the `max_jmx_metrics` option. The value of this metric can be set to any value between 10 and 100000. The default value is 2000. note Datadog may also limit the number of custom metrics that can be received, based on your pricing [plan](https://docs.datadoghq.com/account_management/billing/custom_metrics/?tab=countrate#allocation) . The `max_jmx_metrics` is not exposed in the Aiven Console yet, but you can change the value for it from the [Aiven CLI](https://github.com/aiven/aiven-client) using the following procedure: 1. Find the `SERVICE_INTEGRATION_ID` for your Datadog integration with ``` avn service integration-list --project=PROJECT_NAME SERVICE_NAME ``` 2. Change the value of `max_jmx_metrics` to the new LIMIT: ``` avn service integration-update SERVICE_INTEGRATION_ID --project PROJECT_NAME -c max_jmx_metrics=LIMIT ``` note We recommend you gradually increase the value and monitor the memory usage on the cluster. --- # Use Prometheus with Aiven Discover Prometheus as a tool for monitoring your Aiven services. Check why use it and how it works. Learn how to enable and configure Prometheus on your project. ## About Prometheus[​](#about-prometheus "Direct link to About Prometheus") Prometheus is an open-source systems monitoring and alerting toolkit. It implements an in-memory and persistent storage model for metrics as well as a query language for accessing the metrics. The metrics delivery model of Prometheus is a pull model where the Prometheus server connects to HTTP servers running on the nodes being monitored and pulls the metrics from them. While this makes the service discovery more of a challenge than using the more common push approach, it does have the benefit of making the metrics available not just for Prometheus but for any application that can read the Prometheus format from the HTTP server running on Aiven nodes. ## Check Prometheus support for your service[​](#check-prometheus-support-for-your-service "Direct link to Check Prometheus support for your service") Usually one Prometheus integration endpoint can be used for all services in the same project. To check if Prometheus is supported on your service, verify if the project for this service has a Prometheus integration endpoint created: 1. Log in to [Aiven Console](https://console.aiven.io/), go to **Projects** in the top navigation bar, and select your project. 2. On the **Services** page, select **Integration endpoints** from the left sidebar. 3. On the **Integration endpoints** page, select **Prometheus** from the list available integration endpoints, and check if there is any endpoint available under **Endpoint Name**. If there is a Prometheus endpoint available, your service supports Prometheus. If there's no Prometheus endpoint available, proceed to [Enable Prometheus on your Aiven project](/docs/platform/howto/integrations/prometheus-metrics.md#enable-prometheus) to set up Prometheus for your service (project). ## Enable Prometheus[​](#enable-prometheus "Direct link to Enable Prometheus") Aiven offers Prometheus endpoints for your services. To enable this feature: 1. Log in to [Aiven Console](https://console.aiven.io/), go to **Projects** in the top navigation bar, and select your project. 2. On the **Services** page, select **Integration endpoints** from the left sidebar. 3. On the **Integration endpoints** page, select **Prometheus** from the list available integration endpoints, and select **Add new endpoint**. 4. In the **Create new Prometheus endpoint** window, enter the details for the endpoint, and select **Create**. 5. Select **Services** from the sidebar, and go to the service that you would like to monitor. 6. On the **Overview** page of your service, go to the **Service integrations** section, and select **Manage integrations**. 7. On the **Integrations** page, select **Prometheus**. 8. In the **Prometheus integration** window, select the endpoint name you created, and select **Enable**. note At the top of the **Integrations** page, you will see the Prometheus integration listed and status `active`. 9. Next, go to the service's **Overview** page, and locate the **Connection information** section. 10. Click the **Prometheus** tab. 11. Copy **Service URI**, and use it in your browser to access the Prometheus dashboard. The system initiates an HTTP server on all service nodes, granting access to the metrics. Aiven is exposing endpoints for your services. note There might be a slight delay of approximately one minute before the metrics become available. ### Accessing Prometheus in a VPC[​](#accessing-prometheus-in-a-vpc "Direct link to Accessing Prometheus in a VPC") If you use a VPC in your project, access Prometheus: 1. Access [Aiven Console](https://console.aiven.io/). 2. Select your project, and select the service to monitor using Prometheus. 3. Click **Service settings** from the sidebar. 4. In the **Cloud and network** section, click 5. Choose **More network configurations**. 6. In the **Network configuration** window, select **Add configuration options**. 7. Search for the `public_access.prometheus` property and enable it. 8. Click **Save configuration**. ## Configure Prometheus[​](#configure-prometheus "Direct link to Configure Prometheus") After enabling Prometheus on your project, add a scrape configuration to Prometheus for the servers to pull data from. The examples in this section show how to set up the scrape configuration for both single-node and multi-node services. ### Single-node services[​](#single-node-services "Direct link to Single-node services") For single-node services, configure the following in your `scrape_config` job entry in `prometheus.yml`: * `basic_auth` details: Check your service's the **Overview** page > the **Connection information** section > the **Prometheus** tab. * `PROMETHEUS_SERVICE_URI`: Check **Service URI** on your service's the **Overview** page > the **Connection information** section > the **Prometheus** tab. * `ca_file`: Download the CA certificate from your service's the **Overview** page, and specify its location (the certificates are signed by the Aiven project CA). note You can download the CA certificate using the [Aiven command line client](https://github.com/aiven/aiven-client/) and command `avn project ca-get --target-filepath ca.pem`. Sample configuration ``` scrape_configs: - job_name: aivenmetrics scheme: https basic_auth: username: password: tls_config: ca_file: ca.pem static_configs: - targets: [": password: dns_sd_configs: - names: - type: A port: tls_config: insecure_skip_verify: true ``` note For Aiven services with multiple nodes and a Replica URI, the primary DNS name does not include standby IP addresses. To track those, make sure to include the replica DNS names in the list. If you have `` as `public-example.aivencloud.com`, then you will need to add `public-replica-example.aivencloud.com`. This applies to PostgreSQL®, MySQL®, and Caching services. ### View full list of metrics[​](#view-full-list-of-metrics "Direct link to View full list of metrics") The Prometheus client, provided through the Telegraf plugin, offers access to a comprehensive range of metrics, similar to those previously available through another metrics integration, now accessible via the Prometheus integration. You can preview the full list of metrics in [Prometheus system metrics](/docs/integrations/prometheus-system-metrics.md). note For some services the metrics provided by different hosts may vary depending on the host role. Most notably for Kafka® only one of the nodes provides metrics related to consumer group offsets. Related pages Learn more about integrations with Aiven: * [Aiven integrations](/docs/platform/concepts/service-integration.md) * [Datadog integration](/docs/integrations/datadog.md) * Configure Prometheus for Aiven for Apache Kafka® via Privatelink --- # Authentication methods Users can authenticate to the Aiven platform using a password, single sign-on (SSO), or [tokens](/docs/platform/concepts/authentication-tokens.md). The available authentication methods depend on the organization's [authentication policy](/docs/platform/howto/set-authentication-policies.md). Organization admin set these policies to restrict or require specific authentication methods for all users in an organization. They can also set up SSO through their preferred [identity provider](/docs/platform/howto/list-identity-providers.md). For an additional layer of security, Aiven also supports [two-factor authentication](/docs/platform/howto/set-authentication-policies.md). --- # SAML identity providers and verified domains Set up single sign-on (SSO) access to Aiven through a Security Assertion Markup Language (SAML) compliant identity provider (IdP). This lets you centrally manage your users in your IdP while giving them a seamless login experience. Every IdP must be linked to a domain in Aiven. After you [verify that you own a domain](/docs/platform/howto/manage-domains.md), the users in your organization become managed users, which provides a higher level of security for your organization by controlling things like [how these users log in](/docs/platform/howto/set-authentication-policies.md). With a verified domain you can add an IdP. All users with an email address from the verified domain are automatically authenticated with the linked IdP. With IdP-initiated SSO enabled, users can log in to Aiven directly from the IdP. Aiven also supports System for Cross-domain Identity Management (SCIM) for Okta to automatically provision, update, and deactivate user identities from your IdP. With automatic provisioning you don’t need to manually create organization users. When adding an IdP you link it to the verified domain and can set up SCIM at the same time. ## Limitations[​](#limitations "Direct link to Limitations") If you set up user provisioning with SCIM, only make changes to user details in the IdP. ## Security best practices[​](#security-best-practices "Direct link to Security best practices") It’s recommended to verify your domains in Aiven even if you don’t use SSO. When configuring an IdP it's best to enable the following SAML security settings: * **Require assertion to be signed**: Verifies assertions were issued by a trusted party and have not been tampered with. * **Sign authorization request sent to IdP**: Ensures authenticity and integrity with a digital signature. The [authentication policy](/docs/platform/howto/set-authentication-policies.md) for the organization is also an important component in securing access through an IdP. At a minimum, use these settings for your authentication policy: * Don't allow password authentication * Require log in with this organization's identity provider To limit access further, also consider these authentication policy settings: * **Don't allow third-party authentication**: This combined with the preceding password and organization identity provider settings ensures that users only log in to the Console with your chosen IdP. * **Don't allow users to create personal tokens**: This prevents users from accessing organization resources through the API using a long-lived [personal token](/docs/platform/concepts/authentication-tokens.md) they created. If you allow your users to create personal tokens, you can still make these more secure by enabling **Require users to be logged in with an allowed authentication method**. This means that users cannot access your organization's resources with a token they created when logged in with another organization's allowed authentication methods or a previously allowed method. This setting also gives you the flexibility to change the authentication policy at any time because tokens that are no longer compliant with the new policy cannot be used. --- # Manage virtual private clouds (VPCs) in Aiven Create and manage a virtual private cloud (VPC) for your Aiven project or organization. ## [Manage project VPCs](/docs/platform/howto/manage-project-vpc.md) [Set up or delete a project-wide VPC in your Aiven organization.](/docs/platform/howto/manage-project-vpc.md) ## [Manage organization VPCs](/docs/platform/howto/manage-organization-vpc.md) [Set up or delete an organization-wide VPC on the Aiven Platform.](/docs/platform/howto/manage-organization-vpc.md) --- # Use marketplace subscriptions to pay for Aiven services You can change the payment method for a billing group to a marketplace subscription. This lets you pay for your Aiven services through the AWS, Azure, or Google Cloud marketplaces. To change to a marketplace payment method, contact the support team with your subscription details. Changing from direct billing to a marketplace subscription does not disrupt your services. * AWS Marketplace * Azure Marketplace * Google Cloud Marketplace 1. [Set up your AWS Marketplace for Aiven subscription](/docs/marketplace-setup.md). 2. Collect information about your accounts: * In the Aiven Console, copy your [billing group ID](/docs/platform/reference/get-resource-IDs.md). * In the AWS Marketplace, copy your subscription ID. 3. Send the information to the [support team](/docs/platform/howto/support.md#create-a-support-ticket). 1) [Set up your Azure Marketplace for Aiven subscription](/docs/marketplace-setup.md). 2) Collect information about your accounts: * In the [Aiven Console](https://console.aiven.io/), copy your [billing group ID](/docs/platform/reference/get-resource-IDs.md). * In the Azure Marketplace, copy your subscription ID. 3) Send the information to the [support team](/docs/platform/howto/support.md#create-a-support-ticket). 1. [Set up your Google Cloud Marketplace for Aiven subscription](/docs/marketplace-setup.md). 2. Collect information about your accounts: * In the [Aiven Console](https://console.aiven.io/), copy your [billing group ID](/docs/platform/reference/get-resource-IDs.md). * In the Google Cloud Marketplace, copy your order ID. 3. Send the information to the [support team](/docs/platform/howto/support.md#create-a-support-ticket). --- # Metrics, logs, and alerts Use metrics, logs, alerts, and dashboards to monitor the health of your services and integrations. * Metrics: Real-time information about your services. * Logs: Available for services and integrations. * Alerts: Receive emails and in-app [notifications](/docs/platform/howto/technical-emails.md). tip Use an AI assistant connected to [Aiven MCP](/docs/tools/mcp-server.md) to retrieve service metrics and logs. For example: > Show CPU and memory usage for `my-service` during the past hour, and review the logs for errors during the same period. ## View service metrics[​](#view-service-metrics "Direct link to View service metrics") The service metrics available in [Aiven Console](https://console.aiven.io/) include the following: * **CPU usage:** Shows the percentage of CPU resources consumed by the service. * **Disk space usage:** Represents the percentage of disk space utilized by the service. * **Disk iops (reads):** Indicates the input/output operations per second (IOPS) for disk reads. * **Disk iops (writes):** Indicates the input/output operations per second (IOPS) for disk writes. * **Load average:** Shows the 5-minute average CPU load, indicating the system's computational load. * **Memory usage:** Represents the percentage of memory resources utilized by the service. * **Network received:** Indicates the amount of network traffic received by the service, measured in bytes per second. * **Network transmitted:** Indicates the amount of network traffic transmitted by the service, also measured in bytes per second. To view the metrics: 1. Open the service. 2. Click **Metrics**. To retrieve service logs with the CLI, use [service metrics](/docs/tools/cli/service-cli.md#avn-service-metrics). ## View service logs[​](#view-service-logs "Direct link to View service logs") 1. Open the service. 2. Click **Logs**. Log retention Service logs are retained for 4 days. To control retention time, [set up a log integration with an Aiven for OpenSearch® service](/docs/products/opensearch/howto/opensearch-log-integration.md). The integration allows you to configure longer retention times for your service logs, only limited by the disk space available on the Aiven for OpenSearch® plan you have selected. OpenSearch® together with OpenSearch® Dashboards offers comprehensive logs browsing and analytics platform. To retrieve service logs with the CLI, use [service logs](/docs/tools/cli/service-cli.md#avn-service-logs) ## Export service metrics and logs[​](#export-service-metrics-and-logs "Direct link to Export service metrics and logs") You can export logs and metrics **to an Aiven service**: * Send logs to [Aiven for Opensearch](/docs/products/opensearch/dashboards.md). * Send logs to [Aiven for Metrics](/docs/products/metrics.md). * Visualize logs with [Aiven for Grafana](/docs/products/grafana.md). ## Set up alerts and notifications[​](#set-up-alerts-and-notifications "Direct link to Set up alerts and notifications") See [Manage project and service notifications](/docs/platform/howto/technical-emails.md). Related pages * [Manage notifications](/docs/platform/howto/technical-emails.md) --- # Set up an organization VPC peering [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Establish a peering connection between your Aiven organization VPC and a cloud platform of your choice. ## [AWS peering](/docs/platform/howto/manage-org-vpc-peering-aws.md) [Set up a peering connection between your Aiven organization VPC and an AWS VPC.](/docs/platform/howto/manage-org-vpc-peering-aws.md) ## [Azure peering](/docs/platform/howto/manage-org-vpc-peering-azure.md) [Set up a peering connection between your Aiven organization VPC and a Microsoft Azure virtual network.](/docs/platform/howto/manage-org-vpc-peering-azure.md) ## [Google Cloud peering](/docs/platform/howto/manage-org-vpc-peering-google.md) [Set up a peering connection between your Aiven organization VPC and a Google Cloud VPC.](/docs/platform/howto/manage-org-vpc-peering-google.md) --- # Set up a project VPC peering Establish a peering connection between your Aiven project VPC and a cloud platform of your choice. ## [AWS peering](/docs/platform/howto/vpc-peering-aws.md) [Set up a peering connection between your Aiven project VPC and an AWS VPC.](/docs/platform/howto/vpc-peering-aws.md) ## [Azure peering](/docs/platform/howto/vnet-peering-azure.md) [Set up a peering connection between your Aiven project VPC and a Microsoft Azure virtual network.](/docs/platform/howto/vnet-peering-azure.md) ## [Google Cloud peering](/docs/platform/howto/vpc-peering-gcp.md) [Set up a peering connection between your Aiven project VPC and a Google Cloud VPC.](/docs/platform/howto/vpc-peering-gcp.md) ## [UpCloud peering](/docs/platform/howto/vpc-peering-upcloud.md) [Set up a peering connection between your Aiven project VPC and an UpCloud SDN network.](/docs/platform/howto/vpc-peering-upcloud.md) --- # Service management ## Overview[​](#overview "Direct link to Overview") Aiven manages a number of [services](/docs/products/services.md). These services share similar good practices when it comes to managing them. Understand the fundamentals: * [Service memory limits](/docs/platform/concepts/service-memory-limits.md) * [Service backups](/docs/platform/concepts/service_backups.md) * [Service maintenance](/docs/platform/concepts/maintenance-window.md) Create and configure your service: 1. [Create your service](/docs/platform/howto/create_new_service.md). 2. [Add users](/docs/platform/howto/create_new_service_user.md). 3. See the docs for your [specific service](/docs/products/services.md). Choosing a time series database Aiven offers a wide choice of open source time series databases in its product portfolio. * Aiven for PostgreSQL® with the TimescaleDB extension is your best choice if you already use PostgreSQL, require **SQL compatibility** and have a limited time series use case. * Aiven for ClickHouse® is your best choice when you need a high-performance columnar time series database for OLAP workloads or a data analytics warehouse. See our time series on [our website](https://aiven.io/time-series-databases/what-are-time-series-databases). Related pages ## [Overview](/docs/platform/howto/list-service.md) [Overview](/docs/platform/howto/list-service.md) ## [Create a service](/docs/platform/howto/create_new_service.md) [Create an Aiven service from the Aiven Console.](/docs/platform/howto/create_new_service.md) ## [Create service users](/docs/platform/howto/create_new_service_user.md) [Create additional users to access an Aiven service.](/docs/platform/howto/create_new_service_user.md) ## [Power on/off a service](/docs/platform/concepts/service-power-cycle.md) [Power off an Aiven service to release resources and save credits, power it back on when you need it, or delete it permanently.](/docs/platform/concepts/service-power-cycle.md) ## [Rename a service](/docs/platform/concepts/rename-services.md) [Change the name of an Aiven service by forking it under a new name and deleting the original service.](/docs/platform/concepts/rename-services.md) ## [Use resource tags](/docs/platform/howto/tag-resources.md) [Add key-value tags to an Aiven service to organize services and track ownership, cost allocation, and governance.](/docs/platform/howto/tag-resources.md) ## [Fork a service](/docs/platform/concepts/service-forking.md) [Fork an Aiven service to create an independent copy for testing, debugging, or development without affecting the original service.](/docs/platform/concepts/service-forking.md) ## [Scaling and performance](/docs/platform/howto/scale-services.md) [6 items](/docs/platform/howto/scale-services.md) ## [Maintenance and lifecycle](/docs/platform/concepts/maintenance-window.md) [2 items](/docs/platform/concepts/maintenance-window.md) ## [Backups and migration](/docs/platform/concepts/service_backups.md) [3 items](/docs/platform/concepts/service_backups.md) --- # User profiles Browse through instructions for common Aiven platform tasks related to managing your personal user profile information. --- # Virtual private cloud (VPC) peering in Aiven The VPC peering capability supported on the Aiven Platform improves network connectivity and security. It simplifies architecture, helps reduce network latency, and enhances resource sharing while maintaining isolation and control. [VPC](/docs/platform/concepts/vpcs.md) peering is a networking connection between two VPCs. It allows private and direct communication between the VPCs with no traffic routing over the public internet. ### VPC peering characteristics[​](#vpc-peering-characteristics "Direct link to VPC peering characteristics") * Private communication: Uses private IP addresses for direct communication between VPCs * High performance: Low latency thanks to traffic remaining on the cloud provider's network * Security: Reduces exposure to public networks without using internet gateways, VPNs, or NAT * Scalability: Supports connections across different accounts and regions, depending on a cloud provider ### VPC peering use cases[​](#vpc-peering-use-cases "Direct link to VPC peering use cases") * Multi-tier applications: Secure connection between VPCs hosting different application layers, such as web or database * Resource sharing: Secure sharing between VPCs hosting different resources, for example, datasets or APIs * Data isolation: Enforce access control by using separate VPCs for different projects or teams in an organization ## How it works[​](#how-it-works "Direct link to How it works") Aiven allows you to set up peering connections for [project VPCs](/docs/platform/concepts/vpcs.md#project-vpcs) and for [organization VPCs](/docs/platform/concepts/vpcs.md#vpc-types). ## [Project VPC peering](/docs/platform/howto/list-project-vpc-peering.md) [4 items](/docs/platform/howto/list-project-vpc-peering.md) ## [Organization VPC peering](/docs/platform/howto/list-organization-vpc-peering.md) [3 items](/docs/platform/howto/list-organization-vpc-peering.md) Aiven VPCs can be peered with VPCs in the following cloud platforms: * Google Cloud * Amazon Web Services * Microsoft Azure * UpCloud ## Limitations[​](#limitations "Direct link to Limitations") * Peering is only supported between an Aiven VPC and a VPC hosted with the same cloud provider. For example, an Aiven VPC hosted in AWS can be peered with an AWS VPC, but not with a VPC in Google Cloud, Microsoft Azure, or UpCloud. Peering across different cloud providers is not supported. * A single Aiven VPC can't be peered with VPCs in more than one cloud provider at the same time. To connect to a service from a different cloud provider or region, use [public access](/docs/platform/howto/public-access-in-vpc.md) instead of VPC peering. This lets you reach a VPC-hosted service over the public internet, and you can restrict who connects with an IP allow list. ## Learn more[​](#learn-more "Direct link to Learn more") For information on VPC peering supported by a particular cloud provider, see the following: * AWS: [VPC peering process, lifecycle, and limitations](https://docs.aws.amazon.com/vpc/latest/peering/vpc-peering-basics.html) * Google Cloud: [VPC Network Peering](https://cloud.google.com/vpc/docs/vpc-peering) * Azure: [Virtual network peering](https://learn.microsoft.com/en-us/azure/virtual-network/virtual-network-peering-overview) * UpCloud: [How to configure network peering](https://upcloud.com/docs/guides/configure-network-peering/) --- # Manage application users Application users give non-human users programmatic access to Aiven. You grant them access to organization resources using [roles and permissions](/docs/platform/concepts/permissions.md). You must be an [organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) to access this feature. important Application users can be a security risk if not carefully managed and monitored. Follow [best practices](/docs/platform/concepts/application-users.md#security-best-practices) for mitigating these risks. ## Create an application user[​](#create-an-application-user "Direct link to Create an application user") * Console * Terraform 1. Click **Admin**. 2. Click **Application users**. 3. Click **Create application user**. 4. Enter a name and click **Create application user**. ``` Loading... ``` More information on this resource and its configuration options are available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/organization_application_user). note If you reach the limit of application users, you can request an increase by [contacting Aiven support](/docs/platform/howto/support.md#create-a-support-ticket). ## Create a token for an application user[​](#create-a-token-for-an-application-user "Direct link to Create a token for an application user") * Console * Terraform 1. Click **Admin**. 2. Click **Application users**. 3. Click the application user's name. 4. In the **Authentication tokens** section, click **Generate token**. 5. Optional: Enter a description and session duration. 6. Optional: Add an IP address range to the allowlist for this token. 7. Click **Generate token**. 8. Click the **Copy** icon and save your token somewhere safe. important You cannot view the token after you close this window. 9. Click **Close**. note You cannot change the token's session duration or allowlist after creating it. ``` Loading... ``` More information on this resource and its configuration options are available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/organization_application_user_token). ## Revoke a token for an application user[​](#revoke-a-token-for-an-application-user "Direct link to Revoke a token for an application user") 1. Click **Admin** > **Application users**. 2. Click the name of the application user. 3. In the **Authentication tokens** section, click **Actions**. 4. Select **Revoke**. ## Check last activity of an application user[​](#check-last-activity-of-an-application-user "Direct link to Check last activity of an application user") * Console * Terraform 1. Click **Admin** > **Application users**. 2. Click the name of the application user. 3. To see the date and time that each application token was last used, go to the **Authentication tokens** section. Use the `last_activity_time` attribute in [the `aiven_organization_user_list` data source](https://registry.terraform.io/providers/aiven/aiven/latest/docs/data-sources/organization_user_list#last_activity_time-1) to check when one of the application user's tokens was last used. ## Delete an application user[​](#delete-an-application-user "Direct link to Delete an application user") 1. Click **Admin** > **Application users**. 2. Find the user and click **Actions** > **Delete**. --- # Pay with bank transfers Aiven offers invoice billing and bank transfer payments for customers who have at least 1,000 USD in monthly recurring revenue. Invoices are generated at the end of the month based on actual usage and a PDF is emailed to the [invoice contacts](/docs/platform/howto/use-billing-groups.md#update-a-billing-group) for each billing group. Prices for Aiven services are always in US dollars, but invoices can also be sent in the following currencies: * Australian dollars * Canadian dollars * Swiss francs * Danish kroner * Euros * Pounds sterling * Japanese yen * Norwegian kroner * New Zealand dollars * Swedish kronor Invoices in different currencies are created based on the exchange rates on the date of the invoice. To switch from credit card charges to bank transfers, contact to request invoice billing. --- # Manage billing and shipping addresses Create addresses in the billing section of the Aiven Platform and use them as billing and shipping addresses in your billing groups. You must have the `organization:billing:write` [permission](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) to create and manage billing addresses. note If you pay for your services using a marketplace subscription, you manage your billing address in your marketplace account. The billing address you add in the Aiven Console is not used for billing and they are not synced with the marketplace subscription. The address you add in the Aiven Console only keeps your Aiven organization account details up to date. ## Create an address[​](#create-an-address "Direct link to Create an address") 1. In the organization, click **Billing**. 2. Click **Addresses**. 3. Click **Create address**. 4. Enter the address details and click **Create address**. 5. Optional: Select the billing groups to assign the address to and click **Assign address**. note The address is added as both the billing and shipping address for the billing groups. You can [change the billing or shipping address](/docs/platform/howto/use-billing-groups.md#update-a-billing-group) in the billing groups. ## Update an address[​](#update-an-address "Direct link to Update an address") 1. In the organization, click **Billing**. 2. Click **Addresses**. 3. Find the address and click **Edit**. 4. Edit the address and click **Save changes**. ## Assign addresses to a billing group[​](#assign-addresses-to-a-billing-group "Direct link to Assign addresses to a billing group") To assign an address as a billing or shipping address for a billing group, [edit the billing group](/docs/platform/howto/use-billing-groups.md#update-a-billing-group). ## Delete an address[​](#delete-an-address "Direct link to Delete an address") You cannot delete an address that is assigned to a billing group. To delete an address that is assigned to a billing group, [assign a different address to the billing group](/docs/platform/howto/use-billing-groups.md#update-a-billing-group). 1. In the organization, click **Billing**. 2. Click **Addresses**. 3. Find the address and click **Delete**. --- # Manage domains Adding a verified domain in Aiven adds an extra layer of security to managing your organization's users. When you verify a domain, your organization users automatically become [managed users](/docs/platform/concepts/managed-users.md). There are two ways you can verify a domain: * by adding a DNS TXT record to the domain (recommended) * by uploading an HTML file to your website You can verify a domain in only one Aiven organization. To ensure your domain remains in the verified status, don't remove the verification file from your domain provider. ## Add a domain using a DNS TXT record[​](#add-a-domain-using-a-dns-txt-record "Direct link to Add a domain using a DNS TXT record") 1. In the organization, click **Admin**. 2. Click **Domains**. 3. Click **Add domain**. 4. Enter a **Domain name**. 5. In the **Verification method**, select **Add a DNS TXT record to your domain host**. 6. Click **Add domain**. 7. In the **Verification method** column, click **DNS TXT record**. 8. Copy the TXT record value. 9. In another browser tab or window, log in to your domain hosting provider. 10. Go to the DNS settings. 11. In the DNS settings for your domain provider, create a TXT record with the following: | Field name | Value | | ------------ | ---------------------------------------------------------------------------------- | | Name | `_aiven-challenge.{your domain}` | | Record value | The TXT record value you copied in the format `token=,expiry=never` | | Type | `TXT` | 12. In the Aiven Console, click **Actions** > **Verify**. It can take up to 72 hours for your DNS records to update the domain to be verified. If the domain is still not verified after that time, you can retry it by repeating the last step. ## Add a domain using an HTML file[​](#add-a-domain-using-an-html-file "Direct link to Add a domain using an HTML file") 1. In the organization, click **Admin**. 2. Click **Domains**. 3. Click **Add domain**. 4. Enter a **Domain name**. 5. In the **Verification method**, select **Upload an HTML file to your website**. 6. Click **Add domain**. 7. In the **Verification method** column, click **HTML file upload**. 8. Download the HTML file. 9. Upload the HTML file to your website in the path `/.well-known/aiven`. 10. In the Aiven Console, open the **Actions** > **Verify**. ## Remove a domain[​](#remove-a-domain "Direct link to Remove a domain") important Removing a domain is an irreversible action. 1. In the organization, click **Admin**. 2. Click **Domains**. 3. Find the domain and click **Actions** > **Remove**. --- # Manage groups of users Create groups of users in your organization to make it easier to manage access to your organization's resources. You can [grant permissions](/docs/platform/howto/manage-permissions.md) to groups for projects, giving them the right level of access to the project and its services. ## Create a group[​](#create-a-group "Direct link to Create a group") * Console * Terraform 1. Click **Admin** > **Groups**. 2. Click **Create group**. 3. Enter a unique name for the group. You can also enter a description. 4. Optional: To assign users to the group, click the toggle and choose the users to add. 5. Click **Create group**. ``` Loading... ``` More information on this resource and its configuration options are available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/organization_user_group). ## Add users to a group[​](#add-users-to-a-group "Direct link to Add users to a group") You can only add users that are [part of your organization](/docs/platform/howto/manage-org-users.md) to your groups. * Console * Terraform 1. Click **Admin** > **Groups**. 2. Select the group to add users to. 3. Click **Add users**. 4. Choose the users to add. 5. Click **Add users**. ``` Loading... ``` More information on this resource and its configuration options are available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/organization_user_group_member). ## Rename a group[​](#rename-a-group "Direct link to Rename a group") 1. Click **Admin** > **Groups**. 2. Find the group to rename and click **Actions** > **Rename**. 3. Enter the new name and click **Save changes**. ## Remove a group[​](#remove-a-group "Direct link to Remove a group") When you remove a group, the users in that group will lose access to any resources the group has permissions for unless they are part of another group with that access. 1. Click **Admin** > **Groups**. 2. Find the group to remove and click **Actions** > **Remove** and confirm. --- # Manage users in an organization Adding users to your organization lets you give them access to specific projects and services within that organization. important If you're using an identity provider (IdP), always add or remove users directly through the IdP. This ensures the IdP is the authoritative source for user management, preventing conflicts and simplifying administration. ## Invite users to an organization[​](#invite-users-to-an-organization "Direct link to Invite users to an organization") To add users to your organization, send them an invite: 1. Click **Admin** > **Users**. 2. Click **Invite users**. 3. Enter the email addresses of the people to invite. 4. Click **Invite users**. The users receive an email with instructions to sign up (for new users) and accept the invite. ## Remove users from an organization[​](#remove-users-from-an-organization "Direct link to Remove users from an organization") If you remove a user from an organization, they will also be removed from all groups and projects and no longer have access to any resources in the organization. To remove a user from an organization: 1. Click **Admin** > **Users**. 2. Find the user to remove and click **Actions** > **Remove** and confirm. ## Resend an invite[​](#resend-an-invite "Direct link to Resend an invite") To resend an invite to a user: 1. Click **Admin** > **Users**. 2. Find the email address to resend an invite to and click **Actions** > **Resend invite**. They receive a new email with instructions for signing up or accepting the invite. --- # Manage an organization VPC peering with AWS [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Set up a peering connection between your Aiven organization VPC and an AWS VPC. Before you start, review the [VPC peering limitations](/docs/platform/howto/list-vpc-peering.md#limitations). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Manage organization networking](/docs/platform/concepts/permissions.md#organization-permissions) permissions * Two VPCs to be peered: an [organization VPC](/docs/platform/howto/manage-organization-vpc.md#create-an-organization-vpc) in Aiven and a VPC in your AWS account * Access to the [AWS Management Console](https://console.aws.amazon.com) * One of the following tools for operations on the Aiven Platform: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) ## Create a peering connection[​](#create-a-peering-connection "Direct link to Create a peering connection") ### Collect data from AWS[​](#collect-data-from-aws "Direct link to Collect data from AWS") To [create a peering connection in Aiven](/docs/platform/howto/manage-org-vpc-peering-aws.md#create-a-peering-in-aiven), first collect the required data from AWS: 1. Log in to the [AWS Management Console](https://console.aws.amazon.com) and go to your profile information. 2. Find and save your account ID. 3. Open the navigation menu, and select **All services**. 4. Find **Networking & Content Delivery**, and go to **VPC** > **Your VPCs**. 5. Find a VPC to peer, preview its details, and save its ID and a cloud region that it's located in. ### Create a peering in Aiven[​](#create-a-peering-in-aiven "Direct link to Create a peering in Aiven") With the [data collected from AWS](/docs/platform/howto/manage-org-vpc-peering-aws.md#collect-data-from-aws), create an organization VPC peering connection using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API * Aiven Provider for Terraform 1. Log in to the [Aiven Console](https://console.aiven.io/), and click **Admin** in the top navigation bar. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, select an organization VPC to peer. 4. On the **VPC details** page, go to the **VPC peering connections** section and click **Create peering request**. 5. In the **Create peering request** window: 1. Enter the following: * **AWS account ID** * **AWS VPC region** * **AWS VPC ID** 2. Click **Create**. Run the `avn organization vpc peering-connection create` command: ``` avn organization vpc peering-connection create \ --organization-id AIVEN_ORGANIZATION_ID \ --organization-vpc-id AIVEN_ORGANIZATION_VPC_ID \ --peer-cloud-account AWS_ACCOUNT_ID \ --peer-vpc AWS_VPC_ID ``` Replace `AIVEN_ORGANIZATION_ID`, `AIVEN_ORGANIZATION_VPC_ID`, `AWS_ACCOUNT_ID`, and `AWS_VPC_ID` as needed. Make an API call to the [OrganizationVpcPeeringConnectionCreate](https://api.aiven.io/doc/#tag/Organization_Vpc/operation/OrganizationVpcPeeringConnectionCreate) endpoint: ``` curl --request POST \ --url https://api.aiven.io/v1/organization/ORGANIZATION_ID/vpcs/ORGANIZATION_VPC_ID/peering-connections \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data ' { "peer_cloud_account":"AWS_ACCOUNT_ID", "peer_vpc":"AWS_VPC_ID" } ' ``` Replace the following placeholders with meaningful data: * `ORGANIZATION_ID` * `ORGANIZATION_VPC_ID` * `BEARER_TOKEN` * `AWS_ACCOUNT_ID` * `AWS_VPC_ID` Use the [aiven\_aws\_org\_vpc\_peering\_connection](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/aws_org_vpc_peering_connection) resource. This adds a connection with the **Pending peer** status in the [Aiven Console](https://console.aiven.io/) and a connection pending acceptance in the [AWS Management Console](https://console.aws.amazon.com). ### Accept the peering request in AWS[​](#accept-the-peering-request-in-aws "Direct link to Accept the peering request in AWS") 1. Log in to the [AWS Management Console](https://console.aws.amazon.com), open the navigation menu, and select **All services**. 2. Find **Networking & Content Delivery**, and go to **VPC** > **Peering connections**. 3. Find your peering connection from Aiven pending acceptance, select it, and click **Actions** > **Accept request**. 4. Create or update your [AWS route tables](https://docs.aws.amazon.com/vpc/latest/peering/vpc-peering-routing) to match your Aiven CIDR settings. At this point, your peering connection status should be visible as **Active** both in the [Aiven Console](https://console.aiven.io/) and in the [AWS Management Console](https://console.aws.amazon.com). ## Delete the peering[​](#delete-the-peering "Direct link to Delete the peering") important Once you delete your VPC peering on the Aiven Platform, the cloud-provider side of the peering connection becomes `inactive` or `deleted`, and the traffic between the disconnected VPCs is terminated. Delete an organization VPC peering using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/), and click **Admin** in the top navigation bar. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, select an organization VPC. 4. On the **VPC details** page, go to the **VPC peering connections** section, find the peering to be deleted, and click **Actions** > **Delete**. 5. In the **Confirmation** window, click **Delete VPC peering**. Run the `avn organization vpc peering-connection delete` command: ``` avn organization vpc peering-connection delete \ --organization-id ORGANIZATION_ID \ --organization-vpc-id ORGANIZATION_VPC_ID \ --peering-connection-id ORGANIZATION_VPC_PEERING_ID ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization, for example, `org1a2b3c4d5e6` * `ORGANIZATION_VPC_ID` with the ID of your Aiven organization VPC, for example, `12345678-1a2b-3c4d-5f6g-1a2b3c4d5e6f` * `ORGANIZATION_VPC_PEERING_ID` with your Aiven peering connection ID obtainable in the output of the [avn organization vpc get](/docs/tools/cli/vpc.md#command-avn-organization-vpc-get) command, for example `1a2b3c4d-1234-a1b2-c3d4-1a2b3c4d5e6f` Make an API call to the [OrganizationVpcPeeringConnectionDelete](https://api.aiven.io/doc/#tag/Organization_Vpc/operation/OrganizationVpcPeeringConnectionDeleteById) endpoint: ``` curl --request DELETE \ --url https://api.aiven.io/v1/organization/ORGANIZATION_ID/vpcs/ORGANIZATION_VPC_ID/peering-connections/PEERING_CONNECTION_ID \ --header 'Authorization: Bearer BEARER_TOKEN' ``` Replace the following placeholders with meaningful data: * `ORGANIZATION_ID`: Aiven organization ID * `ORGANIZATION_VPC_ID`: Aiven organization VPC ID * `PEERING_CONNECTION_ID`: Aiven peering connection ID obtainable by calling the [OrganizationVpcGet](https://api.aiven.io/doc/#tag/Organization_Vpc/operation/OrganizationVpcGet) endpoint * `BEARER_TOKEN` --- # Manage an organization VPC peering with Microsoft Azure Set up a peering connection between your [Aiven organization VPC](/docs/platform/howto/manage-organization-vpc.md) and a [Microsoft Azure virtual network](https://learn.microsoft.com/en-us/azure/virtual-network/create-peering-different-subscriptions). Establishing a peering connection between an Aiven organization VPC and an Azure VNet requires creating the peering both from the VPC in Aiven and from the VNet in Azure. To establish the peering from Aiven to Azure, the Aiven Platform's [Active Directory application object](https://learn.microsoft.com/en-us/azure/active-directory/develop/app-objects-and-service-principals) needs permissions in your Azure subscription. Because the peering is between different AD tenants (the Aiven AD tenant and your Azure AD tenant), your Azure AD tenant needs another application object. Once granted permissions, this object allows peering from Azure to Aiven. Before you start, review the [VPC peering limitations](/docs/platform/howto/list-vpc-peering.md#limitations). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Manage organization networking](/docs/platform/concepts/permissions.md#organization-permissions) permissions for the Aiven Platform * Azure account with at least the application administrator role * [Azure CLI](https://learn.microsoft.com/en-us/cli/azure/?view=azure-cli-latest) and, optionally, the [Microsoft Azure portal](https://portal.azure.com/#home) * Access to the [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) installed * Two VPCs to be peered: an [organization VPC](/docs/platform/howto/manage-organization-vpc.md#create-an-organization-vpc) in Aiven and a VNet in your Azure account ## Set up permissions in Azure[​](#set-up-permissions-in-azure "Direct link to Set up permissions in Azure") ### Azure app object permissions[​](#azure-app-object-permissions "Direct link to Azure app object permissions") 1. Log in with an Azure admin account using the Azure CLI: ``` az account clear az login ``` This should open a window in your browser prompting to choose an Azure account to log in with. tip If you manage multiple Azure subscriptions, also configure the Azure CLI to default to the correct subscription for the subsequent commands. This is not needed if there's only one subscription: ``` az account set --subscription SUBSCRIPTION_NAME_OR_ID ``` 2. Create an application object in your AD tenant using the Azure CLI: ``` az ad app create \ --display-name "NAME_OF_YOUR_CHOICE" \ --sign-in-audience AzureADMultipleOrgs \ --key-type Password ``` This creates an application object in Azure AD that can be used to log into multiple AD tenants ( `--sign-in-audience AzureADMultipleOrgs` ), but only the home tenant (the tenant the app was created in) has the credentials to authenticate the app. note Save the `appId` field from the output. It will be referred to as `USER_APP_ID`. 3. Create a service principal for your app object in the Azure subscription where the VNet to be peered is located in: ``` az ad sp create --id USER_APP_ID ``` This creates a service principal in your subscription, which can be assigned permissions to peer your VNet. note Save the `id` field from the JSON output. It will be referred to as `USER_SP_ID`. 4. Set a password for your app object: ``` az ad app credential reset --id USER_APP_ID ``` note Save the `password` field from the output. It will be referred to as `USER_APP_SECRET`. 5. Find properties of your virtual network: * Resource ID * In the [Azure portal](https://portal.azure.com/#home): **Virtual networks** > name of your network > **JSON View** > **Resource ID** * Using the Azure CLI: ``` az network vnet list ``` note The `id` field should have format `/subscriptions/USER_SUBSCRIPTION_ID/ resourceGroups/USER_RESOURCE_GROUP/providers/Microsoft.Network/virtualNetworks/USER_VNET_NAME`. It will be referred to as `USER_VNET_ID`. * Azure Subscription ID (the VNet page in the [Azure portal](https://portal.azure.com/#home) > **Essentials** > **Subscription ID**) or the part after `/subscriptions/` in the resource ID. It will be referred to as `USER_SUBSCRIPTION_ID`. * Resource group name (the VNet page in the [Azure portal](https://portal.azure.com/#home) > **Essentials** > **Resource group**) or the `resourceGroup` field in the output. This will be referred to as `USER_RESOURCE_GROUP`. * VNet name (title of the VNet page), or the `name` field from the output. It will be referred to as `USER_VNET_NAME`. note Save all the properties for later. 6. Grant your service principal permissions to peer. The service principal needs to be assigned a role that includes the `Microsoft.Network/virtualNetworks/virtualNetworkPeerings/write` permission at the scope of your VNet. To limit the permissions granted to the application object and the service principal, you can create a custom role with only this permission. The built-in Network Contributor role also includes this permission. 1. Find the id of the role with the required permission: ``` az role definition list --name "Network Contributor" ``` The `id` field in the output is referred to as `NETWORK_CONTRIBUTOR_ROLE_ID`. 2. Assign the service principal the network contributor role using `NETWORK_CONTRIBUTOR_ROLE_ID`: ``` az role assignment create \ --role NETWORK_CONTRIBUTOR_ROLE_ID \ --assignee-object-id USER_SP_ID \ --scope USER_VNET_ID ``` This allows your application object to manage the network within the specified `--scope`. Since you control the application object, you can also grant it permissions at the scope of an entire resource group or the whole subscription. This enables creating other peerings later without assigning the role to each VNet separately. ### Aiven app object permissions[​](#aiven-app-object-permissions "Direct link to Aiven app object permissions") 1. Create a service principal for the Aiven application object. The Aiven AD tenant contains an application object that the Aiven Platform uses to create a peering from the Aiven organization VPC to the Azure VNet. For this, the Aiven application object needs a service principal in your Azure subscription. To create it, run: ``` az ad sp create --id 55f300d4-fc50-4c5e-9222-e90a6e2187fb ``` The argument to `--id` field is a fixed value that represents the ID of the Aiven application object. note Save the `id` field from the JSON output. It will be referred to as `AIVEN_SP_ID`. important The command might fail for the following reasons: * `When using this permission, the backing application of the service principal being created must in the local tenant`, which means your account doesn't have the required permissions. See [Prerequisites](/docs/platform/howto/vnet-peering-azure.md#prerequisites). * `The service principal cannot be created, updated, or restored because the service principal name 55f300d4-fc50-4c5e-9222-e90a6e2187fb is already in use`, in which case run `az ad sp show --id 55f300d4-fc50-4c5e-9222-e90a6e2187fb` and find `id` in the output. 2. Create a custom role for the Aiven application object. The Aiven application has a service principal that can be granted permissions. To restrict the service principal's permissions to peering, create a custom role with the peering action only allowed: ``` az role definition create --role-definition '{ "Name": "NAME_OF_YOUR_CHOICE", "Description": "Allows creating a peering to vnets in scope (but not from)", "Actions": [ "Microsoft.Network/virtualNetworks/peer/action" ], "AssignableScopes": [ "/subscriptions/USER_SUBSCRIPTION_ID" ] }' ``` `AssignableScopes` includes your Azure subscription ID to restrict scopes that a role assignment can use. note Save the `id` field from the output. It will be referred to as `AIVEN_ROLE_ID`. 3. Assign the custom role to the Aiven service principal. To give the Aiven application object's service principal permissions to peer with your VNet, assign the created role to the Aiven service principal with the scope of your VNet: ``` az role assignment create \ --role AIVEN_ROLE_ID \ --assignee-object-id AIVEN_SP_ID \ --scope USER_VNET_ID ``` 4. Find your AD tenant ID: * In the [Azure portal](https://portal.azure.com/#home): **Settings** > **Directories + subscriptions** > **Directories** > **Directory ID** * Using the Azure CLI: ``` az account list ``` note Save the `tenantId` field from the output. It will be referred to as `USER_TENANT_ID`. ## Create the peering in Aiven[​](#create-the-peering-in-aiven "Direct link to Create the peering in Aiven") * Aiven CLI * Aiven Console * Aiven API * Aiven Provider for Terraform By creating a peering from the Aiven organization VPC to the VNet in your Azure subscription, you also create a service principal for the application object (`--peer-azure-app-id USER_APP_ID`) and grant it the permission to peer with the Aiven organization VPC. The Aiven application object authenticates with your Azure tenant to grant it access to [the service principal of the Aiven application object](/docs/platform/howto/vnet-peering-azure.md#aiven-app-object-permissions) (`--peer-azure-tenant-id USER_TENANT_ID`). 1. [Find your organization ID in the Aiven Console](/docs/platform/reference/get-resource-IDs.md#get-an-organization-id) or retrieve your organization ID from the output of the `avn organization list` command. The organization ID will be referred to as `AIVEN_ORG_ID`. 2. Find your Aiven organization VPC ID using either the [Aiven Console](https://console.aiven.io/) or the [Aiven CLI](/docs/tools/cli.md). * Aiven Console * Aiven CLI 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. Go to your organization, and click **Admin** in the top navigation bar. 3. Click **VPCs** in the sidebar. 4. On the **Virtual private clouds** page, select your organization VPC. 5. On the **VPC details** page, go to the **Overview** section, and copy **ID**. In the [Aiven CLI](/docs/tools/cli.md), run the [avn organization vpc list](/docs/tools/cli/vpc.md) command. The Aiven organization VPC ID will be referred to as `AIVEN_ORGANIZATION_VPC_ID`. 3. Run: ``` avn organization vpc peering-connection create \ --organization-id AIVEN_ORG_ID \ --organization-vpc-id AIVEN_ORGANIZATION_VPC_ID \ --peer-cloud-account USER_SUBSCRIPTION_ID \ --peer-resource-group USER_RESOURCE_GROUP \ --peer-vpc USER_VNET_NAME \ --peer-azure-app-id USER_APP_ID \ --peer-azure-tenant-id USER_TENANT_ID ``` note Use lower case for arguments starting with `USER_`. 4. Run the following command until the state changes from `APPROVED` to `PENDING_PEER`: ``` avn organization vpc peering-connection list \ --organization-id AIVEN_ORG_ID \ --organization-vpc-id AIVEN_ORGANIZATION_VPC_ID ``` tip If the state is `INVALID_SPECIFICATION` or `REJECTED_BY_PEER`, check if the Azure VNet exists and if the Aiven application object has the permission to be peered with. Revise your configuration and recreate the peering connection. Establishing the connection from Aiven to Azure can take a while. When completed, the state changes to `PENDING_PEER` and the output shows details for establishing the peering from your Azure VNet to the Aiven organization VPC. note Save the following from the output: * `to-tenant-id`: It will be referred to as `AIVEN_TENANT_ID`. * `to-network-id`: It will be referred to as `AIVEN_VNET_ID`. 1) Log in to the [Aiven Console](https://console.aiven.io/), and click **Admin** in the top navigation bar. 2) Click **VPCs** in the sidebar. 3) On the **Virtual private clouds** page, select an organization VPC to peer. 4) On the **VPC details** page, go to the **VPC peering connections** section and click **Create peering request**. 5) In the **Create peering request** window: 1. Enter the following: * **Azure subscription ID** * **Resource group** * **Network name** * **Active Directory tenant ID** * **Application object ID** 2. Click **Create**. This adds a connection with the **Pending peer** status in the [Aiven Console](https://console.aiven.io/). 6) While still on the **VPC details** page, make a note of the **ID** of your Aiven VPC. Make an API call to the [OrganizationVpcPeeringConnectionCreate](https://api.aiven.io/doc/#tag/Organization_Vpc/operation/OrganizationVpcPeeringConnectionCreate) endpoint: ``` curl --request POST \ --url https://api.aiven.io/v1/organization/ORGANIZATION_ID/vpcs/ORGANIZATION_VPC_ID/peering-connections \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data ' { "peer_azure_app_id":"USER_APP_ID", "peer_azure_tenant_id":"USER_TENANT_ID", "peer_cloud_account":"USER_SUBSCRIPTION_ID", "peer_resource_group":"USER_RESOURCE_GROUP", "peer_vpc":"USER_VNET_NAME" } ' ``` Replace the following placeholders with meaningful data: * `ORGANIZATION_ID` * `ORGANIZATION_VPC_ID` * `BEARER_TOKEN` * `USER_SUBSCRIPTION_ID` * `USER_RESOURCE_GROUP` * `USER_VNET_NAME` * `USER_APP_ID` * `USER_TENANT_ID` Use the [aiven\_azure\_org\_vpc\_peering\_connection](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/azure_org_vpc_peering_connection) resource. ## Create the peering in Azure[​](#create-the-peering-in-azure "Direct link to Create the peering in Azure") [Establish the peering connection](https://learn.microsoft.com/en-us/azure/virtual-network/create-peering-different-subscriptions) from your Azure VNet to the Aiven organization VPC: 1. Log out the Azure user [you logged in with](/docs/platform/howto/vnet-peering-azure.md#azure-app-object-permissions): ``` az account clear ``` 2. Log in the Azure application object to your AD tenant using the [password](/docs/platform/howto/vnet-peering-azure.md#azure-app-object-permissions): ``` az login \ --service-principal \ -u USER_APP_ID \ -p USER_APP_SECRET \ --tenant USER_TENANT_ID ``` 3. Log in the Azure application object to the Aiven AD tenant: ``` az login \ --service-principal \ -u USER_APP_ID \ -p USER_APP_SECRET \ --tenant AIVEN_TENANT_ID ``` At this point, your application object should have an open session with your Azure AD tenant and the Aiven AD tenant. 4. Create a peering from your Azure VNet to the Aiven organization VPC: ``` az network vnet peering create \ --name PEERING_NAME_OF_YOUR_CHOICE \ --remote-vnet AIVEN_VNET_ID \ --vnet-name USER_VNET_NAME \ --resource-group USER_RESOURCE_GROUP \ --subscription USER_SUBSCRIPTION_ID \ --allow-vnet-access ``` If the peering state in the output is `connected`, the peering is created. tip The command might fail with the following error: ``` The client 'RANDOM_UUID' with object id 'RANDOM_UUID' does not have authorization to perform action 'Microsoft.Network/virtualNetworks/virtualNetworkPeerings/write' over scope 'USER_VNET_ID'. If access was recently granted, refresh your credentials. ``` for two reasons related to the [role assignment](/docs/platform/howto/vnet-peering-azure.md#azure-app-object-permissions): * Role assignment hasn't taken effect yet, in which case try logging in again and recreating the peering. * Role assignment is incorrect, in which case try recreating the role assignment. Wait until the Aiven peering connection is active. The Aiven Platform polls peering connections in state `PENDING_PEER` regularly to see if the peer (your Azure VNet) has created a peering connection to the Aiven organization VPC. Once this is detected, the state changes from `PENDING_PEER` to `ACTIVE`, at which point Aiven services in the organization VPC can be reached through the peering. 5. Check if the status of the peering connection is `ACTIVE`: ``` avn organization vpc get \ --organization-id AIVEN_ORG_ID \ --organization-vpc-id AIVEN_ORGANIZATION_VPC_ID ``` ## Delete the peering[​](#delete-the-peering "Direct link to Delete the peering") important Once you delete your VPC peering on the Aiven Platform, the cloud-provider side of the peering connection becomes `inactive` or `deleted`, and the traffic between the disconnected VPCs is terminated. Delete an organization VPC peering using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/), and click **Admin** in the top navigation bar. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, select an organization VPC. 4. On the **VPC details** page, go to the **VPC peering connections** section, find the peering to be deleted, and click **Actions** > **Delete**. 5. In the **Confirmation** window, click **Delete VPC peering**. Run the `avn organization vpc peering-connection delete` command: ``` avn organization vpc peering-connection delete \ --organization-id ORGANIZATION_ID \ --organization-vpc-id ORGANIZATION_VPC_ID \ --peering-connection-id ORGANIZATION_VPC_PEERING_ID ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization, for example, `org1a2b3c4d5e6` * `ORGANIZATION_VPC_ID` with the ID of your Aiven organization VPC, for example, `12345678-1a2b-3c4d-5f6g-1a2b3c4d5e6f` * `ORGANIZATION_VPC_PEERING_ID` with your Aiven peering connection ID obtainable in the output of the [avn organization vpc get](/docs/tools/cli/vpc.md#command-avn-organization-vpc-get) command, for example `1a2b3c4d-1234-a1b2-c3d4-1a2b3c4d5e6f` Make an API call to the [OrganizationVpcPeeringConnectionDelete](https://api.aiven.io/doc/#tag/Organization_Vpc/operation/OrganizationVpcPeeringConnectionDeleteById) endpoint: ``` curl --request DELETE \ --url https://api.aiven.io/v1/organization/ORGANIZATION_ID/vpcs/ORGANIZATION_VPC_ID/peering-connections/PEERING_CONNECTION_ID \ --header 'Authorization: Bearer BEARER_TOKEN' ``` Replace the following placeholders with meaningful data: * `ORGANIZATION_ID`: Aiven organization ID * `ORGANIZATION_VPC_ID`: Aiven organization VPC ID * `PEERING_CONNECTION_ID`: Aiven peering connection ID obtainable by calling the [OrganizationVpcGet](https://api.aiven.io/doc/#tag/Organization_Vpc/operation/OrganizationVpcGet) endpoint * `BEARER_TOKEN` ## Related pages[​](#related-pages "Direct link to Related pages") * [Manage organization VPCs](/docs/platform/howto/manage-organization-vpc.md) * [Set up an organization VPC peering](/docs/platform/howto/list-organization-vpc-peering.md) * [Manage project VPCs](/docs/platform/howto/manage-project-vpc.md) --- # Manage organization VPC peering with Google Cloud [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Set up a peering connection between your Aiven organization VPC and a Google Cloud VPC. Establishing a peering connection between an Aiven VPC and a Google Cloud VPC requires creating the peering both from the VPC in Aiven and from the VPC in Google Cloud. Before you start, review the [VPC peering limitations](/docs/platform/howto/list-vpc-peering.md#limitations). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Manage organization networking](/docs/platform/concepts/permissions.md#organization-permissions) permissions * Two VPCs to be peered: an [organization VPC](/docs/platform/howto/manage-organization-vpc.md#create-an-organization-vpc) in Aiven and a VPC in your Google Cloud account * Access to the [Google Cloud console](https://console.cloud.google.com/) * One of the following tools for operations on the Aiven Platform: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) ## Create a peering connection[​](#create-a-peering-connection "Direct link to Create a peering connection") ### Collect data from Google Cloud[​](#collect-data-from-google-cloud "Direct link to Collect data from Google Cloud") To [create a peering connection in Aiven](/docs/platform/howto/manage-org-vpc-peering-google.md#create-the-peering-in-aiven), first collect the required data from Google Cloud: 1. Log in to the [Google Cloud console](https://console.cloud.google.com/), open the navigation menu, and select **Cloud overview** > **Dashboard**. 2. Find the **Project info** field, and collect your **Project ID**. 3. Open the navigation menu again, and click **VIEW ALL PRODUCTS** > **Networking** > **VPC Network**. 4. Find a VPC to connect to, and make note of its **Name**. ### Create the peering in Aiven[​](#create-the-peering-in-aiven "Direct link to Create the peering in Aiven") With the [data collected from Google Cloud](/docs/platform/howto/manage-org-vpc-peering-google.md#collect-data-from-google-cloud), create an organization VPC peering connection using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API * Aiven Provider for Terraform 1. Log in to the [Aiven Console](https://console.aiven.io/), and click **Admin** in the top navigation bar. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, select an organization VPC to peer. 4. On the **VPC details** page, go to the **VPC peering connections** section and click **Create peering request**. 5. In the **Create peering request** window: 1. Enter the following: * **GCP project ID** * **GCP VPC network name** 2. Click **Create**. This adds a connection with the **Pending peer** status in the [Aiven Console](https://console.aiven.io/). 6. On the **VPC details** page, go to the **VPC peering connections** section and click **Status details** for the new connection. Make a note of your Aiven VPC network name and project ID. Run the `avn organization vpc peering-connection create` command: ``` avn organization vpc peering-connection create \ --organization-id AIVEN_ORGANIZATION_ID \ --organization-vpc-id AIVEN_ORGANIZATION_VPC_ID \ --peer-cloud-account GOOGLE_CLOUD_PROJECT_ID \ --peer-vpc GOOGLE_CLOUD_VPC_NETWORK_NAME ``` Replace `AIVEN_ORGANIZATION_ID`, `AIVEN_ORGANIZATION_VPC_ID`, `GOOGLE_CLOUD_PROJECT_ID`, and `GOOGLE_CLOUD_VPC_NETWORK_NAME` as needed. Make an API call to the [OrganizationVpcPeeringConnectionCreate](https://api.aiven.io/doc/#tag/Organization_Vpc/operation/OrganizationVpcPeeringConnectionCreate) endpoint: ``` curl --request POST \ --url https://api.aiven.io/v1/organization/ORGANIZATION_ID/vpcs/ORGANIZATION_VPC_ID/peering-connections \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data ' { "peer_cloud_account":"GOOGLE_CLOUD_PROJECT_ID", "peer_vpc":"GOOGLE_CLOUD_VPC_NETWORK_NAME" } ' ``` Replace the following placeholders with meaningful data: * `ORGANIZATION_ID` * `ORGANIZATION_VPC_ID` * `BEARER_TOKEN` * `GOOGLE_CLOUD_PROJECT_ID` * `GOOGLE_CLOUD_VPC_NETWORK_NAME` Use the [aiven\_gcp\_org\_vpc\_peering\_connection](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/gcp_org_vpc_peering_connection) resource. ### Create the peering in Google Cloud[​](#create-the-peering-in-google-cloud "Direct link to Create the peering in Google Cloud") Use the [data collected in the Aiven Console](/docs/platform/howto/manage-org-vpc-peering-google.md#create-the-peering-in-aiven) to create the VPC peering connection in Google Cloud: 1. Log in to the [Google Cloud console](https://console.cloud.google.com/), open the navigation menu, and click **VIEW ALL PRODUCTS** > **Networking** > **VPC Network** > **VPC network peering** > **CREATE PEERING CONNECTION** > **CONTINUE**. 2. Enter a name for the peering connection. 3. Select your Google Cloud VPC network. 4. In the **Peered VPC network** field, select **In another project**. 5. In the **Project ID** field, enter the Aiven project ID collected in the the [Aiven Console](https://console.aiven.io), on the **VPC details** page, in the **VPC peering connections** section, from **Status details** of the new connection. 6. In the **VPC network name** field, enter the name of your Aiven VPC collected in the the [Aiven Console](https://console.aiven.io), on the **VPC details** page, in the **VPC peering connections** section, from **Status details** of the new connection. 7. Click **Create**. As soon as the peering is created, the connection status changes to **Active** both in the [Aiven Console](https://console.aiven.io) and in the [Google Cloud console](https://console.cloud.google.com/). ## Set up multiple organization VPC peerings[​](#set-up-multiple-organization-vpc-peerings "Direct link to Set up multiple organization VPC peerings") To peer multiple Google Cloud VPC networks to your Aiven-managed organization VPC, [add peering connections](/docs/platform/howto/manage-org-vpc-peering-google.md#create-a-peering-connection) one at a time in the [Aiven Console](https://console.aiven.io). For the limit on the number of VPC peering connections allowed to a single VPC network, see the [Google Cloud documentation](https://cloud.google.com/vpc/docs/quota). ## Delete the peering[​](#delete-the-peering "Direct link to Delete the peering") important Once you delete your VPC peering on the Aiven Platform, the cloud-provider side of the peering connection becomes `inactive` or `deleted`, and the traffic between the disconnected VPCs is terminated. Delete an organization VPC peering using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/), and click **Admin** in the top navigation bar. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, select an organization VPC. 4. On the **VPC details** page, go to the **VPC peering connections** section, find the peering to be deleted, and click **Actions** > **Delete**. 5. In the **Confirmation** window, click **Delete VPC peering**. Run the `avn organization vpc peering-connection delete` command: ``` avn organization vpc peering-connection delete \ --organization-id ORGANIZATION_ID \ --organization-vpc-id ORGANIZATION_VPC_ID \ --peering-connection-id ORGANIZATION_VPC_PEERING_ID ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization, for example, `org1a2b3c4d5e6` * `ORGANIZATION_VPC_ID` with the ID of your Aiven organization VPC, for example, `12345678-1a2b-3c4d-5f6g-1a2b3c4d5e6f` * `ORGANIZATION_VPC_PEERING_ID` with your Aiven peering connection ID obtainable in the output of the [avn organization vpc get](/docs/tools/cli/vpc.md#command-avn-organization-vpc-get) command, for example `1a2b3c4d-1234-a1b2-c3d4-1a2b3c4d5e6f` Make an API call to the [OrganizationVpcPeeringConnectionDelete](https://api.aiven.io/doc/#tag/Organization_Vpc/operation/OrganizationVpcPeeringConnectionDeleteById) endpoint: ``` curl --request DELETE \ --url https://api.aiven.io/v1/organization/ORGANIZATION_ID/vpcs/ORGANIZATION_VPC_ID/peering-connections/PEERING_CONNECTION_ID \ --header 'Authorization: Bearer BEARER_TOKEN' ``` Replace the following placeholders with meaningful data: * `ORGANIZATION_ID`: Aiven organization ID * `ORGANIZATION_VPC_ID`: Aiven organization VPC ID * `PEERING_CONNECTION_ID`: Aiven peering connection ID obtainable by calling the [OrganizationVpcGet](https://api.aiven.io/doc/#tag/Organization_Vpc/operation/OrganizationVpcGet) endpoint * `BEARER_TOKEN` --- # Manage organization virtual private clouds (VPCs) in Aiven [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Set up or delete an organization-wide VPC on the Aiven Platform. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Manage organization networking](/docs/platform/concepts/permissions.md#organization-permissions) permissions * One of the following tools for operating organization VPCs: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) ## Create an organization VPC[​](#create-an-organization-vpc "Direct link to Create an organization VPC") Create an organization VPC using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API * Aiven Provider for Terraform 1. Log in to the [Aiven Console](https://console.aiven.io/), and click **Admin** in the top navigation bar. 2. Click **VPCs** in the sidebar and **Create VPC** on the **Virtual private clouds** page. 3. In the **Create VPC** window: 1. Select a cloud provider. 2. Select a cloud region. 3. Specify an IP range. * Use an IP range that does not overlap with any networks to be connected via VPC peering. For example, if your own networks use the range `11.1.1.0/8`, you can set the range for your Aiven organization's VPC to `191.161.1.0/24`. * Use a network prefix that is 20-24 character long. 4. Click **Create VPC**. Your new organization VPC is ready to use as soon as its status visible on the **Virtual private clouds** page changes to **Active**. Run the `avn organization vpc create` command: ``` avn organization vpc create \ --cloud CLOUD_PROVIDER_REGION \ --network-cidr NETWORK_CIDR \ --organization-id ORGANIZATION_ID ``` Replace the following: * `CLOUD_PROVIDER_REGION` with the cloud provider and region to host the VPC, for example `aws-eu-west-1` * `NETWORK_CIDR` with the CIDR block (a range of IP addresses) for the VPC, for example, `10.0.0.0/24` * `ORGANIZATION_ID` with the ID of your Aiven organization where to create the VPC, for example, `org1a2b3c4d5e6` Make an API call to the [OrganizationVpcCreate](https://api.aiven.io/doc/#tag/Organization_Vpc/operation/OrganizationVpcCreate) endpoint: ``` curl --request POST \ --url https://api.aiven.io/v1/organization/ORGANIZATION_ID/vpcs \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data ' { "cloud_name": "CLOUD_PROVIDER_REGION", "network_cidr": "NETWORK_CIDR" } ' ``` Replace the following placeholders with meaningful data: * `ORGANIZATION_ID` * `BEARER_TOKEN` * `CLOUD_PROVIDER_REGION` * `NETWORK_CIDR` Use the [aiven\_organization\_vpc](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/organization_vpc) resource. ## Delete an organization VPC[​](#delete-an-organization-vpc "Direct link to Delete an organization VPC") important Remove all services from your VCP before you delete it. To remove the services from the VCP, either migrate them out of the VCP or delete them. Deleting the VPC terminates its peering connections, if any. Delete an organization VPC using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/), and click **Admin** in the top navigation bar. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, find a VPC to be deleted and click **Actions** > **Delete**. 4. In the **Confirmation** window, click **Delete VPC**. Run the `avn organization vpc delete` command: ``` avn organization vpc delete \ --organization-id ORGANIZATION_ID \ --organization-vpc-id ORGANIZATION_VPC_ID ``` Replace the following: * `ORGANIZATION_ID` with the ID of your Aiven organization, for example, `org1a2b3c4d5e6` * `ORGANIZATION_VPC_ID` with the ID of your Aiven organization VPC, for example, `12345678-1a2b-3c4d-5f6g-1a2b3c4d5e6f` Make an API call to the [OrganizationVpcDelete](https://api.aiven.io/doc/#tag/Organization_Vpc/operation/OrganizationVpcDelete) endpoint: ``` curl --request DELETE \ --url https://api.aiven.io/v1/organization/ORGANIZATION_ID/vpcs/ORGANIZATION_VPC_ID \ --header 'Authorization: Bearer BEARER_TOKEN' \ ``` Replace the following placeholders with meaningful data: * `ORGANIZATION_ID` * `ORGANIZATION_VPC_ID` * `BEARER_TOKEN` Related pages * [VPC peering](/docs/platform/howto/list-vpc-peering.md) * [Manage organization VPC peering connections](/docs/platform/howto/manage-org-vpc-peering-aws.md) --- # Manage organizations Learn how to manage your organizations via the Aiven Console. ## Delete an organization[​](#delete-an-organization "Direct link to Delete an organization") 1. Delete all [projects](/docs/platform/howto/manage-project.md#delete-a-project) in the organization and in the organizational units. 2. If you use a [marketplace subscription](/docs/platform/howto/list-marketplace-payments.md) to pay for your services, cancel the subscription in the marketplace. 3. Click **Admin**. 4. Click **Organization**. 5. Open each organizational unit by clicking its name then click **Delete** to delete it. 6. After all the organizational units have been deleted, on the **Organization** page click **Delete**. 7. To confirm, click **Delete**. note Billing groups with trial credits and the organizations they are assigned to cannot be deleted. You can delete them after the trial period ends. To delete the billing group, organization, or your account during the trial period [contact Aiven support](/docs/platform/howto/support.md). ## Rename an organization[​](#rename-an-organization "Direct link to Rename an organization") 1. In the organization, click **Admin**. 2. Click **Organization**. 3. Click **Rename**. 4. Select **Rename**. 5. Enter the new name. 6. Click **Rename**. ## Leave an organization[​](#leave-an-organization "Direct link to Leave an organization") 1. Go to **User information** > **Organizations**. 2. Find the organization and click **Leave**. --- # Manage credit cards Add credit cards to your organization and use them across different [billing groups](/docs/platform/howto/use-billing-groups.md) to pay for your Aiven services. ## Add a card[​](#add-a-card "Direct link to Add a card") You can add a credit card as a payment method in your organization and assign it to different billing groups. You must be an [organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) to access this feature. 1. Click **Billing**. 2. Click **Payment methods**. 3. Click **Add card**. 4. Enter the credit card details and click **Add card**. 5. Optional: Select billing groups to assign the card to and click **Assign payment card**. The card won't be charged for monthly payments if you don't assign it to a billing group. ## Delete a card[​](#delete-a-card "Direct link to Delete a card") To delete a credit card, [remove it from all billing groups](/docs/platform/howto/use-billing-groups.md) first. 1. Click **Billing**. 2. Click **Payment methods**. 3. On the **Cards** tab, find the card to delete. 4. Click **Delete**. --- # Manage permissions You can grant [organization users](/docs/platform/howto/manage-org-users.md), [application users](/docs/platform/concepts/application-users.md), and [groups](/docs/platform/howto/manage-groups.md) access at the organization, organizational unit, and project level through [roles and permissions](/docs/platform/concepts/permissions.md). If you don't grant any roles or permissions to an organization user, they have the [default access level](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) to the organization. important When you remove permissions from a user or group, service credentials are not changed. Users can still directly access services if they know the service credentials. To prevent this type of access, reset all service passwords. ## Organization and organizational unit permissions[​](#organization-and-organizational-unit-permissions "Direct link to Organization and organizational unit permissions") Grant permissions at the organization level for access to all organizational units and projects in the organization. Grant permissions at the organizational unit level to give users access to all projects in that unit. ### Grant organization or unit permissions to a user or group[​](#grant-organization-or-unit-permissions-to-a-user-or-group "Direct link to Grant organization or unit permissions to a user or group") * Console * Terraform 1. In the organization, click **Admin**. 2. Click **Permissions**. 3. Click **Grant permissions** and select **Grant to users** or **Grant to groups**. 4. Select the users or groups to grant permissions to. 5. In **Resource**, choose an organization or organizational unit. 6. Select the [roles and permissions](/docs/platform/concepts/permissions.md) to grant. 7. Click **Grant permissions**. Use [the `aiven_organization_permission` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/organization_permission) and set the `resource_type` to `organization` or `organization_unit`. ### Change organization or unit permissions for a user or group[​](#change-organization-or-unit-permissions-for-a-user-or-group "Direct link to Change organization or unit permissions for a user or group") You can change the permissions granted at the organization or organizational unit level for a user or group. In the Aiven Console, you cannot change the resource for existing permissions. To change the resource, remove the existing permissions and grant the permissions for the other resource. 1. In the organization, click **Admin**. 2. Click **Permissions**. 3. For the user or group click **Actions** > **Edit permissions**. 4. Add or remove permissions and click **Save changes**. ### Remove all organization or unit-level roles and permissions[​](#remove-all-organization-or-unit-level-roles-and-permissions "Direct link to Remove all organization or unit-level roles and permissions") You can remove all permissions that you granted to a user or group at the organization or organizational unit level. To remove all organization and unit permissions for a user or group: 1. In the organization, click **Admin**. 2. Click **Permissions**. 3. For the user or group click **Actions** > **Remove**. ### Make users super admin[​](#make-users-super-admin "Direct link to Make users super admin") The super admin role is a special role that has unrestricted access to an organization and all its resources and settings. You cannot make application users super admin. important This role should be limited to as few users as possible for organization setup and emergency use. For daily administrative tasks, assign users the [organization admin role](/docs/platform/concepts/permissions.md) instead. Aiven also highly recommends enabling [two-factor authentication](/docs/platform/howto/user-2fa.md) for super admin. 1. In the organization, click **Admin**. 2. Click **Users**. 3. Find the user and click **Actions** > **Make super admin**. To revoke super admin privileges for a user, follow the same steps and select **Revoke super admin**. ## Project permissions[​](#project-permissions "Direct link to Project permissions") You can give users access to a specific project by granting them roles and permissions at the project level. ### Grant project permissions to a user or group[​](#grant-project-permissions-to-a-user-or-group "Direct link to Grant project permissions to a user or group") * Console * Terraform 1. In the project, click **Permissions**. 2. Click **Grant permissions** and select **Grant to users** or **Grant to groups**. 3. Select the users or groups to add to the project. 4. Select the [roles and permissions](/docs/platform/concepts/permissions.md) to grant. 5. Click **Grant permissions**. Use [the `aiven_organization_permission` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/organization_permission) and set the `resource_type` to `project`. ### Change permissions for a user or group[​](#change-permissions-for-a-user-or-group "Direct link to Change permissions for a user or group") 1. In the project, click **Permissions**. 2. For the user or group click **Actions** > **Edit permissions**. 3. Add or remove permissions and click **Save changes**. ### Remove all project-level roles and permissions[​](#remove-all-project-level-roles-and-permissions "Direct link to Remove all project-level roles and permissions") To remove all permissions to a project: 1. In the project, click **Permissions**. 2. For the user or group click **Actions** > **Remove**. --- # Manage projects Projects help you [organize your Aiven services](https://aiven.io/docs/platform/concepts/orgs-units-projects#projects) and centrally configure settings for all services in the project. You can create projects in your organization or its organizational units. ## Create a project[​](#create-a-project "Direct link to Create a project") * Console * Terraform 1. Click **Projects** and select **Create project**. 2. Enter a name for the project. 3. Select an organization or organizational unit to add the project to. 4. Select a [billing group](/docs/platform/howto/use-billing-groups.md). The costs from all services in this project are charged to the payment method for this billing group. ``` Loading... ``` More information on this resource and its configuration options are available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/project). ## Tag a project[​](#tag-a-project "Direct link to Tag a project") * Console * Terraform 1. In the project, click **Settings**. 2. In the **Tags** section, click **Project** or **Billing**. * Billing tags are returned in the invoice API and displayed on PDF invoices for the project. * Project tags are returned for resources in the API and displayed in the list of projects. 3. Click **Add tags**. 4. Enter a key and value for each tag. 5. Click **Save changes**. To add billing and project tags, use the `tag` attribute in [your `aiven_project` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/project#nestedblock--tag). For billing tags, prefix the `key` with `billing:`. For example: `key = "billing:PO"`. ## Rename a project[​](#rename-a-project "Direct link to Rename a project") important The project name in your DNS records will not be updated. 1. Power off all services in the project. note Except for Aiven for Apache Kafka®, all services have backups that are restored when you power them back on. 2. In the project, click **Settings**. 3. In the **Project settings**, edit the **Project name**. 4. Click **Save changes**. ## Move a project[​](#move-a-project "Direct link to Move a project") You can move a project to another organization or organizational unit. Users with the organization admin or project admin [role](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) can move projects within an organization. note You cannot move a project to an Aiven organization on AWS Marketplace, Azure Marketplace, or Google Cloud Marketplace using the Aiven Console. To move to a marketplace, set up your organization and projects in the marketplace and [contact Aiven for help moving your services](/docs/platform/howto/list-marketplace-payments.md). To move a project to a different organization, you must be an organization admin of both organizations. All users with permission to access the project lose the permissions when you move it to a different organization unless they are members of the target organization. Services in the project continue running during the move. * Console * Terraform 1. In the organization with the project, click **Admin**. 2. Click **Projects** and find the project to move. 3. Click > **Move project**. 4. Select the organization or organizational unit to move the project to. 5. Select a **Billing group**. 6. Click **Next** and **Finish**. You can move a project to another organization or organizational unit by [updating your `aiven_project` or `aiven_organization_project` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/guides/move-projects). ## Delete a project[​](#delete-a-project "Direct link to Delete a project") 1. Delete all the services in the project. 2. In the project, click **Settings**. 3. Click **Delete** and **Confirm**. --- # Manage project virtual private clouds (VPCs) in Aiven Set up or delete a project-wide VPC in your Aiven organization. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Manage project networking](/docs/platform/concepts/permissions.md#project-permissions) permissions * One of the following tools for operating project VPCs: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) ## Create a project VPC[​](#create-a-project-vpc "Direct link to Create a project VPC") Create a project VPC using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API * Aiven Provider for Terraform 1. Log in to [Aiven Console](https://console.aiven.io/), go to your project page, and click **VPCs** in the sidebar. 2. On the **Virtual private clouds** page, click **Create VPC**. 3. In the **Create VPC** window: 1. Select a cloud provider and region. 2. Enter the IP range. Use an IP range that does not overlap with any networks that you want to connect via VPC peering. For example, if your own networks use the range `11.1.1.0/8`, you can set the range for your Aiven project's VPC to `191.161.1.0/24`. note Network prefix length must be between 20 and 24 inclusive. 4. Click **Create VPC**. The state of the VPC is shown in the table. Run the [avn vpc create](/docs/tools/cli/vpc.md#create-vpcs) command: ``` avn vpc create \ --cloud CLOUD_PROVIDER_REGION \ --network-cidr NETWORK_CIDR \ --project PROJECT_NAME ``` Replace the following: * `CLOUD_PROVIDER_REGION` with the cloud provider and region to host the VPC, for example `aws-eu-west-1` * `NETWORK_CIDR` with the CIDR block (a range of IP addresses) for the VPC, for example, `10.0.0.0/24` * `PROJECT_NAME` with the name of your Aiven project where to create the VPC Make an API call to the [VpcCreate](https://api.aiven.io/doc/#tag/Project/operation/VpcCreate) endpoint: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_ID/vpcs \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data ' { "cloud_name": "CLOUD_PROVIDER_REGION", "network_cidr": "NETWORK_CIDR" } ' ``` Replace `PROJECT_ID`, `BEARER_TOKEN`, `CLOUD_PROVIDER_REGION`, and `NETWORK_CIDR` with meaningful data. Use the [aiven\_project\_vpc](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/project_vpc) resource. ## Delete a project VPC[​](#delete-a-project-vpc "Direct link to Delete a project VPC") important Remove all services from your VCP before you delete it. To remove the services from the VCP, either migrate them out of the VCP or delete them. Deleting the VPC terminates its peering connections, if any. Delete a project VPC using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API 1. Log in to [Aiven Console](https://console.aiven.io/), and go to your project. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, find a VPC to be deleted and click **Actions** > **Delete**. 4. In the **Confirmation** window, click **Delete VPC**. Run the [avn vpc delete](/docs/tools/cli/vpc.md#delete-vpcs) command: ``` avn vpc delete \ --project-vpc-id PROJECT_VPC_ID ``` Replace `PROJECT_VPC_ID` with the ID of your Aiven project VPC, for example, `12345678-1a2b-3c4d-5f6g-1a2b3c4d5e6f`. Make an API call to the [VpcDelete](https://api.aiven.io/doc/#tag/Project/operation/VpcDelete) endpoint: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_ID/vpcs/PROJECT_VPC_ID \ --header 'Authorization: Bearer BEARER_TOKEN' \ ``` Replace the following placeholders with meaningful data: * `PROJECT_ID` (Aiven project name) * `PROJECT_VPC_ID` (Aiven project VPC ID) * `BEARER_TOKEN` --- # Migrate service to another cloud or region When creating an Aiven service, you are not tied to a cloud provider or region. You can migrate your services later to better match your needs. Services can be moved to another cloud provider, another region within the same provider, or both. When migrating a service, the migration happens in the background and does not affect your service until the service has been rebuilt at the new region. The migration includes the DNS update for the named service to point to the new region. Most services usually have usually no interruptions. However, for services like PostgreSQL®, MySQL®, and Caching, it may cause a short interruption of 5 to 10 seconds in while the DNS changes are propagated. The short interruption does not include potential delays caused by client side library implementation. * Console * Terraform 1. In your service, click **Service settings** from the sidebar. 2. In the **Cloud and network** section, click **Actions** > **Change cloud**. 3. In the **Cloud** section , select a cloud provider and region, and click **Change**. To change the cloud provider or region, update the `cloud_name` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Your service is in the **Rebuilding** state. Once the rebuilding is over, your new cloud provider and region will be in use. important The service's URI remains the same after the migration. note You can also use the [dedicated service update function](/docs/tools/cli/service-cli.md#avn-cli-service-update) to migrate a service via the [Aiven CLI](/docs/tools/cli.md). --- # Event logs Aiven consolidates all event logs for an organization into centralized event logs. This lets you view all events across your organization's units and projects in one place. Events include information on the action, who performed the action, the date and time, and the target resource. The target can be the organization, a unit, project, or service. You can filter by user, organizational unit, project, service, billing group, and time range. Logs are retained for 30 days. Required roles or permissions:role:organization:admin or organization:event\_logs:read To view your organization's event logs in the Aiven Console: 1. Click **Admin**. 2. Click **Event logs**. --- # Prepare services for high load Prepare an Aiven service for higher than usual traffic to avoid service outages. Prepare your services for higher than usual traffic to avoid service outages by doing the following: * **Subscribe to service notifications:** To receive notifications about service health and warnings when resources are low, you can [set service and project contacts](https://aiven.io/docs/platform/howto/technical-emails). You can also view the status of the Aiven Platform and get updates on incidents on the [status page](https://status.aiven.io/). Follow the RSS feed, subscribe to email or SMS notifications, or use the Slack integration to get notifications about incidents. * **Monitor your services:** [Monitor the health of your services](/docs/platform/howto/list-monitoring.md) using metrics, logs, alerts, and dashboards. * **Scale your services:** If you forecast a load that can't be handled by the service, you can scale up your service. * **Set the backup schedule:** To minimize the impact of the higher load during the backup process, schedule backups outside of peak traffic hours. * **Set the maintenance window:** Schedule maintenance updates outside of your peak traffic hours. * **Run load tests on service forks:** To test the impact of high traffic on a production service, fork the service and run your load test on the fork. Additionally, optimizing a service allows it to perform better under stress therefore avoiding the need of an upgrade. The more optimized a service is for your usage, the better you can weather spikes in traffic. Related pages * [Change a service plan](/docs/platform/howto/scale-services.md) * [Service memory limits](/docs/platform/concepts/service-memory-limits.md) * [Fork a service](/docs/platform/concepts/service-forking.md) --- # Handle resolution errors of private IP addresses When an Aiven service is placed in a VPC (Virtual Private Cloud), the DNS hostname of the service will resolve to a private IP address within the VPC's network address range. Any application connecting to the service needs to be on the same VPC. Some DNS resolvers used in office and home networks block the resolution of external hostnames to private IP addresses, known as [DNS-rebinding protection](https://en.wikipedia.org/wiki/DNS_rebinding#Protection). If the hostname of a service in a VPC cannot be resolved, this can be due to DNS-rebinding protection on your network. To verify this assumption: 1. Enable public access to your service. If the `public-` prefixed hostname of the service resolves successfully, then the problem is with the private IP. 2. Request the hostname using a known resolver such as Google Public DNS at `8.8.8.8`. This has no rebinding protection so serves as a good test. You can use the `dig` command: ``` dig +short myservice-myproject.aivencloud.com @8.8.8.8 ``` 3. Compare the output of the above command with the response from your default DNS resolver: ``` dig +short myservice-myproject.aivencloud.com ``` 4. If the response from your default DNS resolver does not return the same IP address as the earlier test, then your default DNS resolver is blocking the resolution. The recommended fix for this issue is to configure your DNS resolver (normally a server for offices, or a home router for home networks) to allow the resolution of hostnames in the Aiven service domain, `aivencloud.com`, to bypass the DNS-rebinding protection. note It is preferable to allow a single domain to bypass the DNS-rebinding if your DNS resolver allows it, instead of disabling DNS-rebinding protection entirely. --- # Enable public access in VPCs To enable public access for a service running within a virtual private cloud (VPC): 1. Log in to [Aiven Console](https://console.aiven.io) and click your service from the **Services** page. 2. On the **Overview** page of your service, click **Service settings** from the sidebar. 3. On the **Service settings** page, in the **Cloud and network** section, click **Actions** > **More network configurations**. 4. In the **Network configuration** window, click **Add configuration options**. In the search field, enter `public_access`. From the displayed parameter names, select a parameter name for your service type and enable it. 5. Click **Save configuration**. The **Overview** page now has an **Access Route** setting inside the **Connection information** section with **Public** and **Dynamic** options. 6. Click **Public** to see the public URL for your service. The connection with the **Dynamic** option is not possible outside the VPC, while the connection with the **Public** option is accessible over the public Internet. **IP Allow-List** applies to all connection types (Dynamic and Public, in this example). note You can change the `public_access` settings without any service downtime. --- # Reactivate suspended accounts If you have bills past due and didn't set up a payment method, your account may be suspended. When you try to log in with a suspended account, you see the error: **Project suspended, access prohibited**. To reactivate your account, email the Aiven billing team at with the following information: * Your full name * The name of your organization * The email address associated with the account note The billing team operates Monday through Friday, 9:00 to 17:00, Eastern European Time (EET/EEST). To avoid suspensions, [add a payment method](/docs/platform/howto/manage-payment-card.md) to your organization. --- # Track service restore progress using the API Track the restore progress of individual nodes in an Aiven service during node replacement, forking, or maintenance, using the Aiven API. You can track restore progress for individual nodes during service node replacement by using the Aiven API. For example, use this endpoint to monitor the restore progress of a forked service or when applying maintenance. The service object exposes restore progress under `node_states[].progress_updates`: * `service.node_states[]` contains per-node state entries. * When a node is restoring or catching up, its `state` is typically `syncing_data`. * When the state is `syncing_data`, the node may include `progress_updates` with one or more phase objects. * Other node states don't include restore progress data. note `progress_updates` may be missing or empty even when a node is in `syncing_data`. This can occur when a restore completes before detailed progress is reported or when the service does not emit detailed progress counters. ## API endpoints[​](#api-endpoints "Direct link to API endpoints") Restore progress fields are part of the standard service response payload. * Get a single service (recommended for polling): `GET /project/{project}/service/{service_name}` * List services in a project: `GET /project/{project}/service` - Request - Response ``` curl -H "Authorization: aivenv1 API_TOKEN" https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME ``` Replace the following placeholders: * `API_TOKEN`: Your Aiven API token. * `PROJECT_NAME`: Your Aiven project name. * `SERVICE_NAME`: The name of your service. ``` { "service": { ... "node_states": [ { "node_name": "...", "state": "syncing_data", "progress_updates": [ { "phase": "basebackup", "completed": false, "current": 3410567, "min": 0, "max": 7569280, "unit": "bytes_uncompressed" } ] } ], ... } } ``` ## Node states[​](#node-states "Direct link to Node states") Common values for `node_states[].state` include: * `setting_up_vm`: The virtual machine is being created or initialized. * `syncing_data`: The node is restoring data or catching up. * `running`: The node is operating normally. * `leaving`: The node is leaving the cluster. * `unknown`: A transient or error state. ## `progress_updates` data model[​](#progress_updates-data-model "Direct link to progress_updates-data-model") `progress_updates` is a list of phase objects. When present, phases appear in the following order: 1. `prepare` 2. `basebackup` 3. `stream` 4. `finalize` Each phase object includes the following fields: ``` { "completed": false, "current": 3410567, "max": 7569280, "min": 0, "phase": "basebackup", "unit": "bytes_uncompressed" } ``` ### Field semantics[​](#field-semantics "Direct link to Field semantics") * `phase`: String, required. The restore phase. Possible values: `prepare`, `basebackup`, `stream`, and `finalize`. * `completed`: Boolean, required. Whether the phase is complete. * `current`: Number or null, optional. The current progress value. This field can be missing or null. * `min`: Number or null, optional. The starting value for the phase. This field can be missing or null. * `max`: Number or null, optional. The expected total value for the phase. This value can be missing, null, or change while the restore is in progress. * `unit`: String or null, optional. The unit for `current`, `min`, and `max`. New unit values can appear over time. Important considerations * Treat `unit` as an opaque identifier. Unknown values can appear. * `max` may change while a restore is in progress. * Not all phases report numeric counters. Some services only indicate phase completion. ## Why `max` values can change[​](#why-max-values-can-change "Direct link to why-max-values-can-change") The `current`, `min`, and `max` values are best-effort progress indicators. They can be based on estimates or on system state that changes over time. Treat `max` as the latest known expected total, not as a fixed guarantee. Common reasons `max` can change include: * The restore process discovers additional work after it starts, such as files, segments, or objects that become visible only after metadata is read. * New data is added on the backend while the node is catching up, which moves the completion point forward. This is common during incremental catch-up phases. * Progress is calculated from system state, such as replication lag, rather than from a fixed work queue. As the system state changes, the value is recalculated. * The service switches restore strategies during the operation, for example from snapshot restore to replication catch-up, which changes what the counters represent. As a result: * Phase percentage can decrease even when the restore operates normally. * Remaining-time estimates based on `max` are unreliable. * Sudden changes in `max` are expected unless the node remains in `syncing_data` longer than expected. ## Restore phase meanings[​](#restore-phase-meanings "Direct link to Restore phase meanings") Phase names are standardized, but the underlying work and the meaning of the counters are service-specific. * `prepare`: Prepares the node for restore. * `basebackup`: Restores the full backup. * `stream`: Applies incremental changes, such as replication or log replay. * `finalize`: Completes final steps before serving traffic. Not all restores include every phase. ## Compute phase progress percentages[​](#compute-phase-progress-percentages "Direct link to Compute phase progress percentages") You cannot reliably compute overall restore progress. You can compute a phase-specific progress percentage when `min`, `max`, and `current` are present and `max != min`. ``` pct = round(((current - min) / (max - min)) * 100, 1) ``` When handling progress values: * If any of `min`, `max`, or `current` is null or missing, display `n/a`. * If `max == min`, treat the percentage as undefined. * Expect the percentage to decrease when `max` changes. * Clamp displayed values to the range `[0, 100]`. ## Polling guidance[​](#polling-guidance "Direct link to Polling guidance") Progress updates are best-effort and refresh every 10 seconds while a node is in `syncing_data`. Poll the service state every 10 to 30 seconds. More frequent polling does not provide additional detail. For each `node_states[]` entry: * If `state` is not `syncing_data`, no restore progress is available. * If `state` is `syncing_data`: * If `progress_updates` is missing or empty, the node is restoring without detailed progress data. * Otherwise, the current phase is the last phase where `completed` is `false`. Stop polling when all nodes reach the `running` state or when a stall is detected. ### Stall detection[​](#stall-detection "Direct link to Stall detection") The API does not provide per-phase timestamps. To detect stalls, use a time-based threshold, such as a node remaining in `syncing_data` longer than expected. Do not rely on counters or `max` values to estimate remaining time. Related pages * [Service backups](/docs/platform/concepts/service_backups.md) * [Backup to another region](/docs/platform/concepts/backup-to-another-region.md) --- # Restrict network access to services Restrict access to your Aiven-managed service to a single IP, an address block, or any combination of both. By default, a connection to an Aiven service can be established from any IP address. To restrict access, you can use the IP filtering capability, which allows you to filter traffic incoming to your services by specifying allowed IP addresses or network ranges. note If your service is within a VPC, the VPC configuration filters incoming traffic before the IP filter is applied. By default, the IP filter is set to `0.0.0.0/0`, which allows all inbound connections. If you remove `0.0.0.0/0` without adding networks or addresses used by clients, no client can connect to your service. tip To access a non-publicly-accessible service from another service, use a [service integration](/docs/platform/concepts/service-integration.md). ## Restrict access[​](#restrict-access "Direct link to Restrict access") 1. Log in to the [Aiven Console](https://console.aiven.io), and select the service to restrict access to. 2. On the **Overview** page of your service, select **Service settings**. 3. On the **Service settings** page, in the **Cloud and network** section: * Set the IP filter for the first time: 1. Click **Actions** > **Set IP address allowlist**. 2. In the **Allowed inbound IP addresses** window, remove `0.0.0.0/0` and enter an IP address or address block using the CIDR notation, for example `10.20.0.0/16`. * Edit the IP filter after the first setup change: 1. Click **Actions** > **Edit IP address allowlist**. 2. In the **Allowed inbound IP addresses** window, enter an IP address or address block using the CIDR notation, for example `10.20.0.0/16`. 4. To add more IP addresses or ranges, click **Add IP address range**. 5. Select **Save changes**. Now your service can be accessed from the specified IP addresses only. Alternative method You can also use the [dedicated service update function](/docs/tools/cli/service-cli.md#avn-cli-service-update) to create or update the IP filter for your service via the [Aiven CLI](/docs/tools/cli.md). Related pages For more ways of securing your service, see: * [Networking with VPC peering](/docs/platform/concepts/cloud-security.md#networking-with-vpc-peering) * [Configure VPC peering](/docs/platform/howto/manage-project-vpc.md#create-a-project-vpc). --- # Add Auth0 as an identity provider Use [Auth0](https://auth0.com/) to give your organization users single sign-on (SSO) access to Aiven. ## Step 1: Add the IdP in the Aiven Console[​](#add-idp-aiven-console "Direct link to Step 1: Add the IdP in the Aiven Console") 1. In the organization, click **Admin**. 2. Click **Identity providers** . 3. Click **Add identity provider**. 4. Select an identity provider and enter a name. 5. Select a [verified domain](/docs/platform/howto/manage-domains.md) to link this IdP to. Users see linked IdPs on the login page. On the **Configuration** step are two parameters that you use to set up the SAML authentication in your IdP: * Metadata URL * ACS URL ## Step 2: Configure SAML on Auth0[​](#step-2-configure-saml-on-auth0 "Direct link to Step 2: Configure SAML on Auth0") 1. Log in to [your Auth0 account](https://manage.auth0.com). 2. Select **Applications**. 3. Click **Create Application**. 4. Enter an application name. 5. Choose **Regular Web Applications** and click **Create**. 6. After your application is created, go to the **Addons** tab. 7. Enable the **SAML 2 WEB APP** option. 8. Click the **SAML 2 WEB APP** option. The **Settings** tab opens. 9. Set the **Application Callback URL** to the **ACS URL** from the Aiven Console. 10. In the **Settings** section for the Application Callback URL, remove the existing configuration and add the following field mapping configuration: ``` { "email": "email", "first_name": "first_name", "identity": "email", "last_name": "last_name", "mapUnknownClaimsAsIs": true } ``` 11. Click **Enable** and **Save**. 12. On the **Usage** tab, make a note of the **Identity Provider Login URL**, **Issuer URN**, and **Identity Provider Certificate**. These are needed for the SAML configuration in Aiven Console. ## Step 3: Finish the configuration in Aiven[​](#step-3-finish-the-configuration-in-aiven "Direct link to Step 3: Finish the configuration in Aiven") Go back to the Aiven Console to complete setting up the IdP. If you saved your IdP as a draft, you can open the settings by clicking the name of the IdP. 1. In the **IDP URL** field, enter the Auth0 **Identity Provider Login URL**. 2. In the **Entity ID** field, enter the Auth0 **Issuer URN**. 3) Paste the certificate from the IdP into the **Certificate** field. 4) Click **Next**. 5) Configure the security options for this IdP and click **Next**. * **Require authentication context**: This lets the IdP enforce stricter security measures to help prevent unauthorized access, such as requiring multi-factor authentication. * **Require assertion to be signed**: The IdP checks for a digital signature. This security measure ensures the integrity and authenticity of the assertions by verifying that they were issued by a trusted party and have not been tampered with. * **Sign authorization request sent to IdP**: A digital signature is added to the request to verify its authenticity and integrity. * **Extend active sessions**: This resets the session duration every time the token is used. * **Enable group syncing**: This syncs the group membership from your IdP to the Aiven Platform. note * The Aiven Platform doesn't create groups with this feature. It only syncs the group membership between the groups in your IdP and the [groups you create in Aiven](/docs/platform/howto/manage-groups.md). For user group provisioning, use SCIM. * User group membership automatically syncs when a user logs in. * The IdP is the single source of truth. If a group in the IdP doesn't exist in Aiven, it will be ignored. Likewise, if a user is added to a group in Aiven Console but not in the IdP, they will be removed from the Aiven group when the group membership syncs. 6) Optional: Select a user group to add all users who sign up with this IdP to. 7) Click **Finish** to complete the setup. note If you set up a SAML authentication method before and are now switching to a new IdP, existing users need to log in with the new account link URL to finish the setup. Related pages * [Troubleshooting for SAML IdPs](/docs/platform/howto/saml/add-identity-providers.md#troubleshooting) --- # Add Microsoft Azure Active Directory as an identity provider Use [Microsoft Azure Active Directory (AD)](https://azure.microsoft.com/en-us/products/active-directory/) to give your organization users single sign-on (SSO) access to Aiven. ## Step 1: Add the IdP in the Aiven Console[​](#add-idp-aiven-console "Direct link to Step 1: Add the IdP in the Aiven Console") 1. In the organization, click **Admin**. 2. Click **Identity providers** . 3. Click **Add identity provider**. 4. Select an identity provider and enter a name. 5. Select a [verified domain](/docs/platform/howto/manage-domains.md) to link this IdP to. Users see linked IdPs on the login page. On the **Configuration** step are two parameters that you use to set up the SAML authentication in your IdP: * Metadata URL * ACS URL ## Step 2: Configure SAML on Microsoft Azure[​](#step-2-configure-saml-on-microsoft-azure "Direct link to Step 2: Configure SAML on Microsoft Azure") ### Set up an Azure application[​](#set-up-an-azure-application "Direct link to Set up an Azure application") 1. Log in to [Microsoft Azure](https://portal.azure.com/). 2. Go to **Enterprise applications**. 3. Click **All applications**. 4. Click **New application**. 5. Click the **Add from the gallery** search bar and use the **Azure AD SAML Toolkit**. 6. Click **Add**. 7. Go back to the **Enterprise applications** list. note The newly created application might not be visible. You can use the **All applications** filter to see the new application. 8. Click the name of the new application. 9. Click **Single sign-on**. 10. Select **SAML** as the single sign-on method. 11. Add the following parameters to the **Basic SAML Configuration**: | Parameter | Value | | ------------------------------------------ | ---------------------------------------------------------------------------------------------------------- | | Identifier (Entity ID) | `https://api.aiven.io/v1/sso/saml/account/{account_id}/method/{account_authentication_method_id}/metadata` | | Reply URL (Assertion Consumer Service URL) | `https://api.aiven.io/v1/sso/saml/account/{account_id}/method/{account_authentication_method_id}/acs` | | Sign on URL | `https://console.aiven.io` | 12. Click **Save**. ### Create a claim and add users[​](#create-a-claim-and-add-users "Direct link to Create a claim and add users") 1. In the **User Attributes & Claims**, click **Add a new claim**. 2. Create an attribute with the following: | Parameter | Value | | ---------------- | --------- | | Name | email | | Source | Attribute | | Source Attribute | user.mail | 3. Download the **Certificate (Base64)** from the **SAML Signing Certificate** section. 4. Go to **Users and groups** and click **Add user**. 5. Select the users that will use Azure AD to log in to Aiven. 6. Click **Assign**. ## Step 3: Finish the configuration in Aiven[​](#step-3-finish-the-configuration-in-aiven "Direct link to Step 3: Finish the configuration in Aiven") Go back to the Aiven Console to complete setting up the IdP. If you saved your IdP as a draft, you can open the settings by clicking the name of the IdP. 1. In the **IDP URL** field, enter the **Login URL** from Azure. 2. In the **Entity ID** field, enter the **Microsoft Entra Identifier** from Azure. 3) Paste the certificate from the IdP into the **Certificate** field. 4) Click **Next**. 5) Configure the security options for this IdP and click **Next**. * **Require authentication context**: This lets the IdP enforce stricter security measures to help prevent unauthorized access, such as requiring multi-factor authentication. * **Require assertion to be signed**: The IdP checks for a digital signature. This security measure ensures the integrity and authenticity of the assertions by verifying that they were issued by a trusted party and have not been tampered with. * **Sign authorization request sent to IdP**: A digital signature is added to the request to verify its authenticity and integrity. * **Extend active sessions**: This resets the session duration every time the token is used. * **Enable group syncing**: This syncs the group membership from your IdP to the Aiven Platform. note * The Aiven Platform doesn't create groups with this feature. It only syncs the group membership between the groups in your IdP and the [groups you create in Aiven](/docs/platform/howto/manage-groups.md). For user group provisioning, use SCIM. * User group membership automatically syncs when a user logs in. * The IdP is the single source of truth. If a group in the IdP doesn't exist in Aiven, it will be ignored. Likewise, if a user is added to a group in Aiven Console but not in the IdP, they will be removed from the Aiven group when the group membership syncs. 6) Optional: Select a user group to add all users who sign up with this IdP to. 7) Click **Finish** to complete the setup. note If you set up a SAML authentication method before and are now switching to a new IdP, existing users need to log in with the new account link URL to finish the setup. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") If you get an error message to contact your administrator: 1. Go to the Microsoft Azure AD user profile for the users. 2. In **Contact Info**, check whether the **Email** field is blank. If it is blank, there are two possible solutions: * In **User Principal Name**, if the **Identity** field is an email address, try changing the **User Attributes & Claims** to `email = user.userprincipalname`. * In **Contact Info**, if none of the **Alternate email** fields are blank, try changing the **User Attributes & Claims** to `email = user.othermail`. Related pages * [Troubleshooting for SAML IdPs](/docs/platform/howto/saml/add-identity-providers.md#troubleshooting) --- # Add FusionAuth as an identity provider Use [FusionAuth](https://fusionauth.io/) to give your organization users single sign-on (SSO) access to Aiven. ## Step 1: Add the IdP in the Aiven Console[​](#add-idp-aiven-console "Direct link to Step 1: Add the IdP in the Aiven Console") 1. In the organization, click **Admin**. 2. Click **Identity providers** . 3. Click **Add identity provider**. 4. Select an identity provider and enter a name. 5. Select a [verified domain](/docs/platform/howto/manage-domains.md) to link this IdP to. Users see linked IdPs on the login page. On the **Configuration** step are two parameters that you use to set up the SAML authentication in your IdP: * Metadata URL * ACS URL ## Step 2: Configure SAML on FusionAuth[​](#step-2-configure-saml-on-fusionauth "Direct link to Step 2: Configure SAML on FusionAuth") The setup on FusionAuth has three parts: * Create an API key * Generate a custom RSA certificate * Create an application ### Create an API key[​](#create-an-api-key "Direct link to Create an API key") 1. In FusionAuth, go to **Settings** > **API Keys**. 2. Click the **Add** icon. 3. Enter a description for the key. 4. In the **Endpoints** list, find **/api/key/import**. 5. Toggle on **POST**. 6. Click the **Save** icon. ![Creating an API key.](/docs/assets/images/create-api-key-51301f2c0bd59d15b6528e5a3bd2714d.png) 7. On the **API Keys** page, find your key and click the value in the **Key** column. 8. Copy the whole key. You'll use this for the script. ![Copying the API key value.](/docs/assets/images/grab-api-key-a3b1f04bc5800af9a1bedc19ab4ea166.png) 9. To clone the [FusionAuth example scripts GitHub repository](https://github.com/FusionAuth/fusionauth-example-scripts), run: ``` git clone git@github.com:FusionAuth/fusionauth-example-scripts.git cd fusionauth-example-scripts/v3-certificate ``` 10. Run the `generate-certificate` script. ``` ./generate-certificate ``` 11. Name the key. 12. Copy the generated certificate created by the script. You now have a certificate in the **Key Master** in your FusionAuth instance. ### Create an application[​](#create-an-application "Direct link to Create an application") 1. In **Applications**, click the **Add** icon. 2. Enter a name for the application. 3. On the **SAML** tab, toggle on **Enabled**. 4. In the **Issuer** field, enter the **Metadata URL** from the Aiven Console. 5. In the **Authorized redirect URLs** field, enter the **ACS URL** from the Aiven Console. 6. In the **Authentication response** section, change the **Signing key** to the API key you created. 7. Click the **Save** icon. 8. On the **Applications** page, click the magnifying glass. 9. In the **SAML v2 Integration details** section, copy the **Entity Id** and **Login URL**. ## Step 3: Finish the configuration in Aiven[​](#step-3-finish-the-configuration-in-aiven "Direct link to Step 3: Finish the configuration in Aiven") Go back to the Aiven Console to complete setting up the IdP. If you saved your IdP as a draft, you can open the settings by clicking the name of the IdP. 1. In the **IDP URL** field, enter the **Login URL** from FusionAuth. 2. In the **Entity ID** field, enter the **Entity ID** from FusionAuth. 3) Paste the certificate from the IdP into the **Certificate** field. 4) Click **Next**. 5) Configure the security options for this IdP and click **Next**. * **Require authentication context**: This lets the IdP enforce stricter security measures to help prevent unauthorized access, such as requiring multi-factor authentication. * **Require assertion to be signed**: The IdP checks for a digital signature. This security measure ensures the integrity and authenticity of the assertions by verifying that they were issued by a trusted party and have not been tampered with. * **Sign authorization request sent to IdP**: A digital signature is added to the request to verify its authenticity and integrity. * **Extend active sessions**: This resets the session duration every time the token is used. * **Enable group syncing**: This syncs the group membership from your IdP to the Aiven Platform. note * The Aiven Platform doesn't create groups with this feature. It only syncs the group membership between the groups in your IdP and the [groups you create in Aiven](/docs/platform/howto/manage-groups.md). For user group provisioning, use SCIM. * User group membership automatically syncs when a user logs in. * The IdP is the single source of truth. If a group in the IdP doesn't exist in Aiven, it will be ignored. Likewise, if a user is added to a group in Aiven Console but not in the IdP, they will be removed from the Aiven group when the group membership syncs. 6) Optional: Select a user group to add all users who sign up with this IdP to. 7) Click **Finish** to complete the setup. note If you set up a SAML authentication method before and are now switching to a new IdP, existing users need to log in with the new account link URL to finish the setup. Related pages * [Troubleshooting for SAML IdPs](/docs/platform/howto/saml/add-identity-providers.md#troubleshooting) --- # Add Google as an identity provider Use Google to give your organization users single sign-on (SSO) access to Aiven. ## Step 1: Add the IdP in the Aiven Console[​](#add-idp-aiven-console "Direct link to Step 1: Add the IdP in the Aiven Console") 1. In the organization, click **Admin**. 2. Click **Identity providers** . 3. Click **Add identity provider**. 4. Select an identity provider and enter a name. 5. Select a [verified domain](/docs/platform/howto/manage-domains.md) to link this IdP to. Users see linked IdPs on the login page. On the **Configuration** step are two parameters that you use to set up the SAML authentication in your IdP: * Metadata URL * ACS URL ## Step 2: Configure SAML on Google[​](#step-2-configure-saml-on-google "Direct link to Step 2: Configure SAML on Google") 1. Log in to Google Admin console. 2. Go to **Menu** > **Apps** > **Web and mobile apps**. 3. Click **Add App** > **Add custom SAML app**. 4. On the **App Details** page, enter a name for the Aiven profile. 5. Click **Continue**. 6. On the **Google Identity Provider** details page, copy the **SSO URL**, **Entity ID**, and the **Certificate**. You'll use these for the SAML configuration in Aiven Console. 7. Click **Continue**. 8. On the **Service Provider Details** page, set the following parameters: | Parameter | Value | | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Entity ID | **Metadata URL** from Aiven Console | | ACS URL | **ACS URL** from Aiven Console | | Start URL | - for Aiven Console
- for Aiven GCP Marketplace Console
- for Aiven AWS Marketplace Console | | Name ID format | EMAIL | | App attributes | email | 9. Click **Finish**. 10. Turn on your SAML app. ## Step 3: Finish the configuration in Aiven[​](#step-3-finish-the-configuration-in-aiven "Direct link to Step 3: Finish the configuration in Aiven") Go back to the Aiven Console to complete setting up the IdP. If you saved your IdP as a draft, you can open the settings by clicking the name of the IdP. 1. In the **IDP URL** field, enter the **SSO URL** from Google. 2. In the **Entity ID** field, enter the **Entity ID** from Google. 3) Paste the certificate from the IdP into the **Certificate** field. 4) Click **Next**. 5) Configure the security options for this IdP and click **Next**. * **Require authentication context**: This lets the IdP enforce stricter security measures to help prevent unauthorized access, such as requiring multi-factor authentication. * **Require assertion to be signed**: The IdP checks for a digital signature. This security measure ensures the integrity and authenticity of the assertions by verifying that they were issued by a trusted party and have not been tampered with. * **Sign authorization request sent to IdP**: A digital signature is added to the request to verify its authenticity and integrity. * **Extend active sessions**: This resets the session duration every time the token is used. * **Enable group syncing**: This syncs the group membership from your IdP to the Aiven Platform. note * The Aiven Platform doesn't create groups with this feature. It only syncs the group membership between the groups in your IdP and the [groups you create in Aiven](/docs/platform/howto/manage-groups.md). For user group provisioning, use SCIM. * User group membership automatically syncs when a user logs in. * The IdP is the single source of truth. If a group in the IdP doesn't exist in Aiven, it will be ignored. Likewise, if a user is added to a group in Aiven Console but not in the IdP, they will be removed from the Aiven group when the group membership syncs. 6) Optional: Select a user group to add all users who sign up with this IdP to. 7) Click **Finish** to complete the setup. note If you set up a SAML authentication method before and are now switching to a new IdP, existing users need to log in with the new account link URL to finish the setup. Related pages * [Troubleshooting for SAML IdPs](/docs/platform/howto/saml/add-identity-providers.md#troubleshooting) --- # Add SAML identity providers You can give your organization users access to Aiven through identity providers (IdPs) that support SAML. You must be an [organization admin](/docs/platform/concepts/permissions.md) to manage IdPs. The following are general steps for setting up single sign-on with an IdP. Setup instructions are also available for these specific providers: * [Auth0](/docs/platform/howto/saml/add-auth0-idp.md) * [FusionAuth](/docs/platform/howto/saml/add-fusionauth-idp.md) * [Google](/docs/platform/howto/saml/add-google-idp.md) * [JumpCloud](/docs/platform/howto/saml/add-jumpcloud-idp.md) * [Microsoft Azure Active Directory](/docs/platform/howto/saml/add-azure-idp.md) * [Okta](/docs/platform/howto/saml/add-okta-idp.md) * [OneLogin](/docs/platform/howto/saml/add-onelogin-idp.md) For all IdPs, Aiven recommends following [security best practices](/docs/platform/howto/list-identity-providers.md). ## Step 1: Add the IdP in the Aiven Console[​](#add-idp-aiven-console "Direct link to Step 1: Add the IdP in the Aiven Console") 1. In the organization, click **Admin**. 2. Click **Identity providers** . 3. Click **Add identity provider**. 4. Select an identity provider and enter a name. 5. Select a [verified domain](/docs/platform/howto/manage-domains.md) to link this IdP to. Users see linked IdPs on the login page. On the **Configuration** step are two parameters that you use to set up the SAML authentication in your IdP: * Metadata URL * ACS URL ## Step 2: Configure SAML on your IdP[​](#step-2-configure-saml-on-your-idp "Direct link to Step 2: Configure SAML on your IdP") Use the metadata URL and ACS URL from the Aiven Console to configure a new application in your IdP. ## Step 3: Finish the configuration in Aiven[​](#step-3-finish-the-configuration-in-aiven "Direct link to Step 3: Finish the configuration in Aiven") Go back to the Aiven Console to complete setting up the IdP. If you saved your IdP as a draft, you can open the settings by clicking the name of the IdP. 1. Enter the **IDP URL** from your identity provider. 2. Enter the **Entity ID** from your identity provider. 3) Paste the certificate from the IdP into the **Certificate** field. 4) Click **Next**. 5) Configure the security options for this IdP and click **Next**. * **Require authentication context**: This lets the IdP enforce stricter security measures to help prevent unauthorized access, such as requiring multi-factor authentication. * **Require assertion to be signed**: The IdP checks for a digital signature. This security measure ensures the integrity and authenticity of the assertions by verifying that they were issued by a trusted party and have not been tampered with. * **Sign authorization request sent to IdP**: A digital signature is added to the request to verify its authenticity and integrity. * **Extend active sessions**: This resets the session duration every time the token is used. * **Enable group syncing**: This syncs the group membership from your IdP to the Aiven Platform. note * The Aiven Platform doesn't create groups with this feature. It only syncs the group membership between the groups in your IdP and the [groups you create in Aiven](/docs/platform/howto/manage-groups.md). For user group provisioning, use SCIM. * User group membership automatically syncs when a user logs in. * The IdP is the single source of truth. If a group in the IdP doesn't exist in Aiven, it will be ignored. Likewise, if a user is added to a group in Aiven Console but not in the IdP, they will be removed from the Aiven group when the group membership syncs. 6) Optional: Select a user group to add all users who sign up with this IdP to. 7) Click **Finish** to complete the setup. note If you set up a SAML authentication method before and are now switching to a new IdP, existing users need to log in with the new account link URL to finish the setup. ## Step 4: Optional: Link your users to the identity provider[​](#step-4-optional-link-your-users-to-the-identity-provider "Direct link to Step 4: Optional: Link your users to the identity provider") You can manually link Aiven organization user accounts using the following URLs. note You don't need to manually link organization users who have an email address that matches a verified domain linked to one of your identity providers. 1. On the **Identity providers** page, click the name of the IdP. 2. In the **Overview** section there are two URLs: * **Signup URL**: Users that don't have an Aiven user account can use this to create an Aiven user linked to this IdP. * **User account link URL**: Users that already have an Aiven user account can link their existing Aiven user with this IdP. 3. Send the appropriate URL to your organization users. If you set up a different IdP before and are now switching to a new IdP, existing users need to log in with the new account link URL to finish the setup. 4. The process for users to link their Aiven user account with the IdP varies for new and existing users. * For existing users that are already logged in to the Aiven Console: 1. Click the account link URL. 2. Click **Link profile** to go to your IdP's authentication page. 3. Log in to the IdP to link the accounts. You can use the IdP for all future logins. * For existing users that are not logged into the Aiven Console: 1. Click the account link URL. 2. Click **Login**. 3. Log in to the Aiven Console. You are redirected to your IdP's authentication page. 4. Log in to the IdP to link the accounts. You can use the IdP for all future logins. * For new users without an Aiven user account: 1. Click the signup URL. 2. Select the identity provider on the signup page. You are redirected to your IdP's authentication page. 3. Log in to the IdP. 4. Complete the sign up process in the Aiven Console. The IdP is automatically linked to your Aiven user account and you can use it for all future logins. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") If you have issues, you can use the [SAML Tracer browser extension](https://addons.mozilla.org/firefox/addon/saml-tracer/) to check the process step by step. ### Authentication failed[​](#authentication-failed "Direct link to Authentication failed") If you get an authentication failed error when launching the Aiven SAML application, ask your organization admin to confirm that **IdP-initiated login** is enabled. ### Invalid relay state[​](#invalid-relay-state "Direct link to Invalid relay state") An invalid relay state error usually means you attempted to log in from the identity provider (IdP), for example, from the IdP dashboard. To avoid this, you can set the default landing page, often called the default relay state or start URL, to the Aiven Console that your organization uses. A more secure and often simpler approach is to log in directly from the Aiven Console's login page. ### The IdP password does not work[​](#the-idp-password-does-not-work "Direct link to The IdP password does not work") Make sure to use the **Account Link URL** to add the IdP to your Aiven user account. You can view all authentication methods for your user account in **User information** > **Authentication**. --- # Add JumpCloud as an identity provider Use [JumpCloud](https://jumpcloud.com/) to give your organization users single sign-on (SSO) access to Aiven. ## Step 1: Add the IdP in the Aiven Console[​](#add-idp-aiven-console "Direct link to Step 1: Add the IdP in the Aiven Console") 1. In the organization, click **Admin**. 2. Click **Identity providers** . 3. Click **Add identity provider**. 4. Select an identity provider and enter a name. 5. Select a [verified domain](/docs/platform/howto/manage-domains.md) to link this IdP to. Users see linked IdPs on the login page. On the **Configuration** step are two parameters that you use to set up the SAML authentication in your IdP: * Metadata URL * ACS URL ## Step 2: Configure SAML on JumpCloud[​](#step-2-configure-saml-on-jumpcloud "Direct link to Step 2: Configure SAML on JumpCloud") 1. In the [JumpCloud admin console](https://console.jumpcloud.com/login), go to **SSO**. 2. Click **Custom SAML App**. 3. Set the **IdP Entity ID**. 4. Set the **Audience URI (SP Entity ID)** to the **Metadata URL** from the Aiven Console. 5. Set the **ACS URL** to the one from the Aiven Console. 6. Set the **Default RelayState** to ****. 7. Add an entry in **Attribute statements** with a **Service Provider Attribute Name** of **email** and **JumpCloud Attribute Name** of **email**. 8. Set the **Login URL** to the **ACS URL** from the Aiven Console. 9. In **User Groups**, assign the application to your user groups. 10. Click **Activate**. 11. Download the certificate. ## Step 3: Finish the configuration in Aiven[​](#step-3-finish-the-configuration-in-aiven "Direct link to Step 3: Finish the configuration in Aiven") Go back to the Aiven Console to complete setting up the IdP. If you saved your IdP as a draft, you can open the settings by clicking the name of the IdP. 1. In the **IDP URL** field, enter **IDP URL** from JumpCloud. 2. In the **Entity ID** field, enter the **IdP Entity ID** from JumpCloud. 3) Paste the certificate from the IdP into the **Certificate** field. 4) Click **Next**. 5) Configure the security options for this IdP and click **Next**. * **Require authentication context**: This lets the IdP enforce stricter security measures to help prevent unauthorized access, such as requiring multi-factor authentication. * **Require assertion to be signed**: The IdP checks for a digital signature. This security measure ensures the integrity and authenticity of the assertions by verifying that they were issued by a trusted party and have not been tampered with. * **Sign authorization request sent to IdP**: A digital signature is added to the request to verify its authenticity and integrity. * **Extend active sessions**: This resets the session duration every time the token is used. * **Enable group syncing**: This syncs the group membership from your IdP to the Aiven Platform. note * The Aiven Platform doesn't create groups with this feature. It only syncs the group membership between the groups in your IdP and the [groups you create in Aiven](/docs/platform/howto/manage-groups.md). For user group provisioning, use SCIM. * User group membership automatically syncs when a user logs in. * The IdP is the single source of truth. If a group in the IdP doesn't exist in Aiven, it will be ignored. Likewise, if a user is added to a group in Aiven Console but not in the IdP, they will be removed from the Aiven group when the group membership syncs. 6) Optional: Select a user group to add all users who sign up with this IdP to. 7) Click **Finish** to complete the setup. note If you set up a SAML authentication method before and are now switching to a new IdP, existing users need to log in with the new account link URL to finish the setup. Related pages * [Troubleshooting for SAML IdPs](/docs/platform/howto/saml/add-identity-providers.md#troubleshooting) --- # Add Okta as an identity provider Use [Okta](https://www.okta.com/) to give your organization users single sign-on (SSO) access to Aiven using SAML. Aiven also supports [user provisioning for Okta](#step-4-optional-configure-user-provisioning) with SCIM. ## Supported features[​](#supported-features "Direct link to Supported features") * Identity provider (IdP) initiated SSO * Service provider (SP) initiated SSO For more information on the listed features, visit the [Okta Glossary](https://help.okta.com/okta_help.htm?type=oie\&id=ext_glossary). ## Step 1: Add the IdP in the Aiven Console[​](#add-idp-aiven-console "Direct link to Step 1: Add the IdP in the Aiven Console") 1. In the organization, click **Admin**. 2. Click **Identity providers** . 3. Click **Add identity provider**. 4. Select an identity provider and enter a name. 5. Select a [verified domain](/docs/platform/howto/manage-domains.md) to link this IdP to. Users see linked IdPs on the login page. On the **Configuration** step are two parameters that you use to set up the SAML authentication in your IdP: * Metadata URL * ACS URL ## Step 2: Configure SAML on Okta[​](#step-2-configure-saml-on-okta "Direct link to Step 2: Configure SAML on Okta") 1. In the [Okta administrator console](https://login.okta.com/), go to **Applications** > **Applications**. 2. Click **Browse App Catalog**. 3. Search for and open the Aiven app. 4. Click **Add Integration** and **Done**. 5. On the **Sign On** tab, click **Edit**. 6. In the **Advanced Sign-on Settings** set the **Metadata URL** and **ACS URL** to the URLs copied from the Aiven Console. 7. Set the **Default Relay State** for the console you use: * For the Aiven Console: * For the Aiven GCP Marketplace Console: * For the Aiven AWS Marketplace Console: 8. Click **Save**. 9. In the **SAML 2.0** section, click **More details**. 10. Copy the **Sign on URL**, **Issuer**, and the **Signing Certificate**. You'll use these to configure the IdP in Aiven. ## Step 3: Finish the configuration in Aiven[​](#step-3-finish-the-configuration-in-aiven "Direct link to Step 3: Finish the configuration in Aiven") Go back to the Aiven Console to complete setting up the IdP. If you saved your IdP as a draft, you can open the settings by clicking the name of the IdP. 1. In the **IDP URL** field, enter the **Sign on URL** from Okta. 2. In the **Entity ID** field, enter the **Issuer** from Okta. 3) Paste the certificate from the IdP into the **Certificate** field. 4) Click **Next**. 5) Configure the security options for this IdP and click **Next**. * **Require authentication context**: This lets the IdP enforce stricter security measures to help prevent unauthorized access, such as requiring multi-factor authentication. * **Require assertion to be signed**: The IdP checks for a digital signature. This security measure ensures the integrity and authenticity of the assertions by verifying that they were issued by a trusted party and have not been tampered with. * **Sign authorization request sent to IdP**: A digital signature is added to the request to verify its authenticity and integrity. * **Extend active sessions**: This resets the session duration every time the token is used. * **Enable group syncing**: This syncs the group membership from your IdP to the Aiven Platform. note * The Aiven Platform doesn't create groups with this feature. It only syncs the group membership between the groups in your IdP and the [groups you create in Aiven](/docs/platform/howto/manage-groups.md). For user group provisioning, use SCIM. * User group membership automatically syncs when a user logs in. * The IdP is the single source of truth. If a group in the IdP doesn't exist in Aiven, it will be ignored. Likewise, if a user is added to a group in Aiven Console but not in the IdP, they will be removed from the Aiven group when the group membership syncs. 6) Optional: Select a user group to add all users who sign up with this IdP to. 7) Click **Finish** to complete the setup. note If you set up a SAML authentication method before and are now switching to a new IdP, existing users need to log in with the new account link URL to finish the setup. ## Step 4: Optional: Configure user provisioning[​](#step-4-optional-configure-user-provisioning "Direct link to Step 4: Optional: Configure user provisioning") You can automate user provisioning with Okta through System for Cross-domain Identity Management (SCIM). This means you can manage your users and their profiles in one place, Okta, and push those changes to the Aiven platform. Aiven's integration with Okta supports these features: * **Push new users**: Users created in Okta are automatically created as managed users in Aiven. * **Push profile updates**: User profile updates in Okta are pushed to Aiven. Profiles for these users cannot be changed in Aiven. * **Push user deactivation**: Users that are deactivated or removed in Okta are deactivated in Aiven. You can manually delete users in Aiven after they are deactivated. * **Push groups**: Groups created or updated in Okta are created and updated in Aiven. * **Sync passwords**: Automatically synchronizes users' Aiven passwords with their Okta passwords. To configure user provisioning for Okta: 1. In Okta, click **Applications** and go to the Aiven application. 2. Click **Provisioning**. 3. Click **Settings** > **Integration** > **Configure API Integration**. 4. Select **Enable API Integration**. 5. In the **Base URL** field, paste the **Base URL** from the Aiven Console. 6. In the **API Token** field, paste the **Access token** from the Aiven Console. 7. Click **Test API Credentials** to confirm the connection is working and save the configuration. important Don't enable **Import Groups**. Aiven groups that aren't managed by SCIM cannot be imported to Okta. 8. Click **Save**. 9. Optional: On the **Provisioning** tab, click **Edit** to enable provisioning settings. Recommended settings Set the following for centralized and secure user management: * Enable **Create users** * Disable **Set password when creating new users** * Enable **Update user attributes** * Enable **Deactivate users** * Disable **Sync password** 10. Click **Save**. 11. Click **Sign On**. 12. In the **Credentials Details** section, for the **Application username format** select **Email**. 13. Click **Save**. Related pages * [Troubleshooting for SAML IdPs](/docs/platform/howto/saml/add-identity-providers.md#troubleshooting) --- # Add OneLogin as an identity provider Use [OneLogin](https://www.onelogin.com/) to give your organization users single sign-on (SSO) access to Aiven. ## Step 1: Add the IdP in the Aiven Console[​](#add-idp-aiven-console "Direct link to Step 1: Add the IdP in the Aiven Console") 1. In the organization, click **Admin**. 2. Click **Identity providers** . 3. Click **Add identity provider**. 4. Select an identity provider and enter a name. 5. Select a [verified domain](/docs/platform/howto/manage-domains.md) to link this IdP to. Users see linked IdPs on the login page. On the **Configuration** step are two parameters that you use to set up the SAML authentication in your IdP: * Metadata URL * ACS URL ## Step 2: Configure SAML on OneLogin[​](#step-2-configure-saml-on-onelogin "Direct link to Step 2: Configure SAML on OneLogin") 1. Log in to the [OneLogin Admin console](https://app.onelogin.com/login). 2. Click **Applications** > **Add App**. 3. Search for and select **SAML Custom Connector (Advanced)**. 4. Change the **Display Name** to **Aiven**. 5. Add any other visual configurations you want and click **Save**. 6. In the **Configuration** section of the menu, set the following parameters: | Parameter | Value | | ------------------ | ---------------------------------------------------------------------------------------- | | ACS URL Validation | `[-a-zA-Z0-9@:%._\+~#=]{2,256}\.[a-z]{2,6}\b([-a-zA-Z0-9@:%_\+.~#?&//=]*)` | | ACS URL | **ACS URL** from Aiven Console | | Login URL | | | SAML Initiator | - Service Provider
- To let your users sign in through OneLogin, enter **OneLogin** | | SAML nameID format | Email | 7. Click **Save**. 8. In the **SSO** section, set **SAML Signature Algorithm** to **SHA-256**. 9. Copy the certificate content, **Issuer URL**, and **SAML 2.0 Endpoint (HTTP)**. You'll use these for the SAML configuration in Aiven Console. 10. Click **Save** 11. Assign users to this application. ## Step 3: Finish the configuration in Aiven[​](#step-3-finish-the-configuration-in-aiven "Direct link to Step 3: Finish the configuration in Aiven") Go back to the Aiven Console to complete setting up the IdP. If you saved your IdP as a draft, you can open the settings by clicking the name of the IdP. 1. In the **IDP URL** field, enter the **SAML 2.0 Endpoint (HTTP)** from OneLogin. 2. In the **Entity ID** field, enter the **Issuer URL** from OneLogin. 3) Paste the certificate from the IdP into the **Certificate** field. 4) Click **Next**. 5) Configure the security options for this IdP and click **Next**. * **Require authentication context**: This lets the IdP enforce stricter security measures to help prevent unauthorized access, such as requiring multi-factor authentication. * **Require assertion to be signed**: The IdP checks for a digital signature. This security measure ensures the integrity and authenticity of the assertions by verifying that they were issued by a trusted party and have not been tampered with. * **Sign authorization request sent to IdP**: A digital signature is added to the request to verify its authenticity and integrity. * **Extend active sessions**: This resets the session duration every time the token is used. * **Enable group syncing**: This syncs the group membership from your IdP to the Aiven Platform. note * The Aiven Platform doesn't create groups with this feature. It only syncs the group membership between the groups in your IdP and the [groups you create in Aiven](/docs/platform/howto/manage-groups.md). For user group provisioning, use SCIM. * User group membership automatically syncs when a user logs in. * The IdP is the single source of truth. If a group in the IdP doesn't exist in Aiven, it will be ignored. Likewise, if a user is added to a group in Aiven Console but not in the IdP, they will be removed from the Aiven group when the group membership syncs. 6) Optional: Select a user group to add all users who sign up with this IdP to. 7) Click **Finish** to complete the setup. note If you set up a SAML authentication method before and are now switching to a new IdP, existing users need to log in with the new account link URL to finish the setup. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") If you get errors, you can try reapplying entitlement mappings: 1. Go to the app in OneLogin and click **Settings**. 2. Click **More Actions** > **Reapply entitlement Mappings**. Related pages * [Troubleshooting for SAML IdPs](/docs/platform/howto/saml/add-identity-providers.md#troubleshooting) --- # Rotate SCIM tokens You can manually rotate the SCIM token for an identity provider to maintain the security of your user provisioning setup. To avoid interruptions to your user provisioning, update the token in your identity provider configuration when you rotate a SCIM token. To generate a SCIM token: 1. In your organization, click **Admin**. 2. Click **Identity providers**. 3. Click the name of the identity provider. 4. In the **SCIM provisioning** section, click **Edit**. 5. To generate a new token, click **Generate new token**. 6. To confirm, click **Generate new token**. 7. Click **Copy** and update it in your identity provider. * To update Okta configuration, follow the steps in the [user provisioning configuration](/docs/platform/howto/saml/add-okta-idp.md#step-4-optional-configure-user-provisioning) guide. 8. Click **Save changes**. Related pages * [Add Okta as an identity provider](/docs/platform/howto/saml/add-okta-idp.md) * [SAML identity providers and verified domains](/docs/platform/howto/list-identity-providers.md) --- # Change a service plan Change the plan of an Aiven service to scale it up or down and optimize costs. Adjust the plan of your services at any time to scale your services as needed and optimize costs. If you can't find a suitable plan, you can [request a custom plan](/docs/platform/concepts/service-pricing.md). tip If you plan to upgrade your service plan, do it immediately after a full backup. This reduces the amount of incremental changes that need to be applied on top of the base backup, which speeds up the upgrade itself. important * When changing a service plan, reserve an additional 25% of disk space. This requirement applies to upgrades and downgrades. * Downgrading to a plan with fewer VMs is supported for most services, including Aiven for Apache Kafka®, Aiven for PostgreSQL®, Aiven for OpenSearch®, Aiven for ClickHouse®, Aiven for MySQL®, Aiven for Metrics, and Aiven for Valkey™. * Changing a service plan triggers a node recycle, service rebuilding, and any pending maintenance updates. - Console - Terraform - CLI 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Change plan**. 3. Select a plan that provides at least 125% of the current disk size and click **Change plan**. Update the `plan` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). To change a service plan in the Aiven CLI, use the [`avn service update --plan `](/docs/tools/cli/service-cli.md#avn-cli-service-update) command. Your service's state becomes **Rebuilding** and remains accessible. When the state switches to **Running**, your new service plan is active. Related pages * [Service memory limits](/docs/platform/concepts/service-memory-limits.md) * [Prepare services for high load](/docs/platform/howto/prepare-for-high-load.md) --- # Search for services On the **Services** page in [Aiven Console](https://console.aiven.io/), you can search for services by keywords and narrow down the results using filters. ## Search by keyword[​](#search-by-keyword "Direct link to Search by keyword") When you search by keyword, [Aiven Console](https://console.aiven.io/) shows all services that have the matching words in the service name, plan, cloud provider, and tags. ## Filter search results with the UI[​](#filter-search-results-with-the-ui "Direct link to Filter search results with the UI") To filter search results: 1. Click **Filter list**. 2. Select the filters of your choice. ## Filter with a query[​](#filter-with-a-query "Direct link to Filter with a query") You can type filter queries. The supported filters are the following: * `service` * `status` * `provider` * `region` Use several filters by separating them by a comma. You can use these filters alongside keyword searches. Examples All running PostgreSQL® services that are hosted on AWS or Google Cloud ``` service:pg status:running provider:aws,google ``` All powered off Kafka® services with 'production' in the name ``` production service:kafka status:poweroff ``` ### Filter by service type[​](#filter-by-service-type "Direct link to Filter by service type") To filter the services by service type, use these filter values: | Service name | Filter value | | --------------------------- | -------------- | | Apache Flink® | `flink` | | Apache Kafka® | `kafka` | | Apache Kafka® Connect | `kafkaconnect` | | Apache Kafka® MirrorMaker 2 | `mirrormaker` | | ClickHouse® | `clickhouse` | | Grafana® | `grafana` | | Metrics | `thanos` | | MySQL® | `mysql` | | OpenSearch® | `opensearch` | | PostgreSQL® | `pg` | ### Filter by status[​](#filter-by-status "Direct link to Filter by status") You can filter the services to show only those that are running, powered off, rebuilding, or rebalancing. Supported `status` values: * `running` * `poweroff` * `rebuilding` * `rebalancing` ### Filter by cloud provider[​](#filter-by-cloud-provider "Direct link to Filter by cloud provider") To filter the services by the cloud provider they are hosted on, use these filter values: | Cloud provider | Filter value | | --------------------------- | ------------ | | Amazon Web Services (AWS) | `aws` | | Azure | `azure` | | Digital Ocean | `do` | | Google Cloud Provider (GCP) | `google` | | UpCloud | `upcloud` | ### Filter by cloud region[​](#filter-by-cloud-region "Direct link to Filter by cloud region") Find the supported values for the `region` filter in the **Cloud** column of the tables in [List of available cloud regions](/docs/platform/reference/list_of_clouds.md). Example All services in the AWS 'eu-central-1' region ``` region:aws-eu-central-1 ``` --- # Set authentication policies for organization users The authentication policy for your organization specifies the ways that users in your organization can access the organization on the Aiven Platform. ## Authentication types[​](#authentication-types "Direct link to Authentication types") When creating an authentication policy, you select the authentication methods to allow for all users in your organization. For increased security, it's a good idea to always [verify your organization's domains](/docs/platform/howto/manage-domains.md). ### Passwords and two-factor authentication[​](#passwords-and-two-factor-authentication "Direct link to Passwords and two-factor authentication") With password authentication enabled, users log in with their email address and password. For an added layer of security, you can enforce two-factor authentication (2FA) for password logins for all users in your organization. When 2FA is required, users can't access any resources in your organization until they set up 2FA. This only applies to logins using email and password. The Aiven Platform cannot enforce 2FA for logins through third-party providers, including identity providers. note Personal tokens are not affected and continue to work when you make 2FA required. However, when users [enable 2FA](/docs/platform/howto/user-2fa.md) their existing tokens might stop working. ### Third-party authentication[​](#third-party-authentication "Direct link to Third-party authentication") Users can choose to log in using Google, Microsoft, or GitHub. ### SSO with an organization identity provider[​](#sso-with-an-organization-identity-provider "Direct link to SSO with an organization identity provider") Users that are part of multiple Aiven organizations can log in using single sign-on (SSO) and access your organization's resources with an [identity provider](/docs/platform/howto/saml/add-identity-providers.md) that is configured for any of those organizations. You can further restrict access by requiring users to log in with one of your organization's identity providers. This means that they cannot log in to your organization using another Aiven organization's identity provider. It's strongly recommended to enable this if you only have one Aiven organization. ### Personal tokens[​](#personal-tokens "Direct link to Personal tokens") Users can generate their own [personal tokens](/docs/platform/howto/create_authentication_token.md) for use with the Aiven API. When you turn off personal tokens, managed users can't create personal tokens. Non-managed users can still create personal tokens, but they can't use them to access the organization's resources. To regularly manage your resources programmatically with the Aiven API, CLI, Terraform Provider, or other tools, it's best to create an [application user](/docs/platform/howto/manage-application-users.md) with its own tokens. Personal tokens are generated with the authentication method that the user logged in with. Tokens are linked to the authentication method they are created with. You can ensure that access to your organization using tokens conforms to the authentication policy by requiring users to be logged in with an allowed authentication method when they use a token. If your authentication policy changes, tokens that don't conform to the new policy stop working. For example, if you have an authentication policy that allows users to log in with a password, a user can log in with their email and password, and create a personal token. This token is tied to the password authentication method they logged in with. If the authentication policy changes later to only allow logging on with an identity provider, then the token generated when the user was logged in with their password will not work. After logging in with an allowed method on the new authentication policy the user can create a token. ### Access from allowed IP addresses[​](#access-from-allowed-ip-addresses "Direct link to Access from allowed IP addresses") You can restrict access to your organization's resources on the Aiven Platform to specific IP address ranges, ensuring connections are coming from trusted networks. This helps you minimize exposure, reduce the risk of breaches, and comply with policies and regulations. This authentication policy setting also applies to access through personal and application tokens. ### MCP connections[​](#mcp-connections "Direct link to MCP connections") Users can connect [Aiven MCP](/docs/tools/mcp-server.md) clients, such as Cursor and Claude Code, to the services and other resources they have access to in your organization. To allow users to connect MCP clients, select **Allow MCP connections**. If you turn off this setting, users cannot connect MCP clients to resources in your organization. When MCP connections are allowed, you can also select **Restrict MCP connections to read-only operations**. MCP clients can then view services and other resources, but they cannot create, modify, or delete them. The restriction applies to all MCP connections in the organization. Users cannot override it in their client configuration. For more information, see [Read-only mode](/docs/tools/mcp-server.md#read-only-mode). ## Set an authentication policy[​](#set-an-authentication-policy "Direct link to Set an authentication policy") 1. In the organization, click **Admin**. 2. Click **Authentication**. 3. Configure the settings for your authentication policy. 4. Click **Save changes**. --- # Support All customers using paid services have access to the Basic support tier. Aiven also offers [paid support tiers](https://aiven.io/support-services) with faster response times, phone support, and other services. Custom [service level agreements](https://aiven.io/sla) are available for the Premium support tier. Customers who use only free services do not have access to support services. ## Change your support tier[​](#change-your-support-tier "Direct link to Change your support tier") To change your organization's support tier, you must be an [organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) and have at least the Basic support tier. 1. In the organization, click **Admin**. 2. In the **Support tier** section, click **View details or change tier**. 3. Click **Change tier** and choose a tier. 4. Select a **Billing group**. important The support costs for all current and future services in the selected organization and its organizational units will be added to the invoice for this billing group. 5. Click **Change tier**. It typically takes 1-2 business days to set up the new support tier. ## Cancel a paid support contract[​](#cancel-a-paid-support-contract "Direct link to Cancel a paid support contract") To cancel your paid support contract, you can [change your support tier](#change-your-support-tier) to Basic. Your paid support will end after the current month. ## Create a support ticket[​](#create-a-support-ticket "Direct link to Create a support ticket") Create a ticket for issues or problems with the platform. For other services included in your support tier like business reviews or disaster recovery planning, contact your account team. 1. In the [Aiven Console](https://console.aiven.io/), click **Support** to open the Aiven Support Center. 2. Click **Create ticket**. 3. Enter email addresses to CC in the support ticket. They receive all new comments and updates. 4. Enter a **Subject**. 5. Select a **Severity** level. 6. Optional: Enter the ID of the affected projects and services. 7. Select the affected **Product** and the reason for creating the ticket. 1) Enter a detailed **Description** of the issue. note Include the following information in the description to help the support team provide timely assistance: * The affected features. For example, networking, metrics, or deployment. * The steps to reproduce the problem. * Any error messages. * Any languages or frameworks you are using. 1. Optional: Upload files such as screenshots, logs, or [HAR files](#create-har-files). important Aiven support will never ask you to provide sensitive data such as passwords or personal information. Remove or replace sensitive data in files that you attach to support tickets. 2. Click **Create ticket**. You can track the status of your tickets on the **My tickets** page. [Response times](https://aiven.io/support-services) vary by case severity and support tier. If you are not satisfied with the processing of your ticket, add `#escalate` in the comments. ## Add participants to a support ticket[​](#add-participants-to-a-support-ticket "Direct link to Add participants to a support ticket") To give every organization user access to all support tickets in your organization contact your account team. To add Aiven users to a support ticket: 1. In the [Aiven Console](https://console.aiven.io/), click **Support** to open the Aiven Support Center. 2. On the **My tickets** page, open the ticket. 3. Click **Add to conversation**. 4. Add the email addresses in the **CC** field separated by a space. These must be the same email addresses they use to log in. 5. Enter a comment and click **Submit**. ## Get notifications for all support tickets[​](#get-notifications-for-all-support-tickets "Direct link to Get notifications for all support tickets") [Organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) can get notifications for updates on all tickets in their organization. 1. In the [Aiven Console](https://console.aiven.io/), click **Support** to open the Aiven Support Center. 2. Click **My tickets**. 3. On the **Tickets in my organization** tab, click **Follow all tickets**. You get email notifications for all updates on both existing and new tickets. You can unfollow them at any time. ## Create HAR files[​](#create-har-files "Direct link to Create HAR files") The support team occasionally needs information about the network requests that are generated in your browser. Browsers can capture a log of these network requests in a HAR (HTTP Archive) file. 1. Use your browser to create the HAR file while you go through the steps to reproduce the problem: * Follow the [instructions for Internet Explorer/Edge, Firefox, and Chrome](https://toolbox.googleapps.com/apps/har_analyzer/). * For Safari, make sure you can access the [developer tools](https://support.apple.com/en-ie/guide/safari/sfri20948/mac) and [export the HAR file](https://webkit.org/web-inspector/network-tab/). 2. Replace sensitive data in the file with placeholders while retaining the JSON structure and format. Examples of sensitive data include: * Personal identifiers such as email addresses and phone numbers * Tokens or passwords * Sensitive URLs * Sensitive cookies or headers 3. Zip the sanitized file and add password protection to it. 4. Follow the support team's instructions to share the file and password with them. --- # Use resource tags Add key-value tags to an Aiven service to organize services and track ownership, cost allocation, and governance. Use tags to add metadata to Aiven services to categorize them or run custom logic on them. Typical uses include: * Tagging for governance to deploy services with specific tags only. * Tagging for internal cost reporting, ownership, allocation, and accountability. A tag is a key/value pair: * **Key**: A case-sensitive string that starts with a letter and consists of letters, numbers, dashes, and underscores. The maximum length for a key is 64 characters. * **Value**: A string value limited to 64 UTF-8 characters. Within a service, the tag keys must be unique. * Console * Terraform 1. In the service, click **Service settings**. 2. In the **Service status** section, click **Actions** > **Add service tags**. 3. Enter a key and value for each tag. 4. Click **Save changes**. Use the `tag` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Create a service](/docs/platform/howto/create_new_service.md) * [Services](/docs/products/services.md) --- # Manage project and service notifications To stay up to date with the latest information about services and projects, you can set service and project contacts to receive email notifications. Notifications include information about plan sizes, performance, outages, and scheduled maintenance. The default contacts for a project are the project admin and operators. You can also set the contacts to specific email addresses. Project contacts receive notifications about the project. They also receive the notifications for all services in the project, unless you set a separate service contact for a service. Service contacts by default are the project contacts. However, if you set other email addresses as service contacts for a service, notifications are sent only to the contacts for that specific service. Project and service contacts cannot unsubscribe from specific project or service notifications. ## Set project contacts[​](#set-project-contacts "Direct link to Set project contacts") * Console * Terraform 1. In the project, click **Settings**. 2. On the **Notifications** tab, select the project contacts that you want to receive email notifications. 3. Click **Save changes**. Use the `technical_emails` attribute in [your `aiven_project` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/project#technical_emails-1). ## Set service contacts[​](#set-service-contacts "Direct link to Set service contacts") * Console * Terraform 1. In the service, click **Service settings**. 2. In the **Service status** section, open the menu in the top right and select **Change service contacts**. 3. Select the contacts that should receive email notifications for this service. 4. Click **Save**. Use the `tech_emails` attribute in [your service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). ## Set up Slack notifications[​](#set-up-slack-notifications "Direct link to Set up Slack notifications") To get notifications in Slack, you can add a Slack channel's or DM email address to the technical contacts for an Aiven project: 1. In Slack, [create an email address for a channel or DM](https://slack.com/help/articles/206819278-Send-emails-to-Slack#h_01F4WDZG8RTCTNAMR4KJ7D419V). note If you don't see the email integrations option, ask the owner or admin of the workspace or organization to [allow incoming emails](https://slack.com/help/articles/360053335433-Manage-incoming-emails-for-your-workspace-or-organization). 2. In the [Aiven Console](https://console.aiven.io/), go to the project or service. 3. Set the Slack email address as a [project contact](/docs/platform/howto/technical-emails.md#set-project-contacts) or [service contact](/docs/platform/howto/technical-emails.md#set-service-contacts). Alternatively, you can [set up a Slackbot forwarding address](https://slack.com/help/articles/206819278-Send-emails-to-Slack#h_01F4WE06MBF06BBHQNZ1G0H2K5) and use to automatically forward Aiven's email notifications from your email client. --- # Change unsafe passwords The Aiven Platform checks your email and password combination against a database of exposed credentials every time you log in and change your password. If Aiven detects an unsafe password, your login is blocked until you [reset your password](/docs/platform/reference/change-password.md#reset-your-password) to keep your account safe. You don't need to do anything else, but Aiven recommends every user [enable two-factor authentication](/docs/platform/howto/user-2fa.md). Two-factor authentication also needs to be re-enabled after you reset or change your password. --- # Use AWS PrivateLink with Aiven services AWS [PrivateLink](https://aws.amazon.com/privatelink/) brings Aiven services to the selected virtual private cloud (VPC) in your AWS account. In a traditional setup that uses [VPC peering](/docs/platform/howto/manage-project-vpc.md#create-a-project-vpc), traffic is routed through an AWS VPC peering connection to your Aiven services. With PrivateLink, you can create a VPC endpoint in your own VPC and access an Aiven service from that. The VPC endpoint creates network interfaces (NIC) to the subnets and availability zones that you choose and receives the private IP addresses that belong to the IP range of your VPC. The VPC endpoint is routed to your Aiven service located in one of Aiven's AWS accounts. You can enable PrivateLink for Aiven services located in project VPC. Before you can set up AWS PrivateLink, [create a VPC](/docs/platform/howto/manage-project-vpc.md#create-a-project-vpc) and launch the services to connect to that VPC. As there is no network routing between the VPC, you can use any private IP range for the VPC, unless you also want to connect to the project VPC using VPC peering connections. This means that overlaps in the IP range are not an issue. To set up AWS PrivateLink, use the [Aiven CLI](/docs/tools/cli.md). You also need [AWS Management Console](https://aws.amazon.com/console) or [CLI](https://aws.amazon.com/cli) to create a VPC endpoint. note AWS PrivateLink is not supported for: * Aiven for Apache Flink® * Aiven for Apache Kafka® MirrorMaker 2 * Aiven for Metrics ## Enable AWS PrivateLink[​](#enable-aws-privatelink "Direct link to Enable AWS PrivateLink") 1. Create an AWS PrivateLink resource on the Aiven service. The Amazon Resource Name (ARN) for the principals that are allowed to connect to the VPC endpoint service and the AWS network load balancer requires your Amazon account ID. In addition, you can set the access scope for an entire AWS account (`root`), a specific AWS user (for example, `user\john`), or a specific role. Only give permissions to roles that you trust, as an allowed role can connect from any VPC. Use the Aiven CLI to run the following command including your AWS account ID, the access scope, and the name of your Aiven service: ``` avn service privatelink aws create --principal arn:aws:iam::$AWS_account_ID:$access_scope $Aiven_service_name ``` For example: ``` avn service privatelink aws create --principal arn:aws:iam::012345678901:user/john my-kafka ``` This creates an AWS network load balancer dedicated to your Aiven service and attaches it to an AWS VPC endpoint service that you can later use to connect to your account's VPC endpoint. The PrivateLink resource stays in the initial `creating` state for up to a few minutes while the load balancer is being launched. After the load balancer and VPC endpoint service have been created, the state changes to `active` and the `aws_service_id` and `aws_service_name` values are set. 2. In the AWS CLI, run the following command to create a VPC endpoint: ``` aws ec2 --region eu-west-1 create-vpc-endpoint --vpc-endpoint-type Interface --vpc-id $your_aws_vpc_id --subnet-ids $space_separated_list_of_subnet_ids --security-group-ids $security_group_ids --service-name com.amazonaws.vpce.eu-west-1.vpce-svc-0b16e88f3b706aaf1 ``` Replace the following placeholders: * `--service-name` with the value shown either in the [Aiven Console](https://console.aiven.io) > **Service settings** page > **Cloud and network** section > **Actions** > **Edit AWS PrivateLink** > **AWS service name** or as an output of: ``` avn service privatelink aws get aiven_service_name ``` * `--security-group-ids` with the IDs of the security groups to associate with the endpoint network interfaces. If this parameter is not specified, the default security group for the VPC is used. For fault tolerance, specify a subnet ID for each availability zone in the region. The security groups determine the instances that are allowed to connect to the endpoint network interfaces created by AWS into the specified subnets. Alternatively, create the VPC endpoint in [AWS Console](https://console.aws.amazon.com) under **VPC** > **Endpoints** > **Create endpoint**. See the [AWS documentation](https://docs.aws.amazon.com/vpc/latest/privatelink/create-interface-endpoint.html) for details. note For Aiven for Apache Kafka® services, the security group for the VPC endpoint must allow ingress in the port range `10000-31000` to accommodate the pool of Kafka broker ports used in our PrivateLink implementation. These are custom TCP ports not included by default rule type `All traffic`. It takes a while before the endpoint is ready to use as AWS provisions network interfaces to each of the subnets and connects them to the Aiven VPC endpoint service. Once the AWS endpoint state changes to `available`, the connection is visible in Aiven. 3. If [your Aiven service is deployed using BYOC](/docs/platform/howto/byoc/aws-privatelink-byoc.md), run the [avn service privatelink aws refresh](/docs/tools/cli/service/privatelink.md#avn_service_privatelink_aws_refresh) command. Otherwise, skip this step. ``` avn service privatelink aws refresh --project $project_name $byoc_service_name ``` tip Check the deployment model of your service in the [Aiven Console](https://console.aiven.io/): Go to your service's **Overview** page > **Network** > **Deployment model**. 4. Enable PrivateLink access for Aiven service components: You can control each service component separately - for example, you can enable PrivateLink access for Kafka while allowing Kafka Connect to connect via VPC peering connections only. * In the Aiven CLI, set `user_config.privatelink_access.` to `true` for the components to enable, for example: ``` # For ClickHouse avn service update -c privatelink_access.clickhouse=true --project $project_name $Aiven_service_name ``` ``` # For PostgreSQL avn service update -c privatelink_access.pg=true --project $project_name $Aiven_service_name ``` ``` # For Kafka avn service update -c privatelink_access.kafka=true $Aiven_service_name avn service update -c privatelink_access.kafka_connect=true $Aiven_service_name avn service update -c privatelink_access.kafka_rest=true $Aiven_service_name avn service update -c privatelink_access.schema_registry=true $Aiven_service_name ``` * In [Aiven Console](https://console.aiven.io): 1. On the **Overview** page of your service, click **Service settings** from the sidebar. 2. On the **Service settings** page, go to the **Cloud and network** section and click **Actions** > **More network configurations** from the menu. 3. In the **Network configuration** window, click **Add configuration options**. In the search field, enter `privatelink_access`. From the displayed component names, select the names of the components to switch on. ![Aiven Console private link configuration](/docs/assets/images/use-aws-privatelink_image1-4492ac9d7d1c6ccea2271d170a7b9280.png) 4. Click the toggle switches for the selected components to switch them on. Click **Save configuration**. As a result, PrivateLink connection details are added to the **Connection information** section on the service **Overview**. ![Screenshot of the configuration](/docs/assets/images/use-aws-privatelink_image2-f3f20eb18ec98fd4ab6494b9cec36da8.png) It takes a couple of minutes before connectivity is available after you enable a service component. This is because AWS requires an AWS load balancer behind each VPC endpoint service, and the target rules on the load balancer for the service nodes need at least two successful heartbeats before they transition from the `initial` state to `healthy` and are included in the active forwarding rules of the load balancer. ## Acquire connection information[​](#h_b6605132ff "Direct link to Acquire connection information") ### One AWS PrivateLink connection[​](#one-aws-privatelink-connection "Direct link to One AWS PrivateLink connection") If you have one private endpoint connected to your Aiven service, you can preview the connection information (URI, hostname, or port required to access the service through the private endpoint) in [Aiven Console](https://console.aiven.io) > the service's **Overview** page > the **Connection information** section, where you'll also find the switch for the `privatelink` access route. `privatelink`-access-route values for `host` and `port` differ from those for the `dynamic` access route used by default to connect to the service. note You can use the same credentials with any access route. ### Multiple AWS PrivateLink connections[​](#multiple-aws-privatelink-connections "Direct link to Multiple AWS PrivateLink connections") Use CLI to acquire connection information for more than one AWS PrivateLink connection. Each endpoint (connection) has a `PRIVATELINK_CONNECTION_ID`, which you can check using the [`avn service privatelink aws connection list`](/docs/tools/cli/service/privatelink.md#avn_service_privatelink_aws_connection_list) command. To acquire connection information for your service component using AWS PrivateLink, run the [avn service connection-info](/docs/tools/cli/service/connection-info.md) command. * For SSL connection information for your service component using AWS PrivateLink, run the following command: ``` avn service connection-info UTILITY_NAME SERVICE_NAME --privatelink-connection-id PRIVATELINK_CONNECTION_ID ``` Where: * UTILITY\_NAME for Aiven for Apache Kafka®, for example, can be `kcat`. * SERVICE\_NAME for Aiven for Apache Kafka®, for example, can be `kafka-12a3b4c5`. * PRIVATELINK\_CONNECTION\_ID can be `plc39413abcdef`. * For SASL connection information for Aiven for Apache Kafka® service components using AWS PrivateLink, run the following command: ``` avn service connection-info UTILITY_NAME SERVICE_NAME --privatelink-connection-id PRIVATELINK_CONNECTION_ID -a sasl ``` Where: * UTILITY\_NAME for Aiven for Apache Kafka®, for example, can be `kcat`. * SERVICE\_NAME for Aiven for Apache Kafka®, for example, can be `kafka-12a3b4c5`. * PRIVATELINK\_CONNECTION\_ID can be `plc39413abcdef`. note SSL certificates and SASL credentials are the same for all the connections. You can use the same credentials with any access route. ## Update the allowed principals list[​](#h_2a1689a687 "Direct link to Update the allowed principals list") To change the list of AWS accounts or IAM users or roles that are allowed to connect a VPC endpoint: * Use the `update` command of the Aiven CLI: ``` avn service privatelink aws update --principal arn:aws:iam::$AWS_account_ID:$access_scope $Aiven_service_name ``` note When you add an entry, also include the `--principal` arguments for existing entries. * In [Aiven Console](https://console.aiven.io): 1. Click your service from the **Services** page. 2. On the **Overview** page, click **Service settings** from the sidebar. 3. On the **Service settings** page, go to the **Cloud and network** section and click **Actions** > **Edit AWS PrivateLink**. 4. In the **Edit AWS PrivateLink** window, enter the principals to include in the **Principal ARNs** field and click **Save**. ## Allow cross-region connections[​](#allow-cross-region-connections "Direct link to Allow cross-region connections") AWS PrivateLink supports connections between different AWS regions. Use this to let a VPC endpoint in another AWS region connect to your Aiven service's PrivateLink endpoint service. important Cross-region connections for AWS PrivateLink are a [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) feature. Contact the [sales team](https://aiven.io/contact) to enable this for your project. ### Limitations[​](#limitations "Direct link to Limitations") * Cross-region connections work only between regions in the same AWS [partition](https://docs.aws.amazon.com/whitepapers/latest/aws-fault-isolation-boundaries/partitions.html). For example, a standard AWS region can't connect to an AWS China region because they're in different partitions. * If your service is deployed with [bring your own cloud (BYOC)](/docs/platform/concepts/byoc.md), [set up the required permissions](/docs/platform/howto/byoc/aws-privatelink-byoc.md#set-up-permissions) before you enable cross-region connections. * You can allow up to 64 additional regions for one PrivateLink connection. * Creating endpoints in additional regions can add to your AWS costs. Check [AWS PrivateLink pricing](https://aws.amazon.com/privatelink/pricing/) before you enable additional regions. ### Set the allowed regions[​](#set-the-allowed-regions "Direct link to Set the allowed regions") Use the `supported_regions` parameter to set the additional AWS regions where a VPC endpoint can connect to your PrivateLink endpoint service. Your service's own AWS region is always included and can't be removed from this list. This parameter isn't available in the Aiven Console. * CLI * Terraform * API Use the `--supported-regions` option with a comma-separated list of AWS regions: ``` avn service privatelink aws create \ --principal arn:aws:iam::012345678901:root \ --supported-regions eu-west-2,us-east-1 \ SERVICE_NAME ``` To update the allowed regions of an existing PrivateLink resource, include the `--principal` arguments for your existing entries too: ``` avn service privatelink aws update \ --principal arn:aws:iam::012345678901:root \ --supported-regions eu-west-2,us-east-1 \ SERVICE_NAME ``` Add `supported_regions` to your `aiven_aws_privatelink` resource: ``` Loading... ``` To set the allowed regions when you create a PrivateLink resource, call the [ServicePrivatelinkAWSCreate](https://api.aiven.io/doc/#tag/Service/operation/ServicePrivatelinkAWSCreate) endpoint: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT/service/SERVICE/privatelink/aws \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "principals": ["arn:aws:iam::012345678901:root"], "supported_regions": ["eu-west-2", "us-east-1"] }' ``` To update the allowed regions of an existing PrivateLink resource, call the [ServicePrivatelinkAWSUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServicePrivatelinkAWSUpdate) endpoint: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT/service/SERVICE/privatelink/aws \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "supported_regions": ["eu-west-2", "us-east-1"] }' ``` Replace the following: * `PROJECT`: your project name. * `SERVICE`: your service name. * `BEARER_TOKEN`: your [Aiven authentication token](/docs/platform/concepts/authentication-tokens.md). ## Deleting a privatelink connection[​](#h_8de68d5894 "Direct link to Deleting a privatelink connection") * Using the Aiven CLI, run the following command: ``` avn service privatelink aws delete $Aiven_service_name ``` ``` AWS_SERVICE_ID AWS_SERVICE_NAME PRINCIPALS STATE ========================== ======================================================= ================================== ======== vpce-svc-0b16e88f3b706aaf1 com.amazonaws.vpce.eu-west-1.vpce-svc-0b16e88f3b ``` * Using [Aiven Console](https://console.aiven.io): 1. Click your service from the **Services** page. 2. On the **Overview** page, click **Service settings** from the sidebar. 3. On the **Service settings** page, go to the **Cloud and network** section and click **Actions** > **Delete AWS PrivateLink** . 4. In the **Confirmation** window, click **Delete**. This deletes the AWS load balancer and VPC service endpoint. Related pages * [Use AWS PrivateLink with BYOC services](/docs/platform/howto/byoc/aws-privatelink-byoc.md) * [Use Azure Private Link with Aiven services](/docs/platform/howto/use-azure-privatelink.md) * [Use Google Private Service Connect with Aiven services](/docs/platform/howto/use-google-private-service-connect.md) * [Manage project VPCs](/docs/platform/howto/manage-project-vpc.md) --- # Use Azure Private Link with Aiven services [Early availability](/docs/platform/concepts/service-and-feature-releases.md) Azure Private Link lets you bring your Aiven services into your virtual network (VNet) over a private endpoint. The endpoint creates a network interface into one of the VNet subnets, and receives a private IP address from its IP range. The private endpoint is routed to your Aiven service. Azure Private Link is supported for the following services: * Aiven for Apache Kafka® * Aiven for Apache Kafka Connect® * Aiven for ClickHouse® * Aiven for Grafana® * Aiven for InfluxDB® * Aiven for Metrics * Aiven for MySQL® * Aiven for OpenSearch® * Aiven for PostgreSQL® * Aiven for Caching note Azure Private Link is not supported for [BYOC](/docs/platform/concepts/byoc.md)-hosted services. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * This feature is in [early availability](/docs/platform/concepts/service-and-feature-releases.md#early-availability-). * [Aiven CLI](/docs/tools/cli.md) is installed. * The Aiven service is in [a project VPC](/docs/platform/howto/manage-project-vpc.md). This ensures the service is not accessible from the public internet. note If you are not using regular VNet peerings, any private IP range can be used for the VPC. There is no network routing between your Azure subscription and the Aiven VPC, so overlapping IP ranges are not an issue. * The Aiven service is using [static IP addresses](/docs/platform/concepts/static-ips.md). note Even though services in a VPC only communicate using private IP addresses, Azure load balancers require [standard SKU IP addresses](https://learn.microsoft.com/en-us/azure/virtual-network/ip-services/public-ip-upgrade-portal) for target virtual machines. Azure sends TCP health probes to load balancer target ports from a public IP address. ## Variables[​](#variables "Direct link to Variables") | Variable | Description | | ----------------- | -------------------------- | | `SUBSCRIPTION_ID` | Azure subscription ID | | `AIVEN_SERVICE` | Name of your Aiven service | ## Set up a Private Link connection[​](#set-up-a-private-link-connection "Direct link to Set up a Private Link connection") There are three steps to setting up an Azure Private Link with your Aiven service: 1. Create a Private Link service 2. Create a private endpoint 3. Enable Private Link access service components ### Step 1: Create a Private Link service[​](#step-1-create-a-private-link-service "Direct link to Step 1: Create a Private Link service") 1. In the Aiven CLI, create a Private Link resource on your Aiven service: ``` avn service privatelink azure create AIVEN_SERVICE --user-subscription-id SUBSCRIPTION_ID ``` This creates an [Azure Standard Internal Load Balancer](https://learn.microsoft.com/en-us/azure/load-balancer/load-balancer-overview) dedicated to your Aiven service and attaches it to an Azure Private Link service. Connections from other subscriptions are automatically rejected. 2. Check the status of the Private Link service: ``` avn service privatelink azure get AIVEN_SERVICE ``` The service is in the `creating` state until Azure provisions a load balancer and Private Link service. 3. When the state changes to `active`, note the `azure_service_alias` and `azure_service_id`: ``` avn service privatelink azure get AIVEN_SERVICE ``` ### Step 2: Create a private endpoint[​](#step-2-create-a-private-endpoint "Direct link to Step 2: Create a private endpoint") Azure resources in the Aiven service are now ready to be connected to your Azure subscription and virtual network. 1. In the Azure web console or Azure CLI, [create a private endpoint](https://learn.microsoft.com/en-us/azure/private-link/create-private-endpoint-portal?tabs=dynamic-ip). If you are using the console, select **Connect to an Azure resource by resource ID or alias** and enter the `azure_service_alias` or `azure_service_id`. 2. Refresh the Aiven Private Link service: ``` avn service privatelink azure refresh AIVEN_SERVICE ``` note Azure does not provide notifications about endpoint connections and the Aiven API will not be aware of new endpoints until it's refreshed. 3. In the Aiven CLI, check that the endpoint is connected to the service: ``` avn service privatelink azure connection list AIVEN_SERVICE ``` The output will look similar to this: ``` PRIVATELINK_CONNECTION_ID PRIVATE_ENDPOINT_ID STATE USER_IP_ADDRESS ========================= ========================================================================================================================================================== ===================== =============== plc35843e8051. /subscriptions/8eefec94-5d63-40c9-983c-03ab083b411d/resourceGroups/test-privatelink/providers/Microsoft.Network/privateEndpoints/my-endpoint pending-user-approval null ``` 4. Check that the endpoint ID matches the one created in your subscription and approve it: ``` avn service privatelink azure connection approve AIVEN_SERVICE PRIVATELINK_CONNECTION_ID ``` The endpoint in your Azure subscription is now connected to the Private Link service in the Aiven service. The state of the endpoint is `pending`. 5. In the Azure web console, go to the private endpoint and select **Network interface**. Copy the private IP address. 6. In the Aiven CLI, add the endpoint's IP address you copied to the connection: ``` avn service privatelink azure connection update \ --endpoint-ip-address IP_ADDRESS \ AIVEN_SERVICE PRIVATELINK_CONNECTION_ID ``` Once the endpoint IP address is added, the connection's status changes to `active`. A DNS name for the service is registered pointing to that IP address. ### Step 3: Enable Private Link access for Aiven service components[​](#step-3-enable-private-link-access-for-aiven-service-components "Direct link to Step 3: Enable Private Link access for Aiven service components") Enable Private Link access on your Aiven services using either the Aiven CLI or [Aiven Console](https://console.aiven.io/). * Aiven CLI * Console To enable Private Link access for your service in the Aiven CLI, set `user_config.privatelink_access.` to true for the components to enable. For example, for PostgreSQL the command is: ``` avn service update -c privatelink_access.pg=true AIVEN_SERVICE ``` To enable Private Link access in [Aiven Console](https://console.aiven.io/): 1. On the **Overview** page of your service, select **Service settings** from the sidebar. 2. On the **Service settings** page, in the **Cloud and network** section, click **Actions** > **More network configurations**. 3. In the **Network configuration** window, take the following actions: 1. Select **Add configuration options**. 2. In the search field, enter `privatelink_access`. 3. From the displayed component names, select the names of the components to enable (`privatelink_access.`). 4. Enable the required components. 5. Select **Save configuration**. tip Each service component can be controlled separately. For example, you can enable Private Link access for your Aiven for Apache Kafka® service, while allowing Kafka® Connect to only be connected via VNet peering. After toggling the values, your Private Link resource will be rebuilt with load balancer rules added for the service component's ports. note For Aiven for Apache Kafka® services, the security group for the VPC endpoint must allow ingress in the port range `10000-31000`. This is to accommodate the pool of Kafka broker ports used in the Private Link implementation. ## Acquire connection information[​](#acquire-connection-information "Direct link to Acquire connection information") ### One Azure Private Link connection[​](#one-azure-private-link-connection "Direct link to One Azure Private Link connection") If you have one private endpoint connected to your Aiven service, you can preview the connection information (URI, hostname, or port required to access the service through the private endpoint) in [Aiven Console](https://console.aiven.io/) > the service's **Overview** page > the **Connection information** section, where you'll also find the switch for the `privatelink` access route. `privatelink`-access-route values for `host` and `port` differ from those for the `dynamic` access route used by default to connect to the service. ### Multiple Azure Private Link connections[​](#multiple-azure-private-link-connections "Direct link to Multiple Azure Private Link connections") Use CLI to acquire connection information for more than one AWS PrivateLink connection. Each endpoint (connection) has PRIVATELINK\_CONNECTION\_ID, which you can check using the [avn service privatelink azure connection list SERVICE\_NAME](/docs/tools/cli/service/privatelink.md) command. To acquire connection information for your service component using Azure Private Link, run the [avn service connection-info](/docs/tools/cli/service/connection-info.md) command. * For SSL connection information for your service component using Azure Private Link, run the following command: ``` avn service connection-info UTILITY_NAME SERVICE_NAME -p PRIVATELINK_CONNECTION_ID ``` Where: * UTILITY\_NAME is `kcat`, for example * SERVICE\_NAME is `kafka-12a3b4c5`, for example * PRIVATELINK\_CONNECTION\_ID is `plc39413abcdef`, for example * For SASL connection information for Aiven for Apache Kafka® service components using Azure Private Link, run the following command: ``` avn service connection-info UTILITY_NAME SERVICE_NAME -p PRIVATELINK_CONNECTION_ID -a sasl ``` Where: * UTILITY\_NAME is `kcat`, for example * SERVICE\_NAME is `kafka-12a3b4c5`, for example * PRIVATELINK\_CONNECTION\_ID is `plc39413abcdef`, for example note SSL certificates and SASL credentials are the same for all the connections. ## Update subscription list[​](#update-subscription-list "Direct link to Update subscription list") In the Aiven CLI, you can update the list of Azure subscriptions that have access to Aiven service endpoints: ``` avn service privatelink azure update AIVEN_SERVICE --user-subscription-id SUBSCRIPTION_ID ``` To update a few subscription IDs, repeat the `SUBSCRIPTION_ID` argument, for example: ``` avn service privatelink azure update AIVEN_SERVICE --user-subscription-id SUBSCRIPTION_ID_1 --user-subscription-id SUBSCRIPTION_ID_2 ``` ## Delete a Private Link service[​](#delete-a-private-link-service "Direct link to Delete a Private Link service") Use the Aiven CLI to delete the Azure Load Balancer and Private Link service: ``` avn service privatelink azure delete AIVEN_SERVICE ``` --- # Manage billing groups Costs associated with services and features in an Aiven project are charged to the payment method assigned to its [billing group](/docs/platform/concepts/billing-and-payment.md#billing-groups). Billing groups let you set up billing profiles and use them across different projects in your organization. A consolidated [invoice](/docs/platform/howto/use-billing-groups.md) is created for each billing group. You must have the `organization:billing:write` [permission](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) to create and manage billing groups. ## Create a billing group[​](#create-a-billing-group "Direct link to Create a billing group") 1. In the organization, click **Billing**. 2. Click **Billing groups**. 3. Click **Create billing group**. 4. Enter a name for the billing group. 5. Select the [addresses](/docs/platform/howto/manage-billing-addresses.md), payment method, and currency. 6. Optional: Enter your **VAT ID**. If you get errors for your VAT ID, you can select **Skip the VAT ID validation**. Your ID will be validated later. 7. Enter the other billing details and click **Create billing group**. ## Update a billing group[​](#update-a-billing-group "Direct link to Update a billing group") To change the name, payment method, billing and shipping addresses, VAT ID, billing contact emails, invoice emails, or other billing details: 1. In the organization, click **Billing**. 2. Click **Billing groups**. 3. Find the billing group to update and click **Edit**. 4. Update the billing group and click **Save changes**. note You cannot change a billing group's payment method to or from a marketplace subscription in the Aiven Console. To use a marketplace subscription as a payment method, [send the subscription details to Aiven support](/docs/platform/howto/list-marketplace-payments.md). ## Assign a billing group to a project[​](#assign-a-billing-group-to-a-project "Direct link to Assign a billing group to a project") You can assign any billing group in your organization to a project. To assign a billing group from another organization, you have to [move the project to that organization](/docs/platform/howto/manage-project.md#move-a-project). You must be an organization admin or have the `organization:projects:write` [permission](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) to assign billing groups. 1. In the project, click **Settings**. 2. In the **Project settings** section, choose a billing group to assign the project to. 3. Click **Save changes**. ## Delete a billing group[​](#delete-a-billing-group "Direct link to Delete a billing group") To delete a billing group, [move all projects](#assign-a-billing-group-to-a-project) assigned to it to a different billing group first. 1. In the organization, click **Billing**. 2. Click **Billing groups**. 3. Find the billing group to delete and click **Details**. 4. Click **Actions** at the top of the page. 5. Click **Delete**. note Billing groups with trial credits and the organizations they are assigned to cannot be deleted. You can delete them after the trial period ends. To delete the billing group, organization, or your account during the trial period [contact Aiven support](/docs/platform/howto/support.md). --- # Use Google Private Service Connect with Aiven services [Early availability](/docs/platform/concepts/service-and-feature-releases.md) Enable Google Private Service Connect and use it with your Aiven-managed services. important * Google Private Service Connect is not supported for [BYOC](/docs/platform/concepts/byoc.md)-hosted services. * To activate Google Private Service Connect for Aiven for PostgreSQL®, [contact us](https://aiven.io/contact?department=1306714). Private Service Connect lets you bring your Aiven services into your networks (virtual private clouds) over a private endpoint. The endpoint receives a private IP address from a range that you assign. Next, connectivity over the private endpoint is routed to your Aiven service. note For consistency, Google Private Service Connect is called *privatelink* in Aiven tools. This applies to all clouds, including Google Cloud. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Google Private Service Connect is an [early availability](/docs/platform/concepts/service-and-feature-releases.md#early-availability-) feature. * Your Aiven service needs to be hosted in [a project virtual private cloud (VPC)](/docs/platform/howto/manage-project-vpc.md) in the region where the connecting endpoint will be created. * [Aiven CLI](/docs/tools/cli.md) * Access to the [Google Cloud console](https://console.cloud.google.com/) * Access to the [Aiven Console](https://console.aiven.io/) note Private Service Connect endpoints are service specific. For each service to connect to, create a separate endpoint. ## Set up a Private Service Connect connection[​](#set-up-a-private-service-connect-connection "Direct link to Set up a Private Service Connect connection") ### Step 1: Enable Private Service Connect for an Aiven service[​](#step-1-enable-private-service-connect-for-an-aiven-service "Direct link to Step 1: Enable Private Service Connect for an Aiven service") Using the Aiven CLI, enable a Private Service Connect for your Aiven service: ``` avn service privatelink google create SERVICE_NAME ``` important For publishing a service over Private Service Connect, a dedicated address range needs to be allocated at the publishing / Aiven end. Aiven reserves network 172.24.0.0/16 for this purpose and forbids creating project VPCs in Google Cloud overlapping with this range. Creating a privatelink usually takes a minute or two. You can use the following command to see the current state: ``` avn service privatelink google get SERVICE_NAME ``` When the state has changed from `creating` to `active`, resources at Aiven end have been allocated, and it's possible to create connections. When the privatelink has been successfully created, you can expect an output similar to the following: ``` GOOGLE_SERVICE_ATTACHMENT STATE ==================================================================================== ====== projects/aivenprod/regions/europe-west1/serviceAttachments/privatelink-s3fd836dfc60 active ``` The `GOOGLE_SERVICE_ATTACHMENT` value is used to connect an endpoint on the client side to the Aiven service. ### Step 2: Create a connection in Google Cloud[​](#step-2-create-a-connection-in-google-cloud "Direct link to Step 2: Create a connection in Google Cloud") Create a Private Service Connect endpoint and connection to your Aiven service: 1. Go to the [Google Cloud console](https://console.cloud.google.com/net-services/psc/addConsumer) (**Networking** > **Network services** > **Private Service Connect** > **CONNECT ENDPOINT**). 2. Select **Published service** as **Target**, and enter the `GOOGLE_SERVICE_ATTACHMENT` value into the **Target service** field. 3. Specify the endpoint name. 4. Select an existing subnet hosting your side of the endpoint. 5. Click **ADD ENDPOINT**. After the endpoint is created, initially it's status is `pending`. To allow connections via the endpoint, it needs to be accepted at the service publisher (Aiven) end. tip If you use an automatically assigned IP address, note the IP address associated with the endpoint to use it later. ### Step 3: Approve the connection in Aiven[​](#step-3-approve-the-connection-in-aiven "Direct link to Step 3: Approve the connection in Aiven") 1. Update the state of Private Service Connect connections for your Aiven service by running: ``` avn service privatelink google refresh SERVICE_NAME ``` 2. Retry the following command until it returns the `pending-user-approval` status: ``` avn service privatelink google connection list SERVICE_NAME ``` ``` PRIVATELINK_CONNECTION_ID PSC_CONNECTION_ID STATE USER_IP_ADDRESS ========================= ================= ===================== =============== plc3fd852bec98 12870921937223780 pending-user-approval null ``` note * `PSC_CONNECTION_ID` comes from Google Cloud and can help you verify that the connection matches your Private Service Connect endpoint. * `PRIVATELINK_CONNECTION_ID` comes from Aiven, and you need it for the final connection approval. 3. To enable the connection, approve it. note By approving the connection, you provide the IP address assigned to your Private Service Connect endpoint - whether automatically assigned or static. Aiven uses this IP address for pointing the service DNS records necessary for the clients to reach the Aiven service through the Private Service Connect connection. Run the following approval command: ``` avn service privatelink google connection approve SERVICE_NAME \ --privatelink-connection-id PRIVATELINK_CONNECTION_ID \ --user-ip-address PSC_ENDPOINT_IP_ADDRESS ``` The connection initially transitions to the `user-approved` state: ``` avn service privatelink google connection list SERVICE_NAME ``` ``` PRIVATELINK_CONNECTION_ID PSC_CONNECTION_ID STATE USER_IP_ADDRESS ========================= ================= ============= =============== plc3fd852bec98 12870921937223780 user-approved 10.0.0.100 ``` 4. You may need to run the `avn service privatelink google refresh` command at this point since updates to service attachment accept lists are not immediately reflected in the states of returned connected endpoints: ``` avn service privatelink google refresh SERVICE_NAME ``` After establishing the connection and populating DNS records, the connection appears as `active`: ``` avn service privatelink google connection list SERVICE_NAME ``` ``` PRIVATELINK_CONNECTION_ID PSC_CONNECTION_ID STATE USER_IP_ADDRESS ========================= ================= ====== =============== plc3fd852bec98 12870921937223780 active 10.0.0.100 ``` The state of your Private Service Connect endpoint in Google Cloud should have transitioned from `pending` to `accepted` at this point. Private Service Connect connectivity has been established now. ### Step 4: Enable the access for service components[​](#step-4-enable-the-access-for-service-components "Direct link to Step 4: Enable the access for service components") Allow connectivity to your Aiven services using the Private Service Connect endpoint. * Console * CLI In the [Aiven Console](https://console.aiven.io/): 1. On the **Overview** page of your service, click **Service settings** in the sidebar. 2. Go to the **Cloud and network** section, and click **Actions** > **More network configurations**. 3. In the **Network configuration** window: 1. Select **Add configuration options**. 2. In the search field, enter `privatelink_access`. 3. From the displayed component names, select the names of the components to enable (`privatelink_access.SERVICE_COMPONENT`). 4. Select the toggle switches for the selected components to enable them. 5. Select **Save configuration**. In the [Aiven CLI](/docs/tools/cli.md), set `user_config.privatelink_access.SERVICE_COMPONENT` to `true` for the components to enable. Take the following command as an example for Aiven for Apache Kafka®: ``` avn service update -c privatelink_access.kafka=true SERVICE_NAME ``` tip Each service component can be controlled separately. For example, you can enable Private Service Connect access for your Aiven for Apache Kafka service while allowing Aiven for Apache Kafka Connect to only be connected via VPC peering. ## Allow cross-region connections[​](#allow-cross-region-connections "Direct link to Allow cross-region connections") Private Service Connect endpoints in Google Cloud can accept traffic from other regions. You configure this setting, called global access, on your own Private Service Connect endpoint. Keep your Private Service Connect endpoint in the same region as your Aiven service. Enabling global access lets clients in other regions, such as Compute Engine VMs, Cloud VPN tunnels, or Cloud Interconnect, reach that endpoint. Enable global access on a new or existing endpoint without disrupting traffic. To do this, run the following command in the [Google Cloud CLI](https://cloud.google.com/sdk/gcloud): ``` gcloud compute forwarding-rules update FORWARDING_RULE_NAME \ --allow-psc-global-access ``` Replace `FORWARDING_RULE_NAME` with the name of your Private Service Connect endpoint from [Step 2: Create a connection in Google Cloud](#step-2-create-a-connection-in-google-cloud). ## Acquire connection information[​](#acquire-connection-information "Direct link to Acquire connection information") ### One Private Service Connect connection[​](#one-private-service-connect-connection "Direct link to One Private Service Connect connection") If you have one private endpoint connected to your Aiven service, preview the connection information (URI, hostname, or port required to access the service through the private endpoint) in the [Aiven Console](https://console.aiven.io/): 1. Go to your Aiven project, and click **Services** in the sidebar. 2. Open your service's **Overview** page, and go to the **Connection information** section. 3. Switch to the `privatelink` access route to preview values for `host` and `port`, which differ from those for the `dynamic` access route used by default to connect to the service. ### Multiple Private Service Connect connections[​](#multiple-private-service-connect-connections "Direct link to Multiple Private Service Connect connections") Use the [Aiven CLI](/docs/tools/cli.md) to acquire connection information for more than one Private Service Connect connection. Each endpoint (connection) has `PRIVATELINK_CONNECTION_ID`, which you can check using the [avn service privatelink google connection list SERVICE\_NAME](/docs/tools/cli/service/privatelink.md) command. To acquire connection information for your service component using Private Service Connect, run the [avn service connection-info](/docs/tools/cli/service/connection-info.md) command. * Get SSL connection information for your service component using Private Service Connect: ``` avn service connection-info UTILITY_NAME SERVICE_NAME -p PRIVATELINK_CONNECTION_ID ``` Where: * `UTILITY_NAME` is `kcat`, for example * `SERVICE_NAME` is `kafka-12a3b4c5`, for example * `PRIVATELINK_CONNECTION_ID` is `plc39413abcdef`, for example * Get SASL connection information for Aiven for Apache Kafka service components using Private Service Connect: ``` avn service connection-info UTILITY_NAME SERVICE_NAME -p PRIVATELINK_CONNECTION_ID -a sasl ``` Where: * `UTILITY_NAME` is `kcat`, for example * `SERVICE_NAME` is `kafka-12a3b4c5`, for example * `PRIVATELINK_CONNECTION_ID` is `plc39413abcdef`, for example note SSL certificates and SASL credentials are the same for all the connections. ## Delete a Private Service Connect connection[​](#delete-a-private-service-connect-connection "Direct link to Delete a Private Service Connect connection") Use the [Aiven CLI](/docs/tools/cli.md) to delete the Private Service Connect connection for your Aiven service: ``` avn service privatelink google delete SERVICE_NAME ``` Related pages * [Use AWS PrivateLink with Aiven services](/docs/platform/howto/use-aws-privatelinks.md) * [Use Azure Private Link with Aiven services](/docs/platform/howto/use-azure-privatelink.md) * [Manage project VPCs](/docs/platform/howto/manage-project-vpc.md) --- # Manage two-factor authentication Two-factor authentication in Aiven provides an extra level of security by requiring a second authentication code in addition to the user password. This only applies to logins using email and password. The Aiven platform cannot enforce 2FA for logins through third-party providers, including identity providers. warning Enabling and disabling two-factor authentication revokes personal tokens that you [created with password authentication](/docs/platform/howto/set-authentication-policies.md#personal-tokens). ## Enable two-factor authentication[​](#enable-2fa "Direct link to Enable two-factor authentication") 1. Click **User information** and select **Authentication**. 2. On the **Aiven Password** method, click the **Two-factor authentication** toggle to the enabled position. 3. Enter your password and click **Next**. 4. On your mobile device, open your authenticator app and scan the QR code shown in Aiven Console. Alternatively, you can enter the TOTP secret from the Aiven Console into your authenticator app. 5. In the Aiven Console enter the **Confirmation code** from the authenticator app. 6. Click **Enable**. To change the mobile device that you use for two-factor authentication, [disable two-factor authentication](/docs/platform/howto/user-2fa.md#disable-2fa) and enable it on the new device. ## Disable two-factor authentication[​](#disable-2fa "Direct link to Disable two-factor authentication") 1. Click **User information** and select **Authentication**. 2. On the **Aiven Password** method, click the **Two-factor authentication** toggle to the disabled position. 3. Enter your password and click **Disable Two-Factor Authentication**. ## Reset two-factor authentication[​](#reset-two-factor-authentication "Direct link to Reset two-factor authentication") If you have lost access to your mobile device or authenticator app, you can regain access to your account by resetting your Aiven password. 1. Log out of Aiven Console. 2. Enter your login email and click **Log in**. 3. Click **Forgot password?**. 4. Enter your login email and click **Reset your password**. 5. [Enable two-factor authentication](/docs/platform/howto/user-2fa.md#enable-2fa) on your new mobile device or authenticator app. --- # View event logs for organizations and organizational units Monitor activity in your Aiven organization with the organization and organizational unit event logs. To view logs at the organization level, you must be an organization admin or have the `organization:event_logs:read` [permission](/docs/platform/concepts/permissions.md). ## View organization logs[​](#view-organization-logs "Direct link to View organization logs") 1. In the organization, click **Admin**. 2. Click **Organization**. 3. In the organization details section, click **Events Log**. ## View organizational unit logs[​](#view-organizational-unit-logs "Direct link to View organizational unit logs") 1. In the organization, click **Admin**. 2. Click **Organization**. 3. In the **Organizational units** section, click the name of a unit. Related pages * [View project logs](/docs/platform/howto/view-project-logs.md) --- # View project logs Monitor activity in your Aiven projects with the project event log. To view the project logs, you must be a project [Admin, Developer, or Operator](/docs/platform/concepts/permissions.md). To view the logs for a project: 1. Click **Projects** and choose a project. 2. Click **Event log**. Related pages * [View organization and unit logs](/docs/platform/howto/view-organization-logs.md) --- # Manage a project VPC peering with Microsoft Azure Set up a peering connection between your [Aiven project VPC](/docs/platform/howto/manage-project-vpc.md) and a [Microsoft Azure virtual network](https://learn.microsoft.com/en-us/azure/virtual-network/create-peering-different-subscriptions). Establishing a peering connection between an Aiven project VPC and an Azure VNet requires creating the peering both from the VPC in Aiven and from the VNet in Azure. To establish the peering from Aiven to Azure, the Aiven Platform's [Active Directory application object](https://learn.microsoft.com/en-us/azure/active-directory/develop/app-objects-and-service-principals) needs permissions in your Azure subscription. Because the peering is between different AD tenants (the Aiven AD tenant and your Azure AD tenant), another application object is needed for your Azure AD tenant to create the peering from Azure to Aiven once granted permissions to do so. Before you start, review the [VPC peering limitations](/docs/platform/howto/list-vpc-peering.md#limitations). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Manage project networking](/docs/platform/concepts/permissions.md#project-permissions) permissions in Aiven * Azure account with at least the application administrator role * [Azure CLI](https://learn.microsoft.com/en-us/cli/azure/?view=azure-cli-latest) and, optionally, the [Microsoft Azure portal](https://portal.azure.com/#home) * [Aiven CLI](/docs/tools/cli.md) installed ## Set up permissions in Azure[​](#set-up-permissions-in-azure "Direct link to Set up permissions in Azure") ### Azure app object permissions[​](#azure-app-object-permissions "Direct link to Azure app object permissions") 1. Log in with an Azure admin account using the Azure CLI: ``` az account clear az login ``` This should open a window in your browser prompting to choose an Azure account to log in with. tip If you manage multiple Azure subscriptions, also configure the Azure CLI to default to the correct subscription for the subsequent commands. This is not needed if there's only one subscription: ``` az account set --subscription SUBSCRIPTION_NAME_OR_ID ``` 2. Create an application object in your AD tenant using the Azure CLI: ``` az ad app create \ --display-name "NAME_OF_YOUR_CHOICE" \ --sign-in-audience AzureADMultipleOrgs \ --key-type Password ``` This creates an application object in Azure AD that can be used to log into multiple AD tenants ( `--sign-in-audience AzureADMultipleOrgs` ), but only the home tenant (the tenant the app was created in) has the credentials to authenticate the app. note Save the `appId` field from the output. It will be referred to as `$user_app_id`. 3. Create a service principal for your app object to the Azure subscription that the VNet to be peered is located in: ``` az ad sp create --id $user_app_id ``` This creates a service principal in your subscription, which can be assigned permissions to peer your VNet. note Save the `id` field from the JSON output. It will be referred to as `$user_sp_id`. 4. Set a password for your app object: ``` az ad app credential reset --id $user_app_id ``` note Save the `password` field from the output. It will be referred to as `$user_app_secret`. 5. Find properties of your virtual network: * Resource ID * In the [Azure portal](https://portal.azure.com/#home): **Virtual networks** > name of your network > **JSON View** > **id** * Using the Azure CLI: ``` az network vnet list ``` tip The `id` field should have format `/subscriptions/$user_subscription_id/ resourceGroups/$user_resource_group/providers/Microsoft.Network/virtualNetworks/$user_vnet_name`. It will be referred to as `$user_vnet_id`. * Azure Subscription ID (the VNet page in the [Azure portal](https://portal.azure.com/#home) > **Essentials** > **Subscription ID**) or the part after `/subscriptions/` in the resource ID. It will be referred to as `$user_subscription_id`. * Resource group name (the VNet page in the [Azure portal](https://portal.azure.com/#home) > **Essentials** > **Resource group**) or the `resourceGroup` field in the output. This will be referred to as `$user_resource_group`. * VNet name (title of the VNet page), or the `name` field from the output. It will be referred to as `$user_vnet_name`. note Save all the properties for later. 6. Grant your service principal permissions to peer. The service principal needs to be assigned a role that has the permission for the `Microsoft.Network/virtualNetworks/virtualNetworkPeerings/write` action on the scope of your VNet. To limit the amount of permissions the application object and the service principal have, you can create a custom role with just that permission. The built-in network contributor role includes that permission. 1. Find the id of the role with the required permission: ``` az role definition list --name "Network Contributor" ``` The `id` field from the output will be referred to as `$network_contributor_role_id`. 2. Assign the service principal the network contributor role using `$network_contributor_role_id`: ``` az role assignment create \ --role $network_contributor_role_id \ --assignee-object-id $user_sp_id \ --scope $user_vnet_id ``` This allows your application object to manage the network in the `--scope`. Since you control the application object, it may also be given permission for the scope of an entire resource group or the whole subscription to allow creating other peerings later without assigning the role for each VNet separately. ### Aiven app object permissions[​](#aiven-app-object-permissions "Direct link to Aiven app object permissions") 1. Create a service principal for the Aiven application object. The Aiven AD tenant contains an application object that the Aiven Platform uses to create a peering from the Aiven project VPC to the Azure VNet. For this, the Aiven application object needs a service principal in your Azure subscription. To create it, run: ``` az ad sp create --id 55f300d4-fc50-4c5e-9222-e90a6e2187fb ``` The argument to `--id` field is a fixed value that represents the ID of the Aiven application object. note Save the `id` field from the JSON output. It will be referred to as `$aiven_sp_id`. important The command might fail for the following reasons: * `When using this permission, the backing application of the service principal being created must in the local tenant`, which means your account doesn't have the required permissions. See [Prerequisites](/docs/platform/howto/vnet-peering-azure.md#prerequisites). * `The service principal cannot be created, updated, or restored because the service principal name 55f300d4-fc50-4c5e-9222-e90a6e2187fb is already in use`, in which case run `az ad sp show --id 55f300d4-fc50-4c5e-9222-e90a6e2187fb` and find `id` in the output. 2. Create a custom role for the Aiven application object. The Aiven application has a service principal that can be granted permissions. To restrict the service principal's permissions to peering, create a custom role with the peering action only allowed: ``` az role definition create --role-definition ' { "Name": "NAME_OF_YOUR_CHOICE", "Description": "Allows creating a peering to vnets in scope (but not from)", "Actions": ["Microsoft.Network/virtualNetworks/peer/action"], "AssignableScopes": ["/subscriptions/'$user_subscription_id'"] }' ``` `AssignableScopes` includes your Azure subscription ID to restrict scopes that a role assignment can use. note Save the `id` field from the output. It will be referred to as `$aiven_role_id`. 3. Assign the custom role to the Aiven service principal. To give the Aiven application object's service principal permissions to peer with your VNet, assign the created role to the Aiven service principal with the scope of your VNet: ``` az role assignment create \ --role $aiven_role_id \ --assignee-object-id $aiven_sp_id \ --scope $user_vnet_id ``` 4. Find your AD tenant ID: * In the [Azure portal](https://portal.azure.com/#home): **Settings** > **Directories + subscriptions** > **Directories** > **Directory ID** * Using the Azure CLI: ``` az account list ``` note Save the `tenantId` field from the output. It will be referred to as `$user_tenant_id`. ## Create the peering in Aiven[​](#create-the-peering-in-aiven "Direct link to Create the peering in Aiven") By creating a peering connection from the Aiven project VPC to the VNet in your Azure subscription, you also create a service principal for the application object (`--peer-azure-app-id $user_app_id`) and grant it the permission to peer with the Aiven project VPC. The Aiven application object authenticates with your Azure tenant to grant it access to [the service principal of the Aiven application object](/docs/platform/howto/vnet-peering-azure.md#aiven-app-object-permissions) (`--peer-azure-tenant-id $user_tenant_id`). Find `$aiven_project_vpc_id` in the [Aiven Console](https://console.aiven.io/) or by running the `avn vpc list` command. 1. Run: ``` avn vpc peering-connection create \ --project-vpc-id $aiven_project_vpc_id \ --peer-cloud-account $user_subscription_id \ --peer-resource-group $user_resource_group \ --peer-vpc $user_vnet_name \ --peer-azure-app-id $user_app_id \ --peer-azure-tenant-id $user_tenant_id ``` note Use lower case for arguments starting with `$user_`. 2. Run the following command until the state changes from `APPROVED` to `PENDING_PEER`: ``` avn vpc peering-connection get -v \ --project-vpc-id $aiven_project_vpc_id \ --peer-cloud-account $user_subscription_id \ --peer-resource-group $user_resource_group \ --peer-vpc $user_vnet_name ``` tip If the state is `INVALID_SPECIFICATION` or `REJECTED_BY_PEER`, check if the Azure VNet exists and if the Aiven application object has the permission to be peered with. Revise your configuration and recreate the peering connection. Establishing the connection from Aiven to Azure can take a while. When completed, the state changes to `PENDING_PEER` and the output shows details for establishing the peering from your Azure VNet to the Aiven project VPC. note Save the following from the output: * `to-tenant-id`: It will be referred to as `$aiven_tenant_id`. * `to-network-id`: It will be referred to as `$aiven_vnet_id`. ## Create the peering in Azure[​](#create-the-peering-in-azure "Direct link to Create the peering in Azure") [Establish the peering connection](https://learn.microsoft.com/en-us/azure/virtual-network/create-peering-different-subscriptions) from your Azure VNet to the Aiven project VPC: 1. Log out the Azure user [you logged in with](/docs/platform/howto/vnet-peering-azure.md#azure-app-object-permissions): ``` az account clear ``` 2. Log in the Azure application object to your AD tenant using the [password](/docs/platform/howto/vnet-peering-azure.md#azure-app-object-permissions): ``` az login \ --service-principal \ -u $user_app_id \ -p $user_app_secret \ --tenant $user_tenant_id ``` 3. Log in the Azure application object to the Aiven AD tenant: ``` az login \ --service-principal \ -u $user_app_id \ -p $user_app_secret \ --tenant $aiven_tenant_id ``` At this point, your application object should have an opened session with your Azure AD tenant and the Aiven AD tenant. 4. Create a peering from your Azure VNet to the Aiven project VPC: ``` az network vnet peering create \ --name PEERING_NAME_OF_YOUR_CHOICE \ --remote-vnet $aiven_vnet_id \ --vnet-name $user_vnet_name \ --resource-group $user_resource_group \ --subscription $user_subscription_id \ --allow-vnet-access ``` If the peering state in the output is `connected`, the peering is created. tip The command might fail with the following error: ``` The client 'RANDOM_UUID' with object id 'RANDOM_UUID' does not have authorization to perform action 'Microsoft.Network/virtualNetworks/virtualNetworkPeerings/write' over scope '$user_vnet_id'. If access was recently granted, refresh your credentials. ``` for two reasons related to the [role assignment](/docs/platform/howto/vnet-peering-azure.md#azure-app-object-permissions): * Role assignment hasn't taken effect yet, in which case try logging in again and recreating the peering. * Role assignment is incorrect, in which case try recreating the role assignment. Wait until the Aiven peering connection is active. The Aiven Platform polls peering connections in state `PENDING_PEER` regularly to see if the peer (your Azure VNet) has created a peering connection to the Aiven project VPC. Once this is detected, the state changes from `PENDING_PEER` to `ACTIVE`, at which point Aiven services in the project VPC can be reached through the peering. 5. Check if the status of the peering connection is `ACTIVE`: ``` avn vpc peering-connection get -v \ --project-vpc-id $aiven_project_vpc_id \ --peer-cloud-account $user_subscription_id \ --peer-resource-group $user_resource_group \ --peer-vpc $user_vnet_name ``` ## Delete the peering[​](#delete-the-peering "Direct link to Delete the peering") important Once you delete your VPC peering on the Aiven Platform, the cloud-provider side of the peering connection becomes `inactive` or `deleted`, and the traffic between the disconnected VPCs is terminated. Delete a project VPC peering using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your project page. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, select a project VPC. 4. On the **VPC details** page, go to the **VPC peering connections** section, find the peering to be deleted, and click **Actions** > **Delete**. 5. In the **Confirmation** window, click **Delete VPC peering**. Run the [avn vpc peering-connection delete](/docs/tools/cli/vpc.md#delete-peering-connections) command: ``` avn vpc peering-connection delete \ --project-vpc-id PROJECT_VPC_ID \ --peer-cloud-account PEER_CLOUD_ACCOUNT \ --peer-vpc PEER_VPC_ID ``` Replace the following with meaningful values: * `PROJECT_VPC_ID`, for example `12345678-1a2b-3c4d-5f6g-1a2b3c4d5e6f` * `PEER_CLOUD_ACCOUNT`, for example `012345678901` * `PEER_VPC_ID`, for example `vpc-abcdef01234567890` Make an API call to the [VpcPeeringConnectionDelete](https://api.aiven.io/doc/#tag/Project/operation/VpcPeeringConnectionDelete) endpoint: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_ID/vpcs/PROJECT_VPC_ID \ --header 'Authorization: Bearer BEARER_TOKEN' \ ``` ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_ID/vpcs/PROJECT_VPC_ID/peering-connections/peer-accounts/PEER_CLOUD_ACCOUNT/peer-vpcs/PEER_VPC \ --header 'Authorization: Bearer BEARER_TOKEN' ``` Replace the following placeholders with meaningful data: * `PROJECT_ID`: Aiven project name * `PROJECT_VPC_ID`: Aiven project VPC ID * `PEER_CLOUD_ACCOUNT`: your cloud provider account ID or name * `PEER_VPC`: your cloud provider VPC ID or name * `BEARER_TOKEN` Related pages * [Manage project VPCs](/docs/platform/howto/manage-project-vpc.md) * [Set up a project VPC peering](/docs/platform/howto/list-project-vpc-peering.md) * [Manage organization VPCs](/docs/platform/howto/manage-organization-vpc.md) --- # Manage a project VPC peering with AWS Set up a peering connection between your Aiven project VPC and an AWS VPC. Before you start, review the [VPC peering limitations](/docs/platform/howto/list-vpc-peering.md#limitations). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Manage project networking](/docs/platform/concepts/permissions.md#project-permissions) permissions * Two VPCs to be peered: a [project VPC](/docs/platform/howto/manage-project-vpc.md) * Access to the [AWS Management Console](https://console.aws.amazon.com) * One of the following tools for operations on the Aiven Platform: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) ## Create a peering connection[​](#create-a-peering-connection "Direct link to Create a peering connection") ### Collect data from AWS[​](#collect-data-from-aws "Direct link to Collect data from AWS") To [create a peering connection in Aiven](/docs/platform/howto/vpc-peering-aws.md#create-a-peering-in-aiven), first collect the required data from AWS: 1. Log in to the [AWS Management Console](https://console.aws.amazon.com) and go to your profile information. 2. Find and save your account ID. 3. Open the navigation menu, and select **All services**. 4. Find **Networking & Content Delivery**, and go to **VPC** > **Your VPCs**. 5. Find a VPC to peer, preview its details, and save its ID and a cloud region that it's located in. ### Create a peering in Aiven[​](#create-a-peering-in-aiven "Direct link to Create a peering in Aiven") With the [data collected from AWS](/docs/platform/howto/vpc-peering-aws.md#collect-data-from-aws), create a project VPC peering connection using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API * Aiven Provider for Terraform 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your project page. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, select a project VPC to peer. 4. On the **VPC details** page, go to the **VPC peering connections** section and click **Create peering request**. 5. In the **Create peering request** window: 1. Enter the following: * **AWS account ID** * **AWS VPC region** * **AWS VPC ID** 2. Click **Create**. Run the [avn vpc peering-connection create](/docs/tools/cli/vpc.md#create-peering-connections) command: ``` avn vpc peering-connection create \ --project-vpc-id AIVEN_PROJECT_VPC_ID \ --peer-cloud-account AWS_ACCOUNT_ID \ --peer-vpc AWS_VPC_ID ``` Replace `AIVEN_PROJECT_VPC_ID`, `AWS_ACCOUNT_ID`, and `AWS_VPC_ID` as needed. Make an API call to the [VpcPeeringConnectionCreate](https://api.aiven.io/doc/#tag/Project/operation/VpcPeeringConnectionCreate) endpoint: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_ID/vpcs/PROJECT_VPC_ID/peering-connections \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data ' { "peer_cloud_account":"AWS_ACCOUNT_ID", "peer_vpc":"AWS_VPC_ID" } ' ``` Replace the following placeholders with meaningful data: * `PROJECT_ID` (Aiven project name) * `PROJECT_VPC_ID` (Aiven project VPC ID) * `BEARER_TOKEN` * `AWS_ACCOUNT_ID` * `AWS_VPC_ID` Use the [aiven\_aws\_vpc\_peering\_connection](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/aws_vpc_peering_connection) resource. This adds a connection with the **Pending peer** status in the [Aiven Console](https://console.aiven.io/) and a connection pending acceptance in the [AWS Management Console](https://console.aws.amazon.com). ### Accept the peering request in AWS[​](#accept-the-peering-request-in-aws "Direct link to Accept the peering request in AWS") 1. Log in to the [AWS Management Console](https://console.aws.amazon.com), open the navigation menu, and select **All services**. 2. Find **Networking & Content Delivery**, and go to **VPC** > **Peering connections**. 3. Find your peering connection from Aiven pending acceptance, select it, and click **Actions** > **Accept request**. 4. Create or update your [AWS route tables](https://docs.aws.amazon.com/vpc/latest/peering/vpc-peering-routing) to match your Aiven CIDR settings. At this point, your peering connection status should be visible as **Active** both in the [Aiven Console](https://console.aiven.io/) and in the [AWS Management Console](https://console.aws.amazon.com). ## Delete the peering[​](#delete-the-peering "Direct link to Delete the peering") important Once you delete your VPC peering on the Aiven Platform, the cloud-provider side of the peering connection becomes `inactive` or `deleted`, and the traffic between the disconnected VPCs is terminated. Delete a project VPC peering using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your project page. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, select a project VPC. 4. On the **VPC details** page, go to the **VPC peering connections** section, find the peering to be deleted, and click **Actions** > **Delete**. 5. In the **Confirmation** window, click **Delete VPC peering**. Run the [avn vpc peering-connection delete](/docs/tools/cli/vpc.md#delete-peering-connections) command: ``` avn vpc peering-connection delete \ --project-vpc-id PROJECT_VPC_ID \ --peer-cloud-account PEER_CLOUD_ACCOUNT \ --peer-vpc PEER_VPC_ID ``` Replace the following with meaningful values: * `PROJECT_VPC_ID`, for example `12345678-1a2b-3c4d-5f6g-1a2b3c4d5e6f` * `PEER_CLOUD_ACCOUNT`, for example `012345678901` * `PEER_VPC_ID`, for example `vpc-abcdef01234567890` Make an API call to the [VpcPeeringConnectionDelete](https://api.aiven.io/doc/#tag/Project/operation/VpcPeeringConnectionDelete) endpoint: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_ID/vpcs/PROJECT_VPC_ID \ --header 'Authorization: Bearer BEARER_TOKEN' \ ``` ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_ID/vpcs/PROJECT_VPC_ID/peering-connections/peer-accounts/PEER_CLOUD_ACCOUNT/peer-vpcs/PEER_VPC \ --header 'Authorization: Bearer BEARER_TOKEN' ``` Replace the following placeholders with meaningful data: * `PROJECT_ID`: Aiven project name * `PROJECT_VPC_ID`: Aiven project VPC ID * `PEER_CLOUD_ACCOUNT`: your cloud provider account ID or name * `PEER_VPC`: your cloud provider VPC ID or name * `BEARER_TOKEN` --- # Manage a project VPC peering with Google Cloud Set up a peering connection between your Aiven project VPC and a Google Cloud VPC. Establishing a peering connection between an Aiven VPC and a Google Cloud VPC requires creating the peering both from the VPC in Aiven and from the VPC in Google Cloud. Before you start, review the [VPC peering limitations](/docs/platform/howto/list-vpc-peering.md#limitations). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Manage project networking](/docs/platform/concepts/permissions.md#project-permissions) permissions * Two VPCs to be peered: a [project VPC](/docs/platform/howto/manage-project-vpc.md) in Aiven and a VPC in your Google Cloud account * Access to the [Google Cloud console](https://console.cloud.google.com/) * One of the following tools for operations on the Aiven Platform: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) ## Create a peering connection[​](#create-a-peering-connection "Direct link to Create a peering connection") ### Collect data from Google Cloud[​](#collect-data-from-google-cloud "Direct link to Collect data from Google Cloud") To [create a peering connection in Aiven](/docs/platform/howto/vpc-peering-gcp.md#create-the-peering-in-aiven), first collect the required data from Google Cloud: 1. Log in to the [Google Cloud console](https://console.cloud.google.com/), open the navigation menu, and select **Cloud overview** > **Dashboard**. 2. Find the **Project info** field, and collect your **Project ID**. 3. Open the navigation menu again, and click **VIEW ALL PRODUCTS** > **Networking** > **VPC Network**. 4. Find a VPC to connect to, and make note of its **Name**. ### Create the peering in Aiven[​](#create-the-peering-in-aiven "Direct link to Create the peering in Aiven") With the [data collected from Google Cloud](/docs/platform/howto/vpc-peering-gcp.md#collect-data-from-google-cloud), create a project VPC peering connection using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API * Aiven Provider for Terraform 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your project page. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, select a project VPC to peer. 4. On the **VPC details** page, go to the **VPC peering connections** section and click **Create peering request**. 5. In the **Create peering request** window: 1. Enter the following: * **GCP project ID** * **GCP VPC network name** 2. Click **Create**. This adds a connection with the **Pending peer** status in the [Aiven Console](https://console.aiven.io/). 6. On the **VPC details** page, go to the **VPC peering connections** section and click **Status details** for the new connection. Make a note of your Aiven VPC network name and project ID. Run the [avn vpc peering-connection create](/docs/tools/cli/vpc.md#create-peering-connections) command: ``` avn vpc peering-connection create \ --project-vpc-id AIVEN_PROJECT_VPC_ID \ --peer-cloud-account GOOGLE_CLOUD_PROJECT_ID \ --peer-vpc GOOGLE_CLOUD_VPC_NETWORK_NAME ``` Replace `AIVEN_PROJECT_VPC_ID`, `GOOGLE_CLOUD_PROJECT_ID`, and `GOOGLE_CLOUD_VPC_NETWORK_NAME` as needed. Make an API call to the [VpcPeeringConnectionCreate](https://api.aiven.io/doc/#tag/Project/operation/VpcPeeringConnectionCreate) endpoint: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_ID/vpcs/PROJECT_VPC_ID/peering-connections \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data ' { "peer_cloud_account":"GOOGLE_CLOUD_PROJECT_ID", "peer_vpc":"GOOGLE_CLOUD_VPC_NETWORK_NAME" } ' ``` Replace the following placeholders with meaningful data: * `PROJECT_ID` (Aiven project name) * `PROJECT_VPC_ID` (Aiven project VPC ID) * `BEARER_TOKEN` * `GOOGLE_CLOUD_PROJECT_ID` * `GOOGLE_CLOUD_VPC_NETWORK_NAME` Use the [aiven\_gcp\_vpc\_peering\_connection](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/gcp_vpc_peering_connection) resource. ### Create the peering in Google Cloud[​](#create-the-peering-in-google-cloud "Direct link to Create the peering in Google Cloud") Use the [data collected in the Aiven Console](/docs/platform/howto/vpc-peering-gcp.md#create-the-peering-in-aiven) to create the VPC peering connection in Google Cloud: 1. Log in to the [Google Cloud console](https://console.cloud.google.com/), open the navigation menu, and click **VIEW ALL PRODUCTS** > **Networking** > **VPC Network** > **VPC network peering** > **CREATE PEERING CONNECTION** > **CONTINUE**. 2. Enter a name for the peering connection. 3. Select your Google Cloud VPC network. 4. In the **Peered VPC network** field, select **In another project**. 5. In the **Project ID** field, enter the Aiven project ID collected in the the [Aiven Console](https://console.aiven.io), on the **VPC details** page, in the **VPC peering connections** section, from **Status details** of the new connection. 6. In the **VPC network name** field, enter the name of your Aiven VPC collected in the the [Aiven Console](https://console.aiven.io), on the **VPC details** page, in the **VPC peering connections** section, from **Status details** of the new connection. 7. Click **Create**. As soon as the peering is created, the connection status changes to **Active** both in the [Aiven Console](https://console.aiven.io) and in the [Google Cloud console](https://console.cloud.google.com/). ## Set up multiple project VPC peerings[​](#set-up-multiple-project-vpc-peerings "Direct link to Set up multiple project VPC peerings") To peer multiple Google Cloud VPC networks to your Aiven-managed project VPC, [add peering connections](/docs/platform/howto/vpc-peering-gcp.md#create-a-peering-connection) one by one in the [Aiven Console](https://console.aiven.io). For the limit on the number of VPC peering connections allowed to a single VPC network, see the [Google Cloud documentation](https://cloud.google.com/vpc/docs/quota). ## Delete the peering[​](#delete-the-peering "Direct link to Delete the peering") important Once you delete your VPC peering on the Aiven Platform, the cloud-provider side of the peering connection becomes `inactive` or `deleted`, and the traffic between the disconnected VPCs is terminated. Delete a project VPC peering using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your project page. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, select a project VPC. 4. On the **VPC details** page, go to the **VPC peering connections** section, find the peering to be deleted, and click **Actions** > **Delete**. 5. In the **Confirmation** window, click **Delete VPC peering**. Run the [avn vpc peering-connection delete](/docs/tools/cli/vpc.md#delete-peering-connections) command: ``` avn vpc peering-connection delete \ --project-vpc-id PROJECT_VPC_ID \ --peer-cloud-account PEER_CLOUD_ACCOUNT \ --peer-vpc PEER_VPC_ID ``` Replace the following with meaningful values: * `PROJECT_VPC_ID`, for example `12345678-1a2b-3c4d-5f6g-1a2b3c4d5e6f` * `PEER_CLOUD_ACCOUNT`, for example `012345678901` * `PEER_VPC_ID`, for example `vpc-abcdef01234567890` Make an API call to the [VpcPeeringConnectionDelete](https://api.aiven.io/doc/#tag/Project/operation/VpcPeeringConnectionDelete) endpoint: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_ID/vpcs/PROJECT_VPC_ID \ --header 'Authorization: Bearer BEARER_TOKEN' \ ``` ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_ID/vpcs/PROJECT_VPC_ID/peering-connections/peer-accounts/PEER_CLOUD_ACCOUNT/peer-vpcs/PEER_VPC \ --header 'Authorization: Bearer BEARER_TOKEN' ``` Replace the following placeholders with meaningful data: * `PROJECT_ID`: Aiven project name * `PROJECT_VPC_ID`: Aiven project VPC ID * `PEER_CLOUD_ACCOUNT`: your cloud provider account ID or name * `PEER_VPC`: your cloud provider VPC ID or name * `BEARER_TOKEN` --- # Manage a project VPC peering with UpCloud Set up a peering connection between your Aiven project VPC and an UpCloud SDN network. Establishing a peering connection between an Aiven VPC and an UpCloud SDN network requires creating the peering both from the VPC in Aiven and from the SDN network in UpCloud. * Setting up the peering from Aiven to UpCloud in the [Aiven Console](https://console.aiven.io/) requires the UpCloud SDN network UUID. To find it, you can use either the [UpCloud Control Panel](https://hub.upcloud.com/) or the [UpCloud API](https://developers.upcloud.com/1.3/). * Setting up the peering from UpCloud to Aiven is possible either in the [UpCloud Control Panel](https://hub.upcloud.com/) or through the [UpCloud API](https://developers.upcloud.com/1.3/). ## Limitations[​](#limitations "Direct link to Limitations") * Peering connections are only supported between networks of type `private`. * You cannot initiate a peering between two networks with overlapping CIDR ranges. * The networks to be peered need to be in the same cloud zone. These are in addition to the platform-wide [VPC peering limitations](/docs/platform/howto/list-vpc-peering.md#limitations). important Make sure you only create peerings between accounts, platforms, or networks you trust. There is no limit on what traffic can flow between the peered components. The server firewall has no effect on `private` type networks. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Manage project networking](/docs/platform/concepts/permissions.md#project-permissions) permissions * Two networks to be peered: a [project VPC](/docs/platform/howto/manage-project-vpc.md) in Aiven and an SDN network in your UpCloud account * Either access to the [UpCloud Control Panel](https://hub.upcloud.com/) or the [UpCloud API](https://developers.upcloud.com/1.3/) * One of the following tools for operations on the Aiven Platform: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) ## Create a peering connection[​](#create-a-peering-connection "Direct link to Create a peering connection") ### Collect data from UpCloud[​](#collect-data-from-upcloud "Direct link to Collect data from UpCloud") To [create a peering connection in Aiven](/docs/platform/howto/vpc-peering-upcloud.md#create-the-peering-in-aiven), first collect the required data from UpCloud using either the [UpCloud Control Panel](https://hub.upcloud.com/) or the [UpCloud API](https://developers.upcloud.com/1.3/): * UpCloud Control Panel * UpCloud API 1. Log in to the [UpCloud Control Panel](https://hub.upcloud.com/), and go to **Network** > **Private networks**. 2. Find the network to peer, and copy its UUID located under its name. Send a request to the [get network details](https://developers.upcloud.com/1.3/13-networks/#get-network-details) UpCloud API endpoint. In the response, you'll get the UpCloud SDN network's UUID. ### Create the peering in Aiven[​](#create-the-peering-in-aiven "Direct link to Create the peering in Aiven") With the [data collected from UpCloud](/docs/platform/howto/vpc-peering-upcloud.md#collect-data-from-upcloud), create an organization VPC peering connection using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your project page. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, select a project VPC to peer. 4. On the **VPC details** page, go to the **VPC peering connections** section and click **Create peering request**. 5. In the **Create peering request** window: 1. Enter your UpCloud SDN network UUID in the **UpCloud Network UUID** field. 2. Click **Create**. This adds a connection with the **Pending peer** status in the [Aiven Console](https://console.aiven.io/). 6. While still on the **VPC details** page, make a note of the **ID** of your Aiven VPC. Run the [avn vpc peering-connection create](/docs/tools/cli/vpc.md#create-peering-connections) command: ``` avn vpc peering-connection create \ --project-vpc-id AIVEN_PROJECT_VPC_ID \ --peer-cloud-account upcloud \ --peer-vpc UPCLOUD_SDN_NETWORK_UUID ``` Replace `AIVEN_PROJECT_VPC_ID` and `UPCLOUD_SDN_NETWORK_UUID` as needed. Make an API call to the [VpcPeeringConnectionCreate](https://api.aiven.io/doc/#tag/Project/operation/VpcPeeringConnectionCreate) endpoint: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_ID/vpcs/PROJECT_VPC_ID/peering-connections \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data ' { "peer_cloud_account":"upcloud", "peer_vpc":"UPCLOUD_SDN_NETWORK_UUID" } ' ``` Replace the following placeholders with meaningful data: * `PROJECT_ID` (Aiven project name) * `PROJECT_VPC_ID` (Aiven project VPC ID) * `BEARER_TOKEN` * `UPCLOUD_SDN_NETWORK_UUID` ### Create the peering in UpCloud[​](#create-the-peering-in-upcloud "Direct link to Create the peering in UpCloud") Use the Aiven VPC network ID [collected in the Aiven Console](/docs/platform/howto/vpc-peering-upcloud.md#create-the-peering-in-aiven) to create the VPC peering connection in UpCloud either in the [UpCloud Control Panel](https://hub.upcloud.com/) or through the [UpCloud API](https://developers.upcloud.com/1.3/): * UpCloud Control Panel * UpCloud API 1. Log in to the [UpCloud Control Panel](https://hub.upcloud.com/), and go to **Network** > **Peering**. 2. Click **Create network peering**, and in the **Create network peering** window: 1. Specify the peering name. 2. Select the source peer network, which is your UpCloud SDN network. 3. Provide the UUID of the target peer network, which is the Aiven network ID. 4. Click **Create**. This creates the peering connection between your Aiven VPC and your UpCloud SDN network. Send a request to the [create network peering](https://developers.upcloud.com/1.3/13-networks/#create-network-peering) UpCloud API endpoint. ``` POST /1.3/network-peering HTTP/1.1 { "network_peering": { "configured_status": "active", "name": "NAME_OF_YOUR_PEERING", "network": { "uuid": "UPCLOUD_SDN_NETWORK_UUID" }, "peer_network": { "uuid": "AIVEN_NETWORK_ID" } } } ``` #### Attributes[​](#attributes "Direct link to Attributes") | Attribute | Accepted value | Default value | Required | Description | Example value | | ------------------- | -------------------------- | ------------- | -------- | ------------------------------------------------------------------------ | -------------------------------------- | | `configured_status` | `active` or `disabled` | `active` | No | Controls whether the peering is administratively up or down. | `active` | | `name` | String of 1-255 characters | None | Yes | Descriptive name for the peering | `peering upcloud->aiven` | | `network.uuid` | Valid network UUID | None | Yes | Sets the local network of the peering. Use the UpCloud SDN network UUID. | `03126dc1-a69f-4bc2-8b24-e31c22d64712` | | `peer_network.uuid` | Valid network UUID | None | Yes | Sets the peer network of the peering. Use the Aiven network ID. | `03585987-bf7d-4544-8e9b-5a1b4d74a333` | #### Expected response[​](#expected-response "Direct link to Expected response") note The sample response provided describes a peering established one way only. If your peering API request is successful, you can expect a response similar to the following: ``` HTTP/1.1 201 Created { "network_peering": { "configured_status": "active", "name": "NAME_OF_YOUR_PEERING", "network": { "ip_networks": { "ip_network": [ { "address": "192.168.0.0/24", "family": "IPv4" }, { "address": "fc02:c4f3::/64", "family": "IPv6" } ] }, "uuid": "UPCLOUD_SDN_NETWORK_UUID" }, "peer_network": { "uuid": "AIVEN_VPC_NETWORK_UUID" }, "state": "pending-peer", "uuid": "PEERING_UUID" } } ``` #### Error responses[​](#error-responses "Direct link to Error responses") | HTTP status | Error code | Description | | ------------- | -------------------------- | -------------------------------- | | 409 Conflict | LOCAL\_NETWORK\_NO\_ROUTER | The local network has no router. | | 404 Not found | NETWORK\_NOT\_FOUND | The local network was not found. | | 404 Not found | PEER\_NETWORK\_NOT\_FOUND | The peer network was not found. | | 409 Conflict | PEERING\_CONFLICT | The peering already exists. | ## Renew a DHCP lease[​](#renew-a-dhcp-lease "Direct link to Renew a DHCP lease") You only need to perform this step if any of your VMs have been created before setting up the network peering. In this case, refresh the Dynamic Host Configuration Protocol (DHCP) lease for a relevant network interface to get new routes. warning A peering connection between an Aiven VPC and VMs created before the peering setup won't work unless you refresh the DHCP lease for a relevant network interface. To refresh the DHCP lease for a network interface, run the following commands: 1. To clear the existing DHCP lease ``` dhclient -r NETWORK_INTERFACE_NAME ``` 2. To request a renewal of the DHCP lease ``` dhclient NETWORK_INTERFACE_NAME ``` ## Delete the peering[​](#delete-the-peering "Direct link to Delete the peering") important Once you delete your VPC peering on the Aiven Platform, the cloud-provider side of the peering connection becomes `inactive` or `deleted`, and the traffic between the disconnected VPCs is terminated. Delete a project VPC peering using a tool of your choice: * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your project page. 2. Click **VPCs** in the sidebar. 3. On the **Virtual private clouds** page, select a project VPC. 4. On the **VPC details** page, go to the **VPC peering connections** section, find the peering to be deleted, and click **Actions** > **Delete**. 5. In the **Confirmation** window, click **Delete VPC peering**. Run the [avn vpc peering-connection delete](/docs/tools/cli/vpc.md#delete-peering-connections) command: ``` avn vpc peering-connection delete \ --project-vpc-id PROJECT_VPC_ID \ --peer-cloud-account PEER_CLOUD_ACCOUNT \ --peer-vpc PEER_VPC_ID ``` Replace the following with meaningful values: * `PROJECT_VPC_ID`, for example `12345678-1a2b-3c4d-5f6g-1a2b3c4d5e6f` * `PEER_CLOUD_ACCOUNT`, for example `012345678901` * `PEER_VPC_ID`, for example `vpc-abcdef01234567890` Make an API call to the [VpcPeeringConnectionDelete](https://api.aiven.io/doc/#tag/Project/operation/VpcPeeringConnectionDelete) endpoint: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_ID/vpcs/PROJECT_VPC_ID \ --header 'Authorization: Bearer BEARER_TOKEN' \ ``` ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_ID/vpcs/PROJECT_VPC_ID/peering-connections/peer-accounts/PEER_CLOUD_ACCOUNT/peer-vpcs/PEER_VPC \ --header 'Authorization: Bearer BEARER_TOKEN' ``` Replace the following placeholders with meaningful data: * `PROJECT_ID`: Aiven project name * `PROJECT_VPC_ID`: Aiven project VPC ID * `PEER_CLOUD_ACCOUNT`: your cloud provider account ID or name * `PEER_VPC`: your cloud provider VPC ID or name * `BEARER_TOKEN` Related pages * [Manage Virtual Private Cloud (VPC) peering](/docs/platform/howto/manage-project-vpc.md) * [Set up Virtual Private Cloud (VPC) peering on AWS](/docs/platform/howto/vpc-peering-aws.md) * [Set up Virtual Private Cloud (VPC) peering on Google Cloud Platform (GCP)](/docs/platform/howto/vpc-peering-gcp.md) * [Set up Azure virtual network peering](/docs/platform/howto/vnet-peering-azure.md) --- # Manage a service in a VPC Manage your Aiven services in a VPC, including setup, migration, and accessing resources securely within your project VPC. Custom domain restrictions in VPCs When you set a custom domain for a service deployed in a VPC, the service certificate is only created for the `public-*` hostname and the custom domain. Certificate verification will fail for the `private-*` hostname and the dynamic service name. To avoid certificate verification issues, ensure your applications connect using either the `public-*` hostname or the custom domain when accessing VPC-deployed services. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") You can manage services either in a project VPC or in an organization VPC. * Project VPC * Organization VPC - [Manage project networking](/docs/platform/concepts/permissions.md#project-permissions) permissions - Tool for operating services in VPCs: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) * [Manage organization networking](/docs/platform/concepts/permissions.md#organization-permissions) permissions * Tool for operating services in VPCs: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) ## Create a service in a VPC[​](#create-a-service-in-a-vpc "Direct link to Create a service in a VPC") You can create a service either in a project VPC or in an organization VPC. * Project VPC * Organization VPC Your project VPC is available as a geolocation (cloud region) for the new service. note You can create a service in a project VPC only if it is in the same project where you are creating the service. Create a service in a project VPC using a tool of your choice: * Aiven Console * CLI * API When you create a service in the Aiven Console, select your project VPC as the cloud region. Run [avn service create](/docs/tools/cli/service-cli.md#avn-cli-service-create): ``` avn service create SERVICE_NAME \ --project PROJECT_NAME \ --project-vpc-id PROJECT_VPC_ID \ --service-type SERVICE_TYPE \ --plan SERVICE_PLAN \ --cloud CLOUD_PROVIDER_REGION ``` Replace the following: * `SERVICE_NAME` with the name of the service to be created, for example, `pg-vpc-test` * `PROJECT_NAME` with the name of the project where to create the service, for example, `pj-test` * `PROJECT_VPC_ID` with the ID of your project VPC, for example, `12345678-1a2b-3c4d-5f6g-1a2b3c4d5e6f` * `SERVICE_TYPE` with the type of the service to be created, for example, `pg` * `SERVICE_PLAN` with the plan of the service to be created, for example, `hobbyist` * `CLOUD_PROVIDER_REGION` with the cloud provider and region to host the service to be created, for example `aws-eu-west-1` Make an API call to the [ServiceCreate endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) endpoint: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data-raw ' { "service_name": "SERVICE_NAME", "cloud": "CLOUD_PROVIDER_REGION", "plan": "SERVICE_PLAN", "service_type": "SERVICE_TYPE", "disk_space_mb": DISK_SIZE, "project_vpc_id":"PROJECT_VPC_ID" } ' ``` Replace the following placeholders with meaningful data: * `PROJECT_NAME`, for example `org-vpc-test` * `BEARER_TOKEN` * `SERVICE_NAME`, for example `org-vpc-test-project` * `CLOUD_PROVIDER_REGION`, for example `google-europe-west10` * `SERVICE_PLAN`, for example `startup-4` * `SERVICE_TYPE`, for example `pg` * `DISK_SIZE` in MiB, for example `81920` * `PROJECT_VPC_ID` Your organization VPC is available as a geolocation (cloud region) for the new service. note You can create a service in an organization VPC only if: * The organization VPC is in the same organization where you are creating the service. * For the service to be created, you use the cloud provider and region that hosts the organization VPC. Create a service in an organization VPC using a tool of your choice: * Console * CLI * API When you create a service in the Aiven Console, select your organization VPC as the cloud region. Run [avn service create](/docs/tools/cli/service-cli.md#avn-cli-service-create): ``` avn service create SERVICE_NAME \ --project PROJECT_NAME \ --project-vpc-id ORGANIZATION_VPC_ID \ --service-type SERVICE_TYPE \ --plan SERVICE_PLAN \ --cloud CLOUD_PROVIDER_REGION ``` Replace the following: * `SERVICE_NAME` with the name of the service to be created, for example, `pg-vpc-test` * `PROJECT_NAME` with the name of the project where to create the service, for example, `pj-test` * `ORGANIZATION_VPC_ID` with the ID of your organization VPC, for example, `12345678-1a2b-3c4d-5f6g-1a2b3c4d5e6f` * `SERVICE_TYPE` with the type of the service to be created, for example, `pg` * `SERVICE_PLAN` with the plan of the service to be created, for example, `hobbyist` * `CLOUD_PROVIDER_REGION` with the cloud provider and region to host the organization VPC, for example `aws-eu-west-1` Make an API call to the [ServiceCreate](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) endpoint: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data-raw ' { "service_name": "SERVICE_NAME", "cloud": "CLOUD_PROVIDER_REGION", "plan": "SERVICE_PLAN", "service_type": "SERVICE_TYPE", "disk_space_mb": DISK_SIZE, "project_vpc_id":"ORGANIZATION_VPC_ID" } ' ``` Replace the following placeholders with meaningful data: * `PROJECT_NAME`, for example `org-vpc-test` * `BEARER_TOKEN` * `SERVICE_NAME`, for example `org-vpc-test-project` * `CLOUD_PROVIDER_REGION`, for example `google-europe-west10` * `SERVICE_PLAN`, for example `startup-4` * `SERVICE_TYPE`, for example `pg` * `DISK_SIZE` in MiB, for example `81920` * `ORGANIZATION_VPC_ID` ## Migrate a service to a VPC[​](#migrate-a-service-to-a-vpc "Direct link to Migrate a service to a VPC") You can migrate a service either to a project VPC or to an organization VPC. * Project VPC * Organization VPC Your project VPC is available as a geolocation (cloud region) for your service. note You can migrate a service to a project VPC only if the project VPC is in the same project running your service. Migrate a service to a project VPC using a tool of your choice: * Console * CLI * API 1. In your service, click **Service settings** from the sidebar. 2. In the **Cloud and network** section, click **Actions** > **Change cloud**. 3. In the **VPCs** tab , select a cloud provider and region, and click **Change**. Run [avn service update](/docs/tools/cli/service-cli.md#avn-cli-service-update): ``` avn service update SERVICE_NAME \ --project-vpc-id PROJECT_VPC_ID ``` Replace the following: * `SERVICE_NAME` with the name of the service to be migrated, for example, `pg-test` * `PROJECT_VPC_ID` with the ID of your project VPC where to migrate the service, for example, `12345678-1a2b-3c4d-5f6g-1a2b3c4d5e6f` Call the [ServiceUpdate endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) to set `project_vpc_id` of the service to the ID of your project VPC: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ -H 'Authorization: Bearer BEARER_TOKEN' \ -H 'content-type: application/json' \ --data '{"project_vpc_id": "PROJECT_VPC_ID"}' ``` Replace the following placeholders with meaningful data: * `PROJECT_NAME`, for example `org-vpc-test` * `SERVICE_NAME`, for example `org-vpc-service` * `BEARER_TOKEN` * `PROJECT_VPC_ID` Your organization VPC is available as a geolocation (cloud region) for your service. note You can only migrate a service to an organization VPC if: * The organization VPC is in the same organization where the service runs. * The service and the organization VPC are hosted using the same cloud provider and region. Migrate a service to an organization VPC using a tool of your choice: * Console * CLI * API 1. In your service, click **Service settings** from the sidebar. 2. In the **Cloud and network** section, click **Actions** > **Change cloud**. 3. In the **VPCs** tab , select a cloud provider and region, and click **Change**. Run [avn service update](/docs/tools/cli/service-cli.md#avn-cli-service-update): ``` avn service update SERVICE_NAME \ --project-vpc-id ORGANIZATION_VPC_ID \ --project PROJECT_NAME ``` Replace the following: * `SERVICE_NAME` with the name of the service to be migrated, for example, `pg-test` * `ORGANIZATION_VPC_ID` with the ID of your organization VPC where to migrate the service, for example, `12345678-1a2b-3c4d-5f6g-1a2b3c4d5e6f` * `PROJECT_NAME` with the name of the project where your service resides, for example, `pj-test` Call the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint to set `vpc_id` of the service to the ID of your organization VPC: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ -H 'Authorization: Bearer BEARER_TOKEN' \ -H 'content-type: application/json' \ --data '{"project_vpc_id": "ORGANIZATION_VPC_ID"}' ``` Replace the following placeholders with meaningful data: * `PROJECT_NAME`, for example `org-vpc-test` * `SERVICE_NAME`, for example `org-vpc-service` * `BEARER_TOKEN` * `ORGANIZATION_VPC_ID` ## Migrate a service deployed in a VPC to another cloud[​](#migrate-a-service-deployed-in-a-vpc-to-another-cloud "Direct link to Migrate a service deployed in a VPC to another cloud") Aiven doesn't natively support automatic migration of a service from a VPC in one cloud provider to another. The migration is possible manually by following these generic instructions, which may need to be adapted to meet specific security or compliance requirements: 1. [Create a service in the destination cloud/VPC](/docs/platform/howto/vpc-service-management.md#create-a-service-in-a-vpc). 2. Set up replication or export/import, depending on the service: 1. Aiven for PostgreSQL®, Aiven for MySQL® or similar: Use `pg_dump`, `pg_restore`, logical replication, or Aiven’s replication features. 2. Aiven for Apache Kafka®: Use [Aiven for Apache Kafka® MirrorMaker 2](/docs/products/kafka/kafka-mirrormaker.md) or Confluent Replicator. 3. Sync data and test the new setup. 4. Cut over traffic to the new service. 5. Decommission the old service. note Reach out to your account team if you need more migration guidance or best practices. ## Access a service deployed in a VPC from the public internet[​](#access-a-service-deployed-in-a-vpc-from-the-public-internet "Direct link to Access a service deployed in a VPC from the public internet") When you move your service to a VPC, access from public networks is blocked by default. If you switch to public access, a separate endpoint is created with a public prefix. You can enable public internet access for your services by following the [Enable public access in a VPC](/docs/platform/howto/public-access-in-vpc.md) instructions. IP filtering is available for a service deployed to a VPC. It's recommended to [use IP filtering](/docs/platform/howto/restrict-access.md#restrict-access) when your VPC service is also exposed to the public internet. note If your service is within a VPC, the VPC configuration filters incoming traffic before the IP filter is applied. Safelisting applies to both internal and external traffic. If you safelist an external IP address and want to keep traffic flowing with the internal (peered) connections, safelist the CIDR blocks of the peered networks as well to avoid disruptions to the service. --- # Change your password You can change your password for the Aiven Platform in your account information. [Two-factor authentication](/docs/platform/howto/user-2fa.md) is disabled when you change your password. Enable it again to keep your account secure. ## Change your password in your profile[​](#change-your-password-in-your-profile "Direct link to Change your password in your profile") 1. Click **User information** > **Authentication**. 2. In the **Aiven Password** section, click **Change password**. 3. Enter your current and new passwords, and click **Change password**. ## Reset your password[​](#reset-your-password "Direct link to Reset your password") If you forgot your password and can't log in: 1. On the login page, enter your email address. 2. Click **Log in**. 3. Click **Forgot password?**. 4. Enter your email and click **Reset your password**. --- # End of life for Aiven services Learn about the upcoming end of life (EOL) for select Aiven services, including timelines, actions after end of life, recommended migration options, and next steps. ## Aiven for M3DB[​](#aiven-for-m3db "Direct link to Aiven for M3DB") **EOL date**: April 30, 2025 After April 30, 2025, all running Aiven for M3DB services are powered off and deleted, making data from these services inaccessible. The recommended alternative for your metrics management and analysis is [Aiven for Metrics](/docs/products/metrics.md). ## Aiven for InfluxDB[​](#aiven-for-influxdb "Direct link to Aiven for InfluxDB") **EOL date**: April 30, 2025 After April 30, 2025, all active Aiven for InfluxDB services are powered off and deleted, making data from these services inaccessible. The recommended alternative for monitoring and observability of your applications and infrastructure is [Aiven for Metrics](/docs/products/metrics.md). ## Aiven for Caching[​](#aiven-for-caching "Direct link to Aiven for Caching") **EOL date**: March 31, 2025 After March 31, 2025, Aiven for Caching services are automatically upgraded to **Aiven for Valkey™** to maintain Redis compatibility. The recommended alternative that offers high performance, scalability, and security is the managed, in-memory NoSQL database service: [Aiven for Valkey™](/docs/products/valkey.md). ## Aiven for Apache Cassandra®[​](#aiven-for-apache-cassandra "Direct link to Aiven for Apache Cassandra®") **EOL date**: January 7, 2026 Since January 7, 2026, Aiven for Apache Cassandra services are no longer available, and data they hosted is inaccessible. ## Aiven for AlloyDB Omni[​](#aiven-for-alloydb-omni "Direct link to Aiven for AlloyDB Omni") **EOL date**: December 5, 2025 Since December 5, 2025, Aiven for AlloyDB Omni services are no longer available, and data they hosted is inaccessible. ### Migration options[​](#migration-options "Direct link to Migration options") The recommended alternatives to Aiven for AlloyDB Omni are: * [Aiven for PostgreSQL®](/docs/products/postgresql.md) * [Aiven for ClickHouse®](/docs/products/clickhouse.md) * [Aiven for MySQL®](/docs/products/mysql.md) ## Aiven for Dragonfly[​](#aiven-for-dragonfly "Direct link to Aiven for Dragonfly") ### Service impact[​](#service-impact "Direct link to Service impact") #### EOA date: June 17, 2026[​](#eoa-date-june-17-2026 "Direct link to EOA date: June 17, 2026") After the end-of-availability (EOA) date, you can **no longer create new services**. Your existing services remain operational until the EOL date. #### EOL date: September 30, 2026[​](#eol-date-september-30-2026 "Direct link to EOL date: September 30, 2026") After the end-of-life (EOL) date, all running **services are powered off and deleted**, making data from these services inaccessible. ### Migration[​](#migration "Direct link to Migration") The recommended alternative that offers high performance, scalability, and security is the managed, in-memory NoSQL database service: [Aiven for Valkey™](/docs/products/valkey.md). See the [guide for migrating from Aiven for Dragonfly to Aiven for Valkey™](/docs/products/valkey/howto/migrate-dragonfly-to-valkey.md). To ensure uninterrupted service, complete your migration before the EOL date. For further assistance, contact the [Aiven support team](mailto:support@aiven.io) or your account team. --- # Aiven service and tool version lifecycle Learn about version lifecycle policies, end of life (EOL) schedules, upgrade procedures, and best practices for Aiven services and tools, including both multi-versioned services and single-versioned services. note EOL is the date after which Aiven services and tools are no longer supported or maintained. ## Aiven version support and upstream EOL[​](#aiven-version-support-and-upstream-eol "Direct link to Aiven version support and upstream EOL") Aiven aims to follow the EOL schedule set by the original authors and maintainers of the open source software (the upstream projects). Once the upstream project retires a specific version, they do not receive security updates and critical bug fixes anymore by the maintainers. Outdated services don't offer the level of protection you need, so Aiven follows the upstream project's EOL schedule to ensure that Aiven services are always running on supported versions. ## Service version numbering[​](#service-version-numbering "Direct link to Service version numbering") Aiven services inherit the upstream project's software versioning scheme. Depending on the service, a major version can be either a single digit (for example, PostgreSQL® 14) or in the format `major.minor` (for example, Kafka® 3.2). The exact version of the service is visible in the [Aiven Console](https://console.aiven.io/) when the service is running. ## Multi-versioned vs single-versioned services[​](#multi-versioned-vs-single-versioned-services "Direct link to Multi-versioned vs single-versioned services") There are two types of Aiven services with respect to versions: * [Multi-versioned services](/docs/platform/reference/eol-for-major-versions.md#aiven-multi-versioned-services-eol) * Multiple service versions supported at a time * Service versions managed by the users: You select a version for your service from the available supported versions. * [Single-versioned services](/docs/platform/reference/eol-for-major-versions.md#aiven-single-versioned-services-eol) * Only one default service version available at a time * Service versions managed by Aiven ## Service version EOL policy[​](#service-version-eol-policy "Direct link to Service version EOL policy") The Aiven service version EOL policy applies only to [multi-versioned services](/docs/platform/reference/eol-for-major-versions.md#aiven-multi-versioned-services-eol), where you select a version. [Single-versioned services](/docs/platform/reference/eol-for-major-versions.md#aiven-single-versioned-services-eol), which run a single version managed by Aiven, are not included. This policy covers both running and powered-off services on affected versions. ## EOL notifications[​](#eol-notifications "Direct link to EOL notifications") When Aiven sets the EOL date for a service major version: * You receive an email notification along with instructions on the next steps. * The [Aiven Console](https://console.aiven.io/) shows an EOL alert for affected services. * You receive email reminders monthly. * In the month of the EOL date, you receive weekly reminders. ## EOL best practices[​](#eol-best-practices "Direct link to EOL best practices") * Fork your service to test the version upgrade before upgrading your production services. * Upgrade to the supported version before the EOL date. This gives you time to test compatibility, resolve any issues, and plan the upgrade on your schedule. After the EOL date: * If the service is powered on, it's automatically upgraded to the latest version when possible, or to another supported version. note If it's not possible to upgrade a powered-on service to a supported version, the service is powered off and ultimately deleted. * If the service is powered off, it's deleted. Exception Aiven for OpenSearch® powered-off services are not deleted after their version EOL. They are upgraded and start running the new version when powered on. ## Aiven multi-versioned services EOL[​](#aiven-multi-versioned-services-eol "Direct link to Aiven multi-versioned services EOL") ### Aiven for MySQL®[​](#aiven-for-mysql "Direct link to Aiven for MySQL®") | Version | Aiven EOL | Service creation supported until | Service creation supported from | | ------- | ---------- | -------------------------------- | ------------------------------- | | 8.0.x | 2026-10-31 | 2026-04-30 | 2018-05-18 | | 8.4.x | 2032-10-30 | 2032-04-30 | 2026-04-30 | ### Aiven for OpenSearch®[​](#aiven-for-opensearch "Direct link to Aiven for OpenSearch®") Aiven for OpenSearch® is the open source continuation of the original Elasticsearch service. The EOL for Aiven for OpenSearch® is generally dependent on the upstream project. Some major versions are designated long-term support (LTS) releases, marked as such in the following table. | Version | Aiven EOL | After EOL | Service creation supported until | Service creation supported from | | ---------- | ------------ | ---------------------------------------- | -------------------------------- | ------------------------------- | | 1.3.x | 2026-07-26 | Automatic upgrade to 2.19 | 2026-07-26 | 2022-05-19 | | 2.17.x | 2026-07-26 | Automatic upgrade to 2.19 | 2026-07-26 | 2024-10-15 | | 2.19.x LTS | Date not set | Automatic upgrade to a supported version | Date not set | 2025-09-15 | | 3.3.x | 2027-02-01 | Automatic upgrade to a supported version | 2027-02-01 | 2026-01-20 | | 3.6.x LTS | Date not set | Automatic upgrade to a supported version | Date not set | 2026-06-23 | ### Aiven for PostgreSQL®[​](#aiven-for-postgresql "Direct link to Aiven for PostgreSQL®") Aiven for PostgreSQL® major versions reach EOL on the same date as the upstream open source project's EOL. | Version | Aiven EOL | Service creation supported until | Service creation supported from | | ------- | ---------- | -------------------------------- | ------------------------------- | | 9.5 | 2021-04-15 | 2021-01-26 | 2015-12-22 | | 9.6 | 2021-11-11 | 2021-05-11 | 2016-09-29 | | 10 | 2022-11-10 | 2022-05-10 | 2017-01-14 | | 11 | 2023-11-09 | 2023-05-09 | 2017-03-06 | | 12 | 2024-11-14 | 2024-05-14 | 2019-11-18 | | 13 | 2025-11-13 | 2025-05-13 | 2021-02-15 | | 14 | 2026-11-12 | 2026-05-12 | 2021-11-11 | | 15 | 2027-11-11 | 2027-05-12 | 2022-12-12 | | 16 | 2028-11-09 | 2028-05-09 | 2024-01-08 | | 17 | 2029-11-08 | 2029-05-08 | 2024-12-09 | | 18 | 2030-11-07 | 2030-05-07 | 2025-09-25 | ### Aiven for Apache Kafka®[​](#aiven-for-kafka "Direct link to Aiven for Apache Kafka®") Aiven for Apache Kafka versions reach EOL one year after they become available on the Aiven Platform. | Version | Aiven EOL | Service creation supported until | Service creation supported from | | ------- | ---------- | -------------------------------- | ------------------------------- | | 3.8.x | 2026-09-30 | 2026-06-30 | 2024-09-06 | | 3.9.x | 2027-09-30 | 2027-06-30 | 2025-03-20 | | 4.0.x | 2026-09-18 | 2026-06-18 | 2025-09-18 | | 4.1.x | 2027-01-31 | 2026-09-10 | 2025-12-10 | | 4.2.x | 2027-06-15 | 2027-03-15 | 2026-06-15 | note Apache Kafka 3.8 is the last version that supports ZooKeeper. Starting with Apache Kafka 3.9, Aiven for Apache Kafka uses KRaft (Kafka Raft) to manage metadata and controllers instead of ZooKeeper. For details about the migration process and rollout limitations, see: * [KRaft in Aiven for Apache Kafka®](/docs/products/kafka/concepts/kraft-mode.md) * [Transitioning to KRaft](/docs/products/kafka/concepts/upgrade-procedure.md#transitioning-to-kraft) To support the transition to KRaft, Aiven supports Apache Kafka 3.8 until the EOL date shown in the table. The EOL date already includes the extended support period. ### Aiven for ClickHouse®[​](#aiven-for-clickhouse "Direct link to Aiven for ClickHouse®") | Version | Aiven EOL | Service creation supported until | Service creation supported from | | ------- | ---------- | -------------------------------- | ------------------------------- | | 25.3 | 2026-09-30 | 2026-08-17 | 2025-12-15 | | 25.8 | 2027-02-28 | 2026-11-30 | 2026-03-25 | | 26.3 | 2027-09-15 | 2027-06-15 | 2026-08-01 | For details, see the [Aiven for ClickHouse version support policy](/docs/products/clickhouse/reference/version-support-policy.md). ### Aiven for Apache Flink®[​](#aiven-for-flink "Direct link to Aiven for Apache Flink®") Service sunset New service creation is no longer available. Existing services remain available during the sunset period. | Version | Aiven EOL | Service creation supported until | Service creation supported from | | ------- | --------------- | -------------------------------- | ------------------------------- | | 1.16 | 2024-11-21 | 2024-08-21 | 2023-01-01 | | 1.19 | N/A | N/A | 2024-05-21 | | 1.20 | To be announced | To be announced | 2024-09-21 | ### Aiven for Valkey™[​](#aiven-for-valkey "Direct link to Aiven for Valkey™") | Version | Aiven EOL | Service creation supported until | Service creation supported from | | ------- | --------------- | -------------------------------- | ------------------------------- | | 8.1.x | To be announced | To be announced | 2025-11-18 | | 9.0.x | 2026-08-31 | 2026-08-31 | 2026-03-09 | | 9.1.x | To be announced | To be announced | 2026-07-15 | ## Aiven single-versioned services EOL[​](#aiven-single-versioned-services-eol "Direct link to Aiven single-versioned services EOL") ### Aiven for Dragonfly®[​](#aiven-for-dragonfly "Direct link to Aiven for Dragonfly®") | Version | Aiven EOL | Service creation supported until | | ------- | ---------- | -------------------------------- | | 1.39.0 | 2026-09-30 | 2026-06-17 | ### Aiven for Grafana®[​](#aiven-for-grafana "Direct link to Aiven for Grafana®") | Version | Aiven EOL | | ------- | --------------- | | 11.6.5 | To be announced | ## Aiven API lifecycle[​](#aiven-api-lifecycle "Direct link to Aiven API lifecycle") The Aiven API endpoints follow a lifecycle that includes the following stages: * **Experimental**: New API endpoints that are still in development and might change without notice. These endpoints are intended for testing and feedback purposes. * **Stable**: API endpoints that are fully supported and maintained. * **Deprecated**: API endpoints that Aiven plans to remove in the future. * **Sunset**: API endpoints that are no longer available. ### API endpoint deprecation[​](#api-endpoint-deprecation "Direct link to API endpoint deprecation") As the Aiven Platform evolves, some API endpoints become outdated or replaced by newer versions. Aiven announces API endpoint deprecations on the [Aiven product updates page](https://aiven.io/changelog). If a replacement endpoint or sunset date is available, this information is included in the deprecation notice and in the API documentation. To allow clients to detect these changes automatically, the API returns specific headers with the deprecation status and sunset date, for example: ``` HTTP/1.1 200 OK Content-Type: application/json Deprecation: @1777248000 Sunset: Wed, 01 Jul 2026 00:00:00 GMT Link: ; rel="sunset" ``` The response headers provide the following information: * **`Deprecation`**: The UTC timestamp when deprecation took effect, in the RFC 9745 `@UNIX-TIMESTAMP` format. * **`Sunset`**: Optional. The date and time when the endpoint becomes unavailable. * **`Link`**: The [product update](https://aiven.io/changelog) URL for the deprecation. Aiven works to reduce the disruptions caused by deprecations. The time between the deprecation and sunset statuses varies based on the endpoint's usage and the migration complexity. During the deprecation period, the endpoint remains fully functional for existing customers, giving you time to migrate to the newer version. Deprecated endpoints may not be available to new customers. ### API endpoint sunset[​](#api-endpoint-sunset "Direct link to API endpoint sunset") After the deprecation period, the endpoint transitions to sunset status. The route remains registered for a period after sunset so clients receive a `410 Gone` response instead of `404 Not Found`. The following is an example of the structured error body: ``` { "errors": [{ "error_code": "retired_api_endpoint", "message": "This API endpoint was deprecated on Tue, 09 Jul 2024 00:00:00 GMT and is no longer available. Use https://api.aiven.io/v1/organization/{organization_id}/user-groups instead." }] } ``` Full route removal happens only after an extended post-sunset period. Migrate before the published sunset date. ## Aiven tools EOL[​](#aiven-tools-eol "Direct link to Aiven tools EOL") Aiven offers multiple tools for interacting with the Aiven Platform and services. These include the Aiven CLI, the Aiven Provider for Terraform, and the Aiven Operator for Kubernetes®. Breaking changes in the Aiven API can result in new major versions of the Aiven tools. While backwards compatibility is typically maintained, certain changes require Aiven to deprecate older versions of the tools. ### Aiven CLI[​](#aiven-cli "Direct link to Aiven CLI") | Version | Aiven EOL | | ------- | --------------- | | 1.x | 2023-12-11 | | 2.x | 2023-12-11 | | 3.x | 2023-12-11 | | 4.x | To be announced | ### Aiven Provider for Terraform[​](#aiven-provider-for-terraform "Direct link to Aiven Provider for Terraform") After an Aiven Provider for Terraform version reaches EOL, it receives no new features or bug fixes but remains functional. | Version | Aiven EOL | | ------- | --------------- | | 1.x | 2023-12-31 | | 2.x | 2023-12-31 | | 3.x | 2023-12-31 | | 4.x | To be announced | ### Aiven Operator for Kubernetes®[​](#aiven-operator-for-kubernetes "Direct link to Aiven Operator for Kubernetes®") | Version | Aiven EOL | | ------- | --------------- | | 0.x | To be announced | --- # Get resource IDs Resource IDs like organization ID or user ID can be useful for working with the developer tools. You can get the IDs for resources in the [Aiven Console](https://console.aiven.io/). To access the IDs in the **Admin** part of the Console, you must be an [organization admin](/docs/platform/concepts/permissions.md#organization-roles-and-permissions) ## Get an organization ID[​](#get-an-organization-id "Direct link to Get an organization ID") Go to **User information** > **Organizations**. You can see the ID for each organization you were invited to or are a member of. ## Get an organizational unit ID[​](#get-an-organizational-unit-id "Direct link to Get an organizational unit ID") 1. Click **Admin**. 2. Click **Organization**. 3. In the list of **Organizational units**, get the **ID**. ## Get a user ID[​](#get-a-user-id "Direct link to Get a user ID") 1. Click **Admin**. 2. Click **Users**. 3. Get the user ID from the URL. User IDs start with `u` and are usually at the end of the URL. For example, the user ID is `u1abc2345678` in the following URL: `https://console.aiven.io/account/a1f23bcdef4a/admin/users/u1abc2345678` ## Get an application user ID[​](#get-an-application-user-id "Direct link to Get an application user ID") 1. Click **Admin**. 2. Click **Application users**. 3. Find the application user and get its **User ID**. ## Get a group ID[​](#get-a-group-id "Direct link to Get a group ID") 1. Click **Admin**. 2. Click **Groups**. 3. Get the group ID from the URL. Group IDs start with `ug` and are usually at the end of the URL. For example, the user ID is `ug1a2bc3d45e6` in the following URL: `https://console.aiven.io/account/a1f23bcdef4a/admin/groups/ug1a2bc3d45e6` ## Get a billing group ID[​](#get-a-billing-group-id "Direct link to Get a billing group ID") 1. Click **Billing**. 2. Click **Billing groups**. The ID is under the billing group names. ## Get an address ID[​](#get-an-address-id "Direct link to Get an address ID") 1. Click **Billing**. 2. Click **Addresses**. The ID for each address is under the address name. --- # Available cloud regions This is a reference list of the default cloud regions available per provider on the Aiven Platform. note * The list of available clouds can differ per project. * Not all Aiven service types are available for all cloud providers and regions. * For availability zone (AZ) support, see [Availability zones](/docs/platform/concepts/availability-zones.md#aiven-services-across-availability-zones). ## OVH[​](#ovh "Direct link to OVH") | Region | Cloud | Description | | ------------- | ------------- | ---------------------------------- | | asia-pacific | avn-ovh-mum | Asia, India: Mumbai | | asia-pacific | avn-ovh-sgp1 | Asia Pacific, Singapore: Singapore | | asia-pacific | avn-ovh-syd1 | Asia Pacific, Australia: Sydney | | europe | avn-ovh-de1 | Europe, Germany: Frankfurt | | europe | avn-ovh-gra | Europe, France: Gravelines | | europe | avn-ovh-gra1 | Europe, France: Gravelines | | europe | avn-ovh-gra11 | Europe, France: Gravelines | | europe | avn-ovh-gra3 | Europe, France: Gravelines | | europe | avn-ovh-gra5 | Europe, France: Gravelines | | europe | avn-ovh-gra7 | Europe, France: Gravelines | | europe | avn-ovh-gra9 | Europe, France: Gravelines | | europe | avn-ovh-mil | Europe, Italy: Milan | | europe | avn-ovh-par | Europe, France: Paris | | europe | avn-ovh-rbx-a | Europe, France: Roubaix | | europe | avn-ovh-sbg5 | Europe, France: Strasbourg | | europe | avn-ovh-sbg7 | Europe, France: Strasbourg | | europe | avn-ovh-uk1 | Europe, United Kingdom: London | | europe | avn-ovh-waw1 | Europe, Poland: Warsaw | | north america | avn-ovh-bhs | North America, Canada: Beauharnois | | north america | avn-ovh-bhs1 | North America, Canada: Beauharnois | | north america | avn-ovh-bhs3 | North America, Canada: Beauharnois | | north america | avn-ovh-bhs5 | North America, Canada: Beauharnois | ## Amazon Web Services[​](#amazon-web-services "Direct link to Amazon Web Services") | Region | Cloud | Description | | ------------- | ------------------ | ---------------------------------------- | | africa | aws-af-south-1 | Africa, South Africa: Cape Town | | asia-pacific | aws-ap-east-1 | Asia, Hong Kong: Hong Kong | | asia-pacific | aws-ap-northeast-1 | Asia, Japan: Tokyo | | asia-pacific | aws-ap-northeast-2 | Asia, Korea: Seoul | | asia-pacific | aws-ap-northeast-3 | Asia, Japan: Osaka | | asia-pacific | aws-ap-south-1 | Asia, India: Mumbai | | asia-pacific | aws-ap-south-2 | Asia, India: Hyderabad | | asia-pacific | aws-ap-southeast-1 | Asia, Singapore: Singapore | | asia-pacific | aws-ap-southeast-3 | Asia, Jakarta: Jakarta | | australia | aws-ap-southeast-2 | Australia, New South Wales: Sydney | | australia | aws-ap-southeast-4 | Australia, Melbourne: Melbourne | | europe | aws-eu-central-1 | Europe, Germany: Frankfurt | | europe | aws-eu-central-2 | Europe, Switzerland: Zurich | | europe | aws-eu-north-1 | Europe, Sweden: Stockholm | | europe | aws-eu-south-1 | Europe, Italy: Milan | | europe | aws-eu-south-2 | Europe, Spain: Madrid | | europe | aws-eu-west-1 | Europe, Ireland: Ireland | | europe | aws-eu-west-2 | Europe, England: London | | europe | aws-eu-west-3 | Europe, France: Paris | | north america | aws-ca-central-1 | Canada, Quebec: Canada Central | | north america | aws-us-east-1 | United States, Virginia: N. Virginia | | north america | aws-us-east-2 | United States, Ohio: Ohio | | north america | aws-us-west-1 | United States, California: N. California | | north america | aws-us-west-2 | United States, Oregon: Oregon | | south america | aws-sa-east-1 | South America, Brazil: São Paulo | ## Azure[​](#azure "Direct link to Azure") | Region | Cloud | Description | | ------------- | ------------------------- | ---------------------------------------------- | | africa | azure-south-africa-north | Africa, South Africa: South Africa North | | asia-pacific | azure-eastasia | Asia, Hong Kong: East Asia | | asia-pacific | azure-india-central | Asia, India: Central India | | asia-pacific | azure-india-south | Asia, India: South India | | asia-pacific | azure-indonesia-central | Asia, Indonesia: Indonesia Central | | asia-pacific | azure-japaneast | Asia, Japan: Japan East | | asia-pacific | azure-japanwest | Asia, Japan: Japan West | | asia-pacific | azure-korea-central | Asia, Korea: Korea Central | | asia-pacific | azure-korea-south | Asia, Korea: Korea South | | asia-pacific | azure-southeastasia | Asia, Singapore: Southeast Asia | | australia | azure-australia-central | Australia, Canberra: Australia Central | | australia | azure-australiaeast | Australia, New South Wales: Australia East | | australia | azure-australiasoutheast | Australia, Victoria: Australia Southeast | | europe | azure-france-central | Europe, France: France Central | | europe | azure-germany-north | Europe, Germany: Germany North | | europe | azure-germany-westcentral | Europe, Germany: Germany West Central | | europe | azure-northeurope | Europe, Ireland: North Europe | | europe | azure-norway-east | Europe, Norway: Norway East | | europe | azure-norway-west | Europe, Norway: Norway West | | europe | azure-sweden-central | Europe, Gävle: Sweden Central | | europe | azure-switzerland-north | Europe, Switzerland: Switzerland North | | europe | azure-uksouth | Europe, England: UK South | | europe | azure-ukwest | Europe, Wales: UK West | | europe | azure-westeurope | Europe, Netherlands: West Europe | | middle east | azure-qatar-central | Middle East, Doha: Qatar Central | | middle east | azure-uae-north | Middle East, United Arab Emirates: Middle East | | north america | azure-canadacentral | Canada, Ontario: Canada Central | | north america | azure-canadaeast | Canada, Quebec: Canada East | | north america | azure-centralus | United States, Iowa: Central US | | north america | azure-eastus | United States, Virginia: East US | | north america | azure-eastus2 | United States, Virginia: East US 2 | | north america | azure-mexico-central | North America, Mexico: Mexico Central | | north america | azure-northcentralus | United States, Illinois: North Central US | | north america | azure-southcentralus | United States, Texas: South Central US | | north america | azure-westcentralus | United States, Wyoming: West Central US | | north america | azure-westus | United States, California: West US | | north america | azure-westus2 | United States, Washington: West US 2 | | north america | azure-westus3 | United States, Phoenix: West US 3 | | south america | azure-brazilsouth | South America, Brazil: Brazil South | ## DigitalOcean[​](#digitalocean "Direct link to DigitalOcean") | Region | Cloud | Description | | ------------- | ------ | ---------------------------------------- | | asia-pacific | do-blr | Asia, India: Bangalore | | asia-pacific | do-sgp | Asia, Singapore: Singapore | | australia | do-syd | Australia, New South Wales: Sydney | | europe | do-ams | Europe, Netherlands: Amsterdam | | europe | do-fra | Europe, Germany: Frankfurt | | europe | do-lon | Europe, England: London | | north america | do-nyc | United States, New York: New York | | north america | do-sfo | United States, California: San Francisco | | north america | do-tor | Canada, Ontario: Toronto | ## Google Cloud[​](#google-cloud "Direct link to Google Cloud") | Region | Cloud | Description | | ------------- | ------------------------------ | --------------------------------------------- | | africa | google-africa-south1 | Africa, South Africa: Johannesburg | | asia-pacific | google-asia-east1 | Asia, Taiwan: Taiwan | | asia-pacific | google-asia-east2 | Asia, Hong Kong: Hong Kong | | asia-pacific | google-asia-northeast1 | Asia, Japan: Tokyo | | asia-pacific | google-asia-northeast2 | Asia, Japan: Osaka | | asia-pacific | google-asia-northeast3 | Asia, Korea: Seoul | | asia-pacific | google-asia-south1 | Asia, India: Mumbai | | asia-pacific | google-asia-south2 | Asia, India: Delhi | | asia-pacific | google-asia-southeast1 | Asia, Singapore: Singapore | | asia-pacific | google-asia-southeast2 | Asia, Indonesia: Jakarta | | australia | google-australia-southeast1 | Australia, New South Wales: Sydney | | australia | google-australia-southeast2 | Australia, Victoria: Melbourne | | europe | google-europe-central2 | Europe, Poland: Warsaw | | europe | google-europe-north1 | Europe, Finland: Finland | | europe | google-europe-north2 | Europe, Sweden: Stockholm | | europe | google-europe-southwest1 | Europe, Madrid: Spain | | europe | google-europe-west1 | Europe, Belgium: Belgium | | europe | google-europe-west10 | Europe, Germany: Berlin | | europe | google-europe-west12 | Europe, Italy: Turin | | europe | google-europe-west2 | Europe, England: London | | europe | google-europe-west3 | Europe, Germany: Frankfurt | | europe | google-europe-west4 | Europe, Netherlands: Netherlands | | europe | google-europe-west6 | Europe, Switzerland: Zürich | | europe | google-europe-west8 | Europe, Italy: Milan | | europe | google-europe-west9 | Europe, France: Paris | | middle east | google-me-central1 | Middle East, Qatar: Doha | | middle east | google-me-central2 | Middle East, Saudi Arabia: Dammam | | middle east | google-me-west1 | Middle East, Israel: Tel Aviv | | north america | google-northamerica-northeast1 | Canada, Quebec: Montréal | | north america | google-northamerica-northeast2 | Canada, Ontario: Toronto | | north america | google-northamerica-south1 | North America, Mexico: Querétaro | | north america | google-us-central1 | United States, Iowa: Iowa | | north america | google-us-east1 | United States, South Carolina: South Carolina | | north america | google-us-east4 | United States, Virginia: Northern Virginia | | north america | google-us-east5 | United States, Ohio: Columbus | | north america | google-us-south1 | United States, Texas: Dallas | | north america | google-us-west1 | United States, Oregon: Oregon | | north america | google-us-west2 | United States, California: Los Angeles | | north america | google-us-west3 | United States, Utah: Salt Lake City | | north america | google-us-west4 | United States, Nevada: Las Vegas | | south america | google-southamerica-east1 | South America, Brazil: Sao Paulo | | south america | google-southamerica-west1 | South America, Chile: Santiago | ## UpCloud[​](#upcloud "Direct link to UpCloud") | Region | Cloud | Description | | ------------- | --------------- | ----------------------------------- | | asia-pacific | upcloud-sg-sin | Asia, Singapore: Singapore | | australia | upcloud-au-syd | Australia, New South Wales: Sydney | | europe | upcloud-de-fra | Europe, Germany: Frankfurt | | europe | upcloud-dk-cph | Europe, Denmark: Copenhagen | | europe | upcloud-es-mad | Europe, Spain: Madrid | | europe | upcloud-fi-hel | Europe, Finland: Helsinki | | europe | upcloud-fi-hel1 | Europe, Finland: Helsinki | | europe | upcloud-fi-hel2 | Europe, Finland: Helsinki | | europe | upcloud-nl-ams | Europe, Netherlands: Amsterdam | | europe | upcloud-no-svg | Europe, Norway: Oslo | | europe | upcloud-pl-waw | Europe, Poland: Warsaw | | europe | upcloud-se-sto | Europe, Sweden: Stockholm | | europe | upcloud-uk-lon | Europe, England: London | | north america | upcloud-us-chi | United States, Illinois: Chicago | | north america | upcloud-us-nyc | United States, New York: New York | | north america | upcloud-us-sjo | United States, California: San Jose | ## Oracle Cloud Infrastructure [Limited availability](/docs/platform/concepts/service-and-feature-releases.md)[​](#oracle-cloud-infrastructure- "Direct link to oracle-cloud-infrastructure-") note Oracle Cloud Infrastructure (OCI) is supported on the Aiven Platform in [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-) for [BYOC](/docs/platform/concepts/byoc.md) only. For more information or access, contact your account team. | Region | Cloud | Description | | ------------- | -------------- | ------------------------------------------ | | Asia-Pacific | ap-melbourne-1 | Australia, Australia Southeast: Melbourne | | Asia-Pacific | ap-mumbai-1 | India, India West: Mumbai | | Asia-Pacific | ap-osaka-1 | Japan, Japan Central: Osaka | | Asia-Pacific | ap-seoul-1 | South Korea, South Korea Central: Seoul | | Asia-Pacific | ap-singapore-1 | Singapore, Singapore: Singapore | | Asia-Pacific | ap-singapore-2 | Singapore, Singapore West: Singapore | | Asia-Pacific | ap-sydney-1 | Australia, Australia East: Sydney | | Asia-Pacific | ap-tokyo-1 | Japan, Japan East: Tokyo | | Europe | eu-frankfurt-1 | Germany, Germany Central: Frankfurt | | Europe | eu-milan-1 | Italy, Italy Northwest: Milan | | Europe | eu-turin-1 | Italy, Italy Northwest: Turin | | Europe | uk-london-1 | United Kingdom, UK South: London | | Middle East | me-dubai-1 | UAE, UAE East: Dubai | | Middle East | me-jeddah-1 | Saudi Arabia, Saudi Arabia West: Jeddah | | Middle East | me-riyadh-1 | Saudi Arabia, Saudi Arabia Central: Riyadh | | North America | ca-montreal-1 | Canada, Canada Southeast: Montreal | | North America | us-ashburn-1 | US East, Virginia: Ashburn | | North America | us-phoenix-1 | US West, Arizona: Phoenix | | South America | sa-saopaulo-1 | Brazil, Brazil East: São Paulo | --- # Password policy Aiven is committed to keeping your data secure. Creating a strong password makes it harder for attackers to gain unauthorized access to your account. Creating a strong password is a first step in securing your account. You can add another layer of security by [enabling two-factor authentication](/docs/platform/howto/user-2fa.md). ## Password requirements[​](#password-requirements "Direct link to Password requirements") Aiven enforces the following rules for password strength: * Minimum length is 8 characters * Cannot contain single repeating characters such as `aaaaaaaa` * Cannot contain your name or email address * Cannot contain common words, phrases, or strings such as password, security, or common names * Cannot contain words that are very similar to common words such as `password1` These rules are also used for service integration passwords. For remote services (for example, sending logs to an external OpenSearch® service), these rules are not enforced, but they are recommended. ## Password tips[​](#password-tips "Direct link to Password tips") The following are some suggestions for creating or resetting your Aiven password: * Use a password manager to create a randomly generated strong password * Use passphrases since these are harder to guess * Do not use the same password for multiple services --- # Refer Aiven and earn credits Invite someone to sign up to Aiven using your referral link and both of you get credits to spend when they start the Aiven trial. To refer someone to Aiven: 1. In the Aiven Console, click **User information** > **Referrals**. 2. Copy your personal referral link. When the user signs up, you get an email with an Aiven credit code that is valid for 6 months. --- # Default service IP address and hostname When a new Aiven service is created, it automatically gets a hostname and one or more public IP addresses. ## Default service IP address[​](#default-service-ip-address "Direct link to Default service IP address") The chosen cloud service provider will dynamically assign one or more public IP address from their connection pool. This IP address is not permanent, and with every service maintenance (in case of failover, maintenance upgrade or cloud migration) the IP address changes since Aiven creates a new node, migrates the existing data to it and retire the old node. note Aiven also offer the ability to define [static IP addresses](/docs/platform/concepts/static-ips.md) if you need them in a service. For more information about obtaining a static IP and assigning it to a particular service, see the [related guide](/docs/platform/concepts/static-ips.md). If you have your own cloud account and want to keep your Aiven services isolated from the public internet, you can however create a VPC and a peering connection to your own cloud account. For more information on how to setup the VPC peering, check [the related article](/docs/platform/howto/manage-project-vpc.md). ## Default service hostname[​](#default-service-hostname "Direct link to Default service hostname") When a new service is being provisioned, its hostname is defined as follows: ``` -.*.aivencloud.com ``` Where: * `` is the name of the service. * `` is the name of the project. * `*` is a variable component consisting of one or more levels of alphanumeric subdomains for the purpose of load balancing between DNS zones. important Always use a fully qualified domain name returned by Aiven API. Make sure your code doesn't put any constraints on the domain part or format of the returned service hostname. note * Second-level domain part of `aivencloud.com` can change to another name in the future if the domain becomes unavailable for updates. * If the `` is too short or was recently used (for example, if you drop and recreate a service with the same name), the hostname format can be `<3RANDOMLETTERS>-.*.aivencloud.com`. --- # Aiven for ClickHouse® Aiven for ClickHouse® is a fully managed distributed columnar database based on open source ClickHouse - a fast, resource effective solution tailored for data warehouse and generation of real-time analytical data reports using advanced SQL queries. Discover Aiven for ClickHouse's key features and attributes which let you focus on turning business data into actionable insights. ClickHouse is a highly scalable fault-tolerant database designed for online analytical processing (OLAP) and data warehousing. Aiven for ClickHouse enables you to execute complex SQL queries on large datasets effectively to process large amounts of data in real time. On top of that, it supports built-in data integrations for [Aiven for Kafka®](/docs/products/kafka.md) and [Aiven for PostgreSQL®](/docs/products/postgresql.md). ## Effortless setup[​](#effortless-setup "Direct link to Effortless setup") With the managed ClickHouse service, you can offload on Aiven multiple time-consuming and laborious operations on your data infrastructure: database initialization and configuration, cluster provisioning and management, or your infrastructure maintenance and monitoring are off your shoulders. **Pre-configured settings:** The managed ClickHouse service is pre-configured with a rational set of parameters and settings appropriate for the plan you have selected. ## Easy management[​](#easy-management "Direct link to Easy management") * **Scalability:** You can seamlessly [scale your ClickHouse cluster](/docs/products/clickhouse/howto/change-service-plan.md) horizontally or vertically as your data and needs change using the pre-packaged plans. You can also [scale disk storage](/docs/products/clickhouse/howto/scale-disk-storage.md) independently of the plan. Aiven for ClickHouse also supports [sharding](/docs/products/clickhouse/howto/use-shards-with-distributed-table.md) as a horizontal cluster scaling strategy. * **Resource tags:** You can assign metadata to your services in the form of tags. They help you organize, search, and filter Aiven resources. You can [tag your service](/docs/products/clickhouse/howto/tag-service.md) by purpose, owner, environment, or any other criteria. * **Forking:** Forking an Aiven for ClickHouse service creates a new database service containing the latest snapshot of an existing service. Forks don't stay up-to-date with the parent database, but you can write to them. It provides a risk-free way of working with your production data and schema. For example, you can use them to test upgrades, new schema migrations, or load test your app with a different plan. Learn how to [fork an Aiven for ClickHouse® service](/docs/products/clickhouse/howto/fork-service.md). ## Effective maintenance[​](#effective-maintenance "Direct link to Effective maintenance") * **Automatic maintenance updates:** With 99.99% SLA, Aiven makes sure that the ClickHouse software and the underlying platform stays up-to-date with the latest patches and updates with zero downtime. You can set [maintenance windows](/docs/products/clickhouse/howto/maintenance-updates.md) for your service to make sure the changes occur during times that do not affect productivity. * **Backups and disaster recovery:** Aiven for ClickHouse has automatic backups taken every 24 hours. The retention period depends on your plan tier. See [disaster recovery](/docs/products/clickhouse/concepts/disaster-recovery.md), [schedule backups](/docs/products/clickhouse/howto/configure-backup.md), and [Plan comparison](https://aiven.io/pricing?product=clickhouse\&tab=plan-comparison). ## Intelligent observability[​](#intelligent-observability "Direct link to Intelligent observability") * **Service health monitoring:** Aiven for ClickHouse provides metrics and logs for your cluster at no additional charge. You can enable pre-integrated Aiven observability services, such as Aiven for Grafana®, Aiven for Metrics, or Aiven for OpenSearch® or push available metrics and logs to external observability tools, such as Prometheus, AWS CloudWatch, or Google Cloud Logging. For more details, see [Monitor Aiven for ClickHouse metrics](/docs/products/clickhouse/howto/monitor-performance.md). * **Notifications and alerts:** The service is pre-configured to alert you on, for example, your disk running out of space or CPU consumption running high when resource usage thresholds are exceeded. Email notifications are sent to admins and technical contacts of the project under which your service is created. Check [Receive technical notifications](/docs/platform/howto/technical-emails.md) to learn how you can sign up for such alerts. ## Security and compliance[​](#security-and-compliance "Direct link to Security and compliance") * **Single tenancy:** Your service runs on dedicated instances. This offers true data isolation that contributes to the optimal protection and an increased security. * **Network isolation:** Aiven platform supports VPC peering as a mechanism for connecting directly to your ClickHouse service via private IP. This provides a more secure network setup. The platform also supports PrivateLink connectivity. * **Regulatory compliance:** ClickHouse runs on Aiven platform that is ISO 27001:2013, SOC2, GDPR, HIPAA, and PCI/DSS compliant. * **Role based Access Control (RBAC)**. To learn what kind of granular access is possible in Aiven for ClickHouse, see [RBAC with Zookeeper](/docs/products/clickhouse/concepts/service-architecture.md#zookeeper). * **Zero lock-in:** Aiven for ClickHouse offers compatibility with open source software (OSS), which protects you from software and vendor lock-in. You can migrate between clouds and regions. See more details on security and compliance in Aiven for ClickHouse in [Secure a managed ClickHouse® service](/docs/products/clickhouse/howto/secure-service.md). ## Devops-friendly tools[​](#devops-friendly-tools "Direct link to Devops-friendly tools") * **Automation:** [Aiven Provider for Terraform](https://registry.terraform.io/providers/aiven/aiven/latest/docs) helps you automate the orchestration of your ClickHouse clusters. * **Command-line tooling:** [Aiven CLI](/docs/tools/cli.md) client provides greater flexibility of use for proficient administrators allowing scripting repetitive actions with ease. * **REST APIs:** [Aiven APIs](/docs/tools/api.md) allow you to manage Aiven resources in a programmatic way using HTTP requests. The whole functionality available via Aiven Console is also available via APIs enabling you to build custom integrations with ClickHouse and the Aiven platform. Related pages * [Aiven.io](https://aiven.io/clickhouse) --- # Choose ORDER BY and partition keys for MergeTree tables In Aiven for ClickHouse®, the `ORDER BY`, `PRIMARY KEY`, and `PARTITION BY` for a `MergeTree` table work together to control how data is sorted, indexed, and grouped on disk. ClickHouse uses this layout to skip irrelevant data during queries and store similar values together for better compression. Use this guidance when you [create MergeTree tables manually](/docs/products/clickhouse/howto/manage-databases-tables.md) or review table definitions before loading data. For background on how the sparse primary index uses the sort order, see [Indexing and data processing](/docs/products/clickhouse/concepts/indexing.md). warning You cannot directly change the `ORDER BY` key after table creation. To use a different key, create a table with the new key and reload the data. ## Choose columns for the ORDER BY key[​](#choose-columns-for-the-order-by-key "Direct link to Choose columns for the ORDER BY key") Choose columns based on your most common query filters. A good `ORDER BY` key usually has three to five columns. Columns that appear first in the key benefit most from data skipping because ClickHouse uses a [sparse primary index](/docs/products/clickhouse/concepts/indexing.md#clickhouse-primary-index). Use these guidelines: 1. **Start with common filter columns.** Include columns that appear often in `WHERE` clauses and remove large amounts of data from queries. 2. **Put lower-cardinality filter columns first.** Good candidates include tenant, organization, region, environment, service name, event type, status, or category. 3. **Add higher-cardinality columns later.** Add columns such as site ID, source ID, device ID, or timestamp only when they help with filtering, grouping, compression, or deduplication. 4. **Include a timestamp for time-series data.** Put the timestamp after the main filtering dimensions, for example `ORDER BY (tenant_id, event_type, event_time)`. 5. **Avoid mostly unique IDs as the first column.** Columns such as event IDs, request IDs, trace IDs, or UUIDs usually do not group related rows together. Do not choose an `ORDER BY` key only because a column is unique. In ClickHouse, the key is used for sorting and data skipping, not for enforcing uniqueness. If you often filter on columns that are not in the `ORDER BY` key, consider adding a [data skipping index](/docs/products/clickhouse/concepts/indexing.md#clickhouse-data-skipping-indexes). ## Example: Event analytics table[​](#example-event-analytics-table "Direct link to Example: Event analytics table") The following table stores event data from multiple tenants: ``` CREATE TABLE events ( tenant_id String, event_type String, event_time DateTime, event_id UUID, payload String ) ENGINE = MergeTree ORDER BY (tenant_id, event_type, event_time); ``` note In Aiven for ClickHouse®, tables created with `MergeTree` are automatically remapped to `ReplicatedMergeTree` to support high availability. Write `CREATE TABLE` statements using `MergeTree`; Aiven applies the required table engine. For details, see [Service architecture](/docs/products/clickhouse/concepts/service-architecture.md). This key works well for queries that filter by tenant, event type, and time range: ``` SELECT count() FROM events WHERE tenant_id = 'tenant-a' AND event_type = 'purchase' AND event_time >= now() - INTERVAL 7 DAY; ``` ClickHouse can use the sort order to skip data ranges that do not match the tenant, event type, or time range. The following key is less effective for this query pattern because `event_id` is almost always unique: ``` ORDER BY (event_id, tenant_id, event_time) ``` ## Separate ORDER BY and PRIMARY KEY[​](#separate-order-by-and-primary-key "Direct link to Separate ORDER BY and PRIMARY KEY") By default, the `PRIMARY KEY` is the same as `ORDER BY`. You can define a shorter `PRIMARY KEY` to keep the sparse index compact while still sorting data by additional columns using `ORDER BY`. The `PRIMARY KEY` must be a prefix of the `ORDER BY` expression. In other words, the `PRIMARY KEY` columns must appear at the start of `ORDER BY` in the same order. For example: ``` ORDER BY (service_name, log_level, timestamp, request_id) PRIMARY KEY (service_name, log_level, timestamp) ``` In this example, ClickHouse sorts data by all four columns, but builds the sparse index from the first three columns. Use a separate `PRIMARY KEY` only when later columns help with compression or deduplication but are not common query filters. For `ReplacingMergeTree` tables, put the unique row identifier at the end of `ORDER BY` and exclude it from `PRIMARY KEY` if it is not used as a filter: ``` ORDER BY (tenant_id, site_id, event_id) PRIMARY KEY (tenant_id, site_id) ``` ## Choose PARTITION BY separately[​](#choose-partition-by-separately "Direct link to Choose PARTITION BY separately") Partitioning complements the `ORDER BY` key. Use `PARTITION BY` for retention, cleanup, and data management. Do not use it as a replacement for `ORDER BY`. For time-series data, start with monthly or weekly partitions: ``` PARTITION BY toYYYYMM(event_time) ORDER BY (tenant_id, event_type, event_time) ``` Avoid high-cardinality partition keys, such as user IDs, request IDs, trace IDs, or event IDs. They can create too many partitions and make table operations slower. Aim for dozens or hundreds of partitions, not thousands. For small tables or tables without a clear retention pattern, you can omit `PARTITION BY`. ## ORDER BY patterns by use case[​](#order-by-patterns-by-use-case "Direct link to ORDER BY patterns by use case") | Use case | Recommended `ORDER BY` pattern | | --------------------------------------- | -------------------------------------------- | | Multi-tenant analytics | `(tenant_id, site_id, event_time)` | | Event analytics | `(tenant_id, event_type, event_time)` | | Application logs | `(service_name, log_level, timestamp)` | | Clickstream or session data | `(session_id, event_type, timestamp)` | | Hierarchical data | `(continent, country, city)` | | Deduplication with `ReplacingMergeTree` | `(..., unique_row_id)`; prefix `PRIMARY KEY` | For clickstream data, `session_id` can lead the key when queries filter by session rather than across all sessions. Related pages * [Manage databases and tables](/docs/products/clickhouse/howto/manage-databases-tables.md) * [Indexing and data processing](/docs/products/clickhouse/concepts/indexing.md) --- # Tiered storage in Aiven for ClickHouse® The tiered storage feature introduces a method of organizing and storing data in two tiers for improved efficiency and cost optimization. The data is automatically moved to an appropriate tier based on your database's disk usage. On top of this default data allocation mechanism, you can control the tier your data is stored in using custom data retention periods. ## Tiered storage architecture[​](#tiered-storage-architecture "Direct link to Tiered storage architecture") The tiered storage in Aiven for ClickHouse® consists of the following two layers: * Network-attached block storage (cloud provider-managed disks) - the first tier: Fast storage with limited capacity, optimized for fresh and frequently queried data, and relatively costly compared to object storage * Object storage - the second tier: Affordable storage with unlimited capacity, better suited for historical and more rarely queried data, and relatively slower. The network-attached block storage implementation depends on the cloud provider: * AWS: Amazon EBS (gp3) * Azure: Azure Managed Disks (Premium SSD v2) * GCP: Google Persistent Disk (Hyperdisk Balanced) Aiven for ClickHouse's tiered storage supports [local on-disk cache for remote files](/docs/products/clickhouse/howto/local-cache-tiered-storage.md), which is enabled by default. You can [disable the cache](/docs/products/clickhouse/howto/local-cache-tiered-storage.md#disable-the-cache) or [drop it](/docs/products/clickhouse/howto/local-cache-tiered-storage.md#free-up-space) to free up the space it occupies. ## Supported cloud platforms[​](#supported-cloud-platforms "Direct link to Supported cloud platforms") On the Aiven tenant (in non-[BYOC](/docs/platform/concepts/byoc.md) environments), Aiven for ClickHouse tiered storage is supported on the following cloud platforms: * Microsoft Azure * Amazon Web Services (AWS) * Google Cloud ## Why use it[​](#why-use-it "Direct link to Why use it") By [enabling](/docs/products/clickhouse/howto/enable-tiered-storage.md) and properly [configuring](/docs/products/clickhouse/howto/configure-tiered-storage.md) the tiered storage feature in Aiven for ClickHouse, you can use storage resources efficiently and, therefore, significantly reduce storage costs of your Aiven for ClickHouse instance. ## How it works[​](#how-it-works "Direct link to How it works") After you [enable](/docs/products/clickhouse/howto/enable-tiered-storage.md) the tiered storage feature, Aiven for ClickHouse by default stores data on network-attached block storage until it reaches 80% of its capacity. After exceeding this size-based threshold, data is stored in object storage. Optionally, you can [configure the time-based threshold](/docs/products/clickhouse/howto/configure-tiered-storage.md) for your storage. Based on the time-based threshold, the data moves from network-attached block storage to object storage after a specified time period. note Aiven backs up data that resides on network-attached block storage and in object storage. ## Access tiered storage details[​](#access-tiered-storage-details "Direct link to Access tiered storage details") When you [enable](/docs/products/clickhouse/howto/enable-tiered-storage.md) tiered storage, you can preview its details in the [Aiven Console](https://console.aiven.io/): * Click **Observe** > **Tiered storage** or * Click **Data** > **Databases and tables** > your table > **Actions** > **View details** > **Tiered storage**. ## Typical use case[​](#typical-use-case "Direct link to Typical use case") In your Aiven for ClickHouse service, a significant amount of data sits unused for a long time and is rarely accessed. That data is stored on network-attached block storage, which is relatively costly. You decide to [enable](/docs/products/clickhouse/howto/enable-tiered-storage.md) tiered storage to make your data storage more efficient and reduce the costs. For that purpose, you [enable](/docs/products/clickhouse/howto/enable-tiered-storage.md) the feature on tables to be optimized. You [configure](/docs/products/clickhouse/howto/configure-tiered-storage.md) the time-based threshold to control how your data is stored between the two layers. ## Limitations[​](#tiered-storage-limitations "Direct link to Limitations") * You can [enable](/docs/products/clickhouse/howto/enable-tiered-storage.md) tiered storage on the Aiven tenant (in non-[BYOC](/docs/platform/concepts/byoc.md) environments) if your Aiven for ClickHouse service is hosted on Azure, AWS, or GCP. * When [enabled](/docs/products/clickhouse/howto/enable-tiered-storage.md), the tiered storage feature cannot be deactivated. tip As a workaround, you can create a table (without enabling tiered storage on it) and copy the data from the original table (with the tiered storage feature [enabled](/docs/products/clickhouse/howto/enable-tiered-storage.md)) to the new table. As soon as the data is copied to the new table, you can remove the original table. * With the tiered storage feature [enabled](/docs/products/clickhouse/howto/enable-tiered-storage.md), it's not possible to connect to an external existing object storage or cloud storage bucket. * In the [Aiven Console](https://console.aiven.io/), there can be a mismatch in the displayed amount of data in object storage between what's shown in [**Observe** > **Tiered storage**](#access-tiered-storage-details) and [Storage details](#access-tiered-storage-details). This is because: * Information in [**Observe** > **Tiered storage**](#access-tiered-storage-details) is updated every hour. tip To check if you successfully transferred data to object storage, display [Storage details](#access-tiered-storage-details) of your table in the [Aiven Console](https://console.aiven.io/). * There can be unused data in object storage, for example before Aiven for ClickHouse performs a merge of parts or when a backup is performed before your table changes. Such unused data is removed once a day. ## What's next[​](#whats-next "Direct link to What's next") * [Enable tiered storage in Aiven for ClickHouse](/docs/products/clickhouse/howto/enable-tiered-storage.md) * [Configure data retention thresholds for tiered storage](/docs/products/clickhouse/howto/configure-tiered-storage.md) Related pages * [Check data volume distribution between different disks](/docs/products/clickhouse/howto/check-data-tiered-storage.md) * [Transfer data between network-attached block storage and object storage](/docs/products/clickhouse/howto/transfer-data-tiered-storage.md) * [Scale disk storage](/docs/products/clickhouse/howto/scale-disk-storage.md) --- # ClickHouse® as a columnar database ClickHouse® is a columnar databases that handles data with specific benefits. ## Fast data reading[​](#fast-data-reading "Direct link to Fast data reading") Compared to traditional row-oriented solutions, columnar database management systems store data tables by columns to provide better performance and efficiency in certain applications. As a truly columnar database, ClickHouse® also stores the values of the same column physically next to each other. This further increases the speed of retrieving the values of a column. However, it also makes it slower to retrieve complete rows, as the values of a single row are stored across different physical locations. ## Enhanced query performance[​](#enhanced-query-performance "Direct link to Enhanced query performance") Storing the data of each column independently minimizes disk access and improves query performance by reading only the data columns that are relevant to a specific query. ## Data compression and queries aggregation[​](#data-compression-and-queries-aggregation "Direct link to Data compression and queries aggregation") This storage approach also provides better options for data compression, for example, by the ability to better utilize similarities between adjacent data. Columnar databases are also better at aggregating queries involving large data sets. ## Massive and complex read operations[​](#massive-and-complex-read-operations "Direct link to Massive and complex read operations") Columnar databases such as ClickHouse are therefore best suited for analytical applications that require big data processing or data warehousing, as these usually involve fewer write operations but more - or more complex - read operations that focus on subsets of the stored data. However, applications where queries mainly affect entire rows in the data tables are less efficient in columnar databases. --- # Aiven for ClickHouse® service integrations Aiven for ClickHouse® supports different types of integration allowing you to efficiently connect with other services or data sources and access the data to be processed. There are a few ways of classifying integration types supported in Aiven for ClickHouse: * [By purpose](/docs/products/clickhouse/concepts/data-integration-overview.md#observability-integrations-vs-data-source-integrations): observability integration vs data source integration * [By location](/docs/products/clickhouse/concepts/data-integration-overview.md#integrations-between-aiven-managed-services-vs-external-integrations): integration between Aiven-managed services vs external integration (between an Aiven-managed service and an external service) * By scope: [managed databases integration](/docs/products/clickhouse/concepts/data-integration-overview.md#managed-databases-integration) and [managed credentials integration](/docs/products/clickhouse/concepts/data-integration-overview.md#managed-credentials-integration) ## Observability integrations vs data source integrations[​](#observability-integrations-vs-data-source-integrations "Direct link to Observability integrations vs data source integrations") Aiven for ClickHouse supports observability integrations and data source integrations, which have different purposes: * [Observability integration](/docs/products/clickhouse/howto/monitor-performance.md) is connecting to other services (either Aiven-managed or external ones) to expose and process logs and metrics. * Data service integration is connecting to other services (either Aiven-managed or external) to use them as data sources. In Aiven for ClickHouse, data service integration is possible with the following data source types: * Apache Kafka® * PostgreSQL® * MySQL® * ClickHouse® * Amazon S3® ## Integrations between Aiven-managed services vs external integrations[​](#integrations-between-aiven-managed-services-vs-external-integrations "Direct link to Integrations between Aiven-managed services vs external integrations") By enabling data service integrations, you create streaming data pipelines across services. Depending on where the services are located, you can have either integrations between Aiven-managed services or external integrations (between an Aiven service and an external data source or application). ## Data source integrations[​](#data-source-integrations "Direct link to Data source integrations") For integrating with external data sources, Aiven for ClickHouse provides two types of data service integrations: * [Managed databases](/docs/products/clickhouse/concepts/data-integration-overview.md#managed-databases-integration) * [Managed credentials](/docs/products/clickhouse/concepts/data-integration-overview.md#managed-credentials-integration) ### Managed credentials integration[​](#managed-credentials-integration "Direct link to Managed credentials integration") The managed credentials integration uses the [ClickHouse named collections](https://clickhouse.com/docs/en/operations/named-collections) logic. It allows integrating with a data source and storing the connection parameters used for the integration. When you use the managed credentials to query the integrated data source, you no longer need to add the connections parameters manually to each data query because the stored credentials are automatically seeded in your queries. For data to be available from Aiven for ClickHouse, you create tables using table engines. Managed credentials integration in Aiven for ClickHouse is supported with the following data source types: * PostgreSQL® * MySQL® * ClickHouse® * Amazon S3® * Azure Blob Storage important The managed credentials integration works one-way: It allows to integrate with data source and access your data from Aiven for ClickHouse, but it doesn’t allow to access Aiven for ClickHouse data from the integrated data source. For information on how table engines work in Aiven for ClickHouse services, preview [Engines: database and table](/docs/products/clickhouse/concepts/service-architecture.md#engines-database-and-table). For the list of table engines available in Aiven for ClickHouse, check [Supported table engines](/docs/products/clickhouse/reference/supported-table-engines.md). ### Managed databases integration[​](#managed-databases-integration "Direct link to Managed databases integration") The managed databases integration allows using a database engine for handling your external data. When enabled, this type of integration provides you with an automatically created database, where the remote data is exposed. Managed databases integration in Aiven for ClickHouse is supported with the following data source types: * PostgreSQL® * Apache Kafka® For information on how database engines work in Aiven for ClickHouse services, preview [Engines: database and table](/docs/products/clickhouse/concepts/service-architecture.md#engines-database-and-table). For more information on ClickHouse database engines, see [Database engines](https://clickhouse.com/docs/en/engines/database-engines). ### Supported data source types[​](#supported-data-source-types "Direct link to Supported data source types") Depending on a data source type, Aiven for ClickHouse supports different integration modes. | Data source type | Data source integration
(with Aiven service
or external source) | Managed databases integration | Managed credentials integration | | ------------------ | ------------------------------------------------------------------------- | ----------------------------- | ------------------------------- | | PostgreSQL | | | | | MySQL | | | | | Apache Kafka | | | | | ClickHouse | | | | | Amazon S3 | | | | | Azure Blob Storage | | | | ## Data flow and residency in integrations[​](#data-flow-and-residency-in-integrations "Direct link to Data flow and residency in integrations") ### How data is pulled in and distributed[​](#how-data-is-pulled-in-and-distributed "Direct link to How data is pulled in and distributed") If you integrate a multi-node Aiven for ClickHouse service with another Aiven-managed service or an external endpoint, one of the Aiven for ClickHouse service nodes connects to the other integrating entity. If the data that is read from this entity is persisted in the Aiven for ClickHouse service, this data is replicated across all the nodes in the cluster. ### Where data resides upon integration[​](#where-data-resides-upon-integration "Direct link to Where data resides upon integration") When you integrate Aiven for ClickHouse with another Aiven-managed service or an external endpoint, by default, data resides outside the Aiven for ClickHouse service in its original source. Queries aggregate data across sources, and the integrated data is accessed as needed in real-time or near-real-time. Related pages * [Set up Aiven for ClickHouse® data service integration](/docs/products/clickhouse/howto/data-service-integration.md) * [Create and manage Aiven for ClickHouse® integration databases](/docs/products/clickhouse/howto/integration-databases.md) --- # Databases, tables, and views in Aiven for ClickHouse® Databases, tables, and views organize data in Aiven for ClickHouse®. A database groups related tables and other database objects. Tables store data using table engines that define how data is stored, replicated, and queried. Views and materialized views help query, transform, or precompute data. You can create and manage databases and tables in the [Aiven Console](https://console.aiven.io/) or with SQL. In addition to regular databases, you can create [integration databases](/docs/products/clickhouse/howto/integration-databases.md#create-integration-databases) to connect with other Aiven services. ## Databases[​](#databases "Direct link to Databases") A database groups related tables, views, and other objects. Aiven for ClickHouse supports multiple database engines depending on the service configuration and use case. ## Tables[​](#tables "Direct link to Tables") A table stores data in rows and columns. Each table uses a table engine that controls storage behavior, indexing, replication, and data lifecycle features. ## Views and materialized views[​](#views-and-materialized-views "Direct link to Views and materialized views") Views and materialized views help query or transform data. A materialized view stores query results and can improve performance for repeated queries or pre-aggregated data. ## Engines[​](#engines "Direct link to Engines") How data is stored and queried depends on the engines you choose: * [Table engines](/docs/products/clickhouse/reference/supported-table-engines.md) define how data is stored on disk or read from external sources. * [Database engines](/docs/products/clickhouse/reference/supported-database-engines.md) control how a database manages its tables and metadata. Engine availability can vary by Aiven for ClickHouse service version. Related pages * [Manage databases and tables](/docs/products/clickhouse/howto/manage-databases-tables.md) * [Supported table engines](/docs/products/clickhouse/reference/supported-table-engines.md) * [Supported database engines](/docs/products/clickhouse/reference/supported-database-engines.md) * [Materialized views](/docs/products/clickhouse/howto/materialized-views.md) --- # Disaster recovery in Aiven for ClickHouse® Aiven for ClickHouse® prevents and mitigates emergencies or crises with multiple disaster recovery methods to keep your data safe and sound. Disaster recovery is a process of coping with emergencies or crises using dedicated methods for protecting resources and/or reestablishing their desired status. In the context of data infrastructure, well-established disaster recovery methods are of a particular importance for preventing data loss or corruption. Software failure, loss of an availability zone, or datacenter outage are only a few examples of emergencies when disaster recovery comes in. ## High availability[​](#high-availability "Direct link to High availability") High availability (HA) is an entity's ability to continuously maintain a certain level of operational performance for a desired period of time. HA is typically achieved by redundancy - securing replicas of databases or services to be highly available. To support disaster recovery technologies, a database service needs to stay highly available, for example, by operating on a few nodes holding the same data. With Aiven, HA for your service is supported in business and premium plans. See [Plan comparison](https://aiven.io/pricing?tab=plan-comparison\&product=clickhouse) for details. Also see [cross-availability-zone data distribution](/docs/platform/concepts/availability-zones.md#cross-zone-data-distro) ## Backup and restore[​](#backup-and-restore "Direct link to Backup and restore") ### Service backup[​](#service-backup "Direct link to Service backup") Backups of Aiven for ClickHouse services happen automatically on a daily basis. An automatic backup is also taken before powering off a service. Backups cover the following: * Access entities (for example, users, roles, passwords, or secrets) stored in Zookeeper * Database definitions * Table schemas * Table content (`part files`) * Dictionaries You can [restore your service from a selected backup](/docs/products/clickhouse/howto/restore-backup.md). Part files With the ClickHouse's ReplicatedMergeTree table engine, each INSERT query results in creating a new file, so-called part, written only once and not modifiable. Using part files allows incremental backups in Aiven for ClickHouse: only changed parts are backed up and files already available in the object storage are left out from the backup. ### Service recovery[​](#service-recovery "Direct link to Service recovery") Regardless of whether your Aiven for ClickHouse service is powered on or powered off, you can create its copy and [restore the data from a selected service backup](/docs/products/clickhouse/howto/restore-backup.md). For this purpose, you create a fork from the original service. This spins up a new service that hosts the data recovered from the selected backup. ## Sharding[​](#sharding "Direct link to Sharding") Essentially, sharding is a technique of splitting database rows across multiple database nodes, which usually significantly increases performance. However, integrating sharding with database replication technologies, data can be replicated across shards of the sharded database. Replication at the shard level provides high availability and helps to achieve disaster recovery. A shard group can be replicated to one or more data centers, which improves the disaster recovery capability. With Aiven for ClickHouse [business and premium plans](https://aiven.io/pricing?tab=plan-comparison\&product=clickhouse), each shard is replicated across three availability zones. The service and the data stay fully available even if an entire availability zone is lost. note Although sharding with replicated nodes can reduce failures, it still cannot save a service from the loss of an entire region. For information on how to work with shards in Aiven for ClickHouse, see [Enable reading and writing data across shards](/docs/products/clickhouse/howto/use-shards-with-distributed-table.md). ## Limitations[​](#limitations "Direct link to Limitations") Aiven for ClickHouse has a few restrictions on the disaster recovery capability. * No backup to another region * No point in time recovery (PITR) For all the restrictions and limits for Aiven for ClickHouse, see [Aiven for ClickHouse limits and limitations](/docs/products/clickhouse/reference/limitations.md). Related pages * [Disaster Recovery testing scenarios](/docs/platform/concepts/disaster-recovery-test-scenarios.md) * [Configure Aiven for ClickHouse® backup settings](/docs/products/clickhouse/howto/configure-backup.md) * [Fork your Aiven for ClickHouse® service](/docs/products/clickhouse/howto/fork-service.md) * [Restore an Aiven for ClickHouse® backup](/docs/products/clickhouse/howto/restore-backup.md) --- # Querying external data in Aiven for ClickHouse® Discover federated queries and their capabilities in Aiven for ClickHouse® and how they simplify and speed up migrating into Aiven from external data sources. Federated queries allow communication between Aiven for ClickHouse and S3-compatible object storages and web resources. The federated queries feature in Aiven for ClickHouse enables you to read and pull data from an external object storage that uses the S3 integration engine, or any web resource accessible over HTTP. note The federated queries feature in Aiven for ClickHouse is enabled by default. ## Why use federated queries[​](#why-use-federated-queries "Direct link to Why use federated queries") There are a few reasons why you might want to use federated queries: * Query remote data from your ClickHouse service. Ingest it into Aiven for ClickHouse or only reference external data sources as part of an analytics query. In the context of an increasing footprint of connected data sources, federated queries can help you better understand how your customers use your products. * Simplify and speed up the import of your data into the Aiven for ClickHouse instance from a legacy data source, avoiding a long and sometimes complex migration path. * Improve the migration of data in Aiven for ClickHouse, and extend analysis over external data sources with a relatively low effort in comparison to enabling distributed tables and [the remote and remoteSecure functionalities](https://clickhouse.com/docs/en/sql-reference/table-functions/remote). note The `remote()` and `remoteSecure()` features are designed to read from remote data sources or provide the ability to create a distributed table across remote data sources but they are not designed to read from an external S3 storage. ## How it works[​](#how-it-works "Direct link to How it works") To run a federated query, the ClickHouse service user connecting to the cluster requires grants to the S3 and/or URL sources. The main service user is granted access to the sources by default, and new users can be allowed to use the sources via the CREATE TEMPORARY TABLE grant, which is required for both sources. For more information on how to enable new users to use the sources, see [Prerequisites](/docs/products/clickhouse/howto/run-federated-queries.md#prerequisites). Federated queries read from external S3-compatible object storage utilizing the ClickHouse S3 engine. Once you read from a remote S3-compatible storage, you can select from that storage and insert into a table in the Aiven local instance, enabling migration of data into Aiven. For more details on how to run federated queries in Aiven for ClickHouse, see [Read and pull data from S3 object storages and web resources over HTTP](/docs/products/clickhouse/howto/run-federated-queries.md). ## Limitations[​](#limitations "Direct link to Limitations") * Federated queries in Aiven for ClickHouse only support S3-compatible object storage providers for the time being. More external data sources coming soon. * Virtual tables are only supported for URL sources, using the URL table engine. Stay tuned for us supporting the S3 table engine in the future. Related pages * [Read and pull data from S3 object storages and web resources over HTTP](/docs/products/clickhouse/howto/run-federated-queries.md) * [Integrating S3 | ClickHouse Docs](https://clickhouse.com/docs/en/integrations/s3) * [remote, remoteSecure | ClickHouse Docs](https://clickhouse.com/docs/en/sql-reference/table-functions/remote) * [Cloud Compatibility | ClickHouse Docs](https://clickhouse.com/docs/en/whats-new/cloud-compatibility#federated-queries) --- # Indexing and data processing in ClickHouse® ClickHouse® processes data differently from other database management systems. ClickHouse uses sparse and skipping indexes and a vector computation engine. ## ClickHouse as a columnar database[​](#clickhouse-as-a-columnar-database "Direct link to ClickHouse as a columnar database") ClickHouse is a columnar database, which means that ClickHouse can read columns selectively, retrieving only necessary information and omitting columns that are not needed for the request. The smaller the number of the columns you read, the faster and more efficient the performance of the request. If you have to read many or all columns, using a columnar database becomes a less effective approach. Read more about characteristics of columnar databases and their features in [Columnar databases](/docs/products/clickhouse/concepts/columnar-databases.md). ## Reading data in blocks[​](#reading-data-in-blocks "Direct link to Reading data in blocks") ClickHouse is designed to process substantial chunks of information for a single request. To make this possible, it shifts the focus from reading individual lines into scanning massive blocks of data for a request. This happens by using what is called the vector computation engine, which reads data in blocks that normally consist of thousands of rows selected for a small set of columns. ## ClickHouse primary index[​](#clickhouse-primary-index "Direct link to ClickHouse primary index") ClickHouse reads data in blocks. To find a specific block, ClickHouse uses its primary index, which has a few distinctive features that make it different from the primary indexes of the other systems. Instead of indexing every row, ClickHouse indexes every 10000th row (or, to be precise, every 8192nd row when using default settings). Such an index type is descriptively called *sparse index*, and the batch of rows is sometimes called *granule*. To make sparse indexing possible, ClickHouse sorts items according to the primary key. The items are intentionally sorted physically on the disk to speed up reading and prevent jumping across the disk when processing data. Therefore, by selecting a primary key, you determine how items are sorted physically on the disk. Such an approach helps ClickHouse effectively work on regular hard drives and depend less on network-attached block storage in comparison to other DBMSs. note Even though ClickHouse sorts data by primary key, it is possible [to choose a primary key that is different from the sorting key](https://clickhouse.com/docs/en/engines/table-engines/mergetree-family/mergetree/#choosing-a-primary-key-that-differs-from-the-sorting-key). Using sparse indexing has significant consequences for capabilities and limitations of ClickHouse. A primary key, as used in ClickHouse, does not ensure uniqueness for a single searched item since only every ten thousandth item is indexed. You need to iterate over thousands of items to find a specific row, which makes this approach inadequate when working with individual rows and suitable for processing millions or trillions of items. Example When analysing error rates based on a server log analysis, you don't focus on individual lines but look at the overall picture to see trends. Such requests allow approximate calculations using only a sample of data to draw conclusions. A good primary index should help limit the number of items you need to read to process a query. tip Learn more about ClickHouse [primary indexes in the official documentation of ClickHouse](https://clickhouse.com/docs/en/engines/table-engines/mergetree-family/mergetree/#choosing-a-primary-key-that-differs-from-the-sorting-key). ## ClickHouse data skipping indexes[​](#clickhouse-data-skipping-indexes "Direct link to ClickHouse data skipping indexes") Although skipping indexes are used in ClickHouse as secondary indexes, they work quite differently to secondary indexes used in other DBMSs. Skipping indexes help boost performance by skipping some irrelevant rows in advance, when it can be predicted that these rows do not satisfy query conditions. Example You have numeric column *number of page visits* and run a query to select all rows where page visits are over 10000. To speed up such a query, you can add a skipping index to store extremes of the field and help ClickHouse to skip in advance values that do not satisfy the request condition. Related pages Learn more about [skipping indexes in the official ClickHouse documentation](https://clickhouse.com/docs/en/engines/table-engines/mergetree-family/mergetree/#table_engine-mergetree-data_skipping-indexes). --- # Online analytical processing Online analytical processing (OLAP) is an approach to producing real-time reports and insights, usually based on large amounts of source data. A major function of OLAP tools is to provide data-based intelligence to inform and support business decisions. The target data usually includes operational activities that are then analyzed for use by various functions, such as sales, marketing, and finance. From a technical perspective, most traditional, row-oriented database systems are better suited to online transactional processing (OLTP), but are not very efficient for OLAP scenarios. Column-oriented databases are more efficient, as they provide quicker access to subsets of the stored data. However, most database systems aim for a hybrid approach that primarily focuses on optimizing either OLAP or OLTP scenarios, while offering solutions to support the alternative scenario. --- # Query Kafka topic data in Aiven for ClickHouse® Query Kafka topic data in Aiven for ClickHouse® by connecting an Aiven for Apache Kafka® topic to a ClickHouse table. Use this managed setup to store Kafka topic data in ClickHouse and analyze live or historical event data with SQL, without manually creating a Kafka engine table, materialized view, or destination ClickHouse table. To set up the integration in the Aiven Console, see [Set up Kafka topic querying in Aiven for ClickHouse®](/docs/products/clickhouse/howto/set-up-kafka-topic-querying.md). ## When to use it[​](#when-to-use-it "Direct link to When to use it") Query Kafka topic data when you want to: * Analyze live or historical Kafka topic data using SQL. * Build dashboards or reports from Kafka topic data stored in ClickHouse. * Reduce the manual setup required to ingest Kafka data into ClickHouse. ## How it works[​](#how-it-works "Direct link to How it works") When you connect a Kafka topic to ClickHouse, Aiven creates a managed integration between the two services. The setup uses the selected Kafka topic as the source and writes the data to a destination ClickHouse table that uses the MergeTree engine. You can start the setup from a Kafka topic in the Aiven Console. During setup, select an existing ClickHouse service or create one, then review and edit the schema and table configuration before deployment. ### Schema mapping[​](#schema-mapping "Direct link to Schema mapping") Automated schema mapping is available for Avro messages with Aiven for Apache Kafka® Schema Registry. The schema must be registered using the topic name strategy, such as `-value`. When a compatible Avro schema is available, Aiven maps the Avro fields to ClickHouse column types and shows them in the schema preview. Before deployment, you can review the mapped columns. If Aiven cannot find a compatible Avro schema, you can define the ClickHouse table columns during setup. ### Ingestion start point[​](#ingestion-start-point "Direct link to Ingestion start point") During setup, choose whether ingestion starts from new messages only or from an earlier point in the topic, when available. If you start from new messages only, Aiven ingests messages produced after you deploy the setup. If you start from an earlier point, Aiven includes existing topic data based on the start point you select. ### Managed ClickHouse resources[​](#managed-clickhouse-resources "Direct link to Managed ClickHouse resources") Kafka topic data ingestion uses the ClickHouse Kafka table engine. Aiven creates and manages the Kafka engine table, materialized view, and destination ClickHouse table. This reduces the manual setup required to store Kafka topic data in ClickHouse. You query the destination table in ClickHouse and do not need to create the ingestion pipeline manually. ## Query data in ClickHouse[​](#query-data-in-clickhouse "Direct link to Query data in ClickHouse") After the setup is deployed and data starts flowing, query the destination table using ClickHouse SQL. You can use the table for real-time analytics, operational dashboards, event analysis, historical reporting, and high-volume log or metrics analysis. ## Limitations[​](#limitations "Direct link to Limitations") * The Kafka service and the ClickHouse service must be in the same cloud region to reduce network latency and data transfer costs. * If data does not appear in the destination table after deployment, review **Observe** > **Logs** for the ClickHouse service for ingestion errors. * Failed messages are not sent to a dead letter queue by default. To change this, set `handle_error_mode` to `dead_letter_queue` in the Kafka engine advanced configuration. In the Aiven Console, this setting is available under **Data** > **Databases and tables**. * You cannot change some table settings, such as the sorting key, after table creation. * Automated schema mapping is available only for Avro messages with Aiven for Apache Kafka® Schema Registry when the schema is registered using the topic name strategy, such as `-value`. Without a compatible Avro schema, define ClickHouse columns during setup. * Automatic schema evolution is not supported. If the Kafka topic schema changes, the ClickHouse resources might need to be updated manually or the integration might need to be recreated. New Avro fields are ignored unless matching columns are added to the ClickHouse table. Removed or incompatible fields can cause ingestion errors. * The integration is not available for free and developer tier Aiven for Apache Kafka® services. Related pages * [Set up Kafka topic querying in Aiven for ClickHouse®](/docs/products/clickhouse/howto/set-up-kafka-topic-querying.md) * [Connect Apache Kafka® to Aiven for ClickHouse®](/docs/products/clickhouse/howto/integrate-kafka.md) --- # Aiven for ClickHouse® service architecture Aiven for ClickHouse® is implemented as a multi-master cluster where data replication is managed by ClickHouse, schema and user replication is managed by ZooKeeper, and backup and restore are managed by Aiven. Understand the technical design of Aiven for ClickHouse. ## Deployment modes[​](#deployment-modes "Direct link to Deployment modes") Aiven for ClickHouse can be deployed either as a single node, a single shard of three nodes, or multiple shards of three nodes each. * With a single shard, all data is available on all nodes and all servers can be used for reads and writes (no main server, leader, or replica). * With multiple shards, the data is split between all shards and the data of each shard is present in all nodes of the shard. Each Aiven for ClickHouse service is exposed as a single server URL pointing to all servers with connections going randomly to any of the servers. ClickHouse is responsible for replicating the writes between the servers. For synchronizing critical low-volume information between servers, Aiven for ClickHouse relies on [ZooKeeper](#zookeeper), which runs on each ClickHouse server. ## Coordinating services[​](#coordinating-services "Direct link to Coordinating services") Each Aiven for ClickHouse node runs ClickHouse and ZooKeeper. ### ZooKeeper[​](#zookeeper "Direct link to ZooKeeper") ZooKeeper handles cross-node coordination and synchronization for these replication processes: * Replication of database changes across the cluster: CREATE, UPDATE, or ALTER TABLE (by ClickHouse's `Replicated` [database engine](/docs/products/clickhouse/concepts/service-architecture.md#replicated-database-engine)) * Replication of table data across the cluster (by ClickHouse's `ReplicatedMergeTree` [table engine](/docs/products/clickhouse/concepts/service-architecture.md#replicated-table-engine)). Data itself is not written to ZooKeeper but transferred directly between ClickHouse servers. * Replication of the storage of Users, Roles, Quotas, Row Policies for the whole cluster (by ClickHouse). note Storing entities such as Users, Roles, Quotas, and Row Policies in ZooKeeper ensures that Role Based Access Control (RBAC) is applied consistently over the entire cluster. This type of entity storage was developed at Aiven and is now part of the upstream ClickHouse. ZooKeeper handles one process per node and is accessible only from within the cluster. ### Backup and restore[​](#backup-and-restore "Direct link to Backup and restore") [Backup and restore](/docs/products/clickhouse/concepts/disaster-recovery.md#service-backup) for the service is implemented and operated by Aiven. Backups are coordinated across the cluster so that restores are applied safely and consistently. ## Data architecture[​](#data-architecture "Direct link to Data architecture") Aiven for ClickHouse enforces * Full schema replication: all databases, tables, users, and grants are the same on all nodes. * Full data replication: table rows are the same on all nodes within a shard. For more information about backups and disaster recovery, see [Service backup](/docs/products/clickhouse/concepts/disaster-recovery.md#service-backup). ## Engines: database and table[​](#engines-database-and-table "Direct link to Engines: database and table") ClickHouse has engines in two flavors: Table engines and database engines. * The database engine manipulates tables and decides what happens when you list, create, or delete a table. It can also restrict a database to specific table engines or manage replication. * The table engine decides how data is stored on disk or how data is read from outside the disk and exposed as a virtual table. ### `Replicated` database engine[​](#replicated-database-engine "Direct link to replicated-database-engine") The default ClickHouse database engine is the Atomic engine, responsible for creating table metadata on the disk and configuring which table engines are allowed in each database. Aiven for ClickHouse uses the `Replicated` database engine, which is a variant of Atomic. With this engine variant, queries for creating, updating, or altering tables are replicated to all other servers using [ZooKeeper](#zookeeper). As a result, all servers can have the same table schema, which makes them an actual data cluster and not multiple independent servers that can talk to each other. ### `Replicated` table engine[​](#replicated-table-engine "Direct link to replicated-table-engine") The table engine handles `INSERT` and `SELECT` queries. From a wide variety of available table engines, the most common ones belong to the `MergeTree` engine family, which Aiven for ClickHouse supports. For a list of all the table engines that you can use in Aiven for ClickHouse, see [Supported table engines in Aiven for ClickHouse](/docs/products/clickhouse/reference/supported-table-engines.md). #### `MergeTree` engine[​](#mergetree-engine "Direct link to mergetree-engine") With the `MergeTree` engine, at least one new file is created for each INSERT query and each new file is written once and never modified. In the background, new files (called *parts*) are re-read, merged, and rewritten into compact form. Writing data in parts determines the performance profile of ClickHouse. * Batch INSERT queries to avoid creating many small parts. * Batch UPDATE and DELETE queries. Removing or updating a single row requires rewriting an entire part with all rows except the one being removed or updated. * SELECT queries are executed rapidly because all the data found in a part is valid and all files can be cached since they never change. #### `ReplicatedMergeTree` engine[​](#replicatedmergetree-engine "Direct link to replicatedmergetree-engine") Each engine of the `MergeTree` family has a matching `ReplicatedMergeTree` engine, which additionally enables the replication of all writes using [ZooKeeper](#zookeeper). The data itself doesn't travel through ZooKeeper and is actually fetched from one ClickHouse server to another. A shared log of update queries is maintained with ZooKeeper. All nodes add entries to the queue and watch for changes to execute the queries. When a query to create a table using the `MergeTree` engine arrives, Aiven for ClickHouse automatically rewrites the query to use the `ReplicatedMergeTree` engine so that all tables are replicated and all servers have the same table data, which makes the deployment a high-availability cluster. --- # Service management in Aiven for ClickHouse® Manage the security, configuration, and lifecycle of your Aiven for ClickHouse® service from the [Aiven Console](https://console.aiven.io/) or the [Aiven API](/docs/tools/api.md). Use this section to manage service operations, control user access, upgrade ClickHouse versions, and review service limits and advanced configuration options. ## Security[​](#security "Direct link to Security") Protect your service by restricting network access, connecting through a [Virtual Private Cloud (VPC)](/docs/platform/concepts/cloud-security.md#networking-with-vpc-peering), and enabling termination protection. See [Secure a managed ClickHouse service](/docs/products/clickhouse/howto/secure-service.md) for details. ## Service operations[​](#service-operations "Direct link to Service operations") Power on or off, rename, tag, fork, or move your service to another cloud or region. See [Power on/off and delete](/docs/products/clickhouse/howto/power-cycle-service.md), [Rename a service](/docs/products/clickhouse/howto/rename-service.md), [Tag a service](/docs/products/clickhouse/howto/tag-service.md), [Fork a service](/docs/products/clickhouse/howto/fork-service.md), and [Change cloud or region](/docs/products/clickhouse/howto/change-cloud-region.md). ## Users and access[​](#users-and-access "Direct link to Users and access") Create users and roles and grant privileges to control who can access databases, tables, and administrative functions. See [Manage users and roles](/docs/products/clickhouse/howto/manage-users-roles.md). ## Versions and configuration[​](#versions-and-configuration "Direct link to Versions and configuration") Choose a ClickHouse version when you create a service and upgrade to a newer supported version later. Review [advanced parameters](/docs/products/clickhouse/reference/advanced-params.md) for available configuration options. Manage [maintenance updates](/docs/products/clickhouse/howto/maintenance-updates.md) and the maintenance window for the service. Related pages * [Secure a managed ClickHouse service](/docs/products/clickhouse/howto/secure-service.md) * [Power on/off and delete](/docs/products/clickhouse/howto/power-cycle-service.md) * [Manage users and roles](/docs/products/clickhouse/howto/manage-users-roles.md) * [Manage versions](/docs/products/clickhouse/howto/manage-clickhouse-versions.md) * [Maintenance and updates](/docs/products/clickhouse/howto/maintenance-updates.md) * [Advanced parameters](/docs/products/clickhouse/reference/advanced-params.md) * [Limits and limitations](/docs/products/clickhouse/reference/limitations.md) * [Scale disk storage](/docs/products/clickhouse/howto/scale-disk-storage.md) * [Change the service plan](/docs/products/clickhouse/howto/change-service-plan.md) --- # String data type in Aiven for ClickHouse® Aiven for ClickHouse® uses ClickHouse® databases, which can store diverse types of data, such as strings, decimals, booleans, or arrays. ## About strings in ClickHouse[​](#about-strings-in-clickhouse "Direct link to About strings in ClickHouse") ClickHouse allows strings of any length. Strings can contain an arbitrary amount of bytes, which are stored and output as-is. The string type replaces the types VARCHAR, BLOB, CLOB, and others from other database management systems (DBMS). When creating tables, numeric parameters for string fields can be set (for example, TEXT(140)) but are ignored. ClickHouse supports the following aliases for strings: LONGTEXT, MEDIUMTEXT, TINYTEXT, TEXT, LONGBLOB, MEDIUMBLOB, TINYBLOB, BLOB, VARCHAR, CHAR. ## String-handling functions[​](#string-handling-functions "Direct link to String-handling functions") * [Functions for working with strings](https://clickhouse.com/docs/en/sql-reference/functions/string-functions/) * [Functions for searching in strings](https://clickhouse.com/docs/en/sql-reference/functions/string-search-functions) note By default, the search is case-sensitive in these functions, but case-insensitive search variants are also available. * [Functions for searching and replacing in strings](https://clickhouse.com/docs/en/sql-reference/functions/string-replace-functions) * [Functions for splitting and merging strings and arrays](https://clickhouse.com/docs/en/sql-reference/functions/splitting-merging-functions) ## String conversions[​](#string-conversions "Direct link to String conversions") Any plain string type can be cast to a different type using functions in [Type Conversion Functions](https://clickhouse.com/docs/en/sql-reference/functions/type-conversion-functions). ## Strings and JSON[​](#strings-and-json "Direct link to Strings and JSON") ClickHouse supports a wide range of functions for working with JSON. With specific functions, you can use strings for extracting JSON. Learn more on [JSON functions in ClickHouse](https://clickhouse.com/docs/en/sql-reference/functions/json-functions/). Examples * `visitParamExtractString(params, name)`: Parse the string in double quotes. * `JSONExtractString(json[, indices_or_keys]…)`: Parse a JSON and extract a string. * `toJSONString`: Convert a value of any data type to its JSON representation. --- # Get started with Aiven for ClickHouse® Start using Aiven for ClickHouse® by creating and configuring a service, connecting to it, and loading sample data. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Console * Terraform * Kubernetes * ClickHouse client - Access to the [Aiven Console](https://console.aiven.io) - [Docker](https://docs.docker.com/desktop/) installed * [Terraform installed](https://www.terraform.io/downloads) * A [personal token](https://docs.aiven.io/docs/platform/howto/create_authentication_token.html) * [Docker](https://docs.docker.com/desktop/) installed - [Aiven Operator for Kubernetes®](https://aiven.github.io/aiven-operator/installation/helm.html) installed - A [personal token](https://docs.aiven.io/docs/platform/howto/create_authentication_token.html) - A [Kubernetes secret](https://aiven.github.io/aiven-operator/authentication.html) storing your token - [Docker](https://docs.docker.com/desktop/) installed * [ClickHouse CLI client](https://clickhouse.com/docs/en/install) installed * [Docker](https://docs.docker.com/desktop/) installed ## Create an Aiven for ClickHouse® service[​](#create-an-aiven-for-clickhouse-service "Direct link to Create an Aiven for ClickHouse® service") * Console * Terraform * Kubernetes 1. In your project, click **Services**. 2. Click **Create service**. 3. Select **ClickHouse®**. 4. Select a **Cloud**. 5. Select a **Plan**. note The plans available can vary between cloud providers and regions for the same service. 6. In the **Service details**, enter a name for your service. 7. Optional: Add service tags. 8. In the **Service summary**, click **Create service**. The status of the service is **Rebuilding** during its creation. When the status is **Running**, you can start using the service. This typically takes a couple of minutes and can vary between cloud providers and regions. In this example, an Aiven for ClickHouse service is used to store IoT sensor data. You create the service, two service users, and assign each user a role: * Give the ETL user permission to insert data. * Give the analyst user access to view data in the measurements database. The following example files are also available in the Aiven Terraform Provider [repository](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/clickhouse) on GitHub. 1. Create a file named `provider.tf` and add the following: ``` Loading... ``` 2. Create a file named `service.tf` and add the following: ``` Loading... ``` 3. Create a file named `service_users.tf` and add the following: ``` Loading... ``` 4. Create a file named `variables.tf` and add the following: ``` Loading... ``` 5. Create the `terraform.tfvars` file and add the values for your token and project name. 6. To output connection details, create a file named `output.tf` and add the following: ``` Loading... ``` To apply your Terraform configuration: 1. Initialize Terraform by running: ``` terraform init ``` The output is similar to the following: ``` Initializing the backend... Initializing provider plugins... - Finding aiven/aiven versions matching ">= 4.0.0, < 5.0.0"... - Installing aiven/aiven v4.9.2... - Installed aiven/aiven v4.9.2 ... Terraform has been successfully initialized! ... ``` 2. To create an execution plan and preview the changes, run: ``` terraform plan ``` 3. To deploy your changes, run: ``` terraform apply --auto-approve ``` Create an Aiven for ClickHouse service using the Aiven Operator for Kubernetes. 1. [Get authenticated and authorized](https://aiven.github.io/aiven-operator/authentication.html). 2. Create file `example.yaml` with the following content: ``` apiVersion: aiven.io/v1alpha1 kind: Clickhouse metadata: name: my-clickhouse spec: authSecretRef: name: aiven-token key: token connInfoSecretTarget: name: my-clickhouse-connection userConfig: service_log: false project: my-aiven-project cloudName: google-europe-west1 plan: startup-16 maintenanceWindowDow: friday maintenanceWindowTime: 23:00:00 ``` 3. Create the service by applying the configuration: ``` kubectl apply -f example.yaml ``` 4. Review the resource you created with the following command: ``` kubectl get clickhouses my-clickhouse ``` The output is similar to the following: ``` Name Project Region Plan State my-clickhouse my-aiven-project google-europe-west1 startup-16 RUNNING ``` The resource might stay in the `BUILDING` state for a couple of minutes. When the state changes to `RUNNING`, you are ready to access it. ## Configure the service[​](#configure-the-service "Direct link to Configure the service") You can change your service settings by updating the service configuration. * Console * Terraform * Kubernetes 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose your new Aiven for ClickHouse service. 2. In the service sidebar, click **Service settings**. 3. In the **Advanced configuration** section, make changes to the service configuration. See the available configuration options in [Advanced parameters for Aiven for ClickHouse®](/docs/products/clickhouse/reference/advanced-params.md). See [the `aiven_clickhouse` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/clickhouse) for the full schema. 1. Update file `example.yaml`: * Add `service_log: true` and `terminationProtection: true`. * Update `maintenanceWindowDow: sunday` and `maintenanceWindowTime: 22:00:00`. ``` apiVersion: aiven.io/v1alpha1 kind: Clickhouse metadata: name: my-clickhouse spec: authSecretRef: name: aiven-token key: token connInfoSecretTarget: name: my-clickhouse-connection userConfig: service_log: true project: my-aiven-project cloudName: google-europe-west1 plan: startup-16 maintenanceWindowDow: sunday maintenanceWindowTime: 22:00:00 terminationProtection: true ``` 2. Update the service by applying the configuration: ``` kubectl apply -f example.yaml ``` 3. Review the resource you updated with the following command: ``` kubectl get clickhouses my-clickhouse ``` The resource might stay in the `BUILDING` state for a couple of minutes. When the state changes to `RUNNING`, you are ready to access it. See the available configuration options in [Aiven Operator for Kubernetes®: ClickHouse](https://aiven.github.io/aiven-operator/resources/clickhouse.html) ## Connect to the service[​](#connect-to-service "Direct link to Connect to the service") * Console * Terraform * ClickHouse client 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose your Aiven for ClickHouse service. 2. On the **Overview** page, click **Quick connect**. 3. In the **Connect** window, click a tool or language to connect to your service, follow the connection instructions, and click **Done**. ``` docker run -it \ --rm clickhouse/clickhouse-server clickhouse-client \ --user avnadmin \ --password admin_password \ --host clickhouse-service-name-project-name.e.aivencloud.com \ --port 12691 \ --secure ``` [Access your new service](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md) with the [ClickHouse client](https://clickhouse.com/docs/en/integrations/sql-clients/cli) using the Terraform outputs. 1. To store the outputs in environment variables, run: ``` CLICKHOUSE_HOST="$(terraform output -raw clickhouse_service_host)" CLICKHOUSE_PORT="$(terraform output -raw clickhouse_service_port)" CLICKHOUSE_USER="$(terraform output -raw clickhouse_service_username)" CLICKHOUSE_PASSWORD="$(terraform output -raw clickhouse_service_password)" ``` 2. To use the environment variables to connect to the service, run: ``` docker run -it \ --rm clickhouse/clickhouse-client \ --user=$CLICKHOUSE_USER \ --password=$CLICKHOUSE_PASSWORD \ --host=$CLICKHOUSE_HOST \ --port=$CLICKHOUSE_PORT \ --secure ``` [Connect to your new service with CLI](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md) using the [ClickHouse client](https://clickhouse.com/docs/en/integrations/sql-clients/cli). tip Discover more tools for connecting to Aiven for ClickHouse in [Connect to Aiven for ClickHouse®](/docs/products/clickhouse/howto/list-connect-to-service.md). ## Load a dataset[​](#load-a-dataset "Direct link to Load a dataset") 1. Download a dataset from [Example Datasets](https://clickhouse.com/docs/en/getting-started/example-datasets/metrica/) using cURL: ``` curl https://datasets.clickhouse.com/hits/tsv/hits_v1.tsv.xz \ | unxz --threads="$(nproc)" > hits_v1.tsv curl https://datasets.clickhouse.com/visits/tsv/visits_v1.tsv.xz \ | unxz --threads="$(nproc)" > visits_v1.tsv ``` note The `nproc` Linux command, which prints the number of processing units, is not available on macOS. To use the command, add an alias for `nproc` into your `~/.zshrc` file: `alias nproc="sysctl -n hw.logicalcpu"`. When the download is complete, you have two files: `hits_v1.tsv` and `visits_v1.tsv`. 2. Create tables `hits_v1` and `visits_v1` in the `default` database, which has been created automatically when you create your Aiven for ClickHouse service. Expand for the `CREATE TABLE default.hits_v1` sample ``` CREATE TABLE default.hits_v1 ( WatchID UInt64, JavaEnable UInt8, Title String, GoodEvent Int16, EventTime DateTime, EventDate Date, CounterID UInt32, ClientIP UInt32, ClientIP6 FixedString(16), RegionID UInt32, UserID UInt64, CounterClass Int8, OS UInt8, UserAgent UInt8, URL String, Referer String, URLDomain String, RefererDomain String, Refresh UInt8, IsRobot UInt8, RefererCategories Array(UInt16), URLCategories Array(UInt16), URLRegions Array(UInt32), RefererRegions Array(UInt32), ResolutionWidth UInt16, ResolutionHeight UInt16, ResolutionDepth UInt8, FlashMajor UInt8, FlashMinor UInt8, FlashMinor2 String, NetMajor UInt8, NetMinor UInt8, UserAgentMajor UInt16, UserAgentMinor FixedString(2), CookieEnable UInt8, JavascriptEnable UInt8, IsMobile UInt8, MobilePhone UInt8, MobilePhoneModel String, Params String, IPNetworkID UInt32, TraficSourceID Int8, SearchEngineID UInt16, SearchPhrase String, AdvEngineID UInt8, IsArtifical UInt8, WindowClientWidth UInt16, WindowClientHeight UInt16, ClientTimeZone Int16, ClientEventTime DateTime, SilverlightVersion1 UInt8, SilverlightVersion2 UInt8, SilverlightVersion3 UInt32, SilverlightVersion4 UInt16, PageCharset String, CodeVersion UInt32, IsLink UInt8, IsDownload UInt8, IsNotBounce UInt8, FUniqID UInt64, HID UInt32, IsOldCounter UInt8, IsEvent UInt8, IsParameter UInt8, DontCountHits UInt8, WithHash UInt8, HitColor FixedString(1), UTCEventTime DateTime, Age UInt8, Sex UInt8, Income UInt8, Interests UInt16, Robotness UInt8, GeneralInterests Array(UInt16), RemoteIP UInt32, RemoteIP6 FixedString(16), WindowName Int32, OpenerName Int32, HistoryLength Int16, BrowserLanguage FixedString(2), BrowserCountry FixedString(2), SocialNetwork String, SocialAction String, HTTPError UInt16, SendTiming Int32, DNSTiming Int32, ConnectTiming Int32, ResponseStartTiming Int32, ResponseEndTiming Int32, FetchTiming Int32, RedirectTiming Int32, DOMInteractiveTiming Int32, DOMContentLoadedTiming Int32, DOMCompleteTiming Int32, LoadEventStartTiming Int32, LoadEventEndTiming Int32, NSToDOMContentLoadedTiming Int32, FirstPaintTiming Int32, RedirectCount Int8, SocialSourceNetworkID UInt8, SocialSourcePage String, ParamPrice Int64, ParamOrderID String, ParamCurrency FixedString(3), ParamCurrencyID UInt16, GoalsReached Array(UInt32), OpenstatServiceName String, OpenstatCampaignID String, OpenstatAdID String, OpenstatSourceID String, UTMSource String, UTMMedium String, UTMCampaign String, UTMContent String, UTMTerm String, FromTag String, HasGCLID UInt8, RefererHash UInt64, URLHash UInt64, CLID UInt32, YCLID UInt64, ShareService String, ShareURL String, ShareTitle String, ParsedParams Nested( Key1 String, Key2 String, Key3 String, Key4 String, Key5 String, ValueDouble Float64 ), IslandID FixedString(16), RequestNum UInt32, RequestTry UInt8 ) ENGINE = MergeTree() PARTITION BY toYYYYMM(EventDate) ORDER BY (CounterID, EventDate, intHash32(UserID)); ``` Expand for the `CREATE TABLE default.visits_v1` sample ``` CREATE TABLE default.visits_v1 ( CounterID UInt32, StartDate Date, Sign Int8, IsNew UInt8, VisitID UInt64, UserID UInt64, StartTime DateTime, Duration UInt32, UTCStartTime DateTime, PageViews Int32, Hits Int32, IsBounce UInt8, Referer String, StartURL String, RefererDomain String, StartURLDomain String, EndURL String, LinkURL String, IsDownload UInt8, TraficSourceID Int8, SearchEngineID UInt16, SearchPhrase String, AdvEngineID UInt8, PlaceID Int32, RefererCategories Array(UInt16), URLCategories Array(UInt16), URLRegions Array(UInt32), RefererRegions Array(UInt32), IsYandex UInt8, GoalReachesDepth Int32, GoalReachesURL Int32, GoalReachesAny Int32, SocialSourceNetworkID UInt8, SocialSourcePage String, MobilePhoneModel String, ClientEventTime DateTime, RegionID UInt32, ClientIP UInt32, ClientIP6 FixedString(16), RemoteIP UInt32, RemoteIP6 FixedString(16), IPNetworkID UInt32, SilverlightVersion3 UInt32, CodeVersion UInt32, ResolutionWidth UInt16, ResolutionHeight UInt16, UserAgentMajor UInt16, UserAgentMinor UInt16, WindowClientWidth UInt16, WindowClientHeight UInt16, SilverlightVersion2 UInt8, SilverlightVersion4 UInt16, FlashVersion3 UInt16, FlashVersion4 UInt16, ClientTimeZone Int16, OS UInt8, UserAgent UInt8, ResolutionDepth UInt8, FlashMajor UInt8, FlashMinor UInt8, NetMajor UInt8, NetMinor UInt8, MobilePhone UInt8, SilverlightVersion1 UInt8, Age UInt8, Sex UInt8, Income UInt8, JavaEnable UInt8, CookieEnable UInt8, JavascriptEnable UInt8, IsMobile UInt8, BrowserLanguage UInt16, BrowserCountry UInt16, Interests UInt16, Robotness UInt8, GeneralInterests Array(UInt16), Params Array(String), Goals Nested( ID UInt32, Serial UInt32, EventTime DateTime, Price Int64, OrderID String, CurrencyID UInt32 ), WatchIDs Array(UInt64), ParamSumPrice Int64, ParamCurrency FixedString(3), ParamCurrencyID UInt16, ClickLogID UInt64, ClickEventID Int32, ClickGoodEvent Int32, ClickEventTime DateTime, ClickPriorityID Int32, ClickPhraseID Int32, ClickPageID Int32, ClickPlaceID Int32, ClickTypeID Int32, ClickResourceID Int32, ClickCost UInt32, ClickClientIP UInt32, ClickDomainID UInt32, ClickURL String, ClickAttempt UInt8, ClickOrderID UInt32, ClickBannerID UInt32, ClickMarketCategoryID UInt32, ClickMarketPP UInt32, ClickMarketCategoryName String, ClickMarketPPName String, ClickAWAPSCampaignName String, ClickPageName String, ClickTargetType UInt16, ClickTargetPhraseID UInt64, ClickContextType UInt8, ClickSelectType Int8, ClickOptions String, ClickGroupBannerID Int32, OpenstatServiceName String, OpenstatCampaignID String, OpenstatAdID String, OpenstatSourceID String, UTMSource String, UTMMedium String, UTMCampaign String, UTMContent String, UTMTerm String, FromTag String, HasGCLID UInt8, FirstVisit DateTime, PredLastVisit Date, LastVisit Date, TotalVisits UInt32, TraficSource Nested( ID Int8, SearchEngineID UInt16, AdvEngineID UInt8, PlaceID UInt16, SocialSourceNetworkID UInt8, Domain String, SearchPhrase String, SocialSourcePage String ), Attendance FixedString(16), CLID UInt32, YCLID UInt64, NormalizedRefererHash UInt64, SearchPhraseHash UInt64, RefererDomainHash UInt64, NormalizedStartURLHash UInt64, StartURLDomainHash UInt64, NormalizedEndURLHash UInt64, TopLevelDomain UInt64, URLScheme UInt64, OpenstatServiceNameHash UInt64, OpenstatCampaignIDHash UInt64, OpenstatAdIDHash UInt64, OpenstatSourceIDHash UInt64, UTMSourceHash UInt64, UTMMediumHash UInt64, UTMCampaignHash UInt64, UTMContentHash UInt64, UTMTermHash UInt64, FromHash UInt64, WebVisorEnabled UInt8, WebVisorActivity UInt32, ParsedParams Nested( Key1 String, Key2 String, Key3 String, Key4 String, Key5 String, ValueDouble Float64 ), Market Nested( Type UInt8, GoalID UInt32, OrderID String, OrderPrice Int64, PP UInt32, DirectPlaceID UInt32, DirectOrderID UInt32, DirectBannerID UInt32, GoodID String, GoodName String, GoodQuantity Int32, GoodPrice Int64 ), IslandID FixedString(16) ) ENGINE = CollapsingMergeTree(Sign) PARTITION BY toYYYYMM(StartDate) ORDER BY (CounterID, StartDate, intHash32(UserID), VisitID) ``` 3. Load data into tables `hits_v1` and `visits_v1`. 1. Go to the folder where you stored the downloaded files for `hits_v1.tsv` and `visits_v1.tsv`. 2. Run the following commands: ``` cat hits_v1.tsv | docker run \ --interactive \ --rm clickhouse/clickhouse-server clickhouse-client \ --user USERNAME \ --password PASSWORD \ --host HOST \ --port PORT \ --secure \ --max_insert_block_size=100000 \ --query="INSERT INTO default.hits_v1 FORMAT TSV" ``` ``` cat visits_v1.tsv | docker run \ --interactive \ --rm clickhouse/clickhouse-server clickhouse-client \ --user USERNAME \ --password PASSWORD \ --host HOST \ --port PORT \ --secure \ --max_insert_block_size=100000 \ --query="INSERT INTO default.visits_v1 FORMAT TSV" ``` ## Query data[​](#query-data "Direct link to Query data") Once the data is loaded, you can run queries against the sample data you imported. * Query the number of items in the `hits_v1` table: ``` SELECT COUNT(*) FROM default.hits_v1 ``` * Find the longest lasting sessions: ``` SELECT StartURL AS URL, MAX(Duration) AS MaxDuration FROM default.visits_v1 GROUP BY URL ORDER BY MaxDuration DESC LIMIT 10 ``` ## Next steps[​](#next-steps "Direct link to Next steps") * [Service architecture](/docs/products/clickhouse/concepts/service-architecture.md) * [Secure an Aiven for ClickHouse® service](/docs/products/clickhouse/howto/secure-service.md) * [Manage Aiven for ClickHouse® users and roles](/docs/products/clickhouse/howto/manage-users-roles.md) * [Manage Aiven for ClickHouse® database and tables](/docs/products/clickhouse/howto/manage-databases-tables.md) * [Aiven for ClickHouse® service integrations](/docs/products/clickhouse/concepts/data-integration-overview.md) --- # Change the cloud or region for your Aiven for ClickHouse® service Move your Aiven for ClickHouse® service to a different cloud provider or region. 1. In your service, click **Service settings** from the sidebar. 2. In the **Cloud and network** section, click **Actions** > **Change cloud**. 3. In the **Cloud** section , select a cloud provider and region, and click **Change**. Your service starts a migration to the new location and remains available during the process. When the migration completes, the service continues running in the new cloud or region. Related pages * [Fork your Aiven for ClickHouse® service](/docs/products/clickhouse/howto/fork-service.md) * [Migrate to another cloud or region](/docs/platform/howto/migrate-services-cloud-region.md) --- # Change the plan for your Aiven for ClickHouse® service Change the service plan for your Aiven for ClickHouse® service to scale resources up or down and optimize costs. Adjust the plan of your services at any time to scale your services as needed and optimize costs. If you can't find a suitable plan, you can [request a custom plan](/docs/platform/concepts/service-pricing.md). tip If you plan to upgrade your service plan, do it immediately after a full backup. This reduces the amount of incremental changes that need to be applied on top of the base backup, which speeds up the upgrade itself. important * When changing a service plan, reserve an additional 25% of disk space. This requirement applies to upgrades and downgrades. * Downgrading to a plan with fewer VMs is supported for most services, including Aiven for Apache Kafka®, Aiven for PostgreSQL®, Aiven for OpenSearch®, Aiven for ClickHouse®, Aiven for MySQL®, Aiven for Metrics, and Aiven for Valkey™. * Changing a service plan triggers a node recycle, service rebuilding, and any pending maintenance updates. - Console - Terraform - CLI 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Change plan**. 3. Select a plan that provides at least 125% of the current disk size and click **Change plan**. Update the `plan` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). To change a service plan in the Aiven CLI, use the [`avn service update --plan `](/docs/tools/cli/service-cli.md#avn-cli-service-update) command. Your service's state becomes **Rebuilding** and remains accessible. When the state switches to **Running**, your new service plan is active. Related pages * [Scale disk storage](/docs/products/clickhouse/howto/scale-disk-storage.md) * [Maintenance and updates](/docs/products/clickhouse/howto/maintenance-updates.md) * [Fork your Aiven for ClickHouse® service](/docs/products/clickhouse/howto/fork-service.md) --- # Check data distribution between storage devices in Aiven for ClickHouse®'s tiered storage Monitor how your data is distributed between the two layers of your tiered storage: Network-attached block storage and object storage. If you have the tiered storage feature [enabled](/docs/products/clickhouse/howto/enable-tiered-storage.md), your data in Aiven for ClickHouse is distributed between two storage devices (tiers). You can check on what storage devices your databases and tables are stored. You can also preview their total sizes as well as part counts, minimum part sizes, median part sizes, and maximum part sizes. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Tiered storage [enabled](/docs/products/clickhouse/howto/enable-tiered-storage.md) * Access to the [Aiven Console](https://console.aiven.io/) * Command line tool ([ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md)) installed ## Check data distribution in Aiven Console[​](#check-data-distribution-in-aiven-console "Direct link to Check data distribution in Aiven Console") You can use the [Aiven Console](https://console.aiven.io/) to check if tiered storage is enabled on a table and, if it is, how much storage is used on each tier (network-attached block storage and object storage) for this particular table. To access tiered storage status information: 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose your Aiven for ClickHouse service. 2. In the service sidebar, click **Data** > **Databases and tables**. 3. Open the database and table that you want to check. 4. Click **Actions** > **View details** > **Tiered storage**. ## Run a data distribution check with the ClickHouse client[​](#run-a-data-distribution-check-with-the-clickhouse-client "Direct link to Run a data distribution check with the ClickHouse client") 1. [Connect to your Aiven for ClickHouse service](/docs/products/clickhouse/howto/list-connect-to-service.md) using, for example, the ClickHouse client. 2. Run the following query: ``` SELECT database, table, disk_name, formatReadableSize(sum(data_compressed_bytes)) AS total_size, count(*) AS parts_count, formatReadableSize(min(data_compressed_bytes)) AS min_part_size, formatReadableSize(median(data_compressed_bytes)) AS median_part_size, formatReadableSize(max(data_compressed_bytes)) AS max_part_size FROM system.parts GROUP BY database, table, disk_name ORDER BY database ASC, table ASC, disk_name ASC ``` You can expect to receive the following output: ``` ┌─database─┬─table─────┬─disk_name─┬─total_size─┬─parts_count─┬─min_part_size─┬─median_part_size─┬─max_part_size─┐ │ datasets │ hits_v1 │ default │ 1.20 GiB │ 6 │ 33.65 MiB │ 238.69 MiB │ 253.18 MiB │ │ datasets │ visits_v1 │ S3 │ 536.69 MiB │ 5 │ 44.61 MiB │ 57.90 MiB │ 317.19 MiB │ │ system │ query_log │ default │ 75.85 MiB │ 102 │ 7.51 KiB │ 12.36 KiB │ 1.55 MiB │ └──────────┴───────────┴───────────┴────────────┴─────────────┴───────────────┴──────────────────┴───────────────┘ ``` The query returns a table with data distribution details for all databases and tables that belong to your service: the storage device they use, their total sizes as well as parts counts and sizing. ## What's next[​](#whats-next "Direct link to What's next") * [Transfer data between network-attached block storage and object storage](/docs/products/clickhouse/howto/transfer-data-tiered-storage.md) * [Configure data retention thresholds for tiered storage](/docs/products/clickhouse/howto/configure-tiered-storage.md) Related pages * [About tiered storage in Aiven for ClickHouse](/docs/products/clickhouse/concepts/clickhouse-tiered-storage.md) * [Enable tiered storage in Aiven for ClickHouse](/docs/products/clickhouse/howto/enable-tiered-storage.md) --- # Use query cache in Aiven for ClickHouse® Aiven for ClickHouse® provides a query cache mechanism that helps improve query performance by caching query results. ## How it works[​](#how-it-works "Direct link to How it works") When the Aiven for [ClickHouse query cache](https://clickhouse.com/docs/en/operations/query-cache) is enabled, multiple identical `SELECT` queries running simultaneously are computed only once. Subsequent executions of the same query are served directly from the cache. important By default, the Aiven for ClickHouse query cache is disabled for all `SELECT` queries. ## Why use it[​](#why-use-it "Direct link to Why use it") Using query cache in your Aiven for ClickHouse services can help reduce latency and resource consumption. Key use case for the Aiven for ClickHouse query cache are the following: * Performance enhancement: Reducing latency and load for frequently executed queries * Analytical workloads: Using complex repetitive queries that involve high data aggregation ## Enable query cache[​](#enable-query-cache "Direct link to Enable query cache") To enable the query cache for a query, set the `use_query_cache` setting for the query to `1`. You can achieve this by appending `SETTINGS use_query_cache = 1` to the end of your query using an SQL client (for example, the [ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md)): ``` SELECT 1 SETTINGS use_query_cache = 1; ``` ## Configure query cache[​](#configure-query-cache "Direct link to Configure query cache") To configure the query cache settings, use an SQL client (for example, the [ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md)) and append a defined setting to your query, for example: ``` SELECT 1 SETTINGS use_query_cache = 1, query_cache_min_query_runs = 5000; ``` You can configure the following query cache settings: * `enable_writes_to_query_cache` * `enable_reads_from_query_cache` * `query_cache_max_entries` * `query_cache_min_query_runs` * `query_cache_min_query_duration` * `query_cache_compress_entries` * `query_cache_squash_partial_results` * `query_cache_ttl` * `query_cache_share_between_users` ## Limitation[​](#limitation "Direct link to Limitation") * Cached results are not invalidated or discarded when the underlying data (the result of a `SELECT` query) changes, which might cause returning stale results. * Maximum query cache size: 64 MiB for each GiB of RAM (for example, 256 MiB for a 4-GiB service or 1 GiB for a 16-GiB service) * Maximum number of query cache entries: 64 entries for each GiB of RAM (for example, 1024 entries for a 16-GiB service) Related pages * [Querying external data in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/federated-queries.md) * [Query Aiven for ClickHouse® databases](/docs/products/clickhouse/howto/query-databases.md) * [Fetch query statistics for Aiven for ClickHouse®](/docs/products/clickhouse/howto/fetch-query-statistics.md) * [Create dictionaries in Aiven for ClickHouse®](/docs/products/clickhouse/howto/create-dictionary.md) --- # Schedule Aiven for ClickHouse® backups Set the time when [backups](/docs/products/clickhouse/concepts/disaster-recovery.md#service-backup) are taken for your Aiven for ClickHouse® service. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * To [configure backup time](/docs/products/clickhouse/howto/configure-backup.md#configure-backup-time), one of the following tools: * [Aiven Console](https://console.aiven.io) * [Aiven API](/docs/tools/api.md) * [Aiven CLI](/docs/tools/cli.md) * To [reset or delete configured backup settings](/docs/products/clickhouse/howto/configure-backup.md#restore-defaults), one of the following tools: * [Aiven API](/docs/tools/api.md) * [Aiven CLI](/docs/tools/cli.md) * At least one Aiven for ClickHouse service backup to be configured ## Configure backup time[​](#configure-backup-time "Direct link to Configure backup time") note A backup process can only start when the previous backup process completes. To edit the backup schedule for your service: * Console * Aiven API * Aiven CLI * Terraform 1. In your service, **Backups**. 2. Click **Actions** > **Configure backup settings**. 3. Click **Add configuration options**. 4. Add `backup_hour` and `backup_minute`, and set their values. 5. Click **Save configuration**. Call the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint, and add the following properties to the `user_config` object: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "backup_hour": BACKUP_HOUR, "backup_minute": BACKUP_MINUTE } }' ``` Replace the following: * `SERVICE_NAME`: the name of your service. * `PROJECT_NAME`: the name of your project. * `BACKUP_HOUR`: the hour when the service backup starts. Accepted values are integers between `0` and `23`. * `BACKUP_MINUTE`: the minute when the service backup starts. Accepted values are integers between `0` and `59`. Run the [avn service update](/docs/tools/cli/service-cli.md#avn-cli-service-update) command, and add the following properties to the `user_config` object: ``` avn service update SERVICE_NAME \ --project PROJECT_NAME \ --user-config '{ "backup_hour": BACKUP_HOUR, "backup_minute": BACKUP_MINUTE }' ``` Replace the following: * `SERVICE_NAME`: the name of your service. * `PROJECT_NAME`: the name of your project. * `BACKUP_HOUR`: the hour when the service backup starts. Accepted values are integers between `0` and `23`. * `BACKUP_MINUTE`: the minute when the service backup starts. Accepted values are integers between `0` and `59`. Use the `backup_hour` and `backup_minute` attributes in [your service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs) to set the start time for backups. If a backup was recently made, it can take another backup cycle before the new backup time takes effect. ## Restore defaults[​](#restore-defaults "Direct link to Restore defaults") * Console * Aiven API * Aiven CLI The [Aiven Console](https://console.aiven.io) doesn't support restoring defaults. Use the Aiven [API](/docs/tools/api.md) or [CLI](/docs/tools/cli.md). Call the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint, and set the backup scheduling parameters to `null` or remove them from `user_config`. Replace placeholders `SERVICE_NAME` and `PROJECT_NAME` with meaningful values. ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "backup_hour": null, "backup_minute": null } }' ``` Run the [avn service update](/docs/tools/cli/service-cli.md#avn-cli-service-update) command, and set the backup scheduling parameters to `null` or remove them from `user_config`. Replace placeholders `SERVICE_NAME` and `PROJECT_NAME` with meaningful values. ``` avn service update SERVICE_NAME \ --project PROJECT_NAME \ --user-config '{ "backup_hour": null, "backup_minute": null }' ``` Related pages * [Disaster Recovery testing scenarios](/docs/platform/concepts/disaster-recovery-test-scenarios.md) * [Fork your Aiven for ClickHouse® service](/docs/products/clickhouse/howto/fork-service.md) * [Restore an Aiven for ClickHouse® backup](/docs/products/clickhouse/howto/restore-backup.md) * [Disaster recovery in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/disaster-recovery.md) --- # Configure data retention thresholds in Aiven for ClickHouse®'s tiered storage Control how your data is distributed between storage devices in the tiered storage of an Aiven for ClickHouse® service. Configure tables so that ClickHouse automatically writes your data to network-attached block storage or object storage as needed. If you have [tiered storage enabled](/docs/products/clickhouse/howto/enable-tiered-storage.md) on your Aiven for ClickHouse service, Aiven distributes your data between two storage devices (tiers). Data is stored either on network-attached block storage or in object storage, depending on whether and how you configure this behavior. By default, ClickHouse moves data from network-attached block storage to object storage when it reaches 80% of its capacity (default size-based data retention policy). To change this default data distribution behavior, [configure your table's schema by adding a TTL (time-to-live) clause](/docs/products/clickhouse/howto/configure-tiered-storage.md#time-based-retention-config). Such a configuration allows ignoring the capacity threshold for network-attached block storage and moving the data from it to object storage based on how long the data has been stored there. To enable this time-based data distribution mechanism, you can set up a retention policy (threshold) on a table level by using the TTL clause. For data retention control purposes, the TTL clause uses the following: * Data item of the `Date` or `DateTime` type as a reference point in time * INTERVAL clause as a time period to elapse between the reference point and the data transfer to object storage ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Tiered storage enabled](/docs/products/clickhouse/howto/enable-tiered-storage.md) * Command line tool installed ([ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md)) ## Configure time-based data retention[​](#time-based-retention-config "Direct link to Configure time-based data retention") 1. [Connect to your Aiven for ClickHouse service](/docs/products/clickhouse/howto/list-connect-to-service.md) using, for example, the ClickHouse client. 2. Select a database for operations you intend to perform. ``` USE DATABASE_NAME ``` ### Add or modify TTL[​](#add-or-modify-ttl "Direct link to Add or modify TTL") * Add TTL to a new table * Modify TTL on an existing table Create a table with the `storage_policy` setting set to `tiered` (to [enable tiered storage](/docs/products/clickhouse/howto/enable-tiered-storage.md)) and TTL (time-to-live) configured to add a time-based data retention threshold on the table. ``` CREATE TABLE example_table ( SearchDate Date, SearchID UInt64, SearchPhrase String ) ENGINE = MergeTree ORDER BY (SearchDate, SearchID) PARTITION BY toYYYYMM(SearchDate) TTL SearchDate + INTERVAL 1 WEEK TO VOLUME 'remote' SETTINGS storage_policy = 'tiered'; ``` Add or update a TTL definition with the `ALTER TABLE ... MODIFY TTL` statement: ``` ALTER TABLE database_name.table_name MODIFY TTL ttl_expression; ``` After TTL is configured, ClickHouse moves data older than the specified time period from network-attached block storage to object storage, regardless of available capacity. ## Best practices for tiered storage TTL[​](#best-practices-for-tiered-storage-ttl "Direct link to Best practices for tiered storage TTL") Follow these recommendations to optimize performance and efficiency when using TTL with tiered storage. ### Optimize part sizes for remote storage[​](#optimize-part-sizes-for-remote-storage "Direct link to Optimize part sizes for remote storage") Avoid creating many small parts on remote storage, as this can negatively impact performance. When writing data that will be immediately moved to remote storage (such as during backfilling of historical data): * **Use large inserts**: Ensure your data inserts are large enough to create substantial parts on remote storage. * **Temporarily disable TTL moves**: Use the following commands to pause data movement while smaller parts merge together: ``` -- Stop TTL-based data moves temporarily SYSTEM STOP MOVES; -- Perform your data operations (inserts, merges) -- ... your operations here ... -- Resume TTL-based data moves SYSTEM START MOVES; ``` warning Remember to run `SYSTEM START MOVES` after your operations to resume normal TTL behavior. Leaving moves disabled will prevent automatic data tiering. ### Configure efficient data deletion[​](#configure-efficient-data-deletion "Direct link to Configure efficient data deletion") Use the `ttl_only_drop_parts` setting when using TTL for data **deletion**, not just for moving between tiers: ``` CREATE TABLE example_table ( SearchDate Date, SearchID UInt64, SearchPhrase String ) ENGINE = MergeTree ORDER BY (SearchDate, SearchID) PARTITION BY toYYYYMM(SearchDate) TTL SearchDate + INTERVAL 1 MONTH DELETE SETTINGS storage_policy = 'tiered', ttl_only_drop_parts = 1; ``` #### How this helps[​](#how-this-helps "Direct link to How this helps") * **Prevents inefficient partial drops**: Instead of repeatedly rewriting parts as individual rows expire, ClickHouse drops entire parts at once. * **Requires matching partition strategy**: Use a `PARTITION BY` expression that aligns with your TTL period so all data in a partition expires simultaneously. * **Improves performance**: Eliminates the overhead of multiple partial rewrites. #### Example of aligned partitioning and TTL[​](#example-of-aligned-partitioning-and-ttl "Direct link to Example of aligned partitioning and TTL") ``` CREATE TABLE example_with_deletion ( SearchDate Date, SearchID UInt64, SearchPhrase String ) ENGINE = MergeTree ORDER BY (SearchDate, SearchID) -- Partition by month, TTL deletes data older than 1 month PARTITION BY toYYYYMM(SearchDate) TTL SearchDate + INTERVAL 1 MONTH DELETE SETTINGS storage_policy = 'tiered', ttl_only_drop_parts = 1; ``` This ensures that when data expires, ClickHouse drops entire monthly partitions rather than removing individual rows from parts. ## What's next[​](#whats-next "Direct link to What's next") * [Check data volume distribution between different disks](/docs/products/clickhouse/howto/check-data-tiered-storage.md) Related pages * [About tiered storage in Aiven for ClickHouse](/docs/products/clickhouse/concepts/clickhouse-tiered-storage.md) * [Enable tiered storage in Aiven for ClickHouse](/docs/products/clickhouse/howto/enable-tiered-storage.md) * [Transfer data between network-attached block storage and object storage](/docs/products/clickhouse/howto/transfer-data-tiered-storage.md) * [Manage Data with TTL (Time-to-live)](https://clickhouse.com/docs/en/guides/developer/ttl) * [Create table statement, TTL documentation](https://clickhouse.com/docs/en/engines/table-engines/mergetree-family/mergetree#mergetree-table-ttl) * [MergeTree - column TTL](https://clickhouse.com/docs/en/engines/table-engines/mergetree-family/mergetree#mergetree-column-ttl) --- # Visualize ClickHouse® data with Grafana® You can visualise your ClickHouse® data using Grafana® and Aiven can help you connect the two services. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") 1. Aiven for ClickHouse® service accessible by HTTPS 2. Aiven for Grafana® service (see how to [get started with Aiven for Grafana®](/docs/products/grafana/get-started.md)) ## Variables[​](#variables "Direct link to Variables") You'll need a few variables for the setup. To get their values, go to [Aiven Console](https://console.aiven.io/) and go to **Overview** of your Aiven for ClickHouse® service (**Connection information** > **ClickHouse HTTPS & JDBC**). | Variable | Description | | ---------------------- | --------------------------------------------- | | `CLICKHOUSE_HTTPS_URI` | HTTPS service URI of your ClickHouse service. | | `CLICKHOUSE_USER` | Username to access ClickHouse service. | | `CLICKHOUSE_PASSWORD` | Password to access ClickHouse service. | ## Integrate ClickHouse® with Grafana®[​](#integrate-clickhouse-with-grafana "Direct link to Integrate ClickHouse® with Grafana®") 1. Log in to Aiven for Grafana® following [the instructions](/docs/products/grafana/get-started.md#log-in-to-grafana). 2. From the **Configuration** menu, select **Data sources** > **Add data source**. 3. Find **Altinity plugin for ClickHouse** in the list and select it. 4. Set **URL** to `CLICKHOUSE_HTTPS_URI`. 5. In **Auth** section, enable **Basic auth** and **With Credentials**. 6. In **Basic Auth Details**, set your `CLICKHOUSE_USER` and `CLICKHOUSE_PASSWORD`. 7. Select **Save & test**. Now you can create a dashboard and panels to work with the data from your Aiven for ClickHouse® service. --- # Connect to Aiven for ClickHouse® with clickhouse-client It's recommended to connect to a ClickHouse® cluster with the ClickHouse® client. ## Use the ClickHouse® client[​](#use-the-clickhouse-client "Direct link to Use the ClickHouse® client") To use the ClickHouse® client across different operating systems, we recommend utilizing [Docker](https://www.docker.com/). You can get the latest image of the ClickHouse server which contains the most recent ClickHouse client directly from [the dedicated page in Docker hub](https://hub.docker.com/r/clickhouse/clickhouse-server). note There are other installation options available for ClickHouse clients for different operating systems. See them in [ClickHouse local](https://clickhouse.com/docs/en/operations/utilities/clickhouse-local) and [Install ClickHouse](https://clickhouse.com/docs/en/install) in the official ClickHouse documentation. ## Connection properties[​](#connection-properties "Direct link to Connection properties") You will need to know the following properties to establish a secure connection with your Aiven for ClickHouse service: **Host**, **Port**, **User** and **Password**. You will find these in the **Connection information** section on the **Overview** page of your service in the [Aiven Console](https://console.aiven.io/). ## Command template[​](#command-template "Direct link to Command template") The command to connect to the service looks like this, substitute the placeholders for `USERNAME`, `PASSWORD`, `HOST` and `PORT`: ``` docker run -it \ --rm clickhouse/clickhouse-server clickhouse-client \ --user USERNAME \ --password PASSWORD \ --host HOST \ --port PORT \ --secure ``` This example includes the `-it` option (a combination of `--interactive` and `--tty`) to take you inside the container and the `--rm` option to automatically remove the container after exiting. The other parameters, such as `--user`, `--password`, `--host`, `--port`, `--secure`, and `--query` are arguments accepted by the ClickHouse client. You can see the full list of command line options in [the ClickHouse CLI documentation](https://clickhouse.com/docs/en/interfaces/cli/#command-line-options). Once you're connected to the server, you can type queries directly within the client, for example, to see the list of existing databases, run ``` SHOW DATABASES ``` Alternatively, sometimes you might want to run individual queries and be able to access the command prompt outside the docker container. In this case you can set `--interactive` and use `--query` parameter without entering the docker container: ``` docker run --interactive \ --rm clickhouse/clickhouse-server clickhouse-client \ --user USERNAME \ --password PASSWORD \ --host HOST \ --port PORT \ --secure \ --query="YOUR SQL QUERY GOES HERE" ``` Similar to above example, you can request the list of present databases directly: ``` docker run --interactive \ --rm clickhouse/clickhouse-server clickhouse-client \ --user USERNAME \ --password PASSWORD \ --host HOST \ --port PORT \ --secure \ --query="SHOW DATABASES" ``` --- # Connect to Aiven for ClickHouse® with Go To connect to your Aiven for ClickHouse® service with Go, you can use the native protocol or the HTTPS protocol in specific cases. This article provides you with instructions for both scenarios. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") [Go 1.17 or later](https://go.dev/dl/) ## Install the ClickHouse Go module[​](#install-the-clickhouse-go-module "Direct link to Install the ClickHouse Go module") To install the ClickHouse Go module, run the following command: ``` go get github.com/ClickHouse/clickhouse-go/v2 ``` note If the version of Go is lower than 1.18.4 (visible via `go version`), install an older version of `clickhouse-go`. For this purpose, use command `go get github.com/ClickHouse/clickhouse-go/v2@v2.2`. ## Connect with the native protocol[​](#connect-with-the-native-protocol "Direct link to Connect with the native protocol") ### Identify connection information[​](#identify-connection-information "Direct link to Identify connection information") To run the code for connecting to your service, first identify values of the following variables: | Variable | Description | | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | | `Host` | **Host** for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse native** | | `Port` | **Port** for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse native** | | `Database` | **Database Name** in your the ClickHouse service available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse native** | | `Username` | **User** for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse native** | | `Password` | **Password** for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse native** | ### Connect to the service[​](#connect-to-the-service "Direct link to Connect to the service") Replace the placeholders in the code with meaningful information on your service connection and run the code. ``` package main import "fmt" import "log" import "crypto/tls" import "github.com/ClickHouse/clickhouse-go/v2" func main() { host := "HOST" native_port := NATIVE_PORT database := "DATABASE_NAME" username := "USERNAME" password := "PASSWORD" tls_config := &tls.Config{} conn, err := clickhouse.Open(&clickhouse.Options{ Addr: []string{fmt.Sprintf("%s:%d", host, native_port)}, Auth: clickhouse.Auth{ Database: database, Username: username, Password: password, }, TLS: tls_config, }) if err != nil { log.Fatal(err) } v, err := conn.ServerVersion() if err != nil { log.Fatal(err) } fmt.Println(v) } ``` ## Connect with HTTPS[​](#connect-with-https "Direct link to Connect with HTTPS") important The HTTPS connection is supported for the database/SQL API only. By default, connections are established over the native protocol. The HTTPS connection needs to be enabled either by modifying the DSN to include the HTTPS protocol or by specifying the protocol in the connection options. ### Identify connection information[​](#identify-connection-information-1 "Direct link to Identify connection information") To run the code for connecting to your service, first identify values of the following variables: | Variable | Description | | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `Host` | **Host** for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `HttpPort` | **Port** for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `Database` | **Database Name** in your the ClickHouse service available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `Username` | **User** for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `Password` | **Password** for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | ### Connect to the service[​](#connect-to-the-service-1 "Direct link to Connect to the service") Replace the placeholders in the code with meaningful information on your service connection and run the code. ``` package main import "database/sql" import "fmt" import "log" import _ "github.com/ClickHouse/clickhouse-go/v2" func main() { host := "HOST" https_port := HTTPS_PORT username := "USERNAME" password := "PASSWORD" conn, err := sql.Open( "clickhouse", fmt.Sprintf( "https://%s:%d?username=%s&password=%s&secure", host, https_port, username, password)) if err != nil { log.Fatal(err) } rows, err := conn.Query("SELECT version()") if err != nil { log.Fatal(err) } defer rows.Close() for rows.Next() { var version string if err := rows.Scan(&version); err != nil { log.Fatal(err) } fmt.Println(version) } } ``` You have your service connection established and configured. You can proceed to [uploading data into your database](/docs/products/clickhouse/get-started.md#load-a-dataset). Related pages * For instructions on how to configure connection settings, see [Connection Details](https://clickhouse.com/docs/en/integrations/go#connection-details). * For information on how to connect to the Aiven for ClickHouse service with the ClickHouse client, see [Connect with the ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md). --- # Connect to Aiven for ClickHouse® with Java Learn how to connect to your Aiven for ClickHouse® service with Java using the ClickHouse JDBC driver and the HTTPS port. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Java 8](https://www.java.com/en/download/) or later * [ClickHouse JDBC driver](https://github.com/ClickHouse/clickhouse-jdbc/tree/master/clickhouse-jdbc) ## Identify connection information[​](#identify-connection-information "Direct link to Identify connection information") To run the code for connecting to your service, first identify values of the following variables: | Variable | Description | | ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | | `CLICKHOUSE_HTTPS_HOST` | `Host` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `CLICKHOUSE_HTTPS_PORT` | `Port` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `CLICKHOUSE_USER` | `User` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `CLICKHOUSE_PASSWORD` | `Password` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | ## Connect to the service[​](#connect-to-the-service "Direct link to Connect to the service") 1. Add the ClickHouse JDBC driver to your Maven dependencies. ``` [](dependency) com.clickhouse clickhouse-jdbc 0.3.2-patch11 all * * [](/dependency) ``` 2. Replace `CLICKHOUSE_HTTPS_HOST` and `CLICKHOUSE_HTTPS_PORT` in the command with your connection values and run the code. ``` jdbc:ch://CLICKHOUSE_HTTPS_HOST:CLICKHOUSE_HTTPS_PORT?ssl=true&sslmode=STRICT ``` 3. Replace `CLICKHOUSE_USER` and `CLICKHOUSE_PASSWORD` in the code with meaningful data and run the code. ``` import com.clickhouse.jdbc.ClickHouseConnection; import com.clickhouse.jdbc.ClickHouseDataSource; import java.sql.ResultSet; import java.sql.SQLException; import java.sql.Statement; public class Main { public static void main(String[] args) throws SQLException { String connString = "jdbc:ch://CLICKHOUSE_HTTPS_HOST:CLICKHOUSE_HTTPS_PORT?ssl=true&sslmode=STRICT"; ClickHouseDataSource database = new ClickHouseDataSource(connString); ClickHouseConnection connection = database.getConnection("CLICKHOUSE_USER", "CLICKHOUSE_PASSWORD"); Statement statement = connection.createStatement(); ResultSet result_set = statement.executeQuery("SELECT 1 AS one"); while (result_set.next()) { System.out.println(result_set.getInt("one")); } } } ``` Now you have your service connection set up and you can proceed to [uploading data into your database](/docs/products/clickhouse/get-started.md#load-a-dataset). Related pages For information on how to connect to the Aiven for ClickHouse service with the ClickHouse client, see [Connect with the ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md). --- # Connect Aiven for ClickHouse® to external databases via JDBC You can use [ClickHouse JDBC driver](https://github.com/ClickHouse/clickhouse-jdbc/tree/master/clickhouse-jdbc) to connect external sources to your Aiven for ClickHouse database. You will need Aiven for ClickHouse® service, accessible by HTTPS. The connection values you need can be found in [Aiven Console](https://console.aiven.io/) > your service's page > **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC**. | Variable | Description | | ----------------------- | ---------------------------------------------- | | `CLICKHOUSE_HTTPS_HOST` | HTTPS service host of your ClickHouse service. | | `CLICKHOUSE_HTTPS_PORT` | HTTPS service port of your ClickHouse service. | | `CLICKHOUSE_USER` | Username to access ClickHouse service. | | `CLICKHOUSE_PASSWORD` | Password to access ClickHouse service. | ## Connection string[​](#connection-string "Direct link to Connection string") Replace `CLICKHOUSE_HTTPS_HOST` and `CLICKHOUSE_HTTPS_PORT` with your connection values: ``` jdbc:ch://CLICKHOUSE_HTTPS_HOST:CLICKHOUSE_HTTPS_PORT?ssl=true&sslmode=STRICT ``` You'll also need to provide user name and password to establish the connection. For example, if you use Java: ``` Connection connection = dataSource.getConnection("CLICKHOUSE_USER", "CLICKHOUSE_PASSWORD"); ``` --- # Connect to Aiven for ClickHouse® with Node.js Learn how to connect to your Aiven for ClickHouse® service with Node.js using the official Node.js client for connecting to ClickHouse and the HTTPS port. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Node.js](https://nodejs.org/en/download/) in your environment * [Node.js client for connecting to ClickHouse](https://clickhouse.com/docs/en/integrations/language-clients/javascript#environment-requirements-nodejs) tip You can install the Node.js client for connecting to ClickHouse using ``` npm i @clickhouse/client ``` ## Identify connection information[​](#identify-connection-information "Direct link to Identify connection information") To run the code for connecting to your service, first identify values of the following variables: | Variable | Description | | --------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `CLICKHOUSE_HOST` | `https://HOST:HTTPS_PORT`, where `Host` and `Port` for the ClickHouse connection are available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `CLICKHOUSE_USER` | `User` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `CLICKHOUSE_PASSWORD` | `Password` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | ## Connect to the service[​](#connect-to-the-service "Direct link to Connect to the service") Replace the placeholders in the code with meaningful information on your service connection and run the code. ``` import { createClient } from '@clickhouse/client' const client = createClient({ host: "CLICKHOUSE_HOST", username: "CLICKHOUSE_USER", password: "CLICKHOUSE_PASSWORD", database: "default", }) const response = await client.query({ query : "SELECT 1", format: "JSONEachRow", wait_end_of_query: 1, }) const data = await response.json() console.log(data) ``` Now you have your service connection set up and you can proceed to [uploading data into your database](/docs/products/clickhouse/get-started.md#load-a-dataset). Related pages For information on how to connect to the Aiven for ClickHouse service with the ClickHouse client, see [Connect with the ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md). --- # Connect to Aiven for ClickHouse® with PHP Learn how to connect to your Aiven for ClickHouse® service with PHP using the PHP ClickHouse client and the HTTPS port. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [PHP 7.4 or later](https://www.php.net/downloads) * `smi2/phpclickhouse` library * [Composer](https://getcomposer.org/) tip You can install `smi2/phpclickhouse` with the following command: ``` composer require smi2/phpclickhouse ``` or ``` php composer.phar require smi2/phpclickhouse ``` ## Identify connection information[​](#identify-connection-information "Direct link to Identify connection information") To run the code for connecting to your service, first identify values of the following variables: | Variable | Description | | ---------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | | `https` | Required to be set to `true` | | `host` | `Host` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `port` | `Port` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `username` | `User` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `password` | `Password` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `database` | `Database Name` in the ClickHouse service available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | ## Connect to the service[​](#connect-to-the-service "Direct link to Connect to the service") Replace the placeholders in the code with meaningful information on your service connection and run the code. ``` true, 'host' => 'HOSTNAME', 'port' => 'HTTPS_PORT', 'username' => 'USERNAME', 'password' => 'PASSWORD' ]); $db->database('DATABASE'); $response = $db->select('SELECT 1'); print_r($response->rows()); ``` Now you have your service connection set up and you can proceed to [uploading data into your database](/docs/products/clickhouse/get-started.md#load-a-dataset). Related pages * [Connect with the ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md). --- # Connect to Aiven for ClickHouse® with Python To connect to your Aiven for ClickHouse® service with Python, you can use either the native protocol or the HTTPS protocol. ## Connect with the native protocol[​](#connect-with-the-native-protocol "Direct link to Connect with the native protocol") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Python 3.5 or later](https://www.python.org/downloads/) * [ClickHouse Python Driver](https://pypi.org/project/clickhouse-driver/) ### Identify connection information[​](#identify-connection-information "Direct link to Identify connection information") To run the code for connecting to your service, first identify values of the following variables: | Variable | Description | | ------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | | `USERNAME` | `User` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse native** | | `PASSOWRD` | `Password` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse native** | | `HOST` | `Host` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse native** | | `NATIVE_PORT` | `Port` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse native** | | `query` | Query to run, for example `SELECT 1` | ### Connect to the service[​](#connect-to-the-service "Direct link to Connect to the service") Replace the placeholders in the code with meaningful information on your service connection and run the code. ``` from clickhouse_driver import Client client = Client(user="USERNAME", password="PASSWORD", host="HOST", port=NATIVE_PORT, secure=True) print(client.execute("SELECT 1")) ``` ## Connect with HTTPS[​](#connect-with-https "Direct link to Connect with HTTPS") ### Prerequisites[​](#prerequisites-1 "Direct link to Prerequisites") * [Python 3.7 or later](https://www.python.org/downloads/) * [Requests HTTP library](https://pypi.org/project/requests/) ### Identify connection information[​](#identify-connection-information-1 "Direct link to Identify connection information") To run the code for connecting to your service, first identify values of the following variables: | Variable | Description | | ------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `https://HOST:HTTPS_PORT` | `Host` and `Port` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `query` | Query to run, for example `SELECT 1` | | `X-ClickHouse-Database` | `Database Name` available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC**, for example `system` | | `X-ClickHouse-User` | `User` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `X-ClickHouse-Key` | `Password` for the ClickHouse connection available in the Aiven console: Service **Overview** > **Connection information** > **ClickHouse HTTPS & JDBC** | | `X-ClickHouse-Format` | Format for the output from your query, for example `JSONCompact` | ### Connect to the service[​](#connect-to-the-service-1 "Direct link to Connect to the service") Replace the placeholders in the code with meaningful information on your service connection and run the code. ``` import requests response = requests.post( "https://HOST:HTTPS_PORT", params={"query": "SELECT 1"}, headers={ "X-ClickHouse-Database": "system", "X-ClickHouse-User": "USERNAME", "X-ClickHouse-Key": "PASSWORD", "X-ClickHouse-Format": "JSONCompact", }) print(response.text) ``` Now you have your service connection set up and you can proceed to [uploading data into your database](/docs/products/clickhouse/get-started.md#load-a-dataset). For information on how to connect to the Aiven for ClickHouse service with the ClickHouse client, see [Connect with the ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md). --- # Controlled upgrade pipelines for your Aiven for ClickHouse® service [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Link Aiven for ClickHouse® services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. Control when your Aiven managed services receive maintenance updates and test maintenance updates in development or staging environments before they reach production. important Controlled upgrade pipeline is a [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-) feature. [Contact Aiven](https://aiven.io/contact) to request access. Aiven performs automatic service maintenance for security fixes, minor software updates, and other platform changes. The controlled upgrade pipeline feature lets you link services of the same type in an ordered sequence to control when each service receives updates. After a maintenance update upgrades a service at the initial pipeline step, you validate that service version before the update proceeds to the service at the next pipeline step. Validating means approving the new version as safe to roll out to the next service. Validation can be manual or automatic after a configurable delay. ## Why use controlled upgrade pipelines[​](#why-use-controlled-upgrade-pipelines "Direct link to Why use controlled upgrade pipelines") Controlled upgrade pipelines prevent production incidents caused by automatic updates reaching production before teams can test the new version in a lower environment. They give you full oversight of the update process: * **Risk mitigation**: Prevents unexpected maintenance updates from breaking your production environment by ensuring they are tested in a non-production setting first. * **Stability**: Keeps destination services (such as production) on a known-good version until you, or the automatic timer, confirm the new version is safe. * **Process control**: Allows platform teams to standardize their deployment and maintenance lifecycle across environments. ## About controlled upgrade pipelines[​](#about-controlled-upgrade-pipelines "Direct link to About controlled upgrade pipelines") ### Upgrade steps[​](#upgrade-steps "Direct link to Upgrade steps") An upgrade step is a pair of services linked by an upgrade constraint: * **Source service**: The service that receives maintenance updates first * **Destination service**: The service that waits for validation before receiving updates Each destination service can have only one source service. A source service can have multiple destination services. ### Upgrade pipelines[​](#upgrade-pipelines "Direct link to Upgrade pipelines") An upgrade pipeline is a chain of upgrade steps that spans multiple environments. For example: * Single chain: development → staging → production * Multiple destinations: development → production-eu and development → production-na ## How validation works[​](#how-validation-works "Direct link to How validation works") When a maintenance update upgrades your source service: 1. The source service receives the update first. 2. Test the updated source service to verify it works as expected. 3. Validate the update manually using the API or CLI, or wait for automatic validation after the configured delay. The default delay is 7 days. 4. After validation, the destination service becomes eligible for the same maintenance update. 5. The destination service receives the update during its next maintenance window. If one source service has multiple destination services, one validation for the source service applies to all connected destination services. ### Validation and maintenance windows[​](#validation-and-maintenance-windows "Direct link to Validation and maintenance windows") Validation and the maintenance window control different things: * **Validation** controls *what* version the destination service upgrades to. * The maintenance window controls *when* the upgrade happens. After you validate an update, or automatic validation applies, the destination service receives the validated version during its next scheduled maintenance window. Validation does not trigger an immediate upgrade outside the maintenance window. Upgrade pipelines add a constraint on what is installed during a maintenance update; they do not change when maintenance runs. Nodes in the destination service maintain the validated version until a newer version is validated, either when you validate it manually or when automatic validation applies after the configured delay. When a node is recycled, it uses the same validated version, not the latest available version. When you create a step, the destination service keeps the newest version that is already validated at that moment. If the destination service is already applying maintenance during step creation, the in-progress target version becomes the initial validated version. warning A powered-off source service cannot receive maintenance updates, so you cannot validate it. If you power off services earlier in the chain, the destination service upgrades regardless. For example, in a development → staging → production chain, if both development and staging are powered off, production upgrades without testing and validation in the earlier environments. Keep services in the chain powered on to preserve the protection that upgrade pipelines provide. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * **Same service type**: You can only link services of the same type. For example, two Aiven for PostgreSQL services. * **Chain length**: The default maximum chain depth is 3 services, which is 2 steps. If you need a longer chain, [contact Aiven](https://aiven.io/contact). * **No cycles**: You cannot create circular dependencies between services. * **Emergency overrides**: Aiven can apply critical security or stability fixes to a destination service before explicit validation. * **Supported services**: This feature supports all Aiven service types except Aiven for Apache Flink® and Aiven for MySQL. * **Automatic maintenance updates only**: Pipelines apply to automatic maintenance updates, such as minor service version updates and node image updates. Major version upgrades, for example Aiven for PostgreSQL® 15 to 16, require manual action and are not promoted automatically through the pipeline. * **No permanent blocking**: You cannot prevent an update indefinitely. Automatic validation applies after the configured delay, up to the maximum delay. * **No validation rollback**: You cannot undo a validation after it is recorded. ## Use controlled upgrade pipelines[​](#use-controlled-upgrade-pipelines "Direct link to Use controlled upgrade pipelines") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") To use controlled upgrade pipelines, you need the following: * The feature enabled by Aiven ([Limited availability](/docs/platform/concepts/service-and-feature-releases.md)) * Dev tool of your choice: * [Aiven CLI](/docs/tools/cli.md) Install the latest version of the Aiven CLI to access the `upgrade-pipeline` commands. * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) * Set `PROVIDER_AIVEN_ENABLE_BETA=true` before running Terraform. * See the [resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) for full schema, import format, and lifecycle behavior. * [Aiven Operator for Kubernetes](/docs/tools/kubernetes.md) Install the operator and create an Aiven token secret named `aiven-token` that the operator uses to authenticate against the Aiven API. * Write access to the source and destination projects * At least two services of the same type (for example, two Aiven for PostgreSQL® services) * Services can be in different projects in the same organization ### Set up an upgrade pipeline[​](#set-up-an-upgrade-pipeline "Direct link to Set up an upgrade pipeline") Use the Aiven CLI or API to create upgrade steps between your services. note The `upgrade-pipeline` CLI commands require Aiven CLI version 4.x or later. Command names and parameters may change before general availability. #### Create an upgrade step[​](#create-an-upgrade-step "Direct link to Create an upgrade step") Create a step to link a source service and a destination service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ [--source-project SOURCE_PROJECT] SOURCE_SERVICE \ [--destination-project DESTINATION_PROJECT] DESTINATION_SERVICE \ [--auto-validation-delay-days DAYS] ``` **Options** * `--organization-id` is required. * `--source-project` and `--destination-project` are optional. If you omit either project option, Aiven CLI uses the current default project set with `avn project switch`. * `--auto-validation-delay-days` is optional. Defaults to 7 days if not specified. ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "SOURCE_PROJECT_NAME", "source_service_name": "SOURCE_SERVICE_NAME", "destination_project_name": "DESTINATION_PROJECT_NAME", "destination_service_name": "DESTINATION_SERVICE_NAME", "auto_validation_delay_days": 7 }' ``` Use the [`aiven_upgrade_step`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) resource: ``` resource "aiven_upgrade_step" "example" { organization_id = "ORGANIZATION_ID" source_project_name = "SOURCE_PROJECT_NAME" source_service_name = "SOURCE_SERVICE_NAME" destination_project_name = "DESTINATION_PROJECT_NAME" destination_service_name = "DESTINATION_SERVICE_NAME" auto_validation_delay_days = 7 } ``` Apply an `UpgradePipelineStep` manifest with `kubectl`: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: upgrade-step-sample spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: SOURCE_PROJECT_NAME sourceServiceName: SOURCE_SERVICE_NAME destinationProjectName: DESTINATION_PROJECT_NAME destinationServiceName: DESTINATION_SERVICE_NAME autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable after the resource is created. Parameters: * `source_project_name`: Name of the project containing the source service * `source_service_name`: Name of the source service * `destination_project_name`: Name of the project containing the destination service * `destination_service_name`: Name of the destination service * `auto_validation_delay_days`: Optional. Number of days before automatic validation. The value must be at least `1`. The default is 7 days. The maximum delay you can configure is 30 days. #### List upgrade steps[​](#list-upgrade-steps "Direct link to List upgrade steps") View all upgrade steps you have access to: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step list --organization-id ORGANIZATION_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) To list managed upgrade steps, use: ``` terraform state list 'aiven_upgrade_step.*' terraform state show 'aiven_upgrade_step.example' ``` List `UpgradePipelineStep` resources in the current namespace: ``` kubectl get upgradepipelinesteps ``` #### View a specific step[​](#view-a-specific-step "Direct link to View a specific step") Get details about a specific upgrade step, including the last validation: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step get \ --organization-id ORGANIZATION_ID \ STEP_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` terraform state show aiven_upgrade_step.example ``` Show the manifest and full status, including `id`, `conditions`, and `lastValidation`: ``` kubectl describe upgradepipelinestep RESOURCE_NAME kubectl get upgradepipelinestep RESOURCE_NAME -o yaml ``` The step details include `last_validation` values such as `validated_at`, `validated_by_user`, and `comment` when validation exists (available through the API). ### Validate an upgrade[​](#validate-an-upgrade "Direct link to Validate an upgrade") After testing your source service with the new update, validate the version to allow the destination service to receive the same update. #### Manual validation[​](#manual-validation "Direct link to Manual validation") Validate the current version of your source service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step validate-for-service \ --project SOURCE_PROJECT \ SERVICE_NAME \ [--comment "COMMENT"] ``` `--comment` is optional. Use it to record a note about the validation, for example `"Tested and verified in development"`. ``` curl -X POST https://api.aiven.io/v1/project/SOURCE_PROJECT/service/SOURCE_SERVICE/upgrade-validation \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "comment": "Tested and verified in development" }' ``` Terraform manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. The operator manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. #### Automatic validation[​](#automatic-validation "Direct link to Automatic validation") If you do not manually validate an update, the system automatically validates the source service version after the configured delay. Auto-validation starts from when the source service receives the update. ### Manage upgrade steps[​](#manage-upgrade-steps "Direct link to Manage upgrade steps") #### Update a step[​](#update-a-step "Direct link to Update a step") Modify the automatic validation delay for an existing step: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step update \ --organization-id ORGANIZATION_ID \ --auto-validation-delay-days 14 \ STEP_ID ``` ``` curl -X PATCH https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "auto_validation_delay_days": 14 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` resource "aiven_upgrade_step" "example" { # ...required fields... auto_validation_delay_days = 14 # Updated from 7 to 14 } ``` Apply the changes: ``` terraform plan terraform apply ``` Edit `autoValidationDelayDays` in your manifest and re-apply: ``` spec: autoValidationDelayDays: 14 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable. To change them, delete the resource and create a new one. #### Delete a step[​](#delete-a-step "Direct link to Delete a step") Remove an upgrade step to allow the destination service to receive updates independently: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step delete --organization-id ORGANIZATION_ID STEP_ID ``` Find `STEP_ID` from the upgrade step list command. ``` curl -X DELETE https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) Remove the resource from configuration and apply, or destroy it directly: ``` terraform apply terraform destroy -target=aiven_upgrade_step.example ``` Delete the `UpgradePipelineStep` resource: ``` kubectl delete upgradepipelinestep RESOURCE_NAME ``` Deleting a step removes all associated validations. ### Example: Three-environment pipeline[​](#example-three-environment-pipeline "Direct link to Example: Three-environment pipeline") Create a pipeline that promotes updates from development to staging to production: * CLI * API * Terraform * Kubernetes 1. Create a step from development to staging: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project dev-project \ --destination-project staging-project \ --auto-validation-delay-days 3 \ pg-dev pg-staging ``` 2. Create a step from staging to production: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project staging-project \ --destination-project prod-project \ --auto-validation-delay-days 7 \ pg-staging pg-prod ``` 1) Create a step from development to staging: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "dev-project", "source_service_name": "pg-dev", "destination_project_name": "staging-project", "destination_service_name": "pg-staging", "auto_validation_delay_days": 3 }' ``` 2) Create a step from staging to production: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "staging-project", "source_service_name": "pg-staging", "destination_project_name": "prod-project", "destination_service_name": "pg-prod", "auto_validation_delay_days": 7 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` # Step 1: Development → Staging resource "aiven_upgrade_step" "dev_to_staging" { organization_id = "ORGANIZATION_ID" source_project_name = "dev-project" source_service_name = "pg-dev" destination_project_name = "staging-project" destination_service_name = "pg-staging" auto_validation_delay_days = 3 } # Step 2: Staging → Production resource "aiven_upgrade_step" "staging_to_prod" { organization_id = "ORGANIZATION_ID" source_project_name = "staging-project" source_service_name = "pg-staging" destination_project_name = "prod-project" destination_service_name = "pg-prod" auto_validation_delay_days = 7 } ``` Apply the configuration: ``` export PROVIDER_AIVEN_ENABLE_BETA=true terraform init terraform plan terraform apply ``` Define both steps in a single manifest and apply it: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: dev-to-staging spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: dev-project sourceServiceName: pg-dev destinationProjectName: staging-project destinationServiceName: pg-staging autoValidationDelayDays: 3 --- apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: staging-to-prod spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: staging-project sourceServiceName: pg-staging destinationProjectName: prod-project destinationServiceName: pg-prod autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-pipeline.yaml ``` When a maintenance update arrives: 1. The development service receives the update. 2. After testing, validate the development version or wait 3 days for auto-validation. 3. The staging service receives the update during its next maintenance window. 4. After testing, validate the staging version or wait 7 days for auto-validation. 5. The production service receives the update during its next maintenance window. Related pages * [Maintenance and updates for your Aiven for ClickHouse® service](/docs/products/clickhouse/howto/maintenance-updates.md) * [Change the service plan](/docs/products/clickhouse/howto/change-service-plan.md) * [Service and feature releases](/docs/platform/concepts/service-and-feature-releases.md) * [Aiven CLI](/docs/tools/cli.md) --- # Copy data between Aiven for ClickHouse® services You can copy data from one ClickHouse® server to another using the `remoteSecure()` function. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Details of the source (remote) server * Hostname * Port * Username * Password ## Copy data[​](#copy-data "Direct link to Copy data") 1. From your target server, use the `remoteSecure()` function to select data from the source server. ``` SELECT * FROM remoteSecure('HOSTNAME:PORT', db.remote_engine_table, 'USERNAME', 'PASSWORD') LIMIT 3; ``` tip If you have the [managed credentials integration](/docs/products/clickhouse/concepts/data-integration-overview.md#managed-credentials-integration) enabled, you can use instead: ``` SELECT * FROM remoteSecure('service_YOUR_REMOTE_CLUSTER', db.remote_engine_table) LIMIT 3; ``` See how to [enable the managed credentials integration](/docs/products/clickhouse/howto/data-service-integration.md#create-managed-credentials-integrations). 2. Insert the selected data into the target server. ``` INSERT INTO [db.]table [(c1, c2, c3)] SELECT ... ``` For details on how to configure and use the INSERT query, see [Inserting the Results of SELECT](https://clickhouse.com/docs/en/sql-reference/statements/insert-into/#inserting-the-results-of-select). Your data has been copied from the remote (source) server to the new (target) server. --- # Create dictionaries in Aiven for ClickHouse® Create dictionaries in Aiven for ClickHouse® to accelerate queries for better efficiency and performance. ## Dictionaries in Aiven for ClickHouse[​](#dictionaries-in-aiven-for-clickhouse "Direct link to Dictionaries in Aiven for ClickHouse") A dictionary is a key-attribute mapping useful for low latency lookup queries, when often looking up attributes for a particular key. Dictionary data resides fully in memory, which is why using a dictionary in JOINs is often much faster than using a MergeTree table. Dictionaries can be an efficient replacement for regular tables in your JOIN clauses. Aiven for ClickHouse supports [backup and restore](/docs/products/clickhouse/concepts/disaster-recovery.md#backup-and-restore) for dictionaries. Also, dictionaries in Aiven for ClickHouse are automatically replicated to all service nodes. Read more on dictionaries in the [upstream ClickHouse documentation](https://clickhouse.com/docs/en/sql-reference/dictionaries). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven for ClickHouse service created * SQL client installed * [Dictionary source](/docs/products/clickhouse/howto/create-dictionary.md#supported-sources) available * Credentials integration for remote ClickHouse, PostgreSQL®, or MySQL® if to be used as sources ## Limitations[​](#limitations "Direct link to Limitations") * Only TLS connections supported * If no host is specified in a dictionary with a ClickHouse source, the local host is assumed, and the dictionary is filled with data from a query against the local ClickHouse, for example: ``` -- users table CREATE TABLE default.users ( id UInt64, username String, email String, country String ) ENGINE = MergeTree() ORDER BY id; CREATE DICTIONARY default.users_dictionary ( id UInt64, username String, email String, country String ) PRIMARY KEY id SOURCE(CLICKHOUSE(DB 'default' TABLE 'users')) LAYOUT(FLAT()) LIFETIME(100); ``` In Aiven for ClickHouse, to fill the dictionary the table users are queried with the permissions of the `avnadmin` user even if another user creates the dictionary. In upstream ClickHouse, the same is true except the `default` user is used. * In Aiven for ClickHouse, the [`dictionaries_lazy_load`](https://clickhouse.com/docs/en/operations/server-configuration-parameters/settings#dictionaries_lazy_load) setting is set to `true`, which means that errors with dictionary source parameters may only become apparent when the dictionary is loaded on the first use, rather than when it is created. ### Supported layouts[​](#supported-layouts "Direct link to Supported layouts") Aiven for ClickHouse supports the same [layouts that the upstream ClickHouse supports](https://clickhouse.com/docs/en/sql-reference/dictionaries#ways-to-store-dictionaries-in-memory) with two exceptions,`ssd_cache` and `complex_key_ssd_cache`, which are not supported. ### Supported sources[​](#supported-sources "Direct link to Supported sources") * HTTP(s) * Remote ClickHouse * Aiven for ClickHouse * Remote MySQL® * Aiven for MySQL * Remote PostgreSQL® * Aiven for PostgreSQL ## Create a dictionary[​](#create-a-dictionary "Direct link to Create a dictionary") To create a dictionary with specified structure (attributes), source, layout, and lifetime, use the following syntax: ``` CREATE [OR REPLACE] DICTIONARY [IF NOT EXISTS] [db.]dictionary_name ( key1 type1 [DEFAULT|EXPRESSION expr1] [IS_OBJECT_ID], key2 type2 [DEFAULT|EXPRESSION expr2], attr1 type2 [DEFAULT|EXPRESSION expr3] [HIERARCHICAL|INJECTIVE], attr2 type2 [DEFAULT|EXPRESSION expr4] [HIERARCHICAL|INJECTIVE] ) PRIMARY KEY key1, key2 SOURCE(SOURCE_NAME([param1 value1 ... paramN valueN])) LAYOUT(LAYOUT_NAME([param_name param_value])) LIFETIME({MIN min_val MAX max_val | max_val}) SETTINGS(setting_name = setting_value, setting_name = setting_value, ...) COMMENT 'Comment' ``` ## Examples[​](#examples "Direct link to Examples") ### Speeding up `JOIN`s[​](#speeding-up-joins "Direct link to speeding-up-joins") 1. Create tables in your ClickHouse database: ``` CREATE TABLE users ( id UInt64, username String, email String, country String ) ENGINE = MergeTree() ORDER BY id; ``` ``` CREATE TABLE transactions ( id UInt64, user_id UInt64, product_id UInt64, quantity Float64, price Float64 ) ENGINE = MergeTree() ORDER BY id; ``` 2. Create a dictionary for the `users` table: ``` CREATE DICTIONARY users_dictionary ( id UInt64, username String, email String, country String ) PRIMARY KEY id SOURCE(CLICKHOUSE(DB 'default' TABLE 'users')) LAYOUT(FLAT()) LIFETIME(100); ``` You can do the same using the `QUERY` parameter: ``` CREATE DICTIONARY users_dictionary ( id UInt64, username String, email String, country String ) PRIMARY KEY id SOURCE(CLICKHOUSE(QUERY 'SELECT id, username, email, country FROM default.users')) LAYOUT(FLAT()) LIFETIME(100); ``` `JOIN`s are much faster as the data is pre-indexed in memory. ``` SELECT t.id, u.username, t.product_id, t.quantity, t.price FROM transactions AS t ANY LEFT JOIN users_dictionary AS u ON t.user_id = u.id; ``` ### Caching data from an external database or URL[​](#caching-data-from-an-external-database-or-url "Direct link to Caching data from an external database or URL") * Create a dictionary for the `pricing` table in your MySQL database using a composite key: ``` CREATE DICTIONARY product_pricing ( product_id UInt64, region String, price Float64 DEFAULT 0.0 ) PRIMARY KEY product_id, region_id SOURCE(MYSQL(NAME mysql_named_collection DB 'product_db' TABLE 'pricing')) LAYOUT(COMPLEX_KEY_HASHED()) LIFETIME(MIN 600 MAX 900); ``` This will periodically query MySQL and store the data in memory. * Create a dictionary for the `pricing` table in your PostgreSQL database using the `FLAT` layout: ``` CREATE DICTIONARY product_pricing ( product_id UInt64, price Float64 DEFAULT 0.0 ) PRIMARY KEY product_id SOURCE(POSTGRESQL(NAME psql_named_collection DB 'product_db' SCHEMA 'schema' TABLE 'pricing')) LAYOUT(FLAT()) LIFETIME(0); ``` Because `LIFETIME` is `0`, it has to be manually refreshed as follows: ``` SYSTEM RELOAD DICTIONARY product_pricing; ``` * Create a dictionary with `HTTP` as a source: ``` CREATE DICTIONARY currency_rates ( currency_code String, rate Float64 DEFAULT 1.0 ) PRIMARY KEY currency_code SOURCE(HTTP(URL 'https://example.com/currency_rates.csv' FORMAT CSV)) LAYOUT(COMPLEX_KEY_HASHED()) LIFETIME(100); ``` * Create a dictionary for the `users` table in a remote ClickHouse database using the `FLAT` layout: ``` CREATE DICTIONARY users_dictionary_remote ( id UInt64, username String, email String, country String ) PRIMARY KEY id SOURCE(CLICKHOUSE(NAME remote_clickhouse_named_collection DB 'default' TABLE 'users')) LAYOUT(FLAT()) LIFETIME(100); ``` --- # Set up Aiven for ClickHouse® data source integrations Connect your Aiven for ClickHouse® service with another Aiven-managed service or external data source to make your data available in the Aiven for ClickHouse service. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * You are familiar with the limitations listed in [About Aiven for ClickHouse® data service integration](/docs/products/clickhouse/concepts/data-integration-overview.md#supported-data-source-types). * You have an organization, a project, and an Aiven for ClickHouse service in Aiven. * You have access to the [Aiven Console](https://console.aiven.io/). ## Create Apache Kafka integrations[​](#create-apache-kafka-integrations "Direct link to Create Apache Kafka integrations") tip Learn about [managed databases integrations](/docs/products/clickhouse/concepts/data-integration-overview.md#managed-databases-integration). Make Apache Kafka data available in Aiven for ClickHouse using the Kafka engine: 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose the Aiven for ClickHouse service that you want to integrate with a data source. 2. In the service sidebar, click **Data** > **Integrations**. 3. On the **Integrations** page, go to the **Data sources** section and click **Apache Kafka**. The **Apache Kafka data source integration** wizard opens and displays a list of external data sources or Aiven-managed data services available for integration. If there are no data sources to integrate with, the wizard allows you to create them either by clicking **Create service** (for Aiven-managed sources) or **Add external endpoint** (for external sources). 4. In the **Apache Kafka data source integration** wizard: 1. Select a data source to integrate with, and click **Continue**. note If a data source to integrate with is not available on the list, click one of the following: * **Create service**: to create an Aiven-managed data service to integrate with * **Create external endpoint**: to make your external data source available for integration 2. Create tables where your Apache Kafka data will be available in Aiven for ClickHouse. Enter **Table name**, **Consumer group name**, **Topics**, **Data format**, and **Table columns**. Click **Save table details**. note * You can have up to 400 such tables for receiving and sending messages from multiple topics. * To query the tables, use the following statement: ``` SELECT * FROM APACHE_KAFKA_RESOURCE_NAME.APACHE_KAFKA_TABLE_NAME ``` * To set up **Data format**, see [Formats for Aiven for ClickHouse® - Aiven for Apache Kafka® data exchange](/docs/products/clickhouse/reference/supported-input-output-formats.md). * For more integration configuration options, see [Update Apache Kafka integration settings](/docs/products/clickhouse/howto/integrate-kafka.md#update-integration-settings). 3. Click **Enable integration** > **Close**. ## Create PostgreSQL integrations[​](#create-postgresql-integrations "Direct link to Create PostgreSQL integrations") tip Learn about [managed databases integrations](/docs/products/clickhouse/concepts/data-integration-overview.md#managed-databases-integration). Make PostgreSQL data available in Aiven for ClickHouse using the PostgreSQL engine: 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose the Aiven for ClickHouse service that you want to integrate with a data source. 2. In the service sidebar, click **Data** > **Integrations**. 3. On the **Integrations** page, go to the **Data sources** section and click **PostgreSQL**. The **PostgreSQL data source integration** wizard opens and displays a list of external data sources or Aiven-managed data services available for integration. If there are no data sources to integrate with, the wizard allows you to create them either by clicking **Create service** (for Aiven-managed sources) or **Add external endpoint** (for external sources). 4. In the **PostgreSQL data source integration** wizard: 1. Select a data source to integrate with, and click **Continue**. note If a data source to integrate with is not available on the list, click one of the following: * **Create service**: to create an Aiven-managed data service to integrate with * **Create external endpoint**: to make your external data source available for integration 2. Optionally, create databases where your PostgreSQL data will be available in Aiven for ClickHouse. Enter **Database name** and **Database schema**. tip You can query the created databases using the following statement: ``` SELECT * FROM POSTGRESQL_RESOURCE_NAME.POSTGRESQL_TABLE_NAME ``` note You can [create such integration databases](/docs/products/clickhouse/howto/integration-databases.md) any time later, for example, by finding your integration on the **Data** > **Integrations** page and clicking **Actions** > **Edit database**. 3. Click **Enable integration** > **Close**. ## Use managed-credentials integrations[​](#use-managed-credentials-integrations "Direct link to Use managed-credentials integrations") tip Learn about [managed credentials integrations](/docs/products/clickhouse/concepts/data-integration-overview.md#managed-credentials-integration). [Set up a managed-credentials integration](/docs/products/clickhouse/howto/data-service-integration.md#create-managed-credentials-integrations) and [create tables](/docs/products/clickhouse/howto/data-service-integration.md#create-tables) for the data to be made available through the integration. [Access your stored credentials](/docs/products/clickhouse/howto/data-service-integration.md#access-credentials-storage). ### Create managed-credentials integrations[​](#create-managed-credentials-integrations "Direct link to Create managed-credentials integrations") 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose the Aiven for ClickHouse service that you want to integrate with a data source. 2. In the service sidebar, click **Data** > **Integrations**. 3. On the **Integrations** page, go to the **Data sources** section and click **ClickHouse Credentials**. The **ClickHouse credentials integration** wizard opens and displays a list of external data sources or Aiven-managed data services available for integration. If there are no data sources to integrate with, the wizard allows you to create them either by clicking **Create service** (for Aiven-managed sources) or **Add external endpoint** (for external sources). 4. In the **ClickHouse credentials integration** wizard: 1. Select a data source to integrate with. note If a data source to integrate with is not available on the list, click one of the following: * **Create service**: to create an Aiven-managed data service to integrate with * **Create external endpoint**: to make your external data source available for integration 2. Click **Enable integration**. 3. Optionally, click **Test connection** > **Open in query editor** > **Execute**. Alternative You can test the connection any time later by going to your Aiven for ClickHouse service's **Data** > **Integrations** page, finding the credentials integration, and clicking **Actions** > **Test connection**. 4. Click **Close**. ### Create tables[​](#create-tables "Direct link to Create tables") Create tables using [table engines](/docs/products/clickhouse/reference/supported-table-engines.md), for example the PostgreSQL engine: ``` CREATE TABLE default.POSTGRESQL_TABLE_NAME ( `float_nullable` Nullable(Float32), `str` String, `int_id` Int32 ) ENGINE = PostgreSQL(postgres_credentials); ``` tip For details on how to use different table engines for integrations with external systems, see the [upstream ClickHouse documentation](https://clickhouse.com/docs/en/engines/table-engines/integrations). ### Access credentials storage[​](#access-credentials-storage "Direct link to Access credentials storage") Depending on the type of data source you are integrated with, you can access your credentials storage by passing your data source name in the following query: PostgreSQL data source ``` SELECT * FROM postgresql( `service_POSTGRESQL_SOURCE_NAME`, database='defaultdb', table='tables', schema='information_schema' ) ``` MySQL data source ``` SELECT * FROM mysql( `service_MYSQL_SOURCE_NAME`, database='mysql', table='slow_log' ) ``` Amazon S3 data source ``` SELECT * FROM s3( `endpoint_S3_SOURCE_NAME`, filename='*.csv', format='CSVWithNames') ``` warning When you try to run a managed credentials query with a typo, the query fails with an error message related to grants. ## View data source integrations[​](#view-data-source-integrations "Direct link to View data source integrations") 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose the Aiven for ClickHouse service whose integrations you want to view. 2. In the service sidebar, click **Data** > **Integrations**. ## Stop data source integrations[​](#stop-data-source-integrations "Direct link to Stop data source integrations") warning By terminating a data source integration, you disconnect from the data source, which erases all databases and configuration information from Aiven for ClickHouse. 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose the Aiven for ClickHouse service whose integrations you want to stop. 2. In the service sidebar, click **Data** > **Integrations**. 3. Find an integration to be stopped, and click **Actions** > **Disconnect**. Your integration is terminated and all the corresponding databases and configuration information are deleted. Related pages * [Aiven for ClickHouse® data service integration](/docs/products/clickhouse/concepts/data-integration-overview.md) * [Managed credentials integration](/docs/products/clickhouse/concepts/data-integration-overview.md#managed-credentials-integration) * [Managed databases integration](/docs/products/clickhouse/concepts/data-integration-overview.md#managed-databases-integration) * [Create and manage Aiven for ClickHouse® integration databases](/docs/products/clickhouse/howto/integration-databases.md) --- # Scale disk storage automatically for your Aiven for ClickHouse® service Automatically increase the disk storage of your Aiven for ClickHouse® service when it's running out of space, instead of resizing it manually. Use the Aiven Autoscaler to automatically increase the storage capacity of a service disk when it's running out of space. Disk autoscaler only increases storage, it doesn't scale storage down. ## Why use disk autoscaling[​](#why-use-disk-autoscaling "Direct link to Why use disk autoscaling") * **Cost efficiency**: Start with a regular-sized disk and let Aiven scale it up only when needed, without the risk of running out of disk space. * **Resiliency**: Avoid a service becoming non-functional because it ran out of disk space, including during unexpected spikes in demand. ## How it works[​](#how-it-works "Direct link to How it works") 1. You create an autoscaler integration endpoint in your project, setting the maximum total disk size to allow. 2. You enable an autoscaler integration for your service using that endpoint. 3. Aiven monitors the disk space usage of your service. 4. When disk usage reaches the threshold for your service type, Aiven increases the available storage by at least 10%, using the current used space as a baseline. note The exact increase depends on the service type and cloud provider. Some providers enforce a minimum increase of 10 GB. Autoscale thresholds per service type The threshold that triggers disk autoscaling is a percentage of the available disk storage capacity: * Aiven for OpenSearch®: 75% of the available disk storage capacity * All other supported service types: 85% of the available disk storage capacity 5. The disk increase is recorded in the project event log, and you receive a notification about the added disk space. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * **Maximum storage**: The maximum storage that the autoscaler can allocate for your service is limited by both the maximum disk size set on the autoscaler endpoint and the maximum disk storage supported for your service plan. * **Timing**: Autoscaling takes a moment to complete. In the meantime, the service disk might fill up and the service might enter read-only mode until autoscaling finishes, unless the autoscaler's disk capacity limit is reached. * **Maintenance updates**: Autoscaling works only on fully running services and can't happen during a maintenance update. * **Manual changes**: Changing disk space manually can delay an autoscaling event. * **Terraform**: Don't manage disk space with the Aiven Terraform Provider on a service that uses the autoscaler, to avoid conflicts between the two. * **Performance**: Disk added through autoscaling is slower than the original disk until the next maintenance update applies. This might affect I/O-intensive workloads. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven organization, project, and service that's up and running * The operator role for the organization, project, and service * Dynamic disk sizing support on your service plan and cloud region * One of the following to manage the autoscaler: * [Aiven Console](https://console.aiven.io/) * [Aiven API](https://api.aiven.io/doc/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) ### Enable disk autoscaling[​](#enable-disk-autoscaling "Direct link to Enable disk autoscaling") To enable disk autoscaling, create an autoscaler integration endpoint, then enable an autoscaler integration on your service using that endpoint. * Console * API * CLI * Terraform Create an autoscaler endpoint: 1. Log in to [Aiven Console](https://console.aiven.io/), and go to your organization and project. 2. On the left sidebar, click **Integration endpoints**. 3. Click **Aiven Autoscaler** > **Add new endpoint**. 4. Set the endpoint name and the maximum total disk storage in GB, and click **Add endpoint**. Enable the autoscaler on a service: 1. On the left sidebar, click **Services**, and open your service. 2. On the left sidebar, click **Integrations**. 3. In **Endpoint integrations**, click **Aiven Autoscaler**. 4. Click the endpoint you created, and click **Enable**. 1) Call [ServiceIntegrationEndpointCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointCreate) to create an autoscaler integration endpoint on your project: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "endpoint_name": "ENDPOINT_NAME", "endpoint_type": "autoscaler", "user_config": { "autoscaling": [ { "type": "autoscale_disk", "cap_gb": 300 } ] } }' ``` 2) Call [ServiceIntegrationCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationCreate) to enable the autoscaler integration on your service, using the endpoint ID from the previous response: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "dest_endpoint_id": "ENDPOINT_ID", "integration_type": "autoscaler", "source_project": "PROJECT_NAME", "source_service": "SERVICE_NAME" }' ``` 1. Create an autoscaler integration endpoint using [avn service integration-endpoint-create](/docs/tools/cli.md): ``` avn service integration-endpoint-create \ --project PROJECT_NAME \ --endpoint-name ENDPOINT_NAME \ --endpoint-type autoscaler \ --user-config-json '{"autoscaling": [{"type": "autoscale_disk", "cap_gb": 300}]}' ``` 2. Find the ID of the new endpoint: ``` avn service integration-endpoint-list --project PROJECT_NAME ``` 3. Enable the autoscaler integration on your service, using the endpoint ID from the previous step: ``` avn service integration-create \ --dest-service SERVICE_NAME \ --integration-type autoscaler \ --source-endpoint-id ENDPOINT_ID ``` Use the [`aiven_service_integration_endpoint`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration_endpoint) and [`aiven_service_integration`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration) resources: ``` resource "aiven_service_integration_endpoint" "autoscaler_endpoint" { project = "PROJECT_NAME" endpoint_name = "ENDPOINT_NAME" endpoint_type = "autoscaler" autoscaler_user_config { autoscaling { type = "autoscale_disk" cap_gb = 300 } } } resource "aiven_service_integration" "autoscaler_integration" { project = "PROJECT_NAME" integration_type = "autoscaler" source_service_name = "SERVICE_NAME" destination_endpoint_id = aiven_service_integration_endpoint.autoscaler_endpoint.id } ``` See the [disk autoscaler guide](https://registry.terraform.io/providers/aiven/aiven/latest/docs/guides/disk-autoscaler) for more details. ### Change the maximum disk space for autoscaling[​](#change-the-maximum-disk-space-for-autoscaling "Direct link to Change the maximum disk space for autoscaling") After you enable disk autoscaling, you can update the maximum total disk size at any time. * Console * API * CLI * Terraform 1. On the left sidebar, click **Integration endpoints**. 2. Click **Aiven Autoscaler**. 3. Find your endpoint, click **Actions**, and click the option to edit it. 4. Set a new maximum disk storage value, and save your changes. Call [ServiceIntegrationEndpointUpdate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointUpdate) with the new `cap_gb` value: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint/ENDPOINT_ID \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "user_config": { "autoscaling": [ { "type": "autoscale_disk", "cap_gb": 500 } ] } }' ``` ``` avn service integration-endpoint-update ENDPOINT_ID \ --user-config-json '{"autoscaling": [{"type": "autoscale_disk", "cap_gb": 500}]}' ``` Update the `cap_gb` value in the `autoscaling` block of your [`aiven_service_integration_endpoint`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration_endpoint) resource, then apply the change. ### Turn off disk autoscaling[​](#turn-off-disk-autoscaling "Direct link to Turn off disk autoscaling") To turn off disk autoscaling, remove the autoscaler integration from your service. You can also delete the integration endpoint if you no longer need it. * Console * API * CLI Disconnect the service from the autoscaler: 1. On the left sidebar, click **Services**, and open your service. 2. On the left sidebar, click **Integrations**. 3. In **Endpoint integrations**, find **Aiven Autoscaler**, click **Actions**, and click the option to disconnect it. Delete the autoscaler endpoint, if you no longer need it: 1. On the left sidebar, click **Integration endpoints**. 2. Click **Aiven Autoscaler**. 3. Find your endpoint, and delete it. 1) Call [ServiceIntegrationDelete](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationDelete) to remove the autoscaler integration from your service: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration/INTEGRATION_ID \ --header "Authorization: Bearer TOKEN" ``` 2) Call [ServiceIntegrationEndpointDelete](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointDelete) to delete the autoscaler integration endpoint, if you no longer need it: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint/ENDPOINT_ID \ --header "Authorization: Bearer TOKEN" ``` 1. Find the ID of the integration to remove: ``` avn service integration-list SERVICE_NAME ``` 2. Remove the autoscaler integration from your service: ``` avn service integration-delete INTEGRATION_ID ``` 3. Find the ID of the integration endpoint to delete, if you no longer need it: ``` avn service integration-endpoint-list --project PROJECT_NAME ``` 4. Delete the autoscaler integration endpoint: ``` avn service integration-endpoint-delete ENDPOINT_ID ``` Related pages * [Scale disk storage manually](/docs/products/clickhouse/howto/scale-disk-storage.md) * [Change the service plan](/docs/products/clickhouse/howto/change-service-plan.md) * [Tiered storage in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/clickhouse-tiered-storage.md) --- # Enable tiered storage in Aiven for ClickHouse® Enable the [tiered storage feature](/docs/products/clickhouse/concepts/clickhouse-tiered-storage.md) on a table in your Aiven for ClickHouse® service. Before you enable tiered storage, review the [limitations](/docs/products/clickhouse/concepts/clickhouse-tiered-storage.md#tiered-storage-limitations). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * You have at least one Aiven for ClickHouse service. * Depending on how to activate tiered storage, you need: * [Aiven Console](https://console.aiven.io) or * SQL and an SQL client (for example, the [ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md)). * All [maintenance updates](/docs/products/clickhouse/howto/maintenance-updates.md) are applied on your service. ## Activate tiered storage on a table[​](#activate-tiered-storage-on-a-table "Direct link to Activate tiered storage on a table") You can enable tiered storage both on new tables and on existing ones. For that purpose, you can use either CLI or the [Aiven Console](https://console.aiven.io). * Console * CLI 1. Log in to the [Aiven Console](https://console.aiven.io) and choose your Aiven for ClickHouse service. 2. In the service sidebar, click **Data** > **Databases and tables**. 3. In the **Databases and tables** view, find a table on which to activate tiered storage, and click **Actions** > **Activate tiered storage** > **Activate**. 1) [Connect to your Aiven for ClickHouse service](/docs/products/clickhouse/howto/list-connect-to-service.md) using, for example, the ClickHouse client. 2) To activate the tiered storage feature on a specific table, set `storage_policy` to `tiered` on this table by executing the following SQL statement: ``` ALTER TABLE database-name.table-name MODIFY SETTING storage_policy = 'tiered' ``` Tiered storage is activated on your table and data in this table is now distributed between two tiers: Network-attached block storage and object storage. ## What's next[​](#whats-next "Direct link to What's next") * [Configure data retention thresholds for tiered storage](/docs/products/clickhouse/howto/configure-tiered-storage.md) * [Check data volume distribution between different disks](/docs/products/clickhouse/howto/check-data-tiered-storage.md) Related pages * [About tiered storage in Aiven for ClickHouse](/docs/products/clickhouse/concepts/clickhouse-tiered-storage.md) * [Transfer data between network-attached block storage and object storage](/docs/products/clickhouse/howto/transfer-data-tiered-storage.md) --- # Fetch query statistics for Aiven for ClickHouse® Usually, query statistics in ClickHouse can be obtained using the `system.query_log` table, which stores statistics of each executed query, including memory usage and duration. In Aiven for ClickHouse®, the `system.query_log` table is currently not accessible for the purpose of obtaining query statistics. To fetch query statistics in Aiven for ClickHouse, you can use either Aiven Console or Aiven API. ## Use Aiven Console[​](#use-aiven-console "Direct link to Use Aiven Console") 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose your Aiven for ClickHouse service. 2. In the service sidebar, click **Observe** > **Query statistics**. 3. View the query statistics in the dashboard. ## Use Aiven API[​](#use-aiven-api "Direct link to Use Aiven API") To access query statistics in Aiven for ClickHouse with Aiven API, use the [ServiceClickHouseQueryStats endpoint](https://api.aiven.io/doc/#tag/Service:_ClickHouse/operation/ServiceClickHouseQueryStats). ``` GET /project//service//clickhouse/query/stats ``` Related pages Learn more on Aiven API in the [Aiven API overview](/docs/tools/api.md). --- # Fork your Aiven for ClickHouse® service Fork your Aiven for ClickHouse® service to create an independent copy for testing, debugging, or development without affecting the original service. Fork an Aiven service to create a complete copy of it from its latest backup. Forked services are independent and don't share resources with or increase the load on the original service. Common use cases for forking include: * Creating a snapshot to analyze an issue. * Creating a development copy of your production environment. * Testing upgrades before applying them to production services. * Creating an instance in a different cloud provider, region, or with a different plan. * Renaming a service. During the forking process, the fork might initially have only one node while backups are being taken. The other nodes appear after the backup process is complete. When you fork a service, its configuration, databases, tables, and access entities are copied to the new service. ## Limitations[​](#limitations "Direct link to Limitations") * You can only fork services that have at least one [backup](/docs/products/clickhouse/concepts/disaster-recovery.md#service-backup). * Service integrations are not copied to the fork. * Cross-project forking is supported only within the same organization. * Point-in-time recovery is not supported. You can restore only to a daily backup state. * You cannot fork Aiven for ClickHouse services to a fewer number of nodes. Reducing the number of nodes is only possible by [changing the service plan](/docs/products/clickhouse/howto/change-service-plan.md) from **Business** to **Startup** on a running service. ## Fork a service[​](#fork-a-service "Direct link to Fork a service") * Console * CLI * API * Terraform 1. In your service, in the **Backups** section, click **Backup management**. 2. Click **Fork & restore**. 3. Choose the backup to fork from. 4. Enter a name, and select the cloud and plan. 5. Click **Create fork**. Use the [create service command](/docs/tools/cli/service-cli.md#avn-cli-service-create) with: * `--service-to-fork-from`: the name of the service to use as the source. * `--project-to-fork-from`: to fork a service in a different project, set this to the project name the source service is in. Use the [`ServiceCreate` endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) and in the `user_config` property set: * `service_to_fork_from`: the name of the source service. * `project_to_fork_from`: to fork a service in a different project, set this to the name of the project the source service is in. Use the `service_to_fork_from` attribute in the user config of your service resource. To fork a service in a different project, set the `project_to_fork_from` attribute. More information on the service resources and their configuration options is available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Once the new fork service is running, you can set up your application's connection settings to point to this new fork service. Related pages * [Schedule Aiven for ClickHouse® backups](/docs/products/clickhouse/howto/configure-backup.md) * [Restore an Aiven for ClickHouse® backup](/docs/products/clickhouse/howto/restore-backup.md) * [Change the service plan](/docs/products/clickhouse/howto/change-service-plan.md) * [Rename your Aiven for ClickHouse® service](/docs/products/clickhouse/howto/rename-service.md) --- # Connect Apache Kafka® to Aiven for ClickHouse® Integrate Aiven for ClickHouse® with either Aiven for Apache Kafka® service located in the same project, or an external Apache Kafka endpoint. For a different use case Need to deliver data from Apache Kafka® topics to a ClickHouse database for efficient querying and analysis? [Use a ClickHouse sink connector](/docs/products/kafka/kafka-connect/howto/clickhouse-sink-connector.md). A single Aiven for ClickHouse instance can connect to multiple Kafka clusters with different authentication mechanism and credentials. Behind the scenes, the integration between Aiven for ClickHouse and Apache Kafka services relies on [ClickHouse Kafka Engine](https://clickhouse.com/docs/en/engines/table-engines/integrations/kafka/). note Aiven for ClickHouse service integrations are available for [Startup plans and higher](https://aiven.io/pricing?product=clickhouse). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Services to integrate using Startup plans and higher: * Aiven for ClickHouse service * Aiven for Apache Kafka service or a self-hosted Apache Kafka service tip If you use the self-hosted Apache Kafka service, [configure an external Apache Kafka endpoint](/docs/products/kafka/howto/integrate-external-kafka-cluster.md). * At least one topic in the Apache Kafka service ### Tools[​](#tools "Direct link to Tools") * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * SQL client ### Variables[​](#variables "Direct link to Variables") Variables used to set up and configure the integration: | Variable | Description | | ------------------------- | ---------------------------------------------------------------------------------------------------------- | | `CLICKHOUSE_SERVICE_NAME` | Name of your Aiven for ClickHouse service. | | `KAFKA_SERVICE_NAME` | Name of the Apache Kafka service you use for the integration. | | `PROJECT` | Name of Aiven project where your services are located. | | `CONNECTOR_TABLE_NAME` | Name of the Kafka engine virtual table that is used as a connector. | | `DATA_FORMAT` | Input/output data format in which data is accepted into Aiven for ClickHouse. See [Reference](#reference). | | `CONSUMER_GROUP_NAME` | Name of the consumer group. Each message is delivered once per consumer group. | ## Create an integration[​](#create-an-integration "Direct link to Create an integration") To connect Aiven for ClickHouse and Aiven for Apache Kafka by enabling a data service integration, see [Create data service integrations](/docs/products/clickhouse/howto/data-service-integration.md#create-apache-kafka-integrations). When you create the integration, a database is automatically added in your Aiven for ClickHouse. Its name is `service_KAFKA_SERVICE_NAME`, where `KAFKA_SERVICE_NAME` is the name of your Apache Kafka service. In this database, you create virtual connector tables, which is also a part of the [integration creation in the Aiven Console](/docs/products/clickhouse/howto/data-service-integration.md#create-apache-kafka-integrations). You can have up to 400 such tables for receiving and sending messages from multiple topics. ## Update integration settings[​](#update-integration-settings "Direct link to Update integration settings") Upon creating the integration and configuring your tables, you can edit both [mandatory integration settings](/docs/products/clickhouse/howto/integrate-kafka.md#mandatory-integration-settings) and [optional integration settings](/docs/products/clickhouse/howto/integrate-kafka.md#optional-integration-settings) on a table level either in the [Aiven Console](https://console.aiven.io/) or the [Aiven CLI](/docs/tools/cli.md). * Console * CLI 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose the Aiven for ClickHouse service that includes a table to edit. 2. In the service sidebar, click **Data** > **Databases and tables**. 3. Find a database including the table to be edited, and expand the database using to display tables inside it. 4. Find the table and click: * For [mandatory settings](/docs/products/clickhouse/howto/integrate-kafka.md#mandatory-integration-settings): **Actions** > **Edit table**. * For [optional settings](/docs/products/clickhouse/howto/integrate-kafka.md#optional-integration-settings): **Actions** > **Advanced configuration**. 5. In the displayed window, update existing settings and / or add new ones. Save your changes. 1) Get your service integration ID by requesting the full list of integrations. Replace `PROJECT`, `CLICKHOUSE_SERVICE_NAME` and `KAFKA_SERVICE_NAME` with the names of your services: ``` avn service integration-list \ --project PROJECT_NAME \ CLICKHOUSE_SERVICE_NAME | grep KAFKA_SERVICE_NAME ``` 2) Run [avn service integration-update](/docs/tools/cli/service/integration.md#avn_service_integration_update) with the service integration id and your integration settings. Replace `SERVICE_INTEGRATION_ID`, `CONNECTOR_TABLE_NAME`, `DATA_FORMAT` and `CONSUMER_NAME` with your values: ``` avn service integration-update SERVICE_INTEGRATION_ID \ --project PROJECT_NAME \ --user-config-json '{ "tables": [ { "name": "CONNECTOR_TABLE_NAME", "columns": [ {"name": "id", "type": "UInt64"}, {"name": "name", "type": "String"} ], "topics": [{"name": "topic1"}, {"name": "topic2"}], "data_format": "DATA_FORMAT", "group_name": "CONSUMER_NAME", "auto_offset_reset": "earliest" } ] }' ``` ## Read and store data[​](#read-and-store-data "Direct link to Read and store data") In Aiven for ClickHouse you can consume messages by running SELECT command. Replace `KAFKA_SERVICE_NAME` and `CONNECTOR_TABLE_NAME` with your values and run: ``` SELECT * FROM service_KAFKA_SERVICE_NAME.CONNECTOR_TABLE_NAME ``` However, the messages are only read once (per consumer group). If you want to store the messages for later, you can send them into a separate ClickHouse table with the help of a materialized view. For example, run to creating a destination table: ``` CREATE TABLE destination (id UInt64, name String) ENGINE = ReplicatedMergeTree() ORDER BY id; ``` Add a materialised view to bring the data from the connector: ``` CREATE MATERIALIZED VIEW materialised_view TO destination AS SELECT * FROM service_KAFKA_SERVICE_NAME.CONNECTOR_TABLE_NAME; ``` Now the messages consumed from the Apache Kafka topic will be read automatically and sent into the destination table directly. For more information on materialized views, see [Create materialized views in ClickHouse®](/docs/products/clickhouse/howto/materialized-views.md). note ClickHouse is strict about allowed symbols in database and table names. You can use backticks around the names when running ClickHouse requests, particularly in the cases when the name contains dashes. ## Write data back to the topic[​](#write-data-back-to-the-topic "Direct link to Write data back to the topic") You can also bring the entries from ClickHouse table into the Apache Kafka topic. Replace `KAFKA_SERVICE_NAME` and `CONNECTOR_TABLE_NAME` with your values: ``` INSERT INTO service_KAFKA_SERVICE_NAME.CONNECTOR_TABLE_NAME(id, name) VALUES (1, 'Michelangelo') ``` warning Writing to more than one topic is not supported. ## Reference[​](#reference "Direct link to Reference") ### Mandatory integration settings[​](#mandatory-integration-settings "Direct link to Mandatory integration settings") Click to see the list * `name` - name of the connector table * `columns` - array of columns with names and types * `topics` - array of topics to pull data from * `data_format` - format for input data ([see supported formats](/docs/products/clickhouse/reference/supported-input-output-formats.md)) * `group_name` - consumer group name to be created on your behalf ### Optional integration settings[​](#optional-integration-settings "Direct link to Optional integration settings") Click to see the list | Name | Type | Description | Default | Example | Allowed values / Range | | --------------------------------------- | ------- | -------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------ | ------------- | ------------------------------------------------------------------------------------------------------------------------- | | `auto_offset_reset` | string | Action to take when there is no initial offset in offset store or the desired offset is out of range | `earliest` | `latest` | `smallest`, `earliest`, `beginning`, `largest`, `latest`, `end` | | `date_time_input_format` | string | Method to read `DateTime` from text input formats | `basic` | `best_effort` | `basic`, `best_effort`, `best_effort_us` | | `handle_error_mode` | string | How to handle errors for Kafka engine | `default` | `stream` | `default`, `stream` | | `max_block_size` | integer | Number of rows collected by polls for flushing data from Kafka | `0` | `100000` | `0` - `1_000_000_000` | | `max_rows_per_message` | integer | Maximum number of rows produced in one Kafka message for row-based formats | `1` | `100000` | `1` - `1_000_000_000` | | `num_consumers` | integer | Number of consumers per table per replica | `1` | `4` | `1` - `10` | | `poll_max_batch_size` | integer | Maximum amount of messages to be polled in a single Kafka poll | `0` | `10000` | `0` - `1_000_000_000` | | `poll_max_timeout_ms` | integer | Timeout in milliseconds for a single poll from Kafka. Defaults to `stream_flush_interval_ms` (`500` ms). | `0` | `1000` | `0` - `30_000` | | `skip_broken_messages` | integer | Skip at least this number of broken messages from Kafka topic per block | `0` | `10000` | `0` - `1_000_000_000` | | `thread_per_consumer` | boolean | Provide an independent thread for each consumer. All consumers run in the same thread by default. | `false` | `true` | `true`, `false` | | `producer_batch_size` | integer | Max size in bytes of a batch of messages sent to Kafka. If exceeded, the batch is sent. | `1000000` | `1000000` | `0` - `2_147_483_647` | | `producer_batch_num_messages` | integer | Max number of messages in a batch sent to Kafka. If exceeded, the batch is sent. | `10000` | `10000` | `1` - `1_000_000` | | `producer_compression_codec` | string | Compression codec to use for Kafka producer | `none` | `zstd` | `none`, `gzip`, `lz4`, `snappy`, `zstd` | | `producer_compression_level` | integer | Compression level for Kafka producer.
The range depends on `producer_compression_codec`. | Codec-dependent value.
`-1` if `producer_compression_codec` set to `none` | `5` | Codec-dependent ranges: - `gzip`: `0` - `9`
- `lz4`: `0` - `12`
- `snappy`: only `0`
- `none`: `-1` - `12` | | `producer_linger_ms` | integer | Time in ms to wait for additional messages before sending a batch. If exceeded, the batch is sent. | `5` | `5` | `0` - `900_000` | | `producer_queue_buffering_max_messages` | integer | Max number of messages to buffer before sending. Max messages in producer queue. | `100000` | `100000` | `0` - `2_147_483_647` | | `producer_queue_buffering_max_kbytes` | integer | Max size of buffer in kilobytes before sending. Max size of producer queue in kB. | `1048576` | `1048576` | `0` - `2_147_483_647` | | `producer_request_required_acks` | integer | Number of acknowledgments required from Kafka brokers for a message to be considered successful | `-1` | `1` | `-1` - `1000` | ### Formats supporting the integration[​](#formats-supporting-the-integration "Direct link to Formats supporting the integration") When connecting ClickHouse® to Kafka® using Aiven integrations, data exchange requires using specific formats. Check the supported formats for input and output data in [Formats for ClickHouse®-Kafka® data exchange](/docs/products/clickhouse/reference/supported-input-output-formats.md). --- # Connect PostgreSQL® to Aiven for ClickHouse® You can integrate Aiven for ClickHouse® with either *Aiven for PostgreSQL* service located in the same project, or *an external PostgreSQL endpoint*. Behind the scenes the integration between Aiven for ClickHouse and PostgreSQL relies on [ClickHouse PostgreSQL Engine](https://clickhouse.com/docs/en/engines/table-engines/integrations/postgresql). note Aiven for ClickHouse service integrations are available for Startup plans and higher. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") To connect Aiven for ClickHouse to PostgreSQL, you need the following: * Aiven for ClickHouse service * Aiven for PostgreSQL service or a self-hosted PostgreSQL service * You have at least one table in your PostgreSQL service. tip If you use the self-hosted PostgreSQL service, an external PostgreSQL endpoint should be configured in **Integration endpoints**. ## Variables[​](#variables "Direct link to Variables") The following variables will be used later in the code snippets: | Variable | Description | | ------------------------- | ----------------------------------------------------------- | | `CLICKHOUSE_SERVICE_NAME` | Name of your Aiven for ClickHouse service. | | `PG_SERVICE_NAME` | Name of the PostgreSQL service you use for the integration. | | `PG_DATABASE` | Name of PostgreSQL database you're integrating. | | `PG_SCHEMA` | Name of PostgreSQL schema. | | `PG_TABLE` | Name of the PostgreSQL table you use for the integration. | ## Create an integration[​](#create-an-integration "Direct link to Create an integration") To connect Aiven for ClickHouse and Aiven for PostgreSQL by enabling a data service integration, see [Create data service integrations](/docs/products/clickhouse/howto/data-service-integration.md#create-postgresql-integrations). The newly created database name has the following format: `service_PG_SERVICE_NAME_PG_DATABASE_PG_SCHEMA`, for example, `service_myPGService_myDatabase_mySchema`. note When connecting to an Aiven for PostgreSQL, we connect as the main service user of that service, which has access to all the PostgreSQL tables. SELECT and INSERT privileges are granted to the main service user (`avnadmin`). It is up to the main service user to grant access to other users. Read more [how to grant privileges](/docs/products/clickhouse/howto/manage-users-roles.md). ## Update PostgreSQL integration settings[​](#update-postgresql-integration-settings "Direct link to Update PostgreSQL integration settings") When connecting to a PostgreSQL service, ClickHouse needs to know the name of the PostgreSQL schema and database to access. By default these settings are set to the `public` schema in the `defaultdb`. However, you can update these values by following next steps. note These configurations can be set only with the CLI command [avn service integration-update](/docs/tools/cli/service/integration.md#avn_service_integration_update). 1. Get *the service integration id* by requesting the full list of integrations. Replace `CLICKHOUSE_SERVICE_NAME` and `PG_SERVICE_NAME` with the names of your services: ``` avn service integration-list --project PROJECT_NAME CLICKHOUSE_SERVICE_NAME | grep PG_SERVICE_NAME ``` 2. Update the configuration settings using the service integration id retrieved in the previous step and your integration settings. Replace `SERVICE_INTEGRATION_ID`, `PG_DATABASE` and `PG_SCHEMA` with your values, you can add more than one combination of database/schema in the object `databases`: ``` avn service integration-update --project PROJECT_NAME SERVICE_INTEGRATION_ID \ --user-config-json '{ "databases":[{"database":"PG_DATABASE","schema":"PG_SCHEMA"}] }' ``` ## Read and store data[​](#read-and-store-data "Direct link to Read and store data") In Aiven for ClickHouse you can read data by running SELECT command. Replace `PG_SERVICE_NAME`, `PG_DATABASE`, `PG_SCHEMA` and `PG_TABLE` with your values and run: ``` SELECT * FROM service_PG_SERVICE_NAME_PG_DATABASE_PG_SCHEMA.PG_TABLE ``` note ClickHouse is strict about allowed symbols in database and table names. You can use backticks around the names when running ClickHouse requests, particularly in the cases when the name contains dashes. For example, ``SELECT * FROM `service_your-kafka-service`.table``. ## Write data to PostgreSQL table[​](#write-data-to-postgresql-table "Direct link to Write data to PostgreSQL table") You can also insert rows from the ClickHouse table into the PostgreSQL table. Replace `PG_SERVICE_NAME`, `PG_DATABASE`, `PG_SCHEMA` and `PG_TABLE` with your values: ``` INSERT INTO service_PG_SERVICE_NAME_PG_DATABASE_PG_SCHEMA.PG_TABLE(id, name) VALUES (1, 'Michelangelo') ``` --- # Create and manage Aiven for ClickHouse® integration databases Create and manage integration databases in Aiven for ClickHouse® to query data from integrated services: * Aiven for Apache Kafka® * Aiven for PostgreSQL® ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to the [Aiven Console](https://console.aiven.io/) * An Aiven for ClickHouse service * A [service integration](/docs/products/clickhouse/howto/data-service-integration.md) with Aiven for Kafka or Aiven for PostgreSQL ## Create integration databases[​](#create-integration-databases "Direct link to Create integration databases") note Your Aiven for ClickHouse service can support up to 400 databases simultaneously. Depending on your data source, select **PostgreSQL** or **Apache Kafka**. * PostgreSQL * Apache Kafka 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your project. 2. On the **Services** page, select an Aiven for ClickHouse service. 3. In the sidebar, click **Data** > **Databases and tables**. 4. Click **Create database** > **PostgreSQL integration database**. 5. In the **Create PostgreSQL integration database** wizard: 1. Select the data source service, and click **Continue**. 2. Enter **Database name** and **Schema name**, and click **Save**. 1) Log in to the [Aiven Console](https://console.aiven.io/), and go to your project. 2) On the **Services** page, select an Aiven for ClickHouse service. 3) In the sidebar, click **Data** > **Integrations**. 4) On the **Integrations** page, find your data source integration and click **Actions** > **Create database**. 5) In the **Create Kafka integration database** wizard: 1. Select the data source service, and click **Continue**. 2. Configure the integration table: * Enter a table name. * Enter a consumer group name. * Select topics. * Select a data format. * Enter table columns. 3. Click **Save table details** > **Save**. ## View integration databases or tables[​](#view-integration-databases-or-tables "Direct link to View integration databases or tables") 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your project. 2. On the **Services** page, select an Aiven for ClickHouse service. 3. In the sidebar, click **Data** > **Databases and tables**. 4. Expand the integration database using . 5. Click **Actions** > **View details** on the table. ## Edit integration databases or tables[​](#edit-integration-databases-or-tables "Direct link to Edit integration databases or tables") Add or delete integration databases or tables. Update table details. note You cannot edit tables inside PostgreSQL integration databases. Depending on what you intend to edit, select **Database** or **Table**. * Database * Table 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your project. 2. On the **Services** page, select an Aiven for ClickHouse service with the database. 3. In the sidebar, click **Data** > **Databases and tables**. 4. Click **Actions** > **Edit database** on the integration database. 5. Use or **Delete** to add or remove integration databases, and click **Save**. Alternative for PostgreSQL On the **Data** > **Integrations** page, find the integration and click **Actions** > **Edit database**. ### Add or delete tables[​](#add-or-delete-tables "Direct link to Add or delete tables") 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your project. 2. On the **Services** page, select an Aiven for ClickHouse service with the database. 3. In the sidebar, click **Data** > **Databases and tables**. 4. Click **Actions** > **Edit database** on the database. 5. Use **Add table** or **Delete** to add or remove tables, and click **Save**. ### Update table details[​](#update-table-details "Direct link to Update table details") 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your project. 2. On the **Services** page, select an Aiven for ClickHouse service with the table. 3. In the sidebar, click **Data** > **Databases and tables**. 4. Expand the database using . 5. Click **Actions** > **Edit table** on the table. 6. Update the table details, and click **Save**. ## Delete integration databases or tables[​](#delete-integration-databases-or-tables "Direct link to Delete integration databases or tables") Depending on what you intend to delete, select **Database** or **Table**. * Database * Table 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your project. 2. On the **Services** page, select an Aiven for ClickHouse service with the database. 3. In the sidebar, click **Data** > **Databases and tables**. 4. Click **Actions** > **Delete database** on the integration database. 5. Review the impact in the confirmation dialog, and click **Confirm** to delete the database and its tables. 1) Log in to the [Aiven Console](https://console.aiven.io/), and go to your project. 2) On the **Services** page, select an Aiven for ClickHouse service with the table. 3) In the sidebar, click **Data** > **Databases and tables**. 4) Expand the integration database using . 5) Click **Actions** > **Delete table** on the table. 6) Review the impact in the confirmation dialog, and click **Confirm** to delete the table. Related pages * [Manage Aiven for ClickHouse® data service integrations](/docs/products/clickhouse/howto/data-service-integration.md) * [Aiven for ClickHouse® service integrations](/docs/products/clickhouse/concepts/data-integration-overview.md) --- # Connect to Aiven for ClickHouse® Connect to the Aiven for ClickHouse® service using various programming languages or tools. ## [Interfaces and drivers](/docs/products/clickhouse/reference/supported-interfaces-drivers.md) [Find out what technologies and tools you can use to interact with Aiven for ClickHouse®.](/docs/products/clickhouse/reference/supported-interfaces-drivers.md) ## [ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md) [It's recommended to connect to a ClickHouse® cluster with the ClickHouse® client.](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md) ## [Go](/docs/products/clickhouse/howto/connect-with-go.md) [To connect to your Aiven for ClickHouse® service with Go, you can use](/docs/products/clickhouse/howto/connect-with-go.md) ## [Python](/docs/products/clickhouse/howto/connect-with-python.md) [To connect to your Aiven for ClickHouse® service with Python, you can](/docs/products/clickhouse/howto/connect-with-python.md) ## [Node.js](/docs/products/clickhouse/howto/connect-with-nodejs.md) [Learn how to connect to your Aiven for ClickHouse® service with Node.js](/docs/products/clickhouse/howto/connect-with-nodejs.md) ## [PHP](/docs/products/clickhouse/howto/connect-with-php.md) [Learn how to connect to your Aiven for ClickHouse® service with PHP using the PHP ClickHouse client and the HTTPS port.](/docs/products/clickhouse/howto/connect-with-php.md) ## [Java](/docs/products/clickhouse/howto/connect-with-java.md) [Learn how to connect to your Aiven for ClickHouse® service with Java](/docs/products/clickhouse/howto/connect-with-java.md) --- # Manage local cache for remote files in Aiven for ClickHouse®'s tiered storage Aiven for ClickHouse®'s tiered storage features local on-disk cache for remote files for improved query performance and reduced latency. To manage data, Aiven for ClickHouse's tiered storage uses local storage and remote storage. When remote storage is used, Aiven for ClickHouse leverages a local on-disk cache to avoid repeated remote fetches. ## How it works[​](#how-it-works "Direct link to How it works") When a query requires parts of a table stored in the remote tier, Aiven for ClickHouse fetches the required parts from the remote storage. The fetched parts are automatically stored in a local cache directory on the disk to avoid repeated downloads for subsequent queries. For future queries, Aiven for ClickHouse checks the local cache first: * If the data is found in the cache, it is read directly from the local disk. * If the data is not found in the cache, it is fetched from the remote storage and stored in the local cache. Local on-disk cache for remote files is enabled by default for Aiven for ClickHouse's tiered storage. You can [disable the cache](/docs/products/clickhouse/howto/local-cache-tiered-storage.md#disable-the-cache) or [drop it](/docs/products/clickhouse/howto/local-cache-tiered-storage.md#free-up-space) to free up the space it occupies. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * At least one Aiven for ClickHouse service using tiered storage * Command line tool ([ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md)) installed ## Disable the cache[​](#disable-the-cache "Direct link to Disable the cache") To disable the local cache for a query, set the `enable_filesystem_cache` setting for the query to `false`. You can achieve this by appending `SETTINGS enable_filesystem_cache = false` to the end of your query using an SQL client (for example, the [ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md)): ``` SELECT 1 SETTINGS enable_filesystem_cache = false; ``` ## Free up space[​](#free-up-space "Direct link to Free up space") To drop the local cache and free up the used space, use the following cache command: ``` SYSTEM DROP FILESYSTEM CACHE 'remote_cache' ``` Related pages * [About tiered storage in Aiven for ClickHouse](/docs/products/clickhouse/concepts/clickhouse-tiered-storage.md) * [Check data distribution between network-attached block storage and object storage](/docs/products/clickhouse/howto/check-data-tiered-storage.md) * [Configure data retention thresholds for tiered storage](/docs/products/clickhouse/howto/configure-tiered-storage.md) --- # Maintenance and updates for your Aiven for ClickHouse® service Manage maintenance updates and set the maintenance window for your Aiven for ClickHouse® service. ## Maintenance updates[​](#maintenance-updates "Direct link to Maintenance updates") Aiven applies some maintenance updates automatically. The following are the types of updates: * **Mandatory updates:** Security updates, quarterly patch releases, and platform updates that affect reliability or stability of the service nodes. * **Optional updates:** All other updates are initially optional. After six months, they become mandatory and are applied in the next week’s maintenance window. * **Periodic infrastructure updates:** Scheduled automatically for services with nodes active for 180 days and more. These updates are mandatory for all services, except those with maintenance turned off. Critical security updates are applied during the next available maintenance window. For other updates, Aiven gives you at least seven days' notice. Maintenance updates are also automatically applied during service upgrades. To view pending updates: * Console * CLI * API 1. In your service, click **Service settings**. 2. Go to the **Service management** section. Use the [`avn service get`](/docs/tools/cli/service-cli.md#avn_service_get) command. Use the [`service`](https://api.aiven.io/doc/#tag/Service/operation/ServiceGet) endpoint. ## Maintenance window[​](#maintenance-window "Direct link to Maintenance window") The maintenance window is the time period when Aiven can automatically apply maintenance updates to a service. When an update becomes available, Aiven schedules it for the next available maintenance window for each service. The update runs in the first window after it becomes available, and can begin any time after the start time. For example, if a service has a maintenance window of Monday 12:00 UTC, and an update becomes available on Tuesday, the update will be applied on the following Monday. During maintenance, Aiven might restart or replace service nodes. This can cause brief connection interruptions, but services are designed to minimize downtime. Aiven performs maintenance in a rolling-forward style, creating new nodes alongside existing ones and retiring the old nodes after the upgrade completes. Major service upgrades are triggered manually. A manually triggered upgrade starts immediately, regardless of the maintenance window. important You cannot control the order in which services are updated. Each service updates according to its own configured maintenance window, and there is no guaranteed way to control the update sequence. Manual updates and maintenance window adjustments only help for non-critical updates. ## Set the maintenance window[​](#set-the-maintenance-window "Direct link to Set the maintenance window") To set the maintenance window for your service: * Console * Terraform 1. In the Aiven Console, open your service. 2. In the **Maintenance** section, click **Actions** > **Change maintenance window**. 3. Set the day and time. 4. Click **Save changes**. Use the `maintenance_window_dow` and `maintenance_window_time` attributes in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Manage versions in Aiven for ClickHouse®](/docs/products/clickhouse/howto/manage-clickhouse-versions.md) * [Change the service plan](/docs/products/clickhouse/howto/change-service-plan.md) * [Aiven for ClickHouse version lifecycle](/docs/products/clickhouse/reference/version-lifecycle.md) * [Scale disk storage](/docs/products/clickhouse/howto/scale-disk-storage.md) --- # Manage versions in Aiven for ClickHouse® Aiven for ClickHouse® supports multiple ClickHouse versions. You can choose a version when you create a service and upgrade to a newer supported version later. ## Supported ClickHouse versions[​](#supported-clickhouse-versions "Direct link to Supported ClickHouse versions") Aiven for ClickHouse supports the following major versions: * `25.3`: default for services created before May 5, 2026 * `25.8`: default for services created on or after May 5, 2026 * `26.3`: available in General Availability from September 1, 2026 If you don't specify a version when you create a service, Aiven uses the default version in effect on the creation date. You can select `26.3` when creating a new service or when upgrading a service running version `25.8`. Upgrade services running `25.3` to `25.8` first. Version `25.8` remains the default for new services. Aiven doesn't automatically upgrade existing `25.3` services to `25.8`. When using infrastructure as code (for example, the [Aiven Provider for Terraform](/docs/tools/terraform.md) or the [Aiven Operator for Kubernetes®](/docs/tools/kubernetes.md)), set `clickhouse_version` to specify the major version for new services. To upgrade the ClickHouse version of an existing service, see [Upgrade your service](#upgrade-your-service). For supported versions and end-of-life timelines, see [Aiven for ClickHouse end-of-life policy](/docs/platform/reference/eol-for-major-versions.md#aiven-for-clickhouse). ## Before you upgrade[​](#before-you-upgrade "Direct link to Before you upgrade") Before upgrading your service, complete the following checks: * Review [Upgrade to Aiven for ClickHouse 26.3](/docs/products/clickhouse/reference/upgrade-to-26-3.md), including the **Requires attention** section and its checks for removed features. * Test the upgrade in a development or staging environment by [forking the service](/docs/products/clickhouse/howto/fork-service.md) and upgrading the fork first. * Verify that your applications and clients support the target version. * Ensure recent backups are available. * Avoid running long-running queries or heavy ingestion during the upgrade. Downgrading to an earlier ClickHouse version is not supported. ## How upgrades work[​](#how-upgrades-work "Direct link to How upgrades work") During an upgrade, the platform replaces service nodes with new nodes running the selected version. * New service nodes are created with the selected ClickHouse version. * The new nodes start alongside the existing nodes. * Data is streamed from the existing nodes to the new nodes. * After the migration completes, the service switches to the upgraded nodes. * The previous nodes are removed. This process avoids modifying the existing nodes directly during the upgrade. During this process: * Queries connected to nodes being replaced can fail and require retries. * New connections are routed to upgraded nodes. ## Service availability during upgrades[​](#service-availability-during-upgrades "Direct link to Service availability during upgrades") The service remains available during the upgrade. * New nodes are added and receive data before old nodes are removed, which helps keep the service reachable while the upgrade runs. * Short interruptions can occur. * Latency can increase and throughput can decrease during the upgrade. ## Upgrade your service[​](#upgrade-your-service "Direct link to Upgrade your service") Upgrade the service to a newer supported ClickHouse version using the Aiven Console, CLI, API, or Terraform. * Console * CLI * API * Terraform 1. In the [Aiven Console](https://console.aiven.io), open your Aiven for ClickHouse service. 2. On the **Overview** page, go to the **Maintenance** section. 3. Click **Actions** > **Upgrade version**. 4. Select a version to upgrade to, and click **Upgrade**. 1) Check the available ClickHouse versions: ``` avn service versions ``` 2) Upgrade the service: ``` avn service update SERVICE_NAME \ --project PROJECT_NAME \ -c clickhouse_version="CLICKHOUSE_VERSION" ``` Parameters: * `SERVICE_NAME`: Name of the service * `PROJECT_NAME`: Name of the project * `CLICKHOUSE_VERSION`: Target ClickHouse version, for example `25.8` Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint and set `clickhouse_version`. ``` curl --request PUT \ --url "https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME" \ --header "Authorization: Bearer API_TOKEN" \ --header "Content-Type: application/json" \ --data '{ "user_config": { "clickhouse_version": "CLICKHOUSE_VERSION" } }' ``` Parameters: * `PROJECT_NAME`: Name of the project * `SERVICE_NAME`: Name of the service * `API_TOKEN`: API authentication token * `CLICKHOUSE_VERSION`: Target ClickHouse version, for example `25.8` 1. Update the [`aiven_clickhouse`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/clickhouse) resource and set `clickhouse_version`. ``` user_config { clickhouse_version = "CLICKHOUSE_VERSION" } ``` Set `CLICKHOUSE_VERSION` to the target ClickHouse version, for example `25.8`. 2. Apply the change: ``` terraform apply ``` --- # Manage Aiven for ClickHouse® databases and tables Create and work with databases and tables in Aiven for ClickHouse®. tip Additionally to regular databases, you can also create [integration databases](/docs/products/clickhouse/howto/integration-databases.md#create-integration-databases). ## Create a database[​](#create-a-clickhouse-database "Direct link to Create a database") You can create a database either in the [Aiven Console](https://console.aiven.io/) or using an SQL client such as the [ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md). note Your Aiven for ClickHouse service can support up to 400 databases simultaneously. * Aiven Console * SQL 1. Log in to the [Aiven Console](https://console.aiven.io/), and select your service from the **Services** page. 2. In the sidebar, click **Data** > **Databases and tables**. 3. Click **Create database** > **ClickHouse database**. 4. In the **Create ClickHouse database** window, enter a name for your database and select **Create database**. The database name appears in the list of databases in the **Databases and tables** page. Aiven applies the required customizations and runs secondary queries to grant access to the admin user. **Limitations** * Only the `avnadmin` user can create databases in SQL. * You can create a database in SQL with the `Replicated` database engine only. To create a database in SQL, run the following SQL command: ``` CREATE DATABASE DATABASE_NAME ENGINE = Replicated ``` For example: ``` CREATE DATABASE transactions ENGINE = Replicated; ``` ## Delete a database[​](#delete-a-database "Direct link to Delete a database") important Deleting a database is irreversible and permanently removes the database along with all its tables and data. You can delete a database either in the [Aiven Console](https://console.aiven.io/) or using an SQL client such as the [ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md): * Aiven Console * SQL 1. Log in to the [Aiven Console](https://console.aiven.io/), and select your service from the **Services** page. 2. In the sidebar, click **Data** > **Databases and tables**. 3. In the **Databases and tables** list, find your database and click **Actions** > **Delete database**. **Limitation**: By default, only the `aiven` user can delete a database. The `aiven` user can grant the permission to delete a database to another user. To delete a database in SQL, run the following SQL command: ``` DROP DATABASE DATABASE_NAME ``` For example: ``` DROP DATABASE transactions; ``` ## Create a table[​](#create-a-table "Direct link to Create a table") Tables can be added with an SQL query, either with the help of the web query editor or with CLI. In both cases, the SQL query looks the same. The example below shows a query to add new table `expenses` to `transactions` database. To keep it simple, this example has an unrealistically small amount of columns: ``` CREATE TABLE transactions.expenses ( Title String, Date DateTime, UserID UInt64, Amount UInt32 ) ENGINE = ReplicatedMergeTree ORDER BY Date; ``` ### Select a table engine[​](#select-a-table-engine "Direct link to Select a table engine") Part of the table definition includes a targeted table engine. See the full [list of supported table engines](/docs/products/clickhouse/reference/supported-table-engines.md) in Aiven for ClickHouse. Aiven for ClickHouse uses `replicated` variants of table engines to ensure high availability. Even if you select `MergeTree` engine, Aiven automatically uses the replicated variant. note A non-replicated table, such as `system.query_log`, can be [queried using `clusterAllReplicas`](/docs/products/clickhouse/howto/query-databases.md#query-a-non-replicated-table). ## Delete a table[​](#delete-a-table "Direct link to Delete a table") You can remove a table of any size if you have the `DROP` permission since parameters `max_table_size_to_drop` and `max_partition_size_to_drop` are disabled for Aiven services. Consider [granting](/docs/products/clickhouse/howto/manage-users-roles.md) only necessary permissions to your database users. * CLI * Console Run the following SQL command to remove your table: ``` DROP TABLE NAME_OF_YOUR_DATABASE.NAME_OF_YOUR_TABLE; ``` To remove your table in the [Aiven Console](https://console.aiven.io/): 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. Go to the table to be removed: organization > project > service > **Data** > **Databases and tables**. 3. In the **Databases and tables** view, go to the table and select **Actions** > **Delete table**. Related pages [Create Aiven for ClickHouse integration databases](/docs/products/clickhouse/howto/integration-databases.md#create-integration-databases) --- # Manage Aiven for ClickHouse® users and roles Create Aiven for ClickHouse® users and roles and grant them specific privileges to control or restrict access to your service. ## Manage users[​](#manage-users "Direct link to Manage users") ### Add a user[​](#add-a-user "Direct link to Add a user") To create a user account for your service: * Console * SQL 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose your Aiven for ClickHouse service. 2. In the service sidebar, click **Users**. This shows a list of all users that are currently available in your service. The default `avnadmin` user has all available access grants to the service. tip To view the roles and grants for any of the listed users, click **View details & grants** for that user. 3. On the **Users** page, click **Add user**. 4. In the **Create a service user** window, enter a name for the new user and click a role. The role that you select defines the access grants that are assigned to the user. For more information on roles, see [Manage roles and privileges](/docs/products/clickhouse/howto/manage-users-roles.md#manage-roles-and-privileges). 5. Click **Add user**. This creates the new user and shows you a summary of the information. 6. Copy the password on screen to a safe place. You can't access it again later, but you can reset it if needed. To create a user by using a hash and salt, run: ``` CREATE USER username IDENTIFIED WITH sha256_hash BY 'hash' SALT 'salt'; ``` To create a user using a password directly, run: ``` CREATE USER username IDENTIFIED BY 'password'; ``` Users' password digests are stored securely in ZooKeeper. Only a super admin can access them. For users created in the console, their passwords are also salted with a random value. The password digests and salts are stored in backup files so they can be recovered. ### Configure user settings[​](#configure-user-settings "Direct link to Configure user settings") Configure user settings to control resource usage and query behavior. You can set limits at the user level (applying to all queries by that user) or configure per-query constraints. #### Set user resource limits[​](#set-user-resource-limits "Direct link to Set user resource limits") Configure user settings, for example, to restrict resource use per user. This can help you: * Prevent single users from monopolizing CPU threads and starving other users. * Ensure fair distribution of system resources among all active users. * Maintain system stability by preventing service overload. 1. Create a user named `limited_user`. ``` CREATE USER limited_user IDENTIFIED WITH SHA256_PASSWORD BY 'PASSWORD'; ``` 2. Set resource limit `max_threads = 1` for this user. ``` ALTER USER limited_user SETTINGS max_threads = 1; ``` 3. Grant read-only access to all databases and tables to `limited_user`. ``` GRANT SELECT ON *.* TO limited_user; ``` 4. Log in as `limited_user` to test the configuration. ``` clickhouse-client \ --host YOUR_HOST \ --port YOUR_PORT \ --user limited_user \ --password PASSWORD \ --secure ``` Replace `YOUR_HOST`, `YOUR_PORT`, and `PASSWORD` with your service connection details. 5. Verify the resource limit setting. ``` SHOW SETTINGS LIKE 'max_threads'; ``` This displays the `max_threads` setting with a value of `1` for the `limited_user`. #### Set per-query resource limits[​](#set-per-query-resource-limits "Direct link to Set per-query resource limits") Constrain resources on a per-query basis for specific users by setting query-level limits in their user profile. 1. Create a user with per-query memory and execution time limits. ``` CREATE USER query_limited_user IDENTIFIED WITH SHA256_PASSWORD BY 'PASSWORD' SETTINGS max_memory_usage = 1000000000, max_execution_time = 30; ``` 2. Grant appropriate permissions. ``` GRANT SELECT ON *.* TO query_limited_user; ``` 3. Test the per-query limits by logging in and running a query. ``` clickhouse-client \ --host YOUR_HOST \ --port YOUR_PORT \ --user query_limited_user \ --password PASSWORD \ --secure ``` 4. Verify the query-level settings. ``` SHOW SETTINGS LIKE '%max_%'; ``` #### Common per-query settings[​](#common-per-query-settings "Direct link to Common per-query settings") * `max_memory_usage`: Maximum memory per query (in bytes) * `max_execution_time`: Maximum query execution time (in seconds) * `max_rows_to_read`: Maximum rows a query can examine * `max_result_rows`: Maximum rows in query result ## Manage roles and privileges[​](#manage-roles-and-privileges "Direct link to Manage roles and privileges") Aiven for ClickHouse has no predefined roles. All roles you create are custom roles you design for your purposes by granting specific privileges to particular roles. For example: you can create a role that allows only reading a single database, table, or column; or you can create another role that allows only inserting data, not deleting it. ClickHouse® supports a **Role Based Access Control** model and allows you to configure access privileges by using SQL statements. You can either [use the query editor](/docs/products/clickhouse/howto/query-databases.md) or rely on [the command-line interface](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md). The upstream ClickHouse documentation includes [detailed documentation for access rights](https://clickhouse.com/docs/en/operations/access-rights/). ### Create a role[​](#create-a-role "Direct link to Create a role") To create a role named **auditor**, run the following command: ``` CREATE ROLE auditor; ``` Find more information on creating roles in the [upstream ClickHouse documentation](https://clickhouse.com/docs/en/sql-reference/statements/create/role/). ### Grant privileges[​](#grant-privileges "Direct link to Grant privileges") You can grant privileges both to specific roles and to individual users. The grants can be also granular, targeting specific databases, tables, columns, or rows. important You cannot grant additional privileges to the main service user. Aiven may grant privileges to the main service user during maintenance updates when adding new features for the service. As an example, the following request grants the `auditor` role privileges to select data from the `transactions` database: ``` GRANT SELECT ON transactions.* TO auditor; ``` You can limit the grant to a specified table: ``` GRANT SELECT ON transactions.expenses TO auditor; ``` Or to particular columns of a table: ``` GRANT SELECT(date,description,amount) ON transactions.expenses TO auditor ``` To grant the `auditor` and `external` roles to several users, run: ``` GRANT auditor, external TO Mary.Anderson, James.Miller; ``` To allow the creation of new users: ``` GRANT CREATE USER ON transactions.* TO administrator ``` There are a variety of privileges that you can grant. Find the full list in the [upstream ClickHouse documentation](https://clickhouse.com/docs/en/sql-reference/statements/grant/#privileges). note You can grant privileges to a table that does not yet exist. note Users can grant privileges according to their privileges. If the user lacks the required privileges for a requested operation, they receive a `Not enough privileges` exception. warning Privileges are not revoked when a table or database is removed. They continue to be active for any new table or database that is created with the same name. Find all details on how the GRANT statement is supported in ClickHouse in the [upstream ClickHouse documentation](https://clickhouse.com/docs/en/sql-reference/statements/grant/). ### Set roles[​](#set-roles "Direct link to Set roles") A single user can be assigned different roles, either individually or simultaneously. ``` SET ROLE auditor; ``` You can also specify a role to be activated by default when the user logs in: ``` SET DEFAULT ROLE auditor, external TO Mary.Anderson, James.Miller; ``` ### Delete a role[​](#delete-a-role "Direct link to Delete a role") If you no longer need a role, you can remove it: ``` DROP ROLE auditor; ``` ### Revoke privileges[​](#revoke-privileges "Direct link to Revoke privileges") Remove all or specific privileges from users or roles: ``` REVOKE SELECT ON transactions.expenses FROM Mary.Anderson; ``` Revoke all privileges to a table or database simultaneously: ``` REVOKE ALL PRIVILEGES ON database.table FROM external; ``` For more information about revoking privileges, see the [ClickHouse documentation](https://clickhouse.com/docs/en/sql-reference/statements/revoke/). ### Check privileges[​](#check-privileges "Direct link to Check privileges") Run the following commands to see all available grants, users, and roles: ``` SHOW GRANTS; ``` ``` SHOW USERS; ``` ``` SHOW ROLES; ``` ### Preview users and roles in the console[​](#preview-users-and-roles-in-the-console "Direct link to Preview users and roles in the console") You can also see the users, their roles, and privileges in the [Aiven Console](https://console.aiven.io/). 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose your Aiven for ClickHouse service. 2. In the service sidebar, click **Users**. 3. Click **View details & grants** next to one of the users listed on the page. This shows you a list of all grants for the selected user. ## Manage using Terraform[​](#manage-using-terraform "Direct link to Manage using Terraform") You can also manage user roles and access using the [Aiven Provider for Terraform](/docs/tools/terraform.md). Learn how to manage multiple grants while avoiding conflicts with the [ClickHouse grants example](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/clickhouse/grants). --- # Create materialized views in ClickHouse® Use materialized views to persist data from the Kafka® table engine. One way of integrating your ClickHouse® service with Kafka® is using the Kafka® table engine, which enables, for example, inserting data into ClickHouse® from Kafka. In such a scenario, ClickHouse can read from a Kafka® topic directly. This is, however, one-time retrieval so the data cannot be re-read. When a new block of data is inserted in a table (whether the table is backed by a MergeTree or Kafka® engine), the table expression defined by the materialized view is triggered with the block of data. With Kafka®, the block of data lives in a buffer in memory and the trigger *flushes* the buffer. For this reason, the data consumed from the topic by the Kafka® engine cannot be read twice. Persisting the data from the Kafka® table engine read requires capturing and inserting it into a different table. A materialized view triggers a read on the table engine. The destination of the data (for example, a Merge Tree family table) is defined by the TO clause. This process is illustrated in the following diagram: ## Create a materialized view[​](#create-a-materialized-view "Direct link to Create a materialized view") To store the Kafka® messages by creating a materialized view on top of the Kafka® table, run: ``` CREATE MATERIALIZED VIEW default.my_view TO destination_name ENGINE = ReplicatedMergeTree ORDER BY x AS SELECT x,y FROM service_kaf.table_name ``` note All ClickHouse® nodes share the same consumer group: Each message is consumed and stored by a single node. However, since materialized views use the ReplicatedMergeTree engine, the stored messages are exchanged and replicated across all the nodes. Related pages For more information on how to integrate Aiven for ClickHouse® with Apache Kafka®, see [Connect Apache Kafka® to Aiven for ClickHouse®](/docs/products/clickhouse/howto/integrate-kafka.md). --- # Monitor Aiven for ClickHouse® metrics with Aiven for Grafana® Push Aiven for ClickHouse® metrics to Aiven for Metrics or Aiven for PostgreSQL®, and integrate with Aiven for Grafana® to monitor your metrics on Grafana dashboards. For more information on the metrics, see [monitoring dashboard metrics shown in Grafana®](/docs/products/clickhouse/reference/metrics-list.md). ## Metrics storage options[​](#metrics-storage-options "Direct link to Metrics storage options") Aiven provides two options for storing your Aiven for ClickHouse metrics: * **Aiven for Metrics**: A managed time-series database service built on Thanos, optimized for long-term metrics storage and querying No external Thanos access While Aiven for Metrics uses the Thanos architecture internally, **you cannot connect your Aiven-managed services to any external (non-Aiven) Thanos endpoint**. Connect to Aiven for Metrics instead. * **Aiven for PostgreSQL**: A relational database that can store metrics data using TimescaleDB extension ## Push ClickHouse® metrics to Metrics or PostgreSQL[​](#push-clickhouse-metrics-to-metrics-or-postgresql "Direct link to Push ClickHouse® metrics to Metrics or PostgreSQL") To collect metrics about your Aiven for ClickHouse service, configure a metrics integration and nominate somewhere to store the collected metrics. 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose your Aiven for ClickHouse service. 2. In the service sidebar, click **Data** > **Integrations**. 3. In the **Aiven services** section, click **Store Metrics**. 4. In the **Metrics integration** window: 1. Choose either a new or existing Aiven for Metrics or Aiven for PostgreSQL service. * **For Aiven for Metrics**: This provides a Thanos-based time-series database optimized for metrics storage and long-term retention * **For Aiven for PostgreSQL**: This stores metrics in a relational format, suitable if you prefer SQL-based querying * If you choose to use a new service, follow instructions on [how to create a service](/docs/products/clickhouse/get-started.md#create-an-aiven-for-clickhouse-service). * If you're already using Aiven for Metrics or Aiven for PostgreSQL, you can submit your Aiven for ClickHouse metrics to the existing service. 2. Click **Enable**. ## Provision and configure Grafana[​](#provision-and-configure-grafana "Direct link to Provision and configure Grafana") 1. In the [Aiven Console](https://console.aiven.io/), choose your Aiven for Metrics or Aiven for PostgreSQL service. 2. Open the service's **Integrations** page. * For Aiven for Metrics, click **Integrations** in the service sidebar. * For Aiven for PostgreSQL, in the service sidebar, click **Connect** > **Integrations**. 3. In the **Aiven services** section, click **Grafana Metrics Dashboard**. 4. In the **Dashboard integration** window: 1. Choose either a new or existing Aiven for Grafana service. * If you choose to use a new service, follow the instructions on [how to create a service](/docs/products/clickhouse/get-started.md#create-an-aiven-for-clickhouse-service). * If you're already using Grafana on Aiven, you can integrate your Aiven for Metrics or Aiven for PostgreSQL as an additional data source for that existing Grafana. 2. Click **Enable**. note Now your Aiven for Grafana service is connected to your metrics storage service as a data source. **After a few minutes**, your metrics are available for visualization on a Grafana dashboard. ## Open ClickHouse metrics dashboard[​](#open-clickhouse-metrics-dashboard "Direct link to Open ClickHouse metrics dashboard") 1. In the Aiven Console, go to the **Overview** page of the integrated Aiven for Grafana service, and click the **Service URI** link. 2. Log in to Grafana using the username and password available on the Aiven for Grafana service **Overview** page > **Connection information**. 3. In Grafana, go to **Dashboards** and open your Aiven for ClickHouse metrics dashboard. 4. Browse prebuilt views or create your own monitoring views. --- # Power on/off and delete your Aiven for ClickHouse® service Power off your Aiven for ClickHouse® service to release resources and save credits, power it back on when you need it, or delete it permanently. ## Power off a service[​](#power-off-a-service "Direct link to Power off a service") When you power off a service: * All virtual machines are removed from the public cloud. * The service configuration is stored on the Aiven Platform. * If there are no backups, all service data is lost. * If the service has time-based or point in time recovery backups, the backups remain on the Aiven Platform. Services powered off for more than 180 days are automatically deleted. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power off service**. To power off a service, run: ``` avn service update SERVICE_NAME --power-off ``` ## Power on a service[​](#power-on-a-service "Direct link to Power on a service") When you power on a service: * New virtual machines are created on the service's public cloud. * The service starts with the stored configuration parameters. * The latest time-based backup is restored. * Maintenance updates are automatically applied. * If a point in time recovery backup is available, the database transaction logs are replayed to recover the service data to a specific point in time. The restoration takes from a few minutes to a few hours, depending on the network bandwidth, the disk IOPS allocated to the service, and the size of the backup. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power on service**. To power on a service, run: ``` avn service update SERVICE_NAME --power-on ``` To see when the service is running, run: ``` avn service wait SERVICE_NAME ``` note When you power on an Aiven for ClickHouse service, your data is restored from the latest available backup. An automatic backup is also taken before the service is powered off. note Static IP addresses are not removed when a service is powered off or deleted. They continue to generate the usual costs. To avoid these costs, [remove the static IP addresses](/docs/platform/concepts/static-ips.md). ## Delete a service[​](#delete-a-service "Direct link to Delete a service") * Console * CLI 1. In your project, click **Services**. 2. Open the service to delete, and click **Actions** > **Delete service**. To delete a service, run: ``` avn service terminate SERVICE_NAME ``` Related pages * [Disaster recovery in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/disaster-recovery.md) * [Fork your Aiven for ClickHouse® service](/docs/products/clickhouse/howto/fork-service.md) --- # Query Aiven for ClickHouse® databases Run a query against an Aiven for ClickHouse® database using a tool of your choice. To ensure data security, stability, and its proper replication, we equip our managed Aiven for ClickHouse® service with specific features, some of them missing from the standard ClickHouse offer. Aiven for ClickHouse® takes care of running queries in the distributed mode over the entire cluster. In the standard ClickHouse, the queries `CREATE`, `ALTER`, `RENAME` and `DROP` only affect the server where they are run. In contrast, we ensure the proper distribution across all cluster machines behind the scenes. You don't need to use `ON CLUSTER` for data definition queries against a `Replicated` database. important There are limitations on the number of concurrent queries and the number of concurrent connections in Aiven for ClickHouse: * `max_concurrent_queries` ranges from `25` to `400`. * `max_concurrent_connections` ranges from `1000` to `4000`. See [Aiven for ClickHouse® limits and limitations](/docs/products/clickhouse/reference/limitations.md) for details. For querying your ClickHouse® databases, you can choose between our query editor, the Play UI, and [the ClickHouse® client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md). ## Query a database with a selected tool[​](#query-a-database-with-a-selected-tool "Direct link to Query a database with a selected tool") ### Query editor[​](#use-query-editor "Direct link to Query editor") Aiven for ClickHouse® includes a web-based query editor. To open it, log in to the [Aiven Console](https://console.aiven.io/), choose your Aiven for ClickHouse service, and click **Data** > **Query editor** in the service sidebar. #### When to use the query editor[​](#when-to-use-the-query-editor "Direct link to When to use the query editor") The query editor is convenient to run queries directly from the console on behalf of the default user. The requests that you run through the query editor rely on the permissions granted to this user. #### Examples of queries[​](#examples-of-queries "Direct link to Examples of queries") Retrieve a list of current databases: ``` SHOW DATABASES ``` Count rows: ``` SELECT COUNT(*) FROM transactions.accounts ``` Create a role: ``` CREATE ROLE accountant ``` ### Play UI[​](#play-iu "Direct link to Play UI") ClickHouse® includes a built-in user interface for running SQL queries. You can access it from a web browser over the HTTPS protocol. #### When to use the play UI[​](#when-to-use-the-play-ui "Direct link to When to use the play UI") Use the play UI to run requests using a non-default user or if you expect a large size of the response. #### Use the play UI[​](#use-the-play-ui "Direct link to Use the play UI") 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose your Aiven for ClickHouse service. 2. On the **Overview** page, in the **Connection information** section, select **ClickHouse HTTPS & JDBC**. 3. Copy **Service URI** and go to `YOUR_SERVICE_URI/play` from a web browser. 4. Set the name and the password of the user on whose behalf you want to run the queries. 5. Enter the body of the query. 6. Select **Run**. note The play interface is only available if you can connect directly to ClickHouse from your browser. If the service is [restricted by IP addresses](/docs/platform/howto/restrict-access.md) or in a [VPC without public access](/docs/platform/howto/public-access-in-vpc.md), you can use the [query editor](/docs/products/clickhouse/howto/query-databases.md#use-query-editor) instead. The query editor can be accessed directly from the console to run requests on behalf of the default user. ## Query a non-replicated table[​](#query-a-non-replicated-table "Direct link to Query a non-replicated table") Behind the DNS name of your Aiven for ClickHouse service, there are multiple nodes. When you query a non-replicated table, for example a log table, requests are routed randomly to one of the nodes regardless of how data is distributed across them. A particular row is found only if your `SELECT` query is directed to the node which executed a `WRITE` on this row. To query a non-replicated table across all the service nodes, use `clusterAllReplicas` as follows: ``` SELECT * FROM clusterAllReplicas(default, system.query_log) WHERE query_id = '1a2b3c4d5e6f7g8h9i0j1a2b3c4d5e6f7g8' ``` --- # Rename your Aiven for ClickHouse® service Change the name of your Aiven for ClickHouse® service by forking it under a new name and deleting the original service. You cannot rename a service after creation. Instead, you can create a fork with the new name and delete the original service. ## Rename a service[​](#rename-a-service "Direct link to Rename a service") 1. Stop writing to the service. 2. Fork the service. 3. Add any integrations or SSO configurations that weren't copied. 4. Connect your clients to the new service. 5. Test the forked service. 6. Delete the original service. Related pages * [Fork your Aiven for ClickHouse® service](/docs/products/clickhouse/howto/fork-service.md) * [Power on/off and delete your Aiven for ClickHouse® service](/docs/products/clickhouse/howto/power-cycle-service.md) --- # Restore an Aiven for ClickHouse® backup Restore an Aiven for ClickHouse® service from a [daily backup](/docs/products/clickhouse/concepts/disaster-recovery.md#service-backup) by forking to a new service. important You cannot restore Aiven for ClickHouse services to a fewer number of nodes. Reducing the number of nodes is only possible by [switching the service plan](/docs/products/clickhouse/howto/change-service-plan.md) from **Business** to **Startup** on a running service. To restore a backup: 1. In your service, in the **Backups** section, click **Backup management**. 2. Click **Fork & restore**. 3. Choose the backup to fork from. 4. Enter a name, and select the cloud and plan. 5. Click **Create fork**. Once the new service is running, change your application's connection settings to point to it and power off the original service. Related pages * [Fork your Aiven for ClickHouse® service](/docs/products/clickhouse/howto/fork-service.md) * [Change the service plan](/docs/products/clickhouse/howto/change-service-plan.md) * [Schedule Aiven for ClickHouse® backups](/docs/products/clickhouse/howto/configure-backup.md) * [Disaster recovery in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/disaster-recovery.md) --- # Read and pull data from S3 object storages and web resources over HTTP With federated queries in Aiven for ClickHouse®, you can read and pull data from an external S3-compatible object storage or any web resource accessible over HTTP. Learn more about capabilities and applications of federated queries in [About querying external data in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/federated-queries.md). ## About running federated queries[​](#about-running-federated-queries "Direct link to About running federated queries") Federated queries are written using specific SQL statements and can be run from CLI, for instance. To run a federated query, just send a query over an external S3-compatible object storage including relevant S3 bucket details. A properly constructed federated query returns a specific output. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") The prerequisites depend on the table function or table engine used in your federated query. ### Access to S3 and URL sources[​](#access-to-s3-and-url-sources "Direct link to Access to S3 and URL sources") To run a federated query, the ClickHouse service user connecting to the cluster requires grants to the S3 and/or URL sources. The main service user is granted access to the sources by default, and new users can be allowed to use the sources with the following query: ``` GRANT CREATE TEMPORARY TABLE, S3, URL ON *.* TO [WITH GRANT OPTION] ``` The CREATE TEMPORARY TABLE grant is required for both sources. Adding WITH GRANT OPTION allows the user to further transfer the privileges. ### Azure Blob Storage access keys[​](#azure-blob-storage-access-keys "Direct link to Azure Blob Storage access keys") To run federated queries using the `azureBlobStorage` table function or the `AzureBlobStorage` table engine, get your Azure Blob Storage keys using one of the following tools: * [Azure portal](https://portal.azure.com/) From the portal menu, select **Storage accounts**, go to your account, and click **Security + Networking** > **Access keys**. View and copy your account access keys and connection strings. * [PowerShell](https://learn.microsoft.com/en-us/powershell/scripting/install/installing-powershell?view=powershell-7.4) * [Azure CLI](https://learn.microsoft.com/en-us/cli/azure/install-azure-cli#install) ### Managed credentials for Azure Blob Storage[​](#managed-credentials-for-azure-blob-storage "Direct link to Managed credentials for Azure Blob Storage") [Managed credentials integration](/docs/products/clickhouse/concepts/data-integration-overview.md#managed-credentials-integration) is: * Required to [run federated queries using the AzureBlobStorage table engine](/docs/products/clickhouse/howto/run-federated-queries.md#query-using-the-azureblobstorage-table-engine) * Optional to [run federated queries using the azureBlobStorage table function](/docs/products/clickhouse/howto/run-federated-queries.md#query-using-the-azureblobstorage-table-function) [Set up a managed credentials integration](/docs/products/clickhouse/howto/data-service-integration.md#create-managed-credentials-integrations) as needed. ## Limitations[​](#limitations "Direct link to Limitations") * Federated queries in Aiven for ClickHouse only support S3-compatible object storage providers for the time being. * Virtual tables are only supported for URL sources, using the URL table engine. ## Run a federated query[​](#run-a-federated-query "Direct link to Run a federated query") See some examples of running federated queries to read and pull data from external S3-compatible object storages. ### Query using the `azureBlobStorage` table function[​](#query-using-the-azureblobstorage-table-function "Direct link to query-using-the-azureblobstorage-table-function") Depending on how you choose to handle passing connection parameters in your queries, you can run federated queries using the `azureBlobStorage` table function: * [With managed credentials integration](/docs/products/clickhouse/howto/run-federated-queries.md#azureblobstorage-table-function-without-managed-credentials) * [Without managed credentials integration](/docs/products/clickhouse/howto/run-federated-queries.md#azureblobstorage-table-function-with-managed-credentials) Before you start, fulfill relevant [prerequisites](/docs/products/clickhouse/howto/run-federated-queries.md#prerequisites), if any. #### `azureBlobStorage` table function without managed credentials[​](#azureblobstorage-table-function-without-managed-credentials "Direct link to azureblobstorage-table-function-without-managed-credentials") ##### SELECT[​](#select "Direct link to SELECT") ``` SELECT * FROM azureBlobStorage( 'DefaultEndpointsProtocol=https;AccountName=ident;AccountKey=secret;EndpointSuffix=core.windows.net', 'ownerresource', 'all_stock_data.csv', 'CSV', 'auto', 'Ticker String, Low Float64, High Float64' ) LIMIT 5 ``` ##### INSERT[​](#insert "Direct link to INSERT") ``` INSERT INTO FUNCTION azureBlobStorage( 'DefaultEndpointsProtocol=https;AccountName=ident;AccountKey=secret;EndpointSuffix=core.windows.net', 'ownerresource', 'test_funcwrite.csv', 'CSV', 'auto', 'key UInt64, data String' ) VALUES ('column1-value', 'column2-value'); ``` #### `azureBlobStorage` table function with managed credentials[​](#azureblobstorage-table-function-with-managed-credentials "Direct link to azureblobstorage-table-function-with-managed-credentials") ``` azureBlobStorage( `named_collection`, blobpath = 'path/to/blob.csv', format = 'CSV' ) ``` ### Query using the `AzureBlobStorage` table engine[​](#query-using-the-azureblobstorage-table-engine "Direct link to query-using-the-azureblobstorage-table-engine") Before you start, fulfill relevant [prerequisites](/docs/products/clickhouse/howto/run-federated-queries.md#prerequisites), if any. 1. Create a table: ``` CREATE TABLE default.test_azure_table ( `Low` Float64, `High` Float64 ) ENGINE = AzureBlobStorage(`endpoint_azure-blob-storage-datasets`, blob_path = 'data.csv', compression = 'auto', format = 'CSV') ``` 2. Query from the `AzureBlobStorage` table engine: ``` SELECT avg(Low) FROM test_azure_table ``` ### Query using the `s3` table function[​](#query-using-the-s3-table-function "Direct link to query-using-the-s3-table-function") Before you start, fulfill relevant [prerequisites](/docs/products/clickhouse/howto/run-federated-queries.md#prerequisites), if any. #### SELECT and `s3`[​](#select-and-s3 "Direct link to select-and-s3") SQL SELECT statements using the S3 and URL functions are able to query public resources using the URL of the resource. For instance, let's explore the network connectivity measurement data provided by the [Open Observatory of Network Interference (OONI)](https://ooni.org/data/). ``` WITH ooni_data_sample AS ( SELECT * FROM s3('https://ooni-data-eu-fra.s3.eu-central-1.amazonaws.com/clickhouse_export/csv/fastpath_202308.csv.zstd') LIMIT 100000 ) SELECT probe_cc AS probe_country_code, test_name, countIf(anomaly = 't') AS total_anomalies FROM ooni_data_sample GROUP BY probe_country_code, test_name HAVING total_anomalies > 10 ORDER BY total_anomalies DESC LIMIT 50 ``` #### INSERT and `s3`[​](#insert-and-s3 "Direct link to insert-and-s3") When executing an INSERT statement into the S3 function, the rows are appended to the corresponding object if the table structure matches: ``` INSERT INTO FUNCTION s3('https://bucket-name.s3.region-name.amazonaws.com/dataset-name/landing/raw-data.csv', 'CSVWithNames') VALUES ('column1-value', 'column2-value'); ``` ### Query a private S3 bucket[​](#query-a-private-s3-bucket "Direct link to Query a private S3 bucket") Before you start, fulfill relevant [prerequisites](/docs/products/clickhouse/howto/run-federated-queries.md#prerequisites), if any. Private buckets can be accessed by providing the access token and secret as function parameters. ``` SELECT * FROM s3( 'https://private-bucket.s3.eu-west-3.amazonaws.com/dataset-prefix/partition-name.csv', 'some_aws_access_key_id', 'some_aws_secret_access_key' ) ``` Depending on the format, the schema can be automatically detected. If it isn't, you may also provide the column types as function parameters. ``` SELECT * FROM s3( 'https://private-bucket.s3.eu-west-3.amazonaws.com/orders-dataset/partition-name.csv', 'access_token', 'secret_token', 'CSVWithNames', "`order_id` UInt64, `quantity` Decimal(9, 18), `order_datetime` DateTime" ) ``` ### Query using the `s3Cluster` table function[​](#query-using-the-s3cluster-table-function "Direct link to query-using-the-s3cluster-table-function") Before you start, fulfill relevant [prerequisites](/docs/products/clickhouse/howto/run-federated-queries.md#prerequisites), if any. The `s3Cluster` function allows all cluster nodes to participate in the query execution. Using `default` for the cluster name parameter, we can compute the same aggregations as above as follows: ``` WITH ooni_clustered_data_sample AS ( SELECT * FROM s3Cluster('default', 'https://ooni-data-eu-fra.s3.eu-central-1.amazonaws.com/clickhouse_export/csv/fastpath_202308.csv.zstd') LIMIT 100000 ) SELECT probe_cc AS probe_country_code, test_name, countIf(anomaly = 't') AS total_anomalies FROM ooni_clustered_data_sample GROUP BY probe_country_code, test_name HAVING total_anomalies > 10 ORDER BY total_anomalies DESC LIMIT 50 ``` ### Query using the `url` table function[​](#query-using-the-url-table-function "Direct link to query-using-the-url-table-function") Before you start, fulfill relevant [prerequisites](/docs/products/clickhouse/howto/run-federated-queries.md#prerequisites), if any. #### SELECT and `url`[​](#select-and-url "Direct link to select-and-url") Let's query the [Growth Projections and Complexity Rankings](https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/XTAQMC\&version=4.0) dataset, courtesy of the [Atlas of Economic Complexity](https://atlas.cid.harvard.edu/) project. ``` WITH economic_complexity_ranking AS ( SELECT * FROM url('https://dataverse.harvard.edu/api/access/datafile/7259657?format=tab', 'TSV') ) SELECT replace(code, '"', '') AS `ISO country code`, growth_proj AS `Forecasted annualized rate of growth`, toInt32(replace(sitc_eci_rank, '"', '')) AS `Economic Complexity Index ranking` FROM economic_complexity_ranking WHERE year = 2021 ORDER BY `Economic Complexity Index ranking` ASC LIMIT 20 ``` #### INSERT and `url`[​](#insert-and-url "Direct link to insert-and-url") With the URL function, INSERT statements generate a POST request, which can be used to interact with APIs having public endpoints. For instance, if your application has a `ingest-csv` endpoint accepting CSV data, you can insert a row using the following statement: ``` INSERT INTO FUNCTION url('https://app-name.company-name.cloud/api/ingest-csv', 'CSVWithNames') VALUES ('column1-value', 'column2-value'); ``` ### Query a virtual table[​](#query-a-virtual-table "Direct link to Query a virtual table") Before you start, fulfill relevant [prerequisites](/docs/products/clickhouse/howto/run-federated-queries.md#prerequisites), if any. Instead of specifying the URL of the resource in every query, it's possible to create a virtual table using the URL table engine. This can be achieved by running a DDL CREATE statement similar to the following: ``` CREATE TABLE trips_export_endpoint_table ( `trip_id` UInt32, `vendor_id` UInt32, `pickup_datetime` DateTime, `dropoff_datetime` DateTime, `trip_distance` Float64, `fare_amount` Float32 ) ENGINE = URL('https://app-name.company-name.cloud/api/trip-csv-export', CSV) ``` Once the table is defined, SELECT and INSERT statements execute GET and POST requests to the URL respectively: ``` SELECT toDate(pickup_datetime) AS pickup_date, median(fare_amount) AS median_fare_amount, max(fare_amount) AS max_fare_amount FROM trips_export_endpoint_table GROUP BY pickup_date INSERT INTO trips_export_endpoint_table VALUES (8765, 10, now() - INTERVAL 15 MINUTE, now(), 50, 20) ``` Related pages * [About querying external data in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/federated-queries.md) * [Cloud Compatibility | ClickHouse Docs](https://clickhouse.com/docs/en/whats-new/cloud-compatibility#federated-queries) * [Integrating S3 with ClickHouse](https://clickhouse.com/docs/en/integrations/s3) * [remote, remoteSecure | ClickHouse Docs](https://clickhouse.com/docs/en/sql-reference/table-functions/remote) --- # Scale disk storage for your Aiven for ClickHouse® service Scale the disk storage of your Aiven for ClickHouse® service up or down without disrupting the running service. note Dynamic disk sizing (DDS) adds network-attached block storage. To move data to object storage instead, see [Tiered storage in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/clickhouse-tiered-storage.md). /eol-for-major-versions#aiven-for-flinkAdding or removing disk storage does not disrupt the running service. You pay only for extra storage instead of upgrading compute resources. You can add extra storage when you create a service or after it is running. When you add storage to a running service, the Aiven Platform provisions the extra disk and adds it to the running instances. For a clustered service such as Aiven for Apache Kafka®, Aiven divides extra storage equally between the nodes. For a shared service, each node receives the full extra capacity. ## Limitations[​](#limitations "Direct link to Limitations") * Disk added for extra storage is slower than the original disk until the next maintenance update. The slower disk can reduce performance for I/O-intensive workloads. * Maximum storage depends on the plan, service type, and cloud provider. It can be up to five times the plan's base storage size. * Cloud providers limit how many times you can increase storage between maintenance updates. If you reach the limit, run a maintenance update to optimize performance. * You cannot add storage during a maintenance update. * Dynamic disk sizing (DDS) is not supported on custom service plans. Pricing If you add storage when you create a service, **Additional disk storage** shows an estimated monthly cost. The **Service summary** lists plan storage plus additional storage. The estimated monthly price includes the additional storage cost. If you add storage to a running service, the Aiven Console shows the cost of the additional storage and related backups. The same costs appear on your invoices. ## Add or remove storage[​](#add-or-remove-storage "Direct link to Add or remove storage") ### Add storage when you create a service[​](#add-storage-when-you-create-a-service "Direct link to Add storage when you create a service") To add storage while you create a service: 1. In **Additional disk storage**, set the size with the slider or enter a value in GB. 2. Review the estimated monthly cost. 3. In the **Service summary**, click **Create service**. Change additional storage later on the running service, or enable automatic disk scaling with Aiven Autoscaler. ### Change storage on a running service[​](#change-storage-on-a-running-service "Direct link to Change storage on a running service") You cannot add or remove storage when service nodes are in the rebuilding state, for example during a maintenance update or a service upgrade. If you are removing disk storage: * Make sure the data in your service does not exceed the allocated storage. If it does, you cannot remove the additional storage. * Plan for the time it takes to rebuild the service. The time depends on the service. - Console - CLI - Terraform 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Actions** > **Manage additional storage**. 3. Change the disk storage. note * The price shown for the additional storage includes backup costs. * You can only remove storage that you previously added using this feature. To downgrade further, you can change your service plan. 4. Click **Save Changes**. Use [Aiven CLI](/docs/tools/cli.md) to add or remove additional storage using [`avn service update`](/docs/tools/cli/service-cli.md#avn-cli-service-update) with the `--disk-space-gib` flag to specify the total disk space to provide to your service. For example, if your service has a 80-GiB disk and you would like to add an extra 10-GiB disk, use: ``` avn service update --disk-space-gib 90 --project PROJECT_NAME SERVICE_NAME ``` note * When you perform a horizontal service upgrade or downgrade, remember to include all additional disks the service uses. For example, when switching from `Startup-4` to `Business-4` or from `Business-4` to `Startup-4`, include all the additional disks available for this service. * When you fork an existing service, include all additional disks the service uses. Use the `additional_disk_space` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). If you added storage, the additional storage is available immediately. If you removed additional storage, the service nodes go through a rolling restart. Depending on the service type and configuration, there might be a short downtime for services with no HA capabilities. note Storage optimization is performed at the next maintenance update after a change to the storage size. Due to cloud provider limitations, there is a limit on how many times storage can be increased between two maintenance updates. When this limit is reached, perform a maintenance update for performance optimization. Plan increases to avoid reaching this limit. Related pages * [Disk autoscaler](/docs/products/clickhouse/howto/disk-autoscaler.md) * [Change the service plan](/docs/products/clickhouse/howto/change-service-plan.md) * [Tiered storage in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/clickhouse-tiered-storage.md) * [Fork your Aiven for ClickHouse® service](/docs/products/clickhouse/howto/fork-service.md) * [Get started with Aiven for ClickHouse®](/docs/products/clickhouse/get-started.md) --- # Secure a managed ClickHouse® service You can secure your Aiven for ClickHouse® service in a few different ways, for example by restricting network access, using Virtual Private Cloud (VPC), and enabling service termination protection. ## Restrict network access to your service[​](#restrict-network-access-to-your-service "Direct link to Restrict network access to your service") One of the most fundamental ways to keep your service secure is managing its network access properly. By default the service is publicly available but you can restrict the access by following the instruction in [Restrict network access to your service](/docs/platform/howto/restrict-access.md). ## Use Virtual Private Cloud (VPC)[​](#use-virtual-private-cloud-vpc "Direct link to Use Virtual Private Cloud (VPC)") With VPC, no public internet-based access is provided to the service and it can only be connected to from the customer's peered VPC using a private network address. Read more on using VPC in [Networking with VPC peering](/docs/platform/concepts/cloud-security.md#networking-with-vpc-peering) and see [Configure VPC peering](/docs/platform/howto/manage-project-vpc.md#create-a-project-vpc). ## Protect a service from termination[​](#protect-a-service-from-termination "Direct link to Protect a service from termination") Aiven services can be protected against accidental deletion or powering off by enabling the Termination Protection feature. note Termination Protection has no effect on service migrations or upgrades. ### Enable the termination protection[​](#enable-the-termination-protection "Direct link to Enable the termination protection") 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose your Aiven for ClickHouse service. 2. On the **Overview** page, click **Service settings** in the service sidebar. 3. On the **Service settings** page, go to the **Service status** section, and click **Actions** > **Enable termination protection**. Termination Protection is enabled for your service: It cannot be terminated or powered down from the Aiven Console, via the Aiven REST API, or by using the Aiven command-line client. ### Terminate a protected service[​](#terminate-a-protected-service "Direct link to Terminate a protected service") Before terminating or powering off a protected service, disable Termination Protection for this service. note Running out of free Aiven sign-up credits causes the service to be powered down unless a credit card has been entered for the project. --- # Set up Kafka topic querying in Aiven for ClickHouse® Send data from an Aiven for Apache Kafka® topic to Aiven for ClickHouse® and query it with SQL. Aiven creates a Kafka-to-ClickHouse integration for the selected topic and opens the ClickHouse query editor with a generated query. For more information about how the integration works, supported schemas, ingestion start points, schema changes, and limitations, see [Query Kafka topic data in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/query-kafka-topic-data.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before you begin, make sure you have: * An Aiven for Apache Kafka® service with at least one topic. * An Aiven for ClickHouse® service in the same cloud region as the Kafka service, or permission to create one during setup. * Optional: [Karapace Schema Registry](/docs/products/kafka/karapace/howto/enable-karapace.md) enabled with Avro, if you want ClickHouse columns to be auto-detected from the topic schema. note Query Kafka topic data is not available for every cloud provider. ## Send topic data to ClickHouse[​](#send-topic-data-to-clickhouse "Direct link to Send topic data to ClickHouse") ### Step 1: Start from a Kafka topic[​](#step-1-start-from-a-kafka-topic "Direct link to Step 1: Start from a Kafka topic") 1. In the [Aiven Console](https://console.aiven.io), select your Aiven for Apache Kafka® service. 2. Click **Topics**. 3. Click the topic name to open the topic info panel. 4. Start the setup in one of the following ways: * If the topic has no active ClickHouse table integration, in **Analyze your data in minutes**, click **Query in ClickHouse**. * To add another ClickHouse table integration, in **ClickHouse tables**, click **Add table**. ### Step 2: Choose a ClickHouse service[​](#step-2-choose-a-clickhouse-service "Direct link to Step 2: Choose a ClickHouse service") Choose the Aiven for ClickHouse® service to receive the topic data. To use an existing service: 1. Choose an Aiven for ClickHouse® service with the **Running** status. 2. Click **Continue**. To create a service during setup: 1. Click **Create ClickHouse service**. 2. Enter a service name. 3. Choose a service plan. 4. Click **Create**. If you need more configuration options, click **Go to the full service creation**. ### Step 3: Configure the ClickHouse table[​](#step-3-configure-the-clickhouse-table "Direct link to Step 3: Configure the ClickHouse table") 1. Choose where ingestion starts in the Kafka topic. You can send all messages from the first offset, only new messages, or messages from a recent time range. The available options depend on your topic and service configuration. 2. Configure the table schema: * **If the Console detects a schema:** Review the schema preview. The Console maps the schema fields to suggested ClickHouse column names and data types. * **If the Console does not detect a schema:** Add the ClickHouse columns manually. Enter each column name, choose the ClickHouse data type, and set **Nullable** as needed. 3. Optional: If Aiven suggested a schema, enable **Override column definitions** and update the columns. 4. In **Order by**, choose the column used to sort the ClickHouse table. important You cannot change the **Order by** column after the table is created. 5. Optional: Expand **Advanced configuration** and review the generated settings. You can review or update settings such as the table name, consumer group name, view name, TTL, TTL column, and local disk TTL. 6. Click **Deploy**. ### Step 4: Query the ClickHouse table[​](#step-4-query-the-clickhouse-table "Direct link to Step 4: Query the ClickHouse table") 1. Wait until deployment completes. Aiven creates the Kafka-to-ClickHouse integration and the required ClickHouse resources. The setup page may show a preview of ingested rows when data starts flowing. 2. Click **Query**. The ClickHouse **Data** > **Query editor** opens with a generated `SELECT` query for the table created during setup. 3. Review the generated SQL query. 4. Click **Execute**. 5. View the integration on the **Data** > **Integrations** page of either the Kafka service or the ClickHouse service. ## Troubleshoot[​](#troubleshoot "Direct link to Troubleshoot") ### Query in ClickHouse option not visible[​](#query-in-clickhouse-option-not-visible "Direct link to Query in ClickHouse option not visible") If **Query in ClickHouse** does not appear in **Analyze your data in minutes**, the topic already has an active ClickHouse table integration. To add another, click **Add table** in **ClickHouse tables**. If **Add table** is also not visible, your Kafka service may be on a cloud provider where this feature is not yet available. Use a Kafka service on a supported cloud provider. ### ClickHouse service not visible[​](#clickhouse-service-not-visible "Direct link to ClickHouse service not visible") If you do not see the ClickHouse service you expect, ensure: * The ClickHouse service has the **Running** status. * The ClickHouse service is in the same cloud region as the Kafka service. * You have access to the ClickHouse service. * ClickHouse is available in the Kafka service cloud region. If ClickHouse is not available in the Kafka service cloud region, migrate the Kafka service to a supported cloud region. ### Data not appearing after deployment[​](#data-not-appearing-after-deployment "Direct link to Data not appearing after deployment") If data does not appear in the ClickHouse table after deployment, review **Observe** > **Logs** for the Aiven for ClickHouse® service. Ingestion errors are reported on the ClickHouse side. Also ensure the selected ingestion start point includes messages from the topic. For example, if you selected **New messages only**, only messages produced after the integration was created are sent to ClickHouse. For more information about ingestion behavior and limitations, see [Limitations](/docs/products/clickhouse/concepts/query-kafka-topic-data.md#limitations). Related pages * [Query Kafka topic data in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/query-kafka-topic-data.md) * [Connect Apache Kafka® to Aiven for ClickHouse®](/docs/products/clickhouse/howto/integrate-kafka.md) * [Set up Aiven for ClickHouse® data service integrations](/docs/products/clickhouse/howto/data-service-integration.md) --- # Use SQL user defined functions in Aiven for ClickHouse® Use SQL user defined functions (UDFs) in Aiven for ClickHouse® to speed up your queries and optimize your application performance. You can define your own SQL UDFs, which are automatically replicated to all nodes in the cluster and contained in backups. note Aiven for ClickHouse supports SQL UDFs only. UDF types other than SQL UDFs, such as [executable UDFs](https://clickhouse.com/docs/en/sql-reference/functions/udf#executable-user-defined-functions), are not supported in Aiven for ClickHouse. ## Create a UDF in SQL[​](#create-a-udf-in-sql "Direct link to Create a UDF in SQL") To create an SQL UDF in Aiven for ClickHouse, run the `CREATE FUNCTION` expression including function parameters, constants, operators, or other function calls. ``` CREATE FUNCTION name AS (parameter0, ...) -> expression ``` ## Limitations[​](#limitations "Direct link to Limitations") * The name of your UDF needs to be unique among other functions. * Recursive functions are not supported. * All variables used by your UDF need to be defined in its parameter list. ## Example of using SQL UDFs[​](#example-of-using-sql-udfs "Direct link to Example of using SQL UDFs") 1. Create an SQL UDF. ``` CREATE FUNCTION is_weekend AS (date) -> toDayOfWeek(date) IN (6, 7); ``` 2. Use your `is_weekend` SQL UDF in a `SELECT` query. ``` SELECT AVG(profit) FROM sales WHERE is_weekend(date) ``` --- # Tag your Aiven for ClickHouse® service Add key-value tags to your Aiven for ClickHouse® service to organize services and track ownership, cost allocation, and governance. Use tags to add metadata to Aiven services to categorize them or run custom logic on them. Typical uses include: * Tagging for governance to deploy services with specific tags only. * Tagging for internal cost reporting, ownership, allocation, and accountability. A tag is a key/value pair: * **Key**: A case-sensitive string that starts with a letter and consists of letters, numbers, dashes, and underscores. The maximum length for a key is 64 characters. * **Value**: A string value limited to 64 UTF-8 characters. Within a service, the tag keys must be unique. * Console * Terraform 1. In the service, click **Service settings**. 2. In the **Service status** section, click **Actions** > **Add service tags**. 3. Enter a key and value for each tag. 4. Click **Save changes**. Use the `tag` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Power on/off and delete your Aiven for ClickHouse® service](/docs/products/clickhouse/howto/power-cycle-service.md) --- # Transfer data between storage devices in Aiven for ClickHouse®'s tiered storage Moving data from network-attached block storage to object storage allows you to size down your block storage by selecting a service plan with less capacity. You can move the data back to network-attached block storage anytime. You can transfer data between storage devices in Aiven for ClickHouse® using SQL statements against your tables directly. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * At least one Aiven for ClickHouse service * Command line tool ([ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md)) installed ## Transfer data from network-attached block storage to object storage[​](#transfer-data-from-network-attached-block-storage-to-object-storage "Direct link to Transfer data from network-attached block storage to object storage") * Automatic data transfer * Manual data transfer If you [enable](/docs/products/clickhouse/howto/enable-tiered-storage.md) the tiered storage feature on your table, by default your data is moved from network-attached block storage to object storage as soon as it reaches 80% of its capacity. 1. [Connect to your Aiven for ClickHouse service](/docs/products/clickhouse/howto/list-connect-to-service.md) using, for example, the ClickHouse client. 2. Run the following query: ``` ALTER TABLE database-name.tablename MODIFY SETTING storage_policy = 'tiered' ``` Now, with the tiered storage feature [enabled](/docs/products/clickhouse/howto/enable-tiered-storage.md), your data is moved from network-attached block storage to object storage when it reaches 80% of its capacity. note You can also [configure your tiered storage](/docs/products/clickhouse/howto/configure-tiered-storage.md) so that data is moved to object storage at a specific time. To move data manually from network-attached block storage to object storage, run ``` ALTER TABLE table_name MOVE PARTITION partition_expr TO VOLUME 'remote' ``` To configure data retention thresholds to automatically move data from network-attached block storage to object storage, see [Configure data retention thresholds in Aiven for ClickHouse®'s tiered storage](/docs/products/clickhouse/howto/configure-tiered-storage.md). ## Transfer data from object storage to network-attached block storage[​](#transfer-data-from-object-storage-to-network-attached-block-storage "Direct link to Transfer data from object storage to network-attached block storage") Use the [MOVE PARTITION|PART](https://clickhouse.com/docs/en/sql-reference/statements/alter/partition#move-partitionpart) statement to transfer data to network-attached block storage. 1. [Connect to your Aiven for ClickHouse service](/docs/products/clickhouse/howto/list-connect-to-service.md) using, for example, the ClickHouse client. 2. Select a database for operations you intend to perform. ``` USE database-name ``` 3. Run the following query: ``` ALTER TABLE table_name MOVE PARTITION partition_expr TO VOLUME 'default' ``` Your data has been moved to network-attached block storage. ## What's next[​](#whats-next "Direct link to What's next") * [Check data distribution between network-attached block storage and object storage](/docs/products/clickhouse/howto/check-data-tiered-storage.md) * [Configure data retention thresholds for tiered storage](/docs/products/clickhouse/howto/configure-tiered-storage.md) Related pages * [About tiered storage in Aiven for ClickHouse](/docs/products/clickhouse/concepts/clickhouse-tiered-storage.md) * [Enable tiered storage in Aiven for ClickHouse](/docs/products/clickhouse/howto/enable-tiered-storage.md) --- # Enable reading and writing data across shards in Aiven for ClickHouse® If your Aiven for ClickHouse® service uses multiple shards, the data is replicated only between nodes of the same shard. When you read from one node, you see the data from the shard of this node. Creating a [distributed table](https://clickhouse.com/docs/en/engines/table-engines/special/distributed/) on top of your replicated table, while the distributed table itself doesn't store data allows you to see data from all the shards but also help spread the data evenly across all the cluster nodes. ## Set up a sharded service with a database[​](#set-up-a-sharded-service-with-a-database "Direct link to Set up a sharded service with a database") 1. [Create an Aiven for ClickHouse® service](/docs/products/clickhouse/get-started.md#create-an-aiven-for-clickhouse-service) with multiple shards. note Shards are created automatically when you create an Aiven for ClickHouse services. The number of shards that your service gets depends on the plan you select for your service. You can calculate the shards number as follows: *number of shards = number of nodes / 3*. 2. [Create database](/docs/products/clickhouse/howto/manage-databases-tables.md#create-a-clickhouse-database) `test_db` in your new service. ## Create a distributed table[​](#create-a-distributed-table "Direct link to Create a distributed table") 1. [Connect to your database](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md). 2. Create a table with the [MergeTree engine](https://clickhouse.com/docs/en/engines/table-engines/mergetree-family/mergetree/) as shown for the `cash_flows` table in the following example: ``` CREATE TABLE test_db.cash_flows ( EventDate DateTime, SourceAccount UInt64, TargetAccount UInt64, Amount Float64 ) ENGINE = MergeTree() PARTITION BY toYYYYMM(EventDate) ORDER BY (EventDate, SourceAccount) ``` note With Aiven for ClickHouse, you can specify `ENGINE` either as `MergeTree` or as `ReplicationMergeTree`. Both of them create a `ReplicationMergeTree` table. 3. Create distributed table `cash_flows_distributed` with the distributed engine: ``` CREATE TABLE test_db.cash_flows_distributed AS test_db.cash_flows ENGINE = Distributed(test_db, test_db, cash_flows, SourceAccount) ``` ## Verify your distributed table[​](#verify-your-distributed-table "Direct link to Verify your distributed table") Check if the distributed table you created is available and if you can use it to access your data from all the shards. 1. Run a read query for the number of table rows: ``` SELECT count() FROM test_db.cash_flows_distributed ``` As a response to this query, you can expect to receive a number of rows from all the shards. This is because when you connect on one node and read from the distributed table, ClickHouse® aggregates the data from all the shards and returns all of it. 2. Run a write query to insert new data into the distributed table: ``` INSERT INTO test_db.cash_flows_distributed ( EventDate, SourceAccount, TargetAccount, Amount ) VALUES ( '2022-01-02 03:04:05', 123, 456, 100.0 ) ``` When you insert data into the distributed table, ClickHouse® decides on which node the data should be stored and write it to the correct node making sure that a similar volume of data is written on all the nodes. --- # Tune the vector similarity index cache in Aiven for ClickHouse® Tune the vector similarity index cache in Aiven for ClickHouse® to improve vector search performance for Hierarchical Navigable Small World (HNSW) indexes. HNSW is a graph-based index that speeds up approximate nearest neighbor search on vector columns. Aiven for ClickHouse uses a segmented least recently used (SLRU) cache to keep HNSW index data in memory during vector search queries. If the cache evicts index data too often, ClickHouse reloads it from disk during queries, which can significantly increase query latency. ## How the cache works[​](#how-the-cache-works "Direct link to How the cache works") The vector similarity index cache uses two advanced configuration settings: * `server_settings.vector_similarity_index_cache_size` sets the total cache size as a fraction of server memory. * `server_settings.vector_similarity_index_cache_size_ratio` controls the fraction of that cache allocated to the protected segment of the SLRU cache. The SLRU cache has two segments: * **Protected segment**: Stores frequently used index entries. * **Probationary segment**: Stores newly loaded or less frequently used index entries. ClickHouse calculates the protected segment size by multiplying the cache ratio by the total cache size. ClickHouse stores HNSW indexes per table part, a chunk of table data on disk. To avoid repeated evictions, the protected segment must be large enough to hold the largest per-part HNSW index used by your queries. note Table part structure affects cache behavior. A table with many smaller parts might cache more efficiently than one with a single large merged part because each smaller per-part HNSW index is more likely to fit within the protected segment. Monitor part structure if you continue to see evictions after increasing the ratio. ## Recommended cache ratio[​](#recommended-cache-ratio "Direct link to Recommended cache ratio") The Aiven default for `server_settings.vector_similarity_index_cache_size_ratio` is `0.1`. For vector search workloads, start with a value of at least `0.4`. caution Vector search performance does not degrade gradually when the ratio is too low. If the protected segment cannot hold the largest per-part HNSW index, ClickHouse evicts and reloads index entries repeatedly, causing a sharp increase in query latency. In Aiven benchmarks, values below `0.4` caused evictions and significantly higher query latency. Increase the value if you still see vector similarity index cache evictions during steady-state queries. The best value depends on: * The size of the HNSW indexes. * The number and size of table parts. * The total vector similarity index cache size. * Query concurrency and access patterns. * Table merges that create larger parts over time. As a rule of thumb, set the ratio so the protected SLRU segment can hold the largest per-part HNSW index without eviction. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven for ClickHouse service running ClickHouse 25.8 or later. * A table with an HNSW vector similarity index and an active vector search workload. * Access to the [Aiven Console](https://console.aiven.io/) to update advanced configuration. * A SQL client, such as the [ClickHouse client](/docs/products/clickhouse/howto/connect-with-clickhouse-cli.md), to run the monitoring queries. ## Configure the cache ratio[​](#configure-the-cache-ratio "Direct link to Configure the cache ratio") You can update `server_settings.vector_similarity_index_cache_size_ratio` without restarting the service. 1. Log in to the [Aiven Console](https://console.aiven.io/) and choose your Aiven for ClickHouse service. 2. In the service sidebar, click **Service settings**. 3. In the **Advanced configuration** section, click **Configure**. 4. Click **Add configuration options**. 5. Search for `server_settings.vector_similarity_index_cache_size_ratio`. 6. Set the value to `0.4` or higher. 7. Click **Save configuration**. Aiven applies the setting without a service restart. The total cache size usually stays at its default. Adjust `server_settings.vector_similarity_index_cache_size` only if evictions continue after raising the ratio. ## Monitor cache health[​](#monitor-cache-health "Direct link to Monitor cache health") Use `system.events` to monitor vector similarity index cache activity: ``` SELECT event, value FROM system.events WHERE event LIKE '%VectorSimilarity%' ORDER BY event; ``` Monitor `VectorSimilarityIndexCacheWeightLost`. If this value is greater than `0` during steady-state vector search queries, ClickHouse evicts vector index data. Increase `server_settings.vector_similarity_index_cache_size_ratio` and run the workload again. note Do not rely on hit rate alone to assess cache health. A high hit rate can coexist with evictions if ClickHouse reloads and reinserts indexes repeatedly. For stable vector search performance, keep `VectorSimilarityIndexCacheWeightLost` at `0` during steady-state queries. ## Troubleshoot slow vector search queries[​](#troubleshoot-slow-vector-search-queries "Direct link to Troubleshoot slow vector search queries") If vector search queries are slower than expected: 1. Verify that the query uses the HNSW vector similarity index. 2. Run the `system.events` query to monitor vector similarity index cache activity. 3. Confirm whether `VectorSimilarityIndexCacheWeightLost` increases during steady-state queries. 4. If evictions occur, increase `server_settings.vector_similarity_index_cache_size_ratio` to at least `0.4`. 5. Monitor query latency and cache events again. If evictions continue after increasing the ratio, the protected segment might still be too small for the largest per-part HNSW index. Review the total vector similarity index cache size and the table part structure. Related pages * [Indexing and data processing in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/indexing.md) * [Advanced parameters for Aiven for ClickHouse®](/docs/products/clickhouse/reference/advanced-params.md) * [Use query cache in Aiven for ClickHouse®](/docs/products/clickhouse/howto/clickhouse-query-cache.md) * [Fetch query statistics for Aiven for ClickHouse®](/docs/products/clickhouse/howto/fetch-query-statistics.md) --- # Aiven for ClickHouse® 25.8 default settings Aiven for ClickHouse® uses a managed configuration that differs from upstream ClickHouse defaults. These differences help keep services reliable, secure, and predictable in Aiven-managed environments. The following sections list the settings where Aiven defaults differ from defaults in version 25.8, grouped by session, table, and server scope. ## How to read the settings tables[​](#how-to-read-the-settings-tables "Direct link to How to read the settings tables") Each setting includes an Aiven default value and indicates whether you can change it: * **Can be changed: No**: Aiven manages the setting value. You cannot override it. * **Can be changed: Yes**: Aiven sets the default value, but you can change it where the setting is configurable. Where Aiven defaults vary by service plan or node resources, the **Aiven default** column shows ranges or placeholders. The **Description** column explains why the default differs or lists any configurable bounds, such as an allowed range or minimum value. Values shown as ranges or placeholders, such as `[3..80, depending on CPU count]`, `{cpu_count}`, or `[~7% of RAM]`, are sized automatically based on your service plan and node resources. note For configurable settings and Aiven-defined limits, see [Advanced parameters for Aiven for ClickHouse®](/docs/products/clickhouse/reference/advanced-params.md) and [Limits and limitations](/docs/products/clickhouse/reference/limitations.md). ## Session settings[​](#session-settings "Direct link to Session settings") These settings apply to sessions and queries. | Setting | Aiven default | Can be changed | Description | | ------------------------------------------------------- | --------------------------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `allow_deprecated_error_prone_window_functions` | `0` | No | Keeps deprecated, error-prone window functions disabled. | | `allow_deprecated_snowflake_conversion_functions` | `0` | No | Keeps deprecated Snowflake conversion functions disabled. | | `allow_experimental_kafka_offsets_storage_in_keeper` | `0` | No | Keeps experimental Kafka offset storage in Keeper disabled. | | `allow_introspection_functions` | `0` | Yes | Keeps sensitive debugging and introspection functions disabled by default. | | `allow_non_default_profile` | `0` | No | Prevents switching to unmanaged profiles. | | `cancel_http_readonly_queries_on_client_close` | `1` | Yes | Cancels read-only HTTP queries when clients disconnect. | | `cluster_function_process_archive_on_multiple_nodes` | `0` | Yes | Disables archive processing on multiple nodes by default. | | `compatibility` | `""` | Yes | Does not force a lower compatibility profile by default. | | `database_replicated_allow_replicated_engine_arguments` | `1` | No | Supports replicated database and table creation in Aiven-managed services. | | `distributed_ddl_entry_format_version` | `7` | No | Pins the distributed DDL entry format to avoid inconsistencies during version upgrades. | | `distributed_ddl_output_mode` | `null_status_on_timeout` | Yes | Makes distributed DDL timeout reporting more resilient to recoverable errors. | | `format_display_secrets_in_show_and_select` | `0` | Yes | Avoids exposing secrets in `SHOW` and `SELECT` output by default. | | `http_write_exception_in_output_format` | `0` | Yes | Keeps exceptions out of the requested output format so client errors are easier to detect. | | `max_autoincrement_series` | `1000` | No | Limits automatic sequence expansion. | | `max_concurrent_queries_for_all_users` | `[100..800, depending on service size]` | No | Enforces a service-wide concurrency limit. | | `max_http_get_redirects` | `10` | Yes | Limits HTTP redirect following. | | `max_insert_threads` | `2` | Yes | Allowed range: `1..{cpu_count}`. Caps insert parallelism so a single query does not use too many shared service resources. | | `max_threads` | `{cpu_count}` | Yes | Allowed range: `1..{cpu_count}`. Caps general query parallelism to the node CPU count. | | `memory_profiler_sample_probability` | `0` | No | Disables random memory profiler sampling. | | `memory_profiler_step` | `4194304` | Yes | Minimum value: `100000`. Keeps the upstream memory profiler granularity while preventing very low values. | | `min_free_disk_ratio_to_perform_insert` | `0.05` | No | Keeps a disk headroom guard before accepting inserts. | | `output_format_json_quote_64bit_integers` | `1` | Yes | Keeps JSON output compatible with clients that cannot safely represent 64-bit integers. This default can change in a future version. | | `postgresql_connection_attempt_timeout` | `10` | Yes | Limits each PostgreSQL connection attempt. | | `postgresql_connection_pool_connect_timeout` | `10` | Yes | Limits pooled PostgreSQL connection setup. | | `push_external_roles_in_interserver_queries` | `0` | No | External authorization is not supported. Roles must be synchronized between nodes. | | `query_profiler_cpu_time_period_ns` | `1000000000` | Yes | Minimum value: `1000000`. Keeps the upstream CPU profiler sampling period while preventing sub-millisecond values. | | `query_profiler_real_time_period_ns` | `1000000000` | Yes | Minimum value: `1000000`. Keeps the upstream real-time profiler sampling period while preventing sub-millisecond values. | | `readonly` | `0` | Yes | Controls read/write access for the session. You can change the value from `0` to `1`, but not from `1` to `0`. | | `stream_like_engine_allow_direct_select` | `1` | Yes | Allows direct `SELECT` queries on stream-like engines, which can have suboptimal performance or behavior. A future release returns this default to the upstream default. | | `write_full_path_in_iceberg_metadata` | `1` | Yes | Writes full Iceberg metadata paths for compatibility with mainstream ClickHouse behavior. | ## Replicated MergeTree table settings[​](#replicated-mergetree-table-settings "Direct link to Replicated MergeTree table settings") These settings apply to the `ReplicatedMergeTree` table engine family. On the Aiven platform, `MergeTree` engines are remapped to their `ReplicatedMergeTree` variants. See [Supported table engines](/docs/products/clickhouse/reference/supported-table-engines.md). | Setting | Aiven default | Can be changed | Description | | --------------------------------------------------------------------- | ------------- | -------------- | -------------------------------------------------------------------------------------------------------------------------- | | `allow_remote_fs_zero_copy_replication` | `1` | No | Enables zero-copy replication for remote filesystem storage. Aiven manages this storage behavior. | | `disable_detach_partition_for_zero_copy_replication` | `0` | No | Keeps `DETACH PARTITION` available for zero-copy replicated tables. | | `disable_fetch_partition_for_zero_copy_replication` | `0` | No | Keeps `FETCH PARTITION` available for zero-copy replicated tables. | | `disable_freeze_partition_for_zero_copy_replication` | `0` | No | Keeps `FREEZE PARTITION` available for zero-copy replicated tables. | | `enable_max_bytes_limit_for_min_age_to_force_merge` | `1` | No | Enforces the byte-size limit when age-based forced merges are considered. This helps prevent creation of very large parts. | | `finished_mutations_to_keep` | `10` | Yes | Limits retained finished mutation metadata so ZooKeeper or Keeper state does not grow unnecessarily. | | `max_parts_in_total` | `10000` | Yes | Allowed range: `100..50000`. Sets a cap on table part count to protect merge performance and metadata size. | | `number_of_free_entries_in_pool_to_execute_mutation` | `20` | Yes | Uses the upstream default while limiting the value below the service background pool capacity. | | `number_of_free_entries_in_pool_to_execute_optimize_entire_partition` | `25` | Yes | Uses the upstream default while limiting the value below the service background pool capacity. | | `number_of_free_entries_in_pool_to_lower_max_size_of_merge` | `8` | Yes | Uses the upstream default while limiting the value below the service background pool capacity. | | `old_parts_lifetime` | `60` | Yes | Reduces the lifetime of outdated merged parts to lower the number of ZooKeeper or Keeper metadata nodes. | ## Server settings[​](#server-settings "Direct link to Server settings") These settings are configured at the server level and are managed by Aiven. They are not configurable as session settings. Most settings are derived from your service plan. Some settings, such as `vector_similarity_index_cache_size`, are exposed through advanced configuration. | Setting | Aiven default | Description | | ----------------------------------------------- | --------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `background_pool_size` | `[8..32, depending on CPU count]` | Sizes the background merge and mutation pool based on service CPU capacity. If an explicit service-level override is configured, Aiven preserves it. | | `cgroups_memory_usage_observer_wait_time` | `0` | Disables the ClickHouse cgroup memory usage observer because Aiven sets explicit memory limits for the managed server process. | | `cluster_database` | `default` | Sets the database used by cluster-related helpers to the managed default database. | | `database_atomic_delay_before_drop_table_sec` | `0` | Removes the Atomic database delayed-drop wait to prevent race conditions in refreshable materialized views. | | `default_database` | `default` | Uses the managed default database unless another database is selected explicitly. | | `default_replica_name` | `{replica}` | Uses the Aiven-provided replica macro in replicated table paths. | | `dictionary_user` | `avnadmin` | Runs dictionary queries through the managed admin user. | | `disable_internal_dns_cache` | `1` | Prevents ClickHouse from keeping stale internal DNS results in a managed environment where node addresses can change. | | `display_secrets_in_show_and_select` | `1` | Keeps server-level secret display available for privileged managed operations. User-facing profile settings hide secrets by default. | | `iceberg_catalog_threadpool_pool_size` | `[3..80, depending on CPU count]` | Scales Iceberg catalog worker concurrency with service CPU capacity. | | `iceberg_metadata_files_cache_size` | `[~7% of RAM]` | Sizes the Iceberg metadata file cache based on the managed memory budget. | | `index_mark_cache_size` | `[~7% of RAM]` | Sizes the secondary index mark cache based on the managed memory budget. | | `keep_alive_timeout` | `10` | Limits idle keep-alive connections to 10 seconds so clients cannot hold server resources indefinitely. | | `mark_cache_size` | `[~7% of RAM]` | Sizes the primary mark cache based on the managed memory budget. | | `max_concurrent_queries` | `[147..847, depending on service size]` | Sets a server-wide concurrency limit based on the service plan, with extra internal capacity reserved for monitoring, backup, and operator queries. | | `max_connections` | `[1000..4000, depending on service size]` | Sets the connection limit based on service size. Smaller nodes have lower limits to prevent overload, and larger service plans allow higher limits. | | `max_database_num_to_throw` | `400` | Enforces a hard limit on the number of databases to protect metadata and reduce operational overhead. | | `max_database_num_to_warn` | `200` | Warns before the hard database-count limit is reached. | | `max_format_parsing_thread_pool_size` | `[3..80, depending on CPU count]` | Scales the format parsing thread pool with service CPU capacity. | | `max_partition_size_to_drop` | `0` | Removes the server-side partition-size guard for drops to avoid common operational issues. | | `max_prefixes_deserialization_thread_pool_size` | `[3..80, depending on CPU count]` | Scales the prefixes deserialization thread pool with service CPU capacity. | | `max_server_memory_usage` | `[65%..70% of host RAM, depending on service size]` | Caps ClickHouse memory use below total host RAM so the node retains memory for the operating system, caching, and service management operations. | | `max_table_size_to_drop` | `0` | Removes the server-side table-size guard for drops to avoid common operational issues. | | `memory_worker_correct_memory_tracker` | `1` | Lets the ClickHouse memory worker correct memory tracking drift. | | `prefetch_threadpool_pool_size` | `[3..80, depending on CPU count]` | Scales the prefetch thread pool with service CPU capacity. | | `prepare_system_log_tables_on_startup` | `1` | Ensures system log tables are ready immediately after startup for managed observability and diagnostics. | | `reserved_replicated_database_prefixes` | `["aiven", "endpoint_", "service_"]` | Reserves internal database name prefixes for Aiven-managed objects. | | `series_keeper_path` | `/clickhouse/series` | Sets the Keeper path for `generateSerialID` counter nodes. | | `show_addresses_in_stack_traces` | `0` | Avoids exposing raw addresses in stack traces returned to users or logs. | | `threadpool_local_fs_reader_pool_size` | `[3..80, depending on CPU count]` | Scales local filesystem reader concurrency with service CPU capacity. | | `threadpool_remote_fs_reader_pool_size` | `[3..80, depending on CPU count]` | Scales remote filesystem reader concurrency with service CPU capacity. | | `threadpool_writer_pool_size` | `[3..80, depending on CPU count]` | Scales filesystem writer concurrency with service CPU capacity. | | `uncompressed_cache_size` | `[~7% of RAM]` | Sizes the uncompressed block cache based on the managed memory budget. | | `user_with_indirect_database_creation` | `avnadmin` | Restricts indirect database creation privileges to the managed admin user. | | `vector_similarity_index_cache_size` | `0.07` | Sizes the vector similarity index cache as a fraction of total server memory. The default is `0.07`, or 7% of RAM. Set to `0` to disable the cache. Maximum value: `0.5`. This setting can be changed in advanced configuration and applies to ClickHouse 25.8 or later. | --- # Aiven for ClickHouse® 26.3 default settings Aiven for ClickHouse® uses a managed configuration that differs from upstream ClickHouse defaults. These differences help keep services reliable, secure, and predictable in Aiven-managed environments. Aiven applies conservative defaults and constraints where upstream behavior can affect query correctness, upgrade compatibility, or shared resource isolation. The following sections list the settings where Aiven defaults or constraints differ from defaults in version 26.3, grouped by session, table, and server scope. ## How to read the settings tables[​](#how-to-read-the-settings-tables "Direct link to How to read the settings tables") Each setting includes an Aiven default value and indicates whether you can change it: * **Can be changed: No**: Aiven manages the setting value. You cannot override it. * **Can be changed: Yes**: Aiven sets the default value, but you can change it where the setting is configurable. Where Aiven defaults vary by service plan or node resources, the **Aiven default** column shows ranges or placeholders. The **Description** column explains why the default differs or lists any configurable bounds, such as an allowed range or minimum value. Values shown as ranges or placeholders, such as `[3..80, depending on CPU count]`, `{cpu_count}`, or `[~7% of RAM]`, are sized automatically based on your service plan and node resources. note For configurable settings and Aiven-defined limits, see [Advanced parameters for Aiven for ClickHouse®](/docs/products/clickhouse/reference/advanced-params.md) and [Limits and limitations](/docs/products/clickhouse/reference/limitations.md). ## Session settings[​](#session-settings "Direct link to Session settings") These settings apply to sessions and queries. | Setting | Aiven default | Can be changed | Description | | ----------------------------------------------------------- | --------------------------------------- | -------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | | `allow_deprecated_error_prone_window_functions` | `0` | No | Keeps deprecated, error-prone window functions disabled. | | `allow_deprecated_snowflake_conversion_functions` | `0` | No | Keeps deprecated Snowflake conversion functions disabled. | | `allow_experimental_alias_table_engine` | `0` | No | Prevents use of an experimental table engine with known data-loss and upgrade-compatibility risks. | | `allow_experimental_database_paimon_rest_catalog` | `0` | No | Prevents unsupported external catalog metadata and credentials from creating security, upgrade, and restore risks. | | `allow_experimental_kafka_offsets_storage_in_keeper` | `0` | No | Keeps experimental Kafka offset storage in Keeper disabled. | | `allow_experimental_nullable_tuple_type` | `0` | No | Prevents persisted `Nullable(Tuple)` data that ClickHouse 25.8 cannot read during rollback or restore. | | `allow_experimental_object_storage_queue_hive_partitioning` | `0` | No | Prevents experimental per-partition Keeper state from affecting coordination, upgrade, and restore reliability. | | `allow_experimental_time_series_aggregate_functions` | `0` | No | Prevents use of experimental time-series aggregate functions until allocation limits are validated. | | `allow_experimental_ts_to_grid_aggregate_function` | `0` | No | Prevents use of the experimental time-series aggregate function alias until allocation limits are validated. | | `allow_fuzz_query_functions` | `0` | No | Keeps test-only query mutation functions from consuming shared resources or destabilizing the service. | | `allow_introspection_functions` | `0` | Yes | Keeps sensitive debugging and introspection functions disabled by default. | | `allow_non_default_profile` | `0` | No | Prevents switching to unmanaged profiles. | | `ast_fuzzer_any_query` | `0` | No | Prevents the test-only AST fuzzer from mutating write and DDL queries. | | `ast_fuzzer_runs` | `0` | No | Prevents test-only randomized query execution from consuming resources or destabilizing the service. | | `cancel_http_readonly_queries_on_client_close` | `1` | Yes | Cancels read-only HTTP queries when clients disconnect. | | `check_named_collection_dependencies` | `1` | No | Prevents removal of named collections that tables use, protecting table availability and managed backup and restore operations. | | `cluster_function_process_archive_on_multiple_nodes` | `0` | Yes | Disables archive processing on multiple nodes by default. | | `compatibility` | `""` | Yes | Does not set an older ClickHouse compatibility version by default. | | `correlated_subqueries_use_in_memory_buffer` | `0` | Yes | Preserves compatibility with join algorithms and prevents valid correlated subqueries from failing. | | `database_replicated_allow_replicated_engine_arguments` | `1` | No | Supports replicated database and table creation in Aiven-managed services. | | `delta_lake_snapshot_end_version` | `-1` | No | Prevents incorrect Delta Lake change data feed results because version 26.3 can silently ignore explicit end bounds on some storage backends. | | `delta_lake_snapshot_start_version` | `-1` | No | Prevents incorrect Delta Lake change data feed results because version 26.3 can silently ignore explicit start bounds on some storage backends. | | `distributed_ddl_entry_format_version` | `7` | No | Pins the distributed DDL entry format to avoid inconsistencies during version upgrades. | | `distributed_ddl_output_mode` | `null_status_on_timeout` | Yes | Makes distributed DDL timeout reporting more resilient to recoverable errors. | | `enable_materialized_cte` | `0` | No | Keeps experimental materialized common table expressions disabled to prevent known crashes and incorrect dependency ordering. | | `enable_producing_buckets_out_of_order_in_aggregation` | `0` | No | Keeps aggregation buckets ordered to prevent failures in multi-layer distributed aggregations. | | `format_display_secrets_in_show_and_select` | `0` | Yes | Avoids exposing secrets in `SHOW` and `SELECT` output by default. | | `http_max_request_header_size` | `131072` | No | Limits HTTP request headers to 128 KiB, protecting service availability by bounding memory allocated before authentication. | | `http_write_exception_in_output_format` | `0` | Yes | Keeps exceptions out of the requested output format so client errors are easier to detect. | | `max_autoincrement_series` | `1000` | No | Limits automatic sequence expansion. | | `max_concurrent_queries_for_all_users` | `[100..800, depending on service size]` | No | Enforces a service-wide concurrency limit. | | `max_http_get_redirects` | `10` | Yes | Limits HTTP redirect following. | | `max_insert_threads` | `2` | Yes | Allowed range: `1..{cpu_count}`. Caps insert parallelism so a single query does not use too many shared service resources. | | `max_threads` | `{cpu_count}` | Yes | Allowed range: `1..{cpu_count}`. Caps general query parallelism to the node CPU count. | | `memory_profiler_sample_probability` | `0` | No | Disables random memory profiler sampling. | | `memory_profiler_step` | `4194304` | Yes | Minimum value: `100000`. Keeps the upstream memory profiler granularity while preventing very low values. | | `min_free_disk_ratio_to_perform_insert` | `0.05` | No | Keeps a disk headroom guard before accepting inserts. | | `optimize_qbit_distance_function_reads` | `0` | Yes | Disables partial reads that can fail queries using nullable, `Variant`, or `Dynamic` vectors. | | `os_thread_priority` | `0` | Yes | Allowed range: `0..19`. Allows normal or lower query priority while preventing workloads from taking priority over shared service operations. | | `os_threads_nice_value_materialized_view` | `0` | Yes | Allowed range: `0..19`. Allows normal or lower materialized-view priority while protecting shared service operations. | | `os_threads_nice_value_query` | `0` | Yes | Allowed range: `0..19`. Allows normal or lower query priority while protecting shared service operations. | | `output_format_json_quote_64bit_integers` | `1` | Yes | Keeps JSON output compatible with clients that cannot safely represent 64-bit integers. | | `postgresql_connection_attempt_timeout` | `10` | Yes | Limits each PostgreSQL connection attempt to 10 seconds. | | `postgresql_connection_pool_connect_timeout` | `10` | Yes | Limits pooled PostgreSQL connection setup to 10 seconds. | | `push_external_roles_in_interserver_queries` | `0` | No | External authorization is not supported. Roles must be synchronized between nodes. | | `query_profiler_cpu_time_period_ns` | `1000000000` | Yes | Minimum value: `1000000`. Keeps the upstream CPU profiler sampling period while preventing sub-millisecond values. | | `query_profiler_real_time_period_ns` | `1000000000` | Yes | Minimum value: `1000000`. Keeps the upstream real-time profiler sampling period while preventing sub-millisecond values. | | `query_plan_convert_any_join_to_semi_or_anti_join` | `0` | Yes | Disables a join optimization that can return incorrect results when a dependent set is not ready. | | `query_plan_direct_read_from_text_index` | `0` | Yes | Disables direct text-index reads because version 26.3 can return incorrect results for nullable columns and updated parts. | | `query_plan_remove_unused_columns` | `0` | Yes | Preserves earlier query-plan behavior and avoids excessive memory use for aggregate queries with `PREWHERE` filters over wide string columns. | | `readonly` | `0` | Yes | Controls read and write access for the session. You can change the value from `0` to `1`, but not from `1` to `0`. | | `s3queue_keeper_fault_injection_probability` | `0` | No | Prevents test-only fault injection from disrupting S3Queue ingestion and managed Keeper operations. | | `stream_like_engine_allow_direct_select` | `1` | Yes | Allows direct `SELECT` queries on stream-like engines for compatibility with existing workloads. | | `write_full_path_in_iceberg_metadata` | `1` | Yes | Writes full Iceberg metadata paths for compatibility with mainstream ClickHouse behavior. | ## Replicated MergeTree table settings[​](#replicated-mergetree-table-settings "Direct link to Replicated MergeTree table settings") These settings apply to the `ReplicatedMergeTree` table engine family. On the Aiven platform, `MergeTree` engines are remapped to their `ReplicatedMergeTree` variants. See [Supported table engines](/docs/products/clickhouse/reference/supported-table-engines.md). | Setting | Aiven default | Can be changed | Description | | --------------------------------------------------------------------- | ------------- | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ | | `allow_remote_fs_zero_copy_replication` | `1` | No | Enables zero-copy replication for remote filesystem storage. Aiven manages this storage behavior. | | `disable_detach_partition_for_zero_copy_replication` | `0` | No | Keeps `DETACH PARTITION` available for zero-copy replicated tables. | | `disable_fetch_partition_for_zero_copy_replication` | `0` | No | Keeps `FETCH PARTITION` available for zero-copy replicated tables. | | `disable_freeze_partition_for_zero_copy_replication` | `0` | No | Keeps `FREEZE PARTITION` available for zero-copy replicated tables. | | `enable_max_bytes_limit_for_min_age_to_force_merge` | `1` | No | Enforces the byte-size limit when age-based forced merges are considered. This helps prevent creation of very large parts. | | `finished_mutations_to_keep` | `10` | Yes | Limits retained finished mutation metadata so ZooKeeper or Keeper state does not grow unnecessarily. | | `max_parts_in_total` | `10000` | Yes | Allowed range: `100..50000`. Sets a cap on table part count to protect merge performance and metadata size. | | `number_of_free_entries_in_pool_to_execute_mutation` | `20` | Yes | Uses the upstream default while limiting the value below the service background pool capacity. | | `number_of_free_entries_in_pool_to_execute_optimize_entire_partition` | `25` | Yes | Uses the upstream default while limiting the value below the service background pool capacity. | | `number_of_free_entries_in_pool_to_lower_max_size_of_merge` | `8` | Yes | Uses the upstream default while limiting the value below the service background pool capacity. | | `old_parts_lifetime` | `60` | Yes | Reduces the lifetime of outdated merged parts to lower the number of ZooKeeper or Keeper metadata nodes. | | `serialization_info_version` | `basic` | Yes | Preserves part compatibility with ClickHouse 25.8 during rolling upgrades, rollback, and restore operations. Keep this value until all nodes run 26.3. | | `table_readonly` | `0` | No | Keeps an unsupported mode disabled because it can block managed operations and distributed DDL without protecting replicated tables. | ## Server settings[​](#server-settings "Direct link to Server settings") These settings are configured at the server level and are managed by Aiven. They are not configurable as session settings. Most settings are derived from your service plan. Some settings, such as `vector_similarity_index_cache_size`, are exposed through advanced configuration. | Setting | Aiven default | Description | | ----------------------------------------------- | --------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | | `aiven_enable_replication_queue_size_limit` | `1` | Preserves replication queue limits that delay or reject inserts before queue growth affects service availability. | | `aiven_enforce_default_replication_path` | `1` | Keeps replicated tables in Aiven-managed Keeper paths to preserve service isolation and reliable lifecycle operations. | | `aiven_prohibit_tmp_table_creation` | `1` | Reserves internal `.tmp` table names so your DDL cannot interfere with managed table operations. | | `aiven_replace_mergetree_with_replicated` | `1` | Converts `MergeTree` engines to replicated variants in `Replicated` databases, preserving managed high availability. | | `aiven_skip_azure_container_creation` | `1` | Skips Azure container probing during DDL replay so backup and restore can proceed without weakening data-access authentication. | | `background_pool_size` | `[8..32, depending on CPU count]` | Sizes the background merge and mutation pool based on service CPU capacity. If an explicit service-level override is configured, Aiven preserves it. | | `background_schedule_pool_size` | `[24..512, depending on CPU count]` | Scales lightweight background scheduling with service CPU capacity and avoids excessive idle threads on smaller plans. | | `cgroups_memory_usage_observer_wait_time` | `0` | Disables the ClickHouse cgroup memory usage observer because Aiven sets explicit memory limits for the managed server process. | | `cluster_database` | `default` | Sets the database used by cluster-related helpers to the managed default database. | | `database_atomic_delay_before_drop_table_sec` | `0` | Removes the Atomic database delayed-drop wait to prevent race conditions in refreshable materialized views. | | `database_replicated.internal_replication` | `1` | Routes `Distributed` inserts once per shard and lets `ReplicatedMergeTree` handle replica distribution, preventing duplicate writes. | | `default_database` | `default` | Uses the managed default database unless another database is selected explicitly. | | `default_replica_name` | `{replica}` | Uses the Aiven-provided replica macro in replicated table paths. | | `dictionary_user` | `avnadmin` | Runs dictionary queries through the managed admin user. | | `disable_internal_dns_cache` | `1` | Prevents ClickHouse from keeping stale internal DNS results in a managed environment where node addresses can change. | | `display_secrets_in_show_and_select` | `1` | Keeps server-level secret display available for privileged managed operations. User-facing profile settings hide secrets by default. | | `enforce_https_for_url_storage` | `1` | Requires encrypted transport for URL storage sources and HTTP dictionaries. | | `iceberg_catalog_threadpool_pool_size` | `[3..80, depending on CPU count]` | Scales Iceberg catalog worker concurrency with service CPU capacity. | | `iceberg_metadata_files_cache_size` | `[~7% of RAM]` | Sizes the Iceberg metadata file cache based on the managed memory budget. | | `index_mark_cache_size` | `[~7% of RAM]` | Sizes the secondary index mark cache based on the managed memory budget. | | `keep_alive_timeout` | `10` | Limits idle keep-alive connections to 10 seconds so clients cannot hold server resources indefinitely. | | `load_marks_threadpool_pool_size` | `[3..80, depending on CPU count]` | Scales mark-loading concurrency with service CPU capacity. | | `mark_cache_size` | `[~7% of RAM]` | Sizes the primary mark cache based on the managed memory budget. | | `max_concurrent_queries` | `[147..847, depending on service size]` | Sets a server-wide concurrency limit based on the service plan, with extra internal capacity reserved for monitoring, backup, and operator queries. | | `max_connections` | `[1000..4000, depending on service size]` | Sets the connection limit based on service size. Smaller nodes have lower limits to prevent overload, and larger service plans allow higher limits. | | `max_database_num_to_throw` | `400` | Enforces a hard limit on the number of databases to protect metadata and reduce operational overhead. | | `max_database_num_to_warn` | `200` | Warns before the hard database-count limit is reached. | | `max_format_parsing_thread_pool_size` | `[3..80, depending on CPU count]` | Scales the format parsing thread pool with service CPU capacity. | | `max_named_collection_num_to_throw` | `5000` | Limits user-managed named collections to 5,000 to protect Keeper metadata, configuration reloads, and backup operations. | | `max_partition_size_to_drop` | `0` | Removes the server-side size limit on partition drops so you can drop large partitions. | | `max_prefixes_deserialization_thread_pool_size` | `[3..80, depending on CPU count]` | Scales the prefixes deserialization thread pool with service CPU capacity. | | `max_server_memory_usage` | `[65%..70% of host RAM, depending on service size]` | Caps ClickHouse memory use below total host RAM so the node retains memory for the operating system, caching, and service management operations. | | `max_table_size_to_drop` | `0` | Removes the server-side size limit on table drops so you can drop large tables. | | `memory_worker_correct_memory_tracker` | `1` | Lets the ClickHouse memory worker correct memory tracking drift. | | `mysql_require_secure_transport` | `1` | Rejects plaintext MySQL protocol connections to protect credentials and query traffic. | | `parquet_metadata_cache_size` | `[~7% of RAM]` | Scales the Parquet metadata cache with service memory to prevent disproportionate use on smaller plans. | | `postgresql_require_secure_transport` | `1` | Rejects plaintext PostgreSQL protocol connections to protect credentials and query traffic. | | `prefetch_threadpool_pool_size` | `[3..80, depending on CPU count]` | Scales the prefetch thread pool with service CPU capacity. | | `prepare_system_log_tables_on_startup` | `1` | Ensures system log tables are ready immediately after startup for managed observability and diagnostics. | | `reserved_replicated_database_prefixes` | `["aiven", "endpoint_", "service_"]` | Reserves internal database name prefixes for Aiven-managed objects. | | `series_keeper_path` | `/clickhouse/series` | Sets the Keeper path for `generateSerialID` counter nodes. | | `show_addresses_in_stack_traces` | `0` | Avoids exposing raw addresses in stack traces returned to users or logs. | | `text_index_header_cache_size` | `[~1.75% of RAM, maximum 1 GiB]` | Assigns 25% of a plan-scaled text-index cache budget to headers, protecting smaller services from disproportionate memory use. | | `text_index_postings_cache_size` | `[~3.5% of RAM, maximum 2 GiB]` | Assigns 50% of a plan-scaled text-index cache budget to posting lists, protecting smaller services from disproportionate memory use. | | `text_index_tokens_cache_size` | `[~1.75% of RAM, maximum 1 GiB]` | Assigns 25% of a plan-scaled text-index cache budget to tokens, protecting smaller services from disproportionate memory use. | | `threadpool_local_fs_reader_pool_size` | `[3..80, depending on CPU count]` | Scales local filesystem reader concurrency with service CPU capacity. | | `threadpool_remote_fs_reader_pool_size` | `[3..80, depending on CPU count]` | Scales remote filesystem reader concurrency with service CPU capacity. | | `threadpool_writer_pool_size` | `[3..80, depending on CPU count]` | Scales filesystem writer concurrency with service CPU capacity. | | `uncompressed_cache_size` | `[~7% of RAM]` | Sizes the uncompressed block cache based on the managed memory budget. | | `user_with_indirect_database_creation` | `avnadmin` | Restricts indirect database creation privileges to the managed admin user. | | `vector_similarity_index_cache_size` | `0.07` | Sizes the vector similarity index cache to 7% of server memory. Set to `0` to disable the cache. Maximum value: `0.5`. | --- # Advanced parameters for Aiven for ClickHouse® See the configuration options available for Aiven for ClickHouse®: | Parameter | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | []()[**ip\_filter**](#ip_filter)`array`- default: `0.0.0.0/0,::/0`IP filterAllow incoming connections from CIDR address block, e.g. '10.20.0.0/16' | | []()[**service\_log**](#service_log)`boolean,null`Service loggingStore logs for the service so that they are available in the HTTP API and console. | | []()[**static\_ips**](#static_ips)`boolean`Use static public IP addresses | | []()[**project\_to\_fork\_from**](#project_to_fork_from)`string,null`Name of another project to fork a service from. This has effect only when a new service is being created. | | []()[**private\_access**](#private_access)`object`Allow access to selected service ports from private networksprivate\_access.clickhouse boolean Allow clients to connect to clickhouse with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.clickhouse\_arrowflight boolean Allow clients to connect to clickhouse\_arrowflight with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.clickhouse\_https boolean Allow clients to connect to clickhouse\_https with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.clickhouse\_mysql boolean Allow clients to connect to clickhouse\_mysql with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.prometheus boolean Allow clients to connect to prometheus with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | []()[**privatelink\_access**](#privatelink_access)`object`Allow access to selected service components through Privatelinkprivatelink\_access.clickhouse boolean Enable clickhouse privatelink\_access.clickhouse\_arrowflight boolean Enable clickhouse\_arrowflight privatelink\_access.clickhouse\_https boolean Enable clickhouse\_https privatelink\_access.clickhouse\_mysql boolean Enable clickhouse\_mysql privatelink\_access.prometheus boolean Enable prometheus | | []()[**public\_access**](#public_access)`object`Allow access to selected service ports from the public Internetpublic\_access.clickhouse boolean Allow clients to connect to clickhouse from the public internet for service nodes that are in a project VPC or another type of private network public\_access.clickhouse\_arrowflight boolean Allow clients to connect to clickhouse\_arrowflight from the public internet for service nodes that are in a project VPC or another type of private network public\_access.clickhouse\_https boolean Allow clients to connect to clickhouse\_https from the public internet for service nodes that are in a project VPC or another type of private network public\_access.clickhouse\_mysql boolean Allow clients to connect to clickhouse\_mysql from the public internet for service nodes that are in a project VPC or another type of private network public\_access.prometheus boolean Allow clients to connect to prometheus from the public internet for service nodes that are in a project VPC or another type of private network | | []()[**service\_to\_fork\_from**](#service_to_fork_from)`string,null`Name of another service to fork from. This has effect only when a new service is being created. | | []()[**recovery\_basebackup\_name**](#recovery_basebackup_name)`string`Name of the basebackup to restore in forked service | | []()[**backup\_hour**](#backup_hour)`integer,null`- max: `23`The hour of day (in UTC) when backup for the service is started. New backup is only started if previous backup has already completed. | | []()[**backup\_minute**](#backup_minute)`integer,null`- max: `59`The minute of an hour when backup for the service is started. New backup is only started if previous backup has already completed. | | []()[**enable\_ipv6**](#enable_ipv6)`boolean`Enable IPv6Register AAAA DNS records for the service, and allow IPv6 packets to service ports | | []()[**clickhouse\_version**](#clickhouse_version)`string,null`ClickHouse major version | | []()[**tiered\_storage\_move\_factor**](#tiered_storage_move_factor)`number`- max: `1`
- default: `0.2`Tiered storage move factorThe percentage of free disk space required on local storage before data is moved to object storage. A value of 0.2 means data is moved when local storage has less than 20% free space. | | []()[**server\_settings**](#server_settings)`object`ClickHouse server settings, which can be found in the `system.server_settings` table.server\_settings.vector\_similarity\_index\_cache\_size number - max: 0.5 - default: 0.07 vector\_similarity\_index\_cache\_size Fraction of total server memory allocated to the vector similarity index cache. 0 disables the cache. Default is 0.07 (7% of server memory). Only effective on ClickHouse 25.8+. | | []()[**session\_settings**](#session_settings)`object`ClickHouse session settings, which can be found in the `system.settings` table.session\_settings.compatibility string,null When set, ClickHouse applies backward-compatible behavior from the specified version. Automatically set to the previous version on major version upgrade. Set to null to disable compatibility mode once all incompatibilities have been resolved. Takes effect after the next service restart/upgrade. | --- # Aiven for ClickHouse® metrics available via Datadog Learn what metrics are available via Datadog for Aiven for ClickHouse® services. ## Get a metrics list for your service[​](#get-a-metrics-list-for-your-service "Direct link to Get a metrics list for your service") The list of Aiven for ClickHouse metrics available in Datadog corresponds to the list of metrics available for the open-source ClickHouse and can be checked in [Metrics](https://docs.datadoghq.com/integrations/clickhouse/?tab=host#metrics). Related pages * Check how to use Datadog with Aiven services in [Datadog and Aiven](/docs/integrations/datadog.md). * Check how to send metrics to Datadog from Aiven services in [Send metrics to Datadog](/docs/integrations/datadog/datadog-metrics.md). --- # Aiven for ClickHouse® metrics available via Prometheus List of all metrics available via Prometheus for Aiven for ClickHouse® services. You can retrieve the complete list of available metrics for your service by requesting the Prometheus endpoint: ``` curl --cacert ca.pem \ --user ':' \ 'https://:/metrics' ``` Where you substitute the following: * Aiven project certificate for `ca.pem` * Prometheus credentials for `:` * Aiven for ClickHouse hostname for `` * Prometheus port for `` tip You can check how to use Prometheus with Aiven in [Prometheus metrics](/docs/platform/howto/integrations/prometheus-metrics.md). ``` # TYPE clickhouse_asynchronous_metrics_number_of_databases untyped clickhouse_asynchronous_metrics_number_of_databases # TYPE clickhouse_asynchronous_metrics_number_of_tables untyped clickhouse_asynchronous_metrics_number_of_tables # TYPE clickhouse_asynchronous_metrics_total_parts_of_merge_tree_tables untyped clickhouse_asynchronous_metrics_total_parts_of_merge_tree_tables # TYPE clickhouse_asynchronous_metrics_total_rows_of_merge_tree_tables untyped clickhouse_asynchronous_metrics_total_rows_of_merge_tree_tables # TYPE clickhouse_events_inserted_bytes untyped clickhouse_events_inserted_bytes # TYPE clickhouse_events_inserted_rows untyped clickhouse_events_inserted_rows # TYPE clickhouse_events_merge untyped clickhouse_events_merge # TYPE clickhouse_events_merged_rows untyped clickhouse_events_merged_rows # TYPE clickhouse_events_merged_uncompressed_bytes untyped clickhouse_events_merged_uncompressed_bytes # TYPE clickhouse_events_query untyped clickhouse_events_query # TYPE clickhouse_events_read_compressed_bytes untyped clickhouse_events_read_compressed_bytes # TYPE clickhouse_events_select_query untyped clickhouse_events_select_query # TYPE clickhouse_metrics_delayed_inserts untyped clickhouse_metrics_delayed_inserts # TYPE clickhouse_metrics_ephemeral_node untyped clickhouse_metrics_ephemeral_node # TYPE clickhouse_metrics_http_connection untyped clickhouse_metrics_http_connection # TYPE clickhouse_metrics_interserver_connection untyped clickhouse_metrics_interserver_connection # TYPE clickhouse_metrics_merge untyped clickhouse_metrics_merge # TYPE clickhouse_metrics_query untyped clickhouse_metrics_query # TYPE clickhouse_metrics_query_preempted untyped clickhouse_metrics_query_preempted # TYPE clickhouse_metrics_readonly_replica untyped clickhouse_metrics_readonly_replica # TYPE clickhouse_metrics_replicated_checks untyped clickhouse_metrics_replicated_checks # TYPE clickhouse_metrics_replicated_fetch untyped clickhouse_metrics_replicated_fetch # TYPE clickhouse_metrics_rw_lock_active_readers untyped clickhouse_metrics_rw_lock_active_readers # TYPE clickhouse_metrics_rw_lock_active_writers untyped clickhouse_metrics_rw_lock_active_writers # TYPE clickhouse_metrics_rw_lock_waiting_readers untyped clickhouse_metrics_rw_lock_waiting_readers # TYPE clickhouse_metrics_rw_lock_waiting_writers untyped clickhouse_metrics_rw_lock_waiting_writers # TYPE clickhouse_metrics_tcp_connection untyped clickhouse_metrics_tcp_connection # TYPE clickhouse_metrics_zoo_keeper_request untyped clickhouse_metrics_zoo_keeper_request # TYPE clickhouse_metrics_zoo_keeper_session untyped clickhouse_metrics_zoo_keeper_session # TYPE clickhouse_metrics_zoo_keeper_watch untyped clickhouse_metrics_zoo_keeper_watch # TYPE clickhouse_replication_queue_num_attach_part untyped clickhouse_replication_queue_num_attach_part # TYPE clickhouse_replication_queue_num_get_part untyped clickhouse_replication_queue_num_get_part # TYPE clickhouse_replication_queue_num_merge_parts untyped clickhouse_replication_queue_num_merge_parts # TYPE clickhouse_replication_queue_num_merge_parts_ttl_delete untyped clickhouse_replication_queue_num_merge_parts_ttl_delete # TYPE clickhouse_replication_queue_num_merge_parts_ttl_recompress untyped clickhouse_replication_queue_num_merge_parts_ttl_recompress # TYPE clickhouse_replication_queue_num_mutate_part untyped clickhouse_replication_queue_num_mutate_part # TYPE clickhouse_replication_queue_num_total untyped clickhouse_replication_queue_num_total # TYPE clickhouse_replication_queue_num_tries_replicas untyped clickhouse_replication_queue_num_tries_replicas # TYPE clickhouse_replication_queue_too_many_tries_replicas untyped clickhouse_replication_queue_too_many_tries_replicas # TYPE cpu_usage_guest gauge cpu_usage_guest # TYPE cpu_usage_guest_nice gauge cpu_usage_guest_nice # TYPE cpu_usage_idle gauge cpu_usage_idle # TYPE cpu_usage_iowait gauge cpu_usage_iowait # TYPE cpu_usage_irq gauge cpu_usage_irq # TYPE cpu_usage_nice gauge cpu_usage_nice # TYPE cpu_usage_softirq gauge cpu_usage_softirq # TYPE cpu_usage_steal gauge cpu_usage_steal # TYPE cpu_usage_system gauge cpu_usage_system # TYPE cpu_usage_user gauge cpu_usage_user # TYPE disk_free gauge disk_free # TYPE disk_inodes_free gauge disk_inodes_free # TYPE disk_inodes_total gauge disk_inodes_total # TYPE disk_inodes_used gauge disk_inodes_used # TYPE disk_total gauge disk_total # TYPE disk_used gauge disk_used # TYPE disk_used_percent gauge disk_used_percent # TYPE diskio_io_time counter diskio_io_time # TYPE diskio_iops_in_progress counter diskio_iops_in_progress # TYPE diskio_merged_reads counter diskio_merged_reads # TYPE diskio_merged_writes counter diskio_merged_writes # TYPE diskio_read_bytes counter diskio_read_bytes # TYPE diskio_read_time counter diskio_read_time # TYPE diskio_reads counter diskio_reads # TYPE diskio_weighted_io_time counter diskio_weighted_io_time # TYPE diskio_write_bytes counter diskio_write_bytes # TYPE diskio_write_time counter diskio_write_time # TYPE diskio_writes counter diskio_writes # TYPE kernel_boot_time counter kernel_boot_time # TYPE kernel_context_switches counter kernel_context_switches # TYPE kernel_entropy_avail counter kernel_entropy_avail # TYPE kernel_interrupts counter kernel_interrupts # TYPE kernel_processes_forked counter kernel_processes_forked # TYPE mem_active gauge mem_active # TYPE mem_available gauge mem_available # TYPE mem_available_percent gauge mem_available_percent # TYPE mem_buffered gauge mem_buffered # TYPE mem_cached gauge mem_cached # TYPE mem_commit_limit gauge mem_commit_limit # TYPE mem_committed_as gauge mem_committed_as # TYPE mem_dirty gauge mem_dirty # TYPE mem_free gauge mem_free # TYPE mem_high_free gauge mem_high_free # TYPE mem_high_total gauge mem_high_total # TYPE mem_huge_page_size gauge mem_huge_page_size # TYPE mem_huge_pages_free gauge mem_huge_pages_free # TYPE mem_huge_pages_total gauge mem_huge_pages_total # TYPE mem_inactive gauge mem_inactive # TYPE mem_low_free gauge mem_low_free # TYPE mem_low_total gauge mem_low_total # TYPE mem_mapped gauge mem_mapped # TYPE mem_page_tables gauge mem_page_tables # TYPE mem_shared gauge mem_shared # TYPE mem_slab gauge mem_slab # TYPE mem_sreclaimable gauge mem_sreclaimable # TYPE mem_sunreclaim gauge mem_sunreclaim # TYPE mem_swap_cached gauge mem_swap_cached # TYPE mem_swap_free gauge mem_swap_free # TYPE mem_swap_total gauge mem_swap_total # TYPE mem_total gauge mem_total # TYPE mem_used gauge mem_used # TYPE mem_used_percent gauge mem_used_percent # TYPE mem_vmalloc_chunk gauge mem_vmalloc_chunk # TYPE mem_vmalloc_total gauge mem_vmalloc_total # TYPE mem_vmalloc_used gauge mem_vmalloc_used # TYPE mem_write_back gauge mem_write_back # TYPE mem_write_back_tmp gauge mem_write_back_tmp # TYPE net_bytes_recv counter net_bytes_recv # TYPE net_bytes_sent counter net_bytes_sent # TYPE net_drop_in counter net_drop_in # TYPE net_drop_out counter net_drop_out # TYPE net_err_in counter net_err_in # TYPE net_err_out counter net_err_out # TYPE net_icmp_inaddrmaskreps untyped net_icmp_inaddrmaskreps # TYPE net_icmp_inaddrmasks untyped net_icmp_inaddrmasks # TYPE net_icmp_incsumerrors untyped net_icmp_incsumerrors # TYPE net_icmp_indestunreachs untyped net_icmp_indestunreachs # TYPE net_icmp_inechoreps untyped net_icmp_inechoreps # TYPE net_icmp_inechos untyped net_icmp_inechos # TYPE net_icmp_inerrors untyped net_icmp_inerrors # TYPE net_icmp_inmsgs untyped net_icmp_inmsgs # TYPE net_icmp_inparmprobs untyped net_icmp_inparmprobs # TYPE net_icmp_inredirects untyped net_icmp_inredirects # TYPE net_icmp_insrcquenchs untyped net_icmp_insrcquenchs # TYPE net_icmp_intimeexcds untyped net_icmp_intimeexcds # TYPE net_icmp_intimestampreps untyped net_icmp_intimestampreps # TYPE net_icmp_intimestamps untyped net_icmp_intimestamps # TYPE net_icmp_outaddrmaskreps untyped net_icmp_outaddrmaskreps # TYPE net_icmp_outaddrmasks untyped net_icmp_outaddrmasks # TYPE net_icmp_outdestunreachs untyped net_icmp_outdestunreachs # TYPE net_icmp_outechoreps untyped net_icmp_outechoreps # TYPE net_icmp_outechos untyped net_icmp_outechos # TYPE net_icmp_outerrors untyped net_icmp_outerrors # TYPE net_icmp_outmsgs untyped net_icmp_outmsgs # TYPE net_icmp_outparmprobs untyped net_icmp_outparmprobs # TYPE net_icmp_outratelimitglobal untyped net_icmp_outratelimitglobal # TYPE net_icmp_outratelimithost untyped net_icmp_outratelimithost # TYPE net_icmp_outredirects untyped net_icmp_outredirects # TYPE net_icmp_outsrcquenchs untyped net_icmp_outsrcquenchs # TYPE net_icmp_outtimeexcds untyped net_icmp_outtimeexcds # TYPE net_icmp_outtimestampreps untyped net_icmp_outtimestampreps # TYPE net_icmp_outtimestamps untyped net_icmp_outtimestamps # TYPE net_icmpmsg_intype3 untyped net_icmpmsg_intype3 # TYPE net_icmpmsg_intype8 untyped net_icmpmsg_intype8 # TYPE net_icmpmsg_outtype0 untyped net_icmpmsg_outtype0 # TYPE net_icmpmsg_outtype3 untyped net_icmpmsg_outtype3 # TYPE net_ip_defaultttl untyped net_ip_defaultttl # TYPE net_ip_forwarding untyped net_ip_forwarding # TYPE net_ip_forwdatagrams untyped net_ip_forwdatagrams # TYPE net_ip_fragcreates untyped net_ip_fragcreates # TYPE net_ip_fragfails untyped net_ip_fragfails # TYPE net_ip_fragoks untyped net_ip_fragoks # TYPE net_ip_inaddrerrors untyped net_ip_inaddrerrors # TYPE net_ip_indelivers untyped net_ip_indelivers # TYPE net_ip_indiscards untyped net_ip_indiscards # TYPE net_ip_inhdrerrors untyped net_ip_inhdrerrors # TYPE net_ip_inreceives untyped net_ip_inreceives # TYPE net_ip_inunknownprotos untyped net_ip_inunknownprotos # TYPE net_ip_outdiscards untyped net_ip_outdiscards # TYPE net_ip_outnoroutes untyped net_ip_outnoroutes # TYPE net_ip_outrequests untyped net_ip_outrequests # TYPE net_ip_reasmfails untyped net_ip_reasmfails # TYPE net_ip_reasmoks untyped net_ip_reasmoks # TYPE net_ip_reasmreqds untyped net_ip_reasmreqds # TYPE net_ip_reasmtimeout untyped net_ip_reasmtimeout # TYPE net_packets_recv counter net_packets_recv # TYPE net_packets_sent counter net_packets_sent # TYPE net_tcp_activeopens untyped net_tcp_activeopens # TYPE net_tcp_attemptfails untyped net_tcp_attemptfails # TYPE net_tcp_currestab untyped net_tcp_currestab # TYPE net_tcp_estabresets untyped net_tcp_estabresets # TYPE net_tcp_incsumerrors untyped net_tcp_incsumerrors # TYPE net_tcp_inerrs untyped net_tcp_inerrs # TYPE net_tcp_insegs untyped net_tcp_insegs # TYPE net_tcp_maxconn untyped net_tcp_maxconn # TYPE net_tcp_outrsts untyped net_tcp_outrsts # TYPE net_tcp_outsegs untyped net_tcp_outsegs # TYPE net_tcp_passiveopens untyped net_tcp_passiveopens # TYPE net_tcp_retranssegs untyped net_tcp_retranssegs # TYPE net_tcp_rtoalgorithm untyped net_tcp_rtoalgorithm # TYPE net_tcp_rtomax untyped net_tcp_rtomax # TYPE net_tcp_rtomin untyped net_tcp_rtomin # TYPE net_udp_ignoredmulti untyped net_udp_ignoredmulti # TYPE net_udp_incsumerrors untyped net_udp_incsumerrors # TYPE net_udp_indatagrams untyped net_udp_indatagrams # TYPE net_udp_inerrors untyped net_udp_inerrors # TYPE net_udp_memerrors untyped net_udp_memerrors # TYPE net_udp_noports untyped net_udp_noports # TYPE net_udp_outdatagrams untyped net_udp_outdatagrams # TYPE net_udp_rcvbuferrors untyped net_udp_rcvbuferrors # TYPE net_udp_sndbuferrors untyped net_udp_sndbuferrors # TYPE net_udplite_ignoredmulti untyped net_udplite_ignoredmulti # TYPE net_udplite_incsumerrors untyped net_udplite_incsumerrors # TYPE net_udplite_indatagrams untyped net_udplite_indatagrams # TYPE net_udplite_inerrors untyped net_udplite_inerrors # TYPE net_udplite_memerrors untyped net_udplite_memerrors # TYPE net_udplite_noports untyped net_udplite_noports # TYPE net_udplite_outdatagrams untyped net_udplite_outdatagrams # TYPE net_udplite_rcvbuferrors untyped net_udplite_rcvbuferrors # TYPE net_udplite_sndbuferrors untyped net_udplite_sndbuferrors # TYPE netstat_tcp_close untyped netstat_tcp_close # TYPE netstat_tcp_close_wait untyped netstat_tcp_close_wait # TYPE netstat_tcp_closing untyped netstat_tcp_closing # TYPE netstat_tcp_established untyped netstat_tcp_established # TYPE netstat_tcp_fin_wait1 untyped netstat_tcp_fin_wait1 # TYPE netstat_tcp_fin_wait2 untyped netstat_tcp_fin_wait2 # TYPE netstat_tcp_last_ack untyped netstat_tcp_last_ack # TYPE netstat_tcp_listen untyped netstat_tcp_listen # TYPE netstat_tcp_none untyped netstat_tcp_none # TYPE netstat_tcp_syn_recv untyped netstat_tcp_syn_recv # TYPE netstat_tcp_syn_sent untyped netstat_tcp_syn_sent # TYPE netstat_tcp_time_wait untyped netstat_tcp_time_wait # TYPE netstat_udp_socket untyped netstat_udp_socket # TYPE processes_blocked gauge processes_blocked # TYPE processes_dead gauge processes_dead # TYPE processes_idle gauge processes_idle # TYPE processes_paging gauge processes_paging # TYPE processes_running gauge processes_running # TYPE processes_sleeping gauge processes_sleeping # TYPE processes_stopped gauge processes_stopped # TYPE processes_total gauge processes_total # TYPE processes_total_threads gauge processes_total_threads # TYPE processes_unknown gauge processes_unknown # TYPE processes_zombies gauge processes_zombies # TYPE service_connections_accepted untyped service_connections_accepted # TYPE service_connections_dropped untyped service_connections_dropped # TYPE service_connections_limit_avg_per_second untyped service_connections_limit_avg_per_second # TYPE service_connections_limit_burst untyped service_connections_limit_burst # TYPE swap_free gauge swap_free # TYPE swap_in counter swap_in # TYPE swap_out counter swap_out # TYPE swap_total gauge swap_total # TYPE swap_used gauge swap_used # TYPE swap_used_percent gauge swap_used_percent # TYPE system_load1 gauge system_load1 # TYPE system_load15 gauge system_load15 # TYPE system_load5 gauge system_load5 # TYPE system_n_cpus gauge system_n_cpus # TYPE system_n_unique_users gauge system_n_unique_users # TYPE system_n_users gauge system_n_users # TYPE system_uptime counter system_uptime # TYPE zookeeper_add_dead_watcher_stall_time untyped zookeeper_add_dead_watcher_stall_time # TYPE zookeeper_approximate_data_size untyped zookeeper_approximate_data_size # TYPE zookeeper_auth_failed_count untyped zookeeper_auth_failed_count # TYPE zookeeper_bytes_received_count untyped zookeeper_bytes_received_count # TYPE zookeeper_cnt_1_ack_latency untyped zookeeper_cnt_1_ack_latency # TYPE zookeeper_cnt_action_create_service_user_done_write_per_namespace untyped zookeeper_cnt_action_create_service_user_done_write_per_namespace # TYPE zookeeper_cnt_action_grant_federated_queries_access_done_write_per_namespace untyped zookeeper_cnt_action_grant_federated_queries_access_done_write_per_namespace # TYPE zookeeper_cnt_action_grant_federated_queries_access_v2_done_write_per_namespace untyped zookeeper_cnt_action_grant_federated_queries_access_v2_done_write_per_namespace # TYPE zookeeper_cnt_action_restore_from_astacus_done_write_per_namespace untyped zookeeper_cnt_action_restore_from_astacus_done_write_per_namespace # TYPE zookeeper_cnt_action_update_service_users_privileges_done_write_per_namespace untyped zookeeper_cnt_action_update_service_users_privileges_done_write_per_namespace # TYPE zookeeper_cnt_action_update_service_users_privileges_v2_done_write_per_namespace untyped zookeeper_cnt_action_update_service_users_privileges_v2_done_write_per_namespace # TYPE zookeeper_cnt_clickhouse_read_per_namespace untyped zookeeper_cnt_clickhouse_read_per_namespace # TYPE zookeeper_cnt_clickhouse_write_per_namespace untyped zookeeper_cnt_clickhouse_write_per_namespace # TYPE zookeeper_cnt_close_session_prep_time untyped zookeeper_cnt_close_session_prep_time # TYPE zookeeper_cnt_commit_commit_proc_req_queued untyped zookeeper_cnt_commit_commit_proc_req_queued # TYPE zookeeper_cnt_commit_process_time untyped zookeeper_cnt_commit_process_time # TYPE zookeeper_cnt_commit_propagation_latency untyped zookeeper_cnt_commit_propagation_latency # TYPE zookeeper_cnt_concurrent_request_processing_in_commit_processor untyped zookeeper_cnt_concurrent_request_processing_in_commit_processor # TYPE zookeeper_cnt_connection_token_deficit untyped zookeeper_cnt_connection_token_deficit # TYPE zookeeper_cnt_dbinittime untyped zookeeper_cnt_dbinittime # TYPE zookeeper_cnt_dead_watchers_cleaner_latency untyped zookeeper_cnt_dead_watchers_cleaner_latency # TYPE zookeeper_cnt_election_leader_read_per_namespace untyped zookeeper_cnt_election_leader_read_per_namespace # TYPE zookeeper_cnt_election_leader_write_per_namespace untyped zookeeper_cnt_election_leader_write_per_namespace # TYPE zookeeper_cnt_election_time untyped zookeeper_cnt_election_time # TYPE zookeeper_cnt_follower_sync_time untyped zookeeper_cnt_follower_sync_time # TYPE zookeeper_cnt_fsynctime untyped zookeeper_cnt_fsynctime # TYPE zookeeper_cnt_health_write_per_namespace untyped zookeeper_cnt_health_write_per_namespace # TYPE zookeeper_cnt_inflight_diff_count untyped zookeeper_cnt_inflight_diff_count # TYPE zookeeper_cnt_inflight_snap_count untyped zookeeper_cnt_inflight_snap_count # TYPE zookeeper_cnt_jvm_pause_time_ms untyped zookeeper_cnt_jvm_pause_time_ms # TYPE zookeeper_cnt_local_write_committed_time_ms untyped zookeeper_cnt_local_write_committed_time_ms # TYPE zookeeper_cnt_netty_queued_buffer_capacity untyped zookeeper_cnt_netty_queued_buffer_capacity # TYPE zookeeper_cnt_node_changed_watch_count untyped zookeeper_cnt_node_changed_watch_count # TYPE zookeeper_cnt_node_children_watch_count untyped zookeeper_cnt_node_children_watch_count # TYPE zookeeper_cnt_node_created_watch_count untyped zookeeper_cnt_node_created_watch_count # TYPE zookeeper_cnt_node_deleted_watch_count untyped zookeeper_cnt_node_deleted_watch_count # TYPE zookeeper_cnt_node_slots_read_per_namespace untyped zookeeper_cnt_node_slots_read_per_namespace # TYPE zookeeper_cnt_node_slots_write_per_namespace untyped zookeeper_cnt_node_slots_write_per_namespace # TYPE zookeeper_cnt_nodes_read_per_namespace untyped zookeeper_cnt_nodes_read_per_namespace # TYPE zookeeper_cnt_nodes_write_per_namespace untyped zookeeper_cnt_nodes_write_per_namespace # TYPE zookeeper_cnt_om_commit_process_time_ms untyped zookeeper_cnt_om_commit_process_time_ms # TYPE zookeeper_cnt_om_proposal_process_time_ms untyped zookeeper_cnt_om_proposal_process_time_ms # TYPE zookeeper_cnt_pending_session_queue_size untyped zookeeper_cnt_pending_session_queue_size # TYPE zookeeper_cnt_prep_process_time untyped zookeeper_cnt_prep_process_time # TYPE zookeeper_cnt_prep_processor_queue_size untyped zookeeper_cnt_prep_processor_queue_size # TYPE zookeeper_cnt_prep_processor_queue_time_ms untyped zookeeper_cnt_prep_processor_queue_time_ms # TYPE zookeeper_cnt_propagation_latency untyped zookeeper_cnt_propagation_latency # TYPE zookeeper_cnt_proposal_ack_creation_latency untyped zookeeper_cnt_proposal_ack_creation_latency # TYPE zookeeper_cnt_proposal_latency untyped zookeeper_cnt_proposal_latency # TYPE zookeeper_cnt_quorum_ack_latency untyped zookeeper_cnt_quorum_ack_latency # TYPE zookeeper_cnt_read_commit_proc_issued untyped zookeeper_cnt_read_commit_proc_issued # TYPE zookeeper_cnt_read_commit_proc_req_queued untyped zookeeper_cnt_read_commit_proc_req_queued # TYPE zookeeper_cnt_read_commitproc_time_ms untyped zookeeper_cnt_read_commitproc_time_ms # TYPE zookeeper_cnt_read_final_proc_time_ms untyped zookeeper_cnt_read_final_proc_time_ms # TYPE zookeeper_cnt_readlatency untyped zookeeper_cnt_readlatency # TYPE zookeeper_cnt_reads_after_write_in_session_queue untyped zookeeper_cnt_reads_after_write_in_session_queue # TYPE zookeeper_cnt_reads_issued_from_session_queue untyped zookeeper_cnt_reads_issued_from_session_queue # TYPE zookeeper_cnt_requests_in_session_queue untyped zookeeper_cnt_requests_in_session_queue # TYPE zookeeper_cnt_server_write_committed_time_ms untyped zookeeper_cnt_server_write_committed_time_ms # TYPE zookeeper_cnt_session_queues_drained untyped zookeeper_cnt_session_queues_drained # TYPE zookeeper_cnt_snapshottime untyped zookeeper_cnt_snapshottime # TYPE zookeeper_cnt_startup_snap_load_time untyped zookeeper_cnt_startup_snap_load_time # TYPE zookeeper_cnt_startup_txns_load_time untyped zookeeper_cnt_startup_txns_load_time # TYPE zookeeper_cnt_startup_txns_loaded untyped zookeeper_cnt_startup_txns_loaded # TYPE zookeeper_cnt_sync_process_time untyped zookeeper_cnt_sync_process_time # TYPE zookeeper_cnt_sync_processor_batch_size untyped zookeeper_cnt_sync_processor_batch_size # TYPE zookeeper_cnt_sync_processor_queue_and_flush_time_ms untyped zookeeper_cnt_sync_processor_queue_and_flush_time_ms # TYPE zookeeper_cnt_sync_processor_queue_flush_time_ms untyped zookeeper_cnt_sync_processor_queue_flush_time_ms # TYPE zookeeper_cnt_sync_processor_queue_size untyped zookeeper_cnt_sync_processor_queue_size # TYPE zookeeper_cnt_sync_processor_queue_time_ms untyped zookeeper_cnt_sync_processor_queue_time_ms # TYPE zookeeper_cnt_time_waiting_empty_pool_in_commit_processor_read_ms untyped zookeeper_cnt_time_waiting_empty_pool_in_commit_processor_read_ms # TYPE zookeeper_cnt_updatelatency untyped zookeeper_cnt_updatelatency # TYPE zookeeper_cnt_write_batch_time_in_commit_processor untyped zookeeper_cnt_write_batch_time_in_commit_processor # TYPE zookeeper_cnt_write_commit_proc_issued untyped zookeeper_cnt_write_commit_proc_issued # TYPE zookeeper_cnt_write_commit_proc_req_queued untyped zookeeper_cnt_write_commit_proc_req_queued # TYPE zookeeper_cnt_write_commitproc_time_ms untyped zookeeper_cnt_write_commitproc_time_ms # TYPE zookeeper_cnt_write_final_proc_time_ms untyped zookeeper_cnt_write_final_proc_time_ms # TYPE zookeeper_cnt_zk_cluster_management_write_per_namespace untyped zookeeper_cnt_zk_cluster_management_write_per_namespace # TYPE zookeeper_cnt_zookeeper_read_per_namespace untyped zookeeper_cnt_zookeeper_read_per_namespace # TYPE zookeeper_cnt_zookeeper_write_per_namespace untyped zookeeper_cnt_zookeeper_write_per_namespace # TYPE zookeeper_commit_count untyped zookeeper_commit_count # TYPE zookeeper_connection_drop_count untyped zookeeper_connection_drop_count # TYPE zookeeper_connection_rejected untyped zookeeper_connection_rejected # TYPE zookeeper_connection_request_count untyped zookeeper_connection_request_count # TYPE zookeeper_connection_revalidate_count untyped zookeeper_connection_revalidate_count # TYPE zookeeper_dead_watchers_cleared untyped zookeeper_dead_watchers_cleared # TYPE zookeeper_dead_watchers_queued untyped zookeeper_dead_watchers_queued # TYPE zookeeper_diff_count untyped zookeeper_diff_count # TYPE zookeeper_digest_mismatches_count untyped zookeeper_digest_mismatches_count # TYPE zookeeper_ensemble_auth_fail untyped zookeeper_ensemble_auth_fail # TYPE zookeeper_ensemble_auth_skip untyped zookeeper_ensemble_auth_skip # TYPE zookeeper_ensemble_auth_success untyped zookeeper_ensemble_auth_success # TYPE zookeeper_ephemerals_count untyped zookeeper_ephemerals_count # TYPE zookeeper_global_sessions untyped zookeeper_global_sessions # TYPE zookeeper_large_requests_rejected untyped zookeeper_large_requests_rejected # TYPE zookeeper_last_client_response_size untyped zookeeper_last_client_response_size # TYPE zookeeper_last_proposal_size untyped zookeeper_last_proposal_size # TYPE zookeeper_leader_uptime untyped zookeeper_leader_uptime # TYPE zookeeper_learner_commit_received_count untyped zookeeper_learner_commit_received_count # TYPE zookeeper_learner_proposal_received_count untyped zookeeper_learner_proposal_received_count # TYPE zookeeper_learners untyped zookeeper_learners # TYPE zookeeper_local_sessions untyped zookeeper_local_sessions # TYPE zookeeper_looking_count untyped zookeeper_looking_count # TYPE zookeeper_max_1_ack_latency untyped zookeeper_max_1_ack_latency # TYPE zookeeper_max_action_create_service_user_done_write_per_namespace untyped zookeeper_max_action_create_service_user_done_write_per_namespace # TYPE zookeeper_max_action_grant_federated_queries_access_done_write_per_namespace untyped zookeeper_max_action_grant_federated_queries_access_done_write_per_namespace # TYPE zookeeper_max_action_grant_federated_queries_access_v2_done_write_per_namespace untyped zookeeper_max_action_grant_federated_queries_access_v2_done_write_per_namespace # TYPE zookeeper_max_action_restore_from_astacus_done_write_per_namespace untyped zookeeper_max_action_restore_from_astacus_done_write_per_namespace # TYPE zookeeper_max_action_update_service_users_privileges_done_write_per_namespace untyped zookeeper_max_action_update_service_users_privileges_done_write_per_namespace # TYPE zookeeper_max_action_update_service_users_privileges_v2_done_write_per_namespace untyped zookeeper_max_action_update_service_users_privileges_v2_done_write_per_namespace # TYPE zookeeper_max_clickhouse_read_per_namespace untyped zookeeper_max_clickhouse_read_per_namespace # TYPE zookeeper_max_clickhouse_write_per_namespace untyped zookeeper_max_clickhouse_write_per_namespace # TYPE zookeeper_max_client_response_size untyped zookeeper_max_client_response_size # TYPE zookeeper_max_close_session_prep_time untyped zookeeper_max_close_session_prep_time # TYPE zookeeper_max_commit_commit_proc_req_queued untyped zookeeper_max_commit_commit_proc_req_queued # TYPE zookeeper_max_commit_process_time untyped zookeeper_max_commit_process_time # TYPE zookeeper_max_commit_propagation_latency untyped zookeeper_max_commit_propagation_latency # TYPE zookeeper_max_concurrent_request_processing_in_commit_processor untyped zookeeper_max_concurrent_request_processing_in_commit_processor # TYPE zookeeper_max_connection_token_deficit untyped zookeeper_max_connection_token_deficit # TYPE zookeeper_max_dbinittime untyped zookeeper_max_dbinittime # TYPE zookeeper_max_dead_watchers_cleaner_latency untyped zookeeper_max_dead_watchers_cleaner_latency # TYPE zookeeper_max_election_leader_read_per_namespace untyped zookeeper_max_election_leader_read_per_namespace # TYPE zookeeper_max_election_leader_write_per_namespace untyped zookeeper_max_election_leader_write_per_namespace # TYPE zookeeper_max_election_time untyped zookeeper_max_election_time # TYPE zookeeper_max_file_descriptor_count untyped zookeeper_max_file_descriptor_count # TYPE zookeeper_max_follower_sync_time untyped zookeeper_max_follower_sync_time # TYPE zookeeper_max_fsynctime untyped zookeeper_max_fsynctime # TYPE zookeeper_max_health_write_per_namespace untyped zookeeper_max_health_write_per_namespace # TYPE zookeeper_max_inflight_diff_count untyped zookeeper_max_inflight_diff_count # TYPE zookeeper_max_inflight_snap_count untyped zookeeper_max_inflight_snap_count # TYPE zookeeper_max_jvm_pause_time_ms untyped zookeeper_max_jvm_pause_time_ms # TYPE zookeeper_max_latency untyped zookeeper_max_latency # TYPE zookeeper_max_local_write_committed_time_ms untyped zookeeper_max_local_write_committed_time_ms # TYPE zookeeper_max_netty_queued_buffer_capacity untyped zookeeper_max_netty_queued_buffer_capacity # TYPE zookeeper_max_node_changed_watch_count untyped zookeeper_max_node_changed_watch_count # TYPE zookeeper_max_node_children_watch_count untyped zookeeper_max_node_children_watch_count # TYPE zookeeper_max_node_created_watch_count untyped zookeeper_max_node_created_watch_count # TYPE zookeeper_max_node_deleted_watch_count untyped zookeeper_max_node_deleted_watch_count # TYPE zookeeper_max_node_slots_read_per_namespace untyped zookeeper_max_node_slots_read_per_namespace # TYPE zookeeper_max_node_slots_write_per_namespace untyped zookeeper_max_node_slots_write_per_namespace # TYPE zookeeper_max_nodes_read_per_namespace untyped zookeeper_max_nodes_read_per_namespace # TYPE zookeeper_max_nodes_write_per_namespace untyped zookeeper_max_nodes_write_per_namespace # TYPE zookeeper_max_om_commit_process_time_ms untyped zookeeper_max_om_commit_process_time_ms # TYPE zookeeper_max_om_proposal_process_time_ms untyped zookeeper_max_om_proposal_process_time_ms # TYPE zookeeper_max_pending_session_queue_size untyped zookeeper_max_pending_session_queue_size # TYPE zookeeper_max_prep_process_time untyped zookeeper_max_prep_process_time # TYPE zookeeper_max_prep_processor_queue_size untyped zookeeper_max_prep_processor_queue_size # TYPE zookeeper_max_prep_processor_queue_time_ms untyped zookeeper_max_prep_processor_queue_time_ms # TYPE zookeeper_max_propagation_latency untyped zookeeper_max_propagation_latency # TYPE zookeeper_max_proposal_ack_creation_latency untyped zookeeper_max_proposal_ack_creation_latency # TYPE zookeeper_max_proposal_latency untyped zookeeper_max_proposal_latency # TYPE zookeeper_max_proposal_size untyped zookeeper_max_proposal_size # TYPE zookeeper_max_quorum_ack_latency untyped zookeeper_max_quorum_ack_latency # TYPE zookeeper_max_read_commit_proc_issued untyped zookeeper_max_read_commit_proc_issued # TYPE zookeeper_max_read_commit_proc_req_queued untyped zookeeper_max_read_commit_proc_req_queued # TYPE zookeeper_max_read_commitproc_time_ms untyped zookeeper_max_read_commitproc_time_ms # TYPE zookeeper_max_read_final_proc_time_ms untyped zookeeper_max_read_final_proc_time_ms # TYPE zookeeper_max_readlatency untyped zookeeper_max_readlatency # TYPE zookeeper_max_reads_after_write_in_session_queue untyped zookeeper_max_reads_after_write_in_session_queue # TYPE zookeeper_max_reads_issued_from_session_queue untyped zookeeper_max_reads_issued_from_session_queue # TYPE zookeeper_max_requests_in_session_queue untyped zookeeper_max_requests_in_session_queue # TYPE zookeeper_max_server_write_committed_time_ms untyped zookeeper_max_server_write_committed_time_ms # TYPE zookeeper_max_session_queues_drained untyped zookeeper_max_session_queues_drained # TYPE zookeeper_max_snapshottime untyped zookeeper_max_snapshottime # TYPE zookeeper_max_startup_snap_load_time untyped zookeeper_max_startup_snap_load_time # TYPE zookeeper_max_startup_txns_load_time untyped zookeeper_max_startup_txns_load_time # TYPE zookeeper_max_startup_txns_loaded untyped zookeeper_max_startup_txns_loaded # TYPE zookeeper_max_sync_process_time untyped zookeeper_max_sync_process_time # TYPE zookeeper_max_sync_processor_batch_size untyped zookeeper_max_sync_processor_batch_size # TYPE zookeeper_max_sync_processor_queue_and_flush_time_ms untyped zookeeper_max_sync_processor_queue_and_flush_time_ms # TYPE zookeeper_max_sync_processor_queue_flush_time_ms untyped zookeeper_max_sync_processor_queue_flush_time_ms # TYPE zookeeper_max_sync_processor_queue_size untyped zookeeper_max_sync_processor_queue_size # TYPE zookeeper_max_sync_processor_queue_time_ms untyped zookeeper_max_sync_processor_queue_time_ms # TYPE zookeeper_max_time_waiting_empty_pool_in_commit_processor_read_ms untyped zookeeper_max_time_waiting_empty_pool_in_commit_processor_read_ms # TYPE zookeeper_max_updatelatency untyped zookeeper_max_updatelatency # TYPE zookeeper_max_write_batch_time_in_commit_processor untyped zookeeper_max_write_batch_time_in_commit_processor # TYPE zookeeper_max_write_commit_proc_issued untyped zookeeper_max_write_commit_proc_issued # TYPE zookeeper_max_write_commit_proc_req_queued untyped zookeeper_max_write_commit_proc_req_queued # TYPE zookeeper_max_write_commitproc_time_ms untyped zookeeper_max_write_commitproc_time_ms # TYPE zookeeper_max_write_final_proc_time_ms untyped zookeeper_max_write_final_proc_time_ms # TYPE zookeeper_max_zk_cluster_management_write_per_namespace untyped zookeeper_max_zk_cluster_management_write_per_namespace # TYPE zookeeper_max_zookeeper_read_per_namespace untyped zookeeper_max_zookeeper_read_per_namespace # TYPE zookeeper_max_zookeeper_write_per_namespace untyped zookeeper_max_zookeeper_write_per_namespace # TYPE zookeeper_min_1_ack_latency untyped zookeeper_min_1_ack_latency # TYPE zookeeper_min_action_create_service_user_done_write_per_namespace untyped zookeeper_min_action_create_service_user_done_write_per_namespace # TYPE zookeeper_min_action_grant_federated_queries_access_done_write_per_namespace untyped zookeeper_min_action_grant_federated_queries_access_done_write_per_namespace # TYPE zookeeper_min_action_grant_federated_queries_access_v2_done_write_per_namespace untyped zookeeper_min_action_grant_federated_queries_access_v2_done_write_per_namespace # TYPE zookeeper_min_action_restore_from_astacus_done_write_per_namespace untyped zookeeper_min_action_restore_from_astacus_done_write_per_namespace # TYPE zookeeper_min_action_update_service_users_privileges_done_write_per_namespace untyped zookeeper_min_action_update_service_users_privileges_done_write_per_namespace # TYPE zookeeper_min_action_update_service_users_privileges_v2_done_write_per_namespace untyped zookeeper_min_action_update_service_users_privileges_v2_done_write_per_namespace # TYPE zookeeper_min_clickhouse_read_per_namespace untyped zookeeper_min_clickhouse_read_per_namespace # TYPE zookeeper_min_clickhouse_write_per_namespace untyped zookeeper_min_clickhouse_write_per_namespace # TYPE zookeeper_min_client_response_size untyped zookeeper_min_client_response_size # TYPE zookeeper_min_close_session_prep_time untyped zookeeper_min_close_session_prep_time # TYPE zookeeper_min_commit_commit_proc_req_queued untyped zookeeper_min_commit_commit_proc_req_queued # TYPE zookeeper_min_commit_process_time untyped zookeeper_min_commit_process_time # TYPE zookeeper_min_commit_propagation_latency untyped zookeeper_min_commit_propagation_latency # TYPE zookeeper_min_concurrent_request_processing_in_commit_processor untyped zookeeper_min_concurrent_request_processing_in_commit_processor # TYPE zookeeper_min_connection_token_deficit untyped zookeeper_min_connection_token_deficit # TYPE zookeeper_min_dbinittime untyped zookeeper_min_dbinittime # TYPE zookeeper_min_dead_watchers_cleaner_latency untyped zookeeper_min_dead_watchers_cleaner_latency # TYPE zookeeper_min_election_leader_read_per_namespace untyped zookeeper_min_election_leader_read_per_namespace # TYPE zookeeper_min_election_leader_write_per_namespace untyped zookeeper_min_election_leader_write_per_namespace # TYPE zookeeper_min_election_time untyped zookeeper_min_election_time # TYPE zookeeper_min_follower_sync_time untyped zookeeper_min_follower_sync_time # TYPE zookeeper_min_fsynctime untyped zookeeper_min_fsynctime # TYPE zookeeper_min_health_write_per_namespace untyped zookeeper_min_health_write_per_namespace # TYPE zookeeper_min_inflight_diff_count untyped zookeeper_min_inflight_diff_count # TYPE zookeeper_min_inflight_snap_count untyped zookeeper_min_inflight_snap_count # TYPE zookeeper_min_jvm_pause_time_ms untyped zookeeper_min_jvm_pause_time_ms # TYPE zookeeper_min_latency untyped zookeeper_min_latency # TYPE zookeeper_min_local_write_committed_time_ms untyped zookeeper_min_local_write_committed_time_ms # TYPE zookeeper_min_netty_queued_buffer_capacity untyped zookeeper_min_netty_queued_buffer_capacity # TYPE zookeeper_min_node_changed_watch_count untyped zookeeper_min_node_changed_watch_count # TYPE zookeeper_min_node_children_watch_count untyped zookeeper_min_node_children_watch_count # TYPE zookeeper_min_node_created_watch_count untyped zookeeper_min_node_created_watch_count # TYPE zookeeper_min_node_deleted_watch_count untyped zookeeper_min_node_deleted_watch_count # TYPE zookeeper_min_node_slots_read_per_namespace untyped zookeeper_min_node_slots_read_per_namespace # TYPE zookeeper_min_node_slots_write_per_namespace untyped zookeeper_min_node_slots_write_per_namespace # TYPE zookeeper_min_nodes_read_per_namespace untyped zookeeper_min_nodes_read_per_namespace # TYPE zookeeper_min_nodes_write_per_namespace untyped zookeeper_min_nodes_write_per_namespace # TYPE zookeeper_min_om_commit_process_time_ms untyped zookeeper_min_om_commit_process_time_ms # TYPE zookeeper_min_om_proposal_process_time_ms untyped zookeeper_min_om_proposal_process_time_ms # TYPE zookeeper_min_pending_session_queue_size untyped zookeeper_min_pending_session_queue_size # TYPE zookeeper_min_prep_process_time untyped zookeeper_min_prep_process_time # TYPE zookeeper_min_prep_processor_queue_size untyped zookeeper_min_prep_processor_queue_size # TYPE zookeeper_min_prep_processor_queue_time_ms untyped zookeeper_min_prep_processor_queue_time_ms # TYPE zookeeper_min_propagation_latency untyped zookeeper_min_propagation_latency # TYPE zookeeper_min_proposal_ack_creation_latency untyped zookeeper_min_proposal_ack_creation_latency # TYPE zookeeper_min_proposal_latency untyped zookeeper_min_proposal_latency # TYPE zookeeper_min_proposal_size untyped zookeeper_min_proposal_size # TYPE zookeeper_min_quorum_ack_latency untyped zookeeper_min_quorum_ack_latency # TYPE zookeeper_min_read_commit_proc_issued untyped zookeeper_min_read_commit_proc_issued # TYPE zookeeper_min_read_commit_proc_req_queued untyped zookeeper_min_read_commit_proc_req_queued # TYPE zookeeper_min_read_commitproc_time_ms untyped zookeeper_min_read_commitproc_time_ms # TYPE zookeeper_min_read_final_proc_time_ms untyped zookeeper_min_read_final_proc_time_ms # TYPE zookeeper_min_readlatency untyped zookeeper_min_readlatency # TYPE zookeeper_min_reads_after_write_in_session_queue untyped zookeeper_min_reads_after_write_in_session_queue # TYPE zookeeper_min_reads_issued_from_session_queue untyped zookeeper_min_reads_issued_from_session_queue # TYPE zookeeper_min_requests_in_session_queue untyped zookeeper_min_requests_in_session_queue # TYPE zookeeper_min_server_write_committed_time_ms untyped zookeeper_min_server_write_committed_time_ms # TYPE zookeeper_min_session_queues_drained untyped zookeeper_min_session_queues_drained # TYPE zookeeper_min_snapshottime untyped zookeeper_min_snapshottime # TYPE zookeeper_min_startup_snap_load_time untyped zookeeper_min_startup_snap_load_time # TYPE zookeeper_min_startup_txns_load_time untyped zookeeper_min_startup_txns_load_time # TYPE zookeeper_min_startup_txns_loaded untyped zookeeper_min_startup_txns_loaded # TYPE zookeeper_min_sync_process_time untyped zookeeper_min_sync_process_time # TYPE zookeeper_min_sync_processor_batch_size untyped zookeeper_min_sync_processor_batch_size # TYPE zookeeper_min_sync_processor_queue_and_flush_time_ms untyped zookeeper_min_sync_processor_queue_and_flush_time_ms # TYPE zookeeper_min_sync_processor_queue_flush_time_ms untyped zookeeper_min_sync_processor_queue_flush_time_ms # TYPE zookeeper_min_sync_processor_queue_size untyped zookeeper_min_sync_processor_queue_size # TYPE zookeeper_min_sync_processor_queue_time_ms untyped zookeeper_min_sync_processor_queue_time_ms # TYPE zookeeper_min_time_waiting_empty_pool_in_commit_processor_read_ms untyped zookeeper_min_time_waiting_empty_pool_in_commit_processor_read_ms # TYPE zookeeper_min_updatelatency untyped zookeeper_min_updatelatency # TYPE zookeeper_min_write_batch_time_in_commit_processor untyped zookeeper_min_write_batch_time_in_commit_processor # TYPE zookeeper_min_write_commit_proc_issued untyped zookeeper_min_write_commit_proc_issued # TYPE zookeeper_min_write_commit_proc_req_queued untyped zookeeper_min_write_commit_proc_req_queued # TYPE zookeeper_min_write_commitproc_time_ms untyped zookeeper_min_write_commitproc_time_ms # TYPE zookeeper_min_write_final_proc_time_ms untyped zookeeper_min_write_final_proc_time_ms # TYPE zookeeper_min_zk_cluster_management_write_per_namespace untyped zookeeper_min_zk_cluster_management_write_per_namespace # TYPE zookeeper_min_zookeeper_read_per_namespace untyped zookeeper_min_zookeeper_read_per_namespace # TYPE zookeeper_min_zookeeper_write_per_namespace untyped zookeeper_min_zookeeper_write_per_namespace # TYPE zookeeper_non_mtls_local_conn_count untyped zookeeper_non_mtls_local_conn_count # TYPE zookeeper_non_mtls_remote_conn_count untyped zookeeper_non_mtls_remote_conn_count # TYPE zookeeper_num_alive_connections untyped zookeeper_num_alive_connections # TYPE zookeeper_open_file_descriptor_count untyped zookeeper_open_file_descriptor_count # TYPE zookeeper_outstanding_changes_queued untyped zookeeper_outstanding_changes_queued # TYPE zookeeper_outstanding_changes_removed untyped zookeeper_outstanding_changes_removed # TYPE zookeeper_outstanding_requests untyped zookeeper_outstanding_requests # TYPE zookeeper_outstanding_tls_handshake untyped zookeeper_outstanding_tls_handshake # TYPE zookeeper_p50_1_ack_latency untyped zookeeper_p50_1_ack_latency # TYPE zookeeper_p50_close_session_prep_time untyped zookeeper_p50_close_session_prep_time # TYPE zookeeper_p50_commit_propagation_latency untyped zookeeper_p50_commit_propagation_latency # TYPE zookeeper_p50_dead_watchers_cleaner_latency untyped zookeeper_p50_dead_watchers_cleaner_latency # TYPE zookeeper_p50_jvm_pause_time_ms untyped zookeeper_p50_jvm_pause_time_ms # TYPE zookeeper_p50_local_write_committed_time_ms untyped zookeeper_p50_local_write_committed_time_ms # TYPE zookeeper_p50_om_commit_process_time_ms untyped zookeeper_p50_om_commit_process_time_ms # TYPE zookeeper_p50_om_proposal_process_time_ms untyped zookeeper_p50_om_proposal_process_time_ms # TYPE zookeeper_p50_prep_processor_queue_time_ms untyped zookeeper_p50_prep_processor_queue_time_ms # TYPE zookeeper_p50_propagation_latency untyped zookeeper_p50_propagation_latency # TYPE zookeeper_p50_proposal_ack_creation_latency untyped zookeeper_p50_proposal_ack_creation_latency # TYPE zookeeper_p50_proposal_latency untyped zookeeper_p50_proposal_latency # TYPE zookeeper_p50_quorum_ack_latency untyped zookeeper_p50_quorum_ack_latency # TYPE zookeeper_p50_read_commitproc_time_ms untyped zookeeper_p50_read_commitproc_time_ms # TYPE zookeeper_p50_read_final_proc_time_ms untyped zookeeper_p50_read_final_proc_time_ms # TYPE zookeeper_p50_readlatency untyped zookeeper_p50_readlatency # TYPE zookeeper_p50_server_write_committed_time_ms untyped zookeeper_p50_server_write_committed_time_ms # TYPE zookeeper_p50_sync_processor_queue_and_flush_time_ms untyped zookeeper_p50_sync_processor_queue_and_flush_time_ms # TYPE zookeeper_p50_sync_processor_queue_flush_time_ms untyped zookeeper_p50_sync_processor_queue_flush_time_ms # TYPE zookeeper_p50_sync_processor_queue_time_ms untyped zookeeper_p50_sync_processor_queue_time_ms # TYPE zookeeper_p50_updatelatency untyped zookeeper_p50_updatelatency # TYPE zookeeper_p50_write_commitproc_time_ms untyped zookeeper_p50_write_commitproc_time_ms # TYPE zookeeper_p50_write_final_proc_time_ms untyped zookeeper_p50_write_final_proc_time_ms # TYPE zookeeper_p95_1_ack_latency untyped zookeeper_p95_1_ack_latency # TYPE zookeeper_p95_close_session_prep_time untyped zookeeper_p95_close_session_prep_time # TYPE zookeeper_p95_commit_propagation_latency untyped zookeeper_p95_commit_propagation_latency # TYPE zookeeper_p95_dead_watchers_cleaner_latency untyped zookeeper_p95_dead_watchers_cleaner_latency # TYPE zookeeper_p95_jvm_pause_time_ms untyped zookeeper_p95_jvm_pause_time_ms # TYPE zookeeper_p95_local_write_committed_time_ms untyped zookeeper_p95_local_write_committed_time_ms # TYPE zookeeper_p95_om_commit_process_time_ms untyped zookeeper_p95_om_commit_process_time_ms # TYPE zookeeper_p95_om_proposal_process_time_ms untyped zookeeper_p95_om_proposal_process_time_ms # TYPE zookeeper_p95_prep_processor_queue_time_ms untyped zookeeper_p95_prep_processor_queue_time_ms # TYPE zookeeper_p95_propagation_latency untyped zookeeper_p95_propagation_latency # TYPE zookeeper_p95_proposal_ack_creation_latency untyped zookeeper_p95_proposal_ack_creation_latency # TYPE zookeeper_p95_proposal_latency untyped zookeeper_p95_proposal_latency # TYPE zookeeper_p95_quorum_ack_latency untyped zookeeper_p95_quorum_ack_latency # TYPE zookeeper_p95_read_commitproc_time_ms untyped zookeeper_p95_read_commitproc_time_ms # TYPE zookeeper_p95_read_final_proc_time_ms untyped zookeeper_p95_read_final_proc_time_ms # TYPE zookeeper_p95_readlatency untyped zookeeper_p95_readlatency # TYPE zookeeper_p95_server_write_committed_time_ms untyped zookeeper_p95_server_write_committed_time_ms # TYPE zookeeper_p95_sync_processor_queue_and_flush_time_ms untyped zookeeper_p95_sync_processor_queue_and_flush_time_ms # TYPE zookeeper_p95_sync_processor_queue_flush_time_ms untyped zookeeper_p95_sync_processor_queue_flush_time_ms # TYPE zookeeper_p95_sync_processor_queue_time_ms untyped zookeeper_p95_sync_processor_queue_time_ms # TYPE zookeeper_p95_updatelatency untyped zookeeper_p95_updatelatency # TYPE zookeeper_p95_write_commitproc_time_ms untyped zookeeper_p95_write_commitproc_time_ms # TYPE zookeeper_p95_write_final_proc_time_ms untyped zookeeper_p95_write_final_proc_time_ms # TYPE zookeeper_p999_1_ack_latency untyped zookeeper_p999_1_ack_latency # TYPE zookeeper_p999_close_session_prep_time untyped zookeeper_p999_close_session_prep_time # TYPE zookeeper_p999_commit_propagation_latency untyped zookeeper_p999_commit_propagation_latency # TYPE zookeeper_p999_dead_watchers_cleaner_latency untyped zookeeper_p999_dead_watchers_cleaner_latency # TYPE zookeeper_p999_jvm_pause_time_ms untyped zookeeper_p999_jvm_pause_time_ms # TYPE zookeeper_p999_local_write_committed_time_ms untyped zookeeper_p999_local_write_committed_time_ms # TYPE zookeeper_p999_om_commit_process_time_ms untyped zookeeper_p999_om_commit_process_time_ms # TYPE zookeeper_p999_om_proposal_process_time_ms untyped zookeeper_p999_om_proposal_process_time_ms # TYPE zookeeper_p999_prep_processor_queue_time_ms untyped zookeeper_p999_prep_processor_queue_time_ms # TYPE zookeeper_p999_propagation_latency untyped zookeeper_p999_propagation_latency # TYPE zookeeper_p999_proposal_ack_creation_latency untyped zookeeper_p999_proposal_ack_creation_latency # TYPE zookeeper_p999_proposal_latency untyped zookeeper_p999_proposal_latency # TYPE zookeeper_p999_quorum_ack_latency untyped zookeeper_p999_quorum_ack_latency # TYPE zookeeper_p999_read_commitproc_time_ms untyped zookeeper_p999_read_commitproc_time_ms # TYPE zookeeper_p999_read_final_proc_time_ms untyped zookeeper_p999_read_final_proc_time_ms # TYPE zookeeper_p999_readlatency untyped zookeeper_p999_readlatency # TYPE zookeeper_p999_server_write_committed_time_ms untyped zookeeper_p999_server_write_committed_time_ms # TYPE zookeeper_p999_sync_processor_queue_and_flush_time_ms untyped zookeeper_p999_sync_processor_queue_and_flush_time_ms # TYPE zookeeper_p999_sync_processor_queue_flush_time_ms untyped zookeeper_p999_sync_processor_queue_flush_time_ms # TYPE zookeeper_p999_sync_processor_queue_time_ms untyped zookeeper_p999_sync_processor_queue_time_ms # TYPE zookeeper_p999_updatelatency untyped zookeeper_p999_updatelatency # TYPE zookeeper_p999_write_commitproc_time_ms untyped zookeeper_p999_write_commitproc_time_ms # TYPE zookeeper_p999_write_final_proc_time_ms untyped zookeeper_p999_write_final_proc_time_ms # TYPE zookeeper_p99_1_ack_latency untyped zookeeper_p99_1_ack_latency # TYPE zookeeper_p99_close_session_prep_time untyped zookeeper_p99_close_session_prep_time # TYPE zookeeper_p99_commit_propagation_latency untyped zookeeper_p99_commit_propagation_latency # TYPE zookeeper_p99_dead_watchers_cleaner_latency untyped zookeeper_p99_dead_watchers_cleaner_latency # TYPE zookeeper_p99_jvm_pause_time_ms untyped zookeeper_p99_jvm_pause_time_ms # TYPE zookeeper_p99_local_write_committed_time_ms untyped zookeeper_p99_local_write_committed_time_ms # TYPE zookeeper_p99_om_commit_process_time_ms untyped zookeeper_p99_om_commit_process_time_ms # TYPE zookeeper_p99_om_proposal_process_time_ms untyped zookeeper_p99_om_proposal_process_time_ms # TYPE zookeeper_p99_prep_processor_queue_time_ms untyped zookeeper_p99_prep_processor_queue_time_ms # TYPE zookeeper_p99_propagation_latency untyped zookeeper_p99_propagation_latency # TYPE zookeeper_p99_proposal_ack_creation_latency untyped zookeeper_p99_proposal_ack_creation_latency # TYPE zookeeper_p99_proposal_latency untyped zookeeper_p99_proposal_latency # TYPE zookeeper_p99_quorum_ack_latency untyped zookeeper_p99_quorum_ack_latency # TYPE zookeeper_p99_read_commitproc_time_ms untyped zookeeper_p99_read_commitproc_time_ms # TYPE zookeeper_p99_read_final_proc_time_ms untyped zookeeper_p99_read_final_proc_time_ms # TYPE zookeeper_p99_readlatency untyped zookeeper_p99_readlatency # TYPE zookeeper_p99_server_write_committed_time_ms untyped zookeeper_p99_server_write_committed_time_ms # TYPE zookeeper_p99_sync_processor_queue_and_flush_time_ms untyped zookeeper_p99_sync_processor_queue_and_flush_time_ms # TYPE zookeeper_p99_sync_processor_queue_flush_time_ms untyped zookeeper_p99_sync_processor_queue_flush_time_ms # TYPE zookeeper_p99_sync_processor_queue_time_ms untyped zookeeper_p99_sync_processor_queue_time_ms # TYPE zookeeper_p99_updatelatency untyped zookeeper_p99_updatelatency # TYPE zookeeper_p99_write_commitproc_time_ms untyped zookeeper_p99_write_commitproc_time_ms # TYPE zookeeper_p99_write_final_proc_time_ms untyped zookeeper_p99_write_final_proc_time_ms # TYPE zookeeper_packets_received untyped zookeeper_packets_received # TYPE zookeeper_packets_sent untyped zookeeper_packets_sent # TYPE zookeeper_pending_syncs untyped zookeeper_pending_syncs # TYPE zookeeper_prep_processor_request_queued untyped zookeeper_prep_processor_request_queued # TYPE zookeeper_proposal_count untyped zookeeper_proposal_count # TYPE zookeeper_quit_leading_due_to_disloyal_voter untyped zookeeper_quit_leading_due_to_disloyal_voter # TYPE zookeeper_quorum_size untyped zookeeper_quorum_size # TYPE zookeeper_request_commit_queued untyped zookeeper_request_commit_queued # TYPE zookeeper_request_throttle_wait_count untyped zookeeper_request_throttle_wait_count # TYPE zookeeper_response_packet_cache_hits untyped zookeeper_response_packet_cache_hits # TYPE zookeeper_response_packet_cache_misses untyped zookeeper_response_packet_cache_misses # TYPE zookeeper_response_packet_get_children_cache_hits untyped zookeeper_response_packet_get_children_cache_hits # TYPE zookeeper_response_packet_get_children_cache_misses untyped zookeeper_response_packet_get_children_cache_misses # TYPE zookeeper_revalidate_count untyped zookeeper_revalidate_count # TYPE zookeeper_sessionless_connections_expired untyped zookeeper_sessionless_connections_expired # TYPE zookeeper_snap_count untyped zookeeper_snap_count # TYPE zookeeper_stale_replies untyped zookeeper_stale_replies # TYPE zookeeper_stale_requests untyped zookeeper_stale_requests # TYPE zookeeper_stale_requests_dropped untyped zookeeper_stale_requests_dropped # TYPE zookeeper_stale_sessions_expired untyped zookeeper_stale_sessions_expired # TYPE zookeeper_sum_1_ack_latency untyped zookeeper_sum_1_ack_latency # TYPE zookeeper_sum_action_create_service_user_done_write_per_namespace untyped zookeeper_sum_action_create_service_user_done_write_per_namespace # TYPE zookeeper_sum_action_grant_federated_queries_access_done_write_per_namespace untyped zookeeper_sum_action_grant_federated_queries_access_done_write_per_namespace # TYPE zookeeper_sum_action_grant_federated_queries_access_v2_done_write_per_namespace untyped zookeeper_sum_action_grant_federated_queries_access_v2_done_write_per_namespace # TYPE zookeeper_sum_action_restore_from_astacus_done_write_per_namespace untyped zookeeper_sum_action_restore_from_astacus_done_write_per_namespace # TYPE zookeeper_sum_action_update_service_users_privileges_done_write_per_namespace untyped zookeeper_sum_action_update_service_users_privileges_done_write_per_namespace # TYPE zookeeper_sum_action_update_service_users_privileges_v2_done_write_per_namespace untyped zookeeper_sum_action_update_service_users_privileges_v2_done_write_per_namespace # TYPE zookeeper_sum_clickhouse_read_per_namespace untyped zookeeper_sum_clickhouse_read_per_namespace # TYPE zookeeper_sum_clickhouse_write_per_namespace untyped zookeeper_sum_clickhouse_write_per_namespace # TYPE zookeeper_sum_close_session_prep_time untyped zookeeper_sum_close_session_prep_time # TYPE zookeeper_sum_commit_commit_proc_req_queued untyped zookeeper_sum_commit_commit_proc_req_queued # TYPE zookeeper_sum_commit_process_time untyped zookeeper_sum_commit_process_time # TYPE zookeeper_sum_commit_propagation_latency untyped zookeeper_sum_commit_propagation_latency # TYPE zookeeper_sum_concurrent_request_processing_in_commit_processor untyped zookeeper_sum_concurrent_request_processing_in_commit_processor # TYPE zookeeper_sum_connection_token_deficit untyped zookeeper_sum_connection_token_deficit # TYPE zookeeper_sum_dbinittime untyped zookeeper_sum_dbinittime # TYPE zookeeper_sum_dead_watchers_cleaner_latency untyped zookeeper_sum_dead_watchers_cleaner_latency # TYPE zookeeper_sum_election_leader_read_per_namespace untyped zookeeper_sum_election_leader_read_per_namespace # TYPE zookeeper_sum_election_leader_write_per_namespace untyped zookeeper_sum_election_leader_write_per_namespace # TYPE zookeeper_sum_election_time untyped zookeeper_sum_election_time # TYPE zookeeper_sum_follower_sync_time untyped zookeeper_sum_follower_sync_time # TYPE zookeeper_sum_fsynctime untyped zookeeper_sum_fsynctime # TYPE zookeeper_sum_health_write_per_namespace untyped zookeeper_sum_health_write_per_namespace # TYPE zookeeper_sum_inflight_diff_count untyped zookeeper_sum_inflight_diff_count # TYPE zookeeper_sum_inflight_snap_count untyped zookeeper_sum_inflight_snap_count # TYPE zookeeper_sum_jvm_pause_time_ms untyped zookeeper_sum_jvm_pause_time_ms # TYPE zookeeper_sum_local_write_committed_time_ms untyped zookeeper_sum_local_write_committed_time_ms # TYPE zookeeper_sum_netty_queued_buffer_capacity untyped zookeeper_sum_netty_queued_buffer_capacity # TYPE zookeeper_sum_node_changed_watch_count untyped zookeeper_sum_node_changed_watch_count # TYPE zookeeper_sum_node_children_watch_count untyped zookeeper_sum_node_children_watch_count # TYPE zookeeper_sum_node_created_watch_count untyped zookeeper_sum_node_created_watch_count # TYPE zookeeper_sum_node_deleted_watch_count untyped zookeeper_sum_node_deleted_watch_count # TYPE zookeeper_sum_node_slots_read_per_namespace untyped zookeeper_sum_node_slots_read_per_namespace # TYPE zookeeper_sum_node_slots_write_per_namespace untyped zookeeper_sum_node_slots_write_per_namespace # TYPE zookeeper_sum_nodes_read_per_namespace untyped zookeeper_sum_nodes_read_per_namespace # TYPE zookeeper_sum_nodes_write_per_namespace untyped zookeeper_sum_nodes_write_per_namespace # TYPE zookeeper_sum_om_commit_process_time_ms untyped zookeeper_sum_om_commit_process_time_ms # TYPE zookeeper_sum_om_proposal_process_time_ms untyped zookeeper_sum_om_proposal_process_time_ms # TYPE zookeeper_sum_pending_session_queue_size untyped zookeeper_sum_pending_session_queue_size # TYPE zookeeper_sum_prep_process_time untyped zookeeper_sum_prep_process_time # TYPE zookeeper_sum_prep_processor_queue_size untyped zookeeper_sum_prep_processor_queue_size # TYPE zookeeper_sum_prep_processor_queue_time_ms untyped zookeeper_sum_prep_processor_queue_time_ms # TYPE zookeeper_sum_propagation_latency untyped zookeeper_sum_propagation_latency # TYPE zookeeper_sum_proposal_ack_creation_latency untyped zookeeper_sum_proposal_ack_creation_latency # TYPE zookeeper_sum_proposal_latency untyped zookeeper_sum_proposal_latency # TYPE zookeeper_sum_quorum_ack_latency untyped zookeeper_sum_quorum_ack_latency # TYPE zookeeper_sum_read_commit_proc_issued untyped zookeeper_sum_read_commit_proc_issued # TYPE zookeeper_sum_read_commit_proc_req_queued untyped zookeeper_sum_read_commit_proc_req_queued # TYPE zookeeper_sum_read_commitproc_time_ms untyped zookeeper_sum_read_commitproc_time_ms # TYPE zookeeper_sum_read_final_proc_time_ms untyped zookeeper_sum_read_final_proc_time_ms # TYPE zookeeper_sum_readlatency untyped zookeeper_sum_readlatency # TYPE zookeeper_sum_reads_after_write_in_session_queue untyped zookeeper_sum_reads_after_write_in_session_queue # TYPE zookeeper_sum_reads_issued_from_session_queue untyped zookeeper_sum_reads_issued_from_session_queue # TYPE zookeeper_sum_requests_in_session_queue untyped zookeeper_sum_requests_in_session_queue # TYPE zookeeper_sum_server_write_committed_time_ms untyped zookeeper_sum_server_write_committed_time_ms # TYPE zookeeper_sum_session_queues_drained untyped zookeeper_sum_session_queues_drained # TYPE zookeeper_sum_snapshottime untyped zookeeper_sum_snapshottime # TYPE zookeeper_sum_startup_snap_load_time untyped zookeeper_sum_startup_snap_load_time # TYPE zookeeper_sum_startup_txns_load_time untyped zookeeper_sum_startup_txns_load_time # TYPE zookeeper_sum_startup_txns_loaded untyped zookeeper_sum_startup_txns_loaded # TYPE zookeeper_sum_sync_process_time untyped zookeeper_sum_sync_process_time # TYPE zookeeper_sum_sync_processor_batch_size untyped zookeeper_sum_sync_processor_batch_size # TYPE zookeeper_sum_sync_processor_queue_and_flush_time_ms untyped zookeeper_sum_sync_processor_queue_and_flush_time_ms # TYPE zookeeper_sum_sync_processor_queue_flush_time_ms untyped zookeeper_sum_sync_processor_queue_flush_time_ms # TYPE zookeeper_sum_sync_processor_queue_size untyped zookeeper_sum_sync_processor_queue_size # TYPE zookeeper_sum_sync_processor_queue_time_ms untyped zookeeper_sum_sync_processor_queue_time_ms # TYPE zookeeper_sum_time_waiting_empty_pool_in_commit_processor_read_ms untyped zookeeper_sum_time_waiting_empty_pool_in_commit_processor_read_ms # TYPE zookeeper_sum_updatelatency untyped zookeeper_sum_updatelatency # TYPE zookeeper_sum_write_batch_time_in_commit_processor untyped zookeeper_sum_write_batch_time_in_commit_processor # TYPE zookeeper_sum_write_commit_proc_issued untyped zookeeper_sum_write_commit_proc_issued # TYPE zookeeper_sum_write_commit_proc_req_queued untyped zookeeper_sum_write_commit_proc_req_queued # TYPE zookeeper_sum_write_commitproc_time_ms untyped zookeeper_sum_write_commitproc_time_ms # TYPE zookeeper_sum_write_final_proc_time_ms untyped zookeeper_sum_write_final_proc_time_ms # TYPE zookeeper_sum_zk_cluster_management_write_per_namespace untyped zookeeper_sum_zk_cluster_management_write_per_namespace # TYPE zookeeper_sum_zookeeper_read_per_namespace untyped zookeeper_sum_zookeeper_read_per_namespace # TYPE zookeeper_sum_zookeeper_write_per_namespace untyped zookeeper_sum_zookeeper_write_per_namespace # TYPE zookeeper_sync_processor_request_queued untyped zookeeper_sync_processor_request_queued # TYPE zookeeper_synced_followers untyped zookeeper_synced_followers # TYPE zookeeper_synced_non_voting_followers untyped zookeeper_synced_non_voting_followers # TYPE zookeeper_synced_observers untyped zookeeper_synced_observers # TYPE zookeeper_tls_handshake_exceeded untyped zookeeper_tls_handshake_exceeded # TYPE zookeeper_unrecoverable_error_count untyped zookeeper_unrecoverable_error_count # TYPE zookeeper_uptime untyped zookeeper_uptime # TYPE zookeeper_watch_count untyped zookeeper_watch_count # TYPE zookeeper_znode_count untyped zookeeper_znode_count ``` --- # System tables in Aiven for ClickHouse® Aiven for ClickHouse® supports multiple types of system tables, which store metadata and system-level information. Querying system tables allows you to check the configuration, performance, and state of your database. ## Supported system tables[​](#supported-system-tables "Direct link to Supported system tables") Aiven for ClickHouse supports the [system tables available with the open-source ClickHouse](https://clickhouse.com/docs/en/operations/system-tables), with the exceptions below. The following system tables are not accessible in 26.3 (all columns are restricted): * `background_schedule_pool` * `build_options` * `certificates` * `database_replicas` * `dns_cache` * `fail_points` * `histogram_metrics` * `instrumentation` * `jemalloc_profile_text` * `jemalloc_stats` * `models` * `primes` * `stack_trace` * `symbols` * `tokenizers` * `unicode` * `user_defined_functions` * `user_directories` * `zookeeper` * `zookeeper_info` A few system tables are accessible but hide specific columns: `asynchronous_inserts`, `azure_queue_metadata_cache`, `distributed_ddl_queue`, `remote_data_paths`, and `s3queue_metadata_cache`. Most open-source system log tables are also not enabled. See [Supported system log tables](/docs/products/clickhouse/reference/clickhouse-system-tables.md#supported-system-log-tables). ## Supported system log tables[​](#supported-system-log-tables "Direct link to Supported system log tables") System log tables store data related to traces, queries, performance metrics, errors, and more. By recording logs and events, they allow you to monitor, debug, and audit your database. Aiven for ClickHouse enables the following system log tables: * `asynchronous_insert_log` * `part_log` * `query_log` * `query_views_log` * `text_log` * `trace_log` Other open-source ClickHouse system log tables (for example `metric_log`, `asynchronous_metric_log`, `query_thread_log`, `session_log`, `blob_storage_log`, `processors_profile_log`, `crash_log`, and `backup_log`) are not enabled. ### System log tables TTL[​](#system-log-tables-ttl "Direct link to System log tables TTL") In Aiven for ClickHouse, time-to-live (TTL) for system log tables is fixed to 1 hour. This means data in system log tables is kept for 1 hour before being automatically deleted. ### Persist data with materialized views[​](#persist-data-with-materialized-views "Direct link to Persist data with materialized views") You can work around the TTL of 1 hour by creating a [materialized view](/docs/products/clickhouse/howto/materialized-views.md) to save system log table data so that you can retrieve and use it later. To achieve this, create a materialized view similar to the following: ``` CREATE MATERIALIZED VIEW query_log ENGINE = MergeTree PARTITION BY event_date ORDER BY event_time AS SELECT * FROM system.query_log; ``` --- # Aiven for ClickHouse® limits and limitations By respecting the Aiven for ClickHouse® restrictions and quotas, you can improve the security and productivity of your service workloads. ## Limitations[​](#limitations "Direct link to Limitations") From the information about restrictions on using Aiven for ClickHouse, you can draw conclusions on how to get your service to operate closer to its full potential. Use **Recommended approach** as guidelines on how to work around specific restrictions. | Name | Description | Recommended approach | | ----------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Backups - one snapshot a day | Since Aiven for ClickHouse service takes a single snapshot a day only:- Point-in-time recovery is not supported. A database can be [restored to one of the daily backup states](/docs/products/clickhouse/howto/restore-backup.md) only.
- When creating a database fork, you can only [create a fork](/docs/products/clickhouse/howto/fork-service.md) that matches the state of one of the backups.
- Any data inserted before the next snapshot is lost if all nodes in a given shard malfunction and need to be replaced. This limitation doesn't apply to patches, migrations, or scaling, which are handled safely and automatically. | N/A | | Table engines support | -Some special table engines are not supported in Aiven for ClickHouse.
-Some engines are remapped to their `Replicated` alternatives, for example, `MergeTree` **>** `ReplicatedMergeTree`. | Use the available table engines listed in [Supported table engines in Aiven for ClickHouse](/docs/products/clickhouse/reference/supported-table-engines.md). | | Log table engine support | Log engine is not supported in Aiven for ClickHouse. | For storing data, use the [Buffer engine](https://clickhouse.com/docs/en/engines/table-engines/special/buffer/) instead of the Log engine. | | Kafka table engine support | The Kafka table engine is supported via [integration](/docs/products/clickhouse/howto/integrate-kafka.md) only, not by creating a table in SQL. | N/A | | Kafka Schema Registry | Kafka Schema Registry is supported with Aiven for Apache Kafka® and not with an external Kafka endpoint. | N/A | | Cloud availability | Available on AWS, GCP, and Azure only | Use the available cloud providers. | | Querying all shards at once | If you have a sharded plan, you must use a distributed table on top of your MergeTree table to query all the shards at the same time, and you should use it for inserts too. | Use a distributed table with sharded plans. See [Query data across shards](/docs/products/clickhouse/howto/use-shards-with-distributed-table.md). | | Creating or deleting a database using SQL | - Only the `avnadmin` user can create databases in SQL.
- You can create a database in SQL with the `Replicated` database engine only.
- By default, only the `aiven` user can delete a database. The `aiven` user can grant the permission to delete a database to another user. | To work around this limitation, use [Aiven's API](https://api.aiven.io/doc/) or the [Aiven Console](https://console.aiven.io/). | | Maximum number of databases per service | Your Aiven for ClickHouse service can support up to 400 databases simultaneously. | Instead of creating multiple databases of the same structure for isolation purposes, it's recommended to create one database where you add an extra column to filter by it. Consider including this column into the primary key or partitioning by it. You can also limit scope of data on which SQL queries can be run by using the `additional_table_filters` query setting. | ## Limits[​](#limits "Direct link to Limits") Service limits are determined by a plan that this service uses. For the number of VMs, CPU per VM, RAM per VM, storage details per plan, see [Plans & pricing](https://aiven.io/pricing?product=clickhouse). | Aiven for ClickHouse | Hobbyist | Startup | Business | Premium | | ------------------------------ | -------------------------- | ---------------------------- | ---------------------------- | ---------------------------- | | Maximum concurrent queries | 25 queries per 4 GB of RAM | 100 queries per 16 GB of RAM | 100 queries per 16 GB of RAM | 100 queries per 16 GB of RAM | | Maximum concurrent connections | 1000 connections per node | 4000 connections per node | 4000 connections per node | 4000 connections per node | Total storage with a plan Total storage represents the maximum amount of data you can insert into a service, and it doesn't depend on the number of nodes. The inserted data is replicated on all available nodes. How many times it's replicated depends on the number of nodes and the number of shards: ``` number_of_data_replication_times = number_of_nodes / number_of_shards ``` **Examples** * Service plan with **one shard** A Startup-16 plan or a Business-16 plan, each has 1150 GB of total storage per VM. Since the Business-16 plan offers three VMs, your total storage is 3450 GB, but effectively it's still 1150 GB because that’s the maximum a single node can hold. * Service plan with **two shards** A Premium-6x-16 plan has two shards and six servers, each server with 1150 GB of storage. The data you insert is replicated three times, held between two shards. tip If you need a custom plan with capacity beyond the listed limits, [contact us](https://aiven.io/contact?department=1306714). Related pages * [Quotas for specific tiers of Business and Premium plans](https://aiven.io/pricing?tab=plan-pricing\&product=clickhouse) * [Plans comparison](https://aiven.io/pricing?tab=plan-comparison\&product=clickhouse) --- # Aiven for ClickHouse® monitoring dashboard metrics shown in Grafana® Browse the Aiven for ClickHouse® service metrics shown in the Grafana® monitoring dashboard. ## Current counts[​](#current-counts "Direct link to Current counts") The following metrics are displayed on the monitoring dashboard as counts reflecting the current status: * `Databases`: Number of databases * `Tables`: Number of tables * `Parts`: Total number of MergeTree table parts * `Rows total`: Total number of rows * `Running Queries`: Total number of queries currently executed * `Running Merges`: Number of merges currently executed in the background * `Parts Merged`: Number of times data parts of ReplicatedMergeTree tables were successfully merged * `Readonly Tables`: Number of replicated tables that are currently in the read-only state due to re-initialization after ZooKeeper session loss or due to startup without ZooKeeper configured ## Charts[​](#charts "Direct link to Charts") The following metrics are displayed on the monitoring dashboard as charts reflecting data changes over time: ### `Connections`[​](#connections "Direct link to connections") * Number of connections to the TCP server (clients with native interface), also included server-server distributed query connections * Number of connections to the HTTP server * Number of connections from other replicas to fetch parts ### `Queries`[​](#queries "Direct link to queries") * Number of executed SELECT queries (includes internal/monitoring queries) * Number of executed INSERT queries * Number of failed queries ### `Delayed Queries`[​](#delayed-queries "Direct link to delayed-queries") * Number of INSERT queries that are throttled due to high number of active data parts for partition in a MergeTree table * Number of queries that are stopped and waiting due to 'priority' setting ### `Replicated Part Fetches`[​](#replicated-part-fetches "Direct link to replicated-part-fetches") * Number of data parts being checked for consistency * Number of data parts being fetched from replica Each fetch consumes disk IO and inter-server network IO. ### Miscellaneous[​](#miscellaneous "Direct link to Miscellaneous") * `Running Merges` Merge operations reduce the number of parts and improve SELECT performance. Each merge consumes disk IO (read and write). A new part will be merged approximately 15 min after being inserted. * `Inserted Rows`: Number of rows INSERTed to all tables * `Inserted Bytes`: Number of bytes (uncompressed) INSERTed to all tables * `Merged Rows`: Rows read for background merges (before merge) * `Merged: Data Read`: Uncompressed bytes that were read for background merges * `Compressed Data Read`: Number of bytes before decompression read from compressed sources * `Rows Delta`: Difference in number of rows between nodes * `Parts`: Total number of MergeTree table parts. Each INSERT creates a new part. * `Replicated Part Merged` --- # S3 table function file formats in Aiven for ClickHouse® The [S3 table function](https://clickhouse.com/docs/en/sql-reference/table-functions/s3) lets you select and insert data in S3-compatible storage. In Aiven for ClickHouse®, the S3 table function supports the following file formats: * `Arrow` * `CSV` * `JSON` * `TSV` * `Parquet` * `ORC` * `Avro` Related pages * [Table functions supported in Aiven for ClickHouse®](/docs/products/clickhouse/reference/supported-table-functions.md) * [Read and pull data from S3 object storage and web resources over HTTP](/docs/products/clickhouse/howto/run-federated-queries.md) --- # Supported database engines in Aiven for ClickHouse® A database engine controls how a database manages tables and metadata. It determines what happens when you list, create, or delete tables and can restrict a database to specific table engines or manage replication. This differs from [table engines](/docs/products/clickhouse/reference/supported-table-engines.md), which determine how data is stored on disk or read from external sources and exposed as virtual tables. Aiven for ClickHouse® supports different database engines depending on the service version. Engine availability can differ from upstream depending on features enabled in Aiven for ClickHouse®. For details about each engine, see the [ClickHouse database engines documentation](https://clickhouse.com/docs/en/engines/database-engines/). ## Supported database engines[​](#supported-database-engines "Direct link to Supported database engines") | Engine | Supported in | | ------------------- | ---------------- | | `Atomic` | 25.3, 25.8, 26.3 | | `Lazy` | 25.3, 25.8 | | `MaterializedMySQL` | 25.3, 25.8, 26.3 | | `Memory` | 25.3, 25.8, 26.3 | | `MySQL` | 25.3, 25.8, 26.3 | | `Ordinary` | 25.3, 25.8, 26.3 | | `PostgreSQL` | 25.3, 25.8, 26.3 | | `Replicated` | 25.3, 25.8, 26.3 | note The `Lazy` database engine is not available in 26.3. It was removed upstream in 26.2. Move tables to a supported database engine and drop `Lazy` databases before upgrading. note On Aiven, new databases can be created in SQL only with the `Replicated` engine. See [Limits and limitations](/docs/products/clickhouse/reference/limitations.md). --- # Formats for Aiven for ClickHouse® - Aiven for Apache Kafka® data exchange When connecting Aiven for ClickHouse® to Aiven for Apache Kafka® using Aiven integrations, data exchange is possible with the following formats only: | Format name | Notes | | --------------------------- | ----------------------------------------------------------------------------------------------------------- | | `Avro` | Binary Avro format with embedded schema. Libraries and documentation: | | `AvroConfluent` | Binary Avro with schema registry. Requires the Karapace Schema Registry to be enabled in the Kafka service. | | `CSV` | Example: `123,"Hello"` | | `JSONAsString` | Example: `{"x":123,"y":"hello"}` | | `JSONCompactEachRow` | Example: `[123,"Hello"]` | | `JSONCompactStringsEachRow` | Example: `["123","Hello"]` | | `JSONEachRow` | Example: `{"x":123,"y":"hello"}` | | `JSONStringsEachRow` | Example: `{"x":"123","y":"hello"}` | | `MsgPack` | Example: `{\\xc4\\x05hello`. Libraries and documentation: | | `Parquet` | Binary parquet format. Libraries and documentation: | | `RawBLOB` | Raw binary data with no delimiters or structure. Reads/writes values as a single blob. | | `TSKV` | Example: `x=123\ty=hello` | | `TSV` | Example: `123\thello` | | `TabSeparated` | Example: `123\thello` | --- # Interfaces and drivers supported in Aiven for ClickHouse® Find out what technologies and tools you can use to interact with Aiven for ClickHouse®. ## Interfaces (protocols)[​](#clickhouse-interfaces "Direct link to Interfaces (protocols)") Aiven for ClickHouse® supports the following fundamental underlying interfaces (protocols): * `HTTPS` * `Native TCP` * `MySQL Interface` * `Arrow Flight` note `Arrow Flight` is available for new services running ClickHouse 25.8 LTS or 26.3 LTS. To get the connection details, use the [Aiven API](/docs/tools/api.md), [Aiven CLI](/docs/tools/cli.md), or [Aiven MCP server](/docs/tools/mcp-server.md). note For security reasons, TLS is required to connect to Aiven for ClickHouse. The following interfaces (protocols) are not supported: * `HTTP` * `gRPC` * `PostgreSQL` For the full list of interfaces and protocols supported in ClickHouse, see [Drivers and Interfaces](https://clickhouse.com/docs/en/interfaces/overview). ## Drivers (libraries)[​](#drivers-libraries "Direct link to Drivers (libraries)") There are a number of drivers (libraries) that use one of [the fundamental underlying interfaces supported in Aiven for ClickHouse](/docs/products/clickhouse/reference/supported-interfaces-drivers.md#clickhouse-interfaces) under the hood. It's up to you to pick up a driver (library) of your choice and use it for connecting to your Aiven for ClickHouse service. * See [how to use different drivers (libraries) for connecting to Aiven for ClickHouse](/docs/products/clickhouse/howto/list-connect-to-service.md). * For the full list of drivers and libraries that support connecting to ClickHouse, see [Drivers and Interfaces](https://clickhouse.com/docs/en/interfaces/overview). note You can connect to Aiven for ClickHouse with any driver that uses TLS and one of the supported protocols. Related pages * [How to connect to Aiven for ClickHouse using different libraries](/docs/products/clickhouse/howto/list-connect-to-service.md) * [Drivers and interfaces supported in ClickHouse](https://clickhouse.com/docs/en/interfaces/overview) --- # Supported table engines in Aiven for ClickHouse® Table engines define how data is stored and which queries a table supports. Aiven for ClickHouse® supports different table engines depending on the service version. Engine availability can differ from upstream depending on features enabled in Aiven for ClickHouse®. For details about each engine, see the [ClickHouse documentation](https://clickhouse.com/docs/en/engines/table-engines/). For database engines, see [Supported database engines](/docs/products/clickhouse/reference/supported-database-engines.md). note To support high availability and node replacement operations on the Aiven platform, some table engines are automatically remapped. In replicated databases, `MergeTree` engines are replaced with their corresponding `ReplicatedMergeTree` variants. ## MergeTree family[​](#mergetree-family "Direct link to MergeTree family") | Engine | Replicated variant | Supported in | | ----------------------------------------- | ---------------------------------------- | ---------------- | | `MergeTree` (remapped) | `ReplicatedMergeTree` | 25.3, 25.8, 26.3 | | `AggregatingMergeTree` (remapped) | `ReplicatedAggregatingMergeTree` | 25.3, 25.8, 26.3 | | `CoalescingMergeTree` (remapped) | `ReplicatedCoalescingMergeTree` | 25.8, 26.3 | | `CollapsingMergeTree` (remapped) | `ReplicatedCollapsingMergeTree` | 25.3, 25.8, 26.3 | | `GraphiteMergeTree` (remapped) | `ReplicatedGraphiteMergeTree` | 25.3, 25.8, 26.3 | | `ReplacingMergeTree` (remapped) | `ReplicatedReplacingMergeTree` | 25.3, 25.8, 26.3 | | `SummingMergeTree` (remapped) | `ReplicatedSummingMergeTree` | 25.3, 25.8, 26.3 | | `VersionedCollapsingMergeTree` (remapped) | `ReplicatedVersionedCollapsingMergeTree` | 25.3, 25.8, 26.3 | ## Integration engines[​](#integration-engines "Direct link to Integration engines") | Engine | Supported in | | ------------------------------------------------------------------------------------------------- | ---------------- | | `AzureBlobStorage` | 25.3, 25.8, 26.3 | | `AzureQueue` | 25.3, 25.8, 26.3 | | `COSN` | 25.3, 25.8, 26.3 | | `GCS` | 25.8, 26.3 | | `Kafka` (via [integration](/docs/products/clickhouse/howto/integrate-kafka.md) only, not via SQL) | 25.3, 25.8, 26.3 | | `MaterializedPostgreSQL` | 25.3, 25.8, 26.3 | | `MySQL` | 25.3, 25.8, 26.3 | | `OSS` | 25.3, 25.8, 26.3 | | `PostgreSQL` (via SQL or [integration](/docs/products/clickhouse/howto/integrate-postgresql.md)) | 25.3, 25.8, 26.3 | | `S3` | 25.3, 25.8, 26.3 | | `S3Queue` | 25.3, 25.8, 26.3 | | `URL` | 25.3, 25.8, 26.3 | ## Data lake engines[​](#data-lake-engines "Direct link to Data lake engines") | Engine | Supported in | | ---------------- | ---------------- | | `DeltaLake` | 25.3, 25.8, 26.3 | | `DeltaLakeAzure` | 25.8, 26.3 | | `DeltaLakeS3` | 25.8, 26.3 | | `Hudi` | 25.3, 25.8, 26.3 | | `Iceberg` | 25.3, 25.8, 26.3 | | `IcebergAzure` | 25.3, 25.8, 26.3 | | `IcebergS3` | 25.3, 25.8, 26.3 | ## View engines[​](#view-engines "Direct link to View engines") | Engine | Supported in | | ------------------ | ---------------- | | `LiveView` | 25.3, 25.8 | | `MaterializedView` | 25.3, 25.8, 26.3 | | `View` | 25.3, 25.8, 26.3 | | `WindowView` | 25.3, 25.8, 26.3 | note `LiveView` was removed upstream and is not available in 26.3. Replace live views before upgrading, for example with refreshable materialized views. ## Special-purpose engines[​](#special-purpose-engines "Direct link to Special-purpose engines") | Engine | Supported in | | ------------- | ---------------- | | `Alias` | 26.3 | | `Buffer` | 25.3, 25.8, 26.3 | | `Dictionary` | 25.3, 25.8, 26.3 | | `Distributed` | 25.3, 25.8, 26.3 | | `Join` | 25.3, 25.8, 26.3 | | `KeeperMap` | 25.3, 25.8, 26.3 | | `Memory` | 25.3, 25.8, 26.3 | | `Merge` | 25.3, 25.8, 26.3 | | `Null` | 25.3, 25.8, 26.3 | | `Set` | 25.3, 25.8, 26.3 | | `TimeSeries` | 25.8, 26.3 | ## Testing and utility engines[​](#testing-and-utility-engines "Direct link to Testing and utility engines") | Engine | Supported in | | ---------------- | ---------------- | | `FuzzJSON` | 25.3, 25.8, 26.3 | | `FuzzQuery` | 25.3, 25.8, 26.3 | | `GenerateRandom` | 25.3, 25.8, 26.3 | | `Loop` | 25.3, 25.8, 26.3 | tip [Managed credentials integrations](/docs/products/clickhouse/concepts/data-integration-overview.md#managed-credentials-integration) simplify authentication for integration engines and dictionaries that access remote sources. Enable them when using these engines. [Learn how to enable managed credentials](/docs/products/clickhouse/howto/data-service-integration.md#create-managed-credentials-integrations). --- # Table functions supported in Aiven for ClickHouse® [Table functions](https://clickhouse.com/docs/en/sql-reference/table-functions) can be used to construct tables, for example, in a FROM clause of a query or in an INSERT INTO TABLE FUNCTION statement. Sample usage of the S3 table function ``` SELECT * FROM deltaLake('s3://bucket/path/to/lake') ``` note Occasionally, you may find specific table functions disabled for security reasons. Aiven for ClickHouse® supports the following table functions. Availability can differ by service version. | Table function | Supported in | | ----------------------------- | ---------------- | | `azureBlobStorage` | 25.3, 25.8, 26.3 | | `azureBlobStorageCluster` | 25.3, 25.8, 26.3 | | `cluster` | 25.3, 25.8, 26.3 | | `clusterAllReplicas` | 25.3, 25.8, 26.3 | | `cosn` | 25.3, 25.8, 26.3 | | `deltaLake` | 25.3, 25.8, 26.3 | | `deltaLakeAzure` | 25.8, 26.3 | | `deltaLakeAzureCluster` | 26.3 | | `deltaLakeCluster` | 25.3, 25.8, 26.3 | | `deltaLakeS3` | 25.8, 26.3 | | `deltaLakeS3Cluster` | 26.3 | | `dictionary` | 25.3, 25.8, 26.3 | | `format` | 25.3, 25.8, 26.3 | | `fuzzJSON` | 25.3, 25.8, 26.3 | | `fuzzQuery` | 25.3, 25.8, 26.3 | | `gcs` | 25.3, 25.8, 26.3 | | `generateRandom` | 25.3, 25.8, 26.3 | | `generateSeries` | 25.3, 25.8, 26.3 | | `generate_series` | 25.3, 25.8, 26.3 | | `hudi` | 25.3, 25.8, 26.3 | | `hudiCluster` | 25.3, 25.8, 26.3 | | `iceberg` | 25.3, 25.8, 26.3 | | `icebergAzure` | 25.3, 25.8, 26.3 | | `icebergAzureCluster` | 25.3, 25.8, 26.3 | | `icebergCluster` | 26.3 | | `icebergS3` | 25.3, 25.8, 26.3 | | `icebergS3Cluster` | 25.3, 25.8, 26.3 | | `input` | 25.3, 25.8, 26.3 | | `loop` | 25.3, 25.8, 26.3 | | `merge` | 25.3, 25.8, 26.3 | | `mergeTreeAnalyzeIndexes` | 26.3 | | `mergeTreeAnalyzeIndexesUUID` | 26.3 | | `mergeTreeIndex` | 25.3, 25.8, 26.3 | | `mergeTreeProjection` | 25.8, 26.3 | | `mergeTreeTextIndex` | 26.3 | | `mysql` | 25.3, 25.8, 26.3 | | `null` | 25.3, 25.8, 26.3 | | `numbers` | 25.3, 25.8, 26.3 | | `numbers_mt` | 25.3, 25.8, 26.3 | | `oss` | 25.3, 25.8, 26.3 | | `paimon` | 26.3 | | `paimonAzure` | 26.3 | | `paimonAzureCluster` | 26.3 | | `paimonCluster` | 26.3 | | `paimonS3` | 26.3 | | `paimonS3Cluster` | 26.3 | | `postgresql` | 25.3, 25.8, 26.3 | | `primes` | 26.3 | | `prometheusQuery` | 25.8, 26.3 | | `prometheusQueryRange` | 25.8, 26.3 | | `remoteSecure` | 25.3, 25.8, 26.3 | | `s3` | 25.3, 25.8, 26.3 | | `s3Cluster` | 25.3, 25.8, 26.3 | | `timeSeriesData` | 25.8, 26.3 | | `timeSeriesMetrics` | 25.8, 26.3 | | `timeSeriesSelector` | 25.8, 26.3 | | `timeSeriesTags` | 25.8, 26.3 | | `url` | 25.3, 25.8, 26.3 | | `urlCluster` | 25.3, 25.8, 26.3 | | `values` | 25.3, 25.8, 26.3 | | `view` | 25.3, 25.8, 26.3 | | `viewExplain` | 25.3, 25.8, 26.3 | | `viewIfPermitted` | 25.3, 25.8, 26.3 | | `zeros` | 25.3, 25.8, 26.3 | | `zeros_mt` | 25.3, 25.8, 26.3 | --- # Upgrade to Aiven for ClickHouse® 26.3 Aiven for ClickHouse® 26.3 is a long-term support (LTS) release. You can create a new service with version 26.3 or upgrade an existing service from version 25.8. Version 25.8 remains the default for new services. Version 26.3 enables full-text search, asynchronous inserts by default, materialized common table expressions (CTEs), the native `Geometry` type, and improvements to JSON processing and query performance. Before upgrading: * Review [Changes that require attention](#changes-that-require-attention). * Review the [26.3 default settings](/docs/products/clickhouse/reference/26-3-default-settings.md). * For production services, test the upgrade on a [service fork](/docs/products/clickhouse/howto/fork-service.md). Aiven does not support downgrades. * If your service runs version 25.3, upgrade it to version 25.8 first. Direct upgrades from version 25.3 to 26.3 are not supported. * Review the upstream changes introduced in versions 25.9 through 26.3. These changes are included when you upgrade from version 25.8. For the complete list of upstream changes, see the [ClickHouse changelog](https://clickhouse.com/docs/whats-new/changelog). ## Changes that require attention[​](#changes-that-require-attention "Direct link to Changes that require attention") Review these changes because they can affect existing workloads. ### Removed features[​](#removed-features "Direct link to Removed features") Migrate or remove these features before upgrading. | Removed feature | Removed in version | Action | | ----------------------------------------- | ------------------ | ----------------------------------------------------------------------------------------------------- | | `Object('json')` column type | 25.11 | Migrate columns to the [`JSON` type](https://clickhouse.com/docs/sql-reference/data-types/newjson). | | `LiveView` | 25.12 | Drop live views or replace them with refreshable materialized views. | | `DEFLATE_QPL` and `ZSTD_QAT` codecs | 26.2 | Recompress affected columns with a supported codec, such as `ZSTD`. | | Text indexes created in version 25.8 | 26.2 | Drop the indexes before upgrading. After the upgrade, recreate them using the current storage format. | | Experimental hypothesis skip indexes | 26.3 | Drop the experimental hypothesis skip indexes. | | `detectProgrammingLanguage()` function | 26.3 | Move language detection to your application or replace calls with a supported alternative. | | `searchAny()` and `searchAll()` functions | 25.10 | Replace calls with `hasAnyTokens()` and `hasAllTokens()`. | ### Automatic pre-upgrade checks[​](#automatic-pre-upgrade-checks "Direct link to Automatic pre-upgrade checks") Aiven automatically scans your service running version 25.8 for columns that use the removed `Object('json')` type, `LiveView` tables, and columns that use removed codecs. Aiven blocks the upgrade if it finds any of these. For `Object('json')` columns, Aiven also provides SQL statements to help you migrate them. The service must be running for Aiven to complete the scan. Aiven does not automatically check text indexes created in version 25.8, hypothesis skip indexes, or uses of `detectProgrammingLanguage()`. Before upgrading, check for and resolve: * Text indexes created in version 25.8 * Experimental hypothesis skip indexes * Uses of `detectProgrammingLanguage()` ### Insert behavior changes[​](#insert-behavior-changes "Direct link to Insert behavior changes") Aiven for ClickHouse 26.3 uses these upstream insert defaults: * **`async_insert` defaults to `1`.** In upstream ClickHouse 25.8, the default was `0`. The server batches small inserts instead of processing each insert synchronously. This change can affect when an insert is acknowledged and when its data becomes visible. To apply the setting defaults recorded for version 25.8, set `compatibility = '25.8'`. This compatibility setting does not restore all version 25.8 behavior or server settings. Introduced in 26.3. * **Insert deduplication applies to all insert paths.** This includes asynchronous inserts and inserts into dependent materialized views. The `deduplicate_insert` setting did not exist in version 25.8. To restore the previous deduplication behavior, set the following settings: ``` SET deduplicate_insert = 'backward_compatible_choice'; SET deduplicate_blocks_in_dependent_materialized_views = 0; ``` Introduced in version 26.2. * **The insert deduplication window is shorter.** The `replicated_deduplication_window_seconds` table setting changes from one week in version 25.8 to one hour. The `replicated_deduplication_window` setting changes from `1000` to `10000`. As a result, ClickHouse does not treat an insert retried more than one hour after the original insert as a duplicate. Introduced in version 25.10. In Aiven for ClickHouse 26.3, `compatibility` defaults to an empty value. During migration, you can set `compatibility = '25.8'` at the session or profile level to apply earlier recorded setting defaults. This setting does not reproduce every previous behavior or server setting. For details, see [Changed default settings and behavior](#changed-default-settings-and-behavior). ### On-disk format changes[​](#on-disk-format-changes "Direct link to On-disk format changes") Aiven for ClickHouse 26.3 can write new part formats that earlier ClickHouse versions cannot read. Because Aiven does not support downgrades, test the upgrade on a service fork before upgrading your production service. * **`object_serialization_version` changes from `v2` to `v3`.** This setting controls `JSON` column serialization compatibility. New parts can use a layout that earlier versions cannot read. Introduced in 25.12. * **`propagate_types_serialization_versions_to_nested_types` defaults to `1`.** This applies serialization versions to nested types. Earlier ClickHouse versions cannot read parts that use this serialization. Introduced in 26.3. * **`String` columns use a new serialization format.** When `serialization_info_version` is set to `with_types`, ClickHouse uses the `with_size_stream` format. Versions earlier than 25.10 cannot read parts written in this format. Introduced in 25.11. * **Bucketed `Map` storage remains disabled by default.** No action is required unless you plan to use the new format. To enable it, set `map_serialization_version = 'with_buckets'`. Introduced in 26.3. * **Filenames are escaped in wide parts.** This applies to tables with `Variant`, `Dynamic`, or `JSON` subcolumns. Introduced in 25.11. * **Rebuild affected column statistics after upgrading.** If you use column statistics for a column changed from `String` to `Nullable(String)`, run `ALTER TABLE ... MATERIALIZE STATISTICS ALL`. Introduced in versions 25.12 through 26.1. ### Query result and privilege changes[​](#query-result-and-privilege-changes "Direct link to Query result and privilege changes") * **`NOT` operator precedence changed.** It follows standard SQL precedence. Review predicates that use `NOT`, especially expressions that do not use parentheses to make the evaluation order explicit. Introduced in 26.3. * **Some `FINAL` queries read more partitions.** These queries cannot apply partition pruning before row merging. This preserves correctness but can increase the number of partitions read. Benchmark affected queries on the upgraded service fork, especially when the partition key is not part of the sorting key. Introduced in 26.3. * **`joinGet()` and `joinGetOrNull()` require the `SELECT` privilege on the underlying `Join` table.** Grant the privilege to users or roles that call these functions. Introduced in 26.2. * **`CREATE TABLE ... AS` requires `SHOW COLUMNS`.** Grant the privilege to users or roles that create tables from existing tables. Introduced in 26.2. * **Nullable-to-non-nullable conversions require an explicit `DEFAULT`.** Add a default expression when you use `ALTER TABLE ... MODIFY COLUMN` to make a nullable column non-nullable. Introduced in 25.12. * **Special MergeTree engines require a non-empty `ORDER BY`.** This applies to engines such as `ReplacingMergeTree` and `CollapsingMergeTree`. Introduced in 25.12. * **Schema inference uses source nullability metadata.** Schema inference now uses the nullability metadata provided by the source format instead of making every inferred column nullable (`schema_inference_make_columns_nullable` defaults to `3`). Review `s3()`, `url()`, and `file()` queries that expected all-nullable columns. Introduced in 25.10. * **`Dynamic` columns are disabled in `JOIN` keys by default.** `allow_dynamic_type_in_join_keys` defaults to `0`. Cast the column to a concrete type, or set `allow_dynamic_type_in_join_keys = 1` to restore the previous behavior. Introduced in 25.10. * **`JOIN ... USING` requires explicit key columns.** Rewrite affected joins to specify one or more join key columns in the `USING` clause. Introduced in 26.1. * **Invalid IP binary operations are rejected.** Update expressions that apply binary operators to `IPv4` or `IPv6` values and non-integer operands. Introduced in 25.9. * **`JSON` `SKIP REGEXP` uses partial matching by default.** `type_json_use_partial_match_to_skip_paths_by_regexp` defaults to `1`, so a pattern skips any path it matches partially. Review `SKIP REGEXP` patterns used in `JSON` column definitions. Introduced in 26.1. * **Formatter alias substitution can affect stored views.** A formatter fix can change stored `CREATE VIEW` statements that use `IN`. After upgrading, recheck those views. Introduced in 26.1. ### Integration behavior changes[​](#integration-behavior-changes "Direct link to Integration behavior changes") Aiven for ClickHouse exposes the `MySQL`, `PostgreSQL`, and `Kafka` table engines, so these integration defaults change when you upgrade: * **MySQL and PostgreSQL date and timestamp columns map to wider ClickHouse types.** `DATE` columns map to `Date32`, and `DATETIME` or timestamp columns with precision map to `DateTime64`. This avoids range errors for pre-1970 dates but changes inferred column types for existing integrations. Review downstream queries that assumed the previous types. For MySQL integrations, this behavior is controlled by `mysql_datatypes_support_level`, which defaults to `decimal,datetime64,date2Date32`. Introduced in 26.2 and 26.3. * **Kafka table-level SASL settings in `CREATE TABLE` take precedence over configuration-file settings.** If the same SASL setting exists in both the table definition and the configuration file, verify that the table-level value is correct before upgrading. Introduced in version 25.11. For a summary, see [Changed default settings and behavior](#changed-default-settings-and-behavior). ## New features and improvements[​](#new-features-and-improvements "Direct link to New features and improvements") Aiven for ClickHouse 26.3 includes the following features and improvements from upstream ClickHouse versions 25.9 through 26.3. ### Query performance[​](#query-performance "Direct link to Query performance") * **Materialized CTEs.** The `WITH x AS MATERIALIZED (...)` syntax evaluates a CTE once and stores the result in a temporary table. This feature is experimental and disabled by default on Aiven (`enable_materialized_cte = 0`). It requires the analyzer, which Aiven enables by default. * **Generally available full-text search.** Text indexes provide native inverted indexes for token-based search and no longer require experimental settings. If you created text indexes in 25.8, follow the steps in [Removed features](#removed-features). The tokenizer syntax also changed in version 25.10. Use the current `tokenizer = '...'` syntax when recreating the indexes. * **Expanded automatic `JOIN` reordering.** Automatic reordering now supports `ANTI`, `SEMI`, and `FULL` joins and requires table statistics. * **More flexible skip indexes.** Skip indexes support more complex query conditions. * **Improved outer join performance.** `RIGHT` and `FULL OUTER` joins have improved performance. * **Lower-memory `TTL DELETE` merges.** `TTL DELETE` merges use less memory on wide tables. ### JSON and data types[​](#json-and-data-types "Direct link to JSON and data types") * **`JSONExtract*` support for `JSON`.** `JSONExtract*` functions work directly on the `JSON` type. * **Lower-memory reads for large `JSON` objects.** Queries that read only a few subcolumns of a large `JSON` object use less memory through more accurate subcolumn size estimation. * **Lazy JSON type hints.** Lazy JSON type hints support metadata-only `ALTER ... MODIFY COLUMN data JSON(...)` operations. This feature is experimental and disabled by default on Aiven. * **Native `Geometry` type.** The `Geometry` type supports WKB and WKT representations. Introduced in 25.11. * **Experimental ALP compression.** The ALP codec compresses floating-point columns and is disabled by default on Aiven. ### Ingestion and integrations[​](#ingestion-and-integrations "Direct link to Ingestion and integrations") * **Bucketed `Map` storage.** Bucketed `Map` storage can improve single-key lookups in `Map` columns. The format is disabled by default. * **S3Queue ordered mode.** S3Queue ordered mode lists only new objects, which reduces the number of `ListObjects` requests. * **Experimental Delta Lake writes.** Delta Lake writes to existing tables are experimental and disabled by default on Aiven. * **Apache Paimon table functions.** Apache Paimon tables can be read through the `paimon*` table functions. Aiven disables Paimon partition pruning by default (`use_paimon_partition_pruning = 0`). * **Azure table-function variants.** Azure table-function variants, such as `deltaLakeAzure` and `icebergAzure`, support Azure-backed data lakes, including Microsoft OneLake. ### SQL and operational improvements[​](#sql-and-operational-improvements "Direct link to SQL and operational improvements") * **Readable query plans.** `EXPLAIN pretty=1, compact=1` produces readable query plans. * **Human-friendly string ordering.** `naturalSortKey()` supports human-friendly string ordering. * **Fractional and negative `LIMIT` and `OFFSET` values.** `LIMIT` and `OFFSET` support fractional and negative values. For example, use `LIMIT 0.25` to select a fraction of rows or a negative offset to count from the end. ### Experimental previews[​](#experimental-previews "Direct link to Experimental previews") These upstream features are experimental and disabled by default on Aiven for ClickHouse 26.3. Test them on a service fork before using them in production. * **`QBit` vector type.** The experimental `QBit` type supports tunable approximate vector search. The transposed distance functions, including `L2DistanceTransposed` and `cosineDistanceTransposed`, are available separately. * **Polyglot SQL transpiler.** The transpiler converts supported SQL dialects to ClickHouse SQL. It is controlled by `allow_experimental_polyglot_dialect`, which defaults to `0`. ## Changed default settings and behavior[​](#changed-default-settings-and-behavior "Direct link to Changed default settings and behavior") The following table lists ClickHouse default changes from 25.8 to 26.3 that Aiven adopts and that can affect workloads. | Setting or behavior | Previous behavior | New behavior | Introduced | Effect | | -------------------------------------------------------- | ------------------------------------------- | ------------------------------------ | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `async_insert` | `0` (25.8) | `1` | 26.3 | Small inserts are batched server-side. Setting `compatibility = '25.8'` applies the earlier default of `0` but does not restore every version 25.8 behavior. | | `deduplicate_insert` | Unavailable | Enabled for all insert paths | 26.2 | Deduplication applies to all inserts, including asynchronous inserts and inserts into dependent materialized views. To restore earlier behavior, set this to `backward_compatible_choice` and `deduplicate_blocks_in_dependent_materialized_views` to `0`. | | `replicated_deduplication_window_seconds` | `604800` (1 week) | `3600` (1 hour) | 25.10 | Retried inserts are deduplicated only within this window. A retry after one hour is no longer treated as a duplicate. | | `schema_inference_make_columns_nullable` | `1` | `3` | 25.10 | Schema inference respects Parquet/ORC/Arrow nullability metadata instead of making all columns nullable. | | `allow_dynamic_type_in_join_keys` | Unavailable | `0` | 25.10 | `Dynamic` columns are rejected in `JOIN` keys unless set to `1`. | | `object_serialization_version` | `v2` | `v3` | 25.12 | Controls `JSON` column serialization compatibility. New parts can use a layout that earlier versions cannot read. | | `propagate_types_serialization_versions_to_nested_types` | Unavailable | `1` | 26.3 | Applies serialization versions to nested types. Earlier ClickHouse versions cannot read parts that use this serialization. | | `apply_row_policy_after_final` | Unavailable | `1` | 26.2 | Row policies are applied after `FINAL`, so row filtering happens after final row merging. | | `cpu_slot_preemption` | `0` | `1` | 26.1 | Enables preemptive CPU slot scheduling. `compatibility` does not control this setting. | | Kafka SASL precedence | Configuration file settings take precedence | Table-level settings take precedence | 25.11 | Verify table-level Kafka engine credentials when the same SASL setting also exists in the configuration file. | ## Upgrade checklist[​](#upgrade-checklist "Direct link to Upgrade checklist") 1. If your service runs version 25.3, upgrade it to 25.8. Only services on 25.8 can upgrade directly to 26.3. 2. Keep the service on version 25.8 powered on until Aiven completes the pre-upgrade scan for removed `Object('json')` columns, `LiveView` tables, and codecs. 3. Check for and remove or migrate text indexes created in version 25.8, experimental hypothesis skip indexes, and uses of `detectProgrammingLanguage()`. 4. [Fork your service](/docs/products/clickhouse/howto/fork-service.md) and upgrade the fork. 5. Validate continuous ingestion and insert acknowledgment behavior on the fork. 6. Test external integrations, including Kafka engine authentication. 7. Test queries that use `NOT`, `FINAL`, dynamic join keys, or schema inference. 8. Verify the privilege changes described in [Query result and privilege changes](#query-result-and-privilege-changes). 9. Compare application results and query performance with the original service. 10. Follow [Manage versions](/docs/products/clickhouse/howto/manage-clickhouse-versions.md) to upgrade the production service. Related pages * [Aiven for ClickHouse® version support policy](/docs/products/clickhouse/reference/version-support-policy.md) * [Fork your Aiven for ClickHouse® service](/docs/products/clickhouse/howto/fork-service.md) * [Supported table engines](/docs/products/clickhouse/reference/supported-table-engines.md) --- # Aiven for ClickHouse® version lifecycle Learn how Aiven manages Aiven for ClickHouse® version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. ## Aiven version support and upstream EOL[​](#aiven-version-support-and-upstream-eol "Direct link to Aiven version support and upstream EOL") Aiven aims to follow the EOL schedule set by the original authors and maintainers of the open source software (the upstream projects). Once the upstream project retires a specific version, they do not receive security updates and critical bug fixes anymore by the maintainers. Outdated services don't offer the level of protection you need, so Aiven follows the upstream project's EOL schedule to ensure that Aiven services are always running on supported versions. ## Service version numbering[​](#service-version-numbering "Direct link to Service version numbering") Aiven services inherit the upstream project's software versioning scheme. Depending on the service, a major version can be either a single digit or in the format `major.minor`. The exact version of the service is visible in the [Aiven Console](https://console.aiven.io/) when the service is running. ## Service version EOL policy[​](#service-version-eol-policy "Direct link to Service version EOL policy") Aiven sets an EOL date for each major version of the service. This policy covers both running and powered-off services on affected versions. ## EOL notifications[​](#eol-notifications "Direct link to EOL notifications") When Aiven sets the EOL date for a service major version: * You receive an email notification along with instructions on the next steps. * The [Aiven Console](https://console.aiven.io/) shows an EOL alert for affected services. * You receive email reminders monthly. * In the month of the EOL date, you receive weekly reminders. ## EOL best practices[​](#eol-best-practices "Direct link to EOL best practices") * Use service forking to test the version upgrade before upgrading your production services. * Upgrade to the supported version before the EOL date. This gives you time to test compatibility, resolve any issues, and plan the upgrade on your schedule. After the EOL date: * If the service is powered on, it's automatically upgraded to the latest version when possible, or to another supported version. note If it's not possible to upgrade a powered-on service to a supported version, the service is powered off and ultimately deleted. * If the service is powered off, it's deleted. ## Version EOL dates[​](#version-eol-dates "Direct link to Version EOL dates") The following table lists the EOL and service creation dates for supported Aiven for ClickHouse versions. For the long-term support release model, lifecycle stages, and security updates, see [Aiven for ClickHouse version support policy](/docs/products/clickhouse/reference/version-support-policy.md). | Version | Aiven EOL | Service creation supported until | Service creation supported from | | ------- | ---------- | -------------------------------- | ------------------------------- | | 25.3 | 2026-09-30 | 2026-08-17 | 2025-12-15 | | 25.8 | 2027-02-28 | 2026-11-30 | 2026-03-25 | | 26.3 | 2027-09-15 | 2027-06-15 | 2026-08-01 | Related pages * [Manage Aiven for ClickHouse® versions](/docs/products/clickhouse/howto/manage-clickhouse-versions.md) * [Maintenance and updates](/docs/products/clickhouse/howto/maintenance-updates.md) * [Fork your Aiven for ClickHouse® service](/docs/products/clickhouse/howto/fork-service.md) * [Upgrade to Aiven for ClickHouse 26.3](/docs/products/clickhouse/reference/upgrade-to-26-3.md) --- # Aiven for ClickHouse® version support policy Aiven for ClickHouse® follows the upstream ClickHouse long-term support, or LTS, release model. This helps you use stable, supported versions and plan upgrades before versions reach end of life. Aiven for ClickHouse versions move through defined lifecycle stages, from availability to end of life. Each stage determines whether you can create new services, receive security updates, or need to plan an upgrade. ## Upstream LTS releases[​](#upstream-lts-releases "Direct link to Upstream LTS releases") Upstream ClickHouse typically releases two LTS versions each year, one in March and one in August. The March release uses the `.3` prefix, and the August release uses the `.8` prefix. Upstream LTS versions receive security updates for 12 months. ## Lifecycle stages[​](#lifecycle-stages "Direct link to Lifecycle stages") Each Aiven for ClickHouse LTS version moves through the following stages: * **Limited Availability (LA):** The version is available for selected early testing. LA versions are not publicly announced. * **Early Availability (EA):** The version is stable and available to all customers. Some settings, features, and UI elements might still change. * **General Availability (GA):** The version is stable and fully available. * **End of Availability (EOA):** You can no longer create new services with the version. Aiven doesn't guarantee security updates after EOA. EOA starts 3 months before End of Life. * **End of Life (EOL):** Aiven removes the version from support and automatically upgrades existing services to the next supported version. ## Support policy[​](#support-policy "Direct link to Support policy") Aiven supports two active generally available LTS versions at a time. New LTS versions become available on Aiven: * As EA approximately three months after the upstream release. * As GA approximately four months after the upstream release. Aiven announces new EA and GA versions in the [Aiven changelog](https://aiven.io/changelog?services=ClickHouse%2CClickHouse%25C2%25AE). When a new version becomes available, new service creation stops for the oldest supported version, and that version enters EOA. EOL follows 3 months later. After EOL, Aiven automatically upgrades services running that version to the next supported version. ## Security updates[​](#security-updates "Direct link to Security updates") Aiven provides security updates for LTS versions until their EOA date. After EOA, Aiven doesn't guarantee security updates. Aiven delivers security updates as maintenance updates that you can apply during the [maintenance window](/docs/products/clickhouse/howto/maintenance-updates.md#set-the-maintenance-window). If you don't apply a patch for a critical vulnerability within 14 days, Aiven applies it automatically. ## Version lifecycle dates[​](#version-lifecycle-dates "Direct link to Version lifecycle dates") For EOA, EOL, and new service creation dates for supported Aiven for ClickHouse versions, see [Aiven for ClickHouse version lifecycle](/docs/products/clickhouse/reference/version-lifecycle.md). Related pages * [Manage Aiven for ClickHouse® versions](/docs/products/clickhouse/howto/manage-clickhouse-versions.md) * [Upgrade to Aiven for ClickHouse 26.3](/docs/products/clickhouse/reference/upgrade-to-26-3.md) * [Service and feature releases](/docs/platform/concepts/service-and-feature-releases.md) * [Fork your Aiven for ClickHouse® service](/docs/products/clickhouse/howto/fork-service.md) * [Maintenance and updates](/docs/products/clickhouse/howto/maintenance-updates.md) --- # Aiven for DataHub Aiven for [DataHub](https://docs.datahub.com/docs/features) is a cost-effective data catalog that integrates seamlessly with Aiven services. It removes complexity by automatically deploying and managing the whole infrastructure. You don’t have to manually create or manage the underlying Aiven Apps, or Aiven for Apache Kafka®, Aiven for OpenSearch® and Aiven for PostgreSQL® services. Additionally, all services are deployed in the same region to simplify networking and reduce operational overhead. You can [add an unlimited number of users](/docs/products/datahub/manage-datahub-users.md) with no per-user licensing costs. ## Features[​](#features "Direct link to Features") Aiven for DataHub is built on DataHub Core and offers the following features: * **Connector setup and ingestion**: Discover and [connect to data sources](/docs/products/datahub/connect-datahub-to-services.md) with simplified ingestion workflows for Aiven services. * **Data discovery**: [Search](https://docs.datahub.com/docs/how/search) across datasets, columns, dashboards, and other metadata. * **Lineage**: See how data flows across systems to understand dependencies and impact. * **AI context**: Use DataHub as a context layer for AI through its [MCP server](/docs/products/datahub/datahub-mcp-server.md). note DataHub Cloud features aren’t supported by Aiven for DataHub. ## Get started with Aiven for DataHub[​](#get-started-with-aiven-for-datahub "Direct link to Get started with Aiven for DataHub") Start exploring Aiven for DataHub by creating your first DataHub service and log in to the DataHub UI with the [get started guide](/docs/products/datahub/get-started.md). --- # Change cloud for Aiven for DataHub You can change the cloud provider or region of an Aiven for DataHub service at any time. Required roles or permissions:role:project:manager, `project:services:write`, `service:configuration:write`, or `role:project:admin` 1. In your project, click **Services**. 2. Open your DataHub service. 3. In the **Cloud and network** section, click **Actions** > **Change cloud or deployment model**. 4. Select a cloud provider and region. 5. Optional: Select a different plan. 6. Click **Change**. --- # Configure Slack notifications for DataHub activity Get activity notifications for your DataHub service in a Slack channel, including new datasets, ownership changes, tags, and glossary updates. You can enable Slack notifications by configuring a Slack app and setting environment variables on the actions app. Required roles or permissions:`project:services:write` or `role:project:admin` ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A Slack workspace where you can create and install apps. * The ID of the Slack channel to send notifications to. ## Create and configure a Slack app[​](#create-and-configure-a-slack-app "Direct link to Create and configure a Slack app") 1. [Create a Slack app](https://docs.slack.dev/app-management/quickstart-app-settings). 2. On the **OAuth & Permissions** page, add these [app scopes](https://docs.slack.dev/app-management/quickstart-app-settings#scopes): * `chat:write`: Post messages as the Slack bot * `chat:write.public`: Post in public channels without being a member * `channels:read`: Look up channel IDs 3. [Install the app](https://docs.slack.dev/app-management/quickstart-app-settings#installing). 4. Get the following app credentials: * [Slack bot token](https://api.slack.com/authentication/token-types#bot): On the **OAuth & Permissions** page. Bot tokens begin with `xoxb-`. * [Signing secret](https://api.slack.com/authentication/verifying-requests-from-slack): From the **Basic Info** section. * The Slack channel ID: In the channel details. These IDs start with `C`. 5. For private channels: To allow the app to post messages, [add it to the channel](https://slack.com/help/articles/201398103-Add-an-app-to-a-channel). For public channels, the `chat:write.public` scope lets the bot post without being a member. ## Enable Slack notifications in DataHub[​](#enable-slack-notifications-in-datahub "Direct link to Enable Slack notifications in DataHub") 1. In your DataHub service, go to the **DataHub resources** section. 2. Open the Aiven App that ends in `-actions`. 3. In the **Environment variables** section, click **Edit**. 4. On the **Secrets** tab, add the following secrets: | Key | Value | | -------------------------------------- | --------------------- | | `DATAHUB_ACTIONS_SLACK_BOT_TOKEN` | Your Slack bot token. | | `DATAHUB_ACTIONS_SLACK_SIGNING_SECRET` | Your signing secret. | 5. On the **Variables** tab, add the following variables: | Key | Value | | ------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------- | | `DATAHUB_ACTIONS_SLACK_ENABLED` | `true` | | `DATAHUB_ACTIONS_SLACK_CHANNEL` | Your Slack channel ID. | | `DATAHUB_ACTIONS_SLACK_DATAHUB_BASE_URL` | The DataHub **Application URL** from the **Connection information**. Adds links in messages. | | `DATAHUB_ACTIONS_SLACK_SUPPRESS_SYSTEM_ACTIVITY` | Optional. To get low-level system activity notifications such as datasets being ingested, set to `false`. Defaults to `true`. | 6. Click **Save**. After setting the variables, the actions app restarts automatically. Related pages * [Configure Teams notifications](/docs/products/datahub/configure-teams-notifications.md) --- # Configure Teams notifications for DataHub activity [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Get activity notifications for your DataHub service in a Microsoft Teams channel, including new datasets, ownership changes, tags, and glossary updates. You can enable Teams notifications by creating a Power Automate flow in Teams and setting environment variables on the actions app. Required roles or permissions:`project:services:write` or `role:project:admin` ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A Microsoft Teams team and channel where you can add workflows. * Permission to create and manage Power Automate flows in that team. ## Create a workflow in Teams[​](#create-a-workflow-in-teams "Direct link to Create a workflow in Teams") 1. In Teams, open the channel to send notifications to. 2. In the channel menu, click **Workflows** > **Post to a channel when a webhook request is received**. 3. Enter a name for the workflow and confirm the team and channel. 4. Open the workflow and configure it: * **Trigger**: When a Teams webhook request is received * **Action**: Post message in a chat or channel * **Message body expression**: `coalesce(triggerBody()?['text'], string(triggerBody()))` 5. Copy the generated webhook URL. important Treat the webhook URL as a secret. If the URL is exposed, rotate it by deleting and recreating the workflow. 6. In the Power Automate flow settings, set **Who can trigger the flow?** to **Anyone**. ## Enable Teams notifications in DataHub[​](#enable-teams-notifications-in-datahub "Direct link to Enable Teams notifications in DataHub") 1. In your DataHub service, go to the **DataHub resources** section. 2. Open the Aiven App that ends in `-actions`. 3. In the **Environment variables** section, click **Edit**. 4. On the **Secrets** tab, add the following secret: | Key | Value | | ----------------------------------- | -------------------------- | | `DATAHUB_ACTIONS_TEAMS_WEBHOOK_URL` | The generated webhook URL. | 5. On the **Variables** tab, add the following variables: | Key | Value | | ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- | | `DATAHUB_ACTIONS_TEAMS_ENABLED` | `true` | | `DATAHUB_ACTIONS_TEAMS_DATAHUB_BASE_URL` | The DataHub **Application URL** from the **Connection information**. Adds links in messages. Defaults to `http://localhost:9002`. | 6. Click **Save**. After setting the variables, the actions app restarts automatically. Related pages * [Configure Slack notifications](/docs/products/datahub/configure-slack-notifications.md) --- # Connect DataHub to services Add connectors to your DataHub service to ingest data. Required roles or permissions:For the DataHub service: `role:project:admin` or `project:integrations:write`. For the service you are connecting, you must have permission to manage service users: `role:project:admin`, `role:project:manager`, or `service:users:write`. ## Connect Aiven services[​](#connect-aiven-services "Direct link to Connect Aiven services") 1. In your DataHub service, click **Connectors**. 2. Click **Add connectors**. 3. Select services from projects in your organization and organizational units. 4. Click **Add connectors**. A service user is created in each connected service to give DataHub read access to the service. To view connected services in the DataHub UI, click **Data Sources**. ## Connect external services to DataHub[​](#connect-external-services-to-datahub "Direct link to Connect external services to DataHub") You can add external data sources to your DataHub service to allow it to automatically ingest metadata from them. You can add connectors through the [DataHub UI](https://docs.datahub.com/docs/ui-ingestion) or the [DataHub CLI](https://docs.datahub.com/docs/metadata-ingestion/cli-ingestion). important Store sensitive information like passwords and API keys used for connectors in DataHub [Secrets](https://docs.datahub.com/docs/ui-ingestion#managing-sensitive-information-with-secrets). ## Remove an Aiven service connector[​](#remove-an-aiven-service-connector "Direct link to Remove an Aiven service connector") To remove an Aiven service connector: 1. In your DataHub service, click **Connectors**. 2. Find the connector and click **Actions** > **Remove**. 3. To confirm, click **Remove**. To remove an external service connector, use the [DataHub UI](https://docs.datahub.com/docs/ui-ingestion) or the [DataHub CLI](https://docs.datahub.com/docs/metadata-ingestion/cli-ingestion). --- # View data lineage in DataHub Data lineage is a map of how each of your data assets moves across your systems. Lineage can help you: * Understand your data flows at a glance, even in complex architectures. * See which downstream systems would be affected before you change a dataset. * Trace data-quality problems or failures back to their source. * Build trust by showing stakeholders and auditors where data came from. When you [connect Aiven services to your Aiven for DataHub service](/docs/products/datahub/connect-datahub-to-services.md), DataHub automatically builds this map of your data assets. In DataHub, data assets are called datasets. note External services that you connect to DataHub are not shown in the data lineage. DataHub UI permissions:`Reader`, `Editor`, or `Admin`
View the [DataHub access control policies documentation](https://docs.datahub.com/docs/authorization/policies) ## View lineage for a dataset[​](#view-lineage-for-a-dataset "Direct link to View lineage for a dataset") 1. In the Aiven Console, go to your DataHub service. 2. [Log in to DataHub](/docs/products/datahub/get-started.md#log-in-to-datahub). 3. Search for and open a dataset. 4. Click **Lineage**. --- # Use the DataHub MCP server Make your data ecosystem visible to AI agents with the DataHub MCP server, enabling natural language search, end-to-end lineage tracking, and context-aware SQL generation. DataHub UI permissions:`Generate Personal Access Tokens` or`Manage All Access Tokens`.
View the [DataHub access control policies documentation](https://docs.datahub.com/docs/authorization/policies) ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * The [uv Python package and project manager installed](https://docs.astral.sh/uv/) * A [DataHub personal access tokens (PAT)](https://docs.datahub.com/docs/authentication/personal-access-tokens) * The [DataHub GMS URL](#get-the-datahub-gms-url) ## Get the DataHub GMS URL[​](#get-the-datahub-gms-url "Direct link to Get the DataHub GMS URL") 1. In the Aiven Console, go to your DataHub service. 2. In the **Connection information** section, copy the **Application URL**. 3. Add `/api/gms` to the end of the copied URL. ## Configure the MCP server[​](#configure-the-mcp-server "Direct link to Configure the MCP server") To run the [open-source MCP server](https://github.com/acryldata/mcp-server-datahub) locally, configure your AI assistant: * Claude Desktop * Cursor * Other AI tools 1. To find the full path to the `uvx` command, run: ``` which uvx ``` 2. Add the following content to your `claude_desktop_config.json` file: ``` { "mcpServers": { "datahub": { "command": "FULL-PATH-TO-UVX", // For example: /Users/hsheth/.local/bin/uvx "args": ["mcp-server-datahub@latest"], "env": { "DATAHUB_GMS_URL": "DATAHUB_APPLICATION_URL", "DATAHUB_GMS_TOKEN": "DATAHUB_PAT" } } } } ``` Where: * `FULL-PATH-TO-UVX` is the full path to the `uvx` command you found. * `DATAHUB_APPLICATION_URL` is the [DataHub GMS URL](#get-the-datahub-gms-url). * `DATAHUB_PAT` is the personal access token you created in the DataHub UI. 1) Add the following content to your `.cursor/mcp.json` file: ``` { "mcpServers": { "datahub": { "command": "uvx", "args": ["mcp-server-datahub@latest"], "env": { "DATAHUB_GMS_URL": "DATAHUB_APPLICATION_URL", "DATAHUB_GMS_TOKEN": "DATAHUB_PAT" } } } } ``` Where: * `DATAHUB_APPLICATION_URL` is the [DataHub GMS URL](#get-the-datahub-gms-url). * `DATAHUB_PAT` is the personal access token you created in the DataHub UI. The model context protocol (MCP) is an open standard supported by many clients. The most common configurations require: * The [DataHub GMS URL](#get-the-datahub-gms-url) * Your DataHub [personal access token](https://docs.datahub.com/docs/authentication/personal-access-tokens) * Command: `uvx` * Args: `mcp-server-datahub@latest` The following standard MCP JSON configuration format works with most clients: ``` "datahub": { "command": "/Users/USER_NAME/.local/bin/uvx", "args": [ "mcp-server-datahub@latest" ], "env": { "DATAHUB_GMS_URL": "DATAHUB_APPLICATION_URL", "DATAHUB_GMS_TOKEN": "DATAHUB_PAT" } ``` Where: * `DATAHUB_APPLICATION_URL` is the [DataHub GMS URL](#get-the-datahub-gms-url). * `DATAHUB_PAT` is the personal access token you created in the DataHub UI. --- # Delete an Aiven for DataHub service When you delete an Aiven for DataHub service, all service data and configuration are permanently deleted. Required roles or permissions:`project:services:write` or `role:project:admin` When you delete an Aiven for DataHub service, all service data and configuration are permanently deleted. The underlying Aiven for Apache Kafka®, Aiven for OpenSearch® and Aiven for PostgreSQL® services are also deleted at the same time. To stop DataHub, you can [power it off](/docs/products/datahub/power-off-service.md) instead. ## Delete a service[​](#delete-a-service "Direct link to Delete a service") * Console * CLI 1. In your project, click **Services**. 2. Open the service to delete, and click **Actions** > **Delete service**. To delete a service, run: ``` avn service terminate SERVICE_NAME ``` --- # Enable OIDC authentication for Aiven for DataHub Use OpenID Connect (OIDC) to configure single sign-on (SSO) to your DataHub service with your identity provider. You can use any OIDC compliant provider such as Auth0, Okta, Google Identity, or Azure AD. When OIDC is enabled, all users are redirected to SSO login by default. Required roles or permissions:`project:services:write` or `role:project:admin` ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * **An application created in your identity provider**: Follow the [DataHub OIDC prerequisites guide](https://docs.datahub.com/docs/authentication/guides/sso/initialize-oidc) to create and register a Google Identity, Okta Identity, or Azure AD application. * To get the domain of your DataHub service for the redirect URIs, open the Aiven App that ends in `-frontend` and copy the **Application URL**. * **The client ID, client secret, and discovery URI for your OIDC provider**: The [DataHub OIDC prerequisites guide](https://docs.datahub.com/docs/authentication/guides/sso/initialize-oidc) has instructions for getting these values from Azure AD, Google Identity, and Okta. * **Your DataHub service URL**: To get the URL, in the **Connection information** copy the **Application URL**. ## Enable OIDC authentication[​](#enable-oidc-authentication "Direct link to Enable OIDC authentication") 1. In your DataHub service, go to the **DataHub resources** section. 2. Open the Aiven App that ends in `-frontend`. 3. In the **Environment variables** section, click **Edit**. 4. On the **Secrets** tab, add a secret. For the **Key** enter `AUTH_OIDC_CLIENT_SECRET` and for the **Value** enter your client secret. 5. On the **Variables** tab, add the following variables: | Key | Value | | ------------------------- | ------------------------- | | `AUTH_OIDC_ENABLED` | `true` | | `AUTH_OIDC_CLIENT_ID` | Your client ID. | | `AUTH_OIDC_DISCOVERY_URI` | Your discovery URI. | | `AUTH_OIDC_BASE_URL` | Your DataHub service URL. | After adding the secrets, wait for the frontend container to redeploy before using the application. To log in using username and password instead, add `/login` to the end of the DataHub application URL. Related pages * [DataHub guide on configuring OIDC Authentication](https://docs.datahub.com/docs/authentication/guides/sso/configure-oidc-react) --- # Enable Prometheus metrics for Aiven for DataHub Enable Prometheus metrics for your Aiven for DataHub service to monitor its performance and health. The service exposes operational metrics including request rates, latencies, and resource usage. You can enable Prometheus metrics and secure them by setting environment variables in the GMS application. Required roles or permissions:`role:project:admin`, `role:project:operator`, or `project:services:write` ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A [DataHub personal access token (PAT)](https://docs.datahub.com/docs/authentication/personal-access-tokens) ## Enable Prometheus metrics[​](#enable-prometheus-metrics "Direct link to Enable Prometheus metrics") 1. In your DataHub service, go to the **DataHub resources** section. 2. Open the Aiven App that ends in `-gms`. 3. In the **Environment variables** section, click **Edit**. 4. On the **Variables** tab, add the following variables: | Key | Value | | ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `MANAGEMENT_ENDPOINT_PROMETHEUS_ENABLED` | `true` | | `MANAGEMENT_SERVER_PORT` | `8080` | | `SPRING_APPLICATION_JSON` | `{"authentication":{"excludedPaths":"/schema-registry/,/health,/health/live,/health/detailed,/config,/config/search/export,/public-iceberg/,/openapi/operations/dev/featureFlags,/openapi/operations/dev/featureFlags/*"}}` | 5. Click **Save**. ## Scrape Prometheus metrics[​](#scrape-prometheus-metrics "Direct link to Scrape Prometheus metrics") Use the following endpoint to scrape metrics: ``` https://GMS_URL/actuator/prometheus ``` Where `GMS_URL` is your [DataHub GMS URL](/docs/products/datahub/datahub-mcp-server.md#get-the-datahub-gms-url). Example Prometheus configuration: ``` scrape_configs: - job_name: datahub-gms metrics_path: /actuator/prometheus scheme: https authorization: type: Bearer credentials: DATAHUB_ACCESS_TOKEN static_configs: - targets: - GMS_URL ``` Where: * `DATAHUB_ACCESS_TOKEN` is your DataHub personal access token. * `GMS_URL` is your DataHub GMS URL Requests without a valid token return an HTTP `401 Unauthorized` status code. ## Access related metrics[​](#access-related-metrics "Direct link to Access related metrics") You can get related metrics in JSON format at the `/actuator/metrics` endpoint. ## Security[​](#security "Direct link to Security") To restrict access to the metrics to authenticated users, do not expose port `4319` and do not include Prometheus in the excluded paths. --- # Fork DataHub services Fork an Aiven for DataHub service to create a complete copy of it from its latest backups. This restores both its PostgreSQL metadata database and OpenSearch search index. Both stores are restored to the latest backup available at or before the recovery time you choose. Forked services are independent and don't share resources with or increase the load on the original service. Common use cases for forking include: * Creating a snapshot to analyze an issue. * Creating a development copy of your production environment. * Testing upgrades before applying them to production services. * Creating an instance in a different cloud provider, region, or with a different plan. * Renaming a service. Required roles or permissions:`project:services:write`, `role:services:recover`, or `role:project:admin` ## Create a fork[​](#create-a-fork "Direct link to Create a fork") 1. In your DataHub service, click **Service settings**. 2. In the **Service management** section, click **Actions** > **Fork service**. 3. Optional: Edit the name and select a project. 4. Optional: Choose a different cloud and plan. By default, the same cloud and plan as the original service are selected. 5. Click **Create fork**. To ensure that the forked service shows all the latest data, [run the restore indices task](/docs/products/datahub/restore-datahub-indices.md) on it. --- # Get started with Aiven for DataHub Start using DataHub by creating and configuring your first service. Required roles or permissions:`project:services:write` or `role:project:admin` important To avoid issues, don’t make any changes to the plans or other settings of the DataHub resources beyond what is documented. ## Create a DataHub service[​](#create-a-datahub-service "Direct link to Create a DataHub service") 1. In your project, click **Services**. 2. Click **Create service**. 3. Select **DataHub**. 4. Choose a **Cloud**. 5. Choose a **Plan**. 6. In the **Service basics**, enter a name for your service. 7. Optional: Add [service tags](/docs/products/datahub/tag-services.md). 8. In the **Service summary**, click **Create service**. While the service is being created, its status is **Rebuilding**. When the status is **Running**, you can start using the service. This typically takes couple of minutes and can vary between cloud providers and regions. ## Log in to DataHub[​](#log-in-to-datahub "Direct link to Log in to DataHub") 1. In the **Connection information** section, copy the **Password**. 2. To open the DataHub UI, click the **Application URL**. 3. For **Username**, enter `datahub`. 4. For **Password**, paste the password you copied. 5. Click **Login**. ## Next steps[​](#next-steps "Direct link to Next steps") * Explore the [DataHub home page](https://docs.datahub.com/docs/features/feature-guides/custom-home-page) * [Give users access](/docs/products/datahub/manage-datahub-users.md) to your DataHub service * [Add connectors](/docs/products/datahub/connect-datahub-to-services.md) and start ingesting data --- # Maintenance updates for Aiven for DataHub Manage maintenance updates and set the maintenance window for your Aiven for DataHub service. Required roles or permissions:`project:services:write`, `role:services:maintenance`, or `role:project:admin` ## Maintenance updates[​](#maintenance-updates "Direct link to Maintenance updates") Aiven applies some maintenance updates automatically. The following are the types of updates: * **Mandatory updates:** Security updates, quarterly patch releases, and platform updates that affect reliability or stability of the service nodes. * **Optional updates:** All other updates are initially optional. After six months, they become mandatory and are applied in the next week’s maintenance window. * **Periodic infrastructure updates:** Scheduled automatically for services with nodes active for 180 days and more. These updates are mandatory for all services, except those with maintenance turned off. Critical security updates are applied during the next available maintenance window. For other updates, Aiven gives you at least seven days' notice. Maintenance updates are also automatically applied during service upgrades. To view pending updates: * Console * CLI * API 1. In your service, click **Service settings**. 2. Go to the **Service management** section. Use the [`avn service get`](/docs/tools/cli/service-cli.md#avn_service_get) command. Use the [`service`](https://api.aiven.io/doc/#tag/Service/operation/ServiceGet) endpoint. ## Maintenance windows[​](#maintenance-windows "Direct link to Maintenance windows") The maintenance window is the time period when Aiven can automatically apply maintenance updates to a service. When an update becomes available, Aiven schedules it for the next available maintenance window for each service. The update runs in the first window after it becomes available, and can begin any time after the start time. For example, if a service has a maintenance window of Monday 12:00 UTC, and an update becomes available on Tuesday, the update will be applied on the following Monday. During maintenance, Aiven might restart or replace service nodes. This can cause brief connection interruptions, but services are designed to minimize downtime. Aiven performs maintenance in a rolling-forward style, creating new nodes alongside existing ones and retiring the old nodes after the upgrade completes. Major service upgrades are triggered manually. A manually triggered upgrade starts immediately, regardless of the maintenance window. important You cannot control the order in which services are updated. Each service updates according to its own configured maintenance window, and there is no guaranteed way to control the update sequence. Manual updates and maintenance window adjustments only help for non-critical updates. ### Set the maintenance window for a DataHub service[​](#set-the-maintenance-window-for-a-datahub-service "Direct link to Set the maintenance window for a DataHub service") To set the maintenance window for your service: * Console * Terraform 1. In the Aiven Console, open your service. 2. In the **Maintenance** section, click **Actions** > **Change maintenance window**. 3. Set the day and time. 4. Click **Save changes**. Use the `maintenance_window_dow` and `maintenance_window_time` attributes in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). You cannot change the maintenance window of the DataHub resource services. When you set the maintenance window for a DataHub service, the same window is set for those Aiven Apps, and the Aiven for Apache Kafka®, Aiven for OpenSearch® and Aiven for PostgreSQL® services. --- # Manage DataHub users Invite users to your DataHub service, giving them access to read or edit metadata. DataHub UI permissions:`Manage Users & Groups`
View the [DataHub access control policies documentation](https://docs.datahub.com/docs/authorization/policies) For programmatic access to your DataHub service, use DataHub’s [personal access tokens (PATs)](https://docs.datahub.com/docs/authentication/personal-access-tokens). ## Security best practices[​](#security-best-practices "Direct link to Security best practices") Don’t share the Aiven for DataHub default password. Instead, invite users to your DataHub service in the DataHub UI, assigning only the level of access they need. ## Invite users to DataHub[​](#invite-users-to-datahub "Direct link to Invite users to DataHub") 1. [Log in to DataHub](/docs/products/datahub/get-started.md#log-in-to-datahub) and click **Settings**. 2. Click **Users & Groups**. 3. Click **Invite Users**. 4. Select a [role](https://docs.datahub.com/docs/authorization/roles). 5. Click **Copy**. 6. Send the link to the users. After clicking the link, the users sign up with their email and create a password. ## Manage user roles[​](#manage-user-roles "Direct link to Manage user roles") The [DataHub documentation on roles](https://docs.datahub.com/docs/authorization/roles#assigning-roles) has instructions on assigning roles to individual users, and details of the privileges for each role. For more fine-grained control over permissions, use [policies](https://docs.datahub.com/docs/authorization/policies). ### Assign roles to groups of users[​](#assign-roles-to-groups-of-users "Direct link to Assign roles to groups of users") You can also create groups of users to more easily manage user access: 1. [Log in to DataHub](/docs/products/datahub/get-started.md#log-in-to-datahub) and click **Settings**. 2. Click **Users & Groups**. 3. Click **Create group**. 4. Enter a name for the group. 5. Click **Create**. 6. Assign a [role](https://docs.datahub.com/docs/authorization/roles) to the group. 7. Click the group name. 8. On the **Members** tab, click **Add Member**. 9. Select the users to add to the group and click **Add**. Related pages * [Resetting user passwords](https://docs.datahub.com/docs/authentication/guides/add-users#resetting-user-passwords) --- # Permissions for Aiven for DataHub features The following roles and permissions are required for specific Aiven for DataHub features. [View all roles and permissions](/docs/platform/concepts/permissions.md) for Aiven organizations and projects. | Action | Required roles and permissions | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Add and remove Aiven service connectors | For the DataHub service: `operator`, `admin`, `role:project:admin`, `role:organization:admin`, or `project:integrations:write`. For the service you are connecting, you must have permission to create service users: `role:project:admin`, `role:project:manager`, or `service:users:write`. | | View DataHub UI connection information including: URL, username, and password | `admin`, `operator`, `developer`, `role:organization:admin`, `service:secrets:read`, `service:users:write`

Users with `read_only`, `role:project:read`, `role:project:admin`, or other service permissions can view the URL, but not the password. | | Edit application environment variables to:
- Enable Slack notifications
- Enable Teams notifications
- Enable OIDC authentication
- Reindex search and graph indices | `admin`, `operator`, `role:project:admin`, `role:organization:admin`, `role:services:maintenance`, `role:services:recover`, `project:services:write`, or `service:configuration:write`

To edit the variables in the Aiven Console, users also need to have access to read secrets through one of the following permissions: `admin`, `operator`, `role:organization:admin`, or `service:secrets:read`.

The `developer` role can view secret values, but cannot change them. | | Rotate secrets | `admin`, `operator`, or `role:organization:admin` | --- # Power an Aiven for DataHub service off or on You can power off an Aiven for DataHub service at any time to stop all processes and reduce costs. Powering off a DataHub service also powers off its component services. Metadata ingestion stops and you cannot access the DataHub UI or API. Ingestion restarts and access is restored when you power the service back on. Required roles or permissions:`project:services:write` or `role:project:admin` ## Power off a service[​](#power-off-a-service "Direct link to Power off a service") When you power off a service: * All virtual machines are removed from the public cloud. * The service configuration is stored on the Aiven Platform. * If there are no backups, all service data is lost. * If the service has time-based or point in time recovery backups, the backups remain on the Aiven Platform. Services powered off for more than 180 days are automatically deleted. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power off service**. To power off a service, run: ``` avn service update SERVICE_NAME --power-off ``` ## Power on a service[​](#power-on-a-service "Direct link to Power on a service") When you power on a service: * New virtual machines are created on the service's public cloud. * The service starts with the stored configuration parameters. * The latest time-based backup is restored. * Maintenance updates are automatically applied. * If a point in time recovery backup is available, the database transaction logs are replayed to recover the service data to a specific point in time. The restoration takes from a few minutes to a few hours, depending on the network bandwidth, the disk IOPS allocated to the service, and the size of the backup. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power on service**. To power on a service, run: ``` avn service update SERVICE_NAME --power-on ``` To see when the service is running, run: ``` avn service wait SERVICE_NAME ``` --- # Reindex Aiven for DataHub search and graph indices [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Rebuild your OpenSearch indices for search and graph data if your search results or relationship graphs differ from the data in your metadata database. This is useful: * After OpenSearch® data loss * When an index is corrupt or inconsistent * After wiping a cluster or re-provisioning the search backend * After a schema or mapping change that requires a full reindexing * For disaster recovery where SQL is intact, but OpenSearch is not To reindex your indices, run the `RestoreIndices` upgrade task. This task rebuilds the indices from the source of truth `metadata_aspect_v2` SQL table. It replays every aspect from the database back into search and graph stores. You can run this at any time. Events are replayed asynchronously and existing reads keep working. Always test reindexing in a staging environment first. Required roles or permissions:`project:services:write` or `role:project:admin` ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Get the URL for the GMS app: 1. In your DataHub service, in the **DataHub resources** section, open the Aiven App that ends in `-gms`. 2. In the **Connection information** section, copy the **Application URL**. ## Run the restore indices task[​](#run-the-restore-indices-task "Direct link to Run the restore indices task") 1. In your DataHub service, go to the **DataHub resources** section. 2. Open the Aiven App that ends in `-upgrade`. 3. In the **Environment variables** section, click **Edit**. 4. On the **Variables** tab, add the following variables: | Key | Value | Description | | -------------------------- | ---------------------------------- | --------------------------------------------------- | | `UPGRADE_JOB` | `RestoreIndices` | The restore indices task. | | `KAFKA_SCHEMAREGISTRY_URL` | `GMS_APP_URL/schema-registry/api/` | Queries Kafka topic schemas for re-emitting events. | The `GMS_APP_URL` is the application URL for the GMS app. 5. Optional: Add `UPGRADE_JOB_ARGS` to include additional arguments: | Arg | Description | | ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- | | `-a clean` | Wipes each index before repopulating. Use when an index has stale documents that you don't want to carry over. | | `-a batchSize` | Number of records per batch. | | `-a urnBasedPagination` | Uses URN-based pagination instead of offset-based. Set to `true` for large datasets. | | `-a aspectNames` | Comma-separated list of aspect names to reindex. Use to speed up partial recoveries. For example, `aspectNames=datasetProperties,ownership`. | | `-a urnLike` | SQL `LIKE` pattern to filter URNs. Use to target specific entity types. For example, `urnLike=urn:li:dataset:%` to reindex all datasets. | 6. Click **Save**. After setting the variables, the upgrade app restarts automatically. It's in the **Building** state until the reindexing completes. 7. When the upgrade app is in the **Powered off** state, remove the variables you added. --- # Rotate Aiven for DataHub secrets Rotate your Aiven for DataHub authentication and token secrets to maintain security and prevent unauthorized access. When you rotate authentication secrets, all logged-in users are signed out and all DataHub resource services restart simultaneously. Inter-service calls may briefly return `401 Unauthorized` errors during the restart window. When you rotate token secrets, every previously issued API access token stops working permanently. This includes personal access tokens and service tokens. Every DataHub application service restarts simultaneously. Browser sessions are not affected, so logged-in users stay logged in. Required roles or permissions:`project:services:write` or `role:project:admin` ## Rotate secrets[​](#rotate-secrets "Direct link to Rotate secrets") 1. In your DataHub service, click **Service settings**. 2. In the **Secret rotations** section, click **Rotate** for the authentication secrets or token secrets. 3. Click **Rotate now**. --- # Scale Aiven for DataHub services Scale your Aiven for DataHub service and its underlying resources to optimize costs and improve performance. Required roles or permissions:`project:services:write`, `role:project:manager`, or `role:project:admin` To scale a DataHub service, you can change its service plan. You can also scale the underlying Aiven for Apache Kafka®, Aiven for PostgreSQL®, and Aiven for OpenSearch® services. ## Scale an Aiven for DataHub service[​](#scale-an-aiven-for-datahub-service "Direct link to Scale an Aiven for DataHub service") Required roles or permissions:`project:services:write` or `role:project:admin` 1. In your project, click **Services**. 2. Open your DataHub service. 3. In the **Cloud and network** section, click **Actions** > **Change cloud or deployment model**. 4. In the **Plan** section, select a new plan. 5. Optional: Select a different **Cloud** provider or region. 6. Click **Change**. ## Scale underlying resources[​](#scale-underlying-resources "Direct link to Scale underlying resources") Beyond changing the DataHub service plan, you can also scale the underlying resources for your DataHub service. You can view service plan usage and metrics for each service on its page. More information on scaling and optimizing the underlying services is available on these pages: * PostgreSQL: [Change service plan](/docs/products/postgresql/howto/change-service-plan.md) * OpenSearch: [Change service plan](/docs/products/opensearch/howto/change-service-plan.md) * Apache Kafka®: [Scaling options](/docs/products/kafka/concepts/horizontal-vertical-scaling.md) * Apache Kafka®: [Optimize performance](/docs/products/kafka/howto/best-practices.md) --- # Tag Aiven for DataHub services Use tags to add metadata to Aiven services to categorize them or run custom logic on them. A tag is a key/value pair: * **Key**: A case-sensitive string that starts with a letter and consists of letters, numbers, dashes, and underscores. The maximum length for a key is 64 characters. * **Value**: A string limited to 64 UTF-8 characters. Within a service, the tag keys must be unique. Required roles or permissions:`project:services:write`, `service:configuration:write`, or `role:project:admin` ## Tag a DataHub service[​](#tag-a-datahub-service "Direct link to Tag a DataHub service") 1. In the service, click **Service settings**. 2. In the **Service status** section, click **Actions** > **Add service tags**. 3. Enter a key and value for each tag. 4. Click **Save changes**. --- # Upgrade Aiven for DataHub [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Upgrade your Aiven for DataHub service to a new version. Aiven for DataHub services are not automatically upgraded to the latest major version and minor patches are not automatically applied. You can, however, manually upgrade your service to a newer major version. Required roles or permissions:`project:services:write` or `role:project:admin` ## Upgrade a DataHub service[​](#upgrade-a-datahub-service "Direct link to Upgrade a DataHub service") 1. In your DataHub service, go to the **Maintenance** section. 2. Click **Actions** > **Upgrade version**. 3. Select a version to upgrade to, and click **Upgrade**. --- # Aiven for Dragonfly® Aiven for Dragonfly is an advanced, **high-scale, and Aiven for Caching compatible in-memory database service** that can be deployed in your preferred cloud environment. It provides lightning-fast data storage and retrieval capabilities, making it ideal for businesses that handle large-scale data operations. Dragonfly is designed to overcome the limitations of Redis Open Source Software (Redis OSS), especially under high-load conditions. While renowned for its speed and adaptability as an in-memory data repository, Redis encounters limitations when handling large-scale data management and achieving high throughput. Dragonfly addresses these challenges by enabling vertical scaling, optimizing hardware resources, and supporting larger memory footprints. With Aiven for Dragonfly, businesses can handle workloads exceeding 1 TB with more than 10 times the throughput performance of Redis OSS, along with reduced latency. This makes it an ideal solution for enterprises with growing data needs. note Aiven for Dragonfly is fully supported by Aiven's service level agreements (SLAs), ensuring its capability to manage production workloads. As the latest addition to Aiven services, we recommend initiating a proof of concept (PoC) with Aiven for Dragonfly to thoroughly evaluate its capabilities and confirm its fit for your production requirements. ## Features and benefits[​](#features-and-benefits "Direct link to Features and benefits") Aiven for Dragonfly offers numerous features and benefits: * **Redis compatibility at scale:** It is a seamless drop-in replacement for Redis, capable of handling extensive workloads with enhanced performance. * **Optimized for large-scale operations:** Dragonfly is specifically built to address the scalability and resource utilization limitations of Redis Open Source Software (Redis OSS). * **Advanced performance:** Dragonfly's unique threading model and shared-nothing architecture allow it to scale vertically, enhancing its performance efficiency, especially in environments with heavy data loads. * **Efficient backup and memory management:** Improved snapshot capabilities lead to more efficient memory usage during backups. * **High availability and replication:** It includes active-passive replication and persistence capabilities, ensuring data reliability and consistency. * **Ease of integration:** Dragonfly integrates smoothly with existing systems, requiring no code changes, simplifying the adoption process. ## Use cases[​](#use-cases "Direct link to Use cases") * **Data-intensive enterprises:** Ideal for businesses that necessitate robust, high-performance in-memory data storage and processing capabilities. * **Scaling and performance needs:** Perfectly suited for situations where the need for greater scalability and higher throughput goes beyond what Redis OSS can handle. Related pages * [Aiven.io](https://aiven.io/dragonfly) * [Dragonfly documentation](https://www.dragonflydb.io/docs) --- # High availability in Aiven for Dragonfly® Aiven for Dragonfly® offers different plans with varying levels of high availability. The available features depend on the selected plan. Refer to the table below for a summary of these plans: | Plan | Node configuration | High availability & Backup features | Backup history | | ------------ | ---------------------------------------- | ----------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ | | **Startup** | Single-node | Limited availability. No automatic failover. | During limited availability, only one latest snapshot stored. | | **Business** | Two-node (primary + standby) | High availability with automatic failover to a standby node if the primary fails. | During limited availability, only one latest snapshot stored. | | **Premium** | Three-node (primary + standby + standby) | Enhanced high availability with automatic failover among multiple standby nodes if the primary fails. | During limited availability, only one latest snapshot stored. | | **Custom** | Custom configurations | Custom high availability and failover features based on user requirements. | During limited availability, only one latest snapshot stored. Custom based on user requirements. | ## Failure handling[​](#failure-handling "Direct link to Failure handling") * **Minor failures**: Aiven automatically handles minor failures, such as service process crashes or temporary loss of network access, without any significant changes to the service deployment. In all plans, the service automatically restores regular operation by restarting crashed processes or restoring network access when available. * **Severe failures**: In case of severe hardware or software problems, such as losing an entire node, more drastic recovery measures are required. Aiven's monitoring infrastructure automatically detects a failing node when it reports problems with its self-diagnostics or stops communicating altogether. The monitoring infrastructure then schedules the creation of a new replacement node. note In case of database failover, your service's **Service URI** remains the same, only the IP address changes to point to the new primary node. ## High availability for business, premium, and custom plans[​](#high-availability-for-business-premium-and-custom-plans "Direct link to High availability for business, premium, and custom plans") If a standby Dragonfly node fails, the primary node continues running. The system prepares the replacement standby node and synchronizes it with the primary for normal operations to resume. In case the primary Dragonfly node fails, the standby node is evaluated for promotion to the new primary based on data from the Aiven monitoring infrastructure. Once promoted, this node starts serving clients, and a new node is scheduled to become the standby. However, during this transition, there may be a brief service interruption. If the primary and standby nodes fail simultaneously, new nodes are created automatically to replace them. However, this may lead to data loss as the primary node is restored from the latest backup. As a result, any database writes made since the last backup can be lost. note The duration for replacing a failed node depends mainly on the **cloud region** and the **amount of data** to be restored. For Business, Premium, and Custom plans with multiple nodes, this process is automatic and requires no administrator intervention, but service interruptions may occur during the recreation of nodes. ## Single-node startup service plans[​](#single-node-startup-service-plans "Direct link to Single-node startup service plans") Losing the only node in the service triggers an automatic process of creating a new replacement node. The new node then restores its state from the latest available backup and resumes serving customers. The service is unavailable for the duration of the restore operation. All the write operations made since the last backup are lost. --- # Get started with Aiven for Dragonfly® Get started with Aiven for Dragonfly by creating your service, integrating it with other services, and connecting to it with your preferred programming language. note Aiven for Dragonfly is fully supported by Aiven's service level agreements (SLAs), ensuring its capability to manage production workloads. As the latest addition to Aiven services, we recommend initiating a proof of concept (PoC) with Aiven for Dragonfly to thoroughly evaluate its capabilities and confirm its fit for your production requirements. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Console * CLI * Terraform - Access to the [Aiven Console](https://console.aiven.io) * [Aiven CLI](https://github.com/aiven/aiven-client#installation) installed * [A personal token](https://docs.aiven.io/docs/platform/howto/create_authentication_token.html) - [Terraform installed](https://www.terraform.io/downloads) - A [personal token](/docs/platform/howto/create_authentication_token.md) ## Create an Aiven for Dragonfly® service[​](#create-an-aiven-for-dragonfly-service "Direct link to Create an Aiven for Dragonfly® service") * Console * CLI * Terraform 1. In your project, click **Services**. 2. Click **Create service**. 3. Select a service type. 4. Select a **Cloud**. 5. Select a **Plan**. note The plans available can vary between cloud providers and regions for the same service. 6. In the **Service details**, enter a name for your service. 7. Optional: Add service tags. 8. In the **Service summary**, click **Create service**. The status of the service is **Rebuilding** during its creation. When the status is **Running**, you can start using the service. This typically takes a couple of minutes and can vary between cloud providers and regions. Use [Aiven CLI](/docs/tools/cli.md) to create your service: 1. Determine the service plan, cloud provider, and region to use for your Aiven for Dragonfly service. 2. Run the following command to create Aiven for Dragonfly service named dragonfly-demo: ``` avn service create dragonfly-demo \ --service-type dragonfly \ --cloud google-europe-north1 \ --plan startup-4 \ --project dev-sandbox ``` To see: * A full list of default flags, run `avn service create -h` * Type-specific options, run `avn service types -v` The following example files are also available in the [Aiven Terraform Provider repository](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/dragonfly) on GitHub. 1. Create a file named `provider.tf` and add the following: ``` Loading... ``` 2. Create a file named `service.tf` and add the following: ``` Loading... ``` 3. Create a file named `variables.tf` and add the following: ``` Loading... ``` 4. Create the `terraform.tfvars` file and add the values for your token and project name. To apply your Terraform configuration: 1. Initialize Terraform by running: ``` terraform init ``` The output is similar to the following: ``` Initializing the backend... Initializing provider plugins... - Finding aiven/aiven versions matching ">= 4.0.0, < 5.0.0"... - Installing aiven/aiven v4.9.2... - Installed aiven/aiven v4.9.2 ... Terraform has been successfully initialized! ... ``` 2. To create an execution plan and preview the changes, run: ``` terraform plan ``` 3. To deploy your changes, run: ``` terraform apply --auto-approve ``` ## Create service integrations[​](#create-service-integrations "Direct link to Create service integrations") Integrate Aiven for Dragonfly® with other Aiven services or third-party tools using the integration wizard available on the [Aiven Console](https://console.aiven.io/), [Aiven CLI](https://github.com/aiven/aiven-client), or [Aiven Terraform Provider](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration). Learn how to [create service integrations](/docs/platform/howto/create-service-integration.md). ## Connect to Aiven for Dragonfly[​](#connect-to-aiven-for-dragonfly "Direct link to Connect to Aiven for Dragonfly") Learn how to connect to Aiven for Dragonfly using different programming languages: * [redis-cli](/docs/products/dragonfly/howto/connect-redis-cli.md) * [Go](/docs/products/dragonfly/howto/connect-go.md) * [Node](/docs/products/dragonfly/howto/connect-node.md) * [Python](/docs/products/dragonfly/howto/connect-python.md) ## Migration limitations[​](#migration-limitations "Direct link to Migration limitations") Aiven for Dragonfly does not support automatic migration of users, access control lists (ACLs), or service configurations from Redis or Valkey. If you customized your Aiven for Caching or Aiven for Valkey service with specific settings, manually apply these configurations to Aiven for Dragonfly. Automatic transfer of custom configurations is unavailable due to differences in service configurations. ## Explore other resources[​](#explore-other-resources "Direct link to Explore other resources") * Learn about how Aiven for Dragonfly supports [high availability](/docs/products/dragonfly/concepts/ha-dragonfly.md). * Migrate data from [Aiven for Caching to Aiven for Dragonfly](/docs/products/dragonfly/howto/migrate-aiven-caching-df-console.md). * Migrate data from [external Dragonfly to Aiven for Dragonfly](/docs/products/dragonfly/howto/migrate-ext-redis-df-console.md). --- # RedisJSON v2 syntax compatibility Learn how to optimize your experience with RedisJSON in Aiven for Dragonfly® with the v2 JSONPath syntax using the $ root node. ## JSONPath syntax versions[​](#jsonpath-syntax-versions "Direct link to JSONPath syntax versions") Aiven for Dragonfly services use the v2 JSONPath syntax, ensuring compatibility with RedisJSON. This syntax designates the dollar sign (`$`) as the root node for JSON paths, moving away from the previously used dot (`.`) notation. This modification improves JSON command standardization for seamless integration between Aiven for Dragonfly and RedisJSON. For a comprehensive list of these JSON commands, see [Dragonfly documentation](https://www.dragonflydb.io/docs/category/json). ## Ensure compatibility with RedisJSON v2[​](#ensure-compatibility-with-redisjson-v2 "Direct link to Ensure compatibility with RedisJSON v2") * **Confirm library support:** Ensure your application uses libraries compatible with RedisJSON's v2 JSONPath syntax. This might require updating to the latest versions of these libraries. * **Adjust JSONPath expressions:** Update your JSONPath expressions to use the `$` root node. Convert dot notation paths (`.path.to.element`) to the v2 syntax (`$.path.to.element`). * **Testing:** Test your application after making these adjustments to confirm that interactions with RedisJSON operate as expected, particularly in areas that rely heavily on JSONPath expressions. Related pages * [RedisJSON documentation](https://redis.io/docs/data-types/json/path/) --- # Connect to Aiven for Dragonfly® with Go This example demonstrates how to connect to Dragonfly® using Go, using the `go-redis/redis` library, which is officially supported with Dragonfly. For more information, see [Dragonfly SDKs](https://www.dragonflydb.io/docs/development/sdks). ## Variables[​](#variables "Direct link to Variables") These are the placeholders you will need to replace in the code sample: | Variable | Description | | --------------- | --------------------------------- | | `DRAGONFLY_URI` | URL for the Dragonfly® connection | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") First, install the `go-redis/redis` library: ``` go get github.com/go-redis/redis/v8 ``` ## Code[​](#code "Direct link to Code") Create a file named `main.go` and add the following content, replacing the `DRAGONFLY_URI` placeholder with your Dragonfly instance's connection URI: ``` package main import ( "context" "fmt" "github.com/go-redis/redis/v8" ) var ctx = context.Background() func main() { dragonflyURI := "DRAGONFLY_URI" opts, err := redis.ParseURL(dragonflyURI) if err != nil { panic(err) } rdb := redis.NewClient(opts) err = rdb.Set(ctx, "key", "hello world", 0).Err() if err != nil { panic(err) } val, err := rdb.Get(ctx, "key").Result() if err != nil { panic(err) } fmt.Println("The value of key is:", val) } ``` This code connects to Dragonfly, sets a key named `key` with the value `hello world` (with no expiration), and retrieves and prints the value of this key. ## Run the code[​](#run-the-code "Direct link to Run the code") To run the code, use the following command in your terminal: ``` go run main.go ``` If everything is set up correctly, the output should be: ``` The value of key is: hello world ``` --- # Connect to Aiven for Dragonfly® with NodeJS This example demonstrates how to connect to Dragonfly® from NodeJS using the `ioredis` library, which is officially supported and compatible with Dragonfly. For more information, see [Dragonfly SDKs](https://www.dragonflydb.io/docs/development/sdks). ## Variables[​](#variables "Direct link to Variables") Replace the following placeholders in the code sample with the appropriate values: | Variable | Description | | --------------- | -------------------------------- | | `DRAGONFLY_URI` | URL for the Dragonfly connection | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Install the `ioredis` library: ``` npm install --save ioredis ``` ## Code[​](#code "Direct link to Code") Create a file named `index.js`, add the following content and replace the `DRAGONFLY_URI` placeholder with your Dragonfly instance's connection URI: ``` const Redis = require('ioredis'); const redis = new Redis('DRAGONFLY_URI'); // Replace with your Dragonfly URI redis.set('key', 'hello world').then(() => { return redis.get('key'); }).then((value) => { console.log('The value of key is:', value); process.exit(); }).catch((err) => { console.error('Error:', err); process.exit(1); }); ``` This code connects to Dragonfly, sets a key named `key` with the value `hello world` (without expiration), then retrieves and prints the value of this key. ## Run the code[​](#run-the-code "Direct link to Run the code") To execute the code, use the following command in your terminal: ``` node index.js ``` If everything is set up correctly, the output should be: ``` The value of key is: hello world ``` --- # Connect to Aiven for Dragonfly® with Python This example demonstrates how to connect to Dragonfly® using Python, using the `redis-py` library, which is officially supported by Dragonfly. For more information, see [Dragonfly SDKs](https://www.dragonflydb.io/docs/development/sdks). ## Variables[​](#variables "Direct link to Variables") Replace the following placeholders in the code sample with the appropriate values: | Variable | Description | | --------------- | -------------------------------- | | `DRAGONFLY_URI` | URL for the Dragonfly connection | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Install the `redis-py` library: ``` pip install redis ``` ## Code[​](#code "Direct link to Code") Create a file named `main.py`, add the following content and replace the `DRAGONFLY_URI` placeholder with your Dragonfly instance's connection URI: ``` import redis # Replace with your Dragonfly URI r = redis.Redis.from_url('DRAGONFLY_URI') r.set('key', 'hello world') value = r.get('key') print(f'The value of key is: {value.decode()}') ``` This code connects to Dragonfly, sets a key named `key` with the value `hello world` (without expiration), and retrieves and prints the value of this key. ## Run the code[​](#run-the-code "Direct link to Run the code") To execute the code, use the following command in your terminal: ``` python main.py ``` note On some systems, you may need to use `python3` instead of `python` to invoke `Python 3`. If everything is set up correctly, the output should be: ``` The value of key is: hello world ``` --- # Connect to Aiven for Dragonfly® with redis-cli This example demonstrates how to connect to Dragonfly® using `redis-cli`, which supports nearly all the same commands as it does for Redis®. For more information, see [Dragonfly CLI](https://www.dragonflydb.io/docs/development/cli). ## Variables[​](#variables "Direct link to Variables") Replace the following placeholders in the code sample with the appropriate values: | Variable | Description | | --------------- | -------------------------------- | | `DRAGONFLY_URI` | URL for the Dragonfly connection | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Ensure the following before proceeding: 1. The `redis-cli` client installed. This can be installed as part of the Redis®\* server installation or as a standalone client. Refer to the [Redis Installation Guide](https://redis.io/docs/getting-started/tutorial/) for more information. ## Code[​](#code "Direct link to Code") To connect to Dragonfly, execute the following command in a terminal window: ``` redis-cli -u DRAGONFLY_URI ``` This command connects you to your Dragonfly instance. To verify the connection is successful, use the `INFO` command: ``` INFO ``` This command should return various Dragonfly parameters similar to Redis: ``` # Server dragonfly_version:1.0.0 dragonfly_git_sha1:0a1b2c3d dragonfly_mode:standalone ... ``` To set a key, use the following command: ``` SET mykey mykeyvalue123 ``` This should return a confirmation `OK`. To retrieve the set key value, use: ``` GET mykey ``` The output will be the value of the key, in this case, `"mykeyvalue123"`. --- # Data eviction policy in Aiven for Dragonfly Aiven for Dragonfly® optimizes cache memory management with a low-overhead data eviction policy. ## Understand data eviction policy[​](#understand-data-eviction-policy "Direct link to Understand data eviction policy") Aiven for Dragonfly® uses a variation of the 2Q key eviction algorithm, which intelligently manages cache by deciding which data to retain or evict based on usage patterns. This system uses two distinct buffers to streamline cache memory management: * **Protected buffer:** Stores frequently accessed items, indicating their higher value for caching due to repeated use. * **Probationary buffer:** Temporarily stores new or less frequently accessed items. These items move to the protected buffer upon repeated access, indicating their growing significance. This dual-buffer approach ensures resilience against fluctuating access patterns, maintaining cache efficiency without additional memory overhead per item. It's a key factor in Dragonfly's ability to keep valuable data accessible while managing less critical data efficiently. ## Enable `cache_mode` via Aiven Console[​](#enable-cache_mode-via-aiven-console "Direct link to enable-cache_mode-via-aiven-console") By default, `cache_mode `in Aiven for Dragonfly is disabled, which might result in out-of-memory errors when the `maxmemory` limit is reached. To prevent these errors, enable `cache_mode` in the advanced settings. 1. Log in to the [Aiven Console](https://console.aiven.io/), select your project, and select your Aiven for Dragonfly service. 2. Select **Service settings** from the sidebar. 3. On the **Service settings** page, scroll to the **Advanced configuration** section. 4. Click **Configure**. 5. In the **Advanced configuration dialog**, set the `cache_mode` toggle to **Enabled**. 6. Click **Save configuration**. --- # Connect to Aiven for Dragonfly® Connect to the Aiven for Dragonfly® service using various programming languages or tools. ## [redis-cli](/docs/products/dragonfly/howto/connect-redis-cli.md) [This example demonstrates how to connect to Dragonfly® using](/docs/products/dragonfly/howto/connect-redis-cli.md) ## [Go](/docs/products/dragonfly/howto/connect-go.md) [This example demonstrates how to connect to Dragonfly® using Go, using](/docs/products/dragonfly/howto/connect-go.md) ## [NodeJS](/docs/products/dragonfly/howto/connect-node.md) [This example demonstrates how to connect to Dragonfly® from NodeJS using](/docs/products/dragonfly/howto/connect-node.md) ## [Python](/docs/products/dragonfly/howto/connect-python.md) [This example demonstrates how to connect to Dragonfly® using Python,](/docs/products/dragonfly/howto/connect-python.md) --- # Migrate from Aiven for Caching or Aiven for Valkey™ to Aiven for Dragonfly Migrate your Aiven for Caching or Aiven for Valkey databases to Aiven for Dragonfly using the Aiven Console migration tool. note Aiven for Dragonfly is fully supported by Aiven's service level agreements (SLAs), ensuring its capability to manage production workloads. As the latest addition to Aiven services, we recommend initiating a proof of concept (PoC) with Aiven for Dragonfly to thoroughly evaluate its capabilities and confirm its fit for your production requirements. ## Compatibility check[​](#compatibility-check "Direct link to Compatibility check") Before migrating an Aiven for Caching or Aiven for Valkey database to Aiven for Dragonfly, review your current database setup. * **Review database setup:** Examine the data structures, storage patterns, and configurations in your Aiven for Caching or Aiven for Valkey database. Identify any unique features, custom settings, and specific configurations. * **API compatibility:** While Dragonfly closely mirrors Redis API commands, some differences exist, especially with newer versions of Redis and Valkey. For information on command compatibility, refer to the [Dragonfly API compatibility documentation](https://www.dragonflydb.io/docs/command-reference/compatibility). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting the migration: * Confirm a target Aiven for Dragonfly service set up and ready. For more information, see [Get started with Aiven for Dragonfly®](/docs/products/dragonfly/get-started.md). * Confirm that your Aiven for Caching or Aiven for Valkey service is accessible over the Internet. For more information, see [Public internet access](/docs/platform/howto/public-access-in-vpc.md). * Note the project and service names for migration in the Aiven Console. The Aiven Console migration tool automatically uses connection details like the hostname, port, and credentials associated with the selected service. ## Migration limitations[​](#migration-limitations "Direct link to Migration limitations") Aiven for Dragonfly does not support automatic migration of users, access control lists (ACLs), or service configurations from Redis or Valkey. If you customized your Aiven for Caching or Aiven for Valkey service with specific settings, manually apply these configurations to Aiven for Dragonfly. Automatic transfer of custom configurations is unavailable due to differences in service configurations. ## Database migration steps[​](#database-migration-steps "Direct link to Database migration steps") 1. Log in to the [Aiven Console](https://console.aiven.io/) and select the Aiven for Dragonfly service for your database migration. 2. On the **Overview** page, click **Service settings** in the sidebar. 3. In the **Service management** section, click **Actions** > **Migrate database**. 4. Follow the wizard to guide you through the database migration process. ### Step 1: Configure[​](#step-1-configure "Direct link to Step 1: Configure") To begin the migration, select **Migrate existing Aiven for Caching/Valkey database**: 1. In the **Project name** drop-down, select the project with the service to migrate. 2. In the **Service name** drop-down, choose the Aiven for Caching or Aiven for Valkey service. 3. Click **Get started** to proceed with the migration. ### Step 2: Validate[​](#step-2-validate "Direct link to Step 2: Validate") The [Aiven Console](https://console.aiven.io/) automatically validates the database configuration for the selected service. Click **Run validation** to check the connection. note * If a validation error occurs during migration, follow the on-screen instructions to fix it. Rerun validation to ensure the database meets migration criteria. * The migration doesn't include service user accounts and commands that are in progress. ### Step 3: Migrate[​](#step-3-migrate "Direct link to Step 3: Migrate") After completing validation, click **Start migration** to begin migrating data to Aiven for Dragonfly. While the migration is in progress: * Click **Close window** to close the migration wizard, and return later to monitor the migration status from the service **Overview** page. * The migration duration depends on the size of your database. During migration, the target database is read-only, and writing to the database is only possible after stopping the migration. * Certain managed database features are disabled during migration. * To stop the migration, click **Stop migration**. Any data already transferred to Aiven for Dragonfly is preserved. note To avoid conflicts during migration: * Avoid actions that may disrupt replication, such as changing replication settings, modifying firewall rules, or altering trusted sources. * Stopping migration halts replication immediately, but any transferred data is preserved. Starting a new migration overwrites the entire database with the latest data from the source. tip If the migration fails, investigate, and resolve the issue. Click **Start over** in the Data migration window to restart the migration. ### Step 5: Close the connection and next steps[​](#step-5-close-the-connection-and-next-steps "Direct link to Step 5: Close the connection and next steps") Once migration completes: * Click **Stop replication** if no further synchronization is required and you are ready to switch to Aiven for Dragonfly after thoroughly testing the service. * Click **Keep replicating** to maintain ongoing data synchronization if further testing or syncing with the source database is needed. warning Avoid system updates or configuration changes during active replication, as these can restart nodes and trigger a new database migration. Ensure replication is complete or stopped before making any modifications. note When replication mode is active, Aiven for Dragonfly continuously synchronizes new writes from the source database, ensuring your data remains up to date. Related pages * [Aiven for Dragonfly overview](/docs/products/dragonfly.md) * [Aiven for Valkey™ overview](/docs/products/valkey.md) --- # Migrate from external Redis®\* or Valkey to Aiven for Dragonfly Migrate external Redis® or Valkey databases to Aiven for Dragonfly® using the Aiven Console migration tool. important The migration of databases from Google Cloud Memorystore for Redis is not supported at this time. note Aiven for Dragonfly is fully supported by Aiven's service level agreements (SLAs), ensuring its capability to manage production workloads. As the latest addition to Aiven services, we recommend initiating a proof of concept (PoC) with Aiven for Dragonfly to thoroughly evaluate its capabilities and confirm its fit for your production requirements. ## Compatibility check[​](#compatibility-check "Direct link to Compatibility check") Before migrating an external Redis or Valkey database to Aiven for Dragonfly, review your current database setup. * **Review database setup:** Examine the data structures, storage patterns, and configurations in your Redis or Valkey database. Identify any unique features, custom settings, or specific configurations. * **API compatibility:** While Dragonfly closely mirrors Redis API commands, some differences exist, especially with newer versions of Redis and Valkey. For information on command compatibility, refer to the [Dragonfly API compatibility documentation](https://www.dragonflydb.io/docs/command-reference/compatibility). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting the migration process, ensure the following: * Confirm a target Aiven for Dragonfly service set up and ready. For more information, see [Get started with Aiven for Dragonfly®](/docs/products/dragonfly/get-started.md). * Source database information: * **Hostname or connection string:** The public hostname, connection string, or IP address used to connect to the database, [accessible from the public Internet](/docs/platform/howto/public-access-in-vpc.md). * **Port:** The port number used for connecting to the database. * **Username:** The username with appropriate permissions to access the data for migration. * **Password:** The password associated with the username. * Firewall rules allow traffic between databases or temporarily disabled firewalls. * An SSL-secured connection is recommended for data transfer during the source Redis database migration. * If the source Redis service is not publicly accessible, establish a VPC peering connection between the private networks. You need the VPC ID and cloud name for the migration. note Instances such as AWS ElastiCache for Redis that do not have public IP addresses will require a VPC and peering connection to establish a migration. ## Migration limitations[​](#migration-limitations "Direct link to Migration limitations") Aiven for Dragonfly does not support automatic migration of users, access control lists (ACLs), or service configurations from Redis or Valkey. If you customized your Aiven for Caching or Aiven for Valkey service with specific settings, manually apply these configurations to Aiven for Dragonfly. Automatic transfer of custom configurations is unavailable due to differences in service configurations. ## Database migration steps[​](#database-migration-steps "Direct link to Database migration steps") To migrate a Redis or Valkey database to Aiven for Dragonfly: 1. Log in to the [Aiven Console](https://console.aiven.io/) and select the Aiven for Dragonfly service for your database migration. 2. On the **Overview** page, click **Service settings** in the sidebar. 3. In the **Service management** section, click **Actions** > **Migrate database**. 4. Follow the wizard to guide you through the database migration process. ### Step 1: Configure[​](#step-1-configure "Direct link to Step 1: Configure") To begin the migration: 1. Select **Migrate an external Redis or Valkey database**. 2. Click **Get started** to begin. ### Step 2: Validation[​](#step-2-validation "Direct link to Step 2: Validation") Enter the required connection details for your source Redis or Valkey database: * **Hostname:** The public hostname, connection string, or IP address for the database connection. * **Port:** The port number used for connections. * **Username:** The username required to access your database. * **Password:** The password for database access. * Select the SSL encryption option for a secure migration and click **Run check** to verify the connection. note Resolve any issues to ensure a smooth migration. Migration does not include all components of your Redis setup, such as user accounts, ACLs, specific settings, and ongoing commands or scripts. It does transfer all database data and contents. ### Step 3: Migration[​](#step-3-migration "Direct link to Step 3: Migration") After completing validation, click **Start migration** to begin migrating data to Aiven for Dragonfly. While the migration is in progress: * Click **Close window** to close the migration wizard, and return later to monitor the migration status from the service **Overview** page. * The migration duration depends on the size of your database. During migration, the target database is read-only, and writing to the database is only possible after stopping the migration. * Certain managed database features are disabled during migration. * To stop the migration, click **Stop migration**. Any data already transferred to Aiven for Dragonfly is preserved. note To avoid conflicts during migration: * Avoid actions that may disrupt replication, such as changing replication settings, modifying firewall rules, or altering trusted sources. * Stopping migration halts replication immediately, but any transferred data is preserved. Starting a new migration overwrites the entire database with the latest data from the source. ### Step 4: Close the connection and next steps[​](#step-4-close-the-connection-and-next-steps "Direct link to Step 4: Close the connection and next steps") Once migration completes: * Click **Stop replication** if no further synchronization is required and you are ready to switch to Aiven for Dragonfly after thoroughly testing the service. * Click **Keep replicating** to maintain ongoing data synchronization if further testing or syncing with the source database is needed. warning Avoid system updates or configuration changes during active replication, as these can restart nodes and trigger a new database migration. Ensure replication is complete or stopped before making any modifications. note When replication mode is active, Aiven for Dragonfly continuously synchronizes new writes from the source database, ensuring your data remains up to date. Related pages * [Aiven for Dragonfly overview](/docs/products/dragonfly.md) --- # Advanced parameters for Aiven for Dragonfly® See the configuration options available for Aiven for Dragonfly®: | Parameter | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | []()[**ip\_filter**](#ip_filter)`array`- default: `0.0.0.0/0,::/0`IP filterAllow incoming connections from CIDR address block, e.g. '10.20.0.0/16' | | []()[**service\_log**](#service_log)`boolean,null`Service loggingStore logs for the service so that they are available in the HTTP API and console. | | []()[**static\_ips**](#static_ips)`boolean`Use static public IP addresses | | []()[**private\_access**](#private_access)`object`Allow access to selected service ports from private networksprivate\_access.prometheus boolean Allow clients to connect to prometheus with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.dragonfly boolean Allow clients to connect to dragonfly with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | []()[**privatelink\_access**](#privatelink_access)`object`Allow access to selected service components through Privatelinkprivatelink\_access.prometheus boolean Enable prometheus privatelink\_access.dragonfly boolean Enable dragonfly | | []()[**public\_access**](#public_access)`object`Allow access to selected service ports from the public Internetpublic\_access.prometheus boolean Allow clients to connect to prometheus from the public internet for service nodes that are in a project VPC or another type of private network public\_access.dragonfly boolean Allow clients to connect to dragonfly from the public internet for service nodes that are in a project VPC or another type of private network | | []()[**migration**](#migration)`object,null`Migrate data from existing servermigration.host string Hostname or IP address of the server where to migrate data from migration.port integer - min: 1 - max: 65535 Port number of the server where to migrate data from migration.password string Password for authentication with the server where to migrate data from migration.ssl boolean - default: true The server where to migrate data from is secured with SSL migration.username string User name for authentication with the server where to migrate data from migration.dbname string Database name for bootstrapping the initial connection migration.ignore\_dbs string Comma-separated list of databases, which should be ignored during migration (supported by MySQL and PostgreSQL only at the moment) migration.ignore\_roles string Comma-separated list of database roles, which should be ignored during migration (supported by PostgreSQL only at the moment) migration.method string The migration method to be used (currently supported only by Redis, Dragonfly, MySQL and PostgreSQL service types) | | []()[**recovery\_basebackup\_name**](#recovery_basebackup_name)`string`Name of the basebackup to restore in forked service | | []()[**dragonfly\_ssl**](#dragonfly_ssl)`boolean`- default: `true`Require SSL to access Dragonfly | | []()[**cache\_mode**](#cache_mode)`boolean`Evict entries when getting close to maxmemory limit | | []()[**service\_to\_fork\_from**](#service_to_fork_from)`string,null`Name of another service to fork from. This has effect only when a new service is being created. | | []()[**project\_to\_fork\_from**](#project_to_fork_from)`string,null`Name of another project to fork a service from. This has effect only when a new service is being created. | | []()[**dragonfly\_persistence**](#dragonfly_persistence)`string`When persistence is 'rdb' or 'dfs', Dragonfly does RDB or DFS dumps every 10 minutes. Dumps are done according to the backup schedule for backup purposes. When persistence is 'off', no RDB/DFS dumps or backups are done, so data can be lost at any moment if the service is restarted for any reason, or if the service is powered off. Also, the service can't be forked. | --- # Aiven for Dragonfly® version lifecycle Learn how Aiven manages the Aiven for Dragonfly® service version, end of life (EOL) dates, and what happens to your service after the version reaches EOL. Aiven for Dragonfly identifies versions in `major.minor.patch` format, for example `1.39.0`. ## Service version management[​](#service-version-management "Direct link to Service version management") Aiven manages the software version of single-versioned services for you. You don't select or change the major version yourself, and only one version is available at a time. The version your service is currently running is visible in the [Aiven Console](https://console.aiven.io/). ## Version updates[​](#version-updates "Direct link to Version updates") Aiven updates the version as part of regular platform maintenance and rolls it out during your service's maintenance window. Because you don't select a version, Aiven doesn't send the EOL email notifications and reminders described for [multi-versioned services](/docs/platform/reference/eol-for-major-versions.md#service-version-eol-policy). ## After a version reaches end of life[​](#after-a-version-reaches-end-of-life "Direct link to After a version reaches end of life") If Aiven sets an EOL date for the version your service is running: * If the service is powered on, it's automatically upgraded to the next supported version. * If the service is powered off, it's deleted. For the EOL dates of a specific single-versioned service, see the [Aiven single-versioned services EOL](/docs/platform/reference/eol-for-major-versions.md#aiven-single-versioned-services-eol) reference. ## If the entire service reaches end of life[​](#if-the-entire-service-reaches-end-of-life "Direct link to If the entire service reaches end of life") In some cases, Aiven retires an entire service rather than upgrading it to a new version. If this happens, Aiven announces the retirement in advance and provides migration guidance. For the list of retired and soon-to-be-retired services, see [End of life for Aiven services](/docs/platform/reference/end-of-life.md). ## Version EOL dates[​](#version-eol-dates "Direct link to Version EOL dates") | Version | Aiven EOL | Service creation supported until | | ------- | ---------- | -------------------------------- | | 1.39.0 | 2026-09-30 | 2026-06-17 | Related pages * [End of life for Aiven services](/docs/platform/reference/end-of-life.md) --- # Aiven for Apache Flink® Support for existing services Existing services remain fully supported. If you need assistance, open a [support ticket](/docs/platform/howto/support.md#create-a-support-ticket) or contact [Aiven Support](mailto:support@aiven.io). For new stream-processing workloads, contact your Aiven account manager to discuss available options. Aiven for Apache Flink® is a fully managed service that leverages the power of the [open-source Apache Flink framework](https://flink.apache.org/) to provide distributed, stateful stream processing capabilities, allowing users to perform real-time computation with SQL efficiently. ## Main features[​](#main-features "Direct link to Main features") ### Flink SQL[​](#flink-sql "Direct link to Flink SQL") Apache Flink allows you to develop streaming applications using standard SQL. Aiven for Apache Flink is a fully managed service that provides various features for developing and running streaming applications using Flink on the Aiven platform. One of these features is the SQL editor, which is a built-in feature of the Aiven Console. The built-in SQL editor allows you to create and test Flink SQL queries, explore the table schema of your data streams and tables, and deploy queries to your streaming application. This makes it easy to develop and maintain streaming applications using Flink SQL on the Aiven platform. ### Build applications to process data[​](#build-applications-to-process-data "Direct link to Build applications to process data") An [Aiven for Apache Flink® Application](/docs/products/flink/concepts/flink-applications.md) is an abstraction layer on top of Apache Flink SQL that includes all the elements related to a Flink job to help build your data processing pipeline. It contains all the components related to a Flink job, including the definition of source and sink tables, data processing logic, deployment parameters, and other relevant metadata. Applications are the starting point for running an Apache Flink job within the Aiven managed service. The [Aiven Console](https://console.aiven.io/) provides a user-friendly, guided wizard to help you build and deploy applications, create source and sink tables, write transformation statements, and validate and ingest data using the interactive query feature. ### Interactive queries[​](#interactive-queries "Direct link to Interactive queries") The [interactive query](/docs/products/flink/concepts/supported-syntax-sql-editor.md) feature in Aiven for Apache Flink grants the ability to preview the data of a Flink table or job without outputting the rows to a sink table like Apache Kafka®. This can be useful for testing and debugging purposes, as it allows you to examine the data being processed by your [Flink application](/docs/products/flink/concepts/flink-applications.md). ### Built-in data flow integration with Aiven for Apache Kafka®[​](#built-in-data-flow-integration-with-aiven-for-apache-kafka "Direct link to Built-in data flow integration with Aiven for Apache Kafka®") Aiven for Apache Flink provides built-in data flow integration with [Aiven for Apache Kafka®](/docs/products/kafka.md), allowing you to connect your Flink streaming applications with Apache Kafka as a source or sink for your data. * When you create data tables in Aiven for Apache Flink, the service provides auto-completion for finding existing topics in a connected Kafka service. * You can also choose the table format when reading data from Kafka, including JSON, Apache Avro, Confluent Avro, and Debezium CDC. * Aiven for Apache Flink also supports upsert Kafka connectors, which allow you to produce a changelog stream where each data record represents an update or delete an event. This can be useful for maintaining data consistency in real-time streaming applications that involve complex data transformations or updates. ### Built-in data flow integration with Aiven for PostgreSQL®[​](#built-in-data-flow-integration-with-aiven-for-postgresql "Direct link to Built-in data flow integration with Aiven for PostgreSQL®") Aiven for Apache Flink provides built-in data flow integration with [Aiven for PostgreSQL](/docs/products/postgresql.md), allowing you to connect your Flink streaming applications with PostgreSQL as a source or sink for your data. When you create data tables in Aiven for Apache Flink, the service provides auto-completion for finding existing databases in a connected PostgreSQL service. This makes it easy to select the appropriate database and table when configuring your Flink streaming application to read or write data from PostgreSQL. ### Automate workflows with Terraform[​](#automate-workflows-with-terraform "Direct link to Automate workflows with Terraform") Aiven for Apache Flink provides integration with the [Aiven Terraform Provider](/docs/tools/terraform.md), which allows you to automate workflows for managing Flink services on the Aiven platform. To use the Aiven Terraform Provider to automate workflows for managing Flink services, you can reference the Flink data source in your Terraform configuration files. ### Disaster recovery[​](#disaster-recovery "Direct link to Disaster recovery") Periodic checkpoints have been configured to be persisted externally in object storage. They allow Flink to recover states and positions in the streams by giving the application the same semantics as a failure-free execution. See [Checkpoints](/docs/products/flink/concepts/checkpoints.md). ## Cluster management[​](#cluster-management "Direct link to Cluster management") ### Cluster deployment mode[​](#cluster-deployment-mode "Direct link to Cluster deployment mode") Aiven for Apache Flink® is configured to use the [HashMap state backend](https://ci.apache.org/projects/flink/flink-docs-stable/api/java/org/apache/flink/runtime/state/hashmap/HashMapStateBackend.html). This means that the [state](https://nightlies.apache.org/flink/flink-docs-stable/docs/concepts/stateful-stream-processing/#what-is-state) is stored in memory, which can impact the performance of jobs that require keeping a very large state. We recommend you provision your platform accordingly. The Flink cluster executes applications in [session mode](https://nightlies.apache.org/flink/flink-docs-stable/docs/deployment/overview/#session-mode) so you can deploy multiple Flink jobs on the same cluster and maximize resource use. ### Cluster scaling[​](#cluster-scaling "Direct link to Cluster scaling") Each node is equipped with a TaskManager and JobManager. We recommend scaling up your cluster to add more CPU and memory for the TaskManager before attempting to scale out, so you make the best use of the resources with a minimum number of nodes. By default, each TaskManager is configured with a single slot for maximum job isolation. It is highly recommended that you modify this option to match your requirements. warning Adjusting the task slots per TaskManager requires a cluster restart. ### Cluster restart strategy[​](#cluster-restart-strategy "Direct link to Cluster restart strategy") The default restart strategy of the cluster is set to `Failure Rate`. This controls how Apache Flink restarts in case of failures during job execution. Administrators can change this setting in the advanced configuration options of the service. For more information on available options, refer to [Apache Flink fault tolerance](https://nightlies.apache.org/flink/flink-docs-master/docs/deployment/config/#fault-tolerance) documentation. ### Cluster logging, metrics, and alerting[​](#cluster-logging-metrics-and-alerting "Direct link to Cluster logging, metrics, and alerting") Log and metrics integration to Aiven services are available for administrators to configure so you can monitor the health of your service. By enabling these integrations, you can: * [Push service logs into an index in Aiven for OpenSearch®](/docs/products/opensearch/howto/opensearch-log-integration.md). * Push service metrics to [Aiven for Metrics](/docs/products/metrics.md) or [PostgreSQL®](/docs/products/postgresql.md) services on Aiven. * Create custom OpenSearch or [Grafana®](/docs/products/grafana.md) dashboards to monitor the service. ### Cluster security considerations[​](#cluster-security-considerations "Direct link to Cluster security considerations") All services run on a dedicated virtual machine with end-to-end encryption, and all nodes are firewall-protected. The credentials used for data flow integrations between Flink and other Aiven services have read/write permissions on the clusters. You can set up separate clusters for writing processed data from Flink and restrict access if you need more strict access management than our default setup offers. This also minimizes the risk of accidental write events to the source cluster. Related pages * [Aiven.io](https://aiven.io/flink) --- # Checkpoints Checkpoints in Aiven for Apache Flink® are a key feature for ensuring resiliency and fault tolerance in stateful functions. By periodically creating snapshots of the data stream and storing them, checkpoints enable Apache Flink to recover the stream's state and position in the event of a failure, ensuring that applications can continue to execute without interruption. ## Recover from failures[​](#recover-from-failures "Direct link to Recover from failures") In the event of a failure, Aiven for Apache Flink uses these checkpoints to restore the application's state and resume processing from the last recorded reading position, allowing the application to continue as if the failure had never occurred. ## Efficient data handling[​](#efficient-data-handling "Direct link to Efficient data handling") Unlike traditional backups that create full data copies, checkpoints act more like recovery logs. They periodically snapshot the data stream and store these snapshots, making data handling more efficient. ## Visualize checkpoints[​](#visualize-checkpoints "Direct link to Visualize checkpoints") Related pages * [Apache Flink® documentation on checkpoints](https://nightlies.apache.org/flink/flink-docs-release-1.19/docs/ops/state/checkpoints/) --- # Custom JARs in Aiven for Apache Flink® [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Aiven for Apache Flink enables you to upload, deploy, and manage your own Java code as custom JARs within a [JAR application](/docs/products/flink/howto/create-jar-application.md). This feature expands the capabilities of Aiven for Apache Flink, allowing you to add custom capabilities that extend beyond the default SQL features. With custom JARs, you can swiftly develop and maximize the potential of your Aiven for Flink application. ## What are custom JARs[​](#what-are-custom-jars "Direct link to What are custom JARs") Custom JARs are specialized Java Archive files containing code and resources for functionalities beyond standard Java or Apache Flink libraries. The capabilities of any custom JAR are defined by your organization and use cases, allowing for tailored solutions that meet specific technical requirements and objectives. ## Why use custom JARs[​](#why-use-custom-jars "Direct link to Why use custom JARs") Using Custom JARs in Aiven for Apache Flink offers several key benefits: * **Enhanced functionality:** Using custom JARs, you can extend the native capabilities of your Aiven for Apache Flink service, incorporating functionalities that go beyond the core Flink APIs. * **Code reuse:** This feature allows you to seamlessly reuse and integrate your existing Java code into your Aiven for Apache Flink applications. * **Simplified management:** Aiven handles the complexities of hosting and operating a Flink service. This streamlined approach requires you only to upload and deploy the JAR file, enabling direct use of its functionalities within the Aiven for Apache Flink application. ## Use cases[​](#use-cases "Direct link to Use cases") Custom JARs can be applied in various scenarios, including but not limited to: * **Custom data processing and enrichment:** Custom JARs facilitate the processing and enrichment of data from sources not natively supported by Flink's core APIs. This includes implementing unique data processing logic and applying custom functions for data transformation and enrichment. Related pages * [How to use custom JARs in Aiven for Apache Flink application](/docs/products/flink/howto/create-jar-application.md). --- # Event and processing times Event time refers to when events occur, and processing time is when a system observes or processes these events. Understanding the difference between these two is essential for data processing and streaming. It affects data handling, analysis, and storage. ## Factors affecting processing[​](#factors-affecting-processing "Direct link to Factors affecting processing") Several factors can cause differences between event time and processing time, including: * Shared hardware resources can lead to variable processing capabilities. * Changes in network traffic can delay data transmission, affecting when data is processed. * The complexities of distributed systems can introduce delays or reorder data. * Variations in the volume of data and the order in which it arrives can further complicate processing times. Ideally, event time and processing time would be the same, but this is rarely the case in practice. ## Use event time[​](#use-event-time "Direct link to Use event time") In Apache Flink®, data might not always arrive in the order in which the events occurred. Relying on processing time can lead to inaccuracies and issues in system behavior. To mitigate these issues, Aiven recommends using **event time** when processing data. This approach maintains the correct sequence of events throughout the data streaming pipeline. Additionally, event time allows for consistent data reprocessing, improving reliability and system resilience. Related pages * [Apache Flink® event time and processing time](https://nightlies.apache.org/flink/flink-docs-release-1.19/docs/concepts/time/) --- # Aiven for Apache Flink® applications An Aiven for Apache Flink® Application is an abstraction layer that simplifies building data processing pipelines in Apache Flink. It supports two types of applications: * **SQL applications:** Built on top of Apache Flink SQL, these applications include elements like source and sink table definitions, data processing logic using SQL, deployment parameters, and other necessary metadata. * **JAR applications:** This type allows users to upload and deploy custom JAR files, enabling the integration of custom code and resources for specialized functionalities beyond the standard Flink capabilities. [Applications](/docs/products/flink/howto/create-flink-applications.md) are the starting point for running an Apache Flink job within the Aiven managed service. The [Aiven Console](https://console.aiven.io/) provides a user-friendly, guided wizard to help you build and deploy applications, create source and sink tables, write transformation statements, and validate and ingest data using the interactive query feature. Each application whether SQL or JAR-based, is designed to perform specific data transformation tasks, allowing you to map and manage different workflows within the Aiven for Apache Flink service. For example, you can create a SQL application to process user IDs from one topic and publish them on another, or use a custom JAR application for more specialized data processing needs. Applications significantly improve the developer experience and simplify the development and deployment of Flink applications. Applications are **automatically versioned** on every edit of the underline definition (tables, transformation SQL, or custom JARs), allowing you to experiment with new transformations and revert to a previously stored definition if the result of the edits doesn't meet expectations. For information on how to create applications, see [Create an Aiven for Apache Flink® application](/docs/products/flink/howto/create-flink-applications.md) ## Application features[​](#application-features "Direct link to Application features") Some of key feature of Aiven for Flink applications include: * **Intuitive interface**: The [Aiven Console](https://console.aiven.io/) provides a user-friendly, guided wizard for building and deploying applications. * **Source and sink tables** definition with **SQL autocomplete**: For SQL applications, create source and sink tables with guidance from SQL autocomplete based on connector types. * **Interactive query**: For SQL applications, the interactive query feature allows you to validate and preview your data in the table or SQL transformation before deploying it. * **Start and stop**: You can start and stop applications stop and start anytime. * **Versioning**: Creating a new application version after editing the table or data transformation definition allows you to track changes and revert to a previous version when needed. * **Savepoints**: You can stop your application with a savepoint, allowing you to restart it from the previous state at a later time. * **Scalability:** You can parallelize the workload into multiple tasks and executing them concurrently in a cluster ## Limitations[​](#limitations "Direct link to Limitations") * **Concurrent applications:** You can create and run up to 4 concurrent applications in the Aiven for Apache Flink service. * **Savepoints:** You can store up to 10 savepoints history per application deployment, allowing you to start from a previous state if necessary. If you exceed this limit, you must clear the savepoints history before creating a new application deployment. ## Application status[​](#application-status "Direct link to Application status") Flink applications can have different statuses based on their execution and savepoint state. Below are the most common statuses you may encounter: * **CANCELED**: Your application has stopped without creating a savepoint. * **CREATED**: Your application has been created but has yet to start running. * **FAILED**: Your application attempted to run but was unsuccessful. * **FINISHED**: Your application has completed its execution and stopped, with a savepoint created. * **INITIALIZING**: Your application is in the process of being initialized and preparing to run. * **RUNNING**: Your application is actively running. You can view its current version. * **RESTARTING**: Your application is being restarted. * **SAVING\_AND\_STOP**: Your application is currently saving its current state in a savepoint and stopping. Other statuses and transient statuses include: * **CANCELLING\_REQUESTED**: A stop request for your application has been made without creating a savepoint. * **CANCELLING**: Your application is currently being stopped. * **DELETE\_REQUESTED**: A request to delete your application has been initiated. * **DELETING**: Your application is in the process of being removed. * **FAILING**: Your application is encountering issues and failing. * **RESTARTING**: Your application is undergoing a restart process. * **SAVING**: Your application is creating a savepoint. * **SAVING\_AND\_STOP\_REQUESTED**: A request has been made to save the current state of your application in a savepoint and stop it. * **SUSPENDED**: Your application has been suspended. --- # Aiven for Apache Flink® architecture At a high level, Flink has a runtime architecture consisting of two types of processes: a **JobManager** and one or more **TaskManager**. ## JobManager[​](#jobmanager "Direct link to JobManager") The JobManager is the central coordination point of Flink and is responsible for managing the execution of Flink jobs. It is responsible for scheduling tasks, managing task execution, and coordinating the overall execution of the Flink application. In other words, Flink provides an exactly once processing guarantee, only if JobManager is always up and running. Some responsibilities of the JobManager include: * **Scheduling tasks:** The JobManager decides when to schedule the next task (or set of tasks) for execution based on the availability of resources and the dependencies between tasks. * **Monitoring task execution:** The JobManager monitors the execution of tasks and responds to finished tasks or execution failures. * **Coordinating checkpoints:** The JobManager coordinates the execution of checkpoints, which are periodic snapshots of the state of the Flink application. Checkpoints are used to enable recovery of the Flink application in the event of a failure. * **Coordinating recovery on failures:** In the event of a failure, the JobManager coordinates recovery by re-executing failed tasks or rolling back to a previous checkpoint. In a high-availability setup, there may be multiple JobManagers running in the cluster, with one JobManager designated as the leader and the others as standby JobManagers. The JobManager in Apache Flink consists of three main components: * **ResourceManager:** The ResourceManager is responsible for managing the allocation and deallocation of resources in the Flink cluster. Additionally, ResourceManger is responsible for managing **Task slots** - the unit of resource scheduling in a Flink cluster. * **Dispatcher:** The Dispatcher in Apache Flink ensures tasks run smoothly on the cluster by scheduling them on available task slots and ensuring that tasks are executed efficiently. * **JobMaster:** The JobMaster in Apache Flink makes sure a specific job runs smoothly on the cluster by coordinating the tasks and executing them correctly and efficiently. ## TaskManager[​](#taskmanager "Direct link to TaskManager") TaskManager is responsible for executing the tasks assigned to them by the JobManager and exchanging data with other TaskManagers as needed. This direct communication between TaskManagers allows for efficient data exchange and helps improve the Flink runtime performance. TaskManagers also communicate with the JobManager to report progress and request necessary resources. This enables the JobManager to monitor the progress of tasks and to allocate resources accordingly to ensure optimal performance. In addition to the JobManager and TaskManager processes, Apache Flink also has a number of other components, including **DataStream API** and **DataSet API** for submitting jobs to the Flink runtime, a configuration system for setting up and tuning the Flink runtime, and a number of libraries and connectors for working with various data sources and sinks. For more information, see [Flink Architecture](https://nightlies.apache.org/flink/flink-docs-master/docs/concepts/flink-architecture/). --- # Settings for Apache Kafka® connectors Explore the necessary settings for standard and upsert Kafka connectors in Aiven for Apache Flink®. note Aiven for Apache Flink® supports the following data formats: JSON (default), Apache Avro, Confluent Avro, Debezium CDC. For more information on these, see the [Apache Flink® documentation on formats](https://ci.apache.org/projects/flink/flink-docs-release-1.19/docs/connectors/table/formats/overview/). | Parameter | Description | Standard connector | Upsert connector | | ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------- | ---------------- | | Key data format | Sets the format that is used to convert the *key* part of Kafka messages. | Optional | Required | | Key fields | Defines the columns from the SQL schema of the data table that are considered keys in the Kafka messages. | Optional (required if a key data format is selected) | Not available | | Value data format | Sets the format that is used to convert the *value* part of Kafka messages. | Required | Required | | Primary key | Defines the column in the SQL schema that is used to identify each message. Flink uses this to determine whether to insert a new message or update or delete an existing message. Defined with the `PRIMARY KEY` entry in the SQL schema for the data table. For example: `PRIMARY KEY (hostname) NOT ENFORCED` | Optional | Required | --- # Standard and upsert connectors for Apache Kafka® In addition to integration with Apache Kafka® through a standard connector, Aiven for Apache Flink® also supports the use of *upsert* connectors, which allows you to create changelog-type data streams. When you integrate a standard Apache Kafka® with your Aiven for Apache Flink® service, you can use the service to read data from one Kafka topic and write the processed data to another Kafka topic. Each message in the source topic that is processed and written to the sink topic is considered unique and is handled as an individual entry. Another integration approach is upsert, or `INSERT/UPDATE`, which is a method of updating or deleting data based on the message key. When used as a source, Flink interprets a message that has an existing key as an update and replaces the value for that key. If the key does not exist, it is inserted as new data, and if the message data is null, it is interpreted as a `DELETE` operation for that key. When used as a sink, upsert provides a method to create Kafka compacted topics (see the Log Compaction section in the [Apache Kafka® documentation](https://kafka.apache.org/documentation/) for details), containing only the latest value for a specific key and pushing a tombstone message on deletion. ## Choose the right connector type[​](#choose-the-right-connector-type "Direct link to Choose the right connector type") For most common purposes, such as filtering data or performing calculations on streamed data, a standard Kafka connector is sufficient. You can also perform data aggregation with standard connectors, but depending on the type of aggregation, it may require more careful planning and more complex SQL compared to using upsert connectors. To use multiple data streams as a source, but have the results for matching identifiers combined into single events within the same target topic, choose upsert connectors. While it varies to some extent depending on your specific use case, this means that using upsert connectors often makes it easier and more efficient to implement data aggregation. In addition, choose upsert connectors to provide the output as a compacted topic and read only the latest value for each message key. For more details on the required information for each connector type, see [Settings for Apache Kafka® connectors](/docs/products/flink/concepts/kafka-connector-requirements.md). --- # Savepoints Savepoints in Aiven for Apache Flink® are snapshots of the current state of your [Flink application](/docs/products/flink/howto/create-flink-applications.md). They are created when stopping an application deployment and allow you to restart it later on without losing any progress or state. Additionally, savepoints play a crucial role in disaster recovery for Aiven for Apache Flink Service. In the event of a failure or interruption, you can use savepoints to resume or restart your application (or Flink job) from the last known state, ensuring that you do not lose any progress. When you use Aiven for Apache Flink, explicitly trigger the creation of savepoints for an application when you stop it. --- # Built-in SQL editor The built-in Table SQL editor in the [Aiven Console](https://console.aiven.io/) for Aiven for Apache Flink® service provides an intuitive interface that allows you to create source and sink tables, write queries, and analyze data in one place. Additionally, the editor includes an auto-completion feature for SQL keywords, which makes the query-writing process faster and more efficient for you and helps with data validation. The interactive query feature also lets you preview the data of a Flink table or job without outputting the rows to a sink table. This is useful for testing and debugging your Flink application since it allows you to view the data that your application is processing. ![Image of the SQL editor in the Aiven for Apache Flink® data table view](/docs/assets/images/flink_sql_editor-66ceb17c973e53074e7a6a1ca1459871.png) --- # Tables in Aiven for Apache Flink® With Aiven for Apache Flink®, you can create and manage data pipelines using [Flink tables](https://nightlies.apache.org/flink/flink-docs-stable/docs/dev/table/sql/create/#create-table). These tables are created within an application on your Aiven for Apache Flink service, allowing you to map source and target data structures, and facilitating the process of data processing and transformation. The interactive query feature on the table editor allows you to review the data structure of the tables. This feature therefore helps validating the data ingested into your application, ensure accurate data transformation, and troubleshoot any issues that may arise. You can [add and manage tables](/docs/products/flink/howto/manage-flink-tables.md) within the Flink application on your Aiven for Apache Flink service. You can also create tables for various data services integrations, such as Aiven for Apache Kafka®, Aiven for PostgreSQL®, and Aiven for OpenSearch®, depending on the data source and the type of analysis to perform. Related pages * [Manage tables in Flink applications](/docs/products/flink/howto/manage-flink-tables.md) --- # Watermarks Apache Flink® uses watermarks to synchronize and process events in data streams accurately. These watermarks are timestamps embedded in the data stream that track the progression of event time. ## Role of watermarks[​](#role-of-watermarks "Direct link to Role of watermarks") Watermarks signal when all events up to a certain time have arrived, allowing Apache Flink operators to synchronize their event time clocks with these timestamps. This mechanism is crucial for timely and accurate event processing. Apache Flink® defines the watermark logic using watermark strategies and watermark generators. For example, you can configure Apache Flink to generate watermarks either periodically at specific intervals or when an event or element with a specific marker triggers it. Related pages For detailed information on watermarks and how to generate them in Apache Flink®, visit the [official documentation](https://ci.apache.org/projects/flink/flink-docs-release-1.19/docs/dev/datastream/event-time/generating_watermarks/). --- # Windows Apache Flink® uses the concept of *windows* to manage the continuous flow of data in streams by segmenting it into manageable chunks. This approach is essential due to the continuous and unbounded nature of data streams, where waiting for all data to arrive is impractical. ## How windows work[​](#how-windows-work "Direct link to How windows work") Windows in Apache Flink® is defined to segment the data stream into subsets for processing and analysis. A window is created when the first element that meets the specified criteria arrives. The trigger of a window determines when the window is ready for processing. Once the window is ready, the processing function determines how the data within the window is analyzed or manipulated. Additionally, each window has an **allowed lateness** value, indicating how long new events are accepted into the window after the trigger has closed it. ### Example scenario[​](#example-scenario "Direct link to Example scenario") Consider setting a time-based window from 15:00 to 15:10 with an allowed lateness of one minute. The window is created upon the arrival of the first event within this interval. Events arriving between 15:10 and 15:11, but still within the window's time range, are included. This mechanism provides flexibility in managing data that is slightly out of order or arrives later than expected. ### Handling late events[​](#handling-late-events "Direct link to Handling late events") Events that arrive after the allowed lateness should be managed separately, such as by logging and discarding them. This is done to ensure the integrity of windowed processing. ## Further reading[​](#further-reading "Direct link to Further reading") * [Apache Flink® windows](https://ci.apache.org/projects/flink/flink-docs-release-1.19/docs/dev/datastream/operators/windows/) --- # Get started with Aiven for Apache Flink® Begin your experience with Aiven for Apache Flink® by setting up a service, configuring data integrations, and building streaming applications. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to the [Aiven Console](https://console.aiven.io) * [Aiven CLI](https://github.com/aiven/aiven-client) ## Create an Aiven for Apache Flink® service[​](#create-an-aiven-for-apache-flink-service "Direct link to Create an Aiven for Apache Flink® service") * Console * CLI 1. In your project, click **Services**. 2. Click **Create service**. 3. Select **Apache Flink®**. 4. Enter a name for your service. important You cannot change the name after you create the service. 5. Optional: Add tags. 6. Select the cloud provider, region, and plan. note Available plans and pricing for the same service can vary between cloud providers and regions. 7. Optional: Add disk storage. 8. In the **Service summary**, click **Create service**. The status of the service is **Rebuilding** during its creation. When the status is **Running**, you can start using the service. This typically takes couple of minutes and can vary between cloud providers and regions. 1. [Create a token ](/docs/platform/howto/create_authentication_token.md). 2. Execute the following command: ``` avn service create flink-demo \ --service-type flink \ --cloud google-europe-west3 \ --plan business-4 \ -c flink_version=1.19 ``` Parameters: * `flink-demo`: The `service_name` for your new service. * `--service-type flink`: The type of service to be created, with `flink` indicating an Aiven for Apache Flink service. * `--cloud google-europe-west3`: Identifies the cloud region where the service will be hosted. * `--plan business-4`: Refers to the subscription plan that dictates the resource allocation for your service. * `-c flink_version=1.19`: This configuration setting specifies the version of Apache Flink that your service will run. ## Configure data service integrations[​](#configure-data-service-integrations "Direct link to Configure data service integrations") * Console * CLI Aiven for Apache Flink® streamlines data processing by enabling integration with various services. Currently, it supports integration with: * **Aiven for Apache Kafka®** and **external Apache Kafka clusters** * **Aiven for PostgreSQL®** * **Aiven for OpenSearch®** * **Google BigQuery®** ### Integration steps[​](#integration-steps "Direct link to Integration steps") 1. Log in to [Aiven Console](https://console.aiven.io) and access your Aiven for Apache Flink service. 2. On the **Overview** page, scroll to **Data pipeline**. 3. Click **Add data source**. 4. Choose the service you wish to integrate. 5. Click **Integrate**. tip For detailed integration steps with specific services, see the [Integrate service](/docs/products/flink/howto/create-integration.md). You can set up your data service integrations using the Aiven CLI with the following command: ``` avn service integration-create \ --project demo-sandbox \ --source-service kafka-demo \ --dest-service flink-demo \ --integration-type flink ``` Parameters: * `--project demo-sandbox`: Defines the project for the integration. * `--source-service kafka-demo`: Sets Aiven for Apache Kafka service as the data source. * `--dest-service flink-demo`: Sets Aiven for Apache Flink service as the data destination. * `--integration-type flink`: Determines the integration type, enabling data transfer from Aiven for Apache Kafka to Aiven for Apache Flink. ## Create an Aiven for Apache Flink® application[​](#create-an-aiven-for-apache-flink-application "Direct link to Create an Aiven for Apache Flink® application") An [Aiven for Apache Flink® application](/docs/products/flink/concepts/flink-applications.md) act as containers that encapsulate all aspects of a Flink job. This includes data source and sink connections, as well as the data processing logic. ### Application types[​](#application-types "Direct link to Application types") * **SQL applications**: These applications are ideal for executing SQL queries. The Aiven Console guides you through selecting source and sink tables, and setting up the SQL queries to process your data. To learn how to create SQL applications, see [Create an SQL application](/docs/products/flink/howto/create-sql-application.md). * **JAR applications**:These applications allow you to deploy custom functionalities. The Aiven Console enables you to upload and manage your JAR files and execute custom Flink jobs that go beyond the standard SQL capabilities. To learn how to create JAR applications, see [Create a JAR application](/docs/products/flink/howto/create-jar-application.md). ## Next steps[​](#next-steps "Direct link to Next steps") * Create source and sink data tables to map the data for [Apache Kafka®](/docs/products/flink/howto/connect-kafka.md), [PostgreSQL®](/docs/products/flink/howto/connect-pg.md) or [OpenSearch®](/docs/products/flink/howto/connect-opensearch.md) services * For details on using the Aiven CLI to create and manage Aiven for Apache Flink® services, see the [Aiven CLI documentation](/docs/tools/cli.md) and the [Flink-specific command reference](/docs/tools/cli/service/flink.md) --- # Integrate Aiven for Apache Flink® with Google BigQuery Connect Aiven for Apache Flink® with Google BigQuery as a sink using the [Aiven client](/docs/tools/cli.md) or the [Aiven Console](https://console.aiven.io/). Aiven for Apache Flink® is a fully managed service that provides distributed stateful stream processing capabilities. Google BigQuery is a cost-effective cloud-based data warehouse that can handle large amounts of data without servers. By connecting Aiven for Apache Flink® with Google BigQuery, you can stream data from Aiven for Apache Flink® to Google BigQuery, where it can be stored and analyzed. Aiven for Apache Flink® uses [BigQuery Connector for Apache Flink](https://github.com/aiven/bigquery-connector-for-apache-flink) as a connector to connect to Google BigQuery. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven for Apache Flink service * Google Cloud Platform (GCP) account * Necessary permissions to create resources and manage integrations in GCP * **Google Project ID**: You have a Google Project ID. For more information, see [Google Cloud documentation](https://cloud.google.com/resource-manager/docs/creating-managing-projects). * **Google Cloud Service Account Credentials**: You have Google Cloud service account credentials in JSON format to authenticate with the Google Cloud Platform. For instructions on how to create and get service account credentials, see [Google Cloud's documentation](https://developers.google.com/workspace/guides/create-credentials). * **Service account permissions**: Your service account is granted the necessary permissions to create log entries. For information on access control with IAM, see [Google Cloud's access control documentation](https://cloud.google.com/logging/docs/access-control). ## Configure integration using Aiven CLI[​](#configure-integration-using-aiven-cli "Direct link to Configure integration using Aiven CLI") ### Step 1: Create or use an Aiven for Apache Flink service[​](#step-1-create-or-use-an-aiven-for-apache-flink-service "Direct link to Step 1: Create or use an Aiven for Apache Flink service") You can use an existing Aiven for Apache Flink service. To get a list of all your existing Flink services, use: ``` avn service list --project --service-type flink ``` Alternatively, to create an Aiven for Apache Flink service, you can use: ``` avn service create -t flink -p --cloud ``` where: * `-t flink`: The type of service to create, which is Aiven for Apache Flink. * `-p `: The name of the Aiven project where the service should be created. * `cloud `: The name of the cloud provider on which the service should be created. * ``: The name of the new Aiven for Apache Flink service to be created. This name must be unique within the specified project. ### Step 2: Configure GCP for a Google BigQuery sink connector[​](#step-2-configure-gcp-for-a-google-bigquery-sink-connector "Direct link to Step 2: Configure GCP for a Google BigQuery sink connector") To be able to sink data from Aiven for Apache Flink to Google BigQuery, go the GCP console to configure GCP for a Google BigQuery sink connector: * Create a [Google service account and generate a JSON service key](https://cloud.google.com/docs/authentication/client-libraries). * Verify that BigQuery API is enabled. * Create a BigQuery dataset or define an existing one where the data is going to be stored. * Grant [dataset access to the service account](https://cloud.google.com/bigquery/docs/control-access-to-resources-iam). ### Step 3: Create an external Google BigQuery endpoint[​](#step-3-create-an-external-google-bigquery-endpoint "Direct link to Step 3: Create an external Google BigQuery endpoint") To integrate Google BigQuery with Aiven for Apache Flink, create an external BigQuery endpoint. You can use the [avn service integration-endpoint-create](/docs/tools/cli/service/integration.md#avn_service_integration_endpoint_create) command with the required parameters. This command will create a integration endpoint that can be used to connect to a BigQuery service. ``` avn service integration-endpoint-create \ --project \ --endpoint-name \ --endpoint-type external_bigquery \ --user-config-json '{ "project_id": "", "service_account_credentials": { "auth_provider_x509_cert_url": "https://www.googleapis.com/oauth2/v1/certs", "auth_uri": "https://accounts.google.com/o/oauth2/auth", "client_email": "", "client_id": "", "client_x509_cert_url": "", "private_key": "", "private_key_id": "", "project_id": "", "token_uri": "https://oauth2.googleapis.com/token", "type": "service_account" } }' ``` where: * `--project`: Specify the name of the project where to create the integration endpoint. * `--endpoint-name`: Set the name of the integration endpoint you are creating. Replace `your_endpoint_name` with your desired endpoint name. * `--endpoint-type`: Specify the type of integration endpoint. For example, if it's an external BigQuery service, enter `external_bigquery`. * `--user-config-json`: This parameter allows you to provide a JSON object with custom configurations for the integration endpoint. The JSON object should include the following fields: * `project_id`: Your actual Google Cloud Platform project ID. * `service_account_credentials`: An object that holds the necessary credentials for authenticating and accessing the external Google BigQuery service. This object should include the following fields: * `auth_provider_x509_cert_url`: The URL where the authentication provider'sx509 certificate can be fetched. * `auth_ur`: The URI used for authenticating requests. * `client_email`: The email address associated with the service account. * `client_id`: The client ID associated with the service account. * `client_x509_cert_url`: The URL to fetch the public x509 certificate for the service account. * `private_key`: The private key content associated with the service account. * `private_key_id`: The ID of the private key associated with the service account. * `project_id`: The project ID associated with the service account. * `token_uri`: The URI used to obtain an access token. * `type`: The type of service account, which is typically set to "service\_account". **Aiven CLI Example: Creating an external BigQuery integration endpoint** ``` avn service integration-endpoint-create --project aiven-test --endpoint-name my-bigquery-endpoint --endpoint-type external_bigquery --user-config-json '{ "project_id": "my-bigquery-project", "service_account_credentials": { "auth_provider_x509_cert_url": "https://www.googleapis.com/oauth2/v1/certs", "auth_uri": "https://accounts.google.com/o/oauth2/auth", "client_email": "bigquery-test@project.iam.gserviceaccount.com", "client_id": "284765298137902130451", "client_x509_cert_url": "https://www.googleapis.com/robot/v1/metadata/x509/bigquery-test%40project.iam.gserviceaccount.com", "private_key": "ADD_PRIVATE_KEY_PATH", "private_key_id": "ADD_PRIVE_KEY_ID_PATH", "project_id": "my-bigquery-project", "token_uri": "https://oauth2.googleapis.com/token", "type": "service_account" } }' ``` ### Step 4: Create an integration for Google BigQuery[​](#step-4-create-an-integration-for-google-bigquery "Direct link to Step 4: Create an integration for Google BigQuery") Now, create an integration between your Aiven for Apache Flink service and your BigQuery endpoint: ``` avn service integration-create --source-endpoint-id --dest-service -t flink_external_bigquery ``` For example, ``` avn service integration-create --source-endpoint-id eb870a84-b91c-4fd7-bbbc-3ede5fafb9a2 --dest-service flink-1 -t flink_external_bigquery ``` where: * `--source-endpoint-id`: The ID of the integration endpoint you want to use as the source. In this case, it is the ID of the external Google BigQuery integration endpoint. In this example, the ID is `eb870a84-b91c-4fd7-bbbc-3ede5fafb9a2`. * `--dest-service`: The name of the Aiven for Apache Flink service to integrate with the external BigQuery endpoint. In this example, the service name is `flink-1`. * `-t`: The type of integration to create. In this case, the `flink_external_bigquery` integration type is used to integrate Aiven for Apache Flink with an external BigQuery endpoint. ### Step 5: Verify integration with service[​](#step-5-verify-integration-with-service "Direct link to Step 5: Verify integration with service") After creating the integration between Aiven for Apache Flink and and Google BigQuery, the next step is to verify that the integration has been created successfully and create Aiven for Apache Flink applications that use the integration. To verify that the integration has been created successfully, run: ``` avn service integration-list --project ``` For example: ``` avn service integration-list --project systest-project flink-1 ``` where: * `--project`: The name of the Aiven project that contains the Aiven service to list integrations for. In this example, the project name is `systest-project`. * `flink-1`: The name of the Aiven service to list integrations for. In this example, the service name is `flink-1`, which is an Aiven for Apache Flink service. To create Aiven for Apache Flink applications, obtain the `integration_id` of the Aiven for Apache Flink service from the integration list. ### Step 6: Create Aiven for Apache Flink applications[​](#step-6-create-aiven-for-apache-flink-applications "Direct link to Step 6: Create Aiven for Apache Flink applications") With the `integration ID` obtained from the previous step, you can now create an application that uses the integration. For information on how to create Aiven for Apache Flink applications, see [avn service flink create-application](/docs/tools/cli/service/flink.md#avn%20service%20flink%20create-application). Following is an example of a Google BigQuery SINK table: ``` CREATE TABLE `table1` ( `name` STRING ) WITH ( 'connector' = 'bigquery', 'Service-account' = '', 'project-id'= '', 'dataset' = 'bqdataset', 'table' = 'bqtable', 'table-create-if-not-exists' = 'true', ) ``` If the integration is successfully created, the service credentials and project id will be automatically populated in the Sink (if you have left them back as shown in the example above). ## Configure integration using Aiven Console[​](#configure-integration-using-aiven-console "Direct link to Configure integration using Aiven Console") If you're using Google BigQuery for your data storage and analysis, you can seamlessly integrate it as a sink for Aiven for Apache Flink streams. To achieve this via the [Aiven Console](https://console.aiven.io/): 1. Log in to [Aiven Console](https://console.aiven.io/) and choose your project. 2. From the **Services** page, you can either create an Aiven for Apache Flink service or select an existing service. 3. Next, configure Google BigQuery service integration endpoint: * Go to the **Projects** page where all the services are listed. * From the left sidebar, select **Integration endpoints**. * Select **Google Cloud BigQuery** from the list, and select **Add new endpoint** or **Create new**. * Enter the following details to set up the integration: * **Endpoint name**: Enter a name for the integration endpoint. For example, `Aiven_BigQuery_Integration`. * **GCP Project ID**: The identifier associated with your Google Cloud Project where BigQuery is set up. For example, `my-gcp-project-12345`. * **Google Service Account Credentials**: The JSON formatted credentials obtained from your Google Cloud Console for service account authentication. For example: ``` { "type": "service_account", "project_id": "my-gcp-project-12345", "private_key_id": "abcd1234", ... } ``` * Select **Create**. 4. Select **Services** and access the Aiven for Apache Flink service where you plan to integrate the Google BigQuery endpoint. 5. If you're integrating with Aiven for Apache Flink for the first time, select **Create data pipeline** on the **Overview** page. Alternatively, you can add a new integration in the **Data Flow** section by using . 6. On the **Data Service integrations** page, select **Create external integration endpoint**. 7. Select the checkbox next to BigQuery, and choose the BigQuery endpoint from the list to integrate. 8. Select **Integrate**. You can now start creating [Aiven for Apache Flink applications](/docs/products/flink/howto/create-flink-applications.md) that use Google BigQuery as a sink. --- # Create an Apache Kafka®-based Apache Flink® table To build data pipelines, Apache Flink® requires you to map source and target data structures as [Flink tables](https://nightlies.apache.org/flink/flink-docs-stable/docs/dev/table/sql/create/#create-table) within an application. You can accomplish this through the [Aiven Console](https://console.aiven.io/) or [Aiven CLI](/docs/tools/cli/service/flink.md). When creating an application to manage streaming data, you can create a Flink table that connects to an existing or new Aiven for Apache Kafka® topic to source or sink streaming data. To define a table over an Apache Kafka® topic, specify the topic name, clearly define the columns' data format, and choose the appropriate connector type. Additionally, enter a clear and meaningful name to the table for reference when building data pipelines. warning To define Flink's tables an [existing integration](/docs/products/flink/howto/create-integration.md) needs to be available between the Aiven for Flink service and one or more Aiven for Apache Kafka services. ## Create Apache Flink® table with Aiven Console[​](#create-apache-flink-table-with-aiven-console "Direct link to Create Apache Flink® table with Aiven Console") To create an Apache Flink® table based on an Aiven for Apache Kafka® topic via the [Aiven Console](https://console.aiven.io/): 1. In the Aiven for Apache Flink service page, select **Application** from the left sidebar. 2. Create an application or select an existing one with Aiven for Apache Kafka integration. note If editing an existing application, create a version to make changes to the source or sink tables. 3. In the **Create new version** screen, select **Add source tables**. 4. Select **Add new table** or select **Edit** to edit an existing source table. 5. In the **Add new source table** or **Edit source table** screen, select the Aiven for Apache Kafka service as the integrated service. 6. In the **Table SQL** section, enter the SQL statement below to create the Apache Kafka-based Apache Flink: ``` CREATE TABLE kafka ( ) WITH ( 'connector' = 'kafka', 'properties.bootstrap.servers' = '', 'scan.startup.mode' = 'earliest-offset', 'topic' = '', 'value.format' = 'json', 'value.format' = 'json' ) ``` The following are the parameters: * `connector`: the **Kafka connector type**, between the **Apache Kafka SQL Connector** (value `kafka`) for standard topic reads/writes and the **Upsert Kafka SQL Connector** (value `upsert-kafka`) for changelog type of integration based on message key. note For more information on the connector types and the requirements for each, see the articles on [Kafka connector types](/docs/products/flink/concepts/kafka-connectors.md) and [the requirements for each connector type](/docs/products/flink/concepts/kafka-connector-requirements.md). * `properties.bootstrap.servers`: this parameter can be left empty since the connection details will be retrieved from the Aiven for Apache Kafka integration definition * `topic`: the topic to be used as a source for the data pipeline. To use a new topic that does not yet exist, write the topic name. warning By default, Flink will not be able to create Apache Kafka topics while pushing the first record automatically. To change this behavior, enable in the Aiven for Apache Kafka target service the `kafka.auto_create_topics_enable` option in **Advanced configuration** section. * `key.format`: specifies the **Key Data Format**. If a value other than **Key not used** is selected, specify the fields from the SQL schema to be used as key. This setting is specifically needed to set message keys for topics acting as target of data pipelines. * `value.format`: specifies the **Value Data Format**. Based on the message format in the Apache Kafka topic. note For Key and Value data format, the following options are available: * `json`: [JSON](https://nightlies.apache.org/flink/flink-docs-master/docs/connectors/table/formats/json/) * `avro`: [Apache Avro](https://nightlies.apache.org/flink/flink-docs-master/docs/connectors/table/formats/avro/) * `avro-confluent`: [Confluent Avro](https://nightlies.apache.org/flink/flink-docs-master/docs/connectors/table/formats/avro-confluent/). For information, see [Create Confluent Avro-based Apache Flink® table](/docs/products/flink/howto/flink-confluent-avro.md). 7. To create a sink table, select **Add sink tables** and repeat steps 4-6 for sink tables. 8. In the **Create statement** section, create a statement that defines the fields retrieved from each message in a topic, additional transformations such as format casting or timestamp extraction, and [watermark settings](/docs/products/flink/concepts/watermarks.md). ## Example: Define a Flink table using the standard connector over topic in JSON format[​](#example-define-a-flink-table-using-the-standard-connector-over-topic-in-json-format "Direct link to Example: Define a Flink table using the standard connector over topic in JSON format") The Aiven for Apache Kafka service named `demo-kafka` contains a topic named `metric-topic` holding a stream of service metrics in JSON format like: ``` {'hostname': 'sleepy', 'cpu': 'cpu3', 'usage': 93.30629927475789, 'occurred_at': 1637775077782} {'hostname': 'dopey', 'cpu': 'cpu4', 'usage': 88.39531418706092, 'occurred_at': 1637775078369} {'hostname': 'happy', 'cpu': 'cpu2', 'usage': 77.90860728236156, 'occurred_at': 1637775078964} {'hostname': 'dopey', 'cpu': 'cpu4', 'usage': 81.17372993952847, 'occurred_at': 1637775079054} ``` We can define a `metrics_in` Flink table by selecting `demo-kafka` as integration service and writing the following as SQL schema: ``` CREATE TABLE metrics_in ( cpu VARCHAR, hostname VARCHAR, usage DOUBLE, occurred_at BIGINT, time_ltz AS TO_TIMESTAMP_LTZ(occurred_at, 3), WATERMARK FOR time_ltz AS time_ltz - INTERVAL '10' SECOND ) WITH ( 'connector' = 'kafka', 'properties.bootstrap.servers' = '', 'topic' = 'metric-topic', 'value.format' = 'json', 'scan.startup.mode' = 'earliest-offset' ) ``` note The SQL schema includes: * the message fields `cpu`, `hostname`, `usage`, `occurred_at` and the related [data type](https://nightlies.apache.org/flink/flink-docs-release-1.16/docs/dev/table/types/#list-of-data-types). The order of fields in the SQL definition doesn't need to follow the order presented in the payload. * the definition of the field `time_ltz` as transformation to `TIMESTAMP(3)` from the `occurred_at` timestamp in Linux format. * the `WATERMARK` definition ## Example: Define a Flink table using the standard connector over topic in Avro format[​](#example-define-a-flink-table-using-the-standard-connector-over-topic-in-avro-format "Direct link to Example: Define a Flink table using the standard connector over topic in Avro format") In cases when target of the Flink data pipeline needs to write in Avro format to a topic named `metric_topic_tgt` within the Aiven for Apache Kafka service named `demo-kafka`. You can define a `metric_topic_tgt` Flink table by selecting the `demo-kafka` as integration service and writing the following SQL schema: ``` CREATE TABLE metric_topic_tgt ( cpu VARCHAR, hostname VARCHAR, usage DOUBLE ) WITH ( 'connector' = 'kafka', 'properties.bootstrap.servers' = '', 'topic' = 'metric-topic', 'value.format' = 'avro', 'scan.startup.mode' = 'earliest-offset' ) ``` note The SQL schema includes the output message fields `cpu`, `hostname`, `usage` and the related [data type](https://nightlies.apache.org/flink/flink-docs-release-1.16/docs/dev/table/types/#list-of-data-types). ## Example: Define a Flink table using the upsert connector over topic in JSON format[​](#example-define-a-flink-table-using-the-upsert-connector-over-topic-in-json-format "Direct link to Example: Define a Flink table using the upsert connector over topic in JSON format") In cases when target of the Flink pipeline needs to write in JSON format and upsert mode to a compacted topic named `metric_topic_tgt` within the Aiven for Apache Kafka service named `demo-kafka`. You can define a `metric_topic_tgt` Flink table by selecting `demo-kafka` as integration service and writing the following SQL schema: ``` CREATE TABLE metric_topic_tgt ( cpu VARCHAR, hostname VARCHAR, max_usage DOUBLE, PRIMARY KEY (cpu, hostname) NOT ENFORCED ) WITH ( 'connector' = 'upsert-kafka', 'properties.bootstrap.servers' = '', 'topic' = 'metric-topic', 'value.format' = 'json', 'scan.startup.mode' = 'earliest-offset' ) ``` note Unlikely the standard Apache Kafka SQL connector, when using the Upsert Kafka SQL connector the key fields are not defined. They are derived by the `PRIMARY KEY` definition in the SQL schema. note The SQL schema includes: * the output message fields `cpu`, `hostname`, `max_usage` and the related [data type](https://nightlies.apache.org/flink/flink-docs-release-1.16/docs/dev/table/types/#list-of-data-types). * the `PRIMARY KEY` definition, driving the key part of the Apache Kafka message --- # Create an OpenSearch®-based Apache Flink® table To build data pipelines, Apache Flink® requires you to map source and target data structures as [Flink tables](https://nightlies.apache.org/flink/flink-docs-stable/docs/dev/table/sql/create/#create-table) within an application. You can accomplish this through the [Aiven Console](https://console.aiven.io/) or [Aiven CLI](/docs/tools/cli/service/flink.md). You can define a Flink table over an existing or new Aiven for OpenSearch® index, to sink streaming data. To define a table over an OpenSearch® index, specify the index name, column data formats, and the Flink table name you want to use as a reference when building data pipelines. warning * To use Flink tables an [existing integration](/docs/products/flink/howto/create-integration.md) must be available between the Aiven for Flink® service and one or more Aiven for OpenSearch® services. * You can use Aiven for OpenSearch® can only as the **target** of a data pipeline. You can create Flink applications that write data to an OpenSearch® index. However, reading data from an OpenSearch® index is currently not possible. ## Create an OpenSearch®-based Apache Flink® table with Aiven Console[​](#create-an-opensearch-based-apache-flink-table-with-aiven-console "Direct link to Create an OpenSearch®-based Apache Flink® table with Aiven Console") To create an Apache Flink table based on an Aiven for OpenSearch® index via Aiven console: 1. In the Aiven for Apache Flink service page, select **Application** from the left sidebar. 2. Create an application or select an existing one with Aiven for OpenSearch® integration. note If editing an existing application, create a version to make changes to the source or sink tables. 3. In the **Create new version** screen, select **Add sink tables**. 4. Select **Add new table** or select **Edit** to edit an existing source table. 5. In the **Add new sink table** or **Edit sink table** screen, select the Aiven for OpenSearch® as the integrated service. 6. In the **Table SQL** section, enter the SQL statement to create the OpenSearch-based Apache Flink table. * Define the **Flink table name**, which will represent the Flink reference to the topic and will be used during the data pipeline definition. 7. In the **Create statement** section, write the SQL schema that defines the fields to be pushed for each message in the OpenSearch index. ## Example: Define an Apache Flink® table to OpenSearch®[​](#example-define-an-apache-flink-table-to-opensearch "Direct link to Example: Define an Apache Flink® table to OpenSearch®") We want to push the result of an Apache Flink® application to an index named `metrics` in an Aiven for OpenSearch® service named `demo-opensearch`. The application result should generate the following data: ``` {'hostname': 'sleepy', 'cpu': 'cpu3', 'usage': 93.30629927475789, 'occurred_at': 1637775077782} {'hostname': 'dopey', 'cpu': 'cpu4', 'usage': 88.39531418706092, 'occurred_at': 1637775078369} {'hostname': 'happy', 'cpu': 'cpu2', 'usage': 77.90860728236156, 'occurred_at': 1637775078964} {'hostname': 'dopey', 'cpu': 'cpu4', 'usage': 81.17372993952847, 'occurred_at': 1637775079054} ``` We can define a `metrics_out` Apache Flink® table with: * `demo-opensearch` as integration service * `metrics` as OpenSearch® index name * `metrics_out` as Flink table name * the following as SQL schema ``` CREATE TABLE metrics_out ( cpu VARCHAR, hostname VARCHAR, usage DOUBLE, occurred_at BIGINT ) WITH ( 'connector' = 'elasticsearch-7', 'hosts' = '', 'index' = 'metrics' ) ``` The `hosts` will be substituted with the appropriate address during runtime. --- # Create a PostgreSQL®-based Apache Flink® table To build data pipelines, Apache Flink® requires source and target data structures to [be mapped as Flink tables](https://ci.apache.org/projects/flink/flink-docs-release-1.19/docs/dev/table/sql/create/#create-table). This functionality can be achieved via the [Aiven console](https://console.aiven.io/) or [Aiven CLI](/docs/tools/cli/service/flink.md). A Flink table can be defined over an existing or new Aiven for PostgreSQL® table to be able to source or sink streaming data. To define a table over an PostgreSQL® table, the table name and columns data format need to be defined, together with the Flink table name to use as reference when building data pipelines. warning To define Flink tables, an [existing integration](/docs/products/flink/howto/create-integration.md) must be available between the Aiven for Flink service and one or more Aiven for PostgreSQL® services. ## Create a PostgreSQL®-based Apache Flink® table with Aiven Console[​](#create-a-postgresql-based-apache-flink-table-with-aiven-console "Direct link to Create a PostgreSQL®-based Apache Flink® table with Aiven Console") To create a Flink table based on Aiven for PostgreSQL® via Aiven console: 1. In the Aiven for Apache Flink service page, click **Application** from the left sidebar. 2. Create new application or select an existing one with Aiven for PostgreSQL® integration. note If editing an existing application, create a version to make changes to the source or sink tables. 3. In the **Create new version** screen, click **Add source tables**. 4. Click **Add new table** or click **Edit** to edit an existing source table. 5. In the **Add new source table** or **Edit source table** screen, select the Aiven for PostgreSQL® service as the integrated service. 6. In the **Table SQL** section, enter the SQL statement to create the PostgreSQL-based Apache Flink table with the following details: * Write the PostgreSQL® table name in the **JDBC table** field with the format `schema_name.table_name` warning Before you create an Apache Flink application, ensure that the PostgreSQL® table used as the target exists; otherwise, the application will fail. * Define the **Flink table name**. This name is used as a reference to the Apache Flink topic and in the definition of the data pipeline. 7. To create a sink table, click **Add sink tables** and repeat steps 4-6 for sink tables. 8. In the **Create statement** section, write the SQL schema that defines the fields retrieved from the PostgreSQL® table and any additional transformations, such as format casting or timestamp extraction. note More details on data types mapping between Apache Flink® and PostgreSQL® are available at the [dedicated JDBC Apache Flink® page](https://nightlies.apache.org/flink/flink-docs-master/docs/connectors/table/jdbc/#data-type-mapping). ## Example: Define a Flink table over a PostgreSQL® table[​](#example-define-a-flink-table-over-a-postgresql-table "Direct link to Example: Define a Flink table over a PostgreSQL® table") The Aiven for PostgreSQL® service named `pg-demo` contains a table named `students` in the `public` schema with the following structure: ``` CREATE TABLE students_tbl ( student_id INT, student_name VARCHAR ) WITH ( 'connector' = 'jdbc', 'url' = 'jdbc:postgresql://', 'table-name' = 'public.students' ) ) ``` The `url` parameter is dynamically updated at runtime. --- # Use Aiven for Apache Flink® applications [Aiven for Flink applications](/docs/products/flink/concepts/flink-applications.md) in Aiven for Apache Flink® servers as a container that includes everything connected to a Flink job, including source and sink connections and data processing logic. Using the [Aiven Console](https://console.aiven.io/), you can create applications that run SQL queries or deploy custom JARs, catering to diverse data processing requirements. The console's guided wizard simplifies the application configuration process, from selecting source and sink tables for SQL applications to uploading and managing JAR files for custom job execution. important Custom JARs for Aiven for Apache Flink is a [limited availability feature](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-). If you're interested in trying out this feature, contact the [sales team](https://aiven.io/contact). ## [Create an SQL application](/docs/products/flink/howto/create-sql-application.md) [Build data processing pipelines in Aiven for Apache Flink® by creating SQL applications using Apache Flink SQL. Set up source and sink tables, define processing logic, and manage your deployments.](/docs/products/flink/howto/create-sql-application.md) ## [Create a JAR application](/docs/products/flink/howto/create-jar-application.md) [Aiven for Apache Flink® enables you to upload and deploy custom code as a JAR file, enhancing your Apache Flink applications with advanced data processing capabilities.](/docs/products/flink/howto/create-jar-application.md) ## [Manage Aiven for Apache Flink® applications](/docs/products/flink/howto/manage-flink-applications.md) [This section provides information on managing your Aiven for Apache Flink® applications.](/docs/products/flink/howto/manage-flink-applications.md) ## [Restart strategy in SQL and JAR applications](/docs/products/flink/howto/restart-strategy-jar-applications.md) [Learn how Aiven for Apache Flink® applications uses restart strategies to recover from job failures, ensuring high availability and fault tolerance for your distributed applications.](/docs/products/flink/howto/restart-strategy-jar-applications.md) ## [Credential management for JAR applications](/docs/products/flink/howto/manage-credentials-jars.md) [Learn how to use the AVNCREDENTIALSDIR environment variable to securely manage credentials for custom JARs in Aiven for Apache Flink®.](/docs/products/flink/howto/manage-credentials-jars.md) --- # Create Apache Flink® data service integrations With Aiven for Apache Flink®, you can create streaming data pipelines to connect various services. Aiven for Apache Flink® enables integration with internal and external services, such as Aiven managed services, Google BigQuery®, and external Apache Kafka and PostgreSQL instances. This flexibility allows you to build data processing pipelines using a wide range of data sources and sinks. ## Create data service integration[​](#create-data-service-integration "Direct link to Create data service integration") Create Aiven for Apache Flink® data service integrations using the [Aiven Console](https://console.aiven.io/). ### Integrate Aiven services[​](#integrate-aiven-services "Direct link to Integrate Aiven services") 1. Log in to [Aiven Console](https://console.aiven.io) and access your Aiven for Apache Flink service. 2. If this is your first integration for the selected Aiven for Apache Flink service, on the **Overview** page, scroll to **Data pipeline**. 3. Click **Add data source** to initiate the integration setup. 4. On the **Data service integrations** page, under **Create service integration** tab, choose the Aiven service to integrate: Aiven for Apache Kafka®, Aiven for PostgreSQL®, or Aiven for OpenSearch®. 5. Click **Integrate**. 6. To include additional integrations, click **Plus** in the **Data pipeline** section. ### Integrate external services[​](#integrate-external-services "Direct link to Integrate external services") 1. On the **Data service integrations** page, click the **Create external integration endpoint** tab. 2. Select the type of external service to integrate: Google BigQuery®, External Apache Kafka, or External PostgreSQL. 3. If no external endpoints are available, go back to the **Projects** page. 4. Click **Integration endpoints** to add and configure external integration points. 5. Once the integration endpoint is added, return to your Aiven for Apache Flink service and the **Data pipeline** section. 6. Click **Add data source**. 7. On the **Data service integrations** page, under **Create external integration endpoint** tab, select the checkbox next to the external data service type and choose the integration endpoint just created. 8. To include additional integrations, click **Plus** in the **Data pipeline** section. --- # Create a JAR application [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Aiven for Apache Flink® enables you to upload and deploy [custom code as a JAR file](/docs/products/flink/concepts/custom-jars.md), enhancing your Apache Flink applications with advanced data processing capabilities. ## Prerequisite[​](#prerequisite "Direct link to Prerequisite") * Custom JARs for Aiven for Apache Flink is a **limited availability** feature. To try this feature, request access by contacting the [sales team](https://aiven.io/contact). * Create an Aiven for Apache Flink service and click **Upload and deploy custom JARs** to enable custom JARs during service creation. note Enabling Custom JARs for existing services is currently not possible. If you did not enable this feature during service creation, you must create a service with Custom JARs enabled. ## Create and deploy a JAR application[​](#create-and-deploy-a-jar-application "Direct link to Create and deploy a JAR application") 1. Access the [Aiven Console](https://console.aiven.io/) and select the Aiven for Apache Flink service to deploy a JAR application to. 2. From the left sidebar, click **Applications** and click **Create application**. 3. In the **Create application** dialog, enter a name for your JAR application, and select **JAR** as the application type from the drop-down. 4. Click **Create application** to proceed. 5. Click **Upload first version** to upload the first version of the application. Custom JAR file size limit Custom JAR files have a maximum size of 100 MB. For more information, contact the sales or support team. 6. In the **Upload new version** dialog: 1. Click **Choose file** to select your custom JAR file. 2. Select the **Terms of Service** checkbox to indicate your agreement. 3. Click **Upload version** to upload your JAR file. 7. After the upload, you are redirected to the application's overview page. 8. To deploy the application, click **Create deployment**. In the **Create new deployment** dialog: 1. Select the application version to deploy. 2. Select a [savepoint](/docs/products/flink/concepts/savepoints.md) if you wish to deploy from a specific state. No savepoints are available for the first application deployment. 3. Toggle **Restart on failure** to automatically restart Flink jobs upon failure. See [Restart strategy in SQL and JAR applications](/docs/products/flink/howto/restart-strategy-jar-applications.md) for details. 4. In the **Program args** field, provide command-line arguments consisting of variables and configurations relevant to your application's logic upon submission. Each argument is limited to 64 characters, with a total limit of 32 separate items. 5. Specify the number of [parallel instances](https://nightlies.apache.org/flink/flink-docs-master/docs/dev/datastream/execution/parallel/) you require for the task. 9. Click **Deploy without a savepoint** to begin the deployment process. 10. While deploying, the application status shows **Initializing**. Once deployed, the status changes to **Running**. Related pages * [Manage Aiven for Apache Flink® applications](/docs/products/flink/howto/manage-flink-applications.md) --- # Create an SQL application Build data processing pipelines in Aiven for Apache Flink® by creating SQL applications using Apache Flink SQL. Set up source and sink tables, define processing logic, and manage your deployments. ## Prerequisite[​](#prerequisite "Direct link to Prerequisite") Before creating applications, configure the [data service integration](/docs/products/flink/howto/create-integration.md) for seamless integration and data management within your Flink applications. ## Create and deploy an SQL application[​](#create-and-deploy-an-sql-application "Direct link to Create and deploy an SQL application") Create an SQL application in Aiven for Apache Flink® using the [Aiven Console](https://console.aiven.io/): 1. In the [Aiven Console](https://console.aiven.io/), select the Aiven for Apache Flink service where to create and deploy a Flink application. 2. From the left sidebar, click **Applications** and click **Create application**. 3. In the **Create application** dialog, enter the name of your application and select **SQL** as the application type. 4. Click **Create application**. 5. Click **Create first version** to create the first version of the application. 6. Click **Add your first source table** to add a source table. note As this is your first application, no other applications are available to import source tables. 7. On the **Add new source table** screen: * Use the **Integrated service** drop-down to select the service. * In the **Table SQL** section, enter the SQL statement to create the source table. * Optionally, click **Run** to test how data is being retrieved from the data source. This may vary in time based on the data volume and connection speed. * Click **Add table**. 8. Click **Next** to proceed to adding a sink table and click **Add your first sink table**. note As this is your first application, no other applications are available to import sink tables. 9. On the **Add new sink table** screen: * Use the **Integrated service** drop-down to select the service. * In the **Table SQL** section, enter the SQL statement to create the sink table. * Click **Add table**. 10. Click **Next** to enter the **SQL statement** that transforms the data from the source stream. Optionally, click **Run** to see how the data is extracted from the source. 11. Click **Save and deploy later** to save the application. You can view and access the application you created on the application overview page. ![Application landing page with a view of the source table, SQL statement, and sink table](/docs/assets/images/application_landingpage_view-7688038d517e905d771ed34a8c18d866.png) 12. To deploy the application, click **Create deployment**. In the **Create new deployment** dialog: * Select the application version to deploy. The default version for the first deployment is **Version: 1**. * Select a [savepoint](/docs/products/flink/concepts/savepoints.md) if you wish to deploy from a specific state. No savepoints are available for the first application deployment. * Toggle **Restart on failure** to automatically restart Flink jobs upon failure. See [Restart strategy in SQL and JAR applications](/docs/products/flink/howto/restart-strategy-jar-applications.md) for details. * Specify the number of [parallel instances](https://nightlies.apache.org/flink/flink-docs-master/docs/dev/datastream/execution/parallel/) you require for the task. 13. Click **Deploy without a savepoint** to begin the deployment process. 14. While deploying, the application status shows **Initializing**. Once deployed, the status changes to **Running**. ## Create SQL applications using Aiven CLI[​](#create-sql-applications-using-aiven-cli "Direct link to Create SQL applications using Aiven CLI") For information on creating and managing Aiven for Apache Flink application using [Aiven CLI](/docs/tools/cli.md), see [Manage Aiven for Apache Flink® applications](/docs/tools/cli/service/flink.md) document. --- # Create a DataGen-based Apache Flink® table The DataGen source table is a built-in connector of the Apache Flink system that generates random data periodically, matching the specified data type of the source table. This section provides you with information on how to connect DataGen as a source table in an Aiven for Apache ®Flink application. ## Configure Datagen as source for Flink application[​](#configure-datagen-as-source-for-flink-application "Direct link to Configure Datagen as source for Flink application") To configure DataGen as the source using the DataGen built-in connector for Apache Flink: 1. In the Aiven for Apache Flink service page, select **Application** from the left sidebar. 2. Create an application or select an existing application for your desired [data service integration](/docs/products/flink/howto/create-integration.md). note If you are editing an existing application, create a version of the application to make changes to the source or sink table. 3. In the **Create new version** screen, select **Add source tables**. 4. Select **Add new table** or select **Edit** to edit an existing source table. 5. In the **Table SQL** section of the **Add new source table** or **Edit source table** screen, set the connector to **datagen** as shown in the example below: ``` CREATE TABLE `gen_me` ( `id` INT, `price` DECIMAL(32,2) ) WITH ( `connector` = 'datagen' ) ``` Where: * `connector`: The connector parameter is set to `datagen`, a built-in connector in Flink that generates random data in memory and is used as a source of data in Flink applications. note For more information on the connector types and the requirements for each, see the articles on [Kafka connector types](/docs/products/flink/concepts/kafka-connectors.md) and [the requirements for each connector type](/docs/products/flink/concepts/kafka-connector-requirements.md). 6. In the **Add sink tables** screen, select the option to add a new sink table or edit an existing one. 7. In the **Create statement** section, write the statement to test your SQL queries using random data. --- # Integrate Aiven for Apache Flink® with Apache Kafka® Integrating external/self-hosted Apache Kafka® with Aiven for Apache Flink® allows users to leverage the power of both technologies to build scalable and robust real-time streaming applications. This section provides instructions on integrating external/self-hosted Apache Kafka with Aiven for Apache Flink® using [Aiven client](/docs/tools/cli.md) and [Aiven Console](https://console.aiven.io/). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven for Apache Flink service * External/self-hosted Apache Kafka service ## Configure integration using CLI[​](#configure-integration-using-cli "Direct link to Configure integration using CLI") To configure integration using Aiven CLI: ### Step 1. Create an Aiven for Apache Flink service[​](#step-1-create-an-aiven-for-apache-flink-service "Direct link to Step 1. Create an Aiven for Apache Flink service") Use the following command to create an Aiven for Apache Flink service: ``` avn service create -t flink -p --cloud ``` where: * `-t flink`: The type of service to create, which is Aiven for Apache Flink. * `-p `: The name of the Aiven project where the service should be created. * `cloud `: The name of the cloud provider on which the service should be created. * ``: The name of the new Aiven for Apache Flink service to be created. This name must be unique within the specified project. ### Step 2: Setup an Apache Kafka service for the integration[​](#step-2-setup-an-apache-kafka-service-for-the-integration "Direct link to Step 2: Setup an Apache Kafka service for the integration") If you currently have a self-hosted or external Apache Kafka instance, you have the choice of either using it to integrate with the Aiven for Apache Flink service or creating a new Apache Kafka service using your preferred method. ### Step 3: Download Apache Kafka certificate[​](#step-3-download-apache-kafka-certificate "Direct link to Step 3: Download Apache Kafka certificate") Download the necessary Apache Kafka credentials certificate using your preferred method and save it in a secure and accessible location on your system. ### Step 4: Create an external Apache Kafka endpoint[​](#step-4-create-an-external-apache-kafka-endpoint "Direct link to Step 4: Create an external Apache Kafka endpoint") To integrate external Apache Kafka with Aiven for Apache Flink, you need to create an external Apache Kafka endpoint. You can use the [avn service integration-endpoint-create](/docs/tools/cli/service/integration.md#avn_service_integration_endpoint_create) command with the required parameters. This command will create an integration endpoint that can be used to connect to an external Apache Kafka service. The available security protocols for Kafka are PLAINTEXT, SSL, SASL\_PLAINTEXT, and SASL\_SSL. #### PLAINTEXT[​](#plaintext "Direct link to PLAINTEXT") To create a PLAINTEXT protocol type endpoint, use the following command: ``` avn service integration-endpoint-create --endpoint-name demo-ext-kafka --endpoint-type external_kafka --user-config-json '{"bootstrap_servers":"servertest:123","security_protocol":"PLAINTEXT"}' ``` Where : * `--endpoint-name`: Name of the endpoint to create. * `--endpoint-type`: The type of endpoint, which should be `external_kafka`. * `--user-config-json`: The configuration for the endpoint in JSON format, which includes the following attributes: * `bootstrap_servers`: List of Apache Kafka broker addresses and ports to connect to. * `security_protocol`: Security protocol to use for the connection. In this example, it's set to **PLAINTEXT**. #### SSL[​](#ssl "Direct link to SSL") To create an SSL protocol type endpoint, use the following command: ``` avn service integration-endpoint-create --endpoint-name demo-ext-kafka \ --endpoint-type external_kafka \ --user-config-json '{ "bootstrap_servers": "10.0.0.1:9092,10.0.0.2:9092,10.0.0.3:9092", "security_protocol": "SSL", "ssl_ca_cert": ssl_cert, "ssl_client_cert": ssl_cert, "ssl_client_key": ssl_key, "ssl_endpoint_identification_algorithm":"https", }' ``` Where: * `--endpoint-name`: Name of the endpoint to create. * `--endpoint-type`: The type of endpoint, which should be `external_kafka`. * `--user-config-json`:The configuration for the endpoint in JSON format, which includes the following attributes: * `bootstrap_servers`: List of Apache Kafka broker addresses and ports to connect to. * `security_protocol`: The type of security protocol to use for the connection, which is `SASL` in this case. * `ssl_ca_cert`: The content of the SSL CA certificate. * `ssl_client_cert`: The content of the SSL client certificate. * `ssl_client_key`: The content of the SSL client key. * `ssl_endpoint_identification_algorithm`: The endpoint identification algorithm to use for SSL verification. For example, `https`. important After downloading your keys or certificates, ensure the cipher is on its own line, and the PEM markers are delimited by a line feed, following the guidelines in [RFC 1421](https://www.rfc-editor.org/rfc/rfc1421#section-4.4). Use the following bash command to format the content correctly: ``` cat $downloaded_cert_or_key | tr -d '\n' | sed 's/\([EY]---[-]*\)\([^-]\)/\1\n\2/g;s/\(=\)\(---[-]*\)/\1\n\2/g' ``` #### SASL\_PLAINTEXT[​](#sasl_plaintext "Direct link to SASL_PLAINTEXT") To create a SASL\_PLAINTEXT protocol type endpoint, use the following command: ``` avn service integration-endpoint-create --endpoint-name demo-ext-kafka \ --endpoint-type external_kafka \ --user-config-json '{ "bootstrap_servers": "10.0.0.1:9092,10.0.0.2:9092,10.0.0.3:9092", "security_protocol": "SASL_PLAINTEXT", "sasl_mechanism": "PLAIN", "sasl_plain_username": sasl_username, "sasl_plain_password": sasl_password }' ``` where: * `--endpoint-name`: Name of the endpoint to create. * `--endpoint-type`: The type of endpoint, which should be `external_kafka`. * `--user-config-json`:The configuration for the endpoint in JSON format, which includes the following attributes: * `bootstrap_servers`: List of Apache Kafka broker addresses and ports to connect to. * `security_protocol`: The type of security protocol to use for the connection, which is `SASL_PLAINTEXT` in this case. * `sasl_mechanism`: The type of SASL mechanism to use for authentication, which is **PLAIN** in this case. * `sasl_plain_username`: The username for SASL authentication. * `sasl_plain_password`: The password for SASL authentication. * `ssl_endpoint_identification_algorithm`: The endpoint identification algorithm to use for SSL verification. For example, `https`. #### SASL\_SSL[​](#sasl_ssl "Direct link to SASL_SSL") To create a SASL\_SSL protocol type endpoint, use the following command: ``` avn service integration-endpoint-create --endpoint-name demo-ext-kafka \ --endpoint-type external_kafka \ --user-config-json '{ "bootstrap_servers": "10.0.0.1:9092,10.0.0.2:9092,10.0.0.3:9092", "security_protocol": "SASL_SSL", "sasl_mechanism": "PLAIN", "sasl_plain_username": sasl_username, "sasl_plain_password": sasl_password, "ssl_ca_cert": ssl_cert, "ssl_endpoint_identification_algorithm": "https" }' ``` where: * `--endpoint-name`: Name of the endpoint to create. * `--endpoint-type`: The type of endpoint, which should be `external_kafka`. * `--user-config-json`:The configuration for the endpoint in JSON format, which includes the following attributes: * `bootstrap_servers`: List of Apache Kafka broker addresses and ports to connect to. * `security_protocol`: The type of security protocol to use for the connection, which is `SASL_SSL` in this case. * `sasl_mechanism`: The type of SASL mechanism to use for authentication, which is **PLAIN** in this case. * `sasl_plain_username`: The username for SASL authentication. * `sasl_plain_password`: The password for SASL authentication. * `ssl_ca_cert`: The path to the SSL CA certificate downloaded for SSL authentication. * `ssl_endpoint_identification_algorithm`: The endpoint identification algorithm to use for SSL verification. For example, `https`. ### Step 5: Integrate Aiven for Apache Flink with endpoints[​](#step-5-integrate-aiven-for-apache-flink-with-endpoints "Direct link to Step 5: Integrate Aiven for Apache Flink with endpoints") To integrate Aiven for Apache Flink with the integration endpoint for external Apache Kafka, use the following command: ``` avn service integration-create --source-endpoint-id --dest-service -t flink_external_kafka ``` For example, ``` avn service integration-create --source-endpoint-id eb870a84-b91c-4fd7-bbbc-3ede5fafb9a2 --dest-service flink-1 -t flink_kafka ``` where: * `--source-endpoint-id`: The ID of the integration endpoint you want to use as the source. In this case, it is the ID of the external Apache Kafka integration endpoint. In this example, the ID is `eb870a84-b91c-4fd7-bbbc-3ede5fafb9a2`. * `--dest-service`: The name of the Aiven for Apache Flink service you want to integrate with the external Apache Kafka endpoint. In this example, the service name is `flink-1`. * `-t`: The type of integration to create. In this case, the `flink_external_kafka` integration type is used to integrate Aiven for Apache Flink with an external Apache Kafka endpoint. ### Step 6: Verify integration with service[​](#step-6-verify-integration-with-service "Direct link to Step 6: Verify integration with service") After creating the integration between Aiven for Apache Flink and external/self-hosted Apache Kafka, the next step is to verify that the integration has been created successfully and create applications that use the integration. To verify that the integration has been created successfully, run the following command: ``` avn service integration-list --project ``` For example: ``` avn service integration-list --project systest-project flink-1 ``` where: * `--project`: The name of the Aiven project that contains the Aiven service to list integrations for. In this example, the project name is `systest-project`. * `flink-1`: The name of the Aiven service to list integrations for. In this example, the service name is `flink-1`, which is an Aiven for Apache Flink service. To create Aiven for Apache Flink applications, you will need the integration ID of the Aiven for Apache Flink service. Obtain the `integration_id` from the integration list. ### Step 7: Create Aiven for Apache Flink applications[​](#step-7-create-aiven-for-apache-flink-applications "Direct link to Step 7: Create Aiven for Apache Flink applications") With the integration ID obtained from the previous step, you can now create an application that uses the integration. For information on how to create Aiven for Apache Flink applications, see [avn service flink create-application](/docs/tools/cli/service/flink.md#avn%20service%20flink%20create-application). ## Configure integration using Aiven Console[​](#configure-integration-using-aiven-console "Direct link to Configure integration using Aiven Console") If you have an external Apache Kafka service already running, you can integrate it with Aiven for Apache Flink using the [Aiven Console](https://console.aiven.io/): 1. Log in to [Aiven Console](https://console.aiven.io/) and choose your project. 2. From the **Services** page, you can either create an Aiven for Apache Flink service or select an existing service. 3. Next, configure an external Apache Kafka service integration endpoint: * Go to the Projects screen where all the services are listed. * From the left sidebar, select **Integration endpoints**. * Select **External Apache Kafka** from the list, and select **Add new endpoint**. * Enter an *Endpoint name* and the *Bootstrap servers*. Then, choose a *Security protocol* and select **Create**. 4. Select **Services** from the left sidebar, and access the Aiven for Apache Flink service where you plan to integrate the external Apache Kafka endpoint. 5. If you're integrating with Aiven for Apache Flink for the first time, on the **Overview** page and select **Get Started**. Alternatively, you can add a new integration in the **Data Flow** section by clicking **Plus (+)**. 6. On the **Data Service integrations** screen, select the **Create external integration endpoint** tab. 7. Select the checkbox next to **Apache Kafka**, and choose the external Apache Kafka endpoint from the list to integrate. 8. Select **Integrate**. The integration is ready, and you can start creating [Aiven for Apache Flink applications](/docs/products/flink/howto/create-flink-applications.md) that use the external Apache Kafka service as either a source or sink. --- # Create Confluent Avro-based Apache Flink® table [Confluent Avro](https://nightlies.apache.org/flink/flink-docs-release-1.16/docs/connectors/table/formats/avro-confluent/) is a serialization format that requires integrating with a schema registry. This enables the serialization and deserialization of data in a format that is language-agnostic, easy to read, and supports schema evolution. Aiven for Apache Flink® simplifies the process of creating an Apache Flink® source table that uses the Confluent Avro data format with Karapace, an open-source schema registry for Apache Kafka® that enables you to store and retrieve Avro schemas. With Aiven for Apache Flink, you can stream data in the Confluent Avro data format and perform real-time transformations. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven for Apache Flink service with Aiven for Apache Kafka® integration. See [Create Apache Flink® data service integrations](/docs/products/flink/howto/create-integration.md) for more information. * Aiven for Apache Kafka® service with Karapace Schema registry enabled. See [Create Apache Flink® data service integrations](/docs/products/flink/howto/create-integration.md) for more information. * By default, Flink cannot create Apache Kafka topics while pushing the first record automatically. To change this behavior, enable in the Aiven for Apache Kafka target service the `kafka.auto_create_topics_enable` option in **Advanced configuration** section. ## Create an Apache Flink® table with Confluent Avro[​](#create-an-apache-flink-table-with-confluent-avro "Direct link to Create an Apache Flink® table with Confluent Avro") 1. In the Aiven for Apache Flink service page, select **Application** from the left sidebar. 2. Create an application or select an existing one with Aiven for Apache Kafka® integration. note If editing an existing application, create a version to make changes to the source or sink tables. 3. In the **Create new version** screen, select **Add source tables**. 4. Select **Add new table** or select **Edit** to edit an existing source table. 5. In the **Add new source table** or **Edit source table** screen, select the Aiven for Apache Kafka® service as the integrated service. 6. In the **Table SQL** section, enter the SQL statement below to create an Apache Kafka®-based Apache Flink® table with Confluent Avro: ``` CREATE TABLE kafka ( -- specify the table columns ) WITH ( 'connector' = 'kafka', 'properties.bootstrap.servers' = '', 'scan.startup.mode' = 'earliest-offset', 'topic' = 'my_test.public.students', 'value.format' = 'avro-confluent', -- the value data format is Confluent Avro 'value.avro-confluent.url' = 'http://localhost:8082', -- the URL of the schema registry 'value.avro-confluent.basic-auth.credentials-source' = 'USER_INFO', -- the source of the user credentials for accessing the schema registry 'value.avro-confluent.basic-auth.user-info' = 'user_info' -- the user credentials for accessing the schema registry ) ``` The following are the parameters: * `connector`: the **Kafka connector type**, between the **Apache Kafka SQL Connector** (value `kafka`) for standard topic reads/writes and the **Upsert Kafka SQL Connector** (value `upsert-kafka`) for changelog type of integration based on message key. note For more information on the connector types and the requirements for each, see the articles on [Kafka connector types](/docs/products/flink/concepts/kafka-connectors.md) and [the requirements for each connector type](/docs/products/flink/concepts/kafka-connector-requirements.md). * `properties.bootstrap.servers`: this parameter can be left empty since the connection details will be retrieved from the Aiven for Apache Kafka integration definition * `topic`: the topic to be used as a source for the data pipeline. To use a new topic that does not yet exist, write the topic name. * `value.format`: indicates that the value data format is in the Confluent Avro format. note The `key.format` parameter can also be set to the `avro-confluent` format. * `avro-confluent.url`: this is the URL for the Karapace schema registry. * `value.avro-confluent.basic-auth.credentials-source`: this specifies the source of the user credentials for accessing the Karapace schema registry. At present, only the `USER_INFO` value is supported for this parameter. * `value.avro-confluent.basic-auth.user-info`: this should be set to the `user_info` string you created earlier. important To access the Karapace schema registry, the user needs to provide the username and password using the `user_info` parameter. The `user_info` parameter is a string formatted as `user_info = f"{username}:{password}"`. Additionally, on the source table, the user only needs read permission to the subject containing the schema. However, on the sink table, if the schema does not exist, the user must have write permission for the schema registry. It is important to provide this information to authenticate and access the Karapace schema registry. 7. To create a sink table, select **Add sink tables** and repeat steps 4-6 for sink tables. 8. In the **Create statement** section, create a statement that defines the fields retrieved from each message in a topic. ## Example: Define a Flink table using the standard connector over topic in Confluent Avro format[​](#example-define-a-flink-table-using-the-standard-connector-over-topic-in-confluent-avro-format "Direct link to Example: Define a Flink table using the standard connector over topic in Confluent Avro format") The Aiven for Apache Kafka service called `demo-kafka` includes a topic called `my_test.public.student` that holds a stream of student data in Confluent Avro format like: ``` {"id": 1, "name": "John", "email": "john@gmail.com"} {"id": 2, "name": "Jane", "email": "jane@yahoo.com"} {"id": 3, "name": "Bob", "email": "bob@hotmail.com"} {"id": 4, "name": "Alice", "email": "alice@gmail.com"} ``` You can define a `students` Flink table by selecting `demo-kafka` as the integration service and writing the following SQL schema: ``` CREATE TABLE students ( id INT, name STRING, email STRING ) WITH ( 'connector' = 'kafka', 'properties.bootstrap.servers' = '', 'scan.startup.mode' = 'earliest-offset', 'topic' = 'my_test.public.students', 'value.format' = 'avro-confluent' 'value.avro-confluent.url' = 'http://localhost:8082', 'value.avro-confluent.basic-auth.credentials-source'= 'USER_INFO', 'value.avro-confluent.basic-auth.user-info" = 'user_info', ) ``` note The SQL schema includes the output message fields `id`, `name`, `email` and the related [data type](https://nightlies.apache.org/flink/flink-docs-release-1.16/docs/dev/table/types/#list-of-data-types). --- # Credential management for JAR applications Learn how to use the `AVN_CREDENTIALS_DIR` environment variable to securely manage credentials for custom JARs in Aiven for Apache Flink®. ## About credential management[​](#about-credential-management "Direct link to About credential management") [Custom JARs](/docs/products/flink/concepts/custom-jars.md) in Aiven for Apache Flink® enable you to connect your Apache Flink jobs with Aiven's supported connectors and any external systems you manage, which is essential for real-time stream processing. Effectively managing credentials is critical to securing and ensuring compliant access to these services. Aiven for Apache Flink® securely manages credentials for your JAR applications using the `AVN_CREDENTIALS_DIR` environment variable. This centralizes credentials for internal and external integrations, ensuring policy-compliant and secure access to sensitive information. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An active Aiven for Apache Flink service * [Integration with services: Aiven for Apache Kafka or PostgreSQL](/docs/products/flink/howto/create-integration.md) or external Apache kafka service * Permission to create a [JAR application for Aiven for Apache Flink](/docs/products/flink/howto/create-jar-application.md) ## Credential provisioning[​](#credential-provisioning "Direct link to Credential provisioning") Aiven streamlines credential management for Aiven-managed and external services using the `AVN_CREDENTIALS_DIR` environment variable. This variable points to a directory with integrated service credentials, providing access for users while abstracting the intricacies of internal storage. note You can have multiple credentials for multiple Apache Kafka services in the same `AVN_CREDENTIALS_DIR` directory. Aiven automatically generates credentials for integrated services and stores them in JSON files named by each service's `integration_id`. For example, credentials for an Apache Kafka service with an `integration_id` such as `my_kafka_service` is stored in a file named `my_kafka_service.json`. This file is in a directory your Aiven for Apache Flink® application can access. The path to this folder is `/AVN_CREDENTIALS_DIR/my_kafka_service.json`. ## Access credentials in JAR applications[​](#access-credentials-in-jar-applications "Direct link to Access credentials in JAR applications") 1. Locate the credentials file 1. Identify the `integration_id` of your service from the [integration list](/docs/tools/cli/service/integration.md#avn_service_integration_list). 2. Retrieve the corresponding credentials file named `{integration_id}.json` located at `/AVN_CREDENTIALS_DIR/`. For example, if your service's `integration_id` is `my_kafka_service`, locate the credentials file `my_kafka_service.json` at `/AVN_CREDENTIALS_DIR/my_kafka_service.json`. 2. Read and parse the JSON file 1. Implement the code within your Aiven for Apache Flink JAR application to read the JSON file. 2. Extract essential details like connection strings, usernames, passwords, and security protocols from the JSON file. ## Example: Parsing credentials in JAR applications[​](#example-parsing-credentials-in-jar-applications "Direct link to Example: Parsing credentials in JAR applications") Aiven for Apache Flink® enables connections to various data sources and sinks, allowing you to create JAR applications. Managing credentials within these applications is crucial for securing data access and ensuring compliance with data protection standards. This section provides examples of how to parse credentials to ensure data security. ### Example 1: Integration with Aiven for Apache Kafka[​](#example-1-integration-with-aiven-for-apache-kafka "Direct link to Example 1: Integration with Aiven for Apache Kafka") To connect to Aiven for Apache Kafka service, the credentials JSON structure is as follows: ``` { "bootstrap_servers": "t-kafka-2137baed0dd94da2.aiven.local:15357", "integration_type": "kafka", "schema_registry_url": "https://t-kafka-2137baed0dd94da2-systest-project-test.avns.net:15361", "security_protocol": "PLAINTEXT", "service_name": "t-kafka-2137baed0dd94da2" } ``` note The JSON structure provided in the code snippet is specific to Aiven services and may differ from standard Apache Kafka configuration formats. **Java application example:** ``` import org.json.simple.JSONObject; import org.json.simple.parser.JSONParser; import java.io.FileReader; import org.apache.flink.api.java.utils.ParameterTool; public class KafkaCredentialsReader { public static void main(String[] args) { // Parse the command-line arguments to get the Kafka source name ParameterTool parameters = ParameterTool.fromArgs(args); String myKafkaSource = parameters.getRequired("myKafkaSource"); // Construct the path to the credentials file dynamically String credentialsFilePath = System.getenv("AVN_CREDENTIALS_DIR") + "/" + myKafkaSource + ".json"; JSONParser jsonParser = new JSONParser(); try (FileReader reader = new FileReader(credentialsFilePath)) { // Parse the JSON file JSONObject kafkaJson = (JSONObject) jsonParser.parse(reader); // Extract necessary details String bootstrapServers = (String) kafkaJson.get("bootstrap_servers"); String securityProtocol = (String) kafkaJson.get("security_protocol"); // Configure Kafka source with the extracted details KafkaSource.builder() .setProperty("bootstrap.servers", bootstrapServers) .setProperty("security.protocol", securityProtocol) // Additional Kafka configuration as needed .build(); } catch (Exception e) { e.printStackTrace(); } } } ``` In this example: * The file named `my-kafka-service.json` contains the credentials in JSON format necessary for connecting to the Aiven for Apache Kafka service. * The Java code shows a dynamic approach to constructing the path to the credentials file, using the `myKafkaSource` argument provided during application execution. This approach allows flexibility and avoids using fixed the file name. * The application reads and parses the JSON file to extract configuration details such as `bootstrap_servers` (the Kafka server address) and `security_protocol` (the communication protocol for security). * These extracted details can be used to configure Apache Kafka sources and sinks in Apache Flink applications. This configuration establishes a connection to the Apache Kafka service and handles data according to the service's specifications. #### Configuration of `myKafkaSource` for application deployment[​](#configuration-of--mykafkasource-for-application-deployment "Direct link to configuration-of--mykafkasource-for-application-deployment") The `myKafkaSource` argument is a dynamic runtime parameter that specifies which Apache Kafka service your application connects to. This flexibility allows you to switch between different Kafka services without recompiling the JAR file each time. For example, you can replace a development Apache Kafka integration with a production Apache Kafka integration using the same JAR. ##### Deploy via Aiven Console[​](#deploy-via-aiven-console "Direct link to Deploy via Aiven Console") When deploying your JAR application using the [Aiven Console](https://console.aiven.io/), you can pass the `myKafkaSource` as **Program Arguments**. 1. [Create a JAR application](/docs/products/flink/howto/create-jar-application.md). 2. In the **Create new deployment** dialog look for the **Program args** field. 3. Insert the following syntax, replacing `integration_id` with the actual [integration ID](/docs/tools/cli/service/integration.md#avn_service_integration_list) of your Aiven for Apache Kafka service: ``` myKafkaSource= ``` ##### Deploy via Java code[​](#deploy-via-java-code "Direct link to Deploy via Java code") Using the `ParameterTool`, you can dynamically parse the `myKafkaSource` argument at runtime in your JAR application. This approach lets you configure the Apache Kafka service connection for your application without needing to recompile the JAR file. ``` ParameterTool parameters = ParameterTool.fromArgs(args); String myKafkaSource = parameters.getRequired("myKafkaSource"); ``` This code snippet shows how `ParameterTool` extracts the `myKafkaSource` value from the command-line arguments. The extracted `myKafkaSource` specifies the integration ID of the Apache Kafka service to be used by the application. Modifying this argument when running the JAR allows you to switch between different Apache Kafka services (For example, from a development environment to a production environment). ### Example 2: Integration with Aiven for PostgreSQL[​](#example-2-integration-with-aiven-for-postgresql "Direct link to Example 2: Integration with Aiven for PostgreSQL") To connect to Aiven for PostgreSQL service, the credentials JSON structure is as follows: ``` { "integration_type": "pg", "service_name": "my-pg-test-1", "url": "postgres://PG_USERNAME:PG_PASSWORD@my-pg-test-1-my-project.aivencloud.com:12691/defaultdb?sslmode=require" } ``` **Java application example:** ``` import java.net.URI; import java.net.URISyntaxException; public class PostgreSQLCredentialsParser { public static void main(String[] args) { // Assuming 'pgJson' is a JSONObject containing the credentials try { URI uri = new URI(pgJson.get("url").toString()); String[] userAndPassword = uri.getUserInfo().split(":"); String user = userAndPassword[0]; String password = userAndPassword[1]; String host = uri.getHost(); int port = uri.getPort(); String database = uri.getPath().substring(1); // Use these details for your PostgreSQL connection } catch (URISyntaxException e) { e.printStackTrace(); // Handle invalid JDBC URL syntax } } } ``` In this example, the application dynamically parses the PostgreSQL connection string to extract the necessary connection parameters, including username, password, host, port, and database name. ### Example 3: External or self-hosted Apache Kafka integration with SASL\_SSL[​](#example-3-external-or-self-hosted-apache-kafka-integration-with-sasl_ssl "Direct link to Example 3: External or self-hosted Apache Kafka integration with SASL_SSL") When integrating with an external Apache Kafka service using `SASL_SSL`, the credentials file is structured based on the configuration details you specify for your external Apache Kafka integration. The following JSON structure represents the minimum required for initial setup, serving as a foundational configuration. Additional parameters may be necessary depending on your security requirements. The credentials JSON structure is as follows: ``` { "bootstrap_servers": "external-kafka-server:port", "security_protocol": "SASL_SSL", "sasl_ssl": { "sasl_mechanism": "SCRAM-SHA-256", "sasl_password": "secure_password", "sasl_username": "kafka_user", "ssl_ca_cert": "certificate content" } } ``` note The structure of this JSON file corresponds to the configuration settings you established when creating an external Apache Kafka integration with Aiven for Apache Kafka service. This includes essential details such as bootstrap servers, security protocol, and SASL\_SSL settings. **Java application example:** ``` import org.json.simple.JSONObject; import org.json.simple.parser.JSONParser; import java.io.FileReader; public class ExternalKafkaCredentialsReader { public static void main(String[] args) { String credentialsFilePath = System.getenv("AVN_CREDENTIALS_DIR") + "/external-kafka-service.json"; JSONParser jsonParser = new JSONParser(); try (FileReader reader = new FileReader(credentialsFilePath)) { JSONObject kafkaCredentials = (JSONObject) jsonParser.parse(reader); JSONObject saslSslConfig = (JSONObject) kafkaCredentials.get("sasl_ssl"); // Configure Kafka properties for SASL_SSL Properties properties = new Properties(); properties.setProperty("bootstrap.servers", (String) kafkaCredentials.get("bootstrap_servers")); properties.setProperty("security.protocol", (String) kafkaCredentials.get("security_protocol")); properties.setProperty("sasl.mechanism", (String) saslSslConfig.get("sasl_mechanism")); properties.setProperty("ssl.truststore.certificates", (String) saslSslConfig.get("ssl_ca_cert")); properties.setProperty("sasl.jaas.config", String.format ("org.apache.kafka.common.security.scram.ScramLoginModule required username=\"%s\" password=\"%s\";", (String) saslSslConfig.get("sasl_username "), (String) saslSslConfig.get("sasl_password"))); // Additional Kafka configuration } catch (Exception e) { e.printStackTrace(); } } } ``` In this example: * The file named `external-kafka-service.json` contains credentials in JSON format for connecting to an external Kafka service that employs the SASL\_SSL security protocol. * The Java code demonstrates how to read the file, parse the JSON content, and extract key details, including SASL and SSL configurations. * These details are then used to configure a Kafka client within the Flink application, ensuring secure communication with the external Kafka service. ## Related page[​](#related-page "Direct link to Related page") * For additional information on integrating external or self-hosted Apache Kafka with Aiven for Apache Flink, see [Integrate Aiven for Apache Flink® with Apache Kafka®](/docs/products/flink/howto/ext-kafka-flink-integration.md) --- # Manage Aiven for Apache Flink® applications This section provides information on managing your Aiven for Apache Flink® applications. ## Creating a new version of an application[​](#creating-a-new-version-of-an-application "Direct link to Creating a new version of an application") To create a version of the application deployed: 1. Log in to the [Aiven Console](https://console.aiven.io/), and select your Aiven for Apache Flink® service. 2. From the left sidebar, select **Applications**. 3. On the **Applications** landing page, click the application name for which to create a version. ### For an SQL application[​](#for-an-sql-application "Direct link to For an SQL application") 1. Click **Create new version**. 2. In the **Create new version** page, modify the create statement, source, or sink tables as needed. 3. Click **Save and deploy later**. You can see the new version listed in the versions drop-down list. 4. To deploy the new version of the application, [stop](/docs/products/flink/howto/manage-flink-applications.md#stop-flink-application) any existing version that is running. 5. Click **Create deployment**, and in the **Create new deployment** dialog: * Select the version to deploy. * Select the savepoint from where to deploy. * Toggle **Restart on failure** to automatically restart Flink jobs upon failure. * Enter the number of [parallel instances](https://nightlies.apache.org/flink/flink-docs-master/docs/dev/datastream/execution/parallel/) for the task. * Click **Deploy from a savepoint** or **Deploy without savepoint** depending on your previous selection. ### For a JAR application[​](#for-a-jar-application "Direct link to For a JAR application") 1. Click **Upload new version**. 2. In the **Upload new version** dialog: * Click **Choose file** to select your custom JAR file. * Review and accept the terms of service by checking the box. * Click **Upload version** to upload your JAR file. 3. In the **Deployment history** you can see the latest version running. ## Stop an application deployment[​](#stop-flink-application "Direct link to Stop an application deployment") To stop a deployment for your Flink application: 1. In your Aiven for Apache Flink service, select **Applications** from the left sidebar. 2. On the **Applications** landing page, click the application name to be stopped. 3. In the application's overview page, click **Stop deployment**. 4. In the **Stop deployment** dialog, enable the option to **Create a savepoint before stopping** to save the current state of the application. To stop a deployment without saving the current state of the application, disable the option for **Create a savepoint before stopping** and click **Stop without creating savepoint**. 5. Click **Create savepoint & stop** to initiate the stopping process. The application status will display `Saving_and_stop_requested` and `Finished` once the stopping process is completed. Additionally, the **Deployment history** provides a record of all the application deployments and statuses. ## Rename an application[​](#rename-an-application "Direct link to Rename an application") To rename an application: 1. In your Aiven for Apache Flink service, select **Applications** from the left sidebar. 2. On the **Applications** landing page, click the application name to rename. 3. In the application's overview page, click the **Application action menu (...)**, and click **Update application** from the menu options. 4. In the **Update Application** dialog, enter the new name for the application and select **Save changes** to confirm the new name and update the application. ## Accessing deployment history[​](#flink-deployment-history "Direct link to Accessing deployment history") The **Deployment History** screen provides the following: * A list of all the deployments for an application * The user who created the application (created by) * Data and time of creation (created at) * Application version * If a savepoint was created or not To view and delete the deployment history of an application: 1. In your Aiven for Apache Flink service, select **Applications** from the left sidebar. 2. On the **Applications** landing page, click the application name for which to view the deployment history. 3. In the application landing page, click **Deployment History** to view the deployment history. 4. To remove a specific deployment from the history, locate it in the deployment history page and click the **Delete** icon next to it. ## Delete an application[​](#delete-an-application "Direct link to Delete an application") Before deleting an application, it is necessary to remove all associated [deployment history](/docs/products/flink/howto/manage-flink-applications.md#flink-deployment-history). 1. In your Aiven for Apache Flink service, select **Applications** from the left sidebar. 2. On the **Applications** landing page, click the application name to delete. 3. In the application's overview page, click the **Application action menu (...)**, and click **Delete application** from the menu options. 4. In the **Delete Confirmation** dialog, enter the name of the application and click **Confirm** to proceed with the deletion. --- # Manage tables in Aiven for Apache Flink® applications Aiven for Apache Flink® allows you to map source and target data structures as [Flink tables](https://nightlies.apache.org/flink/flink-docs-stable/docs/dev/table/sql/create/#create-table) and use transformation statements to reshape, filter or aggregate data. Some of the table operations include: * Import existing tables * Create tables * Clone table definitions from other applications * Edit tables * Delete tables important Before performing any operation on a table in a Flink application, you must **stop** the application. To stop an application, go to the **Applications** from the left sidebar on your Aiven for Apache Flink® service, select the desired application from the list, and select **Stop Deployment**. ## Add a new table[​](#add-a-new-table "Direct link to Add a new table") To add a new table to an application using the [Aiven Console](https://console.aiven.io/): 1. Select **Applications** from the left sidebar on your Aiven for Apache Flink service, and select the application to which you want to add a new table. Make sure the application deployment is stopped. 2. Select **Create new version**. 3. On the **Create new version** screen, go to the **Add source tables** or **Add sink tables** screen within your application. 4. Select **Add new table** to add a new table to your application. note If you already have a sink table listed, you must delete it before adding a new one, only one sink table is allowed per job. 5. Select the **Integrated service** from the drop-down list in the **Add new source table** or **Add new sink table** screen, respectively. 6. In the **Table SQL** section, enter the statement that will create the table. The interactive query feature if the editor will prompt you for error or invalid queries. 7. Select **Add table** to complete the process. ## Import an existing table[​](#import-an-existing-table "Direct link to Import an existing table") To import an existing table from another application: 1. In the **Add source tables** or **Add sink tables** screen, select **Import existing table** to import a table to your application. note If you already have a sink table listed, you must delete it before importing a new one. 2. From the **Import existing source table** or **Import existing sink table** screen: * Select the application from which to import the table. * Select the version of the application. * Select the table to import. 3. Select **Next**. 4. Verify the data on the **Add new source table** or **Add new sink table** screen and select **Add table** to complete the process. ## Clone a table[​](#clone-a-table "Direct link to Clone a table") To clone a table within an application: 1. In the **Add source tables** screen, locate the table to clone and click **Clone** next to it. note Clone option is not available sink tables. 2. Select the **Integrated service** from the drop-down list. 3. In the **Table SQL** section, update the table name. note You will not be able to add the table if there are errors within the statement. 4. Select **Add table** to complete the process. ## Edit a table[​](#edit-a-table "Direct link to Edit a table") To edit an existing table in an application: 1. In the **Add source tables** or **Add sink tables** screen, locate the table to edit and click **Edit** next to it. 2. Make the necessary changes to the table and select **Save changes** to confirm the changes. ## Delete a table[​](#delete-a-table "Direct link to Delete a table") To delete a table in an application: 1. In the **Add source tables** or **Add sink tables** screen, locate the table to delete and click the **Delete** icon next to it. 2. Confirm the deletion by selecting **Confirm** in the pop-up window. --- # Create a PostgreSQL® CDC connector-based Apache Flink® Change Data Capture (CDC) is a technique that enables the tracking and capturing of changes made to data within a PostgreSQL® database. By identifying and capturing changes at the granular row level, CDC enables applications to react and promptly process these changes in real time. This ensures up-to-date data accuracy and minimizes the latency involved in processing updates. When using Aiven for Apache Flink®, you can seamlessly leverage the power of the PostgreSQL CDC Connector to stream and process real-time data changes from PostgreSQL databases. Integrated with the Debezium engine, the CDC connector captures changes at the granular level of each table within an event stream. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before setting up the PostgreSQL CDC Connector with Aiven for Apache Flink, ensure that you have the following prerequisites in place: * Running Aiven for Apache Flink® Service * Running Aiven for PostgreSQL® Service * Running Aiven for Apache Kafka® Service or any other service of your choice as the sink for your data. * [Integration](/docs/products/flink/howto/create-integration.md) between Aiven for Flink and Aiven for PostgreSQL: Establish the necessary integration between your Aiven for Apache Flink service and Aiven for PostgreSQL service. In addition to the above, gather the following information about the source PostgreSQL database: * `Hostname`: The hostname or IP address of the PostgreSQL server where your source database is located. * `Port`: The port number on which the PostgreSQL server is listening for incoming connections. * `Database name`: The name of the source database within the PostgreSQL server. * `Username`: The username used to authenticate and access the PostgreSQL database. * `Password`: The password associated with the provided username for authentication. * `Schema name`: The schema name where the source table is located within the database. * `Table name`: The name of the source table from which to capture data changes. * `Decoding plugin name`: The decoding plugin name to use for capturing the changes. For PostgreSQL CDC, set it as `pgoutput`. important To create a PostgreSQL CDC source connector in Aiven for Apache Flink with Aiven for PostgreSQL using the pgoutput plugin, you need superuser privileges. For more information, see [Troubleshooting](#Troubleshooting). ## Configure the PostgreSQL CDC connector[​](#configure-the-postgresql-cdc-connector "Direct link to Configure the PostgreSQL CDC connector") From the [Aiven Console](https://console.aiven.io/): 1. In the Aiven for Apache Flink service page, select **Application** from the left sidebar. 2. Select **Create new application**, enter the name of your application, and select **Create application**. note If editing an existing application, create a version to make changes to the source or sink tables. 3. Select **Create first version** to create the first version of the application. 4. Select **Add your first source table** to add a source table. 5. In the **Add new source table** screen, select the *Aiven for PostgreSQL®* service as the integrated service. 6. In the **Table SQL** section, enter the SQL statement to create the PostgreSQL-based Apache Flink table with CDC connector. For example: ``` CREATE TABLE test_table ( column1 INT, column2 VARCHAR ) WITH ( 'connector' = 'postgres-cdc', 'hostname' = 'test-project-test.avns.net', 'username' = 'username', 'password' = '12345', 'schema-name' = 'public', 'table-name' = 'test-1', 'port' = '12709', 'database-name' = 'defaultdb', 'decoding.plugin.name' = 'pgoutput' ) ``` Where: * `connector`: The connector type to be used, which is `postgres-cdc` in this case. * `hostname`: The hostname or address of the PostgreSQL database server. * `username`: The username to authenticate with the PostgreSQL database. * `password`: The password for the provided username. * `schema-name`: The name of the schema where the source table is located, which is set to `public` in the example. * `table-name`: The name of the source table to be captured by the CDC connector, which is set to `test-1` in the example. * `port`: The port number of the PostgreSQL database server. * `database-name`: The name of the database where the source table resides, which is set to `defaultdb` in the example. * `decoding.plugin.name`: The decoding plugin to be used by the CDC connector, which is set to `pgoutput` in the example. note The PostgreSQL CDC connector will use or create a publication named `dbz_publication` tracking the changes of one or more tables. Therefore, the publication must already exist in PostgreSQL, or the connecting user must have enough privileges to create it. 7. Select **Next** to add the sink table, and select **Add your first sink table**. Select *Aiven for Apache Kafka®* as the integrated service from the drop-down list. 8. In the **Table SQL** section, input the SQL statement for creating the sink table where the PostgreSQL CDC connector will send the data. Select **Add table**. 9. In the **Create statement** section, write the SQL schema that defines the fields retrieved from the PostgreSQL® table and any additional transformations. 10. Select **Create deployment** to deploy the application, and in the **Create new deployment** screen, choose the desired version to deploy (default: Version 1) and select **Deploy without a savepoint** (as there are no savepoints available for the first application). ## Troubleshooting[​](#Troubleshooting "Direct link to Troubleshooting") If you encounter the `must be superuser to create FOR ALL TABLES publication` error when setting up a PostgreSQL CDC source connector in Aiven for PostgreSQL using the `pgoutput` plugin: 1. Install the `aiven-extras` extension by executing the SQL command: ``` CREATE EXTENSION aiven_extras CASCADE; ``` 2. Create a publication for all tables in the source database: Execute the SQL command: ``` SELECT * FROM aiven_extras.pg_create_publication_for_all_tables( 'dbz_publication', 'INSERT,UPDATE,DELETE' ); ``` note The publication name must be `dbz_publication` for the PostgreSQL CDC connector to work --- # Restart strategy in SQL and JAR applications Learn how Aiven for Apache Flink® applications uses restart strategies to recover from job failures, ensuring high availability and fault tolerance for your distributed applications. ## About restart strategy[​](#about-restart-strategy "Direct link to About restart strategy") A restart strategy is a set of rules that Apache Flink® adheres to when dealing with Flink job failures. These strategies enable the automatic restart of a failed Flink job under specific conditions and parameters, which is crucial for high availability and fault tolerance in distributed and scalable systems. Aiven for Apache Flink® supports restart strategies for both JAR and SQL applications. ## Default restart strategy for JAR and SQL applications[​](#default-restart-strategy-for-jar-and-sql-applications "Direct link to Default restart strategy for JAR and SQL applications") Aiven for Apache Flink® uses the **exponential-delay** strategy as the default restart mechanism for both **JAR and SQL applications**. This strategy incrementally increases the delay between restarts, reaching a configurable maximum. After reaching the maximum, the delay remains constant. The strategy resets the exponential delay after a period of successful restarts, preventing it from staying at the maximum indefinitely. This default strategy is integrated into the Aiven for Apache Flink cluster configuration and automatically applies to JAR and SQL applications. ## View the default strategy[​](#view-the-default-strategy "Direct link to View the default strategy") You can view the default restart strategy configurations for your Aiven for Apache Flink cluster in the Apache Flink Dashboard. Follow these steps to view the current settings: 1. Access the [Aiven Console](https://console.aiven.io/) and select the Aiven for Apache Flink service. 2. From the **Connection information** section on the overview page, copy the **Service URI** and paste it into your web browser's address bar. 3. When prompted, log in using the **User** and **Password** credentials specified in the **Connection information** section. 4. Once in the **Apache Flink Dashboard**, click the **Job Manager** from the menu. 5. Switch to the **Configuration** tab. 6. Review the configurations and parameters related to the restart strategy. ## Disable default restart strategy[​](#disable-default-restart-strategy "Direct link to Disable default restart strategy") While Aiven for Apache Flink® typically recommends following the default restart strategy for high availability and fault tolerance, there might be scenarios, especially during testing or debugging, where disabling automatic restarts might be beneficial. ### JAR applications[​](#jar-applications "Direct link to JAR applications") For JAR applications, you cannot disable the default restart strategy in Aiven for Apache Flink® through configuration files. Instead, directly modify the code of your Jar application to achieve this. ``` StreamExecutionEnvironment env = StreamExecutionEnvironment.getExecutionEnvironment(); env.setRestartStrategy(RestartStrategies.noRestart()); ``` This code sets the restart strategy to `None`, preventing any restart attempts in case of failures. ### SQL applications[​](#sql-applications "Direct link to SQL applications") For SQL applications, you have a simplified approach to restart strategies. You can enable or disable restarts on failure during application deployment, providing a straightforward way to manage applications without complex configurations. For more information, see [Create an SQL application](/docs/products/flink/howto/create-sql-application.md). ## Key considerations when disabling default restarts[​](#key-considerations-when-disabling-default-restarts "Direct link to Key considerations when disabling default restarts") Before disabling the default restart strategy for your applications, consider the following: * **Persistent failures**: Disabling restarts means that if a Flink Job fails, it will not attempt to recover it, leading to permanent job failure. * **Testing and debugging**: Disabling is beneficial when identifying issues in the application code, as it prevents the masking of errors through automatic restarts. * **External factors**: Jobs can fail due to external factors, such as infrastructure changes or maintenance activities. If you disable restarts, your Flink jobs will become vulnerable to failures. * **Operational risks**: In production environments, it is generally advisable to use the default restart strategy to ensure high availability and fault tolerance. Related pages * [Restart strategies in Apache Flink®](https://nightlies.apache.org/flink/flink-docs-release-1.18/docs/ops/state/task_failure_recovery/#restart-strategies) --- # Create a Slack-based Apache Flink® table With Aiven's Slack Connector for Apache Flink®, you can create sink tables in your Flink application and set up statements to send alerts and notifications to your designated Slack channel. This allows for real-time monitoring and tracking of your Flink job progress and any potential issues that may arise. You can access the open-source Slack connect for Apache Flink on Aiven's GitHub repository at [Slack Connector for Apache Flink®](https://github.com/aiven/slack-connector-for-apache-flink). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Slack app created and ready for use. For more information, refer to the [Set-up Slack Application section](https://github.com/aiven/slack-connector-for-apache-flink) on the GitHub repository and the [Slack documentation](https://api.slack.com/start/apps). * Note the **channel ID** and **token value**, as these will be required in the sink connector Table SQL when configuring the connection in your Flink application. ## Configure Slack as sink for Flink application[​](#configure-slack-as-sink-for-flink-application "Direct link to Configure Slack as sink for Flink application") To configure Slack as the target using the Slack connector for Apache Flink: 1. In the Aiven for Apache Flink service page, select **Application** from the left sidebar. 2. Create a application or select an existing application for your desired [data service integration](/docs/products/flink/howto/create-integration.md). note If you are editing an existing application, create new version of the application to make changes to the source or sink table. 3. In the **Add source tables** screen, select the option to add a new source table, edit an existing one, or import a source table. Select your integrated service and in the **Table SQL** section, enter the statement that will be used to create the table. 4. In the **Add sink tables** screen, select the option to add a new sink table or edit an existing one. 5. In the **Table SQL** section, set the connector to **slack** and enter the necessary token as shown in the example below: ``` CREATE TABLE channel_name ( channel_id STRING, message STRING ) WITH ( 'connector' = 'slack', 'token' = '' ) ``` note \* `channel_id` should contain the channel ID parameter. \* Replace the `\` placeholder with the token you created in Slack during the prerequisites. 6. Create the SQL statement to send notifications/alerts to the designated slack channel. You can check the connection by running a query on the sink table. If the data flows into the slack channel, the connection is successful. --- # Define OpenSearch® timestamp data in SQL pipeline Frequently results in Apache Flink® data pipelines include one or more timestamps, either contained in the [source events](/docs/products/flink/concepts/event-processing-time.md) or generated by [window aggregations](/docs/products/flink/concepts/windows.md). When the output of the Apache Flink® data pipeline is an Aiven for OpenSearch® index, convert the Flink timestamps to a format recognizable by OpenSearch®, otherwise they will be interpreted as strings, losing the benefits of time filtering. OpenSearch® recognises the following as correct date/time formats: * `yyyy/MM/dd` for a **date** field * `HH:mm:ss` for a **time** field * `yyyy/MM/dd HH:mm:ss` for a **timestamp** field Structure the data pipeline output to follow one of the recognized formats. ## Define Apache Flink® target tables including timestamps for OpenSearch®[​](#define-apache-flink-target-tables-including-timestamps-for-opensearch "Direct link to Define Apache Flink® target tables including timestamps for OpenSearch®") When the result of the data pipeline contains a timestamp column like the below: ``` EVENT_TIME TIMESTAMP(3), HOSTNAME STRING, CPU DOUBLE ``` To push the data correctly to an OpenSearch® index, you'll need to set the target column format as `STRING` in the Flink table definition, like: ``` EVENT_TIME STRING, HOSTNAME STRING, CPU DOUBLE ``` Assuming the `EVENT_TIME` is a timestamp, you'll need to specify it in the format understood by OpenSearch® using the `DATE_FORMAT` function, like: ``` DATE_FORMAT(EVENT_TIME, 'yyyy/MM/dd HH:mm:ss') ``` Once the pipeline is running, you can check that the `EVENT_TIME` field in OpenSearch® is recognized as a timestamp. --- # Upgrade Aiven for Apache Flink Upgrading to the latest version of Aiven for Apache Flink® allows you to benefit from improved features, enhanced performance, and better security. ## Limitations[​](#limitations "Direct link to Limitations") * Direct upgrades on active Aiven for Apache Flink services are not supported. * Upgrading to Apache Flink 1.19 requires creating a new service and manually transferring applications. ## Migrate to a newer Apache Flink version[​](#migrate-to-a-newer-apache-flink-version "Direct link to Migrate to a newer Apache Flink version") If you are using Aiven for Apache Flink version 1.16, which will soon reach [end-of-life](/docs/platform/reference/eol-for-major-versions.md#aiven-for-flink), create new service for version 1.19 and migrate your applications to continue receiving support and take advantage of the latest enhancements. ### Step 1: Create an Aiven for Apache Flink service[​](#step-1-create-an-aiven-for-apache-flink-service "Direct link to Step 1: Create an Aiven for Apache Flink service") 1. Create an Aiven for Apache Flink service using the [Aiven Console](https://console.aiven.io/), [Aiven CLI](/docs/tools/cli/service/flink.md), or [Aiven Provider for Terraform](/docs/tools/terraform.md) (provider version 4.19.0). 2. Select Apache Flink 1.19 as the deployment version for the Aiven for Apache Flink service. ### Step 2: Transfer your applications[​](#step-2-transfer-your-applications "Direct link to Step 2: Transfer your applications") * Manual * Using Terraform For manually created applications: 1. Stop the application deployment in the existing Aiven for Apache Flink service. 2. Copy the source, sinks, and transformation statements from the existing service for your applications. 3. Create applications in the new service and add the copied source, sinks, and transformation statements. 4. Deploy the applications in the new service. For more information, see [Aiven for Apache Flink applications](/docs/products/flink/howto/create-flink-applications.md). If you used Terraform (ensure you're on provider version 4.19.0) to create your service and applications: 1. Update the `flink_version` property in your Terraform script to 1.19. 2. Run Terraform to create Aiven for Apache Flink service with the updated version and applications. Terraform will automatically handle the recreation and deployment of your applications. For more information, see [`aiven_flink`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/flink). note For information on breaking changes between versions, refer to the official [Apache Flink release notes](https://nightlies.apache.org/flink/flink-docs-release-1.19/release-notes/flink-1.19/). ### Step 3: Verify and power off the service[​](#step-3-verify-and-power-off-the-service "Direct link to Step 3: Verify and power off the service") To complete the migration: 1. Verify that the applications are operating correctly in the new Aiven for Apache Flink service. 2. Power off and delete the old service after confirming the new service is running. --- # Advanced parameters for Aiven for Apache Flink® See the configuration options available for Aiven for Apache Flink®: | Parameter | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | []()[**ip\_filter**](#ip_filter)`array`- default: `0.0.0.0/0,::/0`IP filterAllow incoming connections from CIDR address block, e.g. '10.20.0.0/16' | | []()[**service\_log**](#service_log)`boolean,null`Service loggingStore logs for the service so that they are available in the HTTP API and console. | | []()[**static\_ips**](#static_ips)`boolean`Use static public IP addresses | | []()[**custom\_code**](#custom_code)`boolean`Custom code enabledEnable to upload Custom JARs for Flink applications | | []()[**flink\_version**](#flink_version)`string,null`Flink major version | | []()[**number\_of\_task\_slots**](#number_of_task_slots)`integer`- min: `1`
- max: `1024`Flink taskmanager.numberOfTaskSlotsTask slots per node. For a 3 node plan, total number of task slots is 3x this value | | []()[**pekko\_ask\_timeout\_s**](#pekko_ask_timeout_s)`integer`- min: `5`
- max: `60`Flink pekko.ask.timeoutTimeout in seconds used for all futures and blocking Pekko requests | | []()[**pekko\_framesize\_b**](#pekko_framesize_b)`integer`- min: `1048576`
- max: `52428800`Flink pekko.framesizeMaximum size in bytes for messages exchanged between the JobManager and the TaskManagers | | []()[**privatelink\_access**](#privatelink_access)`object`Allow access to selected service components through Privatelinkprivatelink\_access.flink boolean Enable flink privatelink\_access.prometheus boolean Enable prometheus | | []()[**public\_access**](#public_access)`object`Allow access to selected service ports from the public Internetpublic\_access.flink boolean Allow clients to connect to flink from the public internet for service nodes that are in a project VPC or another type of private network | --- # Aiven for Apache Flink® limitation Because Aiven for Apache Flink is a fully managed service, there are differences between Aiven for Apache Flink and an Apache Flink service you run yourself. Consider the following differences and limitations: * **User-defined functions:** Aiven for Apache Flink does not currently support using user-defined functions (UDFs). * **Apache Flink CLI tool:** The Apache Flink® CLI tool is not currently supported as it requires access to the JobManager in production, which is not exposed to customers. * **Job-level settings:** In Aiven for Apache Flink, each job inherits the cluster-level settings, and job-level settings are not yet supported. You cannot specify separate settings for individual jobs within the same cluster. * **Flame graphs:** Flame graphs are marked as an experimental feature in Apache Flink 1.15 and are not currently enabled in the Aiven for Apache Flink web UI. --- # Aiven for Grafana® Aiven for Grafana® is a fully managed analytics and monitoring solution, deployable in the cloud of your choice, which can bring unlimited scalability and high availability to your monitoring environment and other time series applications. ## Main features[​](#main-features "Direct link to Main features") ### Quick and flexible deployment options[​](#quick-and-flexible-deployment-options "Direct link to Quick and flexible deployment options") With Aiven for Grafana, you can enjoy a quick and flexible deployment process, ensuring production-ready Grafana clusters are available in 10 minutes. You have the flexibility to choose your preferred public cloud platform for deployment from over 100 regions supported. The deployment process also includes high-performance nodes to enhance performance. Aiven supports the Bring-Your-Own-Cloud (BYOC) deployment model, enabling you to meet strict control requirements. ### Integrate with existing Aiven tools and data infrastructure[​](#integrate-with-existing-aiven-tools-and-data-infrastructure "Direct link to Integrate with existing Aiven tools and data infrastructure") Aiven for Grafana seamlessly integrates various services and your existing Aiven tools. You can set up your Aiven services as data sources for Grafana and monitor their health in real time. This integration allows you to leverage the full potential of your data infrastructure, making it easier to manage and monitor. In addition to integration with Aiven services, Aiven also offers pre-built dashboards that allow you to monitor the health of your Aiven services. These dashboards provide valuable insights into the performance of your infrastructure, making it easier to identify and address potential issues. ### Secure network connectivity[​](#secure-network-connectivity "Direct link to Secure network connectivity") Aiven offers various options for secure network connectivity, including [VPC peering](/docs/platform/howto/manage-project-vpc.md), [PrivateLink](/docs/tools/cli/service/privatelink.md), and TransitGateway technologies. These options allow you to securely connect your data infrastructure to Aiven services and ensure the highest level of data security. ### Log management[​](#log-management "Direct link to Log management") You can also integrate with popular logs management solutions such as [OpenSearch®](/docs/products/opensearch/howto/opensearch-log-integration.md) and Google Cloud Logging, which help monitor your system performance. ### Monitoring with dashboards, plugins, and alerting[​](#monitoring-with-dashboards-plugins-and-alerting "Direct link to Monitoring with dashboards, plugins, and alerting") * **Dashboards and plugins:** With Aiven for Grafana, you can monitor your data using ready-made dashboards. Benefit from over 60 advanced panel and [data source plugins](/docs/products/grafana/reference/plugins.md) that provide a high level of customization to your monitoring solution. * **Monitoring and alerting:** Aiven for Grafana allows you to create monitoring solutions for all teams, enabling everyone to keep track of critical metrics. Additionally, you can implement an observability platform with alerting, ensuring that you are promptly notified in case of any issues. ### Automation[​](#automation "Direct link to Automation") You can also automate the process of building, configuring, and managing Aiven services using the [Aiven Provider for Terraform](/docs/tools/terraform.md). Related pages * [Plans and pricing](/docs/products/grafana.md) --- # Backups and migration in Aiven for Grafana® Protect and recover your Aiven for Grafana® service data with cross-region backups and restore tracking. Related pages * [Backup to another region](/docs/products/grafana/howto/backup-to-another-region.md) * [Track restore progress](/docs/products/grafana/howto/track-restore-progress.md) --- # Memory and out-of-memory conditions in Aiven for Grafana® Understand the memory limits and out-of-memory conditions that apply to your Aiven for Grafana® service. ## Service memory limits[​](#service-memory-limits "Direct link to Service memory limits") The practical memory limit will always be less than the service physical memory limit. **All services are subject to operating overhead:** * A small amount of memory is required by the operating system kernel to manage system resources, including networking functions and disk cache. * Aiven's cloud data platform requires memory to monitor availability, provide metrics, logging and manage backups. A server or node's **usable memory** can be calculated as: `usable memory = RAM - overhead` Where: * `overhead` is 350 MiB (≈ 0.34 GiB). Services may utilize optional components, service integrations, connection pooling, or plug-ins, which are not included in overhead calculations. If a service is overcommitted, the operating system, management layer, backups or availability monitoring, may fail status checks or operations due to resource contention. In severe instances, the node may fail completely with an out-of-memory condition. ## Out of memory conditions[​](#out-of-memory-conditions "Direct link to Out of memory conditions") Many processes request more memory from the kernel than they will ever use or need. In these cases, the kernel overallocates memory. This allows it to satisfy multiple processes requesting more memory than is available, which is not used or is freed by the time any other process actually needs it. However, if enough processes start using all their allocated memory simultaneously there may not be enough physical memory available and an `Out Of Memory` (`OOM`) condition occurs. warning This situation is critical and must be resolved immediately. The solution that the Linux kernel employs is to invoke the `Out of Memory Killer` (or `OOM Killer`). This reviews all running processes and kills one or more of them to free up system memory and keep the system running. The `OOM Killer` selects process to kill based on an `oom_score`; a calculation that balances how much memory the process is using with how long the process has been running. Processes that have been running for a long time are less likely to be killed. Subprocesses are summed with parent processes in terms of memory usage, so a process which forks many subprocesses, but itself does not use a lot of memory, may still be killed. In most instances, the hosted data service, or a child process, will have the highest memory footprint and be a prime candidate for termination when the OOM Killer inspects the running processes. Aiven's cloud data platform leverages kernel namespaces or containers to isolate processes from each other. Isolation has several benefits, including: * A smaller footprint for security‑related concerns * A smaller blast radius for failure * Greater control of system resources Left unchecked, the `OOM Killer` may opt to kill the primary service. This is undesirable as unclean termination of the primary service can lead to data loss, inconsistency, or corrupted backups. Further, if Aiven's management platform detects that the primary service is unavailable for , the service will be marked as down and a failover will occur. To mitigate this scenario, namespaces are used, some with additional memory limits, in combination with an `oom_score_adjust` on the primary process, to coax the `OOM Killer` into selection of less critical processes. This will still result in a service restart, but in a more controlled process, where the database is shut down, rather than killed; exposure to data loss is limited and recovery is faster when the service restarts, often avoiding failover. warning Out of Memory conditions can still lead to unexpected behavior, including data unavailable or data loss conditions. ## Avoid running low on memory[​](#avoid-running-low-on-memory "Direct link to Avoid running low on memory") The OOM killer only runs when the system is critically low on memory. To prevent it from running, either reduce your memory usage or increase the available memory. For most databases, the service memory footprint can often be reduced by: * Reducing concurrency or implementing connection pooling * Tuning queries to limit result sets * Tuning indexes for query load * Dropping unused objects from storage In cases where the working set no longer fits into memory, consider scaling your service. Related pages * [Change the service plan](/docs/products/grafana/howto/change-service-plan.md) * [Prepare for high load](/docs/products/grafana/howto/prepare-for-high-load.md) --- # Get started with Aiven for Grafana® To start using Aiven for Grafana, the first step is to create a service. You can do this in the [Aiven Console](https://console.aiven.io/) or with the [Aiven CLI](https://github.com/aiven/aiven-client). ## Create an Aiven for Grafana service[​](#create-an-aiven-for-grafana-service "Direct link to Create an Aiven for Grafana service") * Console * Terraform 1. In your project, click **Services**. 2. Click **Create service**. 3. Select **Grafana®**. 4. Select a **Cloud**. 5. Select a **Plan**. note The plans available can vary between cloud providers and regions for the same service. 6. In the **Service details**, enter a name for your service. 7. Optional: Add service tags. 8. In the **Service summary**, click **Create service**. The status of the service is **Rebuilding** during its creation. When the status is **Running**, you can start using the service. This typically takes a couple of minutes and can vary between cloud providers and regions. The following example files create a Grafana service in your Aiven project. They are part of the Grafana example in the [Aiven Terraform Provider repository](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/grafana) on GitHub. 1. Create a file named `provider.tf` and add the following: ``` Loading... ``` 2. Create a file named `service.tf` and add the following: ``` Loading... ``` 3. Create a file named `variables.tf` and add the following: ``` Loading... ``` 4. Create the `terraform.tfvars` file and add the values for your token and project name. To apply your Terraform configuration: 1. Initialize Terraform by running: ``` terraform init ``` The output is similar to the following: ``` Initializing the backend... Initializing provider plugins... - Finding aiven/aiven versions matching ">= 4.0.0, < 5.0.0"... - Installing aiven/aiven v4.9.2... - Installed aiven/aiven v4.9.2 ... Terraform has been successfully initialized! ... ``` 2. To create an execution plan and preview the changes, run: ``` terraform plan ``` 3. To deploy your changes, run: ``` terraform apply --auto-approve ``` ## Log in to Grafana[​](#log-in-to-grafana "Direct link to Log in to Grafana") After starting the Aiven for Grafana service, you can access Grafana: 1. From the [Aiven Console](https://console.aiven.io/), access your Aiven for Grafana service. 2. In the service **Overview** screen, copy or select the **Service URI** to launch the Grafana login page in a browser. 3. On the login page, enter or copy and paste the **User** and **Password** details from the *Connection information* section, and select **Log in**. You can begin visualizing your data sources using the default dashboards or create your own. tip Your Aiven for OpenSearch® services in the same Aiven project are automatically configured as data sources, so you can start visualizing their data right away. ## Grafana resources[​](#grafana-resources "Direct link to Grafana resources") * [Open source Grafana page](https://grafana.com/oss/grafana/) * [Grafana docs](https://grafana.com/docs/) * [Aiven Terraform Provider - Grafana resource docs](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/grafana) and [Grafana data source docs](https://registry.terraform.io/providers/aiven/aiven/latest/docs/data-sources/grafana) --- # Back up your Aiven for Grafana® service to another region Copy your Aiven for Grafana® service backups to a secondary region for disaster recovery. In addition to the primary service backup, you can have a secondary backup in an alternative location. important This feature is in [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-). Contact your account team to enable it. Backup to another region (BTAR) is a disaster recovery feature that allows backup files to be copied from the service's primary backup region to an additional (secondary) region. BTAR can bolster data resilience and helps improve data protection against disasters in the primary backup region. When the primary region is down, BTAR allows forking the service from an additional copy of the backup residing in a secondary region. ## Limitations[​](#limitations "Direct link to Limitations") * The cloud provider for your additional backup region must match the cloud provider for your service and the primary backup. * Secondary backup can only be restored in the region where it was stored. For a service that has the backup to another region (BTAR) feature enabled, you can check the service backup status, change the backup region, monitor the replication lag, fork and restore using the cross-region backup, or migrate to another cloud or region. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * At least one Aiven service with BTAR enabled * Access to the [Aiven Console](https://console.aiven.io/) * [Aiven API](/docs/tools/api.md) * [Aiven CLI](/docs/tools/cli.md) ## Change a backup region[​](#change-a-backup-region "Direct link to Change a backup region") 1. Log in to the [Aiven Console](https://console.aiven.io/) and go to your project. 2. On the **Services** page, select an Aiven service on which you'd like to enable BTAR. 3. On your service page, **Backups**. 4. On the **Backups** page, click **Actions** > **Edit secondary backup location**. 5. In the **Edit secondary backup location** window, use the **Backup location** menu to select a region for your additional backup. Confirm your choice by selecting **Save**. important You can change the backup region once in 24 hours. ## Monitor a service with BTAR[​](#monitor-a-service-with-btar "Direct link to Monitor a service with BTAR") There are a few things you may want to check for your Aiven service in the context of BTAR: * What is the status of a secondary backup? * Does your service have a backup in another region? * What is the target region of the secondary backup? * What is the replication lag between data availability in the primary region and the secondary region? ### Check BTAR status[​](#check-btar-status "Direct link to Check BTAR status") To see the availability, the status, and the target region of a secondary (BTAR) backup in the [Aiven Console](https://console.aiven.io/), go to your service page > **Backups** > **Secondary backup location**. ### Determine replication lag[​](#determine-replication-lag "Direct link to Determine replication lag") Determine the target region and the replication lag for a secondary (BTAR) backup of your service, call the [ServiceBackupToAnotherRegionReport](https://api.aiven.io/doc/#tag/Service/operation/ServiceBackupToAnotherRegionReport) endpoint. Configure the call as follows: 1. Enter `YOUR-PROJECT-NAME` and `YOUR-SERVICE-NAME` into the URL. 2. Specify `DESIRED-TIME-PERIOD` depending on the time period you need the metrics for: select one of the following values for the `period` key: `hour`, `day`, `week`, `month`, or `year`. ``` curl --request POST \ --url https://api.aiven.io/v1/project/YOUR-PROJECT-NAME/service/YOUR-SERVICE-NAME/backup_to_another_region/report \ --header 'Authorization: Bearer YOUR-BEARER-TOKEN' \ --header 'content-type: application/json' \ --data '{"period":"DESIRED-TIME-PERIOD"}' ``` As output, you get metrics including replication lags at specific points in time. ## Fork and restore a service with BTAR[​](#fork-and-restore "Direct link to Fork and restore a service with BTAR") You can use the [Aiven Console](/docs/tools/aiven-console.md), [API](/docs/tools/api.md), or [CLI](/docs/tools/cli.md) to recover your service from a backup in another region. To restore your service using BTAR, create a fork of the original service in the region where the secondary backup resides. note When you **fork & restore** from the secondary backup, your new fork service is created in the cloud and region where the secondary backup is located. The fork service gets the same plan that the primary service uses. Backups of the fork service are located in the region where this new service is hosted. * Aiven Console * Aiven CLI * Aiven API 1. Open the [Aiven Console](https://console.aiven.io/) and go to your service homepage. 2. Click **Backups**. 3. On the **Backups** page, select **Fork & restore**. 4. In the **New database fork** window: 1. Set **Backup location** to either **Primary location** or **Secondary location**. 2. Set **Backup version** to one of the following: * **Latest transaction** * **Point in time**: Set it up to no earlier than the time of taking the oldest replicated base backup. 3. Specify a name for the new fork service. 4. Select **Create fork**. Run the [avn service create](/docs/tools/cli/service-cli.md#avn-cli-service-create) command with the `--service-to-fork-from` option and the `--recovery-target-time`option. Set `--recovery-target-time` to no earlier than the time of taking the oldest replicated base backup. ``` avn service create FORK_SERVICE_NAME \ --plan SERVICE_PLAN \ --project PROJECT_NAME \ --service-type SERVICE_TYPE \ --cloud SECONDARY_BACKUP_REGION \ --recovery-target-time "YYYY-MM-DDTHH:MM:SS+00:00" \ --service-to-fork-from PRIMARY_SERVICE_NAME ``` Replace the following with meaningful data: * `FORK_SERVICE_NAME` * `SERVICE_PLAN` * `PROJECT_NAME` * `SERVICE_TYPE` * `SECONDARY_BACKUP_REGION` * `PRIMARY_SERVICE_NAME` * `YYYY-MM-DDTHH:MM:SS+00:00` Use the [ServiceCreate](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) API to create a fork service. When constructing the API request, add the `user_config` object to the request body and nest the `service_to_fork_from` field and the `recovery_target_time` field inside. Set `recovery_target_time` to no earlier than the time of taking the oldest replicated base backup. ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "cloud": "SECONDARY_BACKUP_REGION", "plan": "SERVICE_PLAN", "service_name": "FORK_SERVICE_NAME", "service_type": "SERVICE_TYPE", "user_config": { "service_to_fork_from": "PRIMARY_SERVICE_NAME", "recovery_target_time": "YYYY-MM-DDTHH:MM:SS+00:00" } }' ``` Replace the following with meaningful data: * `FORK_SERVICE_NAME` * `SERVICE_PLAN` * `PROJECT_NAME` * `SERVICE_TYPE` * `SECONDARY_BACKUP_REGION` * `PRIMARY_SERVICE_NAME` * `YYYY-MM-DDTHH:MM:SS+00:00` ## Migrate a service with BTAR[​](#migrate-a-service-with-btar "Direct link to Migrate a service with BTAR") You can migrate a service with BTAR the same way you [migrate a service with a regular backup](/docs/platform/howto/migrate-services-cloud-region.md). note When you migrate your service, locations of service backups, both primary and secondary ones, do not change. ## Delete a cross-region backup[​](#delete-a-cross-region-backup "Direct link to Delete a cross-region backup") Delete an additional service backup created in a region different from your primary backup region. You can delete a cross-region backup using the [Aiven Console](/docs/tools/aiven-console.md), [API](/docs/tools/api.md), or [CLI](/docs/tools/cli.md). When you delete the additional cross-region backup, you still have the default backup located in the primary, service-hosting region. * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. From the **Services** view, select an Aiven service on which you'd like to disable BTAR. 3. On your service's page, click **Backups**. 4. On the **Backups** page, click **Actions** > **Secondary backup location**. 5. In the **Edit secondary backup location** window, select **Disable**. Your additional service backup is no longer visible on your service's **Backups** page in the **Secondary backup location** column. To remove secondary backups for your service, use the [avn service update](/docs/tools/cli/service-cli.md) command to remove all target region names from the `additional_backup_regions` array. ``` avn service update your-sevice-name \ -c additional_backup_regions=\[\] ``` To remove secondary backups for your service, update the service configuration. Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint to remove all target regions names from the `additional_backup_regions` array. ``` curl --request PUT \ --url https://api.aiven.io/v1/project/YOUR_PROJECT_NAME/service/YOUR_SERVICE_NAME \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "additional_backup_regions": [] } }' ``` Related pages * [Track restore progress](/docs/products/grafana/howto/track-restore-progress.md) * [Fork your service](/docs/products/grafana/howto/fork-service.md) --- # Change the cloud or region for your Aiven for Grafana® service Move your Aiven for Grafana® service to a different cloud provider or region. 1. In your service, click **Service settings** from the sidebar. 2. In the **Cloud and network** section, click **Actions** > **Change cloud**. 3. In the **Cloud** section , select a cloud provider and region, and click **Change**. Your service starts a migration to the new location and remains available during the process. When the migration completes, the service continues running in the new cloud or region. Related pages * [Fork your service](/docs/products/grafana/howto/fork-service.md) * [Migrate to another cloud or region](/docs/platform/howto/migrate-services-cloud-region.md) --- # Change the plan for your Aiven for Grafana® service Change the service plan for your Aiven for Grafana® service to scale resources up or down and optimize costs. Adjust the plan of your services at any time to scale your services as needed and optimize costs. If you can't find a suitable plan, you can [request a custom plan](/docs/platform/concepts/service-pricing.md). tip If you plan to upgrade your service plan, do it immediately after a full backup. This reduces the amount of incremental changes that need to be applied on top of the base backup, which speeds up the upgrade itself. important * When changing a service plan, reserve an additional 25% of disk space. This requirement applies to upgrades and downgrades. * Downgrading to a plan with fewer VMs is supported for most services, including Aiven for Apache Kafka®, Aiven for PostgreSQL®, Aiven for OpenSearch®, Aiven for ClickHouse®, Aiven for MySQL®, Aiven for Metrics, and Aiven for Valkey™. * Changing a service plan triggers a node recycle, service rebuilding, and any pending maintenance updates. - Console - Terraform - CLI 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Change plan**. 3. Select a plan that provides at least 125% of the current disk size and click **Change plan**. Update the `plan` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). To change a service plan in the Aiven CLI, use the [`avn service update --plan `](/docs/tools/cli/service-cli.md#avn-cli-service-update) command. Your service's state becomes **Rebuilding** and remains accessible. When the state switches to **Running**, your new service plan is active. Related pages * [Prepare for high load](/docs/products/grafana/howto/prepare-for-high-load.md) * [Memory and out-of-memory conditions](/docs/products/grafana/concepts/service-memory.md) --- # Controlled upgrade pipelines for your Aiven for Grafana® service [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Link Aiven for Grafana® services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. Control when your Aiven managed services receive maintenance updates and test maintenance updates in development or staging environments before they reach production. important Controlled upgrade pipeline is a [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-) feature. [Contact Aiven](https://aiven.io/contact) to request access. Aiven performs automatic service maintenance for security fixes, minor software updates, and other platform changes. The controlled upgrade pipeline feature lets you link services of the same type in an ordered sequence to control when each service receives updates. After a maintenance update upgrades a service at the initial pipeline step, you validate that service version before the update proceeds to the service at the next pipeline step. Validating means approving the new version as safe to roll out to the next service. Validation can be manual or automatic after a configurable delay. ## Why use controlled upgrade pipelines[​](#why-use-controlled-upgrade-pipelines "Direct link to Why use controlled upgrade pipelines") Controlled upgrade pipelines prevent production incidents caused by automatic updates reaching production before teams can test the new version in a lower environment. They give you full oversight of the update process: * **Risk mitigation**: Prevents unexpected maintenance updates from breaking your production environment by ensuring they are tested in a non-production setting first. * **Stability**: Keeps destination services (such as production) on a known-good version until you, or the automatic timer, confirm the new version is safe. * **Process control**: Allows platform teams to standardize their deployment and maintenance lifecycle across environments. ## About controlled upgrade pipelines[​](#about-controlled-upgrade-pipelines "Direct link to About controlled upgrade pipelines") ### Upgrade steps[​](#upgrade-steps "Direct link to Upgrade steps") An upgrade step is a pair of services linked by an upgrade constraint: * **Source service**: The service that receives maintenance updates first * **Destination service**: The service that waits for validation before receiving updates Each destination service can have only one source service. A source service can have multiple destination services. ### Upgrade pipelines[​](#upgrade-pipelines "Direct link to Upgrade pipelines") An upgrade pipeline is a chain of upgrade steps that spans multiple environments. For example: * Single chain: development → staging → production * Multiple destinations: development → production-eu and development → production-na ## How validation works[​](#how-validation-works "Direct link to How validation works") When a maintenance update upgrades your source service: 1. The source service receives the update first. 2. Test the updated source service to verify it works as expected. 3. Validate the update manually using the API or CLI, or wait for automatic validation after the configured delay. The default delay is 7 days. 4. After validation, the destination service becomes eligible for the same maintenance update. 5. The destination service receives the update during its next maintenance window. If one source service has multiple destination services, one validation for the source service applies to all connected destination services. ### Validation and maintenance windows[​](#validation-and-maintenance-windows "Direct link to Validation and maintenance windows") Validation and the maintenance window control different things: * **Validation** controls *what* version the destination service upgrades to. * The maintenance window controls *when* the upgrade happens. After you validate an update, or automatic validation applies, the destination service receives the validated version during its next scheduled maintenance window. Validation does not trigger an immediate upgrade outside the maintenance window. Upgrade pipelines add a constraint on what is installed during a maintenance update; they do not change when maintenance runs. Nodes in the destination service maintain the validated version until a newer version is validated, either when you validate it manually or when automatic validation applies after the configured delay. When a node is recycled, it uses the same validated version, not the latest available version. When you create a step, the destination service keeps the newest version that is already validated at that moment. If the destination service is already applying maintenance during step creation, the in-progress target version becomes the initial validated version. warning A powered-off source service cannot receive maintenance updates, so you cannot validate it. If you power off services earlier in the chain, the destination service upgrades regardless. For example, in a development → staging → production chain, if both development and staging are powered off, production upgrades without testing and validation in the earlier environments. Keep services in the chain powered on to preserve the protection that upgrade pipelines provide. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * **Same service type**: You can only link services of the same type. For example, two Aiven for PostgreSQL services. * **Chain length**: The default maximum chain depth is 3 services, which is 2 steps. If you need a longer chain, [contact Aiven](https://aiven.io/contact). * **No cycles**: You cannot create circular dependencies between services. * **Emergency overrides**: Aiven can apply critical security or stability fixes to a destination service before explicit validation. * **Supported services**: This feature supports all Aiven service types except Aiven for Apache Flink® and Aiven for MySQL. * **Automatic maintenance updates only**: Pipelines apply to automatic maintenance updates, such as minor service version updates and node image updates. Major version upgrades, for example Aiven for PostgreSQL® 15 to 16, require manual action and are not promoted automatically through the pipeline. * **No permanent blocking**: You cannot prevent an update indefinitely. Automatic validation applies after the configured delay, up to the maximum delay. * **No validation rollback**: You cannot undo a validation after it is recorded. ## Use controlled upgrade pipelines[​](#use-controlled-upgrade-pipelines "Direct link to Use controlled upgrade pipelines") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") To use controlled upgrade pipelines, you need the following: * The feature enabled by Aiven ([Limited availability](/docs/platform/concepts/service-and-feature-releases.md)) * Dev tool of your choice: * [Aiven CLI](/docs/tools/cli.md) Install the latest version of the Aiven CLI to access the `upgrade-pipeline` commands. * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) * Set `PROVIDER_AIVEN_ENABLE_BETA=true` before running Terraform. * See the [resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) for full schema, import format, and lifecycle behavior. * [Aiven Operator for Kubernetes](/docs/tools/kubernetes.md) Install the operator and create an Aiven token secret named `aiven-token` that the operator uses to authenticate against the Aiven API. * Write access to the source and destination projects * At least two services of the same type (for example, two Aiven for PostgreSQL® services) * Services can be in different projects in the same organization ### Set up an upgrade pipeline[​](#set-up-an-upgrade-pipeline "Direct link to Set up an upgrade pipeline") Use the Aiven CLI or API to create upgrade steps between your services. note The `upgrade-pipeline` CLI commands require Aiven CLI version 4.x or later. Command names and parameters may change before general availability. #### Create an upgrade step[​](#create-an-upgrade-step "Direct link to Create an upgrade step") Create a step to link a source service and a destination service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ [--source-project SOURCE_PROJECT] SOURCE_SERVICE \ [--destination-project DESTINATION_PROJECT] DESTINATION_SERVICE \ [--auto-validation-delay-days DAYS] ``` **Options** * `--organization-id` is required. * `--source-project` and `--destination-project` are optional. If you omit either project option, Aiven CLI uses the current default project set with `avn project switch`. * `--auto-validation-delay-days` is optional. Defaults to 7 days if not specified. ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "SOURCE_PROJECT_NAME", "source_service_name": "SOURCE_SERVICE_NAME", "destination_project_name": "DESTINATION_PROJECT_NAME", "destination_service_name": "DESTINATION_SERVICE_NAME", "auto_validation_delay_days": 7 }' ``` Use the [`aiven_upgrade_step`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) resource: ``` resource "aiven_upgrade_step" "example" { organization_id = "ORGANIZATION_ID" source_project_name = "SOURCE_PROJECT_NAME" source_service_name = "SOURCE_SERVICE_NAME" destination_project_name = "DESTINATION_PROJECT_NAME" destination_service_name = "DESTINATION_SERVICE_NAME" auto_validation_delay_days = 7 } ``` Apply an `UpgradePipelineStep` manifest with `kubectl`: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: upgrade-step-sample spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: SOURCE_PROJECT_NAME sourceServiceName: SOURCE_SERVICE_NAME destinationProjectName: DESTINATION_PROJECT_NAME destinationServiceName: DESTINATION_SERVICE_NAME autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable after the resource is created. Parameters: * `source_project_name`: Name of the project containing the source service * `source_service_name`: Name of the source service * `destination_project_name`: Name of the project containing the destination service * `destination_service_name`: Name of the destination service * `auto_validation_delay_days`: Optional. Number of days before automatic validation. The value must be at least `1`. The default is 7 days. The maximum delay you can configure is 30 days. #### List upgrade steps[​](#list-upgrade-steps "Direct link to List upgrade steps") View all upgrade steps you have access to: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step list --organization-id ORGANIZATION_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) To list managed upgrade steps, use: ``` terraform state list 'aiven_upgrade_step.*' terraform state show 'aiven_upgrade_step.example' ``` List `UpgradePipelineStep` resources in the current namespace: ``` kubectl get upgradepipelinesteps ``` #### View a specific step[​](#view-a-specific-step "Direct link to View a specific step") Get details about a specific upgrade step, including the last validation: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step get \ --organization-id ORGANIZATION_ID \ STEP_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` terraform state show aiven_upgrade_step.example ``` Show the manifest and full status, including `id`, `conditions`, and `lastValidation`: ``` kubectl describe upgradepipelinestep RESOURCE_NAME kubectl get upgradepipelinestep RESOURCE_NAME -o yaml ``` The step details include `last_validation` values such as `validated_at`, `validated_by_user`, and `comment` when validation exists (available through the API). ### Validate an upgrade[​](#validate-an-upgrade "Direct link to Validate an upgrade") After testing your source service with the new update, validate the version to allow the destination service to receive the same update. #### Manual validation[​](#manual-validation "Direct link to Manual validation") Validate the current version of your source service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step validate-for-service \ --project SOURCE_PROJECT \ SERVICE_NAME \ [--comment "COMMENT"] ``` `--comment` is optional. Use it to record a note about the validation, for example `"Tested and verified in development"`. ``` curl -X POST https://api.aiven.io/v1/project/SOURCE_PROJECT/service/SOURCE_SERVICE/upgrade-validation \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "comment": "Tested and verified in development" }' ``` Terraform manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. The operator manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. #### Automatic validation[​](#automatic-validation "Direct link to Automatic validation") If you do not manually validate an update, the system automatically validates the source service version after the configured delay. Auto-validation starts from when the source service receives the update. ### Manage upgrade steps[​](#manage-upgrade-steps "Direct link to Manage upgrade steps") #### Update a step[​](#update-a-step "Direct link to Update a step") Modify the automatic validation delay for an existing step: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step update \ --organization-id ORGANIZATION_ID \ --auto-validation-delay-days 14 \ STEP_ID ``` ``` curl -X PATCH https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "auto_validation_delay_days": 14 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` resource "aiven_upgrade_step" "example" { # ...required fields... auto_validation_delay_days = 14 # Updated from 7 to 14 } ``` Apply the changes: ``` terraform plan terraform apply ``` Edit `autoValidationDelayDays` in your manifest and re-apply: ``` spec: autoValidationDelayDays: 14 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable. To change them, delete the resource and create a new one. #### Delete a step[​](#delete-a-step "Direct link to Delete a step") Remove an upgrade step to allow the destination service to receive updates independently: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step delete --organization-id ORGANIZATION_ID STEP_ID ``` Find `STEP_ID` from the upgrade step list command. ``` curl -X DELETE https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) Remove the resource from configuration and apply, or destroy it directly: ``` terraform apply terraform destroy -target=aiven_upgrade_step.example ``` Delete the `UpgradePipelineStep` resource: ``` kubectl delete upgradepipelinestep RESOURCE_NAME ``` Deleting a step removes all associated validations. ### Example: Three-environment pipeline[​](#example-three-environment-pipeline "Direct link to Example: Three-environment pipeline") Create a pipeline that promotes updates from development to staging to production: * CLI * API * Terraform * Kubernetes 1. Create a step from development to staging: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project dev-project \ --destination-project staging-project \ --auto-validation-delay-days 3 \ pg-dev pg-staging ``` 2. Create a step from staging to production: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project staging-project \ --destination-project prod-project \ --auto-validation-delay-days 7 \ pg-staging pg-prod ``` 1) Create a step from development to staging: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "dev-project", "source_service_name": "pg-dev", "destination_project_name": "staging-project", "destination_service_name": "pg-staging", "auto_validation_delay_days": 3 }' ``` 2) Create a step from staging to production: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "staging-project", "source_service_name": "pg-staging", "destination_project_name": "prod-project", "destination_service_name": "pg-prod", "auto_validation_delay_days": 7 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` # Step 1: Development → Staging resource "aiven_upgrade_step" "dev_to_staging" { organization_id = "ORGANIZATION_ID" source_project_name = "dev-project" source_service_name = "pg-dev" destination_project_name = "staging-project" destination_service_name = "pg-staging" auto_validation_delay_days = 3 } # Step 2: Staging → Production resource "aiven_upgrade_step" "staging_to_prod" { organization_id = "ORGANIZATION_ID" source_project_name = "staging-project" source_service_name = "pg-staging" destination_project_name = "prod-project" destination_service_name = "pg-prod" auto_validation_delay_days = 7 } ``` Apply the configuration: ``` export PROVIDER_AIVEN_ENABLE_BETA=true terraform init terraform plan terraform apply ``` Define both steps in a single manifest and apply it: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: dev-to-staging spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: dev-project sourceServiceName: pg-dev destinationProjectName: staging-project destinationServiceName: pg-staging autoValidationDelayDays: 3 --- apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: staging-to-prod spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: staging-project sourceServiceName: pg-staging destinationProjectName: prod-project destinationServiceName: pg-prod autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-pipeline.yaml ``` When a maintenance update arrives: 1. The development service receives the update. 2. After testing, validate the development version or wait 3 days for auto-validation. 3. The staging service receives the update during its next maintenance window. 4. After testing, validate the staging version or wait 7 days for auto-validation. 5. The production service receives the update during its next maintenance window. Related pages * [Maintenance and updates for your Aiven for Grafana® service](/docs/products/grafana/howto/maintenance-updates.md) * [Change the service plan](/docs/products/grafana/howto/change-service-plan.md) * [Service and feature releases](/docs/platform/concepts/service-and-feature-releases.md) * [Aiven CLI](/docs/tools/cli.md) --- # Dashboard preview for Aiven for Grafana® Grafana's dashboard previews provide a visual overview of your dashboards, displaying each configured dashboard as a graphical thumbnail. Dashboard previews are an optional beta feature available in Grafana 9.0+. By default, this feature is disabled on Aiven for Grafana services. ## Enable dashboard previews[​](#enable-dashboard-previews "Direct link to Enable dashboard previews") 1. In the [Aiven Console](https://console.aiven.io/), select your project and choose your Aiven for Grafana service. 2. Click **Service settings** in the sidebar. 3. Scroll down to **Advanced configuration** and click **Configure**. 4. In the **Advanced configuration** window, click **Add configuration options**. 5. Find and set `dashboard_previews_enabled` to **Enabled**. 6. Click **Save configuration**. The status next to `dashboard_previews_enabled` changes to `synced`. 7. Click **Overview**. From the **Connection information**, copy the **Service URI** into your browser to open the Grafana login page. 8. Enter the username and password from the **Connection information**, and click **Log in**. 9. Click **Dashboards** in the left menu, and select the grid layout to view dashboard previews. Previews are displayed as thumbnails and can be sorted alphabetically. ![Dashboard previews on Grafana](/docs/assets/images/dashboard-previews-on-grafana-0564a09623b3efacaeed5c6ab2780412.png) ## Limitations[​](#limitations "Direct link to Limitations") * Dashboard previews are not available for Hobbyist and Startup-1 plans. * Before downgrading your service plan to Hobbyist or Startup-1, first disable dashboard previews. Related pages For more information on Dashboard previews, see [Grafana documentation](https://grafana.com/docs/grafana/latest/dashboards/). --- # Scale disk storage automatically for your Aiven for Grafana® service Automatically increase the disk storage of your Aiven for Grafana® service when it's running out of space, instead of resizing it manually. Use the Aiven Autoscaler to automatically increase the storage capacity of a service disk when it's running out of space. Disk autoscaler only increases storage, it doesn't scale storage down. ## Why use disk autoscaling[​](#why-use-disk-autoscaling "Direct link to Why use disk autoscaling") * **Cost efficiency**: Start with a regular-sized disk and let Aiven scale it up only when needed, without the risk of running out of disk space. * **Resiliency**: Avoid a service becoming non-functional because it ran out of disk space, including during unexpected spikes in demand. ## How it works[​](#how-it-works "Direct link to How it works") 1. You create an autoscaler integration endpoint in your project, setting the maximum total disk size to allow. 2. You enable an autoscaler integration for your service using that endpoint. 3. Aiven monitors the disk space usage of your service. 4. When disk usage reaches the threshold for your service type, Aiven increases the available storage by at least 10%, using the current used space as a baseline. note The exact increase depends on the service type and cloud provider. Some providers enforce a minimum increase of 10 GB. Autoscale thresholds per service type The threshold that triggers disk autoscaling is a percentage of the available disk storage capacity: * Aiven for OpenSearch®: 75% of the available disk storage capacity * All other supported service types: 85% of the available disk storage capacity 5. The disk increase is recorded in the project event log, and you receive a notification about the added disk space. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * **Maximum storage**: The maximum storage that the autoscaler can allocate for your service is limited by both the maximum disk size set on the autoscaler endpoint and the maximum disk storage supported for your service plan. * **Timing**: Autoscaling takes a moment to complete. In the meantime, the service disk might fill up and the service might enter read-only mode until autoscaling finishes, unless the autoscaler's disk capacity limit is reached. * **Maintenance updates**: Autoscaling works only on fully running services and can't happen during a maintenance update. * **Manual changes**: Changing disk space manually can delay an autoscaling event. * **Terraform**: Don't manage disk space with the Aiven Terraform Provider on a service that uses the autoscaler, to avoid conflicts between the two. * **Performance**: Disk added through autoscaling is slower than the original disk until the next maintenance update applies. This might affect I/O-intensive workloads. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven organization, project, and service that's up and running * The operator role for the organization, project, and service * Dynamic disk sizing support on your service plan and cloud region * One of the following to manage the autoscaler: * [Aiven Console](https://console.aiven.io/) * [Aiven API](https://api.aiven.io/doc/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) ### Enable disk autoscaling[​](#enable-disk-autoscaling "Direct link to Enable disk autoscaling") To enable disk autoscaling, create an autoscaler integration endpoint, then enable an autoscaler integration on your service using that endpoint. * Console * API * CLI * Terraform Create an autoscaler endpoint: 1. Log in to [Aiven Console](https://console.aiven.io/), and go to your organization and project. 2. On the left sidebar, click **Integration endpoints**. 3. Click **Aiven Autoscaler** > **Add new endpoint**. 4. Set the endpoint name and the maximum total disk storage in GB, and click **Add endpoint**. Enable the autoscaler on a service: 1. On the left sidebar, click **Services**, and open your service. 2. On the left sidebar, click **Integrations**. 3. In **Endpoint integrations**, click **Aiven Autoscaler**. 4. Click the endpoint you created, and click **Enable**. 1) Call [ServiceIntegrationEndpointCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointCreate) to create an autoscaler integration endpoint on your project: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "endpoint_name": "ENDPOINT_NAME", "endpoint_type": "autoscaler", "user_config": { "autoscaling": [ { "type": "autoscale_disk", "cap_gb": 300 } ] } }' ``` 2) Call [ServiceIntegrationCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationCreate) to enable the autoscaler integration on your service, using the endpoint ID from the previous response: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "dest_endpoint_id": "ENDPOINT_ID", "integration_type": "autoscaler", "source_project": "PROJECT_NAME", "source_service": "SERVICE_NAME" }' ``` 1. Create an autoscaler integration endpoint using [avn service integration-endpoint-create](/docs/tools/cli.md): ``` avn service integration-endpoint-create \ --project PROJECT_NAME \ --endpoint-name ENDPOINT_NAME \ --endpoint-type autoscaler \ --user-config-json '{"autoscaling": [{"type": "autoscale_disk", "cap_gb": 300}]}' ``` 2. Find the ID of the new endpoint: ``` avn service integration-endpoint-list --project PROJECT_NAME ``` 3. Enable the autoscaler integration on your service, using the endpoint ID from the previous step: ``` avn service integration-create \ --dest-service SERVICE_NAME \ --integration-type autoscaler \ --source-endpoint-id ENDPOINT_ID ``` Use the [`aiven_service_integration_endpoint`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration_endpoint) and [`aiven_service_integration`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration) resources: ``` resource "aiven_service_integration_endpoint" "autoscaler_endpoint" { project = "PROJECT_NAME" endpoint_name = "ENDPOINT_NAME" endpoint_type = "autoscaler" autoscaler_user_config { autoscaling { type = "autoscale_disk" cap_gb = 300 } } } resource "aiven_service_integration" "autoscaler_integration" { project = "PROJECT_NAME" integration_type = "autoscaler" source_service_name = "SERVICE_NAME" destination_endpoint_id = aiven_service_integration_endpoint.autoscaler_endpoint.id } ``` See the [disk autoscaler guide](https://registry.terraform.io/providers/aiven/aiven/latest/docs/guides/disk-autoscaler) for more details. ### Change the maximum disk space for autoscaling[​](#change-the-maximum-disk-space-for-autoscaling "Direct link to Change the maximum disk space for autoscaling") After you enable disk autoscaling, you can update the maximum total disk size at any time. * Console * API * CLI * Terraform 1. On the left sidebar, click **Integration endpoints**. 2. Click **Aiven Autoscaler**. 3. Find your endpoint, click **Actions**, and click the option to edit it. 4. Set a new maximum disk storage value, and save your changes. Call [ServiceIntegrationEndpointUpdate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointUpdate) with the new `cap_gb` value: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint/ENDPOINT_ID \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "user_config": { "autoscaling": [ { "type": "autoscale_disk", "cap_gb": 500 } ] } }' ``` ``` avn service integration-endpoint-update ENDPOINT_ID \ --user-config-json '{"autoscaling": [{"type": "autoscale_disk", "cap_gb": 500}]}' ``` Update the `cap_gb` value in the `autoscaling` block of your [`aiven_service_integration_endpoint`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration_endpoint) resource, then apply the change. ### Turn off disk autoscaling[​](#turn-off-disk-autoscaling "Direct link to Turn off disk autoscaling") To turn off disk autoscaling, remove the autoscaler integration from your service. You can also delete the integration endpoint if you no longer need it. * Console * API * CLI Disconnect the service from the autoscaler: 1. On the left sidebar, click **Services**, and open your service. 2. On the left sidebar, click **Integrations**. 3. In **Endpoint integrations**, find **Aiven Autoscaler**, click **Actions**, and click the option to disconnect it. Delete the autoscaler endpoint, if you no longer need it: 1. On the left sidebar, click **Integration endpoints**. 2. Click **Aiven Autoscaler**. 3. Find your endpoint, and delete it. 1) Call [ServiceIntegrationDelete](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationDelete) to remove the autoscaler integration from your service: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration/INTEGRATION_ID \ --header "Authorization: Bearer TOKEN" ``` 2) Call [ServiceIntegrationEndpointDelete](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointDelete) to delete the autoscaler integration endpoint, if you no longer need it: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint/ENDPOINT_ID \ --header "Authorization: Bearer TOKEN" ``` 1. Find the ID of the integration to remove: ``` avn service integration-list SERVICE_NAME ``` 2. Remove the autoscaler integration from your service: ``` avn service integration-delete INTEGRATION_ID ``` 3. Find the ID of the integration endpoint to delete, if you no longer need it: ``` avn service integration-endpoint-list --project PROJECT_NAME ``` 4. Delete the autoscaler integration endpoint: ``` avn service integration-endpoint-delete ENDPOINT_ID ``` Related pages * [Scale disk storage manually](/docs/products/grafana/howto/scale-disk-storage.md) * [Change the service plan](/docs/products/grafana/howto/change-service-plan.md) * [Memory and out-of-memory conditions](/docs/products/grafana/concepts/service-memory.md) --- # Fork your Aiven for Grafana® service Fork your Aiven for Grafana® service to create an independent copy for testing, debugging, or development without affecting the original service. Forking from a backup also lets you restore your Grafana data to a new service. Fork an Aiven service to create a complete copy of it from its latest backup. Forked services are independent and don't share resources with or increase the load on the original service. Common use cases for forking include: * Creating a snapshot to analyze an issue. * Creating a development copy of your production environment. * Testing upgrades before applying them to production services. * Creating an instance in a different cloud provider, region, or with a different plan. * Renaming a service. During the forking process, the fork might initially have only one node while backups are being taken. The other nodes appear after the backup process is complete. When you fork a service, its configuration, dashboards, and data sources are copied to the new service. ## Limitations[​](#limitations "Direct link to Limitations") * You can only fork services that have at least one backup. * Service integrations are not copied to the fork. * Cross-project forking is supported only within the same organization. ## Fork a service[​](#fork-a-service "Direct link to Fork a service") * Console * CLI * API * Terraform 1. In your service, in the **Backups** section, click **Backup management**. 2. Click **Fork & restore**. 3. Choose the backup to fork from. 4. Enter a name, and select the cloud and plan. 5. Click **Create fork**. Use the [create service command](/docs/tools/cli/service-cli.md#avn-cli-service-create) with: * `--service-to-fork-from`: the name of the service to use as the source. * `--project-to-fork-from`: to fork a service in a different project, set this to the project name the source service is in. Use the [`ServiceCreate` endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) and in the `user_config` property set: * `service_to_fork_from`: the name of the source service. * `project_to_fork_from`: to fork a service in a different project, set this to the name of the project the source service is in. Use the `service_to_fork_from` attribute in the user config of your service resource. To fork a service in a different project, set the `project_to_fork_from` attribute. More information on the service resources and their configuration options is available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs). You are redirected to the **Overview** page of the forked service. The service shows the **Rebuilding** status while it is being created. When the service is ready, the status changes to **Running**. Related pages * [Rename your Aiven for Grafana® service](/docs/products/grafana/howto/rename-service.md) * [Power on/off and delete your Aiven for Grafana® service](/docs/products/grafana/howto/power-cycle-service.md) --- # Maintenance and updates for your Aiven for Grafana® service Manage maintenance updates and set the maintenance window for your Aiven for Grafana® service. ## Maintenance updates[​](#maintenance-updates "Direct link to Maintenance updates") Aiven applies some maintenance updates automatically. The following are the types of updates: * **Mandatory updates:** Security updates, quarterly patch releases, and platform updates that affect reliability or stability of the service nodes. * **Optional updates:** All other updates are initially optional. After six months, they become mandatory and are applied in the next week’s maintenance window. * **Periodic infrastructure updates:** Scheduled automatically for services with nodes active for 180 days and more. These updates are mandatory for all services, except those with maintenance turned off. Critical security updates are applied during the next available maintenance window. For other updates, Aiven gives you at least seven days' notice. Maintenance updates are also automatically applied during service upgrades. To view pending updates: * Console * CLI * API 1. In your service, click **Service settings**. 2. Go to the **Service management** section. Use the [`avn service get`](/docs/tools/cli/service-cli.md#avn_service_get) command. Use the [`service`](https://api.aiven.io/doc/#tag/Service/operation/ServiceGet) endpoint. ## Maintenance window[​](#maintenance-window "Direct link to Maintenance window") The maintenance window is the time period when Aiven can automatically apply maintenance updates to a service. When an update becomes available, Aiven schedules it for the next available maintenance window for each service. The update runs in the first window after it becomes available, and can begin any time after the start time. For example, if a service has a maintenance window of Monday 12:00 UTC, and an update becomes available on Tuesday, the update will be applied on the following Monday. During maintenance, Aiven might restart or replace service nodes. This can cause brief connection interruptions, but services are designed to minimize downtime. Aiven performs maintenance in a rolling-forward style, creating new nodes alongside existing ones and retiring the old nodes after the upgrade completes. Major service upgrades are triggered manually. A manually triggered upgrade starts immediately, regardless of the maintenance window. important You cannot control the order in which services are updated. Each service updates according to its own configured maintenance window, and there is no guaranteed way to control the update sequence. Manual updates and maintenance window adjustments only help for non-critical updates. ## Set the maintenance window[​](#set-the-maintenance-window "Direct link to Set the maintenance window") To set the maintenance window for your service: * Console * Terraform 1. In the Aiven Console, open your service. 2. In the **Maintenance** section, click **Actions** > **Change maintenance window**. 3. Set the day and time. 4. Click **Save changes**. Use the `maintenance_window_dow` and `maintenance_window_time` attributes in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Change the service plan](/docs/products/grafana/howto/change-service-plan.md) --- # Aiven for Grafana® OAuth configuration and security considerations Grafana version 9.5.5 introduced significant changes to the OAuth email lookup behavior to enhance security. However, some users may need to revert to the previous behavior as seen in Grafana 9.5.3. This section provides information on how to revert to the 9.5.3 behavior using the `oauth_allow_insecure_email_lookup` configuration option, its implications, and the associated security threats. ## Security considerations[​](#security-considerations "Direct link to Security considerations") Before reverting to the previous behavior of Grafana version 9.5.3, it is important to consider the security risks involved. ### Authentication bypass vulnerability[​](#authentication-bypass-vulnerability "Direct link to Authentication bypass vulnerability") By enabling the `oauth_allow_insecure_email_lookup` configuration option, the system becomes susceptible to a critical authentication bypass vulnerability using Azure AD OAuth. This vulnerability is officially identified as CVE-2023-3128 and can potentially grant attackers access to sensitive information or unauthorized actions. For more information, refer to the following links: * [Grafana Labs Security Advisory: CVE-2023-3128](https://grafana.com/security/security-advisories/cve-2023-3128/) * [Alternative link for CVE-2023-3128](https://www.cve.org/CVERecord?id=CVE-2023-3128) ## Configuring OAuth email lookup[​](#configuring-oauth-email-lookup "Direct link to Configuring OAuth email lookup") To revert to the OAuth email lookup behavior of Grafana version 9.5.3, you can use the `oauth_allow_insecure_email_lookup` configuration option. ### Enable configuration[​](#enable-configuration "Direct link to Enable configuration") To enable this configuration, include the following line in your Grafana configuration file: ``` [auth] oauth_allow_insecure_email_lookup = true ``` This will restore the behavior to that of Grafana version 9.5.3. However, be aware of the potential security risks if you choose to do so. ## Upgrade to Grafana 9.5.5[​](#upgrade-to-grafana-955 "Direct link to Upgrade to Grafana 9.5.5") In Grafana 9.5.5, the insecure email lookup behavior has been removed to mitigate the security threat. We recommend upgrading to this version to ensure the security of your system. ## Additional resources[​](#additional-resources "Direct link to Additional resources") For more information on configuring authentication in Grafana, refer to the [official Grafana documentation](https://grafana.com/docs/grafana/v9.5/setup-grafana/configure-security/configure-authentication/). --- # Power on/off and delete your Aiven for Grafana® service Power off your Aiven for Grafana® service to release resources and save credits, power it back on when you need it, or delete it permanently. ## Power off a service[​](#power-off-a-service "Direct link to Power off a service") When you power off a service: * All virtual machines are removed from the public cloud. * The service configuration is stored on the Aiven Platform. * If there are no backups, all service data is lost. * If the service has time-based or point in time recovery backups, the backups remain on the Aiven Platform. Services powered off for more than 180 days are automatically deleted. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power off service**. To power off a service, run: ``` avn service update SERVICE_NAME --power-off ``` ## Power on a service[​](#power-on-a-service "Direct link to Power on a service") When you power on a service: * New virtual machines are created on the service's public cloud. * The service starts with the stored configuration parameters. * The latest time-based backup is restored. * Maintenance updates are automatically applied. * If a point in time recovery backup is available, the database transaction logs are replayed to recover the service data to a specific point in time. The restoration takes from a few minutes to a few hours, depending on the network bandwidth, the disk IOPS allocated to the service, and the size of the backup. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power on service**. To power on a service, run: ``` avn service update SERVICE_NAME --power-on ``` To see when the service is running, run: ``` avn service wait SERVICE_NAME ``` note When you power on an Aiven for Grafana service, your dashboards and data sources are restored from the latest available backup. note Static IP addresses are not removed when a service is powered off or deleted. They continue to generate the usual costs. To avoid these costs, [remove the static IP addresses](/docs/platform/concepts/static-ips.md). ## Delete a service[​](#delete-a-service "Direct link to Delete a service") * Console * CLI 1. In your project, click **Services**. 2. Open the service to delete, and click **Actions** > **Delete service**. To delete a service, run: ``` avn service terminate SERVICE_NAME ``` Related pages * [Fork Aiven for Grafana®](/docs/products/grafana/howto/fork-service.md) * [Rename your Aiven for Grafana® service](/docs/products/grafana/howto/rename-service.md) --- # Prepare your Aiven for Grafana® service for high load Prepare your Aiven for Grafana® service for higher than usual traffic to avoid outages and keep performance stable. Prepare your services for higher than usual traffic to avoid service outages by doing the following: * **Subscribe to service notifications:** To receive notifications about service health and warnings when resources are low, you can [set service and project contacts](https://aiven.io/docs/platform/howto/technical-emails). You can also view the status of the Aiven Platform and get updates on incidents on the [status page](https://status.aiven.io/). Follow the RSS feed, subscribe to email or SMS notifications, or use the Slack integration to get notifications about incidents. * **Monitor your services:** [Monitor the health of your services](/docs/platform/howto/list-monitoring.md) using metrics, logs, alerts, and dashboards. * **Scale your services:** If you forecast a load that can't be handled by the service, you can scale up your service. * **Set the backup schedule:** To minimize the impact of the higher load during the backup process, schedule backups outside of peak traffic hours. * **Set the maintenance window:** Schedule maintenance updates outside of your peak traffic hours. * **Run load tests on service forks:** To test the impact of high traffic on a production service, fork the service and run your load test on the fork. Additionally, optimizing a service allows it to perform better under stress therefore avoiding the need of an upgrade. The more optimized a service is for your usage, the better you can weather spikes in traffic. Related pages * [Change the service plan](/docs/products/grafana/howto/change-service-plan.md) * [Maintenance and updates](/docs/products/grafana/howto/maintenance-updates.md) --- # Rename your Aiven for Grafana® service Change the name of your Aiven for Grafana® service by forking it under a new name and deleting the original service. You cannot rename a service after creation. Instead, you can create a fork with the new name and delete the original service. ## Rename a service[​](#rename-a-service "Direct link to Rename a service") 1. Stop writing to the service. 2. Fork the service. 3. Add any integrations or SSO configurations that weren't copied. 4. Connect your clients to the new service. 5. Test the forked service. 6. Delete the original service. Related pages * [Fork Aiven for Grafana®](/docs/products/grafana/howto/fork-service.md) * [Power on/off and delete your Aiven for Grafana® service](/docs/products/grafana/howto/power-cycle-service.md) --- # Replace strings in Grafana® dashboard metric expressions Sometimes, it is useful to replace all occurrences of a string in Grafana metric expressions. Do it with the `aiven-string-replacer-for-grafana` tool available on [GitHub](https://github.com/aiven/aiven-string-replacer-for-grafana). The approach described will work with your own Grafana cluster or with an Aiven for Grafana cluster. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Building the tool requires a Go environment. Follow the [Go installation instructions](https://go.dev/dl/). To build the tool from its [source repository](https://github.com/aiven/aiven-string-replacer-for-grafana), run the following at the terminal prompt: ``` go install github.com/aiven/aiven-string-replacer-for-grafana ``` ## Values you will need[​](#values-you-will-need "Direct link to Values you will need") | Variable | Description | | ----------------------- | --------------------------------------------- | | `GRAFANA_API_KEY` | The API key for accessing Grafana | | `GRAFANA_DASHBOARD_URL` | The URL for the Grafana dashboard | | `GRAFANA_DASHBOARD_UID` | The UID that identifies the Grafana dashboard | | `OLD_STRING` | The old string to replace | | `NEW_STRING` | The new string to use instead | ### Get the Grafana API key[​](#get-the-grafana-api-key "Direct link to Get the Grafana API key") The key needs the Grafana API key to be able to edit the Grafana dashboards. To get your API key (`GRAFANA_API_KEY`): * Go to your Grafana UI. Select the **Configuration** icon and select the **API keys** tab * If you already have an appropriate API key available, save a copy of it. * Otherwise, select **Add API key** and fill in the *Key name*, *Role* and *Time to live*. Click **Add** and save the new API key. tip *Role* must be either *Editor* or *Admin*. ### To get the Grafana dashboard URL and UID[​](#to-get-the-grafana-dashboard-url-and-uid "Direct link to To get the Grafana dashboard URL and UID") To get the dashboard URL and UID (`GRAFANA_DASHBOARD_URL` and `GRAFANA_DASHBOARD_UID`): * Go to the dashboard on which where to change the metric expression strings. * Save the dashboard URL (`GRAFANA_DASHBOARD_URL`). * Select the **Dashboard settings** icon, then select **JSON Model**. Save the dashboard *UID* from the *JSON Model* page (`GRAFANA_DASHBOARD_UID`). ## Perform the replacement[​](#perform-the-replacement "Direct link to Perform the replacement") Use a command like the following to perform the replacement, changing the placeholder variable names to the values you collected above: ``` aiven-string-replacer-for-grafana \ -apikey GRAFANA_API_KEY \ -url GRAFANA_DASHBOARD_URL \ -from OLD_STRING \ -to NEW_STRING \ -uid GRAFANA_DASHBOARD_UID ``` ## Example: changing `elasticsearch_` to `opensearch_`[​](#example-changing-elasticsearch_-to-opensearch_ "Direct link to example-changing-elasticsearch_-to-opensearch_") For instance, use the following command to change all metric expressions that start with `elasticsearch_` to instead start with `opensearch_`: ``` aiven-string-replacer-for-grafana \ -apikey GRAFANA_API_KEY \ -url YOUR_DASHBOARD_URL \ -from elasticsearch_ \ -to opensearch_ \ -uid YOUR_DASHBOARD_UID ``` The change will be visible in the *Query* panel: * Before running the command, metrics start with `elasticsearch_`: ![A screenshot of the Grafana Dashboard query showing metrics with prefix](/docs/assets/images/query-with-elasticsearch-prefix-a8f1f69a22ccf2cd1882d5732ea13311.png) * After running the command, metrics start with `opensearch_`: ![A screenshot of the Grafana Dashboard query showing metrics with prefix](/docs/assets/images/query-with-opensearch-prefix-d05bff7cd4af324c10496c669a8a1461.png) ## Use the version history to revert[​](#use-the-version-history-to-revert "Direct link to Use the version history to revert") If necessary, the *Dashboard changelog* page can be used to revert a change to a specific version. * Go to the dashboard that you modified. * Select the **Dashboard settings** icon, then select **Versions**. * This will show your Dashboard changelog, and you can use this to revert to an earlier version of the dashboard. ![A screenshot of the Grafana Dashboard version changelog after conversion](/docs/assets/images/grafana-version-changelog-16e63227bca81f7b4ae648a1194fcbdc.png) --- # Update Aiven for Grafana® service credentials For improved security, it is recommended to periodically update your credentials. For Grafana, a few steps need to be performed manually to do this. You will need to have access to a web browser, and to have installed `avn`, the [Aiven CLI tool](/docs/tools/cli.md). 1. In the web browser, go to the [Aiven Console](https://console.aiven.io/) page for your Grafana service. 2. Log in to the Grafana instance at the Service URI displayed on that page, using the `avnadmin` credentials. 3. In the bottom-left of Grafana is a small avatar displayed above the help icon. Hover over it, and click Change password. ![Aiven Administrator in Grafana](/docs/assets/images/grafana-credentials-3d970b4af9861d0ac648f02433b08d27.png) 4. Change the password and make a note of it somewhere safe. 5. Log in with `avn` and run the following command to update the stored password in the console: ``` avn service user-password-reset \ --username avnadmin \ --new-password \ ``` For example: ``` avn service user-password-reset \ --username avnadmin \ --new-password my_super_secure_password \ my-grafana-service ``` 6. Refresh the Aiven Console and the new password should now be displayed for the `avnadmin` user. --- # Scale disk storage for your Aiven for Grafana® service Scale the disk storage of your Aiven for Grafana® service up or down without disrupting the running service. /eol-for-major-versions#aiven-for-flinkAdding or removing disk storage does not disrupt the running service. You pay only for extra storage instead of upgrading compute resources. You can add extra storage when you create a service or after it is running. When you add storage to a running service, the Aiven Platform provisions the extra disk and adds it to the running instances. For a clustered service such as Aiven for Apache Kafka®, Aiven divides extra storage equally between the nodes. For a shared service, each node receives the full extra capacity. ## Limitations[​](#limitations "Direct link to Limitations") * Disk added for extra storage is slower than the original disk until the next maintenance update. The slower disk can reduce performance for I/O-intensive workloads. * Maximum storage depends on the plan, service type, and cloud provider. It can be up to five times the plan's base storage size. * Cloud providers limit how many times you can increase storage between maintenance updates. If you reach the limit, run a maintenance update to optimize performance. * You cannot add storage during a maintenance update. * Dynamic disk sizing (DDS) is not supported on custom service plans. Pricing If you add storage when you create a service, **Additional disk storage** shows an estimated monthly cost. The **Service summary** lists plan storage plus additional storage. The estimated monthly price includes the additional storage cost. If you add storage to a running service, the Aiven Console shows the cost of the additional storage and related backups. The same costs appear on your invoices. ## Add or remove storage[​](#add-or-remove-storage "Direct link to Add or remove storage") ### Add storage when you create a service[​](#add-storage-when-you-create-a-service "Direct link to Add storage when you create a service") To add storage while you create a service: 1. In **Additional disk storage**, set the size with the slider or enter a value in GB. 2. Review the estimated monthly cost. 3. In the **Service summary**, click **Create service**. Change additional storage later on the running service, or enable automatic disk scaling with Aiven Autoscaler. ### Change storage on a running service[​](#change-storage-on-a-running-service "Direct link to Change storage on a running service") You cannot add or remove storage when service nodes are in the rebuilding state, for example during a maintenance update or a service upgrade. If you are removing disk storage: * Make sure the data in your service does not exceed the allocated storage. If it does, you cannot remove the additional storage. * Plan for the time it takes to rebuild the service. The time depends on the service. - Console - CLI - Terraform 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Actions** > **Manage additional storage**. 3. Change the disk storage. note * The price shown for the additional storage includes backup costs. * You can only remove storage that you previously added using this feature. To downgrade further, you can change your service plan. 4. Click **Save Changes**. Use [Aiven CLI](/docs/tools/cli.md) to add or remove additional storage using [`avn service update`](/docs/tools/cli/service-cli.md#avn-cli-service-update) with the `--disk-space-gib` flag to specify the total disk space to provide to your service. For example, if your service has a 80-GiB disk and you would like to add an extra 10-GiB disk, use: ``` avn service update --disk-space-gib 90 --project PROJECT_NAME SERVICE_NAME ``` note * When you perform a horizontal service upgrade or downgrade, remember to include all additional disks the service uses. For example, when switching from `Startup-4` to `Business-4` or from `Business-4` to `Startup-4`, include all the additional disks available for this service. * When you fork an existing service, include all additional disks the service uses. Use the `additional_disk_space` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). If you added storage, the additional storage is available immediately. If you removed additional storage, the service nodes go through a rolling restart. Depending on the service type and configuration, there might be a short downtime for services with no HA capabilities. note Storage optimization is performed at the next maintenance update after a change to the storage size. Due to cloud provider limitations, there is a limit on how many times storage can be increased between two maintenance updates. When this limit is reached, perform a maintenance update for performance optimization. Plan increases to avoid reaching this limit. Related pages * [Disk autoscaler](/docs/products/grafana/howto/disk-autoscaler.md) * [Change the service plan](/docs/products/grafana/howto/change-service-plan.md) * [Memory and out-of-memory conditions](/docs/products/grafana/concepts/service-memory.md) --- # Send emails from Aiven for Grafana® Use the Aiven API or the Aiven client to configure the Simple Mail Transfer Protocol (SMTP) server settings and send the following emails from Aiven for Grafana: invite emails, reset password emails, and alert messages. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven client installed * Aiven for Grafana service created * SMTP server IP or hostname * SMTP server port * Username for authentication (if authentication is in use - highly recommended) * Password for authentication * Sender email address (email address shown in **From** field) * Sender name (optional) ## Configure the SMTP server for Grafana[​](#configure-the-smtp-server-for-grafana "Direct link to Configure the SMTP server for Grafana") To configure the Aiven for Grafana service: 1. Open the Aiven client, and log in: ``` avn user login --token ``` 2. Configure the service using your own SMTP values: ``` avn service update --project yourprojectname yourservicename \ -c smtp_server.host=smtp.example.com \ -c smtp_server.port=465 \ -c smtp_server.username=emailsenderuser \ -c smtp_server.password=emailsenderpass \ -c smtp_server.from_address="grafana@yourcompany.com" ``` 3. Optional: Review all available custom options, and configure as needed: ``` avn service types -v ``` You have now set up your Aiven for Grafana to send emails. --- # Tag your Aiven for Grafana® service Add key-value tags to your Aiven for Grafana® service to organize services and track ownership, cost allocation, and governance. Use tags to add metadata to Aiven services to categorize them or run custom logic on them. Typical uses include: * Tagging for governance to deploy services with specific tags only. * Tagging for internal cost reporting, ownership, allocation, and accountability. A tag is a key/value pair: * **Key**: A case-sensitive string that starts with a letter and consists of letters, numbers, dashes, and underscores. The maximum length for a key is 64 characters. * **Value**: A string value limited to 64 UTF-8 characters. Within a service, the tag keys must be unique. * Console * Terraform 1. In the service, click **Service settings**. 2. In the **Service status** section, click **Actions** > **Add service tags**. 3. Enter a key and value for each tag. 4. Click **Save changes**. Use the `tag` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Fork your Aiven for Grafana® service](/docs/products/grafana/howto/fork-service.md) --- # Track restore progress for your Aiven for Grafana® service Track the restore progress of individual nodes in your Aiven for Grafana® service during node replacement, forking, or maintenance, using the Aiven API. You can track restore progress for individual nodes during service node replacement by using the Aiven API. For example, use this endpoint to monitor the restore progress of a forked service or when applying maintenance. The service object exposes restore progress under `node_states[].progress_updates`: * `service.node_states[]` contains per-node state entries. * When a node is restoring or catching up, its `state` is typically `syncing_data`. * When the state is `syncing_data`, the node may include `progress_updates` with one or more phase objects. * Other node states don't include restore progress data. note `progress_updates` may be missing or empty even when a node is in `syncing_data`. This can occur when a restore completes before detailed progress is reported or when the service does not emit detailed progress counters. ## API endpoints[​](#api-endpoints "Direct link to API endpoints") Restore progress fields are part of the standard service response payload. * Get a single service (recommended for polling): `GET /project/{project}/service/{service_name}` * List services in a project: `GET /project/{project}/service` - Request - Response ``` curl -H "Authorization: aivenv1 API_TOKEN" https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME ``` Replace the following placeholders: * `API_TOKEN`: Your Aiven API token. * `PROJECT_NAME`: Your Aiven project name. * `SERVICE_NAME`: The name of your service. ``` { "service": { ... "node_states": [ { "node_name": "...", "state": "syncing_data", "progress_updates": [ { "phase": "basebackup", "completed": false, "current": 3410567, "min": 0, "max": 7569280, "unit": "bytes_uncompressed" } ] } ], ... } } ``` ## Node states[​](#node-states "Direct link to Node states") Common values for `node_states[].state` include: * `setting_up_vm`: The virtual machine is being created or initialized. * `syncing_data`: The node is restoring data or catching up. * `running`: The node is operating normally. * `leaving`: The node is leaving the cluster. * `unknown`: A transient or error state. ## `progress_updates` data model[​](#progress_updates-data-model "Direct link to progress_updates-data-model") `progress_updates` is a list of phase objects. When present, phases appear in the following order: 1. `prepare` 2. `basebackup` 3. `stream` 4. `finalize` Each phase object includes the following fields: ``` { "completed": false, "current": 3410567, "max": 7569280, "min": 0, "phase": "basebackup", "unit": "bytes_uncompressed" } ``` ### Field semantics[​](#field-semantics "Direct link to Field semantics") * `phase`: String, required. The restore phase. Possible values: `prepare`, `basebackup`, `stream`, and `finalize`. * `completed`: Boolean, required. Whether the phase is complete. * `current`: Number or null, optional. The current progress value. This field can be missing or null. * `min`: Number or null, optional. The starting value for the phase. This field can be missing or null. * `max`: Number or null, optional. The expected total value for the phase. This value can be missing, null, or change while the restore is in progress. * `unit`: String or null, optional. The unit for `current`, `min`, and `max`. New unit values can appear over time. Important considerations * Treat `unit` as an opaque identifier. Unknown values can appear. * `max` may change while a restore is in progress. * Not all phases report numeric counters. Some services only indicate phase completion. ## Why `max` values can change[​](#why-max-values-can-change "Direct link to why-max-values-can-change") The `current`, `min`, and `max` values are best-effort progress indicators. They can be based on estimates or on system state that changes over time. Treat `max` as the latest known expected total, not as a fixed guarantee. Common reasons `max` can change include: * The restore process discovers additional work after it starts, such as files, segments, or objects that become visible only after metadata is read. * New data is added on the backend while the node is catching up, which moves the completion point forward. This is common during incremental catch-up phases. * Progress is calculated from system state, such as replication lag, rather than from a fixed work queue. As the system state changes, the value is recalculated. * The service switches restore strategies during the operation, for example from snapshot restore to replication catch-up, which changes what the counters represent. As a result: * Phase percentage can decrease even when the restore operates normally. * Remaining-time estimates based on `max` are unreliable. * Sudden changes in `max` are expected unless the node remains in `syncing_data` longer than expected. ## Restore phase meanings[​](#restore-phase-meanings "Direct link to Restore phase meanings") Phase names are standardized, but the underlying work and the meaning of the counters are service-specific. * `prepare`: Prepares the node for restore. * `basebackup`: Restores the full backup. * `stream`: Applies incremental changes, such as replication or log replay. * `finalize`: Completes final steps before serving traffic. Not all restores include every phase. ## Compute phase progress percentages[​](#compute-phase-progress-percentages "Direct link to Compute phase progress percentages") You cannot reliably compute overall restore progress. You can compute a phase-specific progress percentage when `min`, `max`, and `current` are present and `max != min`. ``` pct = round(((current - min) / (max - min)) * 100, 1) ``` When handling progress values: * If any of `min`, `max`, or `current` is null or missing, display `n/a`. * If `max == min`, treat the percentage as undefined. * Expect the percentage to decrease when `max` changes. * Clamp displayed values to the range `[0, 100]`. ## Polling guidance[​](#polling-guidance "Direct link to Polling guidance") Progress updates are best-effort and refresh every 10 seconds while a node is in `syncing_data`. Poll the service state every 10 to 30 seconds. More frequent polling does not provide additional detail. For each `node_states[]` entry: * If `state` is not `syncing_data`, no restore progress is available. * If `state` is `syncing_data`: * If `progress_updates` is missing or empty, the node is restoring without detailed progress data. * Otherwise, the current phase is the last phase where `completed` is `false`. Stop polling when all nodes reach the `running` state or when a stall is detected. ### Stall detection[​](#stall-detection "Direct link to Stall detection") The API does not provide per-phase timestamps. To detect stalls, use a time-based threshold, such as a node remaining in `syncing_data` longer than expected. Do not rely on counters or `max` values to estimate remaining time. Related pages * [Fork your service](/docs/products/grafana/howto/fork-service.md) * [Back up to another region](/docs/products/grafana/howto/backup-to-another-region.md) --- # Maintenance and lifecycle in Aiven for Grafana® Manage maintenance updates and the maintenance window for your Aiven for Grafana® service. Related pages * [Maintenance and updates](/docs/products/grafana/howto/maintenance-updates.md) --- # Advanced parameters for Aiven for Grafana® See the configuration options available for Aiven for Grafana: | Parameter | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | []()[**additional\_backup\_regions**](#additional_backup_regions)`array`Additional Cloud Regions for Backup Replication | | []()[**custom\_domain**](#custom_domain)`string,null`Serve the web frontend using a custom CNAME pointing to the Aiven DNS name. When you set a custom domain for a service deployed in a VPC, the service certificate is only created for the public-\* hostname and the custom domain. | | []()[**ip\_filter**](#ip_filter)`array`- default: `0.0.0.0/0,::/0`IP filterAllow incoming connections from CIDR address block, e.g. '10.20.0.0/16' | | []()[**service\_log**](#service_log)`boolean,null`Service loggingStore logs for the service so that they are available in the HTTP API and console. | | []()[**static\_ips**](#static_ips)`boolean`Use static public IP addresses | | []()[**external\_image\_storage**](#external_image_storage)`object`External image store settingsexternal\_image\_storage.provider string Provider type External image store provider external\_image\_storage.bucket\_url string Bucket URL for S3 external\_image\_storage.access\_key string S3 access key. Requires permissions to the S3 bucket for the s3:PutObject and s3:PutObjectAcl actions external\_image\_storage.secret\_key string S3 secret key | | []()[**smtp\_server**](#smtp_server)`object`SMTP server settingssmtp\_server.host string Server hostname or IP smtp\_server.port integer - min: 1 - max: 65535 SMTP server port smtp\_server.skip\_verify boolean Skip certificate verification Skip verifying server certificate. Defaults to false smtp\_server.username string,null Username for SMTP authentication smtp\_server.password string,null Password for SMTP authentication smtp\_server.from\_address string From address Address used for sending emails smtp\_server.from\_name string,null From name Name used in outgoing emails, defaults to Grafana smtp\_server.starttls\_policy string StartTLS policy Either OpportunisticStartTLS, MandatoryStartTLS or NoStartTLS. Default is OpportunisticStartTLS. | | []()[**auth\_basic\_enabled**](#auth_basic_enabled)`boolean`Enable or disable basic authentication form, used by Grafana built-in login | | []()[**oauth\_allow\_insecure\_email\_lookup**](#oauth_allow_insecure_email_lookup)`boolean`Allow insecure email lookupEnforce user lookup based on email instead of the unique ID provided by the IdP. This setup introduces significant security risks, such as potential phishing, spoofing, and other data breaches. | | []()[**auth\_generic\_oauth**](#auth_generic_oauth)`object`Generic OAuth integrationauth\_generic\_oauth.allow\_sign\_up boolean Allow sign-up Automatically sign-up users on successful sign-in auth\_generic\_oauth.allowed\_domains array Allowed domains auth\_generic\_oauth.allowed\_organizations array Allowed organizations Require user to be member of one of the listed organizations auth\_generic\_oauth.api\_url string API URL auth\_generic\_oauth.auth\_url string Authorization URL auth\_generic\_oauth.auto\_login boolean Auto login Allow users to bypass the login screen and automatically log in auth\_generic\_oauth.client\_id string Client ID from provider auth\_generic\_oauth.client\_secret string Client secret from provider auth\_generic\_oauth.name string Name of the OAuth integration auth\_generic\_oauth.scopes array OAuth scopes auth\_generic\_oauth.token\_url string Token URL auth\_generic\_oauth.use\_refresh\_token boolean Set to true to use refresh token and check access token expiration. | | []()[**auth\_google**](#auth_google)`object`Google Auth integrationauth\_google.allow\_sign\_up boolean Allow sign-up Automatically sign-up users on successful sign-in auth\_google.client\_id string Client ID from provider auth\_google.client\_secret string Client secret from provider auth\_google.allowed\_domains array Domains allowed to sign-in to this Grafana | | []()[**auth\_github**](#auth_github)`object`Github Auth integrationauth\_github.allow\_sign\_up boolean Allow sign-up Automatically sign-up users on successful sign-in auth\_github.auto\_login boolean Auto login Allow users to bypass the login screen and automatically log in auth\_github.client\_id string Client ID from provider auth\_github.client\_secret string Client secret from provider auth\_github.team\_ids array Require users to belong to one of given team IDs auth\_github.allowed\_organizations array Allowed organizations Require users to belong to one of given organizations auth\_github.skip\_org\_role\_sync boolean Skip organization role sync Stop automatically syncing user roles | | []()[**auth\_gitlab**](#auth_gitlab)`object`GitLab Auth integrationauth\_gitlab.allow\_sign\_up boolean Allow sign-up Automatically sign-up users on successful sign-in auth\_gitlab.api\_url string API URL This only needs to be set when using self hosted GitLab auth\_gitlab.auth\_url string Authorization URL This only needs to be set when using self hosted GitLab auth\_gitlab.client\_id string Client ID from provider auth\_gitlab.client\_secret string Client secret from provider auth\_gitlab.allowed\_groups array Allowed groups Require users to belong to one of given groups auth\_gitlab.token\_url string Token URL This only needs to be set when using self hosted GitLab | | []()[**auth\_azuread**](#auth_azuread)`object`Azure AD OAuth integrationauth\_azuread.allow\_sign\_up boolean Allow sign-up Automatically sign-up users on successful sign-in auth\_azuread.client\_id string Client ID from provider auth\_azuread.client\_secret string Client secret from provider auth\_azuread.auth\_url string Authorization URL auth\_azuread.token\_url string Token URL auth\_azuread.allowed\_groups array Allowed groups Require users to belong to one of given groups auth\_azuread.allowed\_domains array Allowed domains | | []()[**private\_access**](#private_access)`object`Allow access to selected service ports from private networksprivate\_access.grafana boolean Allow clients to connect to grafana with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | []()[**privatelink\_access**](#privatelink_access)`object`Allow access to selected service components through Privatelinkprivatelink\_access.grafana boolean Enable grafana | | []()[**public\_access**](#public_access)`object`Allow access to selected service ports from the public Internetpublic\_access.grafana boolean Allow clients to connect to grafana from the public internet for service nodes that are in a project VPC or another type of private network | | []()[**recovery\_basebackup\_name**](#recovery_basebackup_name)`string`Name of the basebackup to restore in forked service | | []()[**service\_to\_fork\_from**](#service_to_fork_from)`string,null`Name of another service to fork from. This has effect only when a new service is being created. | | []()[**project\_to\_fork\_from**](#project_to_fork_from)`string,null`Name of another project to fork a service from. This has effect only when a new service is being created. | | []()[**user\_auto\_assign\_org**](#user_auto_assign_org)`boolean`Auto-assign new users on signup to main organization. Defaults to false | | []()[**user\_auto\_assign\_org\_role**](#user_auto_assign_org_role)`string`Auto-assign role for new usersSet role for new signups. Defaults to Viewer | | []()[**google\_analytics\_ua\_id**](#google_analytics_ua_id)`string`Google Analytics ID | | []()[**metrics\_enabled**](#metrics_enabled)`boolean`Metrics enabledEnable Grafana's /metrics endpoint | | []()[**cookie\_samesite**](#cookie_samesite)`string`Cookie SameSite attribute: 'strict' prevents sending cookie for cross-site requests, effectively disabling direct linking from other sites to Grafana. 'lax' is the default value. | | []()[**alerting\_error\_or\_timeout**](#alerting_error_or_timeout)`string`Default error or timeout setting for new alerting rules | | []()[**alerting\_nodata\_or\_nullvalues**](#alerting_nodata_or_nullvalues)`string`Default no data or null values settingDefault value for 'no data or null values' for new alerting rules | | []()[**alerting\_enabled**](#alerting_enabled)`boolean`DEPRECATED: setting has no effect with Grafana 11 and onward. Enable or disable Grafana legacy alerting functionality. This should not be enabled with unified\_alerting\_enabled. | | []()[**alerting\_max\_annotations\_to\_keep**](#alerting_max_annotations_to_keep)`integer`- max: `1000000`Max number of alert annotations that Grafana stores. 0 (default) keeps all alert annotations. | | []()[**dashboards\_min\_refresh\_interval**](#dashboards_min_refresh_interval)`string`Minimum refresh intervalSigned sequence of decimal numbers, followed by a unit suffix (ms, s, m, h, d), e.g. 30s, 1h | | []()[**dashboards\_versions\_to\_keep**](#dashboards_versions_to_keep)`integer`- min: `1`
- max: `100`Dashboard versions to keep per dashboard | | []()[**dataproxy\_timeout**](#dataproxy_timeout)`integer`- min: `15`
- max: `90`Timeout for data proxy requests in seconds | | []()[**dataproxy\_send\_user\_header**](#dataproxy_send_user_header)`boolean`Send 'X-Grafana-User' header to data source | | []()[**dashboard\_previews\_enabled**](#dashboard_previews_enabled)`boolean`Enable browsing of dashboards in grid (pictures) mode. This feature is new in Grafana 9 and is quite resource intensive. It may cause low-end plans to work more slowly while the dashboard previews are rendering. | | []()[**dashboard\_scenes\_enabled**](#dashboard_scenes_enabled)`boolean`Enable scenes renderer for dashboardsEnable use of the Grafana Scenes Library as the dashboard engine. i.e. the `dashboardScene` feature flag. Upstream blog post at | | []()[**viewers\_can\_edit**](#viewers_can_edit)`boolean`Viewers can editUsers with view-only permission can edit but not save dashboards | | []()[**editors\_can\_admin**](#editors_can_admin)`boolean`Editors can adminEditors can manage folders, teams and dashboards created by them | | []()[**disable\_gravatar**](#disable_gravatar)`boolean`Set to true to disable gravatar. Defaults to false (gravatar is enabled) | | []()[**allow\_embedding**](#allow_embedding)`boolean`Allow embedding Grafana dashboards with iframe/frame/object/embed tags. Disabled by default to limit impact of clickjacking | | []()[**date\_formats**](#date_formats)`object`Grafana date format specificationsdate\_formats.full\_date string Moment.js style format string for cases where full date is shown date\_formats.interval\_second string Moment.js style format string used when a time requiring second accuracy is shown date\_formats.interval\_minute string Moment.js style format string used when a time requiring minute accuracy is shown date\_formats.interval\_hour string Moment.js style format string used when a time requiring hour accuracy is shown date\_formats.interval\_day string Interval day format Moment.js style format string used when a time requiring day accuracy is shown date\_formats.interval\_month string Moment.js style format string used when a time requiring month accuracy is shown date\_formats.interval\_year string Moment.js style format string used when a time requiring year accuracy is shown date\_formats.default\_timezone string Default time zone for user preferences. Value 'browser' uses browser local time zone. | | []()[**unified\_alerting\_enabled**](#unified_alerting_enabled)`boolean`Enable or disable Grafana unified alerting functionality. By default this is enabled and any legacy alerts will be migrated on upgrade to Grafana 9+. To stay on legacy alerting, set unified\_alerting\_enabled to false and alerting\_enabled to true. See for more details. | | []()[**wal**](#wal)`boolean`Setting to enable/disable Write-Ahead Logging. The default value is false (disabled). | | []()[**grafana\_version**](#grafana_version)`string,null`Grafana major version | --- # Plugins for Aiven for Grafana® Aiven for Grafana includes several pre-installed plugins, updated regularly to ensure optimal performance. For additional plugin requests, contact [Aiven support](mailto:support@aiven.io). ## Panel plugins[​](#panel-plugins "Direct link to Panel plugins") | Plugin Name | Description | Links | | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Alert list | Allows you to display alerts on a dashboard. The list can be configured to show either the current state of your alerts or recent alert state changes | [Grafana](https://grafana.com/grafana/plugins/alertlist/), [Grafana Docs](https://grafana.com/docs/grafana/v7.5/panels/visualizations/alert-list-panel/) | | Annotation list | Shows user annotations in Grafana database | [Grafana](https://grafana.com/grafana/plugins/ryantxu-annolist-panel/), [GitHub](https://github.com/grafana/grafana/tree/main/public/app/plugins/panel/annolist) | | Bar chart | Graphs categorical data | [Grafana Docs](https://grafana.com/docs/grafana/latest/visualizations/bar-chart/) | | Bar gauge | Reduces data fields to a single value and visualizes them as bar gauges | [Grafana Docs](https://grafana.com/docs/grafana/latest/visualizations/bar-gauge-panel/) | | Candlestick | Displays price movements | [Grafana Docs](https://grafana.com/docs/grafana/latest/visualizations/candlestick/) | | Clock | Displays the current time or countdown | [Grafana](https://grafana.com/grafana/plugins/grafana-clock-panel/), [GitHub](https://github.com/grafana/clock-panel) | | Canvas | Provides extensible panel layouts for static and dynamic elements | [Grafana](https://grafana.com/grafana/plugins/canvas/) | | D3 Gauge | Provides D3-based gauge visualizations | [Grafana](https://grafana.com/grafana/plugins/briangann-gauge-panel/), [GitHub](https://github.com/briangann/grafana-gauge-panel) | | Dashboard list | Displays dynamic links to dashboards, filtered by tags or queries | [Grafana](https://grafana.com/grafana/plugins/dashlist/), [Grafana Docs](https://docs.grafana.org/reference/dashlist/) | | Diagram | Creates flowcharts, sequence diagrams, and Gantt charts | [Grafana](https://grafana.com/grafana/plugins/jdbranham-diagram-panel/), [GitHub](https://github.com/jdbranham/grafana-diagram) | | Gauge | Provides standard gauge visualizations | [Grafana Plugin](https://grafana.com/grafana/plugins/gauge), [Grafana Docs](https://grafana.com/docs/grafana/latest/panels-visualizations/visualizations/gauge/) | | Geomap | Displays geographic data | [Grafana Plugin](https://grafana.com/grafana/plugins/geomap/), [Grafana Docs](https://grafana.com/docs/grafana/latest/visualizations/geomap/) | | Graph | Provides a rich set of graphing options | [Grafana Docs](https://grafana.com/docs/grafana/next/panels-visualizations/visualizations/) | | Heatmap | Displays histograms over time | [Grafana](https://grafana.com/grafana/plugins/heatmap-new/), [Grafana Docs](https://docs.grafana.org/features/panels/heatmap/) | | Histogram | Creates histograms for time series data | [Grafana](https://grafana.com/docs/grafana/latest/panels-visualizations/visualizations/histogram/) | | Logs | Displays logs in a built-in panel | [Grafana Docs](https://grafana.com/docs/grafana-cloud/visualizations/panels-visualizations/visualizations/logs/) | | News | Displays an RSS feed | [Grafana Docs](https://grafana.com/docs/grafana/latest/visualizations/news-panel/) | | Node Graph | Visualizes node relationships | [Grafana Plugin](https://grafana.com/grafana/plugins/hamedkarbasi93-nodegraphapi-datasource/), [Grafana Docs](https://grafana.com/docs/grafana/latest/visualizations/node-graph/) | | Plotly (Panel) | Uses Plot.ly to render metrics | [GitHub](https://github.com/NatelEnergy/grafana-plotly-panel) | | Stat | Shows summary statistics for a single series | [Grafana Docs](https://docs.grafana.org/reference/singlestat/) | | State timeline | Displays state changes over time | [Grafana Docs](https://grafana.com/docs/grafana/latest/visualizations/state-timeline/) | | Status history | Displays the history of periodic status changes | [Grafana Docs](https://grafana.com/docs/grafana/latest/visualizations/status-history/) | | Table | Visualizes data in time series, JSON, or annotation formats | [Grafana](https://grafana.com/grafana/plugins/table/) | | Text | Displays markdown-formatted text | [Grafana](https://grafana.com/grafana/plugins/text/) | | Time series | Displays time series data | [Grafana Docs](https://grafana.com/docs/grafana/latest/visualizations/time-series/) | | Worldmap Panel | Displays geohash data or time series on a world map | [GitHub](https://github.com/grafana/worldmap-panel) | | XY Chart | Visualizes X vs Y graphing for arbitrary data points | [Grafana](https://grafana.com/grafana/plugins/xychart/) | ## Data source plugins[​](#data-source-plugins "Direct link to Data source plugins") | Plugin Name | Description | Links | | ------------------------------------- | ----------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Altinity plugin for ClickHouse® | Integrates ClickHouse as a backend database | [GitHub](https://github.com/Altinity/clickhouse-grafana) | | Azure Monitor | Monitors Azure resources | [Grafana](https://grafana.com/grafana/plugins/grafana-azure-monitor-datasource/), [GitHub](https://github.com/grafana/azure-monitor-datasource) | | CloudWatch | Builds dashboards for CloudWatch metrics | [Grafana](https://grafana.com/grafana/plugins/cloudwatch/), [Grafana Docs](https://docs.grafana.org/datasources/cloudwatch/) | | Elasticsearch | Visualizes logs or metrics stored in Elasticsearch | [Grafana](https://grafana.com/grafana/plugins/elasticsearch/), [Grafana Docs](https://docs.grafana.org/datasources/elasticsearch/) | | GitHub | Visualizes GitHub API data in dashboards | [GitHub](https://github.com/grafana/github-datasource) | | Google BigQuery | Integrates BigQuery as a backend database | [GitHub](https://github.com/doitintl/bigquery-grafana) | | Google Sheets | Visualizes Google Sheets data | [Grafana](https://grafana.com/grafana/plugins/grafana-googlesheets-datasource/), [GitHub](https://github.com/grafana/google-sheets-datasource) | | Graphite | Provides various metric functions and navigation | [Grafana](https://grafana.com/grafana/plugins/graphite/), [Grafana Docs](https://docs.grafana.org/datasources/graphite/) | | InfluxDB® | Queries and visualizes InfluxDB data | [Grafana](https://grafana.com/grafana/plugins/influxdb/), [Grafana Docs](https://docs.grafana.org/datasources/influxdb/) | | Instana | Visualizes Instana APM metrics | [Grafana](https://grafana.com/grafana/plugins/instana-datasource/), [GitHub](https://github.com/instana/instana-grafana-datasource) | | Jaeger | Provides distributed tracing | [Grafana Plugin](https://grafana.com/grafana/plugins/jaeger/), [Grafana Docs](https://grafana.com/docs/grafana/latest/datasources/jaeger/) | | Microsoft SQL Server | Queries and visualizes data from Microsoft SQL Server | [Grafana Plugin](https://grafana.com/grafana/plugins/jaeger/), [Grafana Docs](https://grafana.com/docs/grafana/latest/datasources/mssql/) | | MySQL | Queries and visualizes data from MySQL databases | [Grafana](https://grafana.com/grafana/plugins/mysql/), [Grafana Docs](https://docs.grafana.org/features/datasources/mysql/) | | OpenSearch® | Runs OpenSearch queries and visualizes metrics or logs | [Grafana](https://grafana.com/grafana/plugins/grafana-opensearch-datasource/) | | OpenTSDB | Provides a scalable time series database | [Grafana](https://grafana.com/grafana/plugins/opentsdb/), [Grafana Docs](https://docs.grafana.org/datasources/opentsdb/) | | PostgreSQL® | Queries and visualizes data from PostgreSQL databases | [Grafana](https://grafana.com/grafana/plugins/postgres/), [Grafana Docs](https://docs.grafana.org/features/datasources/postgres/) | | Prometheus | Visualizes Prometheus time series data | [Grafana](https://grafana.com/grafana/plugins/prometheus/), [Grafana Docs](https://docs.grafana.org/datasources/prometheus/) | | Prometheus AlertManager | Integrates Prometheus AlertManager API to create dashboards | [GitHub](https://github.com/camptocamp/grafana-prometheus-alertmanager-datasource) | | Stackdriver / Google Cloud Monitoring | Monitors data from Google Cloud Monitoring | [Grafana Plugin](https://grafana.com/grafana/plugins/grafana-stackdriver-datasource/), [Grafana Docs](https://grafana.com/docs/grafana/latest/datasources/google-cloud-monitoring/) | | Tempo | Provides high-volume trace storage | [Grafana Plugin](https://grafana.com/grafana/plugins/tempo/), [Grafana Docs](https://grafana.com/docs/grafana/latest/datasources/tempo/) | | TestData DB | Generates various forms of test data | [Grafana Plugin](https://grafana.com/grafana/plugins/grafana-testdata-datasource/), [Grafana Docs](https://grafana.com/docs/grafana/latest/datasources/testdata/) | | Zipkin | Distributed tracing for visualizing microservices data | [Grafana Plugin](https://grafana.com/grafana/plugins/zipkin/), [Grafana Docs](https://grafana.com/docs/grafana/latest/datasources/zipkin/) | ## Other plugins[​](#other-plugins "Direct link to Other plugins") | Plugin Name | Description | Links | | ---------------------- | -------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | | Grafana Image Renderer | Renders panels and dashboards to PNG images using a headless browser | [Grafana](https://grafana.com/grafana/plugins/grafana-image-renderer/), [GitHub](https://github.com/grafana/grafana-image-renderer) | | Traces (Application) | Grafana Enterprise Traces, based on Tempo, for deploying scalable trace clusters | [Grafana](https://grafana.com/grafana/plugins/grafana-enterprise-traces-app/) | | worldPing | Tests and monitors the global performance of internet applications | [GitHub](https://github.com/raintank/worldping-app) | | Zabbix | Visualizes Zabbix metrics | [Grafana](https://grafana.com/grafana/plugins/alexanderzobnin-zabbix-app/), [GitHub](https://github.com/alexanderzobnin/grafana-zabbix) | *Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries.* --- # Aiven for Grafana® version lifecycle Learn how Aiven manages the Aiven for Grafana® service version, end of life (EOL) dates, and what happens to your service after the version reaches EOL. Aiven for Grafana identifies versions in `major.minor.patch` format, for example `11.6.5`. ## Service version management[​](#service-version-management "Direct link to Service version management") Aiven manages the software version of single-versioned services for you. You don't select or change the major version yourself, and only one version is available at a time. The version your service is currently running is visible in the [Aiven Console](https://console.aiven.io/). ## Version updates[​](#version-updates "Direct link to Version updates") Aiven updates the version as part of regular platform maintenance and rolls it out during your service's maintenance window. Because you don't select a version, Aiven doesn't send the EOL email notifications and reminders described for [multi-versioned services](/docs/platform/reference/eol-for-major-versions.md#service-version-eol-policy). ## After a version reaches end of life[​](#after-a-version-reaches-end-of-life "Direct link to After a version reaches end of life") If Aiven sets an EOL date for the version your service is running: * If the service is powered on, it's automatically upgraded to the next supported version. * If the service is powered off, it's deleted. For the EOL dates of a specific single-versioned service, see the [Aiven single-versioned services EOL](/docs/platform/reference/eol-for-major-versions.md#aiven-single-versioned-services-eol) reference. ## If the entire service reaches end of life[​](#if-the-entire-service-reaches-end-of-life "Direct link to If the entire service reaches end of life") In some cases, Aiven retires an entire service rather than upgrading it to a new version. If this happens, Aiven announces the retirement in advance and provides migration guidance. For the list of retired and soon-to-be-retired services, see [End of life for Aiven services](/docs/platform/reference/end-of-life.md). ## Version EOL dates[​](#version-eol-dates "Direct link to Version EOL dates") | Version | Aiven EOL | | ------- | --------------- | | 11.6.5 | To be announced | ## Newer Grafana versions[​](#newer-grafana-versions "Direct link to Newer Grafana versions") Aiven doesn't publish a release schedule for new Aiven for Grafana major versions. Aiven for Grafana doesn't offer a beta or preview version. If you need a feature that's only available in a newer upstream Grafana release, [contact Aiven](https://aiven.io/contact) to discuss your options. Related pages * [End of life for Aiven services](/docs/platform/reference/end-of-life.md) * [Fork your Aiven for Grafana® service](/docs/products/grafana/howto/fork-service.md) --- # Scaling and performance in Aiven for Grafana® Scale resources and prepare your Aiven for Grafana® service for changing load. Related pages * [Change the service plan](/docs/products/grafana/howto/change-service-plan.md) * [Prepare for high load](/docs/products/grafana/howto/prepare-for-high-load.md) * [Memory and out-of-memory conditions](/docs/products/grafana/concepts/service-memory.md) --- # Aiven for Apache Kafka® Aiven for Apache Kafka® is a fully managed Apache Kafka service for building event-driven applications, data pipelines, and stream processing systems. tip Use an AI assistant connected to [Aiven MCP](/docs/tools/mcp-server.md) to create Kafka services, manage topics, and view cluster details from clients such as Cursor and Claude Code. You create Kafka services as **Standard Kafka** or **Classic Kafka**. New customers create Standard Kafka services. Classic Kafka remains available for existing customers. Free and Developer tier services use Classic Kafka. * [Standard Kafka overview](/docs/products/kafka/standard-kafka-overview.md): diskless topics, object storage, and usage-based sizing on Aiven Cloud * [Classic Kafka overview](/docs/products/kafka/classic-kafka-overview.md): fixed plans, local broker storage, optional tiered storage, and BYOC ## Service tiers and deployment models[​](#service-tiers-and-deployment-models "Direct link to Service tiers and deployment models") * **Free**: Evaluate and experiment with limited throughput and storage. * **Developer**: A paid Classic Kafka tier between Free and Professional, with higher limits than Free and optional Kafka Connect billed separately. * **Professional**: Production workloads, with Kafka Connect on Standard and Classic Kafka services (plan-dependent). Kafka services run on **Aiven Cloud** or **Bring Your Own Cloud (BYOC)**. Standard Kafka is available on Aiven Cloud only. On BYOC, Classic Kafka is available, with diskless topics as an optional add-on. ## Replication with MirrorMaker 2[​](#replication-with-mirrormaker-2 "Direct link to Replication with MirrorMaker 2") Aiven for Apache Kafka® MirrorMaker 2 provides managed replication between Kafka clusters, regions, or cloud providers, including between Standard Kafka and Classic Kafka services. Use it for migration, disaster recovery, and multi-region architectures. ## Data integration with Kafka Connect[​](#data-integration-with-kafka-connect "Direct link to Data integration with Kafka Connect") Aiven for Apache Kafka® Connect provides managed source and sink connectors. On Classic Kafka, Kafka Connect is optional on the Developer tier (billed separately) and supported on the Professional tier. On Standard Kafka, Kafka Connect is available on the Professional tier. ## Get started[​](#get-started "Direct link to Get started") * [Get started with Aiven for Apache Kafka®](/docs/products/kafka/get-started/get-started-kafka.md) * [Create an Aiven for Apache Kafka® Professional tier service](/docs/products/kafka/get-started/create-kafka-service.md) * [Create an Apache Kafka® service with BYOC](/docs/products/kafka/get-started/create-kafka-service-byoc.md) * [Create Kafka topics](/docs/products/kafka/howto/create-topic.md) * [Generate sample data](/docs/products/kafka/howto/generate-sample-data.md) * [Diskless topics overview](/docs/products/kafka/diskless/concepts/diskless-topic-overview.md) --- # Classic Kafka overview Classic Kafka is an Aiven for Apache Kafka® service type that uses fixed plans with local broker storage. You can optionally move older data to object storage with tiered storage when the selected plan and cloud support it. Classic Kafka is compatible with Apache Kafka APIs and clients. It remains available for existing customers. New customers create [Standard Kafka](/docs/products/kafka/standard-kafka-overview.md) services. Free and Developer tier services use Classic Kafka. ## Key characteristics[​](#key-characteristics "Direct link to Key characteristics") Classic Kafka services: * Use **fixed plans** that define compute, memory, and local disk capacity. * Store **classic topics** on local broker disks by default. * Support optional **tiered storage** to offload older data to object storage. * Run on **Aiven Cloud** or **Bring Your Own Cloud (BYOC)**. * Support Free, Developer, and Professional service tiers (availability depends on the deployment model). ## When to use Classic Kafka[​](#when-to-use-classic-kafka "Direct link to When to use Classic Kafka") Use Classic Kafka when you need: * Predictable plan-based capacity and pricing. * Low-latency access to data on local broker storage. * Control over broker size, disk capacity, and scaling. * BYOC deployment, including optional diskless topics as an add-on on supported plans. * Free or Developer tier Kafka services. ## Existing Classic Kafka services[​](#existing-classic-kafka-services "Direct link to Existing Classic Kafka services") Existing Classic Kafka services continue to run unchanged. You cannot upgrade or migrate a Classic Kafka service to Standard Kafka. To use Standard Kafka, create a Standard Kafka service. Classic Kafka and Standard Kafka can run in the same project. To replicate data between them, use [Apache Kafka MirrorMaker 2](/docs/products/kafka/kafka-mirrormaker.md). Related pages * [Create an Aiven for Apache Kafka® Professional tier service](/docs/products/kafka/get-started/create-kafka-service.md) * [Create an Apache Kafka® service with BYOC](/docs/products/kafka/get-started/create-kafka-service-byoc.md) * [Tiered storage for Classic Kafka](/docs/products/kafka/howto/kafka-tiered-storage-get-started.md) * [Standard Kafka overview](/docs/products/kafka/standard-kafka-overview.md) --- # Access Control Lists in Aiven for Apache Kafka® Access Control Lists (ACLs) in Aiven for Apache Kafka® manage access to topics, consumer groups, clusters, and the Schema Registry. Aiven supports two ACL models: * **Aiven ACLs**: Simplified topic-level access control with basic permissions and wildcard support. * **Kafka-native ACLs**: Advanced resource-level access control with fine-grained permissions, including `ALLOW` and `DENY` rules. ## Aiven ACL capabilities[​](#aiven-acl-capabilities "Direct link to Aiven ACL capabilities") Aiven ACLs provide basic permissions and wildcard support, making them suitable for simpler access control scenarios. * **Permissions**: Assign `read`, `write`, `readwrite`, or `admin` permissions to specific users or multiple users using wildcards at the topic level. * **Wildcard support**: Use wildcards in usernames and resource names. * `*` matches zero or more characters. For example, `logs-*` matches `logs-2023` and `logs-error`. * `?` matches exactly one character. For example, `user?` matches `user1` and `userA`, but not `user10`. * **User-specific**: Apply permissions directly to individual users or to multiple users by matching patterns. For example, `username: analyst*` applies to all usernames starting with `analyst`. note By default, a user named `avnadmin` is created with `admin` permissions for all topics. If you use the Aiven Terraform Provider, set `default_acl: false` to disable this behavior. ### Aiven ACL structure[​](#aiven-acl-structure "Direct link to Aiven ACL structure") An Aiven ACL entry consists of the following elements: * **Username**: The Aiven for Apache Kafka service username or a wildcard pattern. * **Permission**: One of `read`, `write`, `readwrite`, or `admin`. * **Associated topics**: Specific Kafka topics or wildcard patterns. Aiven grants access only when an ACL entry matches. If no match is found, access is denied. The order of ACL entries does not affect evaluation. note Aiven ACLs automatically provide access to all consumer groups. You do not need to configure separate ACL entries for consumer group access. ### Examples[​](#examples "Direct link to Examples") * User-specific access: ``` username: abc permission: read topic: xyz ``` Grants user `abc` read access to the topic `xyz`. * Wildcard-based access: ``` username: analyst* permission: read topic: xyz ``` Grants all users with usernames starting with `analyst` read access to the topic `xyz`. * Wildcard in topics: ``` username: developer* permission: read topic: test* ``` Grants all users with usernames starting with `developer` read access to topics starting with `test`. * Wildcard with a single character: ``` username: tester permission: read topic: topic-? ``` Grants `tester` read access to topics such as `topic-1` and `topic-A`. The `?` matches exactly one character. For example, it does not match `topic-10`. To match `topic-10`, use `topic-*` (zero or more characters), `topic-??` (two characters), or `topic-?0`. ## Kafka-native ACL capabilities[​](#kafka-native-acl-capabilities "Direct link to Kafka-native ACL capabilities") Kafka-native ACLs offer precise control with fine-grained permissions and resource-level management. Use them for complex scenarios requiring rules like `ALLOW` and `DENY`. * **Fine-grained permissions**: Support both `ALLOW` and `DENY` rules to provide precise control over access. * **Expanded resource-level control**: Manage access to non-topic resources, such as consumer groups, clusters, and transactional IDs. * **Pattern-based matching**: Use `LITERAL` for exact matches or `PREFIXED` for prefixes to specify how resource names are matched. ### Kafka-native ACL structure[​](#kafka-native-acl-structure "Direct link to Kafka-native ACL structure") A Kafka-native ACL entry consists of the following elements: * **Principal**: The user or service account, such as `User:Alice`. * Use wildcards, such as `User:alice*`, to match all usernames that start with `alice`. * Only `User:` principals are supported. * **Host**: The host to allow. Use `*` to allow all hosts. * **Resource type**: The Apache Kafka resource to control, such as `Topic`, `Group`, `Cluster`, or `TransactionalId`. * **Pattern type**: How the resource value is matched: * **LITERAL**: Matches an exact resource name, such as `my-topic`. * **PREFIXED**: Matches all resources sharing a specified prefix, such as `logs-*`. * **Resource**: The specific Apache Kafka resource, based on the selected pattern type. note When the `pattern_type` is `LITERAL`, setting the `resource` to `*` is a special case that matches all resources. This behavior follows standard Apache Kafka conventions. * **Operation**: The Apache Kafka operation to allow or deny, such as `Read`, `Write`, or `Describe`. * **Permission type**: Specifies whether the action is `ALLOW` or `DENY`. ### Examples[​](#examples-1 "Direct link to Examples") * Granular topic access (prefixed pattern): ``` Principal: User:Alice Resource type: Topic Pattern type: PREFIXED Resource: logs- Operation: Write Permission: ALLOW ``` Grants `User:Alice` write access to all topics starting with the prefix `logs-`. * Granular topic access (literal pattern): ``` Principal: User:Alice Resource type: Topic Pattern type: LITERAL Resource: my-topic Operation: Write Permission: ALLOW ``` Grants `User:Alice` write access to the specific topic `my-topic`. * Restricting sensitive resources (prefixed pattern): ``` Principal: User:Alice Resource type: Topic Pattern type: PREFIXED Resource: logs-sensitive- Operation: Write Permission: DENY ``` Denies `User:Alice` write access to all topics starting with the prefix `logs-sensitive-`. * Restricting sensitive resources (literal pattern): ``` Principal: User:Alice Resource type: Topic Pattern type: LITERAL Resource: logs-sensitive-topic Operation: Write Permission: DENY ``` Denies `User:Alice` write access to the specific topic `logs-sensitive-topic`. ## Rule precedence[​](#rule-precedence "Direct link to Rule precedence") When multiple ACLs apply to the same principal and Kafka resource, `DENY` rules override `ALLOW` rules. Examples where the `DENY` rule takes precedence include: * **Conflicting Kafka-native ACLs**: When multiple Kafka-native ACLs both grant and deny access to the same principal and resource. * **Mixed ACL types**: When both Aiven ACLs and Kafka-native ACLs apply to the same principal and resource with conflicting permissions. **Examples**: * An `ALLOW` rule grants access to resources matching a general pattern, such as: * Topics starting with `test-*` * All consumer groups * A `DENY` rule restricts access to resources matching a specific pattern, such as: * Topics starting with `test-sensitive-*` * Consumer groups with `sensitive` in their names ## ACL permission mapping[​](#acl-permission-mapping "Direct link to ACL permission mapping") The following table summarizes the permissions supported by Aiven ACLs, along with corresponding Apache Kafka actions and Java APIs. | Action | Java API | Admin | Read + Write | Write only | Read only | | ---------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----- | ------------ | ---------- | --------- | | **Cluster** | | | | | | | → `CreateTopics` | [docs](https://kafka.apache.org/42/javadoc/org/apache/kafka/clients/admin/Admin.html#createTopics\(java.util.Collection\)) | ✓ | | | | | **Consumer Groups** | | | | | | | → `Delete` | [docs](https://kafka.apache.org/42/javadoc/org/apache/kafka/clients/admin/Admin.html#deleteConsumerGroups\(java.util.Collection\)) | ✓ | ✓ | | ✓ | | → `Describe` | [docs](https://kafka.apache.org/42/javadoc/org/apache/kafka/clients/admin/Admin.html#describeConsumerGroups\(java.util.Collection\)) | ✓ | ✓ | | ✓ | | → `ListConsumerGroups` | [docs](https://kafka.apache.org/42/javadoc/org/apache/kafka/clients/admin/Admin.html#listConsumerGroups\(org.apache.kafka.clients.admin.ListConsumerGroupsOptions\)) | ✓ | ✓ | | ✓ | | **Topics** | | | | | | | → `Read` | [docs](https://kafka.apache.org/42/javadoc/org/apache/kafka/clients/consumer/KafkaConsumer.html#poll\(java.time.Duration\)) | ✓ | ✓ | | ✓ | | → `Write` | [docs](https://kafka.apache.org/42/javadoc/org/apache/kafka/clients/producer/KafkaProducer.html#send\(org.apache.kafka.clients.producer.ProducerRecord,org.apache.kafka.clients.producer.Callback\)) | ✓ | ✓ | ✓ | | | → `Describe` | [docs](https://kafka.apache.org/42/javadoc/org/apache/kafka/clients/admin/Admin.html#listTransactions\(\)) | ✓ | ✓ | ✓ | ✓ | | → `DescribeConfigs` | [docs](https://kafka.apache.org/42/javadoc/org/apache/kafka/clients/admin/Admin.html#describeConfigs\(java.util.Map\)) | ✓ | | | | | → `AlterConfigs` | [docs](https://kafka.apache.org/42/javadoc/org/apache/kafka/clients/admin/Admin.html#alterConfigs\(java.util.Map\)) | ✓ | | | | | → `Delete` | [docs](https://kafka.apache.org/42/javadoc/org/apache/kafka/clients/admin/Admin.html#deleteTopics\(java.util.Collection\)) | ✓ | | | | | **Transactions** | | | | | | | → `Describe` | [docs](https://kafka.apache.org/42/javadoc/org/apache/kafka/clients/admin/Admin.html#describeTransactions\(java.util.Collection\)) | ✓ | | ✓ | | | → `BeginTransaction` | [docs](https://kafka.apache.org/42/javadoc/org/apache/kafka/clients/producer/KafkaProducer.html#beginTransaction\(\)) | ✓ | | ✓ | | warning Users with `Admin` permissions can create topics with any name because the `CreateTopics` permission is applied at the cluster level. Other permissions, such as `Alter` and `Delete`, apply only to topics that match the specified pattern in the ACL entry. note By default, an Aiven for Apache Kafka service can have up to 50 users. Contact [Aiven support](mailto:support@aiven.io) to request an increase to this limit. Related pages * [Manage access control lists](/docs/products/kafka/howto/manage-acls.md) * [Apache Kafka official documentation](https://kafka.apache.org/42/security/authorization-and-acls/#operations-and-resources-on-protocols) --- # Audit logging for Aiven for Apache Kafka® Audit logging for Aiven for Apache Kafka® records Kafka client activity on your service. Audit logs show which Kafka users were active and can show which Kafka operations were allowed or denied during a configured time period. Use audit logs to support security reviews and compliance workflows. ## How audit logging works[​](#how-audit-logging-works "Direct link to How audit logging works") After a Kafka client authenticates with your Aiven for Apache Kafka service, Aiven checks requested operations and records whether they are allowed or denied. Audit logging writes entries to the service logs with the `AUDIT:` prefix. Each audit entry identifies the Kafka user, also called a principal. In audit logs, the principal is the Kafka user account that a client authenticates with, for example `User:avnadmin`, the default principal in Aiven. Audit logs also include internal client activity on your service. Some internal activity is authenticated, for example `User:karapace_sr`, the principal used by Karapace Schema Registry for accessing the `_schemas` topic. Other activity can be unauthenticated and appear as `User:ANONYMOUS`. Audit logging writes entries to the service logs at regular intervals, based on the configured aggregation period. Each entry shows the first time activity was detected for a Kafka user during that period. Audit logging supports any authentication method and does not require additional access control list (ACL) setup. You can configure audit logging to: * Record Kafka operations, such as read or write requests, or record only that a Kafka user was active. * Include or exclude denied authorization attempts. * Group entries by Kafka user or by Kafka user and source IP address. * Change how often audit entries are written to the service logs. To enable and configure audit logging, see [Configure audit logging for Aiven for Apache Kafka®](/docs/products/kafka/howto/configure-audit-logging.md). ## Audit log examples[​](#audit-log-examples "Direct link to Audit log examples") A `user_operations` entry lists the Kafka user, source IP address, when the user first became active in the aggregation period, and each operation with its result: ``` [2024-07-29 11:02:27,861] AUDIT: User:alice (/192.0.2.10) was active since 2024-07-29T10:57:28.048485122Z: Allow WRITE on TOPIC:orders, Allow READ on GROUP:order-consumers ``` A `user_activity` entry includes only the Kafka user, source IP address, and when the user first became active: ``` [2024-07-29 11:02:27,861] AUDIT: User:alice (/192.0.2.10) was active since 2024-07-29T10:57:28.048485122Z. ``` ## Limitations[​](#limitations "Direct link to Limitations") When using audit logs, note these limitations: * **Audit entries do not include exact operation times.** Audit logging groups activity over an aggregation period, so an entry's timestamp is not the exact time of an individual operation. * **Kafka audit logs do not show which Aiven account made changes using Aiven tools.** Kafka audit logs record only Kafka client activity. Operations performed in the Aiven Console, Aiven CLI, Aiven API, Aiven Provider for Terraform, or Kubernetes operator can appear under an internal Aiven service account, not the individual Aiven account. This includes topic creation and deletion. To find who performed an operation, use the [project event log](/docs/platform/howto/view-project-logs.md). ## Related pages[​](#related-pages "Direct link to Related pages") * [Configure audit logging for Aiven for Apache Kafka®](/docs/products/kafka/howto/configure-audit-logging.md) * [Advanced parameters for Aiven for Apache Kafka®](/docs/products/kafka/reference/advanced-params.md) * [Integrate service logs into an Apache Kafka® topic](/docs/products/kafka/howto/integrate-service-logs-into-kafka-topic.md) * [View project logs](/docs/platform/howto/view-project-logs.md) --- # Authentication types Use modern encryption protocols to protect data in transit with Apache Kafka®. Aiven for Apache Kafka® provides multiple options to secure your data. ## Transport Layer Security[​](#transport-layer-security "Direct link to Transport Layer Security") **Transport Layer Security (TLS)**, also known as Secure Sockets Layer (SSL), is a standard for securing Internet traffic. This method relies on a certificate provided by a Certificate Authority (for example, [letsencrypt.org](https://letsencrypt.org)). With this certificate and the right technical setup, you can use your domain to encrypt traffic to your service. By default, Aiven enables TLS encryption for all **Aiven for Apache Kafka** services and helps with the application, renewal, and configuration of certificates. You can use TLS in two ways: * **TLS encryption** : Your Apache Kafka client validates the certificate for your Apache Kafka broker. * **TLS authentication** : Your Apache Kafka client validates the certificate for your Apache Kafka broker and your broker validates the certificate for your client. ## Simple Authentication and Security Layer[​](#simple-authentication-and-security-layer "Direct link to Simple Authentication and Security Layer") **Simple Authentication and Security Layer (SASL)** allows alternative login methods for your service. Aiven for Apache Kafka supports the following SASL mechanisms: * **PLAIN**: Uses a combination of username and password to log in over a TLS connection, ensuring encrypted traffic. For security reasons, Aiven does not support SASL/PLAIN without TLS, as this would allow anyone to read your credentials when sent. * **SCRAM-SHA-256**: Uses the Salted Challenge Response Authentication Mechanism (SCRAM) with SHA-256 hashing to identify clients without sending plain-text passwords. It does not reveal the password to servers that do not already have it. * **SCRAM-SHA-512**: Similar to SCRAM-SHA-256 but uses SHA-512 hashing for added security. SCRAM works by creating a random `salt`, which is used to create an `identity` that holds: * The `salt` * The number of iterations to use (4096 by default) * `StoredKey` (the hash of the client's key) * `ServerKey` (a key used by the server) By default, this identity is stored in Apache ZooKeeper™. * **OAUTHBEARER**: Uses OAuth 2.0 tokens for authentication. This mechanism is enabled if the `sasl_oauthbearer_jwks_endpoint_url` is specified in the configuration. By default, it is disabled. Schema Registry also supports OAuth 2.0/OIDC authentication and role-based authorization. See [Enable OAuth 2.0/OIDC authentication for Schema Registry](/docs/products/kafka/karapace/howto/enable-oauth-oidc-schema-registry.md). Related pages * [Enable SASL authentication](/docs/products/kafka/howto/kafka-sasl-auth.md) * [Enable OIDC authentication for Aiven for Apache Kafka](/docs/products/kafka/howto/enable-oidc.md) *** *Apache ZooKeeper is a trademark of the Apache Software Foundation in the United States and/or other countries* --- # Configuration backups for Aiven for Apache Kafka® Aiven for Apache Kafka® includes **configuration backups** that automatically back up key service configurations at no additional cost. By eliminating the need for manual reconfiguration, Aiven for Apache Kafka® configuration backups streamline disaster recovery and service restoration. This ensures a reliable way to restore your service configurations after an incident or when powering off/on your service, saving you valuable time and resources. ## What is backed up[​](#what-is-backed-up "Direct link to What is backed up") Configuration backups cover the following components of your Aiven for Apache Kafka service: * **Apache Kafka topics configurations**: Including related configuration parameters such as retention time and number of partitions. * **Apache Kafka users and ACLs**: User permissions and access control lists are included. * **Schema Registry data**: Including schema definitions, schema IDs, configurations (such as compatibility level), subjects, and subject versions. * **Apache Kafka Connect configurations**: Includes all related settings. ## How backups work[​](#how-backups-work "Direct link to How backups work") * **Automatic backups**: Backups are enabled by default and stored in cloud storage every 3 hours. * **Encryption**: Backups are encrypted for security, preventing direct access by users. * **Automatic restoration and configuration retention**: Configuration backups are automatically restored (using the latest backup) after a power off/on cycle. This ensures your service configurations are retained and ready for use. ## Limitations[​](#limitations "Direct link to Limitations") While configuration backups offer significant benefits, there are some limitations: * **No data backups**: Configuration backups do not include the actual data stored in classic topics, consumer groups, and their offsets. Only the cluster configurations are backed up. [Diskless topics](/docs/products/kafka/diskless/concepts/diskless-topic-overview.md) support full restoration of data and configuration because topic data is stored in object storage. * **Latest backup restoration**: Restoration is always from the latest backup available. You cannot choose a specific backup for restoration. ## Additional support[​](#additional-support "Direct link to Additional support") For additional support with configuration backups for your Aiven for Apache Kafka® services, contact . Related pages * [Power on/off and delete your Aiven for Apache Kafka® service](/docs/products/kafka/howto/power-cycle-service.md) * [Diskless topics overview](/docs/products/kafka/diskless/concepts/diskless-topic-overview.md) --- # Consumer lag predictor for Aiven for Apache Kafka® [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) The **consumer lag predictor** for Aiven for Apache Kafka estimates the delay between the time a message is produced and when it's eventually consumed by a consumer group. This information can be used to improve the performance, scalability, and cost-effectiveness of your Kafka cluster. To use the **consumer lag predictor** effectively, set up [Prometheus integration](/docs/platform/howto/integrations/prometheus-metrics.md) with your Aiven for Apache Kafka® service. The Prometheus integration enables the extraction of key metrics necessary for lag prediction and monitoring. ## Consumer lag predictor benefits[​](#consumer-lag-predictor-benefits "Direct link to Consumer lag predictor benefits") By periodically analyzing the Kafka cluster, including topic and consumer group offsets, this *Consumer Lag Predictor* provides insights crucial for scenarios like auto-scaling consumers, ensuring timely message processing. ### Use case[​](#use-case "Direct link to Use case") * **Auto-scaling of consumers**: Leveraging the time lag metric allows users to optimize message processing by dynamically adapting their consumers. ## Limitation[​](#limitation "Direct link to Limitation") Using the **consumer lag predictor** can impact your Apache Kafka cluster's performance. While it offers valuable insights, it also demands more CPU and memory on each Kafka node. To optimize its use, consider the following: * **Resource management:** Enabling the *Consumer Lag Predictor* will lead to higher CPU and memory consumption on each Kafka node. Before enabling it, review your current resource usage to ensure smooth performance. * **Topic selection:** It is advisable to be selective when choosing topics for lag calculation. Starting with a few topics and gradually including more can effectively manage system performance impact. ## Metrics[​](#metrics "Direct link to Metrics") Metrics offer a tangible way to measure and understand consumer lag. The following metrics can be viewed and analyzed using monitoring tools like Prometheus: * `kafka_lag_predictor_topic_produced_records`: Represents the number of records produced, categorized by topic and partition. * `kafka_lag_predictor_group_consumed_records`: Represents the number of records consumed, sorted by group, topic, and partition. * `kafka_lag_predictor_group_lag_predicted_seconds`: Provides the predicted lag time in seconds, organized by group, topic, and partition. ## Next steps[​](#next-steps "Direct link to Next steps") * Learn how to [enable consumer lag predictor](/docs/products/kafka/howto/enabled-consumer-lag-predictor.md) for your Aiven for Apache Kafka® service --- # Follower fetching in Aiven for Apache Kafka® [Follower fetching](/docs/products/kafka/howto/enable-follower-fetching.md) in Aiven for Apache Kafka allows consumers to retrieve data from the nearest replica instead of always fetching from the partition leader. This feature optimizes data fetching by leveraging Apache Kafka's rack awareness, which treats each availability zone (AZ) as a rack. ## Benefits[​](#benefits "Direct link to Benefits") 1. **Reduced network costs:** Fetching data from the closest replica minimizes inter-zone data transfers, reducing costs. 2. **Lower latency:** Fetching from a nearby replica reduces the time it takes to receive data, improving overall performance. note Follower fetching is supported on AWS (Amazon Web Services) and Google Cloud. ## How it works[​](#how-it-works "Direct link to How it works") Aiven for Apache Kafka uses rack awareness to optimize data fetching and maintain availability. Each availability zone (AZ) is treated as a rack. ![Follower fetching](/docs/assets/images/follower-fetching-d893f60c285fbd826ddc053702b81dd0.png) ### Rack awareness[​](#rack-awareness "Direct link to Rack awareness") Rack awareness provides the metadata that follower fetching relies on. In Aiven for Apache Kafka, each availability zone (AZ) is treated as a rack. Each Apache Kafka broker has a `broker.rack` value that corresponds to the AZ where the broker runs: * **AWS:** AZ IDs such as `use1-az1` * **Google Cloud:** AZ names such as `europe-west1-b` Aiven for Apache Kafka automatically manages the `broker.rack` setting. You do not need to configure it manually. ### Follower fetching mechanism[​](#follower-fetching-mechanism "Direct link to Follower fetching mechanism") Follower fetching builds on rack awareness to allow consumers to fetch data from the nearest replica. Apache Kafka consumers use the `client.rack` setting to specify their AZ, ensuring they fetch data from the closest replica when possible. ### Configuration settings[​](#configuration-settings "Direct link to Configuration settings") * `broker.rack`: This setting corresponds to the AZ where each Apache Kafka broker is deployed and helps manage data replication efficiently. Apache Kafka brokers in the same AZ have the same `broker.rack` value, like `use1-az1`. Aiven for Apache Kafka simplifies this process by automatically managing the `broker.rack` setting, eliminating the need for manual configuration. * `client.rack`: This setting on the Apache Kafka consumer indicates the AZ where the consumer is running. It allows you to fetch data from the nearest replica. For example, setting `client.rack` to `use1-az1` on AWS or `europe-west1-b` on Google Cloud ensures that the consumer fetches data from the nearest broker in the same AZ. [Configure](/docs/products/kafka/howto/enable-follower-fetching.md#client-side-configuration) this setting to retrieve data from the closest replica. ## Follower fetching in Kafka Connect and MirrorMaker 2[​](#follower-fetching-in-kafka-connect-and-mirrormaker-2 "Direct link to Follower fetching in Kafka Connect and MirrorMaker 2") Follower fetching is not enabled by default on the Aiven for Apache Kafka service. When it is [enabled](/docs/products/kafka/howto/enable-follower-fetching.md) on the Aiven for Apache Kafka service, Aiven for Apache Kafka® Connect and Aiven for Apache Kafka® MirrorMaker 2 apply rack-aware fetching based on the availability zone (AZ) where each node runs. ### Kafka Connect[​](#kafka-connect "Direct link to Kafka Connect") Kafka Connect assigns a rack value based on the availability zone where each Kafka Connect node runs. Sink connectors use this value when consuming data from Kafka. If rack-aware fetching is supported by the Kafka cluster, sink connectors prefer reading from replicas in the same availability zone. Source connectors do not use follower fetching. ### MirrorMaker 2[​](#mirrormaker-2 "Direct link to MirrorMaker 2") MirrorMaker 2 assigns a rack value based on the availability zone where the MirrorMaker 2 node runs for each replication flow where `follower_fetching_enabled=true`. MirrorMaker always sets a rack value based on the node availability zone when follower fetching is enabled. If the source Kafka cluster does not support follower fetching or uses different rack identifiers, Kafka ignores the rack value and MirrorMaker reads from partition leaders. To disable rack-aware fetching for a specific replication flow, set **Follower fetching enabled** to off when creating or editing the replication flow. For details, see [Configure rack awareness in Aiven for Apache Kafka® MirrorMaker 2](/docs/products/kafka/kafka-mirrormaker/howto/mm2-rack-awareness.md). Related pages * [Enable follower fetching in Aiven for Apache Kafka](/docs/products/kafka/howto/enable-follower-fetching.md) * [Configure rack awareness in Aiven for Apache Kafka® MirrorMaker 2](/docs/products/kafka/kafka-mirrormaker/howto/mm2-rack-awareness.md) --- # Aiven for Apache Kafka® governance [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Governance in Aiven for Apache Kafka® provides a structured and secure way to manage your Aiven for Apache Kafka clusters, ensuring security, compliance, and efficiency. It includes policies, roles, and processes that streamline operations, enforce standards, and centralize control, helping organizations manage their Aiven for Apache Kafka environments effectively. note Governance for Aiven for Apache Kafka is in [limited availability (LA)](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-). Only a subset of governance features are available at this stage. More features and functionalities are incrementally added as development continues. ## Benefits of governance[​](#benefits-of-governance "Direct link to Benefits of governance") * **Enhanced security:** Control access to ensure only authorized users can access and modify Aiven for Apache Kafka resources. * **Increased efficiency:** Decentralize management by allowing teams to manage topics independently, reducing manual effort and streamlining operations. * **Regulatory compliance:** Enforce data retention, audit trails, and data integrity measures to comply with regulatory requirements. * **Centralized control:** Manage Aiven for Apache Kafka resources centrally, simplifying administration across multiple environments. ## Key elements of governance[​](#key-elements-of-governance "Direct link to Key elements of governance") * **Ownership distribution:** Assign topic ownership to user groups, enabling them to manage configurations and access data without direct Apache Kafka cluster access. This promotes shared responsibility, enhances control, and aids in accountability tracking during incidents. * **Request management:** Enable group members to submit and approve requests to create and own topics. The request system follows the Four Eyes Principle, requiring another team member to review and approve changes, thereby maintaining a clear audit trail. * **Self-service:** Provide a user-friendly interface for users to manage Aiven for Apache Kafka resources. Empower team members to claim ownership of Apache Kafka topics proactively, reducing the workload of the administration team and eliminating bottlenecks. * **Standardization:** Establish and enforce standards for topic configurations and partitioning strategies across all Apache Kafka clusters to ensure efficient resource use. * **Improved catalog:** Enhance the topic catalog to make it easier for teams to find and manage topics across the organization. * **Access control:** Implement role-based access control (RBAC) for fine-grained permissions to ensure only authorized users can access sensitive data. Enable organization-wide visibility for Apache Kafka topics, allowing teams to discover topics without direct access to Apache Kafka clusters. * **Monitoring and auditing:** Implement comprehensive monitoring of clusters, topics, and applications. Maintain audit logs for all actions performed on the infrastructure. * **Notifications:** Provide automated notifications to inform users about changes, requests, and approvals. Related pages * [Enable governance for Aiven for Apache Kafka®](/docs/products/kafka/howto/enable-governance.md) * [Aiven for Apache Kafka® topic catalog](/docs/products/kafka/concepts/topic-catalog-overview.md) --- # Scaling options in Apache Kafka® Aiven for Apache Kafka® has a number of predefined plans that specify the number of brokers and the capacity of individual brokers. The predefined plans consist of 3, 6, 9, 15, or 30 brokers, but we can also create larger custom plans based on customer requirements. To increase the capacity of an existing Kafka cluster, two options are available: * **Vertical** scaling the performance of individual brokers * **Horizontal** scaling the cluster by adding more brokers Both scaling options are available for all Aiven for Apache Kafka customers and don't require any downtime, keeping your cluster up and running during the upgrade. note When you change the service plan, Aiven automatically starts adding new brokers with the new specifications to your existing cluster. Once the new brokers are online and the data is replicated to them from the older nodes, the old brokers are retired one by one. Aiven recommends to use both the vertical and horizontal scaling capabilities of Aiven for Apache Kafka to achieve the best possible performance and fault tolerance. tip For production clusters, Aiven advises a minimum of 6 cluster nodes to avoid situations when a failure in a single cluster node causes a sharp increase in load for the remaining nodes. ## Vertical scaling[​](#vertical-scaling "Direct link to Vertical scaling") Scaling vertically a Kafka cluster means keeping the same number of brokers but replacing existing nodes with higher capacity nodes. If you cannot increase the partition or topic count of your Kafka cluster due to application constraints, this is usually the only available option. An example of vertical scaling is when changing the service plan from **Aiven for Apache Kafka Business-4** to **Aiven for Apache Kafka Business-8**. For such case Aiven immediately launches three new brokers with the increased capacity defined by the **Business-8** plan, transfers the data in the new nodes, and once the data is replicated retires the old brokers. ## Horizontal scaling[​](#horizontal-scaling "Direct link to Horizontal scaling") Scaling horizontally means adding more brokers to an existing Kafka cluster. This allows sharing the load in the cluster between a larger number of individual nodes, allowing the cluster to serve more requests as a whole. note Scaling horizontally also makes the cluster more resilient to a failure of a single node: if one broker in a 3-node cluster fails, the remaining two nodes get a 50% load increase, which may cause availability issues in the cluster. If one broker in a 9-node cluster fails, the remaining 8 nodes will only see a load increase of roughly 13%. An example of horizontal scaling is changing the service from the 3-node **Aiven for Apache Kafka Business-8** plan to the 6-node *Aiven for Apache Kafka Premium-6x-8* plan. For such case, Aiven immediately launches six new brokers adding them to the existing cluster. The existing cluster nodes stay online, and once the new brokers are online and included in the cluster configuration, Kafka starts placing partition replicas on them. Once all the data is copied to the new brokers the old nodes are removed. warning Depending on the data volumes, it may take some time until the Apache Kafka cluster is fully balanced. You can also add disk without changing the plan. See [Scale disk storage](/docs/products/kafka/howto/scale-disk-storage.md). Related pages * [Change the service plan](/docs/products/kafka/howto/change-service-plan.md) * [Scale disk storage](/docs/products/kafka/howto/scale-disk-storage.md) * [Prevent full disks](/docs/products/kafka/howto/prevent-full-disks.md) * [Optimizing resource usage for Aiven for Apache Kafka®](/docs/products/kafka/howto/optimizing-resource-usage.md) --- # Pricing for Aiven for Apache Kafka® You pay for the Kafka service plan, retained data, and traffic through your topics. Aiven bills these as compute, storage, and network usage. ## Pricing components[​](#pricing-components "Direct link to Pricing components") Aiven bills the following components separately: * **Compute**: The cost of the selected service plan. * **Storage**: The cost of retained data, based on the amount stored and how long it is retained. * **Network usage**: The cost of data produced to and consumed from Kafka topics. ## Network usage[​](#network-usage "Direct link to Network usage") Network usage depends on Kafka topic traffic. Aiven measures data that producers write to Kafka topics as ingress and data that consumers read from Kafka topics as egress. Classic topics and diskless topics can exist in the same service. Aiven measures network usage separately for each topic type. Ingress and egress can have different rates depending on whether the traffic is for classic topics or diskless topics. Measured ingress and egress are higher than the size of the payload data your application sends and receives. Kafka protocol overhead, including record headers, message framing, and retries, can affect measured traffic, as can client batching behavior. Different client defaults can significantly affect how data is produced and consumed from the cluster. ### Why egress and ingress differ[​](#why-egress-and-ingress-differ "Direct link to Why egress and ingress differ") Aiven measures the ingress and egress independently. Egress is typically higher, because data is usually read by more than one client, or consumed more than once. Reading the same data with multiple consumer groups, repeated reads, retries, or client reconnects all affect cluster egress. Your client library and client configuration can also affect traffic. For example, a client that fetches the same data more than once, uses small fetches, or reconnects frequently can generate more consumed data than expected. ## Cost estimates[​](#cost-estimates "Direct link to Cost estimates") When you create a service, Aiven provides a monthly cost estimate based on your selected configuration. Compute sizing is approximated based on network ingress. This approximation relies on several assumptions, so some workloads can require more compute capacity, such as: * High egress fan-out patterns * A high number of partitions per broker * Uneven distribution of throughput among partitions * A high number of client connections * Enabled integrations, such as Datadog or the consumer lag predictor * Other factors that increase demand on compute For more information about factors that increase broker resource usage, see [Optimizing resource usage](/docs/products/kafka/howto/optimizing-resource-usage.md). Review the estimated monthly cost in **Service summary**. For the Console steps, see [Create an Aiven for Apache Kafka® Professional tier service](/docs/products/kafka/get-started/create-kafka-service.md). ## View usage[​](#view-usage "Direct link to View usage") The [Aiven Console](https://console.aiven.io) shows usage information for the current billing period. To review usage, open the service and go to **Overview** > **Service usage**. You can review: * Ingress and egress usage * Usage split by classic topics and diskless topics * Storage usage * Predicted usage for the billing period To change the compute plan for a Standard Kafka service, see [Change the plan for your Standard Kafka service](/docs/products/kafka/howto/change-standard-kafka-plan.md). note The Aiven Console shows usage values in the unit that best fits the size of the number, for example bytes, KB, MB, GB, or TB. The unit can change as usage grows during the billing period. Aiven bills storage by the GB-month. One GB-month means 1 GB of data stored for one month. The unit combines size and time. Storing 100 GB for a full month is 100 GB-months, and storing 100 GB for half a month is 50 GB-months. Storage usage is prorated by the hour, so the Aiven Console can show usage in smaller units. For example, $0.12 per GB-month is equivalent to $0.000164 per GB-hour, and usage can appear in units such as KB-hours or MB-hours. The Aiven Console shows network usage separately for classic topics and diskless topics. Predicted usage is based on your usage so far in the billing period. Usage information shown during the billing period can change as Aiven processes new usage data. ## Cost drivers[​](#cost-drivers "Direct link to Cost drivers") The following factors affect your estimated or actual cost: * **Cloud and region**: Prices vary by region. * **Service plan**: Determines the Kafka cluster deployed and the compute rate. * **Topic type**: Classic topics and diskless topics have different ingress and egress rates, so the share of traffic that uses each topic type affects network usage costs. For network usage rates, see the [Aiven pricing page](https://aiven.io/pricing). * **Data produced and read**: Data produced and consumed is billed. * **Retention period**: Longer retention keeps data in storage for longer. ## Manage costs[​](#manage-costs "Direct link to Manage costs") To manage costs, review the factors that affect your compute, storage, and network usage: * Review the service plan that matches your workload. Consult an expert for any tuning. * Use diskless topics for high-throughput workloads that do not require low latency. * Adjust topic retention to control storage usage. * Review consumer applications and client configuration if egress is higher than expected. * Monitor usage during the billing period to identify unexpected changes. ## Pricing by service type[​](#pricing-by-service-type "Direct link to Pricing by service type") If **Service type** appears when you create a service, pricing depends on the selected type: * **Standard**: Aiven bills compute, storage, and network usage separately. * **Classic**: Aiven includes network usage in the compute price. Related pages * [Create an Aiven for Apache Kafka® Professional tier service](/docs/products/kafka/get-started/create-kafka-service.md) * [Change the plan for your Standard Kafka service](/docs/products/kafka/howto/change-standard-kafka-plan.md) * [Standard Kafka overview](/docs/products/kafka/standard-kafka-overview.md) * [Classic Kafka overview](/docs/products/kafka/classic-kafka-overview.md) * [Diskless topics overview](/docs/products/kafka/diskless/concepts/diskless-topic-overview.md) * [Compare diskless and classic topics](/docs/products/kafka/diskless/concepts/topics-vs-classic.md) --- # Quotas in Aiven for Apache Kafka® Quotas ensure fair resource allocation, stability, and efficiency in your Kafka cluster. In Aiven for Apache Kafka®, you can [add quotas](/docs/products/kafka/howto/manage-quotas.md) to limit the data or requests exchanged by producers and consumers within a specific period, preventing issues like broker overload, network congestion, and service disruptions caused by excessive or malicious traffic. You can effectively manage resource consumption and ensure optimal user performance by implementing quotas. You can add and manage quotas using [Aiven Console](https://console.aiven.io/), [Aiven CLI](/docs/tools/cli/service/quota.md) and [Aiven API](https://api.aiven.io/doc/). Using quotas offer several benefits: * **Resource management:** Quotas prevent individual clients from consuming excessive resources. This ensures fairness in resource allocation. * **Stability:** Setting limits on network throughput and CPU usage helps maintain stability and prevent performance degradation of the Apache Kafka cluster. * **Efficiency:** Quotas enable you to optimize resource utilization and achieve better overall efficiency within your Kafka deployment. ## Supported quota types[​](#supported-quota-types "Direct link to Supported quota types") Aiven for Apache Kafka provides different quotas to help you manage resources effectively. These quotas offer benefits in controlling network bandwidth and CPU usage: * **Consumer throttle (Network bandwidth quota):** This quota allows you to limit the amount of data a consumer can retrieve from the Kafka cluster per second. Setting a maximum network throughput prevents any single consumer from using excessive network bandwidth. * **Producer throttle (Network bandwidth quota):** Similar to the consumer throttle, this quota limits the amount of data a producer can send to the Kafka cluster per second. It ensures that producers do not overload the system by sending excessive data, thereby maintaining system stability. * **CPU throttle:** This quota is about managing CPU usage. You can manage CPU usage by setting a percentage of the total CPU time. Limiting the CPU resources for specific client IDs or users prevents any individual from monopolizing CPU resources, promoting fairness and efficient resource utilization. ## Client IDs and users[​](#client-ids-and-users "Direct link to Client IDs and users") **Client ID** and **User** are two types of entities that can be used to enforce quotas in Kafka. **Client ID:** A Client ID is a unique identifier assigned to each client application or producer/consumer instance that connects to a Kafka cluster. It helps track the activity and resource usage of individual clients. When configuring quotas, you can set limits based on the Client ID, allowing you to control the amount of resources (such as network bandwidth or CPU) a specific client can utilize. **Users:** A User represents the authenticated identity of a client connecting to a cluster. With authentication mechanisms like SASL, users are associated with specific connections. By setting quotas based on Users, resource limits can be enforced per-user. ## Quota enforcement[​](#quota-enforcement "Direct link to Quota enforcement") Quotas enforcement ensures clients stay within their allocated resources. These quotas are implemented and controlled by the brokers on an individual basis. Each client group is assigned a specific quota for every broker, and when this threshold is reached, throttling mechanisms come into action. When a client exceeds its quota, the broker calculates the necessary delay to bring the client back within its allocated limits. Subsequently, the broker promptly responds to the client, indicating the duration of the delay. Additionally, the broker suspends communication with the client during this delay period. This cooperative approach from both sides ensures the effective enforcement of quotas. Quota violations are swiftly detected using short measurement windows, typically 30 windows of 1 second each. This ensures timely correction and prevents bursts of traffic followed by long delays, providing a better user experience. Related pages * [Quota enforcement in the Apache Kafka® documentation](https://kafka.apache.org/documentation/#design_quotas) * [Manage quotas in Aiven for Apache Kafka®](/docs/products/kafka/howto/manage-quotas.md) * [`avn service quota`](/docs/tools/cli/service/quota.md) * [Introducing Kafka Quotas in Aiven for Apache Kafka](https://aiven.io/blog/introducing-kafka-quotas-in-aiven-for-apache-kafka) --- # Apache Kafka® REST API The Apache Kafka® REST API lets you produce and consume messages over HTTP. Use it from tools such as `curl`, scripts, or services that speak HTTP instead of the Kafka protocol. [Karapace](/docs/products/kafka/karapace.md) serves these APIs through the REST proxy, which forwards HTTP requests to Apache Kafka. Karapace also provides Schema Registry REST APIs to register, update, and delete schemas. Those APIs are separate from the Kafka REST API. ## Enable the REST API[​](#enable-the-rest-api "Direct link to Enable the REST API") The REST API is available after you enable the REST proxy on your service. In **Connection information**, open the **Apache Kafka REST** tab and click **Enable**, or enable **Apache Kafka REST API (Karapace)** from **Service management**. For automation, set the `kafka_rest` parameter. See [Enable schema registry and REST proxy](/docs/products/kafka/karapace/howto/enable-karapace.md). ## Use the REST API[​](#use-the-rest-api "Direct link to Use the REST API") Requests authenticate with a service user's username and password. To produce and consume messages with generated `curl` commands from the Aiven Console, see [Connect with Kafka REST](/docs/products/kafka/howto/connect-with-kafka-rest.md). ## Control access[​](#control-access "Direct link to Control access") REST proxy authorization applies [Kafka ACLs](/docs/products/kafka/concepts/acl.md) to REST requests. See [Enable Apache Kafka® REST proxy authorization](/docs/products/kafka/karapace/howto/enable-kafka-rest-proxy-authorization.md). Related pages * [Enable schema registry and REST proxy](/docs/products/kafka/karapace/howto/enable-karapace.md) * [Connect with Kafka REST](/docs/products/kafka/howto/connect-with-kafka-rest.md) * [Enable Apache Kafka® REST proxy authorization](/docs/products/kafka/karapace/howto/enable-kafka-rest-proxy-authorization.md) --- # Tiered storage in Aiven for Apache Kafka® overview Tiered storage in Aiven for Apache Kafka® helps you manage data more effectively by using two different storage types: local disk and remote cloud storage solutions like AWS S3, Google Cloud Storage, and Azure Blob Storage. This feature lets you allocate frequently accessed data to high-speed local disks while keeping an extended data retention on more cost-effective remote storage solutions. You can store data on specific topics indefinitely without running out of space. Once enabled, you configure tiered storage per topic, which gives you granular control over your data storage. In Classic Kafka services, tiered storage is an optional feature that you enable and configure per topic. In Standard Kafka services, classic topics use managed remote storage by default and you cannot turn it off. note * Aiven for Apache Kafka® supports tiered storage starting from Apache Kafka® version 3.6 or later. It is recommended to upgrade to the latest default version and apply [maintenance updates](/docs/products/kafka/howto/maintenance-updates.md) when using tiered storage for the latest fixes and improvements. * Tiered storage is not available on startup plans. ## Benefits of tiered storage[​](#benefits-of-tiered-storage "Direct link to Benefits of tiered storage") Tiered storage offers multiple benefits, including: * **Scalability:** Tiered storage in Aiven for Apache Kafka decouples storage and computing, allowing them to scale independently. Storage capacity can expand almost infinitely with cloud solutions, and compute resources can adjust based on demand, eliminating concerns about storage or processing limitations. * **Cost efficiency:** By moving less frequently accessed data to a cost-effective storage tier, you can achieve significant financial savings. * **Operational speed:** With the bulk of data offloaded to remote storage, service rebalancing in Aiven for Apache Kafka becomes faster, making for a smoother operational experience. * **Infinite data retention:** Remote storage allows Aiven for Apache Kafka to retain data indefinitely, ensuring long-term storage without impacting local storage limits or speed. ## When and why use it[​](#when-and-why-use-it "Direct link to When and why use it") Understanding when and why to use tiered storage in Aiven for Apache Kafka helps you maximize its benefits, particularly around cost savings and system performance. **Scenarios for use:** * **Long-term data retention**: Many organizations require large-scale data storage for extended periods, either for regulatory compliance or historical data analysis. Cloud services provide an almost limitless storage capacity, making it possible to keep data accessible for as long as required at a reasonable cost. This is where tiered storage becomes especially valuable. * **High-speed data ingestion**: Tiered storage can manage unpredictable or sudden data influxes by supplementing local disks with cloud storage, to ensure optimal performance. * **Unlock unexplored opportunities:** Tiered storage in Aiven for Apache Kafka addresses current storage challenges and enables new and innovative use cases that were previously impractical or too expensive. By eliminating traditional storage limitations, organizations gain the flexibility to support a wide range of applications and workflows, including scenarios where using Apache Kafka was once deemed impractical. This flexibility encourages creative thinking and redefines the experience with Apache Kafka. ## Pricing and charges[​](#pricing-and-charges "Direct link to Pricing and charges") Tiered storage costs depend on the amount of remote storage used, measured in GB/hour. Charges are based on the highest usage level within each hour. ### Tiered storage charges[​](#tiered-storage-charges "Direct link to Tiered storage charges") Aiven charges for tiered storage data based on the data stored in remote tier, such as AWS S3. This means you pay for: * **Data stored in S3:** Billing is based on the highest amount of data stored in S3 within each hour, not the total amount stored over the hour. * **Local storage on the Aiven platform:** Local storage has a fixed cost based on the provisioned capacity. You pay a set amount each month for a fixed amount of local storage. Tiered Storage is designed to transfer **all** topic data to the remote tier, except for the active log segment where the latest data is being appended. This transfer occurs regardless of local storage retention settings. #### Example[​](#example "Direct link to Example") Imagine having a tiered topic with 40 TB of data and with tiered storage enabled. You set a local retention time of 10 TB. Consider the following: * You pay a fixed monthly cost for the local storage allocated to your service, which includes the 10 TB retained locally. * You are charged for storing **approximately 40 TB** of data in S3, even if some data is available locally within the most recent 10 TB. This is because data is automatically transferred to S3, and billing is based on the highest usage level within each hour. note The estimated storage size of 40 TB excludes the active log segments, which are only stored locally. For more information, see [local vs. remote data retention](/docs/products/kafka/concepts/tiered-storage-how-it-works.md#local-vs-remote-data-retention). ### BYOC (Bring your own cloud) billing[​](#byoc-bring-your-own-cloud-billing "Direct link to BYOC (Bring your own cloud) billing") [BYOC](/docs/platform/concepts/byoc.md) billing for tiered storage can vary depending on your specific agreement with Aiven. Potential scenarios include: * **Customer costs:** In all BYOC setups, you are responsible for the full cost of the underlying cloud storage used by tiered storage. This includes the entire storage used, regardless of local retention settings (80 TB in this example). * **Aiven management fee:** In addition to the underlying S3 storage costs, you also pay an Aiven management fee for the data stored in S3. This fee is based on the amount of storage added via tiered storage. Related pages * [How tiered storage works in Aiven for Apache Kafka®](/docs/products/kafka/concepts/tiered-storage-how-it-works.md) * [Guarantees](/docs/products/kafka/concepts/tiered-storage-guarantees.md) * [Limitations](/docs/products/kafka/concepts/tiered-storage-limitations.md) * [Enabled tiered storage for Aiven for Apache Kafka® service](/docs/products/kafka/howto/enable-kafka-tiered-storage.md) --- # KRaft in Aiven for Apache Kafka® Starting with Apache Kafka 3.9, Aiven for Apache Kafka uses **KRaft** (Kafka Raft) to manage metadata and controllers, replacing ZooKeeper. KRaft, like ZooKeeper, is an internal component of Apache Kafka but simplifies metadata management and improves efficiency. note Apache Kafka 3.8 is the last version that supports ZooKeeper. All new Aiven for Apache Kafka services running Apache Kafka 3.9 or later use KRaft by default. ## What is KRaft?[​](#what-is-kraft "Direct link to What is KRaft?") KRaft is the built-in metadata and consensus management system in Aiven for Apache Kafka 3.9 and later, replacing ZooKeeper. Aiven for Apache Kafka services use KRaft to manage metadata internally, eliminating the need for a separate ZooKeeper cluster. Apache Kafka manages metadata and controllers using the **Raft consensus algorithm**. ### Key differences between KRaft and ZooKeeper[​](#key-differences-between-kraft-and-zookeeper "Direct link to Key differences between KRaft and ZooKeeper") KRaft introduces a new way of managing metadata directly within Apache Kafka. Instead of relying on a separate ZooKeeper cluster for metadata storage and coordination, Apache Kafka uses dedicated **controller nodes** that operate using the **Raft consensus algorithm**. | Feature | ZooKeeper | KRaft | | ------------------- | ----------------------------------------------------------------- | --------------------------------------------------------------------------------------- | | Consensus algorithm | ZooKeeper uses the Zab (ZooKeeper Atomic Broadcast) protocol. | KRaft uses the Raft protocol, which is built into Apache Kafka for metadata management. | | Architecture | A separate ZooKeeper cluster is required for metadata management. | KRaft is built into Apache Kafka and uses dedicated controllers. | | Metadata storage | Metadata is stored externally in ZooKeeper. | Metadata is stored internally in Apache Kafka. | Both ZooKeeper and KRaft are fully managed by Aiven. You do not need to handle any operational aspects of either system. Aiven takes care of all maintenance, monitoring, and scaling, regardless of which system is used for metadata management. ## How KRaft works[​](#how-kraft-works "Direct link to How KRaft works") KRaft introduces two key roles within an Aiven for Apache Kafka service: * **Brokers**: Handle message storage and processing. * **Controllers**: Manage cluster metadata using the Raft consensus algorithm. Separating these roles improves metadata management efficiency without affecting broker performance. ## Migration from ZooKeeper to KRaft[​](#migration-from-zookeeper-to-kraft "Direct link to Migration from ZooKeeper to KRaft") The migration occurs as part of the upgrade to Apache Kafka 3.9. It is fully automated and does not require manual steps. After the migration completes, reverting to ZooKeeper is not supported. ### How the migration works[​](#how-the-migration-works "Direct link to How the migration works") * If you upgrade your service to Apache Kafka 3.9, the upgrade starts immediately. After all service nodes are running on Apache Kafka 3.9, the KRaft migration starts automatically. The configured [maintenance window](/docs/products/kafka/howto/maintenance-updates.md#set-the-maintenance-window) does not affect when the upgrade or migration starts. * If Aiven upgrades your service because Apache Kafka 3.8 has reached end of life, the version upgrade starts during the configured maintenance window. The KRaft migration starts after the version upgrade completes. * During the upgrade, the service status changes to **Rebuilding**. * When the upgrade completes, the service status returns to **Running** and the service runs in KRaft mode. * After the migration completes, the service enters a one-week grace period. Further Kafka version upgrades are not supported during this time. #### What happens in the background[​](#what-happens-in-the-background "Direct link to What happens in the background") 1. Aiven replaces the Kafka 3.8 nodes with Kafka 3.9 nodes. The new nodes temporarily run on ZooKeeper. 2. After the old nodes are decommissioned and the new cluster is healthy, Aiven starts the KRaft controllers. The service now runs in dual mode (ZooKeeper + KRaft). 3. Aiven transfers the cluster metadata from ZooKeeper to KRaft. 4. Brokers switch to KRaft mode and no longer use ZooKeeper. 5. During a one-week grace period, the KRaft controllers remain in migration mode. 6. After the grace period ends, the KRaft controllers switch to full KRaft mode. #### Emergency rollback[​](#emergency-rollback "Direct link to Emergency rollback") The KRaft migration can be rolled back by Aiven for up to one week after it completes. Rolling back reverts the service to using ZooKeeper and decommissions the KRaft controller cluster. At this point, the service continues to run on Kafka 3.9, and the migration can be retried. ### Before you migrate[​](#before-you-migrate "Direct link to Before you migrate") * **Service version**: To upgrade to Apache Kafka 4.0 or later, your service must first migrate to KRaft by upgrading to Apache Kafka 3.9. * **Grace period**: After the migration completes, there is a one-week grace period during which the cluster cannot be upgraded to Apache Kafka 4.0. * **Service plan**: If your service is on the **Startup-2** plan, you must upgrade to **Startup-4** or higher during the same maintenance window as the version upgrade. ## Impact of KRaft[​](#impact-of-kraft "Direct link to Impact of KRaft") ### Compatibility and impact[​](#compatibility-and-impact "Direct link to Compatibility and impact") KRaft does not change how Aiven for Apache Kafka services work. Applications, clients, and integrations, such as Apache Kafka brokers, Aiven for Apache Kafka Connect, Aiven for Apache MirrorMaker 2, and Karapace, continue to function as expected. You can run services on Apache Kafka 3.8 or earlier (using ZooKeeper) alongside services on Apache Kafka 3.9 or later (using KRaft) without compatibility issues. The only differences come from features available in Apache Kafka 3.9 that may not exist in previous versions. ### Monitoring and metrics[​](#monitoring-and-metrics "Direct link to Monitoring and metrics") Some ZooKeeper-related controller metrics are not available in KRaft. For a list of removed metrics, see [KRaft and metrics changes](/docs/products/kafka/reference/kafka-metrics-prometheus.md#kraft-mode-and-metrics-changes) Related pages [Transitioning to KRaft](/docs/products/kafka/concepts/upgrade-procedure.md#transitioning-to-kraft) --- # Tiered storage in Aiven for Apache Kafka® [Early availability](/docs/platform/concepts/service-and-feature-releases.md) Discover how tiered storage works in Aiven for Apache Kafka®, explore its use cases, and learn why you might need it and what benefits it offers. --- # Compacted topics One way to reduce the disk space requirements in Apache Kafka® is to use **compacted topics**. This method retains only the newest record for each key on a topic, regardless of whether the retention period of the message has expired or not. Depending on the application, this can significantly reduce the amount of storage required for the topic. To use log compaction, all messages sent to the topic must have an explicit key. To enable log compaction, see [Configure log cleaner for topic compaction](/docs/products/kafka/howto/configure-log-cleaner.md). ## How compacted topics work[​](#how-compacted-topics-work "Direct link to How compacted topics work") An Apache Kafka topic represents a continuous stream of messages that typically get discarded after the message reaches a certain period of time or size. However, for certain use cases it is only needed the most recent value for a certain key. For example, if there is a topic containing a user's home address, on every update, a message is sent using `user_id` as the primary key and home address as the value: ``` 1001 -> "4 Privet Dr" 1002 -> "221B Baker Street" 1003 -> "Milkman Road" 1002 -> "21 Jump St" 1001 -> "Paper St" 1001 -> "Paper Road 21" ``` There are three different options to define for how long to retain the messages: * **infinite message retention**: all changes to user's address are maintained in the logs. This can lead to the log growing in size without a bound. This option involves the risk of outgrowing the disk capacity. * **simple message retention**: older records are deleted after they reach a certain age or size. * **compacted topic**: only latest version of the key's value is kept. In the example above, only the current address for a specific user is kept. With compacted topics, Apache Kafka removes any records from the topics for which there is a newer version (based on the record key) is available in the partition. This retention policy can be set per-topic, so a single cluster can have some topics where retention is enforced by size or time and other topics where retention is enforced by compaction. warning The compaction occurs **per partition**: if two records with the same key land in different partitions, they will not be compacted. This usually doesn't happen since the record key is used to select the partition. However, for custom message routing this might be an issue. ## Compacted topic example[​](#compacted-topic-example "Direct link to Compacted topic example") To understand better how compaction works, we will look at a partition of a compacted topic before and after compaction has been applied. Continuing the example above, the topic records before the compaction would be: | Offset | Key | Value | | ------ | ---- | ----------------- | | 1 | 1001 | 4 Privet Dr | | 2 | 1002 | 221B Baker Street | | 3 | 1003 | Milkman Road | | 4 | 1002 | 21 Jump St | | 5 | 1001 | Paper St | | 6 | 1001 | Paper Road 21 | You can notice that there are some records with duplicate keys (`1001` and `1002`), with the records having offset `4`, `5`, and `6` being the addresses updates. When applying compaction, we only keep records with the latest offset (newest values) and the older ones get discarded. The end result is the following: | Offset | Key | Value | | ------ | ---- | ------------- | | 3 | 1003 | Milkman Road | | 4 | 1002 | 21 Jump St | | 6 | 1001 | Paper Road 21 | ## Compacted topic details[​](#compacted-topic-details "Direct link to Compacted topic details") A compacted topic consists of a **head** and a **tail**: * The **head** is a traditional Apache Kafka topic where new records are appended. The head can contain duplicated keys. * The **tail** contains one record per key. Apache Kafka compaction ensures that keys are unique in the tail. Expanding the example above, let's assume that the **tail** contains the following entries: | Offset | Key | Value | | ------ | ---- | ----------------- | | 1 | 1001 | 4 Privet Dr | | 2 | 1002 | 221B Baker Street | | 3 | 1003 | Milkman Road | And the **head**, the newer records including duplicated keys: | Offset | Key | Value | | ------ | ---- | ------------- | | 4 | 1002 | 21 Jump St | | 5 | 1001 | Paper St | | 6 | 1001 | Paper Road 21 | During the compaction process, Apache Kafka creates a structure called **offset map** for the records in the head section, containing for each key, the latest offset. | Key | Offset | | ---- | ------ | | 1002 | 4 | | 1001 | 6 | The compaction thread then scans the **tail**, removing every record having a key that is also present in the **offset map** with an higher offset. | Offset | Key | Value | | ------ | --------------- | ----------------- | | 1 | 1001 (`delete`) | 4 Privet Dr | | 2 | 1002 (`delete`) | 221B Baker Street | | 3 | 1003 | Milkman Road | Lastly, the records in the offset map are added in the tail. | Offset | Key | Value | | ------ | ---- | ------------- | | 3 | 1003 | Milkman Road | | 4 | 1002 | 21 Jump St | | 6 | 1001 | Paper Road 21 | Related pages * [Configure log cleaner for topic compaction](/docs/products/kafka/howto/configure-log-cleaner.md) --- # Monitoring consumer groups in Aiven for Apache Kafka® With Aiven for Apache Kafka® dashboards and telemetry, you can monitor the performance and system resources of your Aiven for Apache Kafka service. Aiven provides pre-built dashboards and telemetry for your service, allowing you to collect and visualize telemetry data using InfluxDB® and Grafana®. Aiven streamlines the process by automatically configuring the dashboards for each of your Aiven for Apache Kafka instances. This section builds on the [service integrations](/docs/platform/concepts/service-integration.md) documentation and provides an in-depth look at consumer group graphs and related key terminology in Aiven for Apache Kafka®. Consumer group graphs offer valuable insights into the behavior of Apache Kafka consumers, which is crucial for maintaining a continuously running production Kafka system. ## Topics[​](#topics "Direct link to Topics") In Apache Kafka®, a topic serves as a unique channel for discussions. Producers send messages to the topic while consumers read those messages. For instance, in a topic named `soccer`, you can read what others say about soccer (acting as a consumer) or post messages about soccer (acting as a producer). ## Topic partitions[​](#topic-partitions "Direct link to Topic partitions") The storage of messages for an Apache Kafka® topic can be spread across one or more topic partitions. For instance, in a topic that has 100 messages and is set up to have 5 partitions, 20 messages would be assigned to each partition. ## Consumer groups[​](#consumer-groups "Direct link to Consumer groups") Apache Kafka® allows multiple consumers to read messages from a Kafka topic. This improves the message consumption rate and overall performance. Organizing consumers into consumer groups identified by a group ID is common practice. Consumer groups consume messages from a topic with messages spread across multiple partitions. Apache Kafka ensures that each message is consumed by only one consumer, which is essential for certain classes of business applications. For example, with a topic having 100 messages across five partitions and five consumers in a consumer group, each consumer will be allocated a distinct partition, consuming 20 messages each. If the number of consumers exceeds the number of partitions, extra consumers remain idle until an active consumer exits. Also, a consumer cannot read from a partition not assigned to it. ## Consumer group telemetry[​](#consumer-group-telemetry "Direct link to Consumer group telemetry") Aiven for Apache Kafka provides built-in consumer group graphs that offer valuable telemetry to monitor and manage consumer groups effectively. ### Consumer group graph: consumer group replication lag[​](#consumer-group-graph-consumer-group-replication-lag "Direct link to Consumer group graph: consumer group replication lag") Consumer group lag is an important metric in your Apache Kafka dashboard. It shows how far behind the consumers in a group are in consuming messages on the topic. A significant lag can indicate one of two scenarios - terminated consumers or consumers who are alive but unable to keep up with the rate of incoming messages. Persistent lag for long durations may indicate that the system is not behaving according to plan, requiring investigation and follow-up actions to resolve the issue. The terms `Consumer group lag` and `Consumer group replication lag` can be used interchangeably. Consumer Group Lag is typically a metric provided by the client side, while Aiven computes its metric known as Consumer Group Replication Lag (`kafka_consumer_group_rep_lag`) by fetching information about partitions and consumer groups from broker side. This metric captures the difference between the latest published offset (high watermark) and the consumer group offset for the same partition. The consumer group graph below, which is enabled by default, provides valuable insights into consumer behavior. It displays the consumer group replication lag, indicating how far behind the consumers are in consuming messages from a topic. This graph provides information about consumer behavior, enabling you to take appropriate action if necessary. ![Image of consumer group replication lag](/docs/assets/images/consumer-group-graphs-for-kafka-dashboards-c548ba8c8cc6966962ae090018dd3ced.png) ### Consumer group offset telemetry[​](#consumer-group-offset-telemetry "Direct link to Consumer group offset telemetry") In Apache Kafka, messages are written into a partition as append-only logs and each message is assigned a unique incremental number called the offset. These offsets indicate the exact position of messages within the partition. Aiven for Apache Kafka provides offset telemetry, which can help understand message consumption patterns and troubleshoot issues. The `kafka_consumer_group_offset` metric identifies the consumer group's most recent committed offset, which can be used to determine its relative position within the assigned partitions. --- # Partition segments Apache Kafka® divides topics partition data into **segment** files (with `.log` suffix) stored on the file system. Each segment file is named using the offset of the first message (a.k.a. **base offset**) contained. For example, the segment file `04.log` contains the message with offset `4` as first entry. The last segment in the partition is called the **active segment** and it is the only segment to which new messages are appended to. In the example above the topic partition is divided into multiple segments, with `06.log` being the active one: **01.log** | Offset | Key | Value | | ------ | ---- | ----------------- | | 1 | 1001 | 4 Privet Dr | | 2 | 1002 | 221B Baker Street | | 3 | 1003 | Milkman Road | **04.log** | Offset | Key | Value | | ------ | ---- | ---------- | | 4 | 1002 | 21 Jump St | | 5 | 1001 | Paper St | **06.log (active segment)** | Offset | Key | Value | | ------ | ---- | ------------- | | 6 | 1001 | Paper Road 21 | When the segment file reaches a certain size or age, Apache Kafka will create a segment file. This can be controlled by the following settings: * `segment.bytes` : creates a new segment when current segment becomes greater than this size. This setting can be set during topic creation and defaults to 1 GB. * `segment.ms` : forces the segment to roll over and create one when the segment becomes older than this value. note Aiven for Apache Kafka® service is fully manageable, the partitions are rebalanced automatically when the cluster is scaled up or down. --- # Guarantees With Aiven for Apache Kafka®'s tiered storage, there are two primary types of data retention guarantees: **total retention** and **local retention**. **Total retention**: Tiered storage ensures that your data remains available up to the limit defined by the total retention threshold, whether stored locally or remotely. This means your data is not deleted until reaching the total retention threshold, regardless of storage location. **Local retention**: Log segments are only removed from local storage after being successfully uploaded to remote storage, even if the data exceeds the local retention threshold. ## Example[​](#example "Direct link to Example") Let's say you have a topic with a **total retention threshold** of **1000 GB** and a **local retention threshold** of **200 GB**. This means that: * All data for the topic is retained, whether stored locally or remotely, as long as the total size does not exceed 1000 GB. * If tiered storage is enabled per topic, older segments are uploaded immediately to remote storage, regardless of whether the local retention threshold of 200 GB is exceeded. Data is deleted from local storage only after it has been safely transferred to remote storage. * If the total size of the data exceeds 1000 GB, Aiven for Apache Kafka begins deleting the oldest data from remote storage. Related pages * [Tiered storage in Aiven for Apache Kafka® overview](/docs/products/kafka/concepts/kafka-tiered-storage.md) * [How tiered storage works in Aiven for Apache Kafka®](/docs/products/kafka/concepts/tiered-storage-how-it-works.md) * [Enabled tiered storage for Aiven for Apache Kafka® service](/docs/products/kafka/howto/enable-kafka-tiered-storage.md) --- # How tiered storage works in Aiven for Apache Kafka® Aiven for Apache Kafka® tiered storage optimizes data management across two distinct storage tiers: * **Local tier**: Uses faster, typically more expensive storage solutions like solid-state drives (SSDs). * **Remote tier**: Uses slower, cost-effective options like cloud object storage. In Aiven for Apache Kafka's tiered storage architecture, **remote storage** refers to storage options external to the Apache Kafka broker's local disk. This typically includes cloud-based or self-hosted object storage solutions like AWS S3 and Google Cloud Storage. Although network-attached block storage solutions like AWS EBS are technically external to the Apache Kafka broker, Apache Kafka considers them local storage within its tiered storage architecture. Tiered storage operates seamlessly for both Apache Kafka producers and consumers, meaning they interact with Apache Kafka the same way, whether tiered storage is enabled or not. In Classic Kafka services, you can configure tiered storage per topic, including local retention settings that control how much data is kept on broker disks before offload. Standard Kafka services manage storage for classic topics automatically. ## Local vs. remote data retention[​](#local-vs-remote-data-retention "Direct link to Local vs. remote data retention") When tiered storage is enabled, data produced to a topic is initially stored on the local disk of the Apache Kafka broker. All but the active (open) segment is then asynchronously and orderly transferred to remote storage, regardless of the local retention settings. During periods of high data ingestion or transient errors, such as network connectivity issues, local storage might temporarily hold more data than specified by the local retention threshold. This continues until the remote tier is back online, at which point the excess local data is transferred to remote storage, and local segments exceeding the retention threshold are removed. ![Diagram depicting the concept of local vs. remote data retention in a tiered storage system](/docs/assets/images/data-retention-a9143bcfa4270c60b372282c156cb6c3.png) ## Segment management[​](#segment-management "Direct link to Segment management") Data is organized into segments, which are uploaded to remote storage individually. The active (newest) segment remains in local storage, so the segment size can also influence local data retention. For example, if the local retention threshold is 1 GB, but the segment size is 2 GB, the local storage exceeds the 1 GB limit until the active segment is rolled over and uploaded to remote storage. ## Asynchronous uploads and replication[​](#asynchronous-uploads-and-replication "Direct link to Asynchronous uploads and replication") Data is transferred to remote storage asynchronously and does not interfere with the producer activity. While the Apache Kafka broker aims to move data as swiftly as possible, certain conditions, such as high ingestion rate or connectivity issues can temporarily cause local storage to exceed the limits set by the local retention policy. The log cleaner does not purge any data exceeding the local retention threshold until it is successfully uploaded to remote storage. The replication factor is not considered during the upload process, and only one copy of each segment is uploaded to the remote storage. Most remote storage options have their own measures, including data replication to ensure data durability. ## Data retrieval[​](#data-retrieval "Direct link to Data retrieval") When consumers fetch records stored in remote storage, the Apache Kafka broker downloads and caches these records locally. This allows for quicker access in subsequent retrieval operations. Aiven allocates a small amount of disk space, ranging from 2 GB to 16 GB, equivalent to 5% of the Apache Kafka broker's total available disk for the temporary storage of fetched records. ## Security[​](#security "Direct link to Security") Segments are encrypted with 256-bit AES encryption before being uploaded to the remote storage. The encryption keys are not shared with the cloud storage provider and generally do not leave Aiven management machines and the Kafka brokers. Related pages * [Tiered storage in Aiven for Apache Kafka® overview](/docs/products/kafka/concepts/kafka-tiered-storage.md) * [Guarantees](/docs/products/kafka/concepts/tiered-storage-guarantees.md) * [Limitations](/docs/products/kafka/concepts/tiered-storage-limitations.md) * [Enabled tiered storage for Aiven for Apache Kafka® service](/docs/products/kafka/howto/enable-kafka-tiered-storage.md) --- # Trade-offs and limitations The main trade-off of tiered storage is the higher latency when accessing and reading data from remote storage compared to local disk storage. Adding local caching can partially mitigate this issue, but it cannot eliminate the latency entirely. ## Limitations[​](#limitations "Direct link to Limitations") * Tiered storage does not support compacted topics. * After tiered storage is enabled, you cannot disable it. As a workaround to keep all data locally, set the local retention to `-2` (the default value).For further assistance, contact [Aiven support](mailto:support@aiven.io). * Increasing the local retention threshold does not move segments already uploaded to remote storage back to local storage. This change only affects new data segments. * If tiered storage is enabled, you cannot migrate the service to a different region or cloud, except to a virtual cloud in the same region. For migration assistance, contact [Aiven support](mailto:support@aiven.io). * Tiered storage is only available on AWS, GCP, and Azure. * If you power off your service with tiered storage enabled, you will permanently lose all remote data. Charges for tiered storage are not applicable while the service is off. * Tiered storage is not available on all plans and regions. Check the [plans and pricing page](https://aiven.io/pricing?product=kafka) for supported plans and regions. Related pages * [Tiered storage in Aiven for Apache Kafka® overview](/docs/products/kafka/concepts/kafka-tiered-storage.md) * [Enabled tiered storage for Aiven for Apache Kafka® service](/docs/products/kafka/howto/enable-kafka-tiered-storage.md) * [Power on/off and delete your Aiven for Apache Kafka® service](/docs/products/kafka/howto/power-cycle-service.md) --- # Aiven for Apache Kafka® topic catalog [Early availability](/docs/platform/concepts/service-and-feature-releases.md) The Aiven for Apache Kafka® topic catalog provides a centralized interface within the Aiven console to view and manage Apache Kafka topics across different projects and services. This unified view consolidates all topics into one accessible location, streamlining the management of your Aiven for Apache Kafka infrastructure. ## Key features and benefits[​](#key-features-and-benefits "Direct link to Key features and benefits") * **Centralized management**: Access all Apache Kafka topics across your organization and projects from a single interface. * **Search topics**: Find specific Apache Kafka topics using the search bar. * **Topic information at a glance**: View topic names, associated services, and projects in either table view or tile view. The owner column is displayed when governance is enabled and shows the group owning the topic. ## Governance on topic catalog[​](#governance-on-topic-catalog "Direct link to Governance on topic catalog") With [governance enabled](/docs/products/kafka/howto/enable-governance.md) for your organization, you can efficiently manage topic ownership and creation requests. You can also access and govern topics across your organization, even if you aren't a member of all projects. * **Claim topic ownership**: Request to claim ownership of individual topics not currently owned by your group. Additionally, select multiple topics and submit a batch claim request for ownership transfer. * **Topic creation requests and approvals**: Submit requests for creating new topics, which go through an approval process to ensure proper governance. ## Related page[​](#related-page "Direct link to Related page") * [Manage Apache Kafka topics with topic catalog](/docs/products/kafka/howto/view-kafka-topic-catalog.md) * [Aiven for Apache Kafka governance](/docs/products/kafka/concepts/governance-overview.md) --- # Apache Kafka® upgrade procedure Aiven for Apache Kafka® provides an automated upgrade process during the following operations: * Maintenance updates * Plan changes * Cloud region migrations * Manual node replacements by an Aiven operator Aiven creates new broker nodes to replace the existing ones. ## How upgrades work[​](#how-upgrades-work "Direct link to How upgrades work") The following steps show the upgrade process for a 3-node Apache Kafka service: ![3-node Kafka service](/docs/assets/images/kafka-cluster-overview-d0fca5f8ed9bec3617419b7220ca2b02.png) 1. **Start new nodes:** New Apache Kafka® nodes start alongside existing ones. 2. **Join the cluster:** The new nodes join the Apache Kafka cluster once they are running, and the cluster temporarily contains a mix of old and new nodes. 3. **Transfer data and leadership:** The partition data and leadership are transferred to new nodes. ![Kafka cluster illustration](/docs/assets/images/kafka-cluster-overview-upgraded-ced043667d57e430dbf15677cd3d4a72.png) warning This step is CPU intensive because of the additional data movement. 4. **Retire old nodes:** Old nodes are removed after their data is transferred. note The number of new nodes added depends on the cluster size. By default, up to 6 nodes are replaced at a time during the upgrade. 5. **Finish upgrade**: The process completes when all old nodes are removed. ![Kafka cluster new node illustration](/docs/assets/images/kafka-cluster-overview-final-2e3a411bfa23b4dc78590743de3c20da.png) ## Service availability during upgrades[​](#service-availability-during-upgrades "Direct link to Service availability during upgrades") Your Aiven for Apache Kafka service remains available during upgrades. All active nodes stay operational, and clients can continue to connect. During upgrades, expect: * Reduced performance: Data transfers between nodes can temporarily lower cluster performance, especially on heavily loaded clusters. * Temporary client warnings: You might see `leader not found` warnings in application logs during partition leadership changes. * Automatic recovery: Most Kafka client libraries retry automatically and handle these warnings without manual action. These effects are temporary and resolve as the upgrade completes. For troubleshooting, see [NOT\_LEADER\_FOR\_PARTITION errors](/docs/products/kafka/troubleshooting/non-leader-for-partition.md). ## Upgrade duration[​](#upgrade-duration "Direct link to Upgrade duration") Upgrade duration depends on several factors: * **Data volume:** Larger datasets take longer to process. * **Number of partitions:** Each partition adds processing overhead. * **Cluster load:** Heavily loaded clusters have fewer resources available for upgrades. To reduce upgrade times: * Schedule upgrades during low-traffic periods. * Pause non-essential workloads to free up resources. ## Rollback options[​](#rollback-options "Direct link to Rollback options") Rollback (reverting to a previous Kafka version) is not available because old nodes are removed during the upgrade. note Nodes holding data are not removed until data transfer is complete, preventing data loss. If the upgrade does not progress, old nodes remain in the cluster. If sufficient disk capacity is available, you can downgrade to a smaller plan. Use the [Aiven Console](/docs/products/kafka/howto/change-service-plan.md) or the [Aiven CLI](/docs/tools/cli/service-cli.md#avn-cli-service-update) to perform the downgrade. The upgrade process remains the same when changing the node type during a service plan change. For instance, when downgrading to a plan with fewer resources, such as fewer CPUs, memory, or disk space, the latest system software versions are applied to all new nodes, and data is transferred accordingly. ## Upgrade impact and risks[​](#upgrade-impact-and-risks "Direct link to Upgrade impact and risks") Upgrades can increase CPU usage because of partition leadership changes and data transfers to new nodes. To reduce the risk of disruptions: * Schedule upgrades during low-traffic periods. * Pause non-essential producers and consumers to minimize cluster load. If you change to a smaller plan, the disk may reach the [maximum allowed limit](https://aiven.io/docs/products/kafka/howto/prevent-full-disks), which can block the upgrade. Check disk usage before starting the upgrade and make sure there is enough free space. note In critical situations, Aiven's operations team can temporarily add extra storage to the old nodes. ## Upgrade to Apache Kafka® 4.0[​](#upgrade-to-apache-kafka-40 "Direct link to Upgrade to Apache Kafka® 4.0") To upgrade to Apache Kafka® 4.0 or later, your service must first upgrade to Kafka 3.9 and migrate to [KRaft mode](/docs/products/kafka/concepts/kraft-mode.md). * Services running Apache Kafka 3.8 or earlier (ZooKeeper-based) cannot upgrade directly to 4.0. They must first upgrade to 3.9, which performs the ZooKeeper-to-KRaft migration. * After migrating to KRaft in 3.9, the service can upgrade to 4.0 or later. * Once a service migrates to KRaft, rollback to ZooKeeper is not possible. note The `message.format.version` configuration is deprecated in Kafka 3.x and removed in Kafka 4.0. Remove this configuration from your topics before upgrading to Kafka 4.0. ### Configuration changes in Kafka 4.0[​](#configuration-changes-in-kafka-40 "Direct link to Configuration changes in Kafka 4.0") Kafka 4.0 removes and replaces some configuration settings. Update these before starting the upgrade: * `message.format.version` (topic-level): Remove this setting from all topic and service-level configurations. It is deprecated in Kafka 3.x and removed in Kafka 4.0. * `log.message.timestamp.difference.max.ms` (service-level): Use `log.message.timestamp.before_max_ms` and `log.message.timestamp.after_max_ms` instead. These settings define the acceptable timestamp range for messages. Update your configurations before upgrading to avoid validation errors. ## Transitioning to KRaft[​](#transitioning-to-kraft "Direct link to Transitioning to KRaft") With the release of Apache Kafka® 3.9, Aiven introduces support for Apache Kafka Raft (KRaft), the new consensus protocol for Kafka metadata management. KRaft simplifies the architecture while keeping compatibility with existing features and integrations, including Aiven for Apache Kafka Connect, Aiven for Apache Kafka MirrorMaker 2, and Aiven for Karapace. Kafka 3.9 includes all features from Kafka 3.8. Some controller metrics are no longer available due to the transition to KRaft mode. For details, see [Apache Kafka controller metrics](/docs/products/kafka/reference/kafka-metrics-prometheus.md#kraft-mode-and-metrics-changes). ACL permissions and governance behaviors remain unchanged. For a detailed overview of how KRaft mode works and how it differs from ZooKeeper-based metadata management, see [KRaft in Aiven for Apache Kafka®](/docs/products/kafka/concepts/kraft-mode.md). ### Availability and migration[​](#availability-and-migration "Direct link to Availability and migration") #### New services[​](#new-services "Direct link to New services") * All new Aiven for Apache Kafka services with Kafka 3.9 run KRaft as the default metadata management protocol. * Startup-4 replaces Startup-2 plans in Kafka 3.9 and later. All feature restrictions from Startup-2 also apply to Startup-4, including Datadog restrictions. * Available on all cloud providers. #### Existing services[​](#existing-services "Direct link to Existing services") * Migration from ZooKeeper to KRaft is part of the upgrade from Apache Kafka 3.x to 3.9. This migration will be available soon. * Aiven will notify you when your service becomes eligible for migration. For details, see [Migration from ZooKeeper to KRaft](/docs/products/kafka/concepts/kraft-mode.md#migration-from-zookeeper-to-kraft). * After migrating to Kafka 3.9 in KRaft mode, you can upgrade to Kafka 4.0 or later. * Aiven has extended support for Apache Kafka 3.8 to give you more time to complete the migration. For the EOL date, see [End of life for major versions](/docs/platform/reference/eol-for-major-versions.md#aiven-for-kafka). #### Performance impact[​](#performance-impact "Direct link to Performance impact") Performance testing across different cloud providers and plan sizes has not shown significant changes. --- # Create an Aiven for Apache Kafka® Developer tier service Create an Aiven for Apache Kafka® Developer tier service in the [Aiven Console](https://console.aiven.io) or with Skills. Use it when the [Free tier](/docs/products/kafka/free-tier/kafka-free-tier.md) no longer meets your needs, or when you require capabilities not available on the Free tier, such as [service integrations](/docs/platform/concepts/service-integration.md) or [Aiven for Apache Kafka® Connect](/docs/products/kafka/kafka-connect/get-started.md). For quotas, pricing, and features, see [Aiven for Apache Kafka® Developer tier](/docs/products/kafka/dev-tier/kafka-dev-tier.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven account. Sign up or sign in using the [Aiven Console](https://console.aiven.io). * A project where you can create paid services. * If you use paid services, set up billing for your organization. See [Billing and payment](/docs/platform/concepts/billing-and-payment.md). * To create the service with Skills, install the [Aiven CLI](/docs/tools/cli.md) and authenticate. You also need [Node.js](https://nodejs.org/) to run `npx` commands. ## Choose how to create the service[​](#choose-how-to-create-the-service "Direct link to Choose how to create the service") Use one of the following: * **Aiven Console**: Guided setup in the browser. * **Skills**: Command-line setup using `npx` and the Aiven CLI. ## Create a Developer tier service using the Aiven Console[​](#create-a-developer-tier-service-using-the-aiven-console "Direct link to Create a Developer tier service using the Aiven Console") 1. In your project, click **Services**. 2. Click **Create service**. 3. Select **Aiven for Apache Kafka®**. 4. In **Service tier**, select **Developer**. 5. In **Cloud**, select your geographical region from the broad areas shown, such as North America or Europe. The **Service summary** shows this under **Cloud**. note On the Developer tier, **Cloud** lists broad geographical areas only. To select a specific cloud provider and region, use the **Professional** tier. For details, see the message in the **Cloud** section of the Aiven Console. 6. In **Retention**, select a retention period between **1 day** and **3 days**. note Your estimated monthly cost depends on **retention**, **Cloud**, and geographical region. Check the **Service summary** for the price that applies to your selections. Retention options are 1 day, 2 days, or 3 days, with pricing adjusted accordingly. 7. In **Service basics**, enter a **Service name**. 8. Review the **Service summary** and estimated monthly cost. 9. Click **Create service**. When the status is **Running**, your Kafka service is ready. ## Create a Developer tier service using Skills[​](#create-a-developer-tier-service-using-skills "Direct link to Create a Developer tier service using Skills") Use Skills to create and configure the Kafka service from the command line. 1. Install the Skills bundle: ``` npx skills add Aiven-Open/aiven-skills-bundle ``` If your environment supports Skills, you can invoke them from compatible tools, such as AI-assisted editors or agents, by describing the task, for example: `Create a test Kafka service in Aiven`. 2. Run the Kafka service creation Skill: ``` npx skills run kafka-create-service ``` 3. Follow the prompts to select your project, **Developer** tier, and **Cloud**. The Skill might ask for cloud values that the Aiven CLI needs. Select the values that the Skill lists for the Developer tier. The Skill creates and configures the service. When the run completes, the Kafka service is ready to use. For the full Skills workflow, see [Set up Aiven for Apache Kafka® using Skills](/docs/products/kafka/howto/set-up-kafka-with-skills.md). ## After you create the service[​](#after-you-create-the-service "Direct link to After you create the service") When the service status is **Running**, you can: * **Connect a client**: On your service's **Overview** page, open **Quick connect** for connection details and sample code. See [Connect to Aiven for Apache Kafka®](/docs/products/kafka/howto/list-code-samples.md). * **Run a sample workload**: Use the sample data generator to test producing and consuming messages. See [Generate sample data for Aiven for Apache Kafka®](/docs/products/kafka/howto/generate-sample-data.md). * **Create topics**: Keep topic and partition counts within the [Developer tier limits](/docs/products/kafka/dev-tier/kafka-dev-tier.md#limits-and-specifications). See [Create a Kafka topic](/docs/products/kafka/howto/create-topic.md). * **Run connectors**: Add an [Aiven for Apache Kafka® Connect](/docs/products/kafka/kafka-connect/get-started.md) service and integrate it with your Kafka service. See the [list of available connectors](/docs/products/kafka/kafka-connect/concepts/list-of-connector-plugins.md). ## Upgrade the service[​](#upgrade-the-service "Direct link to Upgrade the service") Upgrade to the Professional tier when the Developer tier no longer meets your throughput, storage, or scaling needs. For upgrade considerations, including Kafka Connect compatibility, see [Upgrade path](/docs/products/kafka/dev-tier/kafka-dev-tier.md#upgrade-path). You can upgrade from the service overview or from service settings. ### Before you upgrade[​](#before-you-upgrade "Direct link to Before you upgrade") If your Kafka service is integrated with a Kafka Connect service: * Power off the Kafka Connect service before upgrading Kafka. You cannot upgrade Kafka while an integrated Kafka Connect service is running. * After upgrading Kafka, upgrade the Kafka Connect service to the same tier before powering it on again. Run both services on the same tier. For details, see [Kafka and Kafka Connect tier compatibility](/docs/products/kafka/dev-tier/kafka-dev-tier.md#kafka-and-kafka-connect-tier-compatibility). ### Upgrade from the service overview[​](#upgrade-from-the-service-overview "Direct link to Upgrade from the service overview") 1. Open your service's **Overview** page. 2. In **Service usage**, click **Upgrade**. 3. Select **Professional** tier. 4. In **Cloud**, select a cloud provider, geographical region, and plan. 5. Click **Upgrade service**. ### Upgrade from service settings[​](#upgrade-from-service-settings "Direct link to Upgrade from service settings") 1. In your service, click **Service settings**. 2. In **Service summary**, click **Upgrade**. 3. Select **Professional** tier. 4. In **Cloud**, select a cloud provider, geographical region, and plan. 5. Click **Upgrade service**. ### Limitations[​](#limitations "Direct link to Limitations") * Downgrading from Developer tier to the Free tier is not supported. * After upgrading to the Professional tier, available downgrade options depend on the selected plan. Review supported changes in the [Aiven Console](https://console.aiven.io). Related pages * [Developer tier overview](/docs/products/kafka/dev-tier/kafka-dev-tier.md) * [Connect to Aiven for Apache Kafka®](/docs/products/kafka/howto/list-code-samples.md) * [Generate sample data for Aiven for Apache Kafka®](/docs/products/kafka/howto/generate-sample-data.md) * [Create a Kafka topic](/docs/products/kafka/howto/create-topic.md) --- # Aiven for Apache Kafka® Developer tier Aiven for Apache Kafka® Developer tier is a paid service tier for Classic Kafka that sits between the [Free tier](/docs/products/kafka/free-tier/kafka-free-tier.md) and Professional tier. Use it for development, prototyping, testing, or production workloads that need more throughput, more topics, longer retention, more storage, or [service integrations](/docs/platform/concepts/service-integration.md) than the Free tier provides, without committing to a Professional plan. [Aiven for Apache Kafka® Connect](/docs/products/kafka/kafka-connect/get-started.md) is available as a separately billed service. For capacity, pricing, and service limits, see [Limits and specifications](#limits-and-specifications). ## When to use the Developer tier[​](#when-to-use-the-developer-tier "Direct link to When to use the Developer tier") Use the Developer tier when: * You need higher throughput, more topics, longer retention, or more storage than the [Free tier](/docs/products/kafka/free-tier/kafka-free-tier.md) offers. * You need a paid Aiven for Apache Kafka® Connect service or [service integrations](/docs/platform/concepts/service-integration.md), neither of which is available on the Free tier. * You need a paid environment for development, prototyping, testing, or smaller-scale production workloads without committing to a Professional plan. ## What the Developer tier includes[​](#what-the-developer-tier-includes "Direct link to What the Developer tier includes") A Developer tier service includes: * Apache Kafka® on two nodes with fixed compute and storage. * Karapace Schema Registry and REST Proxy. * A [sample data generator](/docs/products/kafka/howto/generate-sample-data.md) in the console for testing producers, consumers, and Karapace Schema Registry. * Quick connect from the Kafka service overview page in the Aiven Console using **Set up your Stream** or **Connection information** to get connection parameters and sample client code. For details, see [Connect to Aiven for Apache Kafka®](/docs/products/kafka/howto/list-code-samples.md). * Supported [service integrations](/docs/platform/concepts/service-integration.md), including Aiven for Apache Kafka® Connect and Prometheus. ### Automate workflows with Skills[​](#automate-workflows-with-skills "Direct link to Automate workflows with Skills") Skills are ready-to-run workflows that automate common Kafka tasks, such as: * Creating a Kafka service and configuring it end-to-end. * Creating topics, users, and permissions. * Generating sample data and validating the deployment. * Configuring connectors and integrations. Skills are available on Developer tier and Professional tier services. Run them with [`npx skills`](/docs/products/kafka/howto/set-up-kafka-with-skills.md) or in another environment that supports Skills, such as an editor or AI assistant. Skills use [Aiven CLI](/docs/tools/cli.md) internally and require an authenticated CLI session. For details, see [Set up Aiven for Apache Kafka® using Skills](/docs/products/kafka/howto/set-up-kafka-with-skills.md). Manage Developer tier services from the [Aiven Console](https://console.aiven.io), the [Aiven CLI](/docs/tools/cli.md), or the [Aiven API](/docs/tools/api.md). ## Limits and specifications[​](#limits-and-specifications "Direct link to Limits and specifications") The following specifications apply to Developer tier services. | Specification | Developer tier | | ---------------------- | -------------------------------------------------------------------------------------------- | | Price | Starts at $35 per month for 1 day retention; pricing varies by cloud and geographical region | | Nodes | 2 | | Throughput | 1 MB/s ingress, 2 MB/s egress | | Topics | Up to 20 | | Partitions | Up to 100 per topic and 2000 max across the service | | Replication factor | 1 | | Metadata mode | KRaft | | Retention | 1, 2, or 3 days. Pricing varies by option and is shown in the Aiven Console | | Storage | Fixed local storage per node | | Kafka Connect | Optional service, billed separately | | Cloud | Geographical region in the Aiven Console. Pricing varies by cloud | | Service integrations | Kafka Connect and Prometheus | | Upgrade | To a Professional tier from the Aiven Console | | Downgrade to Free tier | Not supported | Developer tier pricing depends on retention, cloud, and geographical region. Review current pricing during service creation in the [Aiven Console](https://console.aiven.io) and on [Aiven for Apache Kafka® pricing](https://aiven.io/pricing?product=kafka). For higher throughput, storage, retention, or connector limits, upgrade your service to a Professional plan in the [Aiven Console](https://console.aiven.io). ## Compare service tiers[​](#compare-service-tiers "Direct link to Compare service tiers") | Feature | Free | Developer | Professional | | ----------------------------- | --------------------------------- | ----------------------------- | ---------------------------------------- | | Price | $0 | Starts at $35 per month | Varies by plan | | Throughput | Up to 250 KB/s ingress and egress | 1 MB/s ingress, 2 MB/s egress | Higher, plan-dependent | | Topics | Up to 5 | Up to 20 | Plan-dependent | | Partitions | 1 per topic | Up to 100 per topic | Plan-dependent | | Nodes | 1 | 2 | Plan-dependent | | Replication factor | 1 | 1 | Plan-dependent | | Metadata mode | KRaft | KRaft | KRaft | | Retention | Fixed | 1, 2, or 3 days | Plan-dependent | | Storage | Fixed | Fixed local storage per node | Plan-dependent | | SLA | None | 99% | Up to 99.99%, plan-dependent | | Kafka Connect | Not supported | Optional, billed separately | Full support, plan-dependent | | Service integrations | Not supported | Kafka Connect, Prometheus | Yes, plan-dependent | | Cloud and geographical region | Fixed | Geographical region only | Cloud and geographical region selectable | For replication factor 3 or similar production redundancy, use a Professional plan. Starting with Apache Kafka 3.9, new services use [KRaft](/docs/products/kafka/concepts/kraft-mode.md) for metadata instead of ZooKeeper. ## Upgrade path[​](#upgrade-path "Direct link to Upgrade path") Upgrade to the Professional tier when the Developer tier no longer meets your throughput, storage, or scaling needs. Professional tier supports higher throughput, storage, topic limits, and SLA. ### Kafka and Kafka Connect tier compatibility[​](#kafka-and-kafka-connect-tier-compatibility "Direct link to Kafka and Kafka Connect tier compatibility") Aiven for Apache Kafka® and Aiven for Apache Kafka® Connect services that are integrated with each other must run on the same service tier. Both services must be on the Developer tier or both on the Professional tier. This requirement affects how you upgrade integrated services. To upgrade a Developer tier Kafka service with an integrated Kafka Connect service to the Professional tier: 1. Power off the Kafka Connect service. This preserves the service and its configuration. 2. Upgrade the Kafka service to the Professional tier. 3. Upgrade the Kafka Connect service to the same tier. 4. Power on the Kafka Connect service. For step-by-step instructions, see [Upgrade the service](/docs/products/kafka/dev-tier/create-dev-tier-kafka-service.md#upgrade-the-service). note Developer tier services cannot be downgraded to the Free tier. Professional tier downgrade options depend on the selected plan. Review supported changes in the [Aiven Console](https://console.aiven.io). Related pages * [Create an Aiven for Apache Kafka® Developer tier service](/docs/products/kafka/dev-tier/create-dev-tier-kafka-service.md) * [Manage Apache Kafka® parameters](/docs/products/kafka/howto/set-kafka-parameters.md) * [Apache Kafka® upgrade procedure](/docs/products/kafka/concepts/upgrade-procedure.md) --- # Batching and delivery in diskless topics Diskless topics use a batching-based delivery model designed for cloud-native environments. Instead of writing messages to local disks, brokers batch messages in memory and upload them to object storage. This approach improves scalability, reduces cost, and simplifies the Kafka architecture while preserving ordering and delivery guarantees. ## Batching behavior[​](#batching-behavior "Direct link to Batching behavior") Diskless topics create batches based on time or size thresholds: * **Batch interval**: The broker flushes a batch after a configurable delay (typically 250 milliseconds). * **Batch size**: The broker flushes a batch once it reaches a specified size (commonly 8 MB). These defaults are optimized for throughput and cost-efficiency. You can adjust them based on your message rate, latency tolerance, and resource constraints. tip Low batch sizes or short flush intervals may increase object storage costs and degrade performance. Test custom configurations with realistic workloads before using them in production. ## Benefits of batching[​](#benefits-of-batching "Direct link to Benefits of batching") Writing individual messages is costly and inefficient. Object storage is optimized for large, infrequent writes. Batching enables diskless topics to: * Reduce the number of storage write operations * Minimize per-message storage overhead * Improve throughput * Lower total cost of ownership (TCO) ## Upload process[​](#upload-process "Direct link to Upload process") After a batch is finalized: 1. The broker uploads the batch to cloud storage as an object. 2. The broker registers the object's offset ranges and metadata with the internal Batch Coordinator. 3. The Batch Coordinator tracks partition ordering and ensures metadata consistency. Objects may contain messages from multiple partitions. Messages within an object are not ordered across partitions. Ordering is managed at the partition level. This separation of storage and coordination allows brokers to continue accepting new messages while finalizing storage operations in the background. ## Consumer delivery[​](#consumer-delivery "Direct link to Consumer delivery") Consumers use the standard Kafka fetch protocol to retrieve data from diskless topics: 1. The consumer requests data from a partition. 2. The broker uses metadata from the Batch Coordinator to locate the object containing the requested offset range. 3. The broker retrieves the object from the local object cache or fetches it from storage. 4. The broker extracts and returns the offset range to the consumer. Each object is cached by one broker per availability zone (AZ). Other brokers in the same AZ can access this cached data through inter-broker communication, enabling any broker to serve consumer requests efficiently. This caching and distribution model allows brokers to cooperate when serving consumers, which helps balance the load across the cluster. As the number of producers and consumers increases, the system continues to scale by distributing object fetches and cache usage across brokers. ## Delivery guarantees[​](#delivery-guarantees "Direct link to Delivery guarantees") Diskless topics provide the same delivery guarantees as classic Kafka topics configured with `acks=all`: * **Ordering**: Messages are delivered in the order they are produced within a partition. * **Durability**: Messages are stored in object storage, which provides built-in replication and durability. * **Integrity**: Messages are validated before being delivered to consumers. * **Consistency**: The Batch Coordinator manages total ordering and ensures metadata consistency. Regardless of the producer `acks` setting, diskless topics always acknowledge writes only after the data is successfully stored in object storage. ## Performance and tuning[​](#performance-and-tuning "Direct link to Performance and tuning") You can tune batching behavior to suit your workload: * Increase the batch size to reduce write frequency and storage API calls. * Decrease the batch interval to reduce end-to-end latency. * Monitor cache hit rates and consumer lag to optimize read performance. Tuning for very low latency (for example, under 10 milliseconds) or small object sizes (under 100 KB) can significantly increase costs and reduce performance benefits. ## Components involved in batching and delivery[​](#components-involved-in-batching-and-delivery "Direct link to Components involved in batching and delivery") Diskless topics use several internal components to manage batching and delivery efficiently: * **Batch Coordinator**: Manages partition-level metadata and total ordering. It tracks which messages are stored in which objects and their offset ranges. * **Object cache**: Caches recently accessed objects in memory or ephemeral disk to improve consumer fetch performance. * **Compaction jobs**: Merges smaller or older objects into larger ones. This improves trailing read efficiency and reduces the number of required storage operations. --- # Diskless topics for Apache Kafka® Diskless topics are a feature of Standard Kafka services for Aiven for Apache Kafka®. They store topic data in cloud object storage without writing to local disk. Diskless topics are available in Standard Kafka services on Aiven Cloud, and on Business and Premium `-inkless` plans in Bring Your Own Cloud (BYOC). In BYOC deployments, Aiven manages the Kafka service in your cloud account, while you retain control over your infrastructure and data. note Diskless topics are **limited availability** on Aiven Cloud. ## About diskless topics[​](#about-diskless-topics "Direct link to About diskless topics") Diskless topics store topic data in cloud object storage, such as Amazon S3, Google Cloud Storage (GCS), or Azure Blob Storage, instead of on broker disks. This design simplifies operations, reduces cross-availability zone (AZ) traffic, and supports cost-effective scaling. Data is batched and written to object storage. Partition metadata and message ordering are managed by an internal coordination layer that Aiven deploys and operates to support diskless topics. For details, see [Batching and delivery](/docs/products/kafka/diskless/concepts/batching-and-delivery.md). Diskless topics work with standard Kafka APIs and clients, and most applications do not require any changes to use them. If you use the Kafka CLI (`kafka-topics.sh`) to create or configure diskless topics, use the script from the [Inkless repository](https://github.com/aiven/inkless). For architectural details, see [Diskless topics architecture](/docs/products/kafka/diskless/concepts/diskless-topics-architecture.md). ## Benefits of using diskless topics[​](#benefits-of-using-diskless-topics "Direct link to Benefits of using diskless topics") Diskless topics are well suited for workloads that require high throughput and rapid scaling. They provide: * **Elastic scaling**: Supports high throughput and scales in seconds. * **No disk overruns**: Shifting to object storage removes broker disk capacity limits. * **Lower storage and network costs**: Reduces cross-availability zone traffic by offloading data to cloud object storage. * **Lower latency for hot data**: Frequently accessed data is cached on brokers to improve fetch performance. * **Simplified storage management**: No need to manage broker disks, rebalance partitions, or manually provision storage. * **Faster scaling and node replacement**: Removing large local disks reduces data movement during scaling and node rotation. * **Compliance and security**: In BYOC deployments, the service runs entirely within your own cloud account. note Internal and ecosystem topics (such as consumer offsets, Kafka Connect, MirrorMaker 2, and Schema Registry topics) are managed by the service. You cannot change their storage type or configuration. ## Diskless vs. classic Kafka topics[​](#diskless-vs-classic-kafka-topics "Direct link to Diskless vs. classic Kafka topics") Diskless topics store data in cloud object storage and do not rely on broker-managed replication or partition leadership. In Classic Kafka services, classic topics store data on broker-local disks and use standard Kafka replication. You can use both diskless and classic Kafka topics in the same Standard Kafka service. This allows you to: * Adopt diskless topics gradually. * Continue running workloads that require features not yet supported by diskless topics. * Maintain flexibility in your deployment strategy. For a detailed comparison, see [Compare diskless and classic topics](/docs/products/kafka/diskless/concepts/topics-vs-classic.md). For information about how compute, storage, and network usage are billed, see [Pricing for Aiven for Apache Kafka®](/docs/products/kafka/concepts/kafka-pricing.md). Related pages * [Create an Aiven for Apache Kafka® Professional tier service](/docs/products/kafka/get-started/create-kafka-service.md) * [Diskless topics architecture](/docs/products/kafka/diskless/concepts/diskless-topics-architecture.md) * [Diskless topics limitations](/docs/products/kafka/diskless/concepts/limitations.md) * [Create Apache Kafka® topics](/docs/products/kafka/howto/create-topic.md) --- # Diskless topics architecture Diskless topics extend the Apache Kafka® storage model by replacing local disk storage with cloud object storage. This approach reduces broker responsibilities and avoids inter-broker replication for diskless topics. The architecture introduces an internal metadata service (the Batch Coordinator) and relies on object storage for durability and scalability. ## How diskless topics work[​](#how-diskless-topics-work "Direct link to How diskless topics work") In diskless topics, brokers write messages in batches to object storage. Each batch is registered with the Batch Coordinator, which tracks its offset range and location. Consumers use this metadata to fetch message data directly from storage. Frequently accessed data may be cached on brokers to reduce latency. Apache Kafka’s partition and topic model remains unchanged. Diskless topics continue to support parallelism and ordering guarantees, using object storage for the data path and the coordinator for metadata. ![Producer and consumer flow in diskless topics](/docs/assets/images/schema-producer-consumer-9d80e33d678979daabc4a534db56df39.png) ## Leaderless data layer[​](#leaderless-data-layer "Direct link to Leaderless data layer") In diskless topics, partitions do not have leaders. Any broker in the cluster can read data from any diskless topic partition because all brokers access the same underlying object storage. This design eliminates inter-broker replication for diskless topic data and reduces operational complexity. Although the data path is leaderless, metadata still requires coordination. Diskless topics use the Batch Coordinator to manage this metadata. It tracks which data batches map to which offsets and where they are stored. To ensure consistency, the Batch Coordinator has a single leader that handles updates and preserves message order across the cluster. To reduce costs and latency, brokers are designed to access object storage and serve cached data within the same availability zone (AZ). This AZ-aware affinity minimizes cross-zone traffic and improves performance, especially in cloud environments where network costs between zones can be high. ## Role of object storage[​](#role-of-object-storage "Direct link to Role of object storage") For diskless topics, object storage replaces the local disk storage used in Classic Kafka services. Instead of writing log segments to disk, brokers batch messages and upload them to cloud object storage. This design shifts durability and replication responsibilities from Kafka to the cloud provider, reducing broker-to-broker data transfer and operational complexity. To improve performance, brokers cache recently read data. When a broker fetches an object, it temporarily stores the data to speed up future reads—especially within the same availability zone. ## Batch Coordinator and metadata[​](#batch-coordinator-and-metadata "Direct link to Batch Coordinator and metadata") The Batch Coordinator manages metadata for diskless topics. It assigns offsets to message batches, tracks where each batch is stored, and preserves message order within partitions. It does not handle message data directly. When you schedule or start maintenance for an Aiven Standard Kafka service that uses diskless topics, the internal Batch Coordinator is updated in the same maintenance window as the Kafka service. Maintenance details for both components are shown together in the service’s maintenance view in the Aiven Console. ## Services with mixed topics: classic and diskless[​](#services-with-mixed-topics-classic-and-diskless "Direct link to Services with mixed topics: classic and diskless") Standard Kafka services can include both classic and diskless topics in the same deployment: * In Classic Kafka services, classic topics store data on local disks managed by brokers. * In Standard Kafka services, classic topics use managed remote storage. * Diskless topics store data in cloud object storage. * Metadata for both topic types is shared using Kafka’s KRaft protocol. This setup supports gradual adoption. You can run classic and diskless topics side by side, depending on your use case. Some features, such as transactions, are supported only on classic Kafka topics. --- # Diskless topic limitations and behavior Diskless topics are compatible with Kafka APIs and clients, with some limitations: * Transactions are not supported for produce or consume operations. * Compacted topics are not supported. * Kafka Streams state stores are not supported. Stream processing can read from diskless topics, but use classic topics for writes at this time. * You cannot switch a diskless topic back to a classic topic. To switch a classic topic to diskless, see [Switch a classic topic to a diskless topic](/docs/products/kafka/howto/switch-topic-to-diskless.md). ## Internal metadata service behavior[​](#internal-metadata-service "Direct link to Internal metadata service behavior") Diskless topics rely on an internal metadata service that is managed automatically as part of the Kafka service. This service stores the metadata required for diskless topics to function. If a service has no diskless topics for two days, Aiven automatically disables diskless topics and turns off the metadata service. When you re-enable diskless topics, you might experience a short delay before you can create or use them while the metadata service starts. ### Aiven Cloud deployments[​](#aiven-cloud-deployments "Direct link to Aiven Cloud deployments") In Aiven Cloud deployments, this service does not appear as a separate service in the console or billing. ### Bring Your Own Cloud (BYOC) deployments[​](#bring-your-own-cloud-byoc-deployments "Direct link to Bring Your Own Cloud (BYOC) deployments") In BYOC deployments, enabling diskless topics automatically creates an Aiven for PostgreSQL® service in the project. This PostgreSQL service: * Is required for diskless topics to function. * Stores metadata required for diskless topics. * Appears as a separate service in the project. * Is created and managed automatically by Aiven. * Do not configure or manage independently. ### Maintenance behavior[​](#maintenance-behavior "Direct link to Maintenance behavior") Maintenance for this internal service occurs in the same maintenance window as the Kafka service. In the Aiven Console, references to internal components can appear during maintenance or upgrade flows, but you cannot manage them independently. For more information about how diskless topics work, see [Diskless topics architecture](/docs/products/kafka/diskless/concepts/diskless-topics-architecture.md). --- # Partitions and objects in Diskless Topics Diskless topics use the standard Kafka partitioning model but store data in cloud object storage instead of broker-local disks. Brokers batch messages and upload them as objects to the storage layer. ## Partitions[​](#partitions "Direct link to Partitions") Partitions in diskless topics behave the same as in classic Kafka topics. Each partition is an ordered append-only log of messages that supports message ordering, parallelism, and horizontal scalability. * Producers write to partitions based on a key or round-robin logic. * Consumers read from partitions independently, enabling concurrent processing. * The number of partitions controls how many producers or consumers can operate in parallel. ## Objects in diskless topics[​](#objects-in-diskless-topics "Direct link to Objects in diskless topics") In classic Kafka, partitioned data is stored in ordered segment files on broker disks. Diskless topics replace these segments with cloud-stored objects. Each object is a batch of messages that a broker uploads to cloud object storage. Unlike classic Kafka segments, an object is not limited to a single partition. It can include messages from multiple partitions. Messages within an object are not ordered across partitions. | Storage detail | Classic Kafka segment | Diskless topics object | | -------------- | ------------------------------ | ------------------------------------------------------- | | Location | Local disk on broker | Cloud object storage | | Structure | Ordered messages per partition | Batches containing messages from one or more partitions | | Management | Via broker | Via internal Batch Coordinator metadata | | Replication | Kafka-based | Storage-provider-based | Message ordering is preserved at the partition level using metadata. When a broker uploads a batch, it registers the offset range and object reference with the internal **Batch Coordinator**. Consumers use this metadata to fetch messages in the correct order, even when data spans multiple objects. To reduce latency, each broker may cache frequently accessed objects in memory or on ephemeral disk, typically within the same availability zone. Related pages [Batching and delivery in diskless topics](/docs/products/kafka/diskless/concepts/batching-and-delivery.md) --- # Compare diskless and classic Apache Kafka® topics Diskless topics are Apache Kafka®-compatible topics that store data in cloud object storage instead of broker-managed local disks. Classic and diskless topics can coexist within the same Standard Kafka service. In Standard Kafka services, classic topics use managed remote storage by default and you cannot change this setting. ## Compare classic and diskless topics[​](#compare-classic-and-diskless-topics "Direct link to Compare classic and diskless topics") The table below highlights the key differences between classic Kafka topics and diskless topics. | Feature | Classic topic | Diskless topic | | -------------------- | ------------------------------ | ---------------------------------------------- | | Storage | Managed remote storage | Cloud object storage | | Replication | Managed by Kafka brokers | Handled by the storage provider | | Partition leadership | Required | Not required (leaderless data path) | | Data path | Brokers write and serve data | Brokers batch data and write to object storage | | Rebalancing | Required when scaling brokers | Not required | | Segment format | Ordered files per partition | Objects tracked via metadata | | Retention policies | Time- and size-based retention | Limited retention options | | Compacted topics | Supported | Not supported | | Transactions | Supported | Not supported | | Topic creation | Auto-creation or manual | Manual or API-based only | ## When to use classic or diskless topics[​](#when-to-use-classic-or-diskless-topics "Direct link to When to use classic or diskless topics") Use **classic Kafka topics** when: * Your application requires transactions, compaction, or low-latency delivery. * You need compatibility with tooling that depends on classic Kafka features. * You rely on time- or size-based retention policies that are not yet available for diskless topics. Use **diskless topics** when: * High throughput and cost-efficient durable storage are required. * Workload tolerates batching and slightly higher latency. * Simplified broker operations and reduced infrastructure costs are priorities, especially for storage and cross-availability zone (AZ) traffic. ## Use classic and diskless topics in the same cluster[​](#use-classic-and-diskless-topics-in-the-same-cluster "Direct link to Use classic and diskless topics in the same cluster") Apache Kafka clients can produce to and consume from diskless topics using the same APIs as for classic topics. No client-side changes are needed. A single Kafka cluster can contain both diskless and classic topics. This enables: * Gradual adoption of diskless topics by migrating workloads incrementally * Using classic topics for workloads that depend on unsupported features * Consolidating streaming pipelines while optimizing for cost, durability, or throughput where needed --- # Create diskless topics automatically using regular expressions Configure your Aiven for Apache Kafka® service to automatically create new topics as diskless topics when their names match configured regular expressions. This lets clients and connectors use diskless topics without setting `diskless.enable=true` in each create-topic request. When a client creates a topic, the service compares the topic name against the configured regular expressions. If the name matches any expression, the service creates the topic as diskless. ## When to use regular expressions[​](#when-to-use-regular-expressions "Direct link to When to use regular expressions") Use regular expressions to automatically create new topics as diskless topics based on their names, without requiring every client or workflow to set `diskless.enable=true`. This is useful when: * Clients do not support setting `diskless.enable=true` in the topic configuration, or you cannot change the client code to include it. * Kafka Connect frameworks, especially CDC connectors, create and alter topics dynamically and might not let you add diskless settings to the workflow. * Some clients reject unknown topic configurations before the request reaches the broker. * MirrorMaker 2 replicates source topics with their existing configurations. With this configuration, the service creates diskless topics based only on topic names. ## Behavior and limitations[​](#behavior-and-limitations "Direct link to Behavior and limitations") The service applies configured regular expressions only when it creates a new topic. Existing topics do not change. ### Topic matching[​](#topic-matching "Direct link to Topic matching") For each new topic, the service checks the create-topic request and topic name in this order: * If the request sets `diskless.enable=true`, the service creates the topic as a diskless topic. * If the topic name matches any configured regular expression, the service creates the topic as a diskless topic. * Otherwise, the service creates the topic as a classic topic. ### Excluded topics[​](#excluded-topics "Direct link to Excluded topics") The service does not apply regular expressions to: * Internal Kafka topics, such as topics that start with `__`. * Compacted topics. If a topic is created with `cleanup.policy=compact` or `cleanup.policy=compact,delete`, the service creates it as a classic topic even if its name matches a regular expression. ### Classic topic conflicts[​](#classic-topic-conflicts "Direct link to Classic topic conflicts") If a non-compacted topic name matches a configured regular expression, the service creates it as a diskless topic. If you set `diskless.enable=false` for the same topic, the request fails with a conflict error. To create a classic topic, use a topic name that does not match any configured regular expression. In the Aiven Console, the same conflict occurs if the topic name matches a configured regular expression. ### Other limitations[​](#other-limitations "Direct link to Other limitations") * Invalid regular expressions are ignored and do not match any topics. Other valid regular expressions continue to be evaluated. * The regular expressions have no effect when `kafka_diskless.enabled` is `false`. * Topic type is set when the topic is created and cannot be changed later. * Do not set a replication factor. If you must set it when creating diskless topics, use `1` or `-1`. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") To configure default diskless topic regular expressions, you need: * An Aiven for Apache Kafka® service with diskless topics enabled. * A project role or permission that lets you update service configuration. For details, see [Project roles and permissions](/docs/platform/concepts/permissions.md#project-roles-and-permissions). note For existing services, automatic diskless topic creation by regular expression is available after scheduled maintenance updates complete. ## Configure default diskless topic regular expressions[​](#configure-default-diskless-topic-regular-expressions "Direct link to Configure default diskless topic regular expressions") Set `auto_diskless_topic_regexes` in the `kafka_diskless` section of the Kafka service configuration. Add each topic-name regular expression as a separate value. Before you add regular expressions, review these requirements: * You can add up to 32 regular expressions. * Each regular expression can contain up to 128 characters. * Regular expressions use Java [`java.util.regex.Pattern`](https://docs.oracle.com/javase/8/docs/api/java/util/regex/Pattern.html) syntax. * Supported characters are alphanumeric characters and `^ $ - _ . * + ? \ [ ] | { } ( )`. * Unsupported characters, such as `#` or whitespace, are rejected. * Regex features such as look-ahead assertions are not supported. Test each regular expression before you apply it. Use specific regular expressions where possible. Avoid broad regular expressions such as `.*` unless most new non-internal, non-compacted topics must be diskless. note The `auto_diskless_topic_regexes` option is not available in the Aiven Console for services with diskless topics enabled. Set it using the API or CLI. * API * CLI Use the Aiven API to update the service configuration. ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "user_config": { "kafka_diskless": { "enabled": true, "auto_diskless_topic_regexes": ["events\\..*", "cdc\\..*"] } } }' ``` Replace: * `PROJECT_NAME` with the name of your Aiven project. * `SERVICE_NAME` with the name of your Aiven for Apache Kafka service. * `TOKEN` with your Aiven authentication token. * `["events\\..*", "cdc\\..*"]` with the JSON-escaped regular expressions for topic names to create as diskless topics. Use the [Aiven CLI](/docs/tools/cli.md) to set the regular expressions with the `-c` option. Pass the value as a JSON array. Set `kafka_diskless.enabled=true` to keep diskless topics enabled when you update the configuration. ``` avn service update SERVICE_NAME \ --project PROJECT_NAME \ -c kafka_diskless.enabled=true \ -c kafka_diskless.auto_diskless_topic_regexes='["events\\..*", "cdc\\..*"]' ``` Replace: * `SERVICE_NAME` with the name of your Aiven for Apache Kafka service. * `PROJECT_NAME` with the name of your Aiven project. * `["events\\..*", "cdc\\..*"]` with the JSON-escaped regular expressions for topic names to create as diskless topics. If the CLI rejects the array value, use the API. Changing the list of regular expressions might briefly interrupt topic creation requests. Kafka brokers remain online. If you manage services or topics with infrastructure as code, ensure your topic naming conventions and regular expressions stay aligned. If you use the Aiven Terraform Provider, verify that your provider version supports `auto_diskless_topic_regexes` before adding this setting to your configuration. After you save the configuration, Aiven applies the regular expressions to new topic creation requests. Existing topics keep their current type. ## Regular expression examples[​](#regular-expression-examples "Direct link to Regular expression examples") | Goal | Regular expression | | ---------------------------------------------------------- | ------------------ | | Create topics with the `events.` prefix as diskless topics | `events\..*` | | Create topics with the `cdc.` prefix as diskless topics | `cdc\..*` | | Create topics ending in `.diskless` as diskless topics | `.*\.diskless` | | Create topics that contain `.events.` as diskless topics | `.*\.events\..*` | Regular expression matching uses full-match semantics, so leading `^` and trailing `$` anchors are optional. Use `.*` to match variable text before or after a fixed value. For example, `events` matches only the topic name `events`. To match `events.orders`, use `events\..*`. ## Verify the topic type[​](#verify-the-topic-type "Direct link to Verify the topic type") After you create a topic that matches one of the regular expressions, verify that it was created as a diskless topic. * Console * API * CLI 1. In the [Aiven Console](https://console.aiven.io/), click your Aiven for Apache Kafka service. 2. Click **Manage stream** > **Topics**. 3. Check the **Topic type** column. Topics that match a configured regular expression show **Diskless**. ``` BASE_URL="https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME" curl --request GET \ --url "$BASE_URL/topic/TOPIC_NAME" \ --header "Authorization: Bearer TOKEN" ``` In the output, verify that `diskless_enable` is set to `true`. ``` avn service topic-get SERVICE_NAME TOPIC_NAME \ --project PROJECT_NAME ``` In the output, verify that `diskless_enable` is set to `true`. ``` { "topic_name": "events.orders", "diskless_enable": true, "state": "ACTIVE" } ``` ## Remove default diskless topic regular expressions[​](#remove-default-diskless-topic-regular-expressions "Direct link to Remove default diskless topic regular expressions") To stop creating matching topics as diskless topics by default, set `auto_diskless_topic_regexes` to an empty array or `null`. note The `auto_diskless_topic_regexes` option is not available in the Aiven Console for services with diskless topics enabled. Update it using the API or CLI. * API * CLI ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "user_config": { "kafka_diskless": { "enabled": true, "auto_diskless_topic_regexes": [] } } }' ``` To clear the value, set `auto_diskless_topic_regexes` to `null`. Set `kafka_diskless.enabled=true` to keep diskless topics enabled when you update the configuration. ``` avn service update SERVICE_NAME \ --project PROJECT_NAME \ -c kafka_diskless.enabled=true \ -c kafka_diskless.auto_diskless_topic_regexes='[]' ``` To clear the value, set `auto_diskless_topic_regexes` to `null`: ``` avn service update SERVICE_NAME \ --project PROJECT_NAME \ -c kafka_diskless.enabled=true \ -c kafka_diskless.auto_diskless_topic_regexes=null ``` Removing the regular expressions applies only to new topics. Existing topics do not change. The change does not disable diskless topics for the service. It only removes automatic matching. Related pages * [Diskless topics for Apache Kafka®](/docs/products/kafka/diskless/concepts/diskless-topic-overview.md) * [Create Apache Kafka® topics](/docs/products/kafka/howto/create-topic.md) * [Compare diskless and classic topics](/docs/products/kafka/diskless/concepts/topics-vs-classic.md) * [Limitations of diskless topics](/docs/products/kafka/diskless/concepts/limitations.md) --- # Create a free tier Aiven for Apache Kafka® service You can create a free tier Aiven for Apache Kafka® service to learn Kafka, test producers and consumers, or run small proof-of-concept workloads. For details about throughput, topic limits, and feature restrictions, see [Free tier service limitations](/docs/products/kafka/free-tier/kafka-free-tier.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven account (sign up or sign in at ) * An Aiven project note Each organization can have only one free tier Kafka service. ## Create a free tier Kafka service[​](#create-a-free-tier-kafka-service "Direct link to Create a free tier Kafka service") The free tier includes fixed throughput, topic limits, and Karapace Schema Registry and REST Proxy. The cloud provider is managed automatically; you can select only a region group. * Console * CLI 1. In the [Aiven Console](https://console.aiven.io), open the project and select **Services**. 2. Click **Create service**. 3. Select **Aiven for Apache Kafka®**. 4. In **Service tier**, select **Free**. The free tier includes: * Up to 250 KiB/s total throughput * Up to 5 topics, each with two partitions * Karapace Schema Registry and REST Proxy 5. Review the **Cloud** section. The cloud provider is managed automatically for the free tier. You can select only a region group. 6. In **Service basics**, enter a **Service name**. 7. Review the **Service summary**. 8. Click **Create service**. Create a free tier Kafka service using the Aiven CLI. ``` avn service create SERVICE_NAME \ --service-type kafka \ --plan free-0 \ --project PROJECT_NAME \ -c kafka_version=4.0 ``` Parameters: * `SERVICE_NAME`: Name of the Kafka service. * `PROJECT_NAME`: Aiven project name. When the status changes to **Running**, your free tier Kafka service is ready to use. To stream sample data, use the sample data generator in the Aiven Console. It produces test messages to your Kafka service. For details, see [Generate sample data for Aiven for Apache Kafka®](/docs/products/kafka/howto/generate-sample-data.md). To connect your application, use **Quick connect** to get setup steps and sample code for common programming languages. ## After service creation[​](#after-service-creation "Direct link to After service creation") * If no data is produced or consumed for 24 hours, the service is powered off automatically. You receive a notification before shutdown, and you can power on the service from the Aiven Console. * You can upgrade the service at any time to enable: * Cloud and region selection * Higher throughput and storage * Kafka Connect and service integrations * Support for tiered storage ## Upgrade the free tier service[​](#upgrade-the-free-tier-service "Direct link to Upgrade the free tier service") You can upgrade from the free tier at any time to remove throughput, storage, and feature limits. Before upgrading, [add a payment card](/docs/platform/howto/manage-payment-card.md) to the project's billing group. * Console * CLI Upgrade from the service overview: 1. Open your service’s **Overview** page. 2. In **Service usage**, click **Upgrade**. 3. Select a service tier, cloud, region, and plan. 4. Click **Upgrade service**. Or upgrade from service settings: 1. In your service, click **Service settings**. 2. In the **Service summary** section, click **Upgrade**. 3. Select a service tier, cloud, region, and plan. 4. Click **Upgrade service**. Upgrade a free tier Kafka service using the Aiven CLI. 1. List available Kafka plans for the target cloud and region: ``` avn service plans --service-type kafka --cloud CLOUD_REGION ``` 2. Upgrade the service to a paid plan: ``` avn service update SERVICE_NAME \ --project PROJECT_NAME \ --plan NEW_PLAN_NAME \ --cloud CLOUD_REGION ``` When upgrading from the free tier, specify `--cloud` to select the cloud provider and region for the paid plan. Parameters: * `SERVICE_NAME`: Name of the Kafka service. * `PROJECT_NAME`: Aiven project name. * `CLOUD_REGION`: Cloud provider and region for the upgraded service, such as `aws-eu-central-1`. * `NEW_PLAN_NAME`: Paid Kafka plan returned by `avn service plans`. Upgrades apply immediately. The service state becomes **Rebuilding** and remains accessible. When the state switches to **Running**, the new plan is active. Paid services cannot be downgraded to the free tier. Related pages * [Kafka free tier overview](/docs/products/kafka/free-tier/kafka-free-tier.md) * [Create a Kafka topic](/docs/products/kafka/howto/create-topic.md) * [Change the service plan](/docs/products/kafka/howto/change-service-plan.md) --- # Aiven for Apache Kafka® free tier Get started with Apache Kafka® at no cost. The Aiven for Apache Kafka® free tier is a fully managed service for learning, prototyping, and evaluation. No credit card is required. ## When to use the free tier[​](#when-to-use-the-free-tier "Direct link to When to use the free tier") Use the free tier to: * Explore Apache Kafka concepts with a managed service * Build or test small event-driven applications * Evaluate Aiven for Apache Kafka before choosing a paid plan * Test end-to-end message flow during early development * Run small proofs of concept or demonstrations The free tier supports limited-scale workloads. For production use, longer retention, or features such as Aiven for Apache Kafka® Connect, choose a paid plan. ## What the free tier includes[​](#what-the-free-tier-includes "Direct link to What the free tier includes") The free tier includes: * A managed Kafka cluster with a fixed configuration * Karapace Schema Registry * Streaming throughput up to 250 KiB/s ingress and 250 KiB/s egress * Sample data generation for testing message flow * Basic monitoring for metrics and logs Standard Kafka clients can connect the same way as they do to paid Aiven for Apache Kafka plans. ## Limitations[​](#limitations "Direct link to Limitations") Free tier services have the following restrictions. ### Performance and capacity[​](#performance-and-capacity "Direct link to Performance and capacity") * Fixed throughput limits for produce and consume traffic * Up to 5 topics, each with two partitions * Fixed data retention settings * Limited number of users and ACLs ### Features not available[​](#features-not-available "Direct link to Features not available") * Aiven for Apache Kafka® Connect * Aiven for Apache Kafka® MirrorMaker 2 * Tiered storage * Service integrations such as logs, metrics, and authentication * Custom configuration for certain Kafka settings ### Service restrictions[​](#service-restrictions "Direct link to Service restrictions") * One free Kafka service per organization * Paid services cannot be reverted to the free tier * Cloud provider is fixed with a limited set of regions * Free tier services are not covered by an SLA * Maintenance window is fixed * Additional disk storage cannot be added * Creation available only through the Aiven Console ## How free tier services operate[​](#how-free-tier-services-operate "Direct link to How free tier services operate") Free tier Kafka services operate as follows: * **Idle shutdown:** The service powers off automatically if there is no ongoing Kafka activity, such as producing or consuming messages. You receive a notification before shutdown and can power on the service from the Aiven Console. * **First-use shutdown:** A new free tier service with no initial usage can power off within the first few hours after the service is running. You can power the service back on from the Aiven Console. * **Leader election:** The service continues operating after a node failure by using a simplified leader election mode. This ensures availability for small, non-production workloads. * **Alerts:** Platform alerts are not routed to Aiven operators. * **Configuration updates:** Aiven may change the cloud provider, region availability, or configuration of free services. Related pages [Create a free tier Aiven for Apache Kafka® service](/docs/products/kafka/free-tier/create-free-tier-kafka-service.md) --- # Create an Aiven for Apache Kafka® Professional tier service Create an Aiven for Apache Kafka® Professional tier service on Aiven Cloud. Choose the deployment location and configuration for your workload. To create a Free or Developer tier service, see: * [Create a free tier Aiven for Apache Kafka® service](/docs/products/kafka/free-tier/create-free-tier-kafka-service.md) * [Create an Aiven for Apache Kafka® Developer tier service](/docs/products/kafka/dev-tier/create-dev-tier-kafka-service.md) To run Kafka in your own cloud account, see [Create an Apache Kafka® service with BYOC](/docs/products/kafka/get-started/create-kafka-service-byoc.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to the [Aiven Console](https://console.aiven.io) * An Aiven project where you can create Kafka services ## Create a Professional tier service on Aiven Cloud[​](#create-the-service "Direct link to Create a Professional tier service on Aiven Cloud") 1. In the [Aiven Console](https://console.aiven.io), open the project and click **Services**. 2. Click **Create service**. 3. Click **Apache Kafka®**. 4. In **Service tier**, click **Professional**. 5. In **Deployment mode**, click **Aiven cloud**. 6. If **Service type** appears, click **Standard** or **Classic**: * **Standard**: Usage-based pricing, with classic and diskless topics. * **Classic**: A predefined service plan and classic topics. You cannot change the service type after you create the service. note Accounts created after August 2026 create Standard Kafka services. **Service type** appears only for accounts created before August 2026. 7. In **Cloud**, click a cloud provider and region. * Standard * Classic 8. In **Average ingress**, click your expected average ingress rate. Aiven uses this value to estimate compute demand and assumes egress is 3x ingress. If you click **Custom**, enter your expected average ingress rate. 9. If **Cost optimization** appears, set the estimated share of traffic in diskless topics. This section appears when you click **10 MB/s** or **Custom**. The slider previews the estimated network cost for different shares of diskless topic traffic. It does not change the service configuration. 10. In **Retention**, click a default topic retention period. Aiven uses this value for the storage estimate. If you click **Custom**, enter a period between 1 and 30 days. 11. In **Service basics**, enter a **Name**. You cannot change the name after you create the service. 12. Optional: Click **Add tag to this service** to add [resource tags](/docs/products/kafka/howto/tag-service.md). 13. In **Service summary**, review the estimated monthly cost. 14. Click **Create service**. 15. Wait until the service status is **Running**. ### Review the cost estimate[​](#review-the-cost-estimate "Direct link to Review the cost estimate") The estimate includes compute, storage, and network usage. It is based on the selected cloud, region, expected traffic, and retention. The estimated monthly cost is based on your selected configuration. Your invoice reflects actual usage during the billing period. After the service is **Running**, open **Overview** > **Service usage** to view ingress, egress, and storage usage. For more information, see [Pricing for Aiven for Apache Kafka®](/docs/products/kafka/concepts/kafka-pricing.md). 8. In **Plan**, click a plan. 9. Optional: If **Additional disk storage** is available, adjust the disk size. 10. Optional: If **Enable tiered storage** is available, click it to enable [tiered storage](/docs/products/kafka/concepts/kafka-tiered-storage.md). 11. In **Service basics**, enter: * **Name**: Enter a name for the service. You cannot change the name after you create the service. * **Version**: Click a Kafka version. The default version is preselected. * **Tags**: Optional. Add [resource tags](/docs/products/kafka/howto/tag-service.md) to organize your services. 12. In **Service summary**, review the estimated monthly price. 13. Click **Create service**. 14. Wait until the service status is **Running**. For information about Classic Kafka plans and capabilities, see [Classic Kafka overview](/docs/products/kafka/classic-kafka-overview.md). ## Create a service with the Aiven CLI[​](#create-a-service-with-the-aiven-cli "Direct link to Create a service with the Aiven CLI") * Standard * Classic In the Aiven CLI and advanced configuration, Standard Kafka is identified as `inkless`. 1. List the Standard Kafka offerings available for the project: ``` avn inkless offering list \ --organization-id ORGANIZATION_ID \ --project PROJECT_NAME ``` The command returns the available offerings and their maximum ingress and egress throughput. 2. Optional: Filter offerings by required ingress throughput: ``` avn inkless offering list \ --organization-id ORGANIZATION_ID \ --project PROJECT_NAME \ --ingress REQUIRED_MBPS ``` 3. View pricing rates for the available offerings: ``` avn inkless offering rates \ --organization-id ORGANIZATION_ID \ --project PROJECT_NAME \ --cloud-provider CLOUD_PROVIDER ``` Optional: Filter rates by offering with `--offering-name OFFERING_NAME` or by region with `--cloud-name CLOUD_NAME`. 4. Create the service using an offering as the plan: ``` avn service create SERVICE_NAME \ --project PROJECT_NAME \ --service-type kafka \ --cloud CLOUD_REGION \ --plan OFFERING_NAME \ -c kafka_version=4.1 \ -c tiered_storage.enabled=true \ -c inkless.enabled=true ``` 5. Optional: To enable diskless topics when you create the service, also set: ``` -c kafka_diskless.enabled=true ``` You can also enable diskless topics later in the service configuration. Replace the following: * `ORGANIZATION_ID`: organization ID that owns the project * `PROJECT_NAME`: Aiven project name * `REQUIRED_MBPS`: minimum ingress throughput in megabits per second * `CLOUD_PROVIDER`: cloud provider for rate listings: `aws`, `google`, or `azure` * `CLOUD_NAME`: cloud or region identifier returned by the rates listing * `CLOUD_REGION`: cloud region for the service, such as `aws-us-east-1` * `OFFERING_NAME`: Standard Kafka offering returned by `avn inkless offering list` * `SERVICE_NAME`: name of the Kafka service Create a Classic Kafka service: ``` avn service create SERVICE_NAME \ --project PROJECT_NAME \ --service-type kafka \ --cloud CLOUD_REGION \ --plan PLAN_NAME ``` Replace the following: * `SERVICE_NAME`: name of the Kafka service * `PROJECT_NAME`: Aiven project name * `CLOUD_REGION`: cloud provider and region * `PLAN_NAME`: Classic Kafka plan ## Create a Classic Kafka service with Terraform[​](#create-a-classic-kafka-service-with-terraform "Direct link to Create a Classic Kafka service with Terraform") Terraform examples apply to Classic Kafka services. 1. Create a file named `provider.tf` and add the following: ``` Loading... ``` 2. Create a file named `service.tf` and add the following: ``` Loading... ``` 3. Create a file named `variables.tf` and add the following: ``` Loading... ``` 4. Create a `terraform.tfvars` file and add the values for your token and project name. 5. Optional: To output connection details, create a file named `output.tf` and add the following: ``` Loading... ``` To apply your Terraform configuration: 1. Initialize Terraform by running: ``` terraform init ``` The output is similar to the following: ``` Initializing the backend... Initializing provider plugins... - Finding aiven/aiven versions matching ">= 4.0.0, < 5.0.0"... - Installing aiven/aiven v4.9.2... - Installed aiven/aiven v4.9.2 ... Terraform has been successfully initialized! ... ``` 2. To create an execution plan and preview the changes, run: ``` terraform plan ``` 3. To deploy your changes, run: ``` terraform apply --auto-approve ``` ## Next steps[​](#next-steps "Direct link to Next steps") * [Create a Kafka topic](/docs/products/kafka/howto/create-topic.md) * [Pricing for Aiven for Apache Kafka®](/docs/products/kafka/concepts/kafka-pricing.md) * [Generate sample data in the console](/docs/products/kafka/howto/generate-sample-data.md) Related pages * [Standard Kafka overview](/docs/products/kafka/standard-kafka-overview.md) * [Classic Kafka overview](/docs/products/kafka/classic-kafka-overview.md) * [Change the plan for your Standard Kafka service](/docs/products/kafka/howto/change-standard-kafka-plan.md) --- # Create an Apache Kafka® service with bring your own cloud (BYOC) Create an Aiven for Apache Kafka® service in your own cloud account with Bring Your Own Cloud (BYOC). The service runs in your cloud account while Aiven manages the Kafka infrastructure and operations. BYOC uses Classic Kafka with fixed plans. Classic topics are available by default. You can optionally enable diskless topics on supported custom clouds. To create a Kafka service on Aiven Cloud, see [Create an Aiven for Apache Kafka® Professional tier service](/docs/products/kafka/get-started/create-kafka-service.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to the [Aiven Console](https://console.aiven.io) * An Aiven project where you can create Kafka services ## Create a Kafka service with BYOC[​](#create-a-kafka-service-with-byoc "Direct link to Create a Kafka service with BYOC") BYOC Kafka services are available on the Professional service tier. * Console * CLI 1. In the [Aiven Console](https://console.aiven.io), open the project and click **Services**. 2. Click **Create service**. 3. Click **Apache Kafka®**. 4. In **Service tier**, click **Professional**. 5. In **Deployment mode**, click **Bring your own cloud (BYOC)**. If **No custom clouds available** appears, [request BYOC access](#request-byoc-access). If **Set up your custom cloud first** appears, [set up a custom cloud](#set-up-a-custom-cloud). 6. Optional: Click **Enable diskless topics**. If **Diskless topics with BYOC aren't enabled for your account yet** appears, [request diskless topics for a custom cloud](#request-diskless-topics-for-a-custom-cloud). 7. In **Cloud**, click a cloud provider and a custom cloud. 8. In **Plan**, click a plan. 9. Optional: If **Additional disk storage** is available, adjust the disk size. 10. Optional: If **Enable tiered storage** appears, click it to enable [tiered storage](/docs/products/kafka/concepts/kafka-tiered-storage.md). note Enabling diskless topics also enables tiered storage for classic topics on the service. 11. In **Service basics**, enter: * **Name**: Enter a name for the service. You cannot change the name after you create the service. * **Version**: Click a Kafka version. The default version is preselected. * **Tags**: Optional. Add [resource tags](/docs/products/kafka/howto/tag-service.md) to organize your services. 12. Review the **Service summary**, then click **Create service**. 13. Wait until the service status is **Running**. Create a BYOC Kafka service using the Aiven CLI: ``` avn service create SERVICE_NAME \ --project PROJECT_NAME \ --service-type kafka \ --cloud CUSTOM_CLOUD_REGION \ --plan PLAN_NAME ``` Replace the following: * `SERVICE_NAME`: name of the Kafka service * `PROJECT_NAME`: Aiven project name * `CUSTOM_CLOUD_REGION`: custom cloud region, such as `custom-aws-eu-central-1` * `PLAN_NAME`: Kafka plan available in the custom cloud, such as `business-4` Optional: Enable diskless topics when you create the service. Use a Business or Premium `-inkless` plan and a custom cloud that supports diskless topics: ``` avn service create SERVICE_NAME \ --project PROJECT_NAME \ --service-type kafka \ --cloud CUSTOM_CLOUD_REGION \ --plan INKLESS_PLAN_NAME \ -c kafka_version=4.1 \ -c tiered_storage.enabled=true \ -c kafka_diskless.enabled=true ``` Replace `INKLESS_PLAN_NAME` with a plan such as `business-8-inkless`. For more information about diskless topics, see [Diskless topics for Apache Kafka®](/docs/products/kafka/diskless/concepts/diskless-topic-overview.md). ## Request BYOC access[​](#request-byoc-access "Direct link to Request BYOC access") If **No custom clouds available** appears after you click **Bring your own cloud (BYOC)**: * Click **Request access**. Aiven reviews the request and contacts you about enabling BYOC for your organization. After BYOC is enabled, set up a custom cloud and return to create the service. For eligibility and the organization-admin process, see [Enable BYOC](/docs/platform/howto/byoc/enable-byoc.md). ## Set up a custom cloud[​](#set-up-a-custom-cloud "Direct link to Set up a custom cloud") If **Set up your custom cloud first** appears: 1. Click **Go to Admin**. 2. Create a custom cloud and assign it to the project. For more information, see [Create a custom cloud](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md). 3. Return to create the service after the custom cloud is ready. ## Request diskless topics for a custom cloud[​](#request-diskless-topics-for-a-custom-cloud "Direct link to Request diskless topics for a custom cloud") If **Diskless topics with BYOC aren't enabled for your account yet** appears after you click **Enable diskless topics**: 1. Click **Request access**. 2. In **Select a cloud to enable diskless topics**, click one or more custom clouds. 3. Click **Send request**. Aiven reviews the request and contacts you with the next steps. After diskless topics are enabled for those custom clouds, return to create the service and click **Enable diskless topics**. Related pages * [Create an Aiven for Apache Kafka® Professional tier service](/docs/products/kafka/get-started/create-kafka-service.md) * [Bring your own cloud (BYOC)](/docs/platform/concepts/byoc.md) * [Enable BYOC](/docs/platform/howto/byoc/enable-byoc.md) * [Create a custom cloud](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) * [Diskless topics for Apache Kafka®](/docs/products/kafka/diskless/concepts/diskless-topic-overview.md) * [Create a Kafka topic](/docs/products/kafka/howto/create-topic.md) --- # Get started with Aiven for Apache Kafka® Create a managed Apache Kafka® service on Aiven. Choose the tier and deployment model that fit your workload, then send sample data to verify end-to-end streaming. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Create an Aiven account and sign in to the [Aiven Console](https://console.aiven.io). * Create or select an Aiven project with permission to create services. ## Choose your path[​](#choose-your-path "Direct link to Choose your path") Choose a path based on your workload and deployment needs: * Start with [Free tier](#free-tier) for no-cost, low-throughput Kafka workloads. * Use [Developer tier](#developer-tier) for paid development and smaller production workloads. * Use [Professional tier](#professional-tier) for production workloads on Aiven Cloud or Bring Your Own Cloud. * Use [Skills](#set-up-a-kafka-service-using-skills) for command-line setup and configuration. ## Free tier[​](#free-tier "Direct link to Free tier") The Free tier provides a fully managed Kafka service with limited resources. For limits, regions, and supported features, see the [Free tier overview](/docs/products/kafka/free-tier/kafka-free-tier.md). * No payment method is required. * Fixed resource limits support low-throughput workloads. * Basic Kafka workflows are supported, including service creation and sample data generation. **Continue with:** [Create a free tier Aiven for Apache Kafka® service](/docs/products/kafka/free-tier/create-free-tier-kafka-service.md). ## Developer tier[​](#developer-tier "Direct link to Developer tier") The Developer tier is a paid service tier for Classic Kafka. It sits between the [Free tier](/docs/products/kafka/free-tier/kafka-free-tier.md) and Professional tier. Use it for development, prototyping, testing, or production workloads. It provides more capacity and paid-only features than the Free tier. For limits, pricing, Karapace, Connect, integrations, and upgrades, see [Aiven for Apache Kafka® Developer tier](/docs/products/kafka/dev-tier/kafka-dev-tier.md). Manage the service in the console, CLI, API, or with [Skills](/docs/products/kafka/howto/set-up-kafka-with-skills.md). **Continue with:** [Create an Aiven for Apache Kafka® Developer tier service](/docs/products/kafka/dev-tier/create-dev-tier-kafka-service.md). ## Professional tier[​](#professional-tier "Direct link to Professional tier") The Professional tier is for production and workload-heavy environments. Create the service on Aiven Cloud, or in your own cloud account with BYOC. For deployment options and service types, see [Aiven for Apache Kafka® Professional tier](/docs/products/kafka/get-started/professional-tier.md). **Continue with:** * [Create an Aiven for Apache Kafka® Professional tier service](/docs/products/kafka/get-started/create-kafka-service.md) * [Create an Apache Kafka® service with BYOC](/docs/products/kafka/get-started/create-kafka-service-byoc.md) ## Set up a Kafka service using Skills[​](#set-up-a-kafka-service-using-skills "Direct link to Set up a Kafka service using Skills") Use Skills to create and configure a Kafka service from the command line. Install the Aiven Skills bundle: ``` npx skills add Aiven-Open/aiven-skills-bundle ``` To run this command, make sure you have the following: * [Aiven CLI](/docs/tools/cli.md) installed and authenticated. * Node.js installed so `npm` provides `npx`. note Skills operate on **Developer** and **Professional** tier services. Create [Free tier](/docs/products/kafka/free-tier/kafka-free-tier.md) Kafka services in the console. **Continue with:** [Set up Kafka using Skills](/docs/products/kafka/howto/set-up-kafka-with-skills.md). ## Generate sample data using the console[​](#generate-sample-data-using-the-console "Direct link to Generate sample data using the console") Use the built-in sample data generator to send events to a Kafka topic and confirm that the service is working. * Start the generator from the Kafka service overview in the console. * Consume messages from the topic to verify connectivity. * Validate access control, networking, and client configuration. **After your service is running, continue with:** [Generate sample data in the console](/docs/products/kafka/howto/generate-sample-data.md). ## Generate data manually (optional)[​](#generate-data-manually-optional "Direct link to Generate data manually (optional)") Produce data manually to test application integrations or custom workloads. * Create topics and configure Kafka clients. * Produce and consume messages using client libraries, command-line tools, or container-based utilities. * Use this approach for development, automation, or advanced testing scenarios. **After your service is running, continue with:** [Generate sample data manually with Docker](/docs/products/kafka/howto/generate-sample-data-manually.md). ## Create and manage Kafka using AI assistants[​](#create-and-manage-kafka-using-ai-assistants "Direct link to Create and manage Kafka using AI assistants") Use the [Aiven MCP server](/docs/tools/mcp-server.md) to create Kafka services, manage topics, produce and consume messages, and configure connectors from MCP-compatible clients such as Cursor, Claude Code, VS Code, and Gemini CLI. * Configure the Aiven MCP server in your AI assistant. * Describe the service or operation you want in natural language. * The assistant creates and manages Kafka resources through the Aiven API. **Continue with:** [Set up the Aiven MCP server](/docs/tools/mcp-server.md). --- # Aiven for Apache Kafka® Professional tier The Professional tier is for production and workload-heavy Aiven for Apache Kafka® services. Create the service on Aiven Cloud or in your own cloud account with Bring Your Own Cloud (BYOC). ## When to use the Professional tier[​](#when-to-use-the-professional-tier "Direct link to When to use the Professional tier") Use the Professional tier when you need: * Higher throughput, storage, topic limits, or SLA than the [Developer tier](/docs/products/kafka/dev-tier/kafka-dev-tier.md) * Aiven Cloud or BYOC deployment * [Standard Kafka](/docs/products/kafka/standard-kafka-overview.md) or [Classic Kafka](/docs/products/kafka/classic-kafka-overview.md) on Aiven Cloud * Kafka Connect on the same service tier Free and Developer tier services use Classic Kafka. ## Deployment[​](#deployment "Direct link to Deployment") * **Aiven Cloud**: Default path. New customers create Standard Kafka. If **Service type** appears, click **Standard** or **Classic**. You cannot change the service type after you create the service. * **BYOC**: Classic Kafka with fixed plans. You can optionally enable diskless topics on supported custom clouds. Standard Kafka is available on Aiven Cloud only. ## Create a service[​](#create-a-service "Direct link to Create a service") * [Create an Aiven for Apache Kafka® Professional tier service](/docs/products/kafka/get-started/create-kafka-service.md) * [Create an Apache Kafka® service with BYOC](/docs/products/kafka/get-started/create-kafka-service-byoc.md) For compute, storage, and network billing, see [Pricing for Aiven for Apache Kafka®](/docs/products/kafka/concepts/kafka-pricing.md). Related pages * [Get started with Aiven for Apache Kafka®](/docs/products/kafka/get-started/get-started-kafka.md) * [Standard Kafka overview](/docs/products/kafka/standard-kafka-overview.md) * [Classic Kafka overview](/docs/products/kafka/classic-kafka-overview.md) * [Create an Aiven for Apache Kafka® Professional tier service](/docs/products/kafka/get-started/create-kafka-service.md) * [Create an Apache Kafka® service with BYOC](/docs/products/kafka/get-started/create-kafka-service-byoc.md) --- # Manage service users in Aiven for Apache Kafka® Create and manage service users in Aiven for Apache Kafka® to enable secure access and interaction with your service. note Users with `Admin` permission can create topics with any name because the `CreateTopics` permission applies at the cluster level. Other permissions, such as `Alter` and `Delete`, apply only to topics that match the specified pattern. ## Add a user[​](#add-a-user "Direct link to Add a user") * Aiven Console * Aiven CLI * Aiven API 1. In your service, **Users**. 2. Click **Add service user** or **Create user**. 3. Enter a name for your service user. 4. Set up all the other configuration options. If a password is required, a random password is generated automatically. You can change it later. 5. Click **Add service user**. After creating a user, download their access key and certificate from the **Users** page. Run the following command to create a service user: ``` avn service user-create SERVICE_NAME --username USER_NAME ``` Replace the following: * `SERVICE_NAME`: the name of the Aiven service * `USER_NAME`: the username for the new service user Use the [ServiceUserCreate](https://api.aiven.io/doc/#operation/ServiceUserCreate) API endpoint to create a service user: ``` curl -X POST https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/user \ -H "Authorization: Bearer API_TOKEN" \ -H "Content-Type: application/json" \ -d '{"username": "USER_NAME"}' ``` Replace the following: * `PROJECT_NAME`: the name of the Aiven project * `SERVICE_NAME`: the name of the Aiven service * `USER_NAME`: the username for the new service user * `API_TOKEN`: Aiven API [token](https://aiven.io/docs/platform/howto/create_authentication_token) for authentication ## Manage users[​](#manage-users "Direct link to Manage users") * Aiven Console * Aiven CLI * Aiven API 1. Open your Aiven for Apache Kafka service in the [Aiven Console](https://console.aiven.io). 2. Click **Access & Control** > **Users** in the sidebar to view the list of users. 3. To view the password, click **Show password** in the password field for the respective user. 4. Click **Actions** in the user row and choose an action: * Click **Reset credentials** to reset the credentials. * Click **Delete user** to delete the user. * View users: ``` avn service user-list SERVICE_NAME ``` Replace `SERVICE_NAME` with the name of your Aiven service. * Reset user credentials: ``` avn service user-password-reset SERVICE_NAME --username USER_NAME ``` Replace the following: * `SERVICE_NAME`: the name of the Aiven service * `USER_NAME`: the username of the service user * Delete a user: ``` avn service user-delete SERVICE_NAME --username USER_NAME ``` Replace the following: * `SERVICE_NAME`: the name of the Aiven service * `USER_NAME`: the username of the service user - View user details: Use the username-specific endpoint to get details for a service user. ``` curl -X GET https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/user/USER_NAME \ -H "Authorization: Bearer API_TOKEN" ``` Replace the following: * `PROJECT_NAME`: the name of the Aiven project * `SERVICE_NAME`: the name of the Aiven service * `USER_NAME`: the username of the service user * `API_TOKEN`: Aiven API token for authentication - Reset user credentials: ``` curl -X POST https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/user/USER_NAME/reset-credentials \ -H "Authorization: Bearer API_TOKEN" ``` Replace the following: * `PROJECT_NAME`: the name of the Aiven project * `SERVICE_NAME`: the name of the Aiven service * `USER_NAME`: the username of the service user * `API_TOKEN`: Aiven API token for authentication - Delete a user: ``` curl -X DELETE https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/user/USER_NAME \ -H "Authorization: Bearer API_TOKEN" ``` Replace the following: * `PROJECT_NAME`: the name of the Aiven project * `SERVICE_NAME`: the name of the Aiven service * `USER_NAME`: the username of the service user * `API_TOKEN`: Aiven API token for authentication Related pages * [Access Control Lists in Aiven for Apache Kafka®](/docs/products/kafka/concepts/acl.md) * [Manage access control lists in Aiven for Apache Kafka®](/docs/products/kafka/howto/manage-acls.md) --- # Add client-side Apache Kafka® producer and consumer Datadog metrics When you enable the [Datadog integration](/docs/integrations/datadog/datadog-metrics.md) in Aiven for Apache Kafka®, the service supports all of the broker-side metrics listed in the [Datadog Kafka integration documentation](https://docs.datadoghq.com/integrations/kafka/?tab=host#data-collected) and allows you to send additional [custom metrics](/docs/products/kafka/howto/datadog-customised-metrics.md). Additionally, you can collect client-side metrics directly from the producer or consumer and send them to Datadog. For guidance, refer to the *Missing producer and consumer metrics* section in the [Datadog documentation](https://docs.datadoghq.com/integrations/faq/troubleshooting-and-deep-dive-for-kafka), which outlines the process for integrating missing metrics natively for Java-based producers and consumers. For clients using languages other than Java, incorporating these metrics can be achieved through [DogStatsD](https://docs.datadoghq.com/developers/dogstatsd/). --- # Manage approvals [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) The **Approvals** page allows you to manage requests for Aiven for Apache Kafka® resources owned by your group. These requests can include claims for existing topics or requests to create new topics. To manage approvals using Terraform and GitHub, see [Manage approvals using Terraform and GitHub](/docs/products/kafka/howto/terraform-governance-approvals.md). ## Requests flow[​](#requests-flow "Direct link to Requests flow") Managing requests involves several key steps: * **Request initiation**: Users submit a request from the Apache Kafka topic catalog to claim ownership of an existing topic or create a topic. * **Request visibility**: All members of the topic's owning group can see the request on the **Approvals** page. * **Review and approval**: A group member reviews the request and either approves or declines it. * **Completion**: If approved, the requested change of topic ownership or topic creation is completed. If declined, the request is not implemented. ## Key elements[​](#key-elements "Direct link to Key elements") The approvals page displays the following key elements: * **Topic**: Name of the topic * **Type**: Type of request * **Status**: Current status of the request * **Approved**: Request approved and processed * **Declined**: Request declined. The topic is now available for another group to claim * **Deleted**: Request deleted by the requester * **Failed**: Request cannot complete the required action (for example, due to Apache Kafka cluster issues or connectivity problems) * **Pending**: Request awaiting approval or decline * **Provisioning**: Request in progress. The topic is being created on the Apache Kafka cluster * **Requesting group**: Group that made the request * **Requester**: User who submitted the request * **Requested on**: Date and time the request was made * **Filters**: Filter the list by type, status, requester, and group * **Search**: Search for specific topics by name ## Approve or decline requests[​](#approve-or-decline-requests "Direct link to Approve or decline requests") 1. Access the [Aiven console](https://console.aiven.io/), and click **Tools** > **Governance** > ****Approvals****. 2. Click a topic to open the **Review request** pane. 3. Review the request details. 4. Do one of the following: * To **approve** a request: 1. Click **Approve**. 2. Optional: Enter a message for the requester in the **Approve create request** pop-up window. 3. Click **Approve** to confirm. * To **decline** a request: 1. Click **Decline**. 2. Enter a reason for declining the request in the **Decline create request** pop-up window. 3. Click **Decline** to confirm. Related pages * [Aiven for Apache Kafka® topic catalog](/docs/products/kafka/concepts/topic-catalog-overview.md) * [Manage approvals using Terraform and GitHub](/docs/products/kafka/howto/terraform-governance-approvals.md) * [Claim topic ownership](/docs/products/kafka/howto/claim-topic.md) * [Group requests](/docs/products/kafka/howto/group-requests.md) --- # Avoid OutOfMemoryError errors in Aiven for Apache Kafka® When a node in an Aiven for Apache Kafka® or Aiven for Apache Kafka® Connect cluster runs low on memory, the Java virtual machine (JVM) running the service may not be able to allocate the memory, and will raise a `java.lang.OutOfMemoryError` exception. As a result, Apache Kafka may stop processing messages in topics, and Apache Kafka Connect may be unable to manage connectors. For example, you can have a [Kafka Connect S3 sink connector](/docs/products/kafka/kafka-connect/howto/s3-sink-connector-aiven.md) with a `TimeBasedPartitioner` that is configured to generate data directories with a `path.format=YYYY/MM` format. This means that each output S3 object contains the data for a full month of the year. Internally, the connector uses a buffer in memory for each S3 object before it is flushed to AWS. If the total amount of data for any single month exceeds the amount of available RAM at the time of data ingestion, the JVM executing the Kafka Connect worker, throws a `java.lang.OutOfMemoryError` exception and cannot manage the workload. ## How to avoid `OutOfMemoryError` issues[​](#how-to-avoid-outofmemoryerror-issues "Direct link to how-to-avoid-outofmemoryerror-issues") There are three options for handling heavy workloads and potentially avoiding `java.lang.OutOfMemoryError` exceptions: * Decrease the maximum amount of data simultaneously kept in memory * Increase the available memory for Kafka services or Kafka Connect workers * Create a dedicated Kafka Connect cluster The strategy to apply depends on a number of factors, and may combine more than one option. ### Keep memory peak usage low[​](#keep-memory-peak-usage-low "Direct link to Keep memory peak usage low") The first option is to minimise the amount of memory needed to process the data. note This approach usually requires some configuration changes in the data pipeline. In the S3 example above, you can change the settings for the S3 sink connector to limit each S3 object to the data for one day, rather than one month, by using the data directory format `path.format=YYYY/MM/dd`. You can also fine tune the connector settings such as `rotate.schedule.interval.ms`, `rotate.interval.ms`, and `partition.duration.ms` and set them to smaller values enabling the connector to commit S3 objects at a higher frequency. All the above options will decrease peak memory usage by forcing the connector to write fewer data items to S3 at shorter intervals, allowing the JVM to release smaller memory buffers more often. ### Upgrade to a larger plan[​](#upgrade-to-a-larger-plan "Direct link to Upgrade to a larger plan") The above approach requires configuration changes in the data pipeline. If this is not possible, the alternative is to upgrade your Aiven for Apache Kafka service to a bigger plan with more memory. You can do this in the [Aiven Console](https://console.aiven.io/) or with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn-cli-service-update). ### Create a dedicated Kafka Connect cluster[​](#create-a-dedicated-kafka-connect-cluster "Direct link to Create a dedicated Kafka Connect cluster") Another possible scenario preventing you from handling large workloads is if you are running Kafka Connect on the same nodes as your Aiven for Apache Kafka cluster. This forces the connectors to share available memory with the Apache Kafka brokers and reduces the total amount of memory available for each connector. If this is the case, you can separate the services and create a dedicated Aiven for Apache Kafka Connect cluster using an appropriate service plan. This separation provides the added advantage of allowing you to scale the Kafka brokers and Kafka Connect services independently. note When upgrading or expanding the cluster, Aiven automatically launches new nodes and transfers the data from the old nodes to the new ones. To know more about horizontal and vertical scaling options, see the [Scaling options in Apache Kafka®](/docs/products/kafka/concepts/horizontal-vertical-scaling.md). --- # Optimize Apache Kafka® performance Follow these best practices to optimize the performance and reliability of your Aiven for Apache Kafka® service. ## Check your topic replication factors[​](#check-your-topic-replication-factors "Direct link to Check your topic replication factors") Apache Kafka uses replication between brokers to protect data in case of node failures. The replication factor (RF) determines how many copies of each partition are maintained across the cluster. Evaluate the importance of each topic and set a replication factor that balances durability requirements with cost and performance. An RF of 3 is recommended for production because it improves durability and availability. In multi-AZ deployments, replication traffic across availability zones can increase network costs, especially for high-throughput workloads. For Diskless Topics architecture and considerations, see [Diskless Topics overview](/docs/products/kafka/diskless/concepts/diskless-topic-overview.md). Set the replication factor when creating or editing a [topic](/docs/products/kafka/howto/create-topic.md) in the [Aiven Console](https://console.aiven.io/). note Replication factors below 2 are not allowed to prevent data loss from unexpected node terminations. ## Choose a reasonable number of partitions for a topic[​](#choose-a-reasonable-number-of-partitions-for-a-topic "Direct link to Choose a reasonable number of partitions for a topic") Too few partitions can create processing bottlenecks. A single partition processes messages sequentially, which limits throughput. Too many partitions increase overhead and reduce cluster efficiency. Because partition counts cannot be reduced, start with a number that supports parallel processing and increase it as needed. A maximum of 4,000 partitions per broker and 200,000 per cluster is recommended. For details, see this [Apache Kafka blog post](https://blogsarchive.apache.org/kafka/entry/apache-kafka-supports-more-partitions). Keep the total number of topics under 7,000. note Ordering is guaranteed only within a partition. To maintain ordering of related records, place them in the same partition. ## Check entity-based partitions for imbalances[​](#check-entity-based-partitions-for-imbalances "Direct link to Check entity-based partitions for imbalances") Partitioning messages by an entity identifier, such as a user ID, can create imbalanced partitions. This results in uneven load distribution and reduces parallel processing efficiency. You can view the size of each partition by selecting the [topic](/docs/products/kafka/howto/create-topic.md) in the **Manage stream** > **Topics** list and opening the **Partitions** tab in the [Aiven Console](https://console.aiven.io/). ## Balance between throughput and latency[​](#balance-between-throughput-and-latency "Direct link to Balance between throughput and latency") Adjust producer and consumer batch sizes to balance throughput and latency. Larger batches increase throughput but add latency. Smaller batches reduce latency but increase overhead, which can lower throughput. Settings such as `batch.size` and `linger.ms` can be configured in the producer. For more details, refer to the [Apache Kafka documentation](https://kafka.apache.org/documentation/). ## Configure acknowledgments for received data[​](#configure-acknowledgments-for-received-data "Direct link to Configure acknowledgments for received data") The `acks` parameter in the producer configuration controls how write operations are acknowledged. Choose a setting that matches your reliability requirements: * **`acks=0`**: The producer does not wait for confirmation. This minimizes latency but increases the risk of data loss if the broker fails during transmission. Use this setting only when some data loss is acceptable. * **`acks=1`** (default and recommended): The producer waits for the leader broker to confirm receipt. This reduces the risk of data loss but does not protect against leader failure before replication completes. * **`acks=all`**: The producer waits for acknowledgment from the leader and all in-sync replicas. This prevents data loss but increases latency. ## Configure single availability zone (AZ) for BYOC[​](#configure-single-availability-zone-az-for-byoc "Direct link to Configure single availability zone (AZ) for BYOC") Deploying Aiven for Apache Kafka in a single availability zone (AZ) reduces inter-zone data transfer costs. Single-AZ deployment places all brokers and replicas in one failure domain, so the cluster cannot tolerate an AZ outage. If the zone becomes unavailable, the service cannot recover until the zone is restored. note Before enabling this configuration, contact your account team to discuss your use case and agree on the reduced SLA. The standard uptime SLA does not apply to services deployed in a single AZ. ### Replication factor considerations in a single AZ[​](#replication-factor-considerations-in-a-single-az "Direct link to Replication factor considerations in a single AZ") **Replication factor 1 (RF=1):** Creates a single copy of each partition. In a single-AZ deployment, a broker or AZ failure results in data loss. Use RF=1 only when losing data is acceptable. **Replication factor 3 (RF=3):** Protects against individual broker failures. It does not protect against an AZ failure when all replicas are in the same zone. If the AZ becomes unavailable, all replicas can be lost, and the cluster cannot recover until the zone is restored. ### When to use single AZ[​](#when-to-use-single-az "Direct link to When to use single AZ") Avoid single-AZ deployment for production workloads or any data that cannot be recreated. Use single AZ only for: * Development, QA, or test workloads * Temporary proof-of-concept environments * Workloads where data can be recreated ### Risks and considerations[​](#risks-and-considerations "Direct link to Risks and considerations") * All brokers and replicas exist within one failure domain, increasing the impact of an AZ outage. * Recovery options are limited because the cluster cannot fail over to another zone. * Service downtime may increase during an AZ failure because no cross-zone redundancy exists. * SLA terms for single-AZ deployments must be agreed with your account team. ### Enable single-AZ allocation[​](#enable-single-az-allocation "Direct link to Enable single-AZ allocation") Single-AZ allocation must be configured during service creation. It cannot be enabled for existing Kafka services. To enable this option for your project, contact [Aiven support](mailto:support@aiven.io) or your account team. To create a single-AZ Kafka service using the [Aiven CLI](/docs/tools/cli.md), set `single_zone.enabled=true`: ``` avn service create SERVICE_NAME \ --service-type kafka \ --plan PLAN_NAME \ --cloud CLOUD_REGION \ -c single_zone.enabled=true \ --disk-space-gib DISK_SIZE \ --project PROJECT_NAME ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Kafka service. * `--service-type SERVICE_TYPE`: Service type. * `--plan PLAN_NAME`: Aiven service plan. * `--cloud CLOUD_REGION`: Cloud region and provider. * `-c single_zone.enabled=true`: Enables single AZ allocation. * `--disk-space-gib DISK_SIZE`: Specifies total disk space for the service (in GiB). * `--project PROJECT_NAME`: Name of your Aiven project. --- # Change data retention period To avoid running out of disk space, by default, Apache Kafka® drops the oldest messages from the beginning of each log after their retention period expires. **Aiven for Apache Kafka®** allows you to configure the retention period for each topic. The retention period can be configured at both the service and topic levels. If no retention period is specified for a particular topic, the service-level setting will be applied, with a default value of 168 hours. When modifying the service retention period, it will override the retention period of any previously created topics. ## For a single topic[​](#for-a-single-topic "Direct link to For a single topic") To change the retention period for a single topic: 1. In the [Aiven Console](https://console.aiven.io/), select your project and choose your Aiven for Apache Kafka® service. 2. Select **Manage stream** > **Topics** from the left sidebar. 3. Select the topic to modify. 4. In the **Topic info** screen, select **Modify**. 5. In the modify topic screen, update the value of **Retention ms** to the desired retention length in milliseconds. If you cannot find **Retention ms**, use the search bar to locate it note The **Retention ms** option is displayed in the modify topic screen for topics where advanced configuration was enabled during topic creation. 6. Select **Update** to save your changes. 7. In the *Advanced configuration* view find **Retention ms**. 8. Change the value of **Retention ms** value to the desired retention length in milliseconds. tip You can also change **Retention bytes** setting to limit amount of data retained based on the storage usage. ## At a service level[​](#at-a-service-level "Direct link to At a service level") 1. In the [Aiven Console](https://console.aiven.io/), select your project and choose your Aiven for Apache Kafka® service. 2. Click **Service settings**, scroll to **Advanced configuration**, and click **Configure**. 3. In the **Advanced configuration** dialog, click **Add Advanced Configuration**. 4. Configure the retention period for Apache Kafka® logs: * Find `kafka.log_retention_hours` or `kafka.log_retention_ms`, and set the retention period. * To limit retention by storage usage, set `kafka.log_retention_bytes`. 5. Click **Save configuration**. ## Unlimited retention[​](#unlimited-retention "Direct link to Unlimited retention") Aiven does not limit the maximum retention period. To turn off time-based content expiration, set the retention value to `-1`. warning Using high retention periods without monitoring the available storage space can cause your service to run out of disk space. These situations are not covered by the Aiven SLA. --- # Change the plan for your Aiven for Apache Kafka® service Change the service plan for your Aiven for Apache Kafka® service to scale resources up or down and optimize costs. Plan changes apply to Kafka services that use a service plan, including [Classic Kafka](/docs/products/kafka/classic-kafka-overview.md), Developer, and Free tier services. To change the plan for a Standard Kafka service, see [Change the plan for your Standard Kafka service](/docs/products/kafka/howto/change-standard-kafka-plan.md). Adjust the plan of your services at any time to scale your services as needed and optimize costs. If you can't find a suitable plan, you can [request a custom plan](/docs/platform/concepts/service-pricing.md). tip If you plan to upgrade your service plan, do it immediately after a full backup. This reduces the amount of incremental changes that need to be applied on top of the base backup, which speeds up the upgrade itself. important * When changing a service plan, reserve an additional 25% of disk space. This requirement applies to upgrades and downgrades. * Downgrading to a plan with fewer VMs is supported for most services, including Aiven for Apache Kafka®, Aiven for PostgreSQL®, Aiven for OpenSearch®, Aiven for ClickHouse®, Aiven for MySQL®, Aiven for Metrics, and Aiven for Valkey™. * Changing a service plan triggers a node recycle, service rebuilding, and any pending maintenance updates. - Console - Terraform - CLI 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Change plan**. 3. Select a plan that provides at least 125% of the current disk size and click **Change plan**. Update the `plan` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). To change a service plan in the Aiven CLI, use the [`avn service update --plan `](/docs/tools/cli/service-cli.md#avn-cli-service-update) command. Your service's state becomes **Rebuilding** and remains accessible. When the state switches to **Running**, your new service plan is active. Related pages * [Change the plan for your Standard Kafka service](/docs/products/kafka/howto/change-standard-kafka-plan.md) * [Scaling options in Apache Kafka®](/docs/products/kafka/concepts/horizontal-vertical-scaling.md) * [Scale disk storage](/docs/products/kafka/howto/scale-disk-storage.md) * [Kafka upgrade procedure](/docs/products/kafka/concepts/upgrade-procedure.md) --- # Change the plan for your Standard Kafka service Change the service plan for your Standard Kafka service to scale resources up or down and optimize costs. 1. In the [Aiven Console](https://console.aiven.io), open your Standard Kafka service. 2. On the **Overview** page, in **Service usage**, click **Change plan**. 3. In **Average ingress**, click a rate. Aiven uses the rate to estimate compute demand. Aiven assumes egress is three times ingress. Optional: Click **Custom** and enter a rate. 4. If **Cost optimization** is shown, set the estimated share of traffic in diskless topics. **Cost optimization** is shown when you click **10 MB/s** or **Custom**. The slider previews estimated network cost for diskless topic traffic. The slider does not change the service configuration. 5. In **Retention**, click a retention period. Aiven uses the period for the storage estimate. Optional: Click **Custom** and enter a period of 1 to 30 days. 6. In **Service summary**, review the estimated monthly cost. The estimate is based on the selected configuration. The invoice reflects actual usage in the billing period. 7. Click **Upgrade service**. 8. Wait until the service status is **Running**. Related pages * [Change the plan for your Aiven for Apache Kafka® service](/docs/products/kafka/howto/change-service-plan.md) * [Pricing for Aiven for Apache Kafka®](/docs/products/kafka/concepts/kafka-pricing.md) * [Standard Kafka overview](/docs/products/kafka/standard-kafka-overview.md) --- # Claim topic ownership [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) To take ownership of a topic that your group does not currently own, you can submit a claim request. Once you send the request, the current owner can approve or decline it. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") [Governance](/docs/products/kafka/howto/enable-governance.md) enabled for your organization. ## Claim a single topic[​](#claim-a-single-topic "Direct link to Claim a single topic") 1. Access the [Aiven console](https://console.aiven.io/), and click **Tools** > **Apache Kafka topic catalog**. 2. Click the topic name to open the **Topic details** pane. 3. Click **Claim topic**. 4. In the **User group** field, select the group to assign as the owner. 5. (Optionally) Enter a message for the approver explaining why you are claiming ownership. 6. Click **Claim**. Alternatively: 1. Click **Open topic page** in the **Topic details** pane. 2. From the **Overview** tab, click **Claim topic**. 3. In the **User group** field, select the group to assign as the owner. 4. (Optionally) Enter a message for the approver explaining why you are claiming ownership. 5. Click **Claim**. note * If another group has already requested ownership of the topic, you cannot claim ownership. * You receive notifications when your request is approved or declined. ## Claim multiple topics[​](#claim-multiple-topics "Direct link to Claim multiple topics") 1. Select the checkbox next to the **Topic** heading to select all claimable topics, or individually select the checkboxes for the topics to claim. 2. Click **Claim ownership**. 3. In the **User group** field, select the group to be assigned as the owner. note Users can belong to multiple groups. Ensure you select the appropriate group. 4. (Optionally) Enter a message for the approver explaining why you are claiming ownership. 5. Click **Claim**. ## View and track topic claim requests[​](#view-and-track-topic-claim-requests "Direct link to View and track topic claim requests") 1. Click **Tools** > **Apache Kafka topic catalog**. 2. Click the topic. 3. In the **Topic details** pane, under the **Topic owner** section, click **Open Group requests**. Alternatively: 1. Click **Tools** > **Governance** > ****Group requests****. 2. Search for your topic to view its status. ## Delete a topic claim request[​](#delete-a-topic-claim-request "Direct link to Delete a topic claim request") To delete your topic ownership claim request: 1. Click **Tools** > **Governance** > ****Group requests****. 2. Search for and click the topic to delete the ownership claim request for. 3. In the **Topic details** pane, click **Delete**. The review status updates to **Deleted** in the pane and on the **Group requests** page. Related pages * [Governance in Aiven for Apache Kafka](/docs/products/kafka/concepts/governance-overview.md) * [Apache Kafka topic catalog](/docs/products/kafka/concepts/topic-catalog-overview.md) * [Approvals](/docs/products/kafka/howto/approvals.md) * [Group requests](/docs/products/kafka/howto/group-requests.md) --- # Configure audit logging for Aiven for Apache Kafka® Turn audit logging on for your Aiven for Apache Kafka® service, change what it records, and manage audit log volume. For what audit logging captures and its limitations, see [Audit logging for Aiven for Apache Kafka®](/docs/products/kafka/concepts/audit-logging.md). important Before you turn on audit logging, note the following: * Turning audit logging on or changing its settings restarts the Kafka brokers in your service one at a time. Make these changes during a maintenance window or a period of low traffic. * After you turn on audit logging, you cannot remove audit logging settings or turn off audit logging yourself. To turn off audit logging, [contact Aiven support](/docs/platform/howto/support.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") To configure audit logging, you need one of the following [project roles or permissions](/docs/platform/concepts/permissions.md#project-roles-and-permissions): * `admin`: Full access to services in the project. * `operator`: Full service management access. * `project:services:write`: Broad services write access. * `service:configuration:write`: Least-privilege access for changing service configuration. The `developer` and `read_only` roles cannot configure audit logging. ## Enable audit logging[​](#enable-audit-logging "Direct link to Enable audit logging") To enable audit logging, add at least one `kafka.audit_log` setting to your service configuration. Any setting you add must have a valid value. * Aiven Console * Aiven CLI * Aiven API * Terraform 1. In the [Aiven Console](https://console.aiven.io/), open your Aiven for Apache Kafka service. 2. Click **Service settings**. 3. In the **Advanced configuration** section, click **Configure**. 4. Click **Add configuration options** and enter `audit` to find the audit logging settings. 5. Add `kafka.audit_log.record_type` and select `user_operations`. 6. Optional: Add other audit logging settings and set their values. 7. Click **Save configuration**. Set `kafka.audit_log.record_type` with the [`avn service update`](/docs/tools/cli/service-cli.md) command and the `-c` flag: ``` avn service update SERVICE_NAME \ --project PROJECT_NAME \ -c kafka.audit_log.record_type=user_operations ``` Replace `SERVICE_NAME` and `PROJECT_NAME` with your service and project names. Send a `PUT` request to the [service update](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint: ``` curl -s -X PUT \ --url "https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME" \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "user_config": { "kafka": { "audit_log": { "record_type": "user_operations" } } } }' ``` Replace `PROJECT_NAME`, `SERVICE_NAME`, and `TOKEN` with your project name, service name, and authentication token. Add an `audit_log` block to the `kafka_user_config` block of your `aiven_kafka` resource: ``` resource "aiven_kafka" "example_kafka" { project = var.project_name cloud_name = "google-europe-west1" plan = "business-4" service_name = "example-kafka" kafka_user_config { kafka { audit_log { record_type = "user_operations" } } } } ``` ## Audit logging settings[​](#audit-logging-settings "Direct link to Audit logging settings") Use these advanced configuration settings to customize audit logging. In the service configuration, add these settings under `kafka.audit_log`, for example `kafka.audit_log.record_type`. | Setting | Type | Default | Description | | ------------------------ | ------- | ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `record_type` | string | `user_operations` | The type of activity to record. Use `user_operations` for detailed operation entries, or `user_activity` to record only that a Kafka user was active. | | `aggregation_period_sec` | integer | `300` | How long, in seconds, to group entries before writing them to the service logs. A higher value produces fewer, larger entries. Accepts a value from 1 to 1800. | | `include_denials` | boolean | `false` | Whether to include denied authorization attempts in audit log entries. When false, audit log entries include only allowed operations. | | `group_by` | string | `user_and_ip` | How to group entries: by Kafka user only (`user`), or by Kafka user and IP address (`user_and_ip`). Applies only when `record_type` is `user_operations`. | ## Change audit logging settings[​](#change-audit-logging-settings "Direct link to Change audit logging settings") To change what audit logging records, set new values for the `kafka.audit_log` settings with any of the preceding methods. Services that already use audit logging keep their current settings until you change them. ## View audit logs[​](#view-audit-logs "Direct link to View audit logs") Audit entries appear in the service logs with the `AUDIT:` prefix. To view them, use one of the following methods: * In the [Aiven Console](https://console.aiven.io/), open your service and click **Observe** > **Logs**. * With the Aiven CLI, run: ``` avn service logs SERVICE_NAME \ --project PROJECT_NAME \ | grep AUDIT: ``` * Send the service logs to another system through a [log integration](/docs/products/kafka/howto/integrate-service-logs-into-kafka-topic.md). ## Manage audit log volume[​](#manage-audit-log-volume "Direct link to Manage audit log volume") Audit logging can produce many log entries. To manage the volume: * Set `group_by` to `user` instead of `user_and_ip` to combine a Kafka user's activity across IP addresses. * Increase `aggregation_period_sec` to group entries over a longer time window. * Keep `include_denials` set to `false` unless you need denied attempts in audit log entries. * Use `user_activity` instead of `user_operations` when you only need to know which Kafka users were active. ## Related pages[​](#related-pages "Direct link to Related pages") * [Audit logging for Aiven for Apache Kafka®](/docs/products/kafka/concepts/audit-logging.md) * [Advanced parameters for Aiven for Apache Kafka®](/docs/products/kafka/reference/advanced-params.md) * [Integrate service logs into an Apache Kafka® topic](/docs/products/kafka/howto/integrate-service-logs-into-kafka-topic.md) * [Monitor and alert logs for denied ACL](/docs/products/kafka/howto/monitor-logs-acl-failure.md) --- # Configure a custom domain for Kafka REST API, Schema Registry, and Kafka Connect Configure a custom domain to replace the default Aiven service hostname for Kafka REST API, Schema Registry, and Kafka Connect. ## About custom domains[​](#about-custom-domains "Direct link to About custom domains") Custom domains let you use your own domain name, such as `kafka.example.com`, instead of the default Aiven service hostname for REST-based Kafka components. Supported services include Kafka REST API, Schema Registry, and Kafka Connect. Custom domains are not supported for Kafka broker endpoints that use the native Kafka protocol. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running Aiven for Apache Kafka® service * Permission to manage DNS records for your domain * Access to the Aiven CLI or Aiven API ## Configure a custom domain[​](#configure-a-custom-domain "Direct link to Configure a custom domain") ### Step 1: Create a DNS CNAME record[​](#step-1-create-a-dns-cname-record "Direct link to Step 1: Create a DNS CNAME record") Create a `CNAME` record for your custom domain in your DNS provider. Point the record to the Aiven hostname that matches custom domain certificate provisioning. Use the following target: * `SUBDOMAIN.example.com` → `public-PROJECT_NAME-SERVICE_NAME.aivencloud.com`: Use this target for public access. Use this target for VPC or other private-network access when you configure `custom_domain`. Aiven creates the certificate for the `public-*` hostname and for your custom domain. If your service uses a VPC or another private network path, clients can still use the custom domain. The `CNAME` target remains `public-PROJECT_NAME-SERVICE_NAME.aivencloud.com`. ### Step 2: Configure the custom domain[​](#step-2-configure-the-custom-domain "Direct link to Step 2: Configure the custom domain") Set the custom domain in the Kafka service configuration. * CLI * API Configure the custom domain using Aiven CLI: ``` avn service update SERVICE_NAME \ --user-config '{"custom_domain": "SUBDOMAIN.example.com"}' ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `custom_domain`: Custom domain for Kafka REST API, Schema Registry, and Kafka Connect. Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) API to configure the custom domain for the service: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer API_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "custom_domain": "SUBDOMAIN.example.com" } }' ``` Parameters: * `PROJECT_NAME`: Name of your project. * `SERVICE_NAME`: Name of your service. * `API_TOKEN`: API token for authentication. * `custom_domain`: Custom domain for Kafka REST API, Schema Registry, and Kafka Connect. ### Step 3: Wait for certificate provisioning[​](#step-3-wait-for-certificate-provisioning "Direct link to Step 3: Wait for certificate provisioning") After you configure the custom domain, Aiven automatically requests a TLS certificate from Let's Encrypt. Certificate provisioning typically completes within a few minutes. No additional action is required. ### Step 4: Update client configuration[​](#step-4-update-client-configuration "Direct link to Step 4: Update client configuration") Update client applications to use the custom domain and the existing service port. ``` https://SUBDOMAIN.example.com:PORT ``` Use the same port as the default Aiven endpoint. --- # Configure the log cleaner for topic compaction The log cleaner serves the purpose of preserving only the latest value associated with a specific message key in a partition for [compacted topics](/docs/products/kafka/concepts/log-compaction.md). In Aiven for Apache Kafka®, the log cleaner is enabled by default, while log compaction remains disabled. ## Enable log compaction for all topics[​](#enable-log-compaction-for-all-topics "Direct link to Enable log compaction for all topics") 1. In the [Aiven Console](https://console.aiven.io/), select your project and choose your Aiven for Apache Kafka® service. 2. In the service page, select **Service settings** from the sidebar. 3. On the **Service settings** page, scroll down to the **Advanced configuration** section, and click **Configure**. 4. In the **Advanced configuration** dialog, click **Add configuration options**. 5. Find `log.cleanup.policy` in the list and select it. 6. Set the value to `compact`. 7. Click **Save configuration**. warning This change will affect all topics in the cluster that do not have a configuration override in place. ## Enable log compaction for a specific topic[​](#enable-log-compaction-for-a-specific-topic "Direct link to Enable log compaction for a specific topic") 1. In the [Aiven Console](https://console.aiven.io/), select your project and choose your Aiven for Apache Kafka® service. 2. Select **Manage stream** > **Topics** from the left sidebar. 3. Select the topic to modify and select **Modify** in the context menu. 4. From the drop-down options for the **Cleanup policy**, select the value `compact`. 5. Select **Update**. ## Configure log cleaning frequency and delay[​](#configure-log-cleaning-frequency-and-delay "Direct link to Configure log cleaning frequency and delay") Before the cleaning begins, the cleaner thread will inspect the logs to find those with highest **dirty ratio** calculated as the number of bytes in the head vs the total number of bytes in the log (tail + head). Read more about head and tail definition in the [compacted topic documentation](/docs/products/kafka/concepts/log-compaction.md). The ratio provides an estimation of how many duplicated keys are present in a topic, and therefore needs to be compacted. tip For the log cleaner to start compacting a topic, the dirty ratio needs to be bigger than a threshold set to **50% by default**. You can change this value: * Globally for the cluster: In the **Advanced configuration** section of the service overview, modify the value of the `kafka.log_cleaner_min_cleanable_ratio` property. * For a specific topic: Modify the value of `min_cleanable_ratio` property. The log cleaner can be configured to leave some amount of uncompacted data in the head of the log by setting **compaction time lag**. To do so, 1. Open In the **Advanced configuration** of your service or an individual topic: 2. Set the following properties: * `log.cleaner.min.compaction.lag.ms`: Setting to a value greater than 0 will prevent the log cleaner from compacting messages with an age newer than a minimum message age. This delays compacting records. * `log.cleaner.max.compaction.lag.ms`: The maximum amount of time a message will remain uncompacted. tip The compaction lag can be bigger than the `log.cleaner.max.compaction.lag.ms` setting since it directly depends on the time to complete the actual compaction process and can be delayed by the log cleaner threads availability. ## Tombstone records[​](#tombstone-records "Direct link to Tombstone records") During the cleanup process, the log cleaner also remove records that have a null value, also known as **tombstone** records. To delay tombstone records from being deleted, set the `delete.retention.ms` property for the compacted topic. Consumers can read all tombstone messages as long as they reach the head of the topic before the period defined in `delete.retention.ms` is passed. Related pages * [Compacted topics](/docs/products/kafka/concepts/log-compaction.md) * [Kafka advanced parameters](/docs/products/kafka/reference/advanced-params.md) --- # Configure preferred availability zones Configure preferred availability zones for Aiven for Apache Kafka®, Aiven for Apache Kafka® Connect, and Aiven for Apache Kafka® MirrorMaker 2. Use preferred zones to control where service nodes are placed within a cloud region. By default, Aiven distributes Kafka service nodes across the available zones in a cloud region. With preferred zones, you can limit node placement to specific availability zones (AZs) while keeping the service highly available. Use preferred zones to: * Reduce cross-AZ data transfer costs by placing Kafka nodes closer to your applications. * Reduce latency between service nodes and client applications. * Meet requirements that restrict workloads to specific zones within a region. * Align Kafka broker placement with consumer locations when using follower fetching. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven for Apache Kafka®, Aiven for Apache Kafka® Connect, or Aiven for Apache Kafka® MirrorMaker 2 service on AWS, Google Cloud, or Azure. * The zone IDs for your cloud region. * [Aiven CLI](/docs/tools/cli.md), to configure preferred zones from the command line. ## Zone ID formats[​](#zone-id-formats "Direct link to Zone ID formats") Zone ID formats vary by cloud provider. | Cloud provider | Format | Examples | | -------------- | ------------------------ | -------------------------------------- | | AWS | Zone ID | `use1-az1`, `use1-az2`, `euc1-az1` | | Google Cloud | Zone name | `europe-west1-a`, `us-central1-b` | | Azure | Location and zone number | `germanywestcentral/1`, `westeurope/2` | AWS zone IDs AWS availability zone names, such as `us-east-1a`, can map to different physical locations in different accounts. Use zone IDs, such as `use1-az1`, because they are consistent across all accounts. For more information, see [AWS documentation on AZ IDs](https://docs.aws.amazon.com/ram/latest/userguide/working-with-az-ids). ## Configure preferred zones[​](#configure-preferred-zones "Direct link to Configure preferred zones") * Console * CLI * API * Terraform 1. In the [Aiven Console](https://console.aiven.io), open your Aiven for Apache Kafka®, Aiven for Apache Kafka® Connect, or Aiven for Apache Kafka® MirrorMaker 2 service. 2. Click **Service settings**. 3. In the **Advanced configuration** section, click **Configure**. 4. In **`preferred_zones`**, enter the zone IDs, separated by commas. Example: ``` use1-az1,use1-az2,use1-az3 ``` 5. Click **Save changes**. To configure preferred zones for an existing Kafka, Kafka Connect, or MirrorMaker 2 service, run: ``` avn service update SERVICE_NAME \ -c preferred_zones='["use1-az1", "use1-az2", "use1-az3"]' ``` To configure preferred zones when you create a service, use the service type for your service: ``` avn service create SERVICE_NAME \ --service-type SERVICE_TYPE \ --plan business-4 \ --cloud aws-us-east-1 \ -c preferred_zones='["use1-az1", "use1-az2", "use1-az3"]' ``` Use one of the following service types: * `kafka`: Aiven for Apache Kafka® * `kafka_connect`: Aiven for Apache Kafka® Connect * `kafka_mirrormaker`: Aiven for Apache Kafka® MirrorMaker 2 Replace the following: * `SERVICE_NAME`: Name of your Aiven service. * `SERVICE_TYPE`: Type of service to create. * `preferred_zones`: JSON array of zone IDs where nodes can be placed. To configure preferred zones for a Kafka, Kafka Connect, or MirrorMaker 2 service, use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) API operation: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer API_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "preferred_zones": ["use1-az1", "use1-az2", "use1-az3"] } }' ``` Replace the following: * `PROJECT_NAME`: Name of your Aiven project. * `SERVICE_NAME`: Name of your Aiven for Apache Kafka®, Kafka Connect, or MirrorMaker 2 service. * `API_TOKEN`: Your Aiven API token. Use the `preferred_zones` attribute in your `aiven_kafka` resource: ``` resource "aiven_kafka" "example" { project = "my-project" service_name = "my-kafka" cloud_name = "aws-us-east-1" plan = "business-4" kafka_user_config { preferred_zones = ["use1-az1", "use1-az2", "use1-az3"] } } ``` Use the `preferred_zones` attribute in your `aiven_kafka_connect` resource: ``` resource "aiven_kafka_connect" "example" { project = "my-project" service_name = "my-kafka-connect" cloud_name = "aws-us-east-1" plan = "business-4" kafka_connect_user_config { preferred_zones = ["use1-az1", "use1-az2", "use1-az3"] } } ``` Use the `preferred_zones` attribute in your `aiven_kafka_mirrormaker` resource: ``` resource "aiven_kafka_mirrormaker" "example" { project = "my-project" service_name = "my-mirrormaker" cloud_name = "aws-us-east-1" plan = "business-4" kafka_mirrormaker_user_config { preferred_zones = ["use1-az1", "use1-az2", "use1-az3"] } } ``` ## How preferred zones work[​](#how-preferred-zones-work "Direct link to How preferred zones work") When you configure preferred zones: * New nodes use the specified zones when capacity is available. * The service validates zone IDs when you save the configuration. * If a preferred zone is unavailable, Aiven can use another zone in the same region to keep the service available. See [Automatic node rebalancing](#automatic-node-rebalancing). * Existing nodes do not move immediately. Preferred zones take effect when Aiven recreates nodes, such as during maintenance or a plan change. ### Minimum zone count[​](#minimum-zone-count "Direct link to Minimum zone count") For high availability, configure at least three preferred zones. Configuring fewer than three zones reduces fault tolerance and requires special account permissions. ### Interaction with single-zone configuration[​](#interaction-with-single-zone-configuration "Direct link to Interaction with single-zone configuration") If you configure both `preferred_zones` and [`single_zone.availability_zone`](/docs/products/kafka/reference/advanced-params-standard.md#single_zone.availability_zone) settings and set `single_zone.enabled` to `true`, the `single_zone` setting takes precedence. ## Automatic node rebalancing[​](#automatic-node-rebalancing "Direct link to Automatic node rebalancing") When Aiven creates or replaces a node, it uses one of your preferred zones if capacity is available. If none of the preferred zones have capacity, Aiven places the node in another availability zone in the same region to keep your service available. For Kafka plans that support automatic rebalancing, Aiven regularly checks for nodes running outside their preferred zones. When capacity is available, Aiven automatically moves those nodes back to a preferred zone. For Kafka plans that do not support automatic rebalancing, Aiven can move a node back to a preferred zone when Aiven recreates the node, such as during maintenance or a plan change. Automatic rebalancing is supported on: * `inkless-professional` plans * `kafka-professional` plans * Business and Premium `-inkless` plans on BYOC * All Kafka Connect and MirrorMaker 2 plans ## Example: Optimize follower fetching[​](#example-optimize-follower-fetching "Direct link to Example: Optimize follower fetching") To reduce cross-AZ network costs with [follower fetching](/docs/products/kafka/howto/enable-follower-fetching.md), align preferred zones with the availability zones where your Kafka consumers run. 1. Identify the availability zones where your Kafka consumers run. 2. Configure preferred zones to match those zones. 3. Enable [follower fetching](/docs/products/kafka/howto/enable-follower-fetching.md) on your Kafka service. 4. Configure `client.rack` on your consumers to match their availability zone. With this configuration, consumers can fetch data from replicas in the same zone when local replicas are available. ## Example: Reduce cross-AZ costs for diskless topics[​](#example-reduce-cross-az-costs-for-diskless-topics "Direct link to Example: Reduce cross-AZ costs for diskless topics") For Kafka services that use diskless topics, you can reduce cross-AZ data transfer costs by routing requests to brokers in the same availability zone as the client application. This configuration uses the `client.id` pattern to communicate the client's availability zone to the broker. ### How it works[​](#how-it-works "Direct link to How it works") Diskless topics store data in object storage, which is accessible from all availability zones. Any broker can serve any partition. By configuring client rack awareness, clients can send requests to brokers in their local availability zone. ### Configuration[​](#configuration "Direct link to Configuration") 1. Identify the availability zones where your producers run. 2. Configure preferred zones to match those availability zones. 3. Configure each producer's `client.id` to include the `diskless_az` pattern: ``` # Producer in use1-az1 client.id=my-producer,diskless_az=use1-az1 # Producer in use1-az2 client.id=my-producer,diskless_az=use1-az2 ``` Set `diskless_az` to match the broker's `broker.rack` configuration, which uses the zone IDs in your `preferred_zones` setting. 4. Apply the same pattern to consumers to route requests consistently: ``` # Consumer in use1-az1 client.id=my-consumer,diskless_az=use1-az1 ``` ### Benefits[​](#benefits "Direct link to Benefits") * Eliminates most cross-AZ data transfer costs for diskless topics. * Improves cache hit rates by routing clients to consistent brokers per partition. * Works with any Kafka client version because it uses the standard `client.id` property. note This `client.id` pattern is specific to Kafka with diskless topics. For classic Kafka topics, use [follower fetching](#example-optimize-follower-fetching) with the standard `client.rack` configuration instead. ## Remove preferred zones[​](#remove-preferred-zones "Direct link to Remove preferred zones") To return to automatic zone distribution across all available zones, remove the `preferred_zones` configuration. * CLI * API ``` avn service update KAFKA_SERVICE_NAME --remove-option preferred_zones ``` ``` avn service update KAFKA_CONNECT_SERVICE_NAME --remove-option preferred_zones ``` ``` avn service update MIRRORMAKER_SERVICE_NAME --remove-option preferred_zones ``` ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer API_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "preferred_zones": null } }' ``` Related pages * [Availability zones](/docs/platform/concepts/availability-zones.md) * [Aiven for Apache Kafka® Connect](/docs/products/kafka/kafka-connect.md) * [Aiven for Apache Kafka® MirrorMaker 2](/docs/products/kafka/kafka-mirrormaker.md) * [Enable follower fetching](/docs/products/kafka/howto/enable-follower-fetching.md) * [Follower fetching concepts](/docs/products/kafka/concepts/follower-fetching.md) --- # Enable and configure tiered storage for topics Aiven for Apache Kafka® allows you to configure tiered storage and set retention policies for individual topics. ## Prerequisite[​](#prerequisite "Direct link to Prerequisite") [Tiered storage enabled for the Aiven for Apache Kafka service](/docs/products/kafka/howto/enable-kafka-tiered-storage.md). ## Tiered storage status for topics[​](#tiered-storage-status-for-topics "Direct link to Tiered storage status for topics") When you activate tiered storage for a service, any new topics created in that service using the Aiven Console have tiered storage enabled by default. * **Disable tiered storage for new topics**: You can disable tiered storage by setting **Remote storage enable** to false during the creation of a new topic in the topic advanced configurations. Once you activate tiered storage for a topic, you cannot disable it. Contact [Aiven support](mailto:support@aiven.io) for assistance. * **Tiered storage for existing topics**: Disabled for existing topics. ## Enable and configure tiered storage for topics[​](#enable-and-configure-tiered-storage-for-topics "Direct link to Enable and configure tiered storage for topics") * Console * CLI 1. Access the [Aiven Console](https://console.aiven.io/), select your project, and select your Aiven for Apache Kafka service. 2. Click **Manage stream** > **Topics**. 3. Either add a new topic with tiered storage or modify an existing one: * To add a new topic, click **Add topic**. * To modify an existing topic, select the topic to modify, then either: * Click the topic name to open the **Topic info** screen, and click **Modify**. * Alternatively, in the topic row, click **Actions** > **Edit topic **. 4. The **Enable advanced configuration?** option is set to **Yes** by default. 5. Ensure the **Remote storage enable** option is set to **True** to activate tiered storage for the topic. To disable it, set it to **False**. 6. Optionally, adjust the default values for `local_retention_ms` and `local_retention_bytes`. important By default, `local_retention_bytes` and `local_retention_ms` are set to `-2` or inherit the configuration from the service level. When set to `-2`, the retention in local storage matches the total retention. In this scenario, the data segments sent to remote storage are also retained locally. Remote storage contains older data segments only if the total retention exceeds the local retention. 7. Click **Create topic** or **Update** to save your changes and activate tiered storage. Using the [Aiven CLI](/docs/tools/cli.md), you can enable tiered storage for specific Apache Kafka® topics and set retention policies. To create a topic with tiered storage enabled: ``` avn service topic-create \ --project demo-kafka-project \ --partitions 2 \ --replication 2 \ --remote-storage-enable \ demo-kafka-service exampleTopic ``` In this example: * `demo-kafka-project` is the name of your project. * `demo-kafka-service` is the name of your Aiven for Apache Kafka® service. * `exampleTopic` is the name of the topic you are creating with tiered storage enabled. * The topic has the partition set to `2` and the replication factor set to `2` ### Configure retention policies for a topic[​](#configure-retention-policies-for-a-topic "Direct link to Configure retention policies for a topic") After enabling tiered storage, you can configure the retention policies for local storage: ``` avn service topic-update \ --project demo-kafka-project \ --partitions 2 \ --local-retention-ms 100 \ --local-retention-bytes 10 \ demo-kafka-service exampleTopic ``` This command sets the local retention time to 100 milliseconds and the local retention size to 10 bytes for the `exampleTopic` topic in the `demo-kafka-service` of the `demo-kafka-project`. important By default, `local_retention_bytes` and `local_retention_ms` are set to `-2` or inherit the configuration from the service level. When set to `-2`, the retention in local storage matches the total retention. In this scenario, the data segments sent to remote storage are also retained locally. Remote storage contains older data segments only if the total retention exceeds the local retention. ## Optional: Configure the client-side parameter[​](#optional-configure-the-client-side-parameter "Direct link to Optional: Configure the client-side parameter") For optimal performance and reduced risk of broker interruptions when using tiered storage, it is recommended to update the client-side parameter `fetch.max.wait.ms` from its default value of 500 ms to 5000 ms. This consumer configuration is no longer necessary starting from Apache Kafka version 3.6.2. Consider upgrading to Apache Kafka version 3.6.2 or later before enabling tiered storage. --- # Manage configurations with Apache Kafka® CLI tools Aiven for Apache Kafka® services are fully manageable and customizable via the [Aiven CLI](/docs/tools/cli.md). To guarantee the service stability, direct Apache ZooKeeper™ access isn't available, but our tooling provides you all the options that you need - whether your Apache Kafka version has Apache ZooKeeper™ in it or not. Some of the configuration changes can be made using the standard client tools shipped with the Apache Kafka® binaries. The example below shows how to create a topic using one of these tools, `kafka-topics.sh`. ## Example: Create a topic with retention time to 30 minutes with `kafka-topics.sh`[​](#example-create-a-topic-with-retention-time-to-30-minutes-with-kafka-topicssh "Direct link to example-create-a-topic-with-retention-time-to-30-minutes-with-kafka-topicssh") Each topic in Apache Kafka can have a different retention time, defining for the messages' time to live. The Aiven Console and API offer the ability to set the retention time as part of the topic [creation](/docs/tools/cli/service/topic.md#avn_cli_service_topic_create) or [update](/docs/tools/cli/service/topic.md#avn-cli-topic-update). The same can be achieved using the `kafka-topics.sh` script included in the [Apache Kafka binaries](https://kafka.apache.org/downloads): 1. Download the [Apache Kafka binaries](https://kafka.apache.org/downloads) and unpack the archive. 2. Go to the `bin` folder containing the Apache Kafka client tools. 3. Create a [Java keystore and truststore](/docs/products/kafka/howto/keystore-truststore.md) to authenticate with the Aiven for Apache Kafka service. 4. Create a [client configuration file](/docs/products/kafka/howto/kafka-tools-config-file.md) with the necessary authentication details. 5. Run the following command to check the connectivity to the Aiven for Apache Kafka service, replacing the `` with the URI of the service available in the [Aiven Console](https://console.aiven.io/). ``` ./kafka-topics.sh \ --bootstrap-server \ --command-config consumer.properties \ --list ``` If successful, the above command lists all the available topics 6. Run the following command to create a topic named `new-test-topic` with a retention rate of 30 minutes. Use the kafka-topics script for this and set the retention value in milliseconds `((100 * 60) * 30 = 180000)`. ``` ./kafka-topics.sh \ --bootstrap-server \ --command-config consumer.properties \ --topic new-test-topic \ --create \ --config retention.ms=180000 ``` 7. Run the same command as step 5 to check the topic creation. Optionally, you can also run the `kafka-topics.sh` command with the `--describe` flag to check the details of your topic and the retention rate. note It is currently not possible to change the configurations for an existing topic via `kafka-topics.sh` as that requires a connection to ZooKeeper. --- # Connect to Aiven for Apache Kafka® with command-line tools Use **Quick connect** to set up Apache Kafka® command-line tools for Aiven for Apache Kafka®. The guided flow helps you select or create a topic, choose an authentication method, grant permissions, and copy generated producer and consumer commands. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md). * Python 3.7 or later and the `pip` package manager to install the [Aiven CLI](https://github.com/aiven/aiven-client). * The Apache Kafka command-line tools. The `kafka-console-producer.sh` and `kafka-console-consumer.sh` scripts are included in the [Apache Kafka® binary downloads](https://kafka.apache.org/downloads) in the `bin` directory. ## Open Quick connect[​](#open-quick-connect "Direct link to Open Quick connect") 1. In the [Aiven Console](https://console.aiven.io/), open your Aiven for Apache Kafka service. 2. On the service **Overview** page, in the **Set up your stream** section, click **Quick connect**. 3. At the top of the page, click **CLI**. ## Step 1: Set up a topic[​](#step-1-set-up-a-topic "Direct link to Step 1: Set up a topic") Topics organize and store the events that you stream to Apache Kafka. 1. In **Topic name**, do one of the following: * Select an existing topic. * Click **Create new topic**, enter a name, and create the topic. The new topic is selected automatically for the next steps. note If your service has **Diskless topics** enabled, you can create a **Classic** or **Diskless** topic, or select an existing topic of either type. You can't change the topic type after creation. For more information, see [Create Apache Kafka® topics](/docs/products/kafka/howto/create-topic.md). ## Step 2: Set up an authentication method[​](#step-2-set-up-an-authentication-method "Direct link to Step 2: Set up an authentication method") 1. Select an authentication method: * **SASL**: Recommended. Use SASL/SCRAM-SHA-256 for username and password authentication. SASL is enabled by default for new Aiven for Apache Kafka services. For existing services, check whether SASL is enabled. If **Enable SASL** appears, click **Enable SASL**, then continue. * **Client certificate**: Use certificate-based authentication with mTLS instead of a password. 2. Select a service user: * Select an existing service user from the list. * To create one, click **Create new service user**, enter a username, and click **Add service user**. 3. Check the permission status shown for the selected user: * If the user has all the required permissions, the granted permissions are shown (for example, `read, write`). To change them, click **Manage access in ACLs**. * If the user has some or no permissions, click **Grant permissions**. Select **Produce**, **Consume**, or both, and click **Save**. Permissions the user already has are selected and cannot be removed in Quick connect. To remove them, click **Manage access in ACLs**. note The `avnadmin` user has permissions by default, and you cannot change them from Quick connect. For other service users, you can add permissions from Quick connect, but you cannot remove them from Quick connect. To remove permissions, [remove ACL entries for the service user](/docs/products/kafka/howto/manage-acls.md#delete-acl-entries). ## Step 3: Copy the command snippets[​](#step-3-copy-the-command-snippets "Direct link to Step 3: Copy the command snippets") 1. Select the **Prerequisites** tab, install the Aiven CLI, and sign in to your Aiven account using the generated commands. 2. Download and extract the Apache Kafka command-line tools, then open the `bin` directory. To run Kafka commands from any directory, add the absolute path of the `bin` directory to your `PATH`. 3. Run the generated Aiven CLI command from your Kafka `bin` directory to download the service credentials. The command downloads the certificates to your Kafka `bin` directory. For SASL, the keystore file isn't required. For **Client certificate**, both the keystore and truststore files are needed. When you reference the truststore in `configuration.properties`, use its absolute path. 4. Create a `configuration.properties` file in your home directory and copy the generated configuration into it. If you move the certificate files or run Kafka commands from another directory, use absolute paths in `configuration.properties`. For a description of the configuration properties, see [Apache Kafka toolbox properties](/docs/products/kafka/howto/kafka-tools-config-file.md). warning For SASL, the configuration includes the service user password in plaintext. Store the file securely and do not commit it to source control. 5. Select the **Producer** tab, copy the generated `kafka-console-producer.sh` command, and run it from your Kafka `bin` directory. Type messages in the terminal and press Enter to send them to the topic. 6. Select the **Consumer** tab, copy the generated `kafka-console-consumer.sh` command, and run it from your Kafka `bin` directory to read messages from the topic. After you start the producer and consumer, use the producer terminal to send messages and the consumer terminal to read them. --- # Connect to Aiven for Apache Kafka® with C++ Use **Quick connect** to set up a C++ client for Aiven for Apache Kafka®. The guided flow helps you choose or create a topic, choose an authentication method, grant permissions, and copy generated producer and consumer code. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md). * A C++ development environment. * The [`librdkafka`](https://github.com/confluentinc/librdkafka) and `librdkafka-devel` packages. note For SASL/SCRAM authentication with Apache Kafka® 4.x, use `librdkafka` version 2.6.1 or later. For more information, see [Enable and configure SASL authentication for Apache Kafka®](/docs/products/kafka/howto/kafka-sasl-auth.md#configure-sasl-mechanisms). ## Open Quick connect[​](#open-quick-connect "Direct link to Open Quick connect") 1. In the [Aiven Console](https://console.aiven.io/), open your Aiven for Apache Kafka service. 2. On the service **Overview** page, in the **Set up your stream** section, click **Quick connect**. 3. At the top of the page, click **C++**. ## Step 1: Set up a topic[​](#step-1-set-up-a-topic "Direct link to Step 1: Set up a topic") Topics organize and store the events that you stream to Apache Kafka. 1. In **Topic name**, do one of the following: * Choose an existing topic. * Click **Create new topic**, enter a name, and create the topic. The new topic is selected automatically for the next steps. note If your service has **Diskless topics** enabled, you can create a **Classic** or **Diskless** topic, or choose an existing topic of either type. You can't change the topic type after creation. For more information, see [Create Apache Kafka® topics](/docs/products/kafka/howto/create-topic.md). ## Step 2: Set up an authentication method[​](#step-2-set-up-an-authentication-method "Direct link to Step 2: Set up an authentication method") 1. Choose an authentication method: * **SASL**: Recommended. Use SASL/SCRAM-SHA-256 for username and password authentication. SASL is enabled by default for new Aiven for Apache Kafka services. For existing services, check whether SASL is enabled. If **Enable SASL** appears, click **Enable SASL**, then continue. * **Client certificate**: Use certificate-based authentication with mTLS instead of a password. 2. Choose a service user: * Choose an existing service user from the list. * To create one, click **Create new service user**, enter a username, and click **Add service user**. 3. Check the permission status shown for the selected user: * If the user has all the required permissions, the granted permissions are shown (for example, `read, write`). To change them, click **Manage access in ACLs**. * If the user has some or no permissions, click **Grant permissions**. Choose **Produce**, **Consume**, or both, and click **Save**. note The `avnadmin` user has permissions by default. For other service users, you can add permissions from Quick connect. To change or remove permissions, click **Manage access in ACLs**. For more information, see [Manage Apache Kafka® ACLs](/docs/products/kafka/howto/manage-acls.md#delete-acl-entries). ## Step 3: Copy the code snippets[​](#step-3-copy-the-code-snippets "Direct link to Step 3: Copy the code snippets") tip **Download template** is an optional shortcut for testing the connection without a local C++ project. The ZIP file includes ready-to-run producer and consumer code, certificates, and dependencies. To run it, follow the included `README.md`. If you already have a C++ project, copy the snippet from the **Producer** or **Consumer** tab. 1. Under **Prerequisites**, install `librdkafka` and `librdkafka-devel`. For installation instructions, open **librdkafka on GitHub**. 2. Under **Downloads**, download the certificate files for your authentication method: * For **SASL**, click **Download CA certificate (ca.pem)**. * For **Client certificate**, download **CA certificate (ca.pem)**, **service certificate (service.cert)**, and **service key (service.key)**. The generated snippet loads the certificates directly, so you don't need to create a truststore. 3. Click the **Producer** or **Consumer** tab to view the generated producer or consumer code. 4. Copy the generated code into `producer.cpp` or `consumer.cpp`. 5. Compile and run the producer or consumer using the generated command: ``` gcc -o producer.o -Wall -lrdkafka producer.cpp && ./producer.o ``` ``` gcc -o consumer.o -Wall -lrdkafka consumer.cpp && ./consumer.o ``` warning For SASL, copied snippets include the service user password in plaintext. Store the code securely, and do not commit it to source control. After you add the code to your project, update the certificate file paths to match where you saved the files, then run your producer or consumer to start streaming events. --- # Connect to Aiven for Apache Kafka® with C# Use **Quick connect** to set up a C# client for Aiven for Apache Kafka®. The guided flow helps you choose or create a topic, choose an authentication method, grant permissions, and copy generated producer and consumer code. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md). * .NET SDK 9 or later. * The [`Confluent.Kafka`](https://www.nuget.org/packages/Confluent.Kafka) package. ## Open Quick connect[​](#open-quick-connect "Direct link to Open Quick connect") 1. In the [Aiven Console](https://console.aiven.io/), open your Aiven for Apache Kafka service. 2. On the service **Overview** page, in the **Set up your stream** section, click **Quick connect**. 3. At the top of the page, click **C#**. ## Step 1: Set up a topic[​](#step-1-set-up-a-topic "Direct link to Step 1: Set up a topic") Topics organize and store the events that you stream to Apache Kafka. 1. In **Topic name**, do one of the following: * Choose an existing topic. * Click **Create new topic**, enter a name, and create the topic. The new topic is selected automatically for the next steps. note If your service has **Diskless topics** enabled, you can create a **Classic** or **Diskless** topic, or choose an existing topic of either type. You can't change the topic type after creation. For more information, see [Create Apache Kafka® topics](/docs/products/kafka/howto/create-topic.md). ## Step 2: Set up an authentication method[​](#step-2-set-up-an-authentication-method "Direct link to Step 2: Set up an authentication method") 1. Choose an authentication method: * **SASL**: Recommended. Use SASL/SCRAM-SHA-256 for username and password authentication. SASL is enabled by default for new Aiven for Apache Kafka services. For existing services, verify whether SASL is enabled. If **Enable SASL** appears, click **Enable SASL**, then continue. * **Client certificate**: Use certificate-based authentication with mTLS instead of a password. Downloads include a CA certificate, service certificate, and access key. 2. Choose a service user: * Choose an existing service user from the list. * To create one, click **Create new service user**, enter a username, and click **Add service user**. 3. Check the permission status shown for the selected user: * If the user has all the required permissions, the granted permissions are shown (for example, `read, write`). To change them, click **Manage access in ACLs**. * If the user has some or no permissions, click **Grant permissions**. Choose **Produce**, **Consume**, or both, and click **Save**. note The `avnadmin` user has permissions by default. For other service users, you can add permissions from Quick connect. To change or remove permissions, click **Manage access in ACLs**. For more information, see [Manage Apache Kafka® ACLs](/docs/products/kafka/howto/manage-acls.md#delete-acl-entries). ## Step 3: Copy the code snippets[​](#step-3-copy-the-code-snippets "Direct link to Step 3: Copy the code snippets") tip **Download template** is an optional shortcut for testing the connection. The ZIP file includes ready-to-run producer and consumer code, certificates, and a multi-stage `Containerfile` using .NET SDK 9 and .NET runtime 9. To run it, follow the included `README.md`. If you already have a C# project, copy the snippet from the **Producer** or **Consumer** tab. 1. Under **Prerequisites**, install the `Confluent.Kafka` package in your .NET project: ``` dotnet add package Confluent.Kafka ``` 2. Under **Downloads**, download the certificate files for your authentication method: * For **SASL**, click **Download CA certificate**. * For **Client certificate**, click **Download CA certificate**, **Download service certificate**, and **Download service access key**. The generated snippet loads the certificates directly. You don't need to create a truststore. 3. Click the **Producer** or **Consumer** tab to view the generated producer or consumer code. 4. Copy the code into `Producer.cs` or `Consumer.cs`. warning For SASL, copied snippets include the service user password in plaintext. Store the code securely, and do not commit it to source control. After you add the code to your project, update the certificate file paths to match where you saved the files, then run your producer or consumer to start streaming events. --- # Connect to Aiven for Apache Kafka® with Go Use **Quick connect** to set up a Go client for Aiven for Apache Kafka®. The guided flow helps you choose or create a topic, choose an authentication method, grant permissions, and copy generated producer and consumer code. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md). * A Go development environment. * The [`kafka-go`](https://github.com/segmentio/kafka-go) library. ## Open Quick connect[​](#open-quick-connect "Direct link to Open Quick connect") 1. In the [Aiven Console](https://console.aiven.io/), open your Aiven for Apache Kafka service. 2. On the service **Overview** page, in the **Set up your stream** section, click **Quick connect**. 3. At the top of the page, click **Go**. ## Step 1: Set up a topic[​](#step-1-set-up-a-topic "Direct link to Step 1: Set up a topic") Topics organize and store the events that you stream to Apache Kafka. 1. In **Topic name**, do one of the following: * Choose an existing topic. * Click **Create new topic**, enter a name, and create the topic. The new topic is selected automatically for the next steps. note If your service has **Diskless topics** enabled, you can create a **Classic** or **Diskless** topic, or choose an existing topic of either type. You can't change the topic type after creation. For more information, see [Create Apache Kafka® topics](/docs/products/kafka/howto/create-topic.md). ## Step 2: Set up an authentication method[​](#step-2-set-up-an-authentication-method "Direct link to Step 2: Set up an authentication method") 1. Choose an authentication method: * **SASL**: Recommended. Use SASL/SCRAM-SHA-512 for username and password authentication. Go code snippets use `SCRAM-SHA-512` for SASL authentication. SASL is enabled by default for new Aiven for Apache Kafka services. For existing services, check whether SASL is enabled. If **Enable SASL** appears, click **Enable SASL**, then continue. * **Client certificate**: Use certificate-based authentication with mTLS instead of a password. 2. Select a service user: * Select an existing service user from the list. * To create one, click **Create new service user**, enter a username, and click **Add service user**. 3. Check the permission status shown for the selected user: * If the user has all the required permissions, the granted permissions are shown (for example, `read, write`). To change them, click **Manage access in ACLs**. * If the user has some or no permissions, click **Grant permissions**. Select **Produce**, **Consume**, or both, and click **Save**. note The `avnadmin` user has permissions by default. For other service users, you can add permissions from Quick connect. To change or remove permissions, click **Manage access in ACLs**. For more information, see [Manage Apache Kafka® ACLs](/docs/products/kafka/howto/manage-acls.md#delete-acl-entries). ## Step 3: Copy the code snippets[​](#step-3-copy-the-code-snippets "Direct link to Step 3: Copy the code snippets") tip **Download template** is an optional shortcut for testing the connection without a local Go environment. The ZIP file includes ready-to-run producer and consumer code, certificates, and dependencies. To run it, follow the included `README.md`. If you already have a Go project, copy the snippet from the **Producer** or **Consumer** tab. 1. Under **Prerequisites**, initialize a Go module and install the required libraries: ``` go mod init ``` ``` go get github.com/segmentio/kafka-go ``` For **SASL**, also install the SCRAM package: ``` go get github.com/segmentio/kafka-go/sasl/scram@v0.4.51 ``` 2. Under **Downloads**, download the certificate files for your authentication method: * For **SASL**, click **Download CA certificate**. * For **Client certificate**, click **Download CA certificate**, **Download service certificate**, and **Download service access key**. The generated snippet loads the certificates directly, so you don't need to create a truststore. 3. Click the **Producer** or **Consumer** tab to view the generated producer or consumer code. 4. Copy the code. warning For SASL, copied snippets include the service user password in plaintext. Store the code securely, and do not commit it to source control. After you add the code to your project, update the certificate file paths to match where you saved the files, then run your producer or consumer to start streaming events. --- # Connect to Aiven for Apache Kafka® with Java Use **Quick connect** to set up a Java client for Aiven for Apache Kafka®. The guided flow helps you choose or create a topic, choose an authentication method, grant permissions, and copy generated producer and consumer code. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md). * A Java development environment with [Maven](https://maven.apache.org/install.html). ## Open Quick connect[​](#open-quick-connect "Direct link to Open Quick connect") 1. In the [Aiven Console](https://console.aiven.io/), open your Aiven for Apache Kafka service. 2. On the service **Overview** page, in the **Set up your stream** section, click **Quick connect**. 3. At the top of the page, click **Java**. ## Step 1: Set up a topic[​](#step-1-set-up-a-topic "Direct link to Step 1: Set up a topic") Topics organize and store the events that you stream to Apache Kafka. 1. In **Topic name**, do one of the following: * Choose an existing topic. * Click **Create new topic**, enter a name, and create the topic. The new topic becomes the topic for the next steps. note If your service has **Diskless topics** enabled, you can create a **Classic** or **Diskless** topic, or choose an existing topic of either type. You can't change the topic type after creation. For more information, see [Create Apache Kafka® topics](/docs/products/kafka/howto/create-topic.md). ## Step 2: Set up an authentication method[​](#step-2-set-up-an-authentication-method "Direct link to Step 2: Set up an authentication method") 1. Choose an authentication method: * **SASL**: Recommended. Use SASL/SCRAM-SHA-256 for username and password authentication. SASL is enabled by default for new Aiven for Apache Kafka services. For existing services, check whether SASL is enabled. If **Enable SASL** appears, click **Enable SASL**, then continue. * **Client certificate**: Use certificate-based authentication with mTLS instead of a password. 2. Choose a service user: * Choose an existing service user from the list. * To create one, click **Create new service user**, enter a username, and click **Add service user**. 3. Check the permission status shown for the selected user: * If the user has all the required permissions, the granted permissions are shown (for example, `read, write`). To change them, click **Manage access in ACLs**. * If the user has some or no permissions, click **Grant permissions**. Choose **Produce**, **Consume**, or both, and click **Save**. note The `avnadmin` user has permissions by default. For other service users, you can add permissions from Quick connect. To change or remove permissions, click **Manage access in ACLs**. For more information, see [Manage Apache Kafka® ACLs](/docs/products/kafka/howto/manage-acls.md#delete-acl-entries). ## Step 3: Copy the code snippets[​](#step-3-copy-the-code-snippets "Direct link to Step 3: Copy the code snippets") tip **Download template** is an optional shortcut for when you want no local setup. The ZIP file includes ready-to-run producer and consumer code, certificates, and dependencies, so you can verify the connection using the included `README.md` without a Java project or build tool. If you already have a Java project, copy the snippet from the **Producer** or **Consumer** tab. 1. Under **Prerequisites**, add the [Kafka clients dependency](https://mvnrepository.com/artifact/org.apache.kafka/kafka-clients) from your preferred artifact repository. 2. Under **Downloads**, download the certificate files for your authentication method: * For **SASL**, click **Download CA certificate**. * For **Client certificate**, click **Download CA certificate** and **Download service certificate and access key**. The generated snippet loads the PEM files directly. You don't need to create a truststore. 3. Click the **Producer** or **Consumer** tab to view the generated producer or consumer code. warning For SASL, code snippets include the service user password in plaintext. Store the code securely, and do not commit it to source control. 4. Copy the code. After you add the code to your project, update the certificate file paths to match where you saved the files, then run your producer or consumer to start streaming events. --- # Connect to Aiven for Apache Kafka® with Kafka REST Use **Quick connect** to produce and consume messages with the Apache Kafka® REST API for Aiven for Apache Kafka®. The guided flow helps you select or create a topic, select a service user, and copy generated `curl` commands. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md). * [Kafka REST enabled](/docs/products/kafka/karapace/howto/enable-karapace.md) on the service. * `curl` installed in your terminal. ## Open Quick connect[​](#open-quick-connect "Direct link to Open Quick connect") 1. In the [Aiven Console](https://console.aiven.io/), open your Aiven for Apache Kafka service. 2. On the service **Overview** page, in the **Set up your stream** section, click **Quick connect**. 3. At the top of the page, click **Kafka REST**. ## Step 1: Set up a topic[​](#step-1-set-up-a-topic "Direct link to Step 1: Set up a topic") Topics organize and store the events that you stream to Apache Kafka. 1. In **Topic name**, do one of the following: * Select an existing topic. * Click **Create new topic**, enter a name, and create the topic. The new topic is selected automatically for the next steps. note If your service has **Diskless topics** enabled, you can create a **Classic** or **Diskless** topic, or select an existing topic of either type. You can't change the topic type after creation. For more information, see [Create Apache Kafka® topics](/docs/products/kafka/howto/create-topic.md). ## Step 2: Set up authentication[​](#step-2-set-up-authentication "Direct link to Step 2: Set up authentication") Kafka REST authenticates using the service user's username and password. 1. Select a service user: * Select an existing service user from the list. * To create one, click **Create new service user**, enter a username, and click **Add service user**. 2. Check the permission status shown for the selected user: * If the user has all the required permissions, the granted permissions are shown. To change them, click **Manage access in ACLs**. * If the user has some or no permissions, click **Grant permissions**. Select **Produce**, **Consume**, or both, and click **Save**. note The `avnadmin` user has permissions by default. For other service users, you can add permissions from Quick connect. To change or remove permissions, click **Manage access in ACLs**. For more information, see [Manage Apache Kafka® ACLs](/docs/products/kafka/howto/manage-acls.md#delete-acl-entries). ## Step 3: Copy the code snippets[​](#step-3-copy-the-code-snippets "Direct link to Step 3: Copy the code snippets") 1. Copy the generated command for producing a message and run it in your terminal. A successful request returns the message partition and offset. 2. Copy the generated command for creating a consumer and run it in your terminal. The response includes `base_uri`. Copy this value to use in the subscribe and consume commands. warning Generated Kafka REST snippets include the service user password in plaintext. Store the commands securely, and do not commit them to source control. 3. Paste the `base_uri` value without `https://` into the field in Quick connect. 4. Copy and run the generated command for subscribing the consumer to the topic. A successful request returns HTTP status `204`. 5. Copy and run the generated command for consuming messages. Returns messages waiting to be consumed. After you run the generated commands, use the **Preview the messages in the console** link to inspect the data sent to the Kafka topic. --- # Connect to Aiven for Apache Kafka® with Klaw Use **Quick connect** to set up [Klaw](https://www.klaw-project.io/) for Aiven for Apache Kafka®. The guided flow helps you choose or create a topic, choose an authentication method, grant permissions, and copy the setup steps for Klaw. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md). * Access to an existing Klaw installation, or a local environment where you can install Klaw. * Java `keytool` installed in your terminal. * For client certificate authentication, OpenSSL installed in your terminal. ## Open Quick connect[​](#open-quick-connect "Direct link to Open Quick connect") 1. In the [Aiven Console](https://console.aiven.io/), open your Aiven for Apache Kafka service. 2. On the service **Overview** page, in the **Set up your stream** section, click **Quick connect**. 3. At the top of the page, click **Klaw**. ## Step 1: Set up a topic[​](#step-1-set-up-a-topic "Direct link to Step 1: Set up a topic") Topics organize and store the events that you stream to Apache Kafka. 1. In **Topic name**, do one of the following: * Choose an existing topic. * Click **Create new topic**, enter a name, and create the topic. The new topic is selected automatically for the next steps. note If your service has **Diskless topics** enabled, you can create a **Classic** or **Diskless** topic, or choose an existing topic of either type. You can't change the topic type after creation. For more information, see [Create Apache Kafka® topics](/docs/products/kafka/howto/create-topic.md). ## Step 2: Set up an authentication method[​](#step-2-set-up-an-authentication-method "Direct link to Step 2: Set up an authentication method") 1. Choose an authentication method: * **SASL**: Recommended. Use SASL/SCRAM-SHA-256 for username and password authentication. SASL is enabled by default for new Aiven for Apache Kafka services. For existing services, check whether SASL is enabled. If **Enable SASL** appears, click **Enable SASL**, then continue. * **Client certificate**: Use certificate-based authentication with mTLS instead of a password. 2. Choose a service user: * Choose an existing service user from the list. * To create one, click **Create new service user**, enter a username, and click **Add service user**. 3. Check the permission status shown for the selected user: * If the user has all the required permissions, the granted permissions are shown (for example, `read, write`). To change them, click **Manage access in ACLs**. * If the user has some or no permissions, click **Grant permissions**. Choose **Produce**, **Consume**, or both, and click **Save**. note The `avnadmin` user has permissions by default. For other service users, you can add permissions from Quick connect. To change or remove permissions, click **Manage access in ACLs**. For more information, see [Manage Apache Kafka® ACLs](/docs/products/kafka/howto/manage-acls.md#delete-acl-entries). ## Step 3: Copy the code snippets[​](#step-3-copy-the-code-snippets "Direct link to Step 3: Copy the code snippets") 1. Under **Downloads**, download the certificate files for your authentication method: * For **SASL**, click **Download CA certificate**. * For **Client certificate**, click **Download CA certificate**, **Download service certificate**, and **Download service access key**. 2. Create the certificate stores required by Klaw: * For **SASL**, use the generated `keytool` command to create `client.truststore.jks` from `ca.pem`. * For **Client certificate**, use the generated `openssl` command to create `client.keystore.p12`, then use the generated `keytool` command to create `client.truststore.jks`. When prompted, enter a password for the store. For the truststore, enter `yes` to trust the CA certificate. 3. Under **Klaw setup**, choose the tab that matches your setup: * **Existing Klaw installation**: Configure the existing installation by connecting Klaw Core and Klaw Cluster APIs, connecting Klaw to Aiven for Apache Kafka using SSL, and synchronizing topics from the cluster. * **New Klaw installation**: Install Klaw from source with Docker, then connect Klaw Core and Klaw Cluster APIs, connect Klaw to Aiven for Apache Kafka using SSL, and synchronize topics from the cluster. Related pages * [Connect Aiven for Apache Kafka® with Klaw](/docs/products/kafka/howto/kafka-klaw.md) * [Install Klaw with Docker](https://www.klaw-project.io/docs/getting-started/quickstart/) * [Connect Klaw Core and Klaw Cluster APIs](https://www.klaw-project.io/docs/cluster-connectivity-setup/klaw-core-with-clusterapi/) * [Connect Klaw with Aiven for Apache Kafka® using SSL](https://www.klaw-project.io/docs/cluster-connectivity-setup/aiven-kafka-cluster-ssl-protocol/) * [Synchronize topics from the cluster](https://www.klaw-project.io/docs/cluster-management/kafka-cluster-sync/sync-topics-from-cluster/) --- # Connect to Aiven for Apache Kafka® with Node.js Use **Quick connect** to set up a Node.js client for Aiven for Apache Kafka®. The guided flow helps you select or create a topic, choose an authentication method, grant permissions, and copy the generated producer and consumer code. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md). * A Node.js development environment. * The [`node-rdkafka`](https://www.npmjs.com/package/node-rdkafka) library. note For SASL/SCRAM authentication with Apache Kafka® 4.x, use a `node-rdkafka` version built on `librdkafka` version 2.6.1 or later. For more information, see [Enable and configure SASL authentication for Apache Kafka®](/docs/products/kafka/howto/kafka-sasl-auth.md#configure-sasl-mechanisms). ## Open Quick connect[​](#open-quick-connect "Direct link to Open Quick connect") 1. In the [Aiven Console](https://console.aiven.io/), open your Aiven for Apache Kafka service. 2. On the service **Overview** page, in the **Set up your stream** section, click **Quick connect**. 3. At the top of the page, click **Node.js**. ## Step 1: Set up a topic[​](#step-1-set-up-a-topic "Direct link to Step 1: Set up a topic") Topics organize and store the events that you stream to Apache Kafka. 1. In **Topic name**, do one of the following: * Select an existing topic. * Click **Create new topic**, enter a name, and create the topic. The new topic is selected automatically for the next steps. note If your service has **Diskless topics** enabled, you can create a **Classic** or **Diskless** topic, or select an existing topic of either type. You can't change the topic type after creation. For more information, see [Create Apache Kafka® topics](/docs/products/kafka/howto/create-topic.md). ## Step 2: Set up an authentication method[​](#step-2-set-up-an-authentication-method "Direct link to Step 2: Set up an authentication method") 1. Select an authentication method: * **SASL**: Recommended. Use SASL/SCRAM-SHA-256 for username and password authentication. SASL is enabled by default for new Aiven for Apache Kafka services. For existing services, check whether SASL is enabled. If **Enable SASL** appears, click **Enable SASL**, then continue. * **Client certificate**: Use certificate-based authentication with mTLS instead of a password. 2. Select a service user: * Select an existing service user from the list. * To create one, click **Create new service user**, enter a username, and click **Add service user**. 3. Check the permission status shown for the selected user: * If the user has all the required permissions, the granted permissions are shown (for example, `read, write`). To change them, click **Manage access in ACLs**. * If the user has some or no permissions, click **Grant permissions**. Select **Produce**, **Consume**, or both, and click **Save**. note The `avnadmin` user has permissions by default. For other service users, you can add permissions from Quick connect. To change or remove permissions, click **Manage access in ACLs**. For more information, see [Manage Apache Kafka® ACLs](/docs/products/kafka/howto/manage-acls.md#delete-acl-entries). ## Step 3: Copy the code snippets[​](#step-3-copy-the-code-snippets "Direct link to Step 3: Copy the code snippets") tip **Download template** is an optional shortcut for testing the connection without a local Node.js environment. The ZIP file includes ready-to-run producer and consumer code, certificates, and dependencies. To run it, follow the included `README.md`. If you already have a Node.js project, copy the snippet from the **Producer** or **Consumer** tab. 1. Install the [`node-rdkafka`](https://www.npmjs.com/package/node-rdkafka) library: ``` npm install node-rdkafka ``` 2. Under **Downloads**, download the certificate files for your authentication method: * For **SASL**, click **Download CA certificate**. * For **Client certificate**, click **Download CA certificate**, **Download service certificate**, and **Download service access key**. For **Client certificate**, the generated snippet loads the certificates directly. You don't need to create a truststore. 3. Select the **Producer** or **Consumer** tab to view the generated producer or consumer code. 4. Copy the code. note For SASL, the snippet includes your service user's password. The console masks it by default; click the reveal icon to view it. The copied code contains the password in plaintext, so store it securely and don't commit it to source control. After you add the code to your project, update the certificate file paths to match where you saved the files, then run your producer or consumer to start streaming events. --- # Connect to Aiven for Apache Kafka® with PHP Use **Quick connect** to set up a PHP client for Aiven for Apache Kafka®. The guided flow helps you choose or create a topic, choose an authentication method, grant permissions, and copy generated producer and consumer code. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md). * PHP 8.3 or later. * The [`rdkafka` PECL extension](https://github.com/arnaud-lb/php-rdkafka), which requires the [`librdkafka`](https://github.com/confluentinc/librdkafka) system library. note For SASL/SCRAM authentication with Apache Kafka® 4.x, use a `php-rdkafka` version built on `librdkafka` version 2.6.1 or later. For more information, see [Enable and configure SASL authentication for Apache Kafka®](/docs/products/kafka/howto/kafka-sasl-auth.md#configure-sasl-mechanisms). ## Open Quick connect[​](#open-quick-connect "Direct link to Open Quick connect") 1. In the [Aiven Console](https://console.aiven.io/), open your Aiven for Apache Kafka service. 2. On the service **Overview** page, in the **Set up your stream** section, click **Quick connect**. 3. At the top of the page, click **PHP**. ## Step 1: Set up a topic[​](#step-1-set-up-a-topic "Direct link to Step 1: Set up a topic") Topics organize and store the events that you stream to Apache Kafka. 1. In **Topic name**, do one of the following: * Choose an existing topic. * Click **Create new topic**, enter a name, and create the topic. The new topic is selected automatically for the next steps. note If your service has **Diskless topics** enabled, you can create a **Classic** or **Diskless** topic, or choose an existing topic of either type. You can't change the topic type after creation. For more information, see [Create Apache Kafka® topics](/docs/products/kafka/howto/create-topic.md). ## Step 2: Set up an authentication method[​](#step-2-set-up-an-authentication-method "Direct link to Step 2: Set up an authentication method") 1. Choose an authentication method: * **SASL**: Recommended. Use SASL/SCRAM-SHA-256 for username and password authentication. SASL is enabled by default for new Aiven for Apache Kafka services. For existing services, verify whether SASL is enabled. If **Enable SASL** appears, click **Enable SASL**, then continue. * **Client certificate**: Use certificate-based authentication with mTLS instead of a password. Downloads include a CA certificate, service certificate, and access key. 2. Choose a service user: * Choose an existing service user from the list. * To create one, click **Create new service user**, enter a username, and click **Add service user**. 3. Check the permission status shown for the selected user: * If the user has all the required permissions, the granted permissions are shown (for example, `read, write`). To change them, click **Manage access in ACLs**. * If the user has some or no permissions, click **Grant permissions**. Choose **Produce**, **Consume**, or both, and click **Save**. note The `avnadmin` user has permissions by default. For other service users, you can add permissions from Quick connect. To change or remove permissions, click **Manage access in ACLs**. For more information, see [Manage Apache Kafka® ACLs](/docs/products/kafka/howto/manage-acls.md#delete-acl-entries). ## Step 3: Copy the code snippets[​](#step-3-copy-the-code-snippets "Direct link to Step 3: Copy the code snippets") tip **Download template** is an optional shortcut for testing the connection. The ZIP file includes ready-to-run producer and consumer code, certificates, and a `Containerfile` based on `php:8.3-cli`. To run it, follow the included `README.md`. If you already have a PHP project, copy the snippet from the **Producer** or **Consumer** tab. 1. Install `librdkafka`. For installation instructions, see [`librdkafka` on GitHub](https://github.com/confluentinc/librdkafka). 2. Install the `rdkafka` PECL extension: ``` pecl install rdkafka ``` 3. If the extension is not already enabled, add `extension=rdkafka.so` to your `php.ini` file. The generated examples require PHP CLI with the `pcntl` extension enabled for signal handling. 4. Under **Downloads**, download the certificate files for your authentication method: * For **SASL**, click **Download CA certificate**. * For **Client certificate**, click **Download CA certificate**, **Download service certificate**, and **Download service access key**. The generated snippet loads the certificates directly. You don't need to create a truststore. 5. Click the **Producer** or **Consumer** tab to view the generated producer or consumer code. 6. Copy the code into `producer.php` or `consumer.php`. warning For SASL, copied snippets include the service user password in plaintext. Store the code securely, and do not commit it to source control. After you add the code to your project, update the certificate file paths to match where you saved the files, then run your producer or consumer to start streaming events. --- # Connect to Aiven for Apache Kafka® with Python Use **Quick connect** to set up a Python client for Aiven for Apache Kafka®. The guided flow helps you select or create a topic, choose an authentication method, grant permissions, and copy generated producer and consumer code. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md). * A Python development environment. * The [`kafka-python`](https://pypi.org/project/kafka-python/) library. ## Open Quick connect[​](#open-quick-connect "Direct link to Open Quick connect") 1. In the [Aiven Console](https://console.aiven.io/), open your Aiven for Apache Kafka service. 2. On the service **Overview** page, in the **Set up your stream** section, click **Quick connect**. 3. At the top of the page, click **Python**. ## Step 1: Set up a topic[​](#step-1-set-up-a-topic "Direct link to Step 1: Set up a topic") Topics organize and store the events that you stream to Apache Kafka. 1. In **Topic name**, do one of the following: * Select an existing topic. * Click **Create new topic**, enter a name, and create the topic. The new topic is selected automatically for the next steps. note If your service has **Diskless topics** enabled, you can create a **Classic** or **Diskless** topic, or select an existing topic of either type. You can't change the topic type after creation. For more information, see [Create Apache Kafka® topics](/docs/products/kafka/howto/create-topic.md). ## Step 2: Set up an authentication method[​](#step-2-set-up-an-authentication-method "Direct link to Step 2: Set up an authentication method") 1. Select an authentication method: * **SASL**: Recommended. Use SASL/SCRAM-SHA-256 for username and password authentication. SASL is enabled by default for new Aiven for Apache Kafka services. For existing services, check whether SASL is enabled. If **Enable SASL** appears, click **Enable SASL**, then continue. * **Client certificate**: Use certificate-based authentication with mTLS instead of a password. 2. Select a service user: * Select an existing service user from the list. * To create one, click **Create new service user**, enter a username, and click **Add service user**. 3. Check the permission status shown for the selected user: * If the user has all the required permissions, the granted permissions are shown (for example, `read, write`). To change them, click **Manage access in ACLs**. * If the user has some or no permissions, click **Grant permissions**. Select **Produce**, **Consume**, or both, and click **Save**. note The `avnadmin` user has permissions by default. For other service users, you can add permissions from Quick connect. To change or remove permissions, click **Manage access in ACLs**. For more information, see [Manage Apache Kafka® ACLs](/docs/products/kafka/howto/manage-acls.md#delete-acl-entries). ## Step 3: Copy the code snippets[​](#step-3-copy-the-code-snippets "Direct link to Step 3: Copy the code snippets") tip **Download template** is an optional shortcut for testing the connection without a local Python environment. The ZIP file includes ready-to-run producer and consumer code, certificates, and dependencies. To run it, follow the included `README.md`. If you already have a Python project, copy the snippet from the **Producer** or **Consumer** tab. For more Python examples, under **Useful resources**, open **Kafka + Python: Getting Started**. 1. Under **Prerequisites**, install the [`kafka-python`](https://pypi.org/project/kafka-python/) library: ``` python3 -m pip install kafka-python ``` warning For SASL, copied snippets include the service user password in plaintext. Store the code securely, and do not commit it to source control. 2. Under **Downloads**, download the certificate files for your authentication method: * For **SASL**, click **Download CA certificate**. * For **Client certificate**, click **Download CA certificate**, **Download service certificate**, and **Download service access key**. The generated snippet loads the certificates directly. You don't need to create a truststore. 3. Select the **Producer** or **Consumer** tab to view the generated producer or consumer code. 4. Copy the code. After you add the code to your project, update the certificate file paths to match where you saved the files, then run your producer or consumer to start streaming events. --- # Controlled upgrade pipelines for your Aiven for Apache Kafka® service [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Link Aiven for Apache Kafka® services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. Control when your Aiven managed services receive maintenance updates and test maintenance updates in development or staging environments before they reach production. important Controlled upgrade pipeline is a [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-) feature. [Contact Aiven](https://aiven.io/contact) to request access. Aiven performs automatic service maintenance for security fixes, minor software updates, and other platform changes. The controlled upgrade pipeline feature lets you link services of the same type in an ordered sequence to control when each service receives updates. After a maintenance update upgrades a service at the initial pipeline step, you validate that service version before the update proceeds to the service at the next pipeline step. Validating means approving the new version as safe to roll out to the next service. Validation can be manual or automatic after a configurable delay. ## Why use controlled upgrade pipelines[​](#why-use-controlled-upgrade-pipelines "Direct link to Why use controlled upgrade pipelines") Controlled upgrade pipelines prevent production incidents caused by automatic updates reaching production before teams can test the new version in a lower environment. They give you full oversight of the update process: * **Risk mitigation**: Prevents unexpected maintenance updates from breaking your production environment by ensuring they are tested in a non-production setting first. * **Stability**: Keeps destination services (such as production) on a known-good version until you, or the automatic timer, confirm the new version is safe. * **Process control**: Allows platform teams to standardize their deployment and maintenance lifecycle across environments. ## About controlled upgrade pipelines[​](#about-controlled-upgrade-pipelines "Direct link to About controlled upgrade pipelines") ### Upgrade steps[​](#upgrade-steps "Direct link to Upgrade steps") An upgrade step is a pair of services linked by an upgrade constraint: * **Source service**: The service that receives maintenance updates first * **Destination service**: The service that waits for validation before receiving updates Each destination service can have only one source service. A source service can have multiple destination services. ### Upgrade pipelines[​](#upgrade-pipelines "Direct link to Upgrade pipelines") An upgrade pipeline is a chain of upgrade steps that spans multiple environments. For example: * Single chain: development → staging → production * Multiple destinations: development → production-eu and development → production-na ## How validation works[​](#how-validation-works "Direct link to How validation works") When a maintenance update upgrades your source service: 1. The source service receives the update first. 2. Test the updated source service to verify it works as expected. 3. Validate the update manually using the API or CLI, or wait for automatic validation after the configured delay. The default delay is 7 days. 4. After validation, the destination service becomes eligible for the same maintenance update. 5. The destination service receives the update during its next maintenance window. If one source service has multiple destination services, one validation for the source service applies to all connected destination services. ### Validation and maintenance windows[​](#validation-and-maintenance-windows "Direct link to Validation and maintenance windows") Validation and the maintenance window control different things: * **Validation** controls *what* version the destination service upgrades to. * The maintenance window controls *when* the upgrade happens. After you validate an update, or automatic validation applies, the destination service receives the validated version during its next scheduled maintenance window. Validation does not trigger an immediate upgrade outside the maintenance window. Upgrade pipelines add a constraint on what is installed during a maintenance update; they do not change when maintenance runs. Nodes in the destination service maintain the validated version until a newer version is validated, either when you validate it manually or when automatic validation applies after the configured delay. When a node is recycled, it uses the same validated version, not the latest available version. When you create a step, the destination service keeps the newest version that is already validated at that moment. If the destination service is already applying maintenance during step creation, the in-progress target version becomes the initial validated version. warning A powered-off source service cannot receive maintenance updates, so you cannot validate it. If you power off services earlier in the chain, the destination service upgrades regardless. For example, in a development → staging → production chain, if both development and staging are powered off, production upgrades without testing and validation in the earlier environments. Keep services in the chain powered on to preserve the protection that upgrade pipelines provide. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * **Same service type**: You can only link services of the same type. For example, two Aiven for PostgreSQL services. * **Chain length**: The default maximum chain depth is 3 services, which is 2 steps. If you need a longer chain, [contact Aiven](https://aiven.io/contact). * **No cycles**: You cannot create circular dependencies between services. * **Emergency overrides**: Aiven can apply critical security or stability fixes to a destination service before explicit validation. * **Supported services**: This feature supports all Aiven service types except Aiven for Apache Flink® and Aiven for MySQL. * **Automatic maintenance updates only**: Pipelines apply to automatic maintenance updates, such as minor service version updates and node image updates. Major version upgrades, for example Aiven for PostgreSQL® 15 to 16, require manual action and are not promoted automatically through the pipeline. * **No permanent blocking**: You cannot prevent an update indefinitely. Automatic validation applies after the configured delay, up to the maximum delay. * **No validation rollback**: You cannot undo a validation after it is recorded. ## Use controlled upgrade pipelines[​](#use-controlled-upgrade-pipelines "Direct link to Use controlled upgrade pipelines") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") To use controlled upgrade pipelines, you need the following: * The feature enabled by Aiven ([Limited availability](/docs/platform/concepts/service-and-feature-releases.md)) * Dev tool of your choice: * [Aiven CLI](/docs/tools/cli.md) Install the latest version of the Aiven CLI to access the `upgrade-pipeline` commands. * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) * Set `PROVIDER_AIVEN_ENABLE_BETA=true` before running Terraform. * See the [resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) for full schema, import format, and lifecycle behavior. * [Aiven Operator for Kubernetes](/docs/tools/kubernetes.md) Install the operator and create an Aiven token secret named `aiven-token` that the operator uses to authenticate against the Aiven API. * Write access to the source and destination projects * At least two services of the same type (for example, two Aiven for PostgreSQL® services) * Services can be in different projects in the same organization ### Set up an upgrade pipeline[​](#set-up-an-upgrade-pipeline "Direct link to Set up an upgrade pipeline") Use the Aiven CLI or API to create upgrade steps between your services. note The `upgrade-pipeline` CLI commands require Aiven CLI version 4.x or later. Command names and parameters may change before general availability. #### Create an upgrade step[​](#create-an-upgrade-step "Direct link to Create an upgrade step") Create a step to link a source service and a destination service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ [--source-project SOURCE_PROJECT] SOURCE_SERVICE \ [--destination-project DESTINATION_PROJECT] DESTINATION_SERVICE \ [--auto-validation-delay-days DAYS] ``` **Options** * `--organization-id` is required. * `--source-project` and `--destination-project` are optional. If you omit either project option, Aiven CLI uses the current default project set with `avn project switch`. * `--auto-validation-delay-days` is optional. Defaults to 7 days if not specified. ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "SOURCE_PROJECT_NAME", "source_service_name": "SOURCE_SERVICE_NAME", "destination_project_name": "DESTINATION_PROJECT_NAME", "destination_service_name": "DESTINATION_SERVICE_NAME", "auto_validation_delay_days": 7 }' ``` Use the [`aiven_upgrade_step`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) resource: ``` resource "aiven_upgrade_step" "example" { organization_id = "ORGANIZATION_ID" source_project_name = "SOURCE_PROJECT_NAME" source_service_name = "SOURCE_SERVICE_NAME" destination_project_name = "DESTINATION_PROJECT_NAME" destination_service_name = "DESTINATION_SERVICE_NAME" auto_validation_delay_days = 7 } ``` Apply an `UpgradePipelineStep` manifest with `kubectl`: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: upgrade-step-sample spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: SOURCE_PROJECT_NAME sourceServiceName: SOURCE_SERVICE_NAME destinationProjectName: DESTINATION_PROJECT_NAME destinationServiceName: DESTINATION_SERVICE_NAME autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable after the resource is created. Parameters: * `source_project_name`: Name of the project containing the source service * `source_service_name`: Name of the source service * `destination_project_name`: Name of the project containing the destination service * `destination_service_name`: Name of the destination service * `auto_validation_delay_days`: Optional. Number of days before automatic validation. The value must be at least `1`. The default is 7 days. The maximum delay you can configure is 30 days. #### List upgrade steps[​](#list-upgrade-steps "Direct link to List upgrade steps") View all upgrade steps you have access to: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step list --organization-id ORGANIZATION_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) To list managed upgrade steps, use: ``` terraform state list 'aiven_upgrade_step.*' terraform state show 'aiven_upgrade_step.example' ``` List `UpgradePipelineStep` resources in the current namespace: ``` kubectl get upgradepipelinesteps ``` #### View a specific step[​](#view-a-specific-step "Direct link to View a specific step") Get details about a specific upgrade step, including the last validation: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step get \ --organization-id ORGANIZATION_ID \ STEP_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` terraform state show aiven_upgrade_step.example ``` Show the manifest and full status, including `id`, `conditions`, and `lastValidation`: ``` kubectl describe upgradepipelinestep RESOURCE_NAME kubectl get upgradepipelinestep RESOURCE_NAME -o yaml ``` The step details include `last_validation` values such as `validated_at`, `validated_by_user`, and `comment` when validation exists (available through the API). ### Validate an upgrade[​](#validate-an-upgrade "Direct link to Validate an upgrade") After testing your source service with the new update, validate the version to allow the destination service to receive the same update. #### Manual validation[​](#manual-validation "Direct link to Manual validation") Validate the current version of your source service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step validate-for-service \ --project SOURCE_PROJECT \ SERVICE_NAME \ [--comment "COMMENT"] ``` `--comment` is optional. Use it to record a note about the validation, for example `"Tested and verified in development"`. ``` curl -X POST https://api.aiven.io/v1/project/SOURCE_PROJECT/service/SOURCE_SERVICE/upgrade-validation \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "comment": "Tested and verified in development" }' ``` Terraform manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. The operator manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. #### Automatic validation[​](#automatic-validation "Direct link to Automatic validation") If you do not manually validate an update, the system automatically validates the source service version after the configured delay. Auto-validation starts from when the source service receives the update. ### Manage upgrade steps[​](#manage-upgrade-steps "Direct link to Manage upgrade steps") #### Update a step[​](#update-a-step "Direct link to Update a step") Modify the automatic validation delay for an existing step: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step update \ --organization-id ORGANIZATION_ID \ --auto-validation-delay-days 14 \ STEP_ID ``` ``` curl -X PATCH https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "auto_validation_delay_days": 14 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` resource "aiven_upgrade_step" "example" { # ...required fields... auto_validation_delay_days = 14 # Updated from 7 to 14 } ``` Apply the changes: ``` terraform plan terraform apply ``` Edit `autoValidationDelayDays` in your manifest and re-apply: ``` spec: autoValidationDelayDays: 14 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable. To change them, delete the resource and create a new one. #### Delete a step[​](#delete-a-step "Direct link to Delete a step") Remove an upgrade step to allow the destination service to receive updates independently: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step delete --organization-id ORGANIZATION_ID STEP_ID ``` Find `STEP_ID` from the upgrade step list command. ``` curl -X DELETE https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) Remove the resource from configuration and apply, or destroy it directly: ``` terraform apply terraform destroy -target=aiven_upgrade_step.example ``` Delete the `UpgradePipelineStep` resource: ``` kubectl delete upgradepipelinestep RESOURCE_NAME ``` Deleting a step removes all associated validations. ### Example: Three-environment pipeline[​](#example-three-environment-pipeline "Direct link to Example: Three-environment pipeline") Create a pipeline that promotes updates from development to staging to production: * CLI * API * Terraform * Kubernetes 1. Create a step from development to staging: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project dev-project \ --destination-project staging-project \ --auto-validation-delay-days 3 \ pg-dev pg-staging ``` 2. Create a step from staging to production: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project staging-project \ --destination-project prod-project \ --auto-validation-delay-days 7 \ pg-staging pg-prod ``` 1) Create a step from development to staging: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "dev-project", "source_service_name": "pg-dev", "destination_project_name": "staging-project", "destination_service_name": "pg-staging", "auto_validation_delay_days": 3 }' ``` 2) Create a step from staging to production: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "staging-project", "source_service_name": "pg-staging", "destination_project_name": "prod-project", "destination_service_name": "pg-prod", "auto_validation_delay_days": 7 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` # Step 1: Development → Staging resource "aiven_upgrade_step" "dev_to_staging" { organization_id = "ORGANIZATION_ID" source_project_name = "dev-project" source_service_name = "pg-dev" destination_project_name = "staging-project" destination_service_name = "pg-staging" auto_validation_delay_days = 3 } # Step 2: Staging → Production resource "aiven_upgrade_step" "staging_to_prod" { organization_id = "ORGANIZATION_ID" source_project_name = "staging-project" source_service_name = "pg-staging" destination_project_name = "prod-project" destination_service_name = "pg-prod" auto_validation_delay_days = 7 } ``` Apply the configuration: ``` export PROVIDER_AIVEN_ENABLE_BETA=true terraform init terraform plan terraform apply ``` Define both steps in a single manifest and apply it: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: dev-to-staging spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: dev-project sourceServiceName: pg-dev destinationProjectName: staging-project destinationServiceName: pg-staging autoValidationDelayDays: 3 --- apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: staging-to-prod spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: staging-project sourceServiceName: pg-staging destinationProjectName: prod-project destinationServiceName: pg-prod autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-pipeline.yaml ``` When a maintenance update arrives: 1. The development service receives the update. 2. After testing, validate the development version or wait 3 days for auto-validation. 3. The staging service receives the update during its next maintenance window. 4. After testing, validate the staging version or wait 7 days for auto-validation. 5. The production service receives the update during its next maintenance window. Related pages * [Maintenance and updates for your Aiven for Apache Kafka® service](/docs/products/kafka/howto/maintenance-updates.md) * [Change the service plan](/docs/products/kafka/howto/change-service-plan.md) * [Service and feature releases](/docs/platform/concepts/service-and-feature-releases.md) * [Aiven CLI](/docs/tools/cli.md) --- # Create Apache Kafka® topics Create topics in your Aiven for Apache Kafka® service to organize message streams between producers and consumers. ## Understand Kafka topics[​](#understand-kafka-topics "Direct link to Understand Kafka topics") A topic in Aiven for Apache Kafka is a named stream of messages. Producers write data to topics, and consumers read from them. Available topic types depend on the Kafka service type. ### Topic types[​](#topic-types "Direct link to Topic types") * **Classic topics** * They are available in **Classic Kafka** and **Standard Kafka** services. * Storage behavior varies by service type: * In **Classic Kafka** services, data is stored on local broker disks unless tiered storage is enabled. * In **Standard Kafka** services, classic topics use remote storage by default. You cannot change the storage mode or local retention settings. * Support log compaction. * **Diskless topics** * Available only in **Standard Kafka** services. * Store topic data directly in cloud object storage. * Require a service type and plan that support diskless topics. * Select the topic type when creating the topic. * You cannot change the topic type after creation. In Standard Kafka services, classic and diskless topics can coexist. To choose between them, see [Compare diskless and classic topics](/docs/products/kafka/diskless/concepts/topics-vs-classic.md#compare-classic-and-diskless-topics). For diskless topic limitations, see [Limitations of diskless topics](/docs/products/kafka/diskless/concepts/limitations.md). ## Before you begin[​](#before-you-begin "Direct link to Before you begin") Before you create topics in your Aiven for Apache Kafka service, review the following: * **Tiered storage** is available in **Classic Kafka** services, where you can enable and configure it per topic. See [Tiered storage](/docs/products/kafka/concepts/kafka-tiered-storage.md). * **Manual topic creation** gives you control over partitions, replication, and retention. This approach is preferred in production to prevent accidental topic creation. * **Automatic topic creation** is disabled by default. Enable it only when needed (for example, for development clusters). In production, leave it disabled to govern topic creation and avoid accidental topics. On Standard Kafka, auto-created topics are classic topics. Create diskless topics manually. For more information, see [Create topics automatically](/docs/products/kafka/howto/create-topics-automatically.md). ## Create an Apache Kafka topic[​](#create-an-apache-kafka-topic "Direct link to Create an Apache Kafka topic") The topic creation options shown in the Aiven Console depend on the Kafka service type. In Classic Kafka services, you create classic topics. Standard Kafka services let you choose between classic and diskless topics. tip Use an AI assistant connected to [Aiven MCP](/docs/tools/mcp-server.md) to create a topic and test producing and consuming messages. For example: > Create a topic named `orders` on `my-kafka-service` with three partitions. Produce a test message to the topic, and consume the message. Kafka CLI and client APIs on Standard Kafka * When you create a topic with the Kafka CLI (`kafka-topics.sh`), the Kafka Admin API, or other Kafka clients without specifying a topic type, the topic is created as a classic topic with remote storage by default. * To create a diskless topic, use the Aiven Console or Aiven CLI and explicitly select the diskless topic type. If you use `kafka-topics.sh` for diskless topics, use the script from the [Inkless repository](https://github.com/aiven/inkless). - Console - CLI - Terraform * Classic Kafka service * Standard Kafka service 1. In the [Aiven Console](https://console.aiven.io/), select the Aiven for Apache Kafka service. 2. In the sidebar, click **Manage stream** > **Topics**. 3. Click **Create topic**. 4. Enter a name in the **Topic** field. 5. Optional: Turn on **Enable advanced configurations** to set values such as replication, partitions, retention, and cleanup policy. 6. Optional: Click **Add configuration** to add other Kafka topic parameters. 7. Optional: Under **Tiered storage**, enable tiered storage for the topic. note You cannot disable tiered storage after you enable it. 8. Click **Create topic**. 1) In the [Aiven Console](https://console.aiven.io/), select the Aiven for Apache Kafka service. 2) In the sidebar, click **Manage stream** > **Topics**. 3) Click **Create topic**. 4) Enter a name in the **Topic** field. 5) In **Topic type**, select one of the following: * **Classic topic** * **Diskless topic** note In **Standard Kafka** services, classic topics use managed remote storage by default. You cannot change storage mode or local retention settings. 6) Optional: Turn on **Keep default settings** to configure advanced topic settings such as retention, cleanup policy, and message size limits. note Diskless topics do not support compaction. Create a **classic topic** if your workload requires compaction. 7) Click **Create topic**. After creation, the topic appears on the **Manage stream** > **Topics** page. The **Topic type** column shows **Classic** or **Diskless**. tip Analyze your topic data with SQL by sending it to Aiven for ClickHouse®. From the topic's **Actions** menu, click **Query in ClickHouse** to set up a managed integration. See [Query Kafka topic data in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/query-kafka-topic-data.md). When creating topics using the CLI, the required options depend on the topic type. * Classic topic * Diskless topic To create a classic topic: ``` avn service topic-create \ SERVICE_NAME \ TOPIC_NAME \ --project PROJECT_NAME \ --partitions PARTITION_COUNT \ --replication REPLICATION_FACTOR ``` To create a diskless topic (Standard Kafka services only): ``` avn service topic-create \ SERVICE_NAME \ TOPIC_NAME \ --project PROJECT_NAME \ --partitions PARTITION_COUNT \ --replication 1 \ --diskless-enable ``` For diskless topics, set `--replication 1`. Use the `--diskless-enable` flag to create a diskless topic. Enable diskless topics on the service before creating them. On Standard Kafka services, omitting the flag creates a classic topic with remote storage. Common parameters: * `PROJECT_NAME`: Aiven project that contains the Kafka service. * `SERVICE_NAME`: Aiven for Apache Kafka service name. * `TOPIC_NAME`: Name of the topic. * `PARTITION_COUNT`: Number of partitions to distribute messages. * `--replication REPLICATION_FACTOR`: Number of replicas for each partition. - Classic topic - Diskless topic Define topics using the [`aiven_kafka_topic`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/kafka_topic) resource. Diskless topics are supported only in Standard Kafka services. Enable diskless in the [`aiven_kafka`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/kafka) resource before defining topics with [`aiven_kafka_topic`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/kafka_topic). Related pages * [Compare diskless and classic topics](/docs/products/kafka/diskless/concepts/topics-vs-classic.md) * [Automatically create topics](/docs/products/kafka/howto/create-topics-automatically.md) * [Tiered storage](/docs/products/kafka/concepts/kafka-tiered-storage.md) * [Query Kafka topic data in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/query-kafka-topic-data.md) --- # Create Apache Kafka® topics automatically When you send a message to a topic that does not exist, Apache Kafka® can automatically create that topic. Automatic topic creation is off by default on both Classic Kafka and Standard Kafka services. Turn it on when you want topics to be created automatically (for example, in development). In production, leave it off to avoid creating topics by mistake (for example, due to typos or incorrectly configured clients). If automatic topic creation is disabled and a producer sends to a non-existent topic, you may see errors such as: ``` UNKNOWN_TOPIC_OR_PARTITION ``` For example: ``` WARN [Producer clientId=console-producer] The metadata response from the cluster reported a recoverable issue with correlation id 18: {example-topic=UNKNOWN_TOPIC_OR_PARTITION} (org.apache.kafka.clients.NetworkClient) ``` You can: 1. **Create topics in advance**: [Create the topics](/docs/products/kafka/howto/create-topic.md) before use. Prefer this in production so you control partition count, replication factor, and retention. 2. **Enable automatic topic creation**: This is simpler, but you cannot set partitions, replication, or topic type. The topic uses [default configuration values](/docs/products/kafka/howto/set-kafka-parameters.md). note If [tiered storage is enabled](/docs/products/kafka/howto/enable-kafka-tiered-storage.md) on your Classic Kafka service, new topics use tiered storage by default. On Standard Kafka services, automatically created topics are classic topics with remote storage. Diskless topics are not auto-created. Create them manually. ## Enable automatic topic creation[​](#enable-automatic-topic-creation "Direct link to Enable automatic topic creation") * Console * CLI Enable automatic topic creation in the Aiven Console: 1. In the [Aiven Console](https://console.aiven.io/), open your project and your Aiven for Apache Kafka® service. 2. Click **Service settings**. 3. Scroll to **Advanced configuration** and click **Configure**. 4. In the **Advanced configuration** dialog, click **Add Advanced Configuration**. 5. Find `auto_create_topics_enable` and set it to `true`. 6. Click **Save configuration**. warning With automatic topic creation enabled, the user who sends the message must have admin permissions. To change permissions, open the **Users** tab on the service page in the [Aiven Console](https://console.aiven.io/). To enable automatic topic creation, use the [Aiven CLI `service update` command](/docs/tools/cli/service-cli.md#avn-cli-service-update). Replace `SERVICE_NAME` with your service name: ``` avn service update SERVICE_NAME -c kafka.auto_create_topics_enable=true ``` Parameters: * `auto_create_topics_enable=true`: Enable automatic topic creation. * `SERVICE_NAME`: The name of your Aiven for Apache Kafka service. Related pages * [Manage Aiven for Apache Kafka® topics via CLI](/docs/tools/cli/service/topic.md#avn_cli_service_topic_create) * [Create an Apache Kafka® topic](/docs/products/kafka/howto/create-topic.md) --- # Apache Kafka® metrics sent to Datadog When you configure a [Datadog service integration](https://docs.datadoghq.com/integrations/kafka/?tab=host#kafka-consumer-integration) for Aiven for Apache Kafka®, Aiven sends Kafka metrics to Datadog. You can also customize selected metrics using the [Aiven CLI](/docs/tools/cli.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running Aiven for Apache Kafka® service * A Datadog account * A Datadog [API key](https://docs.datadoghq.com/account_management/api-app-keys/) * A [Datadog integration endpoint](/docs/integrations/datadog/datadog-metrics.md#add-a-datadog-metrics-integration-to-an-aiven-service) note Datadog integration is not available for new Startup-2 plans in Aiven for Apache Kafka®. Existing customers who already use Startup-2 with Datadog integration can continue to create Startup-2 services with Datadog integration and use their existing services without upgrading to a higher plan. Aiven recommends Business-4 or higher for Aiven for Apache Kafka® services with Datadog integration to avoid resource pressure on Startup-2 plans. If you are an existing customer and cannot create a Startup-2 service with Datadog integration in a new project, contact [Aiven Support](/docs/platform/howto/support.md). ## Metrics sent to Datadog[​](#metrics-sent-to-datadog "Direct link to Metrics sent to Datadog") When you configure a Datadog integration for your Aiven for Apache Kafka® service, Aiven collects Kafka metrics covering broker throughput, request handling, replication, the controller, consumers, producers, the JVM, quotas, and tiered storage. Datadog uses a different naming convention from [Prometheus metrics](/docs/products/kafka/reference/kafka-metrics-prometheus.md). Metrics in this section are grouped by category to help you find the relevant Datadog metric. note The available metrics depend on your Apache Kafka version, service configuration, and metadata mode: * KRaft controller metrics apply to services running in KRaft mode (Apache Kafka 3.9 and later). * Group coordinator metrics apply to the new group coordinator (Apache Kafka 4.x and later). * ZooKeeper metrics apply only to services running in ZooKeeper mode. ### Broker throughput metrics[​](#broker-throughput-metrics "Direct link to Broker throughput metrics") | Metric | Description | | --------------------------------------- | ------------------------------------------------------ | | `kafka.messages_in.rate` | Rate of messages received by the broker | | `kafka.net.bytes_in.rate` | Rate of bytes received by the broker across all topics | | `kafka.net.bytes_out.rate` | Rate of bytes sent by the broker across all topics | | `kafka.net.bytes_rejected.rate` | Rate of bytes rejected by the broker | | `kafka.net.processor.avg.idle.pct.rate` | Average idle percentage of network processor threads | ### Per-topic throughput metrics[​](#per-topic-throughput-metrics "Direct link to Per-topic throughput metrics") | Metric | Description | | ------------------------------------- | ----------------------------------------------------------------------- | | `kafka.topic.messages_in.rate` | Rate of incoming messages per topic (tagged by `topic` and `partition`) | | `kafka.topic.net.bytes_in.rate` | Rate of incoming bytes per topic | | `kafka.topic.net.bytes_out.rate` | Rate of outgoing bytes per topic | | `kafka.topic.net.bytes_rejected.rate` | Rate of rejected bytes per topic | ### Request handling metrics[​](#request-handling-metrics "Direct link to Request handling metrics") | Metric | Description | | ------------------------------------------------- | ------------------------------------------------------------------ | | `kafka.request.channel.queue.size` | Number of requests in the request queue | | `kafka.request.handler.avg.idle.pct.rate` | Average idle percentage of request handler threads (1-minute rate) | | `kafka.request.produce.rate` | Rate of Produce requests | | `kafka.request.produce.failed.rate` | Rate of failed Produce requests | | `kafka.request.produce.time.avg` | Average total time for Produce requests (ms) | | `kafka.request.produce.time.99percentile` | 99th-percentile total time for Produce requests (ms) | | `kafka.request.fetch_consumer.rate` | Rate of FetchConsumer requests | | `kafka.request.fetch_consumer.time.avg` | Average total time for FetchConsumer requests (ms) | | `kafka.request.fetch_consumer.time.99percentile` | 99th-percentile total time for FetchConsumer requests (ms) | | `kafka.request.fetch_follower.rate` | Rate of FetchFollower requests | | `kafka.request.fetch_follower.time.avg` | Average total time for FetchFollower requests (ms) | | `kafka.request.fetch_follower.time.99percentile` | 99th-percentile total time for FetchFollower requests (ms) | | `kafka.request.fetch.failed.rate` | Rate of failed Fetch requests | | `kafka.request.metadata.time.avg` | Average total time for Metadata requests (ms) | | `kafka.request.metadata.time.99percentile` | 99th-percentile total time for Metadata requests (ms) | | `kafka.request.offsets.time.avg` | Average total time for Offsets requests (ms) | | `kafka.request.offsets.time.99percentile` | 99th-percentile total time for Offsets requests (ms) | | `kafka.request.update_metadata.time.avg` | Average total time for UpdateMetadata requests (ms) | | `kafka.request.update_metadata.time.99percentile` | 99th-percentile total time for UpdateMetadata requests (ms) | | `kafka.request.producer_request_purgatory.size` | Number of requests waiting in the producer purgatory | | `kafka.request.fetch_request_purgatory.size` | Number of requests waiting in the fetch purgatory | ### Replication metrics[​](#replication-metrics "Direct link to Replication metrics") | Metric | Description | | ------------------------------------------------- | -------------------------------------------------------------------------------- | | `kafka.replication.active_controller_count` | Number of active controllers (should be 1) | | `kafka.replication.leader_count` | Number of partitions for which this broker is the leader | | `kafka.replication.partition_count` | Number of partitions on this broker | | `kafka.replication.under_replicated_partitions` | Number of under-replicated partitions | | `kafka.replication.offline_partitions_count` | Number of offline partitions | | `kafka.replication.isr_expands.rate` | Rate of in-sync replica (ISR) expand events | | `kafka.replication.isr_shrinks.rate` | Rate of ISR shrink events | | `kafka.replication.leader_elections.rate` | Rate of leader elections | | `kafka.replication.unclean_leader_elections.rate` | Rate of unclean leader elections. Unclean leader elections can lead to data loss | | `kafka.replication.max_lag` | Maximum replication lag across all followers | ### KRaft controller metrics[​](#kraft-controller-metrics "Direct link to KRaft controller metrics") | Metric | Description | | ----------------------------------------------------------------- | ----------------------------------------------- | | `kafka.controller.active_broker_count` | Number of active brokers | | `kafka.controller.global_topic_count` | Total number of topics | | `kafka.controller.global_partition_count` | Total number of partitions | | `kafka.controller.topics_to_delete_count` | Number of topics pending deletion | | `kafka.controller.replicas_to_delete_count` | Number of replicas pending deletion | | `kafka.controller.election_from_eligible_leader_replicas_per_sec` | Rate of leader elections from eligible replicas | ### Group coordinator metrics[​](#group-coordinator-metrics "Direct link to Group coordinator metrics") | Metric | Description | | ------------------------------------------------------------------------ | ----------------------------------------------------------------- | | `kafka.server.group_coordinator_metrics.group_count` | Total number of consumer groups | | `kafka.server.group_coordinator_metrics.consumer_group_count` | Number of consumer groups in the new protocol | | `kafka.server.group_coordinator_metrics.consumer_group_rebalance_count` | Total count of consumer group rebalances | | `kafka.server.group_coordinator_metrics.consumer_group_rebalance_rate` | Rate of consumer group rebalances | | `kafka.server.group_coordinator_metrics.group_completed_rebalance_count` | Total count of completed group rebalances | | `kafka.server.group_coordinator_metrics.group_completed_rebalance_rate` | Rate of completed group rebalances | | `kafka.server.group_coordinator_metrics.streams_group_count` | Number of Kafka Streams groups | | `kafka.server.group_coordinator_metrics.streams_group_rebalance_count` | Total count of Kafka Streams group rebalances | | `kafka.server.group_coordinator_metrics.streams_group_rebalance_rate` | Rate of Kafka Streams group rebalances | | `kafka.server.group_coordinator_metrics.offset_commit_count` | Total count of offset commits | | `kafka.server.group_coordinator_metrics.offset_commit_rate` | Rate of offset commits | | `kafka.server.group_coordinator_metrics.offset_deletion_count` | Total count of offset deletions | | `kafka.server.group_coordinator_metrics.offset_deletion_rate` | Rate of offset deletions | | `kafka.server.group_coordinator_metrics.offset_expiration_count` | Total count of offset expirations | | `kafka.server.group_coordinator_metrics.offset_expiration_rate` | Rate of offset expirations | | `kafka.server.group_coordinator_metrics.num_partitions` | Number of partitions managed by the group coordinator | | `kafka.server.group_coordinator_metrics.partition_load_time_avg` | Average partition load time for the group coordinator (ms) | | `kafka.server.group_coordinator_metrics.partition_load_time_max` | Maximum partition load time for the group coordinator (ms) | | `kafka.server.group_coordinator_metrics.batch_flush_rate` | Rate of batch flushes in the group coordinator | | `kafka.server.group_coordinator_metrics.batch_flush_time_ms_max` | Maximum batch flush time in the group coordinator (ms) | | `kafka.server.group_coordinator_metrics.batch_linger_time_ms_max` | Maximum batch linger time in the group coordinator (ms) | | `kafka.server.group_coordinator_metrics.event_processing_time_ms_max` | Maximum event processing time in the group coordinator (ms) | | `kafka.server.group_coordinator_metrics.event_purgatory_time_ms_max` | Maximum time events spent in the group coordinator purgatory (ms) | | `kafka.server.group_coordinator_metrics.event_queue_size` | Size of the group coordinator event queue | | `kafka.server.group_coordinator_metrics.event_queue_time_ms_max` | Maximum time events waited in the group coordinator queue (ms) | | `kafka.server.group_coordinator_metrics.thread_idle_ratio_avg` | Average idle ratio of group coordinator threads | ### Group metadata manager metrics[​](#group-metadata-manager-metrics "Direct link to Group metadata manager metrics") | Metric | Description | | -------------------------------------------------------------------- | ------------------------------------------------- | | `kafka.server.group_metadata_manager.num_groups` | Number of consumer groups managed by this broker | | `kafka.server.group_metadata_manager.num_groups_preparing_rebalance` | Number of consumer groups preparing for rebalance | | `kafka.server.group_metadata_manager.num_offsets` | Number of committed offsets stored by this broker | ### Log metrics[​](#log-metrics "Direct link to Log metrics") | Metric | Description | | --------------------------- | ---------------------------- | | `kafka.log.flush_rate.rate` | Rate of log flush operations | note Per-partition log size and offset metrics (`kafka.log.log_size`, `kafka.log.log_start_offset`, and `kafka.log.log_end_offset`) are not collected by default. You can enable them by configuring the Datadog integration. See [Configurable custom metrics](#configurable-custom-metrics). ### JVM metrics[​](#jvm-metrics "Direct link to JVM metrics") | Metric | Description | | --------------------------------- | ---------------------------------------------- | | `jvm.heap_memory` | Heap memory used (bytes) | | `jvm.heap_memory_committed` | Heap memory committed (bytes) | | `jvm.heap_memory_init` | Heap memory initially requested (bytes) | | `jvm.heap_memory_max` | Maximum heap memory (bytes) | | `jvm.non_heap_memory` | Non-heap memory used (bytes) | | `jvm.non_heap_memory_committed` | Non-heap memory committed (bytes) | | `jvm.non_heap_memory_init` | Non-heap memory initially requested (bytes) | | `jvm.non_heap_memory_max` | Maximum non-heap memory (bytes) | | `jvm.gc.cms.count` | Number of concurrent (CMS) garbage collections | | `jvm.gc.parnew.time` | Time spent in ParNew garbage collection (ms) | | `jvm.gc.eden_size` | Size of the eden space (bytes) | | `jvm.gc.survivor_size` | Size of the survivor space (bytes) | | `jvm.gc.old_gen_size` | Size of the old generation (bytes) | | `jvm.gc.metaspace_size` | Size of the metaspace (bytes) | | `jvm.buffer_pool.direct.capacity` | Capacity of direct buffer pools (bytes) | | `jvm.buffer_pool.direct.count` | Number of direct buffers in the pool | | `jvm.buffer_pool.direct.used` | Memory used by direct buffer pools (bytes) | | `jvm.buffer_pool.mapped.capacity` | Capacity of mapped buffer pools (bytes) | | `jvm.buffer_pool.mapped.count` | Number of mapped buffers in the pool | | `jvm.buffer_pool.mapped.used` | Memory used by mapped buffer pools (bytes) | | `jvm.cpu_load.process` | JVM process CPU load | | `jvm.cpu_load.system` | System CPU load | | `jvm.thread_count` | Number of live JVM threads | | `jvm.loaded_classes` | Number of currently loaded classes | | `jvm.unloaded_classes` | Number of classes unloaded since JVM start | | `jvm.os.open_file_descriptors` | Number of open file descriptors | ### Quota metrics[​](#quota-metrics "Direct link to Quota metrics") | Metric | Description | | ------------------------------------- | ----------------------------------------------------- | | `kafka.bandwidth.quota.byte.rate` | Bandwidth quota byte rate per client or user | | `kafka.bandwidth.quota.throttle.time` | Bandwidth quota throttle time per client or user (ms) | | `kafka.request.quota.request.time` | Request quota request time per client or user (ms) | | `kafka.request.quota.throttle.time` | Request quota throttle time per client or user (ms) | ### Producer metrics[​](#producer-metrics "Direct link to Producer metrics") | Metric | Description | | --------------------------------------- | ------------------------------------------------------------------ | | `kafka.producer.request_rate` | Producer request rate | | `kafka.producer.response_rate` | Producer response rate | | `kafka.producer.request_latency_avg` | Average producer request latency (ms) | | `kafka.producer.request_latency_max` | Maximum producer request latency (ms) | | `kafka.producer.requests_in_flight` | Number of producer requests in flight | | `kafka.producer.message_rate` | Rate of messages sent by producers | | `kafka.producer.bytes_out` | Rate of bytes sent by producers | | `kafka.producer.record_send_rate` | Rate of records sent per topic (tagged by `topic` and `partition`) | | `kafka.producer.records_send_rate` | Rate of records sent by producers | | `kafka.producer.records_per_request` | Average records per producer request | | `kafka.producer.record_error_rate` | Rate of errored producer records | | `kafka.producer.record_retry_rate` | Rate of retried producer records | | `kafka.producer.record_size_avg` | Average producer record size (bytes) | | `kafka.producer.record_size_max` | Maximum producer record size (bytes) | | `kafka.producer.record_queue_time_avg` | Average record queue time for producers (ms) | | `kafka.producer.record_queue_time_max` | Maximum record queue time for producers (ms) | | `kafka.producer.batch_size_avg` | Average batch size for producers (bytes) | | `kafka.producer.batch_size_max` | Maximum batch size for producers (bytes) | | `kafka.producer.compression_rate` | Compression rate per topic (tagged by `topic` and `partition`) | | `kafka.producer.compression_rate_avg` | Average compression rate for producers | | `kafka.producer.available_buffer_bytes` | Available producer buffer memory (bytes) | | `kafka.producer.buffer_bytes_total` | Total producer buffer memory (bytes) | | `kafka.producer.bufferpool_wait_time` | Time producer threads blocked on the buffer pool (ms) | | `kafka.producer.waiting_threads` | Number of waiting producer threads | | `kafka.producer.io_wait` | Average I/O wait time for producers (ns) | | `kafka.producer.metadata_age` | Age of producer metadata (seconds) | | `kafka.producer.throttle_time_avg` | Average producer throttle time (ms) | | `kafka.producer.throttle_time_max` | Maximum producer throttle time (ms) | ### Consumer metrics[​](#consumer-metrics "Direct link to Consumer metrics") | Metric | Description | | ---------------------------------------- | ------------------------------------------------------------------------------- | | `kafka.consumer.messages_in` | Rate of messages consumed | | `kafka.consumer.bytes_in` | Rate of bytes consumed | | `kafka.consumer.bytes_consumed` | Rate of bytes consumed per topic (tagged by `topic` and `partition`) | | `kafka.consumer.records_consumed` | Rate of records consumed per topic (tagged by `topic` and `partition`) | | `kafka.consumer.records_per_request_avg` | Average records per fetch request per topic (tagged by `topic` and `partition`) | | `kafka.consumer.fetch_rate` | Consumer fetch rate | | `kafka.consumer.fetch_size_avg` | Average fetch size per topic (tagged by `topic` and `partition`) | | `kafka.consumer.fetch_size_max` | Maximum fetch size per topic (tagged by `topic` and `partition`) | | `kafka.consumer.max_lag` | Maximum consumer lag | | `kafka.consumer.kafka_commits` | Rate of offset commits through Kafka (legacy) | | `kafka.consumer.zookeeper_commits` | Rate of offset commits through ZooKeeper (legacy) | note Some client-side producer and consumer metrics require additional configuration to appear in Datadog. See [Add client-side Apache Kafka® producer and consumer Datadog metrics](/docs/products/kafka/howto/add-missing-producer-consumer-metrics.md). ### Consumer lag and offset metrics[​](#consumer-lag-and-offset-metrics "Direct link to Consumer lag and offset metrics") | Metric | Description | | ----------------------- | ------------------------------------------------------------------------------------- | | `kafka.broker_offset` | Latest offset on the broker for a topic-partition (tagged by `topic` and `partition`) | | `kafka.consumer_offset` | Committed consumer offset for a topic-partition (tagged by `topic` and `partition`) | | `kafka.consumer_lag` | Consumer lag in messages (tagged by `topic` and `partition`) | ### ZooKeeper metrics[​](#zookeeper-metrics "Direct link to ZooKeeper metrics") These metrics apply only to services running in ZooKeeper mode. | Metric | Description | | ----------------------------------------- | --------------------------------------- | | `kafka.session.zookeeper.disconnect.rate` | Rate of ZooKeeper disconnections | | `kafka.session.zookeeper.expire.rate` | Rate of ZooKeeper session expirations | | `kafka.session.zookeeper.readonly.rate` | Rate of ZooKeeper read-only connections | | `kafka.session.zookeeper.sync.rate` | Rate of ZooKeeper sync connections | ### Tiered storage metrics[​](#tiered-storage-metrics "Direct link to Tiered storage metrics") For services with [tiered storage](/docs/products/kafka/concepts/kafka-tiered-storage.md) enabled, the following metrics are collected automatically to monitor the health and performance of tiered storage operations. #### Throughput metrics[​](#throughput-metrics "Direct link to Throughput metrics") | Metric | Description | | --------------------------------------------------------------- | ------------------------------------------- | | `kafka.tiered_storage.remote_copy_bytes.rate` | Rate of bytes copied to remote storage | | `kafka.tiered_storage.remote_copy_requests.rate` | Rate of copy requests to remote storage | | `kafka.tiered_storage.remote_fetch_bytes.rate` | Rate of bytes fetched from remote storage | | `kafka.tiered_storage.remote_fetch_requests.rate` | Rate of fetch requests from remote storage | | `kafka.tiered_storage.remote_delete_requests.rate` | Rate of delete requests to remote storage | | `kafka.tiered_storage.build_remote_log_aux_state_requests.rate` | Rate of remote log aux-state build requests | #### Error metrics[​](#error-metrics "Direct link to Error metrics") | Metric | Description | | ------------------------------------------------------------- | --------------------------------------------------------- | | `kafka.tiered_storage.remote_copy_errors.rate` | Rate of errors when copying segments to remote storage | | `kafka.tiered_storage.remote_fetch_errors.rate` | Rate of errors when fetching segments from remote storage | | `kafka.tiered_storage.remote_delete_errors.rate` | Rate of errors when deleting segments from remote storage | | `kafka.tiered_storage.build_remote_log_aux_state_errors.rate` | Rate of errors rebuilding remote log auxiliary state | #### Lag metrics[​](#lag-metrics "Direct link to Lag metrics") | Metric | Description | | ------------------------------------------------- | -------------------------------------------------- | | `kafka.tiered_storage.remote_copy_lag_bytes` | Bytes eligible for tiering but not yet copied | | `kafka.tiered_storage.remote_copy_lag_segments` | Segments eligible for tiering but not yet copied | | `kafka.tiered_storage.remote_delete_lag_bytes` | Bytes eligible for deletion but not yet deleted | | `kafka.tiered_storage.remote_delete_lag_segments` | Segments eligible for deletion but not yet deleted | #### Storage metrics[​](#storage-metrics "Direct link to Storage metrics") | Metric | Description | | ------------------------------------------------------- | ------------------------------------------ | | `kafka.tiered_storage.remote_log_size_bytes` | Total size of remote log in bytes | | `kafka.tiered_storage.remote_log_size_computation_time` | Time taken to compute remote log size (ms) | | `kafka.tiered_storage.remote_log_metadata_count` | Number of remote log metadata entries | #### Thread pool metrics[​](#thread-pool-metrics "Direct link to Thread pool metrics") | Metric | Description | | ---------------------------------------------------------------- | ---------------------------------------------------------- | | `kafka.tiered_storage.remote_log_manager_tasks_avg_idle_percent` | Average idle percentage of remote log manager task threads | | `kafka.tiered_storage.remote_log_reader_avg_idle_percent` | Average idle percentage of remote log reader threads | | `kafka.tiered_storage.remote_log_reader_task_queue_size` | Size of the remote log reader task queue | | `kafka.tiered_storage.remote_log_reader_fetch.rate` | Rate of remote log reader fetch operations | | `kafka.tiered_storage.remote_log_reader_fetch_time_avg` | Average time for remote log reader fetch operations (ms) | | `kafka.tiered_storage.remote_log_reader_fetch_time_99percentile` | 99th-percentile time for remote log reader fetch (ms) | | `kafka.tiered_storage.delayed_remote_fetch_expires.rate` | Rate of expired delayed remote fetch operations | #### Throttling metrics[​](#throttling-metrics "Direct link to Throttling metrics") | Metric | Description | | ----------------------------------------------------- | --------------------------------------------------- | | `kafka.tiered_storage.remote_copy_throttle_time_avg` | Average copy throttle time for remote storage (ms) | | `kafka.tiered_storage.remote_copy_throttle_time_max` | Maximum copy throttle time for remote storage (ms) | | `kafka.tiered_storage.remote_fetch_throttle_time_avg` | Average fetch throttle time for remote storage (ms) | | `kafka.tiered_storage.remote_fetch_throttle_time_max` | Maximum fetch throttle time for remote storage (ms) | #### Cache metrics[​](#cache-metrics "Direct link to Cache metrics") | Metric | Description | | ------------------------------------------------------------- | ------------------------------------------ | | `kafka.tiered_storage.cache.chunk_cache_size` | Total size of the chunk cache | | `kafka.tiered_storage.cache.chunk_cache_hits` | Chunk cache hit count | | `kafka.tiered_storage.cache.chunk_cache_misses` | Chunk cache miss count | | `kafka.tiered_storage.cache.chunk_cache_evictions` | Chunk cache eviction count | | `kafka.tiered_storage.cache.chunk_cache_eviction_weight` | Total eviction weight from the chunk cache | | `kafka.tiered_storage.cache.segment_manifest_cache_size` | Total size of the segment manifest cache | | `kafka.tiered_storage.cache.segment_manifest_cache_hits` | Segment manifest cache hit count | | `kafka.tiered_storage.cache.segment_manifest_cache_misses` | Segment manifest cache miss count | | `kafka.tiered_storage.cache.segment_manifest_cache_evictions` | Segment manifest cache eviction count | | `kafka.tiered_storage.cache.segment_indexes_cache_size` | Total size of the segment indexes cache | | `kafka.tiered_storage.cache.segment_indexes_cache_hits` | Segment indexes cache hit count | | `kafka.tiered_storage.cache.segment_indexes_cache_misses` | Segment indexes cache miss count | | `kafka.tiered_storage.cache.segment_indexes_cache_evictions` | Segment indexes cache eviction count | Metrics specific to your cloud storage provider and to the Aiven remote storage manager are listed by backend below. #### Amazon S3 tiered storage metrics[​](#amazon-s3-tiered-storage-metrics "Direct link to Amazon S3 tiered storage metrics") | Metric | Description | | ------------------------------------------------------------ | ----------------------------------------------------------- | | `kafka.tiered_storage.s3.get_object_requests_rate` | Rate of S3 GetObject requests | | `kafka.tiered_storage.s3.get_object_time_avg` | Average latency of S3 GetObject requests (ms) | | `kafka.tiered_storage.s3.delete_object_requests_rate` | Rate of S3 DeleteObject requests | | `kafka.tiered_storage.s3.upload_part_requests_rate` | Rate of S3 UploadPart requests | | `kafka.tiered_storage.s3.create_multipart_upload_time_avg` | Average latency of S3 CreateMultipartUpload requests (ms) | | `kafka.tiered_storage.s3.complete_multipart_upload_time_avg` | Average latency of S3 CompleteMultipartUpload requests (ms) | #### Google Cloud Storage tiered storage metrics[​](#google-cloud-storage-tiered-storage-metrics "Direct link to Google Cloud Storage tiered storage metrics") | Metric | Description | | --------------------------------------------------------- | ---------------------------------------- | | `kafka.tiered_storage.gcs.object_get_rate` | Rate of GCS object get requests | | `kafka.tiered_storage.gcs.object_delete_rate` | Rate of GCS object delete requests | | `kafka.tiered_storage.gcs.resumable_upload_initiate_rate` | Rate of GCS resumable upload initiations | | `kafka.tiered_storage.gcs.resumable_chunk_upload_rate` | Rate of GCS resumable chunk uploads | #### Azure Blob Storage tiered storage metrics[​](#azure-blob-storage-tiered-storage-metrics "Direct link to Azure Blob Storage tiered storage metrics") | Metric | Description | | --------------------------------------------------- | --------------------------------------- | | `kafka.tiered_storage.azure.blob_get_rate` | Rate of Azure Blob get requests | | `kafka.tiered_storage.azure.blob_upload_rate` | Rate of Azure Blob upload requests | | `kafka.tiered_storage.azure.blob_delete_rate` | Rate of Azure Blob delete requests | | `kafka.tiered_storage.azure.block_upload_rate` | Rate of Azure Block upload requests | | `kafka.tiered_storage.azure.block_list_upload_rate` | Rate of Azure BlockList upload requests | #### Aiven remote storage manager (RSM) tiered storage metrics[​](#aiven-remote-storage-manager-rsm-tiered-storage-metrics "Direct link to Aiven remote storage manager (RSM) tiered storage metrics") | Metric | Description | | ------------------------------------------------------ | --------------------------------------------------------- | | `kafka.tiered_storage.aiven.segment_copy_bytes_rate` | Rate of bytes copied to remote storage (Aiven RSM) | | `kafka.tiered_storage.aiven.segment_copy_time_avg` | Average time to copy a segment to remote storage (ms) | | `kafka.tiered_storage.aiven.segment_copy_time_max` | Maximum time to copy a segment to remote storage (ms) | | `kafka.tiered_storage.aiven.segment_delete_bytes_rate` | Rate of bytes deleted from remote storage (Aiven RSM) | | `kafka.tiered_storage.aiven.segment_delete_time_avg` | Average time to delete a segment from remote storage (ms) | | `kafka.tiered_storage.aiven.segment_delete_time_max` | Maximum time to delete a segment from remote storage (ms) | ## Configurable custom metrics[​](#configurable-custom-metrics "Direct link to Configurable custom metrics") The following per-partition log metrics are not collected by default. You can enable them by configuring the Datadog integration. These metrics are tagged with `topic` and `partition`, enabling independent monitoring of each topic and partition: * `kafka.log.log_size` * `kafka.log.log_start_offset` * `kafka.log.log_end_offset` ## Variables[​](#variables "Direct link to Variables") Replace the following placeholders in the code samples: | Variable | Description | | ---------------- | ------------------------------------------------------------------------- | | `SERVICE_NAME` | Aiven for Apache Kafka® service name | | `INTEGRATION_ID` | ID of the integration between Aiven for Apache Kafka® service and Datadog | To find the `INTEGRATION_ID` parameter, run: ``` avn service integration-list SERVICE_NAME ``` ## Customize metrics for Datadog[​](#customize-metrics-for-datadog "Direct link to Customize metrics for Datadog") Before customizing metrics, configure and enable a Datadog endpoint in your Aiven for Apache Kafka® service. For setup instructions, see [Send metrics to Datadog](/docs/integrations/datadog/datadog-metrics.md). Format any listed parameters as a comma-separated list: `['value0', 'value1', 'value2', ...]`. To customize Datadog metrics, use the `service integration-update` command with the `kafka_custom_metrics` parameter. Specify a comma-separated list of custom metrics, such as `kafka.log.log_size`, `kafka.log.log_start_offset`, and `kafka.log.log_end_offset`. For example, to send the `kafka.log.log_size` and `kafka.log.log_end_offset` metrics, run: ``` avn service integration-update \ -c 'kafka_custom_metrics=["kafka.log.log_size","kafka.log.log_end_offset"]' \ INTEGRATION_ID ``` After updating settings, view the collected metrics in the Datadog Metrics Explorer. ## Customize consumer metrics for Datadog[​](#customize-consumer-metrics-for-datadog "Direct link to Customize consumer metrics for Datadog") [Apache Kafka Consumer Integration](https://docs.datadoghq.com/integrations/kafka/?tab=host#kafka-consumer-integration) collects metrics for message offsets. To customize the metrics sent from this Datadog integration to Datadog, use the `service integration-update` command with the following parameters: * `include_topics`: A comma-separated list of topics to include. note By default, all topics are included. * `exclude_topics`: A comma-separated list of topics to exclude. note To use `exclude_topics`, specify at least one `include_consumer_groups` value. Otherwise, `exclude_topics` does not take effect. * `include_consumer_groups`: A comma-separated list of consumer groups to include. * `exclude_consumer_groups`: A comma-separated list of consumer groups to exclude. For example, to include topics `topic1` and `topic2`, run: ``` avn service integration-update \ -c 'kafka_custom_metrics=["kafka.log.log_size","kafka.log.log_end_offset"]' \ -c 'include_topics=["topic1","topic2"]' \ INTEGRATION_ID ``` After updating settings, view the collected metrics in the Datadog Metrics Explorer. Related pages * [Datadog and Aiven](/docs/integrations/datadog.md) * [Add client-side Apache Kafka® producer and consumer Datadog metrics](/docs/products/kafka/howto/add-missing-producer-consumer-metrics.md) * [Aiven for Apache Kafka® metrics available via Prometheus](/docs/products/kafka/reference/kafka-metrics-prometheus.md) --- # Scale disk storage automatically for your Aiven for Apache Kafka® service Automatically increase the disk storage of your Aiven for Apache Kafka® service when it's running out of space, instead of resizing it manually. Use the Aiven Autoscaler to automatically increase the storage capacity of a service disk when it's running out of space. Disk autoscaler only increases storage, it doesn't scale storage down. ## Why use disk autoscaling[​](#why-use-disk-autoscaling "Direct link to Why use disk autoscaling") * **Cost efficiency**: Start with a regular-sized disk and let Aiven scale it up only when needed, without the risk of running out of disk space. * **Resiliency**: Avoid a service becoming non-functional because it ran out of disk space, including during unexpected spikes in demand. ## How it works[​](#how-it-works "Direct link to How it works") 1. You create an autoscaler integration endpoint in your project, setting the maximum total disk size to allow. 2. You enable an autoscaler integration for your service using that endpoint. 3. Aiven monitors the disk space usage of your service. 4. When disk usage reaches the threshold for your service type, Aiven increases the available storage by at least 10%, using the current used space as a baseline. note The exact increase depends on the service type and cloud provider. Some providers enforce a minimum increase of 10 GB. Autoscale thresholds per service type The threshold that triggers disk autoscaling is a percentage of the available disk storage capacity: * Aiven for OpenSearch®: 75% of the available disk storage capacity * All other supported service types: 85% of the available disk storage capacity 5. The disk increase is recorded in the project event log, and you receive a notification about the added disk space. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * **Maximum storage**: The maximum storage that the autoscaler can allocate for your service is limited by both the maximum disk size set on the autoscaler endpoint and the maximum disk storage supported for your service plan. * **Timing**: Autoscaling takes a moment to complete. In the meantime, the service disk might fill up and the service might enter read-only mode until autoscaling finishes, unless the autoscaler's disk capacity limit is reached. * **Maintenance updates**: Autoscaling works only on fully running services and can't happen during a maintenance update. * **Manual changes**: Changing disk space manually can delay an autoscaling event. * **Terraform**: Don't manage disk space with the Aiven Terraform Provider on a service that uses the autoscaler, to avoid conflicts between the two. * **Performance**: Disk added through autoscaling is slower than the original disk until the next maintenance update applies. This might affect I/O-intensive workloads. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven organization, project, and service that's up and running * The operator role for the organization, project, and service * Dynamic disk sizing support on your service plan and cloud region * One of the following to manage the autoscaler: * [Aiven Console](https://console.aiven.io/) * [Aiven API](https://api.aiven.io/doc/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) ### Enable disk autoscaling[​](#enable-disk-autoscaling "Direct link to Enable disk autoscaling") To enable disk autoscaling, create an autoscaler integration endpoint, then enable an autoscaler integration on your service using that endpoint. * Console * API * CLI * Terraform Create an autoscaler endpoint: 1. Log in to [Aiven Console](https://console.aiven.io/), and go to your organization and project. 2. On the left sidebar, click **Integration endpoints**. 3. Click **Aiven Autoscaler** > **Add new endpoint**. 4. Set the endpoint name and the maximum total disk storage in GB, and click **Add endpoint**. Enable the autoscaler on a service: 1. On the left sidebar, click **Services**, and open your service. 2. On the left sidebar, click **Integrations**. 3. In **Endpoint integrations**, click **Aiven Autoscaler**. 4. Click the endpoint you created, and click **Enable**. 1) Call [ServiceIntegrationEndpointCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointCreate) to create an autoscaler integration endpoint on your project: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "endpoint_name": "ENDPOINT_NAME", "endpoint_type": "autoscaler", "user_config": { "autoscaling": [ { "type": "autoscale_disk", "cap_gb": 300 } ] } }' ``` 2) Call [ServiceIntegrationCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationCreate) to enable the autoscaler integration on your service, using the endpoint ID from the previous response: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "dest_endpoint_id": "ENDPOINT_ID", "integration_type": "autoscaler", "source_project": "PROJECT_NAME", "source_service": "SERVICE_NAME" }' ``` 1. Create an autoscaler integration endpoint using [avn service integration-endpoint-create](/docs/tools/cli.md): ``` avn service integration-endpoint-create \ --project PROJECT_NAME \ --endpoint-name ENDPOINT_NAME \ --endpoint-type autoscaler \ --user-config-json '{"autoscaling": [{"type": "autoscale_disk", "cap_gb": 300}]}' ``` 2. Find the ID of the new endpoint: ``` avn service integration-endpoint-list --project PROJECT_NAME ``` 3. Enable the autoscaler integration on your service, using the endpoint ID from the previous step: ``` avn service integration-create \ --dest-service SERVICE_NAME \ --integration-type autoscaler \ --source-endpoint-id ENDPOINT_ID ``` Use the [`aiven_service_integration_endpoint`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration_endpoint) and [`aiven_service_integration`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration) resources: ``` resource "aiven_service_integration_endpoint" "autoscaler_endpoint" { project = "PROJECT_NAME" endpoint_name = "ENDPOINT_NAME" endpoint_type = "autoscaler" autoscaler_user_config { autoscaling { type = "autoscale_disk" cap_gb = 300 } } } resource "aiven_service_integration" "autoscaler_integration" { project = "PROJECT_NAME" integration_type = "autoscaler" source_service_name = "SERVICE_NAME" destination_endpoint_id = aiven_service_integration_endpoint.autoscaler_endpoint.id } ``` See the [disk autoscaler guide](https://registry.terraform.io/providers/aiven/aiven/latest/docs/guides/disk-autoscaler) for more details. ### Change the maximum disk space for autoscaling[​](#change-the-maximum-disk-space-for-autoscaling "Direct link to Change the maximum disk space for autoscaling") After you enable disk autoscaling, you can update the maximum total disk size at any time. * Console * API * CLI * Terraform 1. On the left sidebar, click **Integration endpoints**. 2. Click **Aiven Autoscaler**. 3. Find your endpoint, click **Actions**, and click the option to edit it. 4. Set a new maximum disk storage value, and save your changes. Call [ServiceIntegrationEndpointUpdate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointUpdate) with the new `cap_gb` value: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint/ENDPOINT_ID \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "user_config": { "autoscaling": [ { "type": "autoscale_disk", "cap_gb": 500 } ] } }' ``` ``` avn service integration-endpoint-update ENDPOINT_ID \ --user-config-json '{"autoscaling": [{"type": "autoscale_disk", "cap_gb": 500}]}' ``` Update the `cap_gb` value in the `autoscaling` block of your [`aiven_service_integration_endpoint`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration_endpoint) resource, then apply the change. ### Turn off disk autoscaling[​](#turn-off-disk-autoscaling "Direct link to Turn off disk autoscaling") To turn off disk autoscaling, remove the autoscaler integration from your service. You can also delete the integration endpoint if you no longer need it. * Console * API * CLI Disconnect the service from the autoscaler: 1. On the left sidebar, click **Services**, and open your service. 2. On the left sidebar, click **Integrations**. 3. In **Endpoint integrations**, find **Aiven Autoscaler**, click **Actions**, and click the option to disconnect it. Delete the autoscaler endpoint, if you no longer need it: 1. On the left sidebar, click **Integration endpoints**. 2. Click **Aiven Autoscaler**. 3. Find your endpoint, and delete it. 1) Call [ServiceIntegrationDelete](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationDelete) to remove the autoscaler integration from your service: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration/INTEGRATION_ID \ --header "Authorization: Bearer TOKEN" ``` 2) Call [ServiceIntegrationEndpointDelete](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointDelete) to delete the autoscaler integration endpoint, if you no longer need it: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint/ENDPOINT_ID \ --header "Authorization: Bearer TOKEN" ``` 1. Find the ID of the integration to remove: ``` avn service integration-list SERVICE_NAME ``` 2. Remove the autoscaler integration from your service: ``` avn service integration-delete INTEGRATION_ID ``` 3. Find the ID of the integration endpoint to delete, if you no longer need it: ``` avn service integration-endpoint-list --project PROJECT_NAME ``` 4. Delete the autoscaler integration endpoint: ``` avn service integration-endpoint-delete ENDPOINT_ID ``` Related pages * [Scale disk storage manually](/docs/products/kafka/howto/scale-disk-storage.md) * [Change the service plan](/docs/products/kafka/howto/change-service-plan.md) * [Prevent full disks](/docs/products/kafka/howto/prevent-full-disks.md) --- # Enable follower fetching in Aiven for Apache Kafka® Enabling follower fetching in Aiven for Apache Kafka® allows your consumers to fetch data from the nearest replica instead of the leader, optimizing data fetching and enhancing performance. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven for Apache Kafka service version 3.6 or later. * [Availability zone (AZ)](#identify-availability-zone) information for your Aiven for Apache Kafka service. * [Aiven CLI](/docs/tools/cli.md) client. * [Aiven Provider for Terraform](https://registry.terraform.io/providers/aiven/aiven/latest/docs). note Follower fetching is supported on AWS (Amazon Web Services) and Google Cloud. ## Identify availability zone[​](#identify-availability-zone "Direct link to Identify availability zone") Before configuring client-side rack awareness, identify the AZs where your Kafka brokers run. * **AWS**: Availability zone (AZ) names can vary across different accounts. The same physical location might have different AZ names in different accounts. To ensure consistency when configuring `client.rack`, use the AZ ID, which remains the same across accounts. To find the AZ ID used by your brokers, check the broker rack information in your client logs or service connection details. To map AZ names to AZ IDs, see the [AWS Knowledge Center article](https://repost.aws/knowledge-center/vpc-map-cross-account-availability-zones) and the [AWS documentation on AZ IDs](https://docs.aws.amazon.com/ram/latest/userguide/working-with-az-ids). * **Google Cloud**: Use the AZ name directly as the `client.rack` value. For more information, see [Google Cloud's regions and zones documentation](https://cloud.google.com/compute/docs/regions-zones/). ## Enable follower fetching[​](#enable-follower-fetching "Direct link to Enable follower fetching") Use one of the following methods to enable follower fetching on your Aiven for Apache Kafka service. * Console * CLI * API * Terraform 1. Access the [Aiven Console](https://console.aiven.io), and select your Aiven for Apache Kafka service. 2. Click **Service settings**. 3. Scroll to **Advanced configuration** and click **Configure**. 4. Click **Add configuration options**. 5. Select `follower_fetching.enabled` from the list and set the value to **Enabled**. 6. Click **Save configurations**. Enabling follower fetching at the service level allows Kafka clients and Aiven-managed services to use rack-aware fetching. You must still configure `client.rack` on Kafka consumers. Aiven for Apache Kafka® Connect and Aiven for Apache Kafka® MirrorMaker 2 configure this automatically based on the availability zone where each node runs. Enable follower fetching on an existing service with the Aiven CLI: ``` avn service update -c follower_fetching.enabled=true ``` Parameters: * ``: Name of your Aiven for Apache Kafka service. * `follower_fetching.enabled=true`: Enables the follower fetching feature. Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) API to enable follower fetching on an existing service: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/YOUR_PROJECT_NAME/service/YOUR_SERVICE_NAME \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "follower_fetching": { "enabled": true } } }' ``` Parameters: * `YOUR_PROJECT_NAME`: Name of your project. * `YOUR_SERVICE_NAME`: Name of your service. * `YOUR_BEARER_TOKEN`: API token for authentication. * `follower_fetching={"enabled": true}`: Enables the follower fetching feature. Use [the `follower_fetching` attribute](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/kafka#nested-schema-for-kafka_user_configfollower_fetching) in your `aiven_kafka` resource. ## Client-side configuration[​](#client-side-configuration "Direct link to Client-side configuration") To enable follower fetching at the client level, configure the `client.rack` setting in the Apache Kafka client. Set the `client.rack` value to the corresponding AZ ID for AWS or AZ name for Google Cloud for each client. This ensures the client fetches data from the nearest replica. Example configuration for your consumer properties file: ``` client.rack=use1-az1 # AWS example client.rack=europe-west1-b # Google Cloud example ``` ### Example scenario: follower fetching in different AZs[​](#example-scenario-follower-fetching-in-different-azs "Direct link to Example scenario: follower fetching in different AZs") Assume you have an Aiven for Apache Kafka cluster running in two AZs in the `us-east-1` region for AWS and in the `europe-west1` region for Google Cloud: #### Cluster setup and consumer distribution[​](#cluster-setup-and-consumer-distribution "Direct link to Cluster setup and consumer distribution") | Cloud | Region | AZs for brokers | AZs for consumers | | ------------ | -------------- | ---------------------------------- | ---------------------------------------------------- | | AWS | `us-east-1` | `use1-az1`, `use1-az2` | `use1-az1`, `use1-az2`, `use1-az3` | | Google Cloud | `europe-west1` | `europe-west1-b`, `europe-west1-c` | `europe-west1-b`, `europe-west1-c`, `europe-west1-d` | #### Consumer configuration[​](#consumer-configuration "Direct link to Consumer configuration") Set the `client.rack` value to the respective AZ ID for AWS or AZ name for Google Cloud for each consumer: ``` # AWS consumers in use1-az1 client.rack=use1-az1 # AWS consumers in use1-az2 client.rack=use1-az2 # AWS consumers in use1-az3 client.rack=use1-az3 # Google Cloud consumers in europe-west1-b client.rack=europe-west1-b # Google Cloud consumers in europe-west1-c client.rack=europe-west1-c # Google Cloud consumers in europe-west1-d client.rack=europe-west1-d ``` #### Fetching behavior[​](#fetching-behavior "Direct link to Fetching behavior") | Cloud | Consumer location | Fetching behavior | Notes | | ------------ | ----------------- | ------------------------------------------------- | --------------------------------- | | AWS | `use1-az1` | Fetch from the nearest replica in their AZ | Reduced latency and network costs | | AWS | `use1-az2` | Fetch from the nearest replica in their AZ | Reduced latency and network costs | | AWS | `use1-az3` | Fetch from the leader (no matching `broker.rack`) | No follower fetching possible | | Google Cloud | `europe-west1-b` | Fetch from the nearest replica in their AZ | Reduced latency and network costs | | Google Cloud | `europe-west1-c` | Fetch from the nearest replica in their AZ | Reduced latency and network costs | | Google Cloud | `europe-west1-d` | Fetch from the leader (no matching `broker.rack`) | No follower fetching possible | ## Use follower fetching with Kafka Connect and MirrorMaker 2[​](#use-follower-fetching-with-kafka-connect-and-mirrormaker-2 "Direct link to Use follower fetching with Kafka Connect and MirrorMaker 2") Aiven for Apache Kafka® Connect and Aiven for Apache Kafka® MirrorMaker 2 use follower fetching when it is enabled on your Aiven for Kafka service. ### Kafka Connect[​](#kafka-connect "Direct link to Kafka Connect") When follower fetching is enabled on the Aiven for Apache Kafka® service, rack-aware fetching is enabled by default for Kafka Connect sink connectors. Kafka Connect sets `consumer.client.rack` based on each node’s availability zone. Sink connectors use this value when consuming data from Kafka. Source connectors do not use follower fetching. To disable rack awareness for a specific sink connector, set: ``` { "consumer.override.client.rack": "noop" } ``` ### MirrorMaker 2[​](#mirrormaker-2 "Direct link to MirrorMaker 2") When follower fetching is enabled for a replication flow, MirrorMaker 2 reads from in-sync follower replicas in the same availability zone as the MirrorMaker 2 node. Follower fetching is enabled by default for new replication flows. You can disable it per replication flow if needed. For details about how rack awareness works in MirrorMaker 2, see [Configure rack awareness in MirrorMaker 2](/docs/products/kafka/kafka-mirrormaker/howto/mm2-rack-awareness.md). ## Verify follower fetching[​](#verify-follower-fetching "Direct link to Verify follower fetching") After configuring follower fetching, monitor for a decrease in cross-availability zone network costs to verify its effectiveness. Related pages * [Follower fetching in Aiven for Apache Kafka®](/docs/products/kafka/concepts/follower-fetching.md) --- # Enable governance for Aiven for Apache Kafka® [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Enable governance in Aiven for Apache Kafka® to create a secure and compliant framework to manage your Aiven for Apache Kafka services efficiently. By default, governance applies at the organization level. Users with either the [super admin](/docs/platform/howto/manage-permissions.md#make-users-super-admin) or [organization admin](/docs/platform/concepts/permissions.md#organization-roles) role can manually select which services to govern through the Apache Kafka governance settings. ## Impact of enabling governance[​](#impact-of-enabling-governance "Direct link to Impact of enabling governance") * **Existing topics**: * The default group owns all existing Apache Kafka resources. * Ownership details for Apache Kafka resources are visible in the [Apache Kafka topic catalog](/docs/products/kafka/concepts/topic-catalog-overview.md). * Users in different groups can request ownership of individual resources. * **Topic creation workflow**: * Existing topics remain unaffected. * You can continue to [create topics](/docs/products/kafka/howto/create-topic.md) in your Aiven for Apache Kafka service. Governance adds a request-and-approval process for claiming topic ownership through the Apache Kafka topic catalog. * All topics align with your organization's data management policies. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * This is a [limited availability feature](/docs/platform/concepts/service-and-feature-releases.md). To try it out, contact the [sales team](https://aiven.io/contact). * Ensure you have either [super admin role](/docs/platform/howto/manage-permissions.md#make-users-super-admin) or [organization admin role](/docs/platform/concepts/permissions.md#organization-roles) to enable governance. ## Enable governance[​](#enable-governance "Direct link to Enable governance") 1. Access the [Aiven Console](https://console.aiven.io/) and click **Admin**. 2. On the organization page, click **Apache Kafka governance**. 3. Click **Enable governance**. 4. Select a user group to manage topics. This group is the default group and owns all topics in the organization. * If no user group exists, click **Create new user group**. 1. Enter the group name and description. 2. Click **Create group**. 3. Select the created user group. 5. Click **Next**. 6. Set global default topic configurations: * To use the default settings, click **Keep defaults** and review the default values on the confirmation window. * Click **Enable governance**. To customize the settings: * Enter your global default topic configurations. * Click **Enable governance** to apply your changes. note Global topic configurations apply only to new topics created after the policies are updated. Existing topics are not affected by these changes. ## Select services for governance[​](#select-services-for-governance "Direct link to Select services for governance") Choose which Aiven for Apache Kafka services to include in governance. Topics from these services will appear in the Topic Catalog, allowing users to claim ownership of existing topics. 1. On the **Administration** page, click **Apache Kafka governance**. 2. In the **Governed services** section, click **Change**. 3. In the **Select services for governance** dialog, select the services. 4. Click **Save**. ## Select the governance method[​](#select-the-governance-method "Direct link to Select the governance method") You can manage governance for Aiven for Apache Kafka® service using either the Aiven Console or the Aiven Terraform Provider. Depending on your workflow, choose between: * **Aiven Console**: Manage governance tasks visually through the Aiven Console. View and claim ownership of resources in the Topic Catalog. This method is selected by default. * **Aiven Terraform Provider**: Automate governance with the Terraform Provider. It integrates with GitOps workflows and allows governance management across multiple projects. note When using the Terraform method, all governance actions must be performed through Aiven Terraform Provider. The Aiven Console provides an audit log of created requests. 1. On the **Administration** page, click **Apache Kafka governance**. 2. In the **Governance method** section, click **Change**. 3. In the **Select governance method** dialog, choose either **Aiven Console** or **Terraform**. warning Switching from the Aiven Console to the Terraform method results in the loss of all pending requests. 4. Click **Save**. ## Change the default user group[​](#change-the-default-user-group "Direct link to Change the default user group") To change the default user group after enabling governance: 1. On the **Administration** page, click **Apache Kafka governance**. 2. Click **Change** next to **Default user group**. 3. Select a new user group from the list. 4. If no suitable group exists, [create a group](/docs/platform/howto/manage-groups.md#create-a-group) and select it. 5. Click **Save**. ## Update global topic configurations[​](#update-global-topic-configurations "Direct link to Update global topic configurations") To change global topic configurations after enabling governance: 1. On the **Administration** page, click **Apache Kafka governance**. 2. Click **Change** next to **Global default topic configurations**. 3. Update the global default topic configurations as needed. 4. Modify retention policies, partition strategies, or other settings. 5. Click **Save**. ## Disable governance[​](#disable-governance "Direct link to Disable governance") 1. On the **Administration** page, click **Apache Kafka governance**. 2. Expand **Governance is enabled for your organization**. 3. Click **Disable**. ### Impact of disabling governance[​](#impact-of-disabling-governance "Direct link to Impact of disabling governance") * Existing ownership assigned to the Apache Kafka resources remains unchanged. * Re-enabling governance later preserves the Apache Kafka resources ownership from the last time it was disabled. * Apache Kafka resources claimed by specific groups retain their ownership. * If you select a different user group when re-enabling governance, Apache Kafka resources under the previous default group are assigned to the new default governance group. Related pages * [Aiven for Apache Kafka® governance overview](/docs/products/kafka/concepts/governance-overview.md) * [Project member roles and permissions](/docs/platform/concepts/permissions.md) --- # Enable tiered storage for Aiven for Apache Kafka® Tiered storage significantly improves the storage efficiency of your Aiven for Apache Kafka® service. You can enable tiered storage for topics in a Classic Kafka service. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to an Aiven organization and at least one project. * Aiven for Apache Kafka® service with Apache Kafka version 3.6 or later. Upgrade to the latest default version and apply [maintenance updates](/docs/products/kafka/howto/maintenance-updates.md) for the latest fixes and improvements when using tiered storage. * [Aiven CLI](/docs/tools/cli.md). note Review the [trade-offs and limitations](/docs/products/kafka/concepts/tiered-storage-limitations.md) of tiered storage before enabling it. ## Steps to enable tiered storage[​](#steps-to-enable-tiered-storage "Direct link to Steps to enable tiered storage") * Console * CLI 1. In the [Aiven Console](https://console.aiven.io/), select your project. 2. Create an Aiven for Apache Kafka service or select an existing one. * For [a new service](/docs/products/kafka/get-started/create-kafka-service.md): 1. On the **Create Apache Kafka® service** page, scroll down to the **Tiered storage** section. 2. Click **Enable tiered storage**. 3. View the pricing for tiered storage in the **Service summary**. * For an existing service: 1. Go to the service's **Overview** page and click **Service settings** from the sidebar. 2. In the **Service plan** section, click **Enable tiered storage** to activate it. 3. Click **Activate tiered storage** in the confirmation window. When tiered storage is activated and configured for your [topics](/docs/products/kafka/howto/configure-topic-tiered-storage.md), monitor usage and costs in [Storage usage and settings in Aiven Console](/docs/products/kafka/howto/view-kafka-storage-in-console.md) section. note Alternatively, if tiered storage is not yet active for your service, you can enable it by selecting **Observe** > **Tiered storage** from the sidebar. Enable tiered storage for your Aiven for Apache Kafka service using the [Aiven CLI](/docs/tools/cli.md): 1. Retrieve the project information with: ``` avn project details ``` For a specific project, use: ``` avn project details --project ``` 2. Find the name of the Aiven for Apache Kafka service for enabling tiered storage with: ``` avn service list ``` Make a note of the `SERVICE_NAME` corresponding to your Aiven for Apache Kafka service. 3. Enable tiered storage with: ``` avn service update \ --project demo-kafka-project \ demo-kafka-service \ -c tiered_storage.enabled=true ``` In this command: * `--project demo-kafka-project`: Replace `demo-kafka-project` with your project name. * `demo-kafka-service`: Specify the Aiven for Apache Kafka service you intend to update. * `-c tiered_storage.enabled=true`: Activates tiered storage. ## Configure default retention policies at service-level[​](#configure-default-retention-policies-at-service-level "Direct link to Configure default retention policies at service-level") To manage data retention, set default policies for tiered storage at the service level. 1. In the [Aiven Console](https://console.aiven.io/), select your project and your Aiven for Apache Kafka service. 2. Click **Service settings** from the sidebar. 3. Scroll to the **Advanced configuration** section, and click **Configure**. 4. In the **Advanced configuration** dialog, click **Add Advanced Configuration**. 5. Define the retention policy: * Find `kafka.log_local_retention_ms` and set the value to define the retention period in milliseconds for time-based retention. * Find `kafka.log_local_retention_bytes` and set the value to define the retention limit in bytes for size-based retention. 6. Click **Save configuration** to apply your changes. You can configure retention policies from the [Storage usage and settings in Aiven Console](/docs/products/kafka/howto/view-kafka-storage-in-console.md#modify-retention-polices) page. ## Optional: Configure client-side parameter[​](#optional-configure-client-side-parameter "Direct link to Optional: Configure client-side parameter") For optimal performance and reduced risk of broker interruptions when using tiered storage, update the client-side parameter `fetch.max.wait.ms` from its default value of 500 ms to 5000 ms. This consumer configuration is no longer necessary starting from Apache Kafka version 3.6.2. Consider upgrading to Apache Kafka version 3.6.2 or later before enabling tiered storage. Related pages * [Tiered storage in Aiven for Apache Kafka® overview](/docs/products/kafka/concepts/kafka-tiered-storage.md) * [How tiered storage works in Aiven for Apache Kafka®](/docs/products/kafka/concepts/tiered-storage-how-it-works.md) * [Enable and configure tiered storage for topics](/docs/products/kafka/howto/configure-topic-tiered-storage.md) --- # Enable OAuth 2.0/OIDC authentication for Apache Kafka® Aiven for Apache Kafka® supports OAuth 2.0/OIDC authentication for Kafka clients. Use OAuth 2.0/OIDC authentication to let clients authenticate with tokens issued by an identity provider or by an identity broker that supports Outbound Identity Federation. For AWS IAM, see [OAuth 2.0/OIDC with AWS IAM](/docs/products/kafka/howto/kafka-oauth2-aws-iam.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before you begin, make sure you have: * An [Aiven for Apache Kafka](/docs/products/kafka/get-started/create-kafka-service.md) service. * [SASL authentication](/docs/products/kafka/howto/kafka-sasl-auth.md#enable-sasl-authentication) enabled on the service. OAuth 2.0/OIDC uses the `OAUTHBEARER` SASL mechanism. * Access to an OIDC provider, such as Auth0, Okta, Google Identity Platform, Azure, or another OIDC-compliant provider. * Configuration details from your OIDC provider: * **JWKS endpoint URL:** Required. HTTPS URL to retrieve the JSON Web Key Set, or JWKS. * **Issuer URL or identifier:** Required by most OIDC providers. Identifies and verifies the JWT issuer. * **Audience identifiers:** Required by most OIDC providers. Validates the JWT's intended recipients. For multiple audiences, note each value. * **Subject claim name:** Optional. Typically `sub`, but this can vary depending on your OIDC provider. Configuration steps vary by identity provider. See your provider's documentation for JWKS URL, issuer, and audience values. note OAuth 2.0/OIDC authentication works independently of Aiven service users. If you use **Aiven ACLs** instead of Kafka-native ACLs to control access, create a [matching Aiven service user](/docs/products/kafka/howto/add-manage-service-users.md) with the same identifier as the OIDC principal. This identifier can be the identity provider's service principal ID. Aiven enforces the ACL only when this matching service user exists. Kafka-native ACLs do not have this requirement. For more information, see [Manage access control lists](/docs/products/kafka/howto/manage-acls.md). ## Configure OAuth 2.0/OIDC settings[​](#configure-oauth-20oidc-settings "Direct link to Configure OAuth 2.0/OIDC settings") Set `kafka.sasl_oauthbearer_jwks_endpoint_url` to enable `OAUTHBEARER`. To use only OAuth 2.0/OIDC authentication, enable SASL authentication, set `kafka.sasl_oauthbearer_jwks_endpoint_url`, and [disable PLAIN, SCRAM-SHA-256, and SCRAM-SHA-512](/docs/products/kafka/howto/kafka-sasl-auth.md#configure-sasl-mechanisms). note When SASL authentication is enabled, at least one SASL mechanism must be available. `OAUTHBEARER` satisfies this requirement when `kafka.sasl_oauthbearer_jwks_endpoint_url` is set. Configure OAuth 2.0/OIDC authentication using one of the following methods. * Aiven Console * CLI 1. In the [Aiven Console](https://console.aiven.io/), select your project and choose your Aiven for Apache Kafka service. 2. Click **Service settings**. 3. Scroll to **Advanced configuration** and click **Configure**. 4. In the **Advanced configuration** window, click **Add configuration options**. 5. Enable SASL authentication by setting `kafka_authentication_methods.sasl` to **Enabled**. 6. Configure the JWKS endpoint by setting `kafka.sasl_oauthbearer_jwks_endpoint_url` to your provider's JWKS URL. This enables the `OAUTHBEARER` mechanism. `PLAIN`, `SCRAM-SHA-256`, and `SCRAM-SHA-512` remain enabled by default. 7. Optional: Configure other OIDC parameters, such as expected issuer, expected audience, and subject claim. See [OIDC parameters](#oidc-parameters) for details. 8. Optional: To use only OAuth 2.0/OIDC authentication, set `kafka_sasl_mechanisms.plain`, `kafka_sasl_mechanisms.scram_sha_256`, and `kafka_sasl_mechanisms.scram_sha_512` to **Disabled**. 9. Click **Save configurations**. To configure OAuth 2.0/OIDC authentication for your Aiven for Apache Kafka service using the [Aiven CLI](/docs/tools/cli.md): Each `avn service update` that changes OIDC or SASL settings triggers a rolling restart of Apache Kafka brokers. Combine the `-c` flags you need in a single command when applying multiple changes. 1. Get the name of your Aiven for Apache Kafka service: ``` avn service list ``` Note the `SERVICE_NAME` corresponding to your Aiven for Apache Kafka service. 2. Enable SASL authentication and configure the JWKS endpoint: Run a single `avn service update` command. Include the required flags below, and add any optional flags to the same command. **Required:** ``` avn service update SERVICE_NAME \ -c kafka_authentication_methods.sasl=true \ -c kafka.sasl_oauthbearer_jwks_endpoint_url="https://my-jwks-endpoint.example.com/jwks" ``` This enables the `OAUTHBEARER` mechanism. `PLAIN`, `SCRAM-SHA-256`, and `SCRAM-SHA-512` remain enabled by default. **Optional:** Add issuer, audience, and subject claim verification. To use only OAuth 2.0/OIDC authentication, set `kafka_sasl_mechanisms.plain`, `kafka_sasl_mechanisms.scram_sha_256`, and `kafka_sasl_mechanisms.scram_sha_512` to `false`. Example with issuer, audience, subject claim verification, and OAuth-only SASL configuration: ``` avn service update SERVICE_NAME \ -c kafka_authentication_methods.sasl=true \ -c kafka.sasl_oauthbearer_jwks_endpoint_url="https://my-jwks-endpoint.example.com/jwks" \ -c kafka.sasl_oauthbearer_expected_issuer="https://my-issuer.example.com" \ -c kafka.sasl_oauthbearer_expected_audience="my-audience" \ -c kafka.sasl_oauthbearer_sub_claim_name="sub" \ -c kafka_sasl_mechanisms.plain=false \ -c kafka_sasl_mechanisms.scram_sha_256=false \ -c kafka_sasl_mechanisms.scram_sha_512=false ``` Omit optional flags you do not need. Do not run the required and optional examples as separate commands. Replace the following: * `SERVICE_NAME`: name of your Aiven for Apache Kafka service. For details about the OIDC parameters, see [OIDC parameters](#oidc-parameters). ## OIDC parameters[​](#oidc-parameters "Direct link to OIDC parameters") Configure the following OIDC parameters: * `kafka.sasl_oauthbearer_jwks_endpoint_url` * **Description**: Endpoint for retrieving the JSON Web Key Set, or JWKS, which enables OIDC authentication. Corresponds to the Apache Kafka parameter `sasl.oauthbearer.jwks.endpoint.url`. * **Value**: Enter the HTTPS JWKS endpoint URL provided by your OIDC provider. note Starting with Apache Kafka 4.0, the broker verifies that the JWKS endpoint URL for OAuth authentication matches an entry in the system property `org.apache.kafka.sasl.oauthbearer.allowed.urls`. Aiven sets this property from the value of `kafka.sasl_oauthbearer_jwks_endpoint_url`. You do not need additional configuration. * `kafka.sasl_oauthbearer_sub_claim_name` * **Optional** * **Description**: Name of the JWT's subject claim for broker verification. It is typically set to `sub`. Corresponds to the Apache Kafka parameter `sasl.oauthbearer.sub.claim.name`. * **Value**: Enter `sub` or the specific claim name provided by your OIDC provider if different. note The claim must be a string. Claims that contain arrays, such as `groups`, are not supported. * `kafka.sasl_oauthbearer_expected_issuer` * **Optional** * **Description**: Specifies the JWT's issuer for the broker to verify. Corresponds to the Apache Kafka parameter `sasl.oauthbearer.expected.issuer`. * **Value**: Enter the issuer URL or identifier provided by your OIDC provider. * `kafka.sasl_oauthbearer_expected_audience` * **Optional** * **Description**: Validates the intended JWT audience for the broker. Corresponds to the Apache Kafka parameter `sasl.oauthbearer.expected.audience`. Use this parameter when your OIDC provider specifies an audience. * **Value**: Enter the audience identifiers given by your OIDC provider. If there are multiple audiences, separate them with commas. For more information about each corresponding Apache Kafka parameter, see [Apache Kafka documentation](https://kafka.apache.org/documentation/) on configuration options starting with `sasl.oauthbearer`. warning Changing OIDC settings triggers a rolling restart of Apache Kafka brokers. As a result, the brokers temporarily operate with different configurations. To reduce operational impact, apply these changes during a maintenance window. Related pages * [Enable OAuth 2.0/OIDC support for Apache Kafka REST proxy](/docs/products/kafka/karapace/howto/enable-oauth-oidc-kafka-rest-proxy.md) * [Enable OAuth 2.0/OIDC authentication for Aiven for Apache Kafka® Schema Registry](/docs/products/kafka/karapace/howto/enable-oauth-oidc-schema-registry.md) * [Enable and configure SASL authentication](/docs/products/kafka/howto/kafka-sasl-auth.md) * [OAuth 2.0/OIDC with AWS IAM](/docs/products/kafka/howto/kafka-oauth2-aws-iam.md) --- # Enable the consumer lag predictor for Aiven for Apache Kafka® [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) The [consumer lag predictor](/docs/products/kafka/concepts/consumer-lag-predictor.md) in Aiven for Apache Kafka® provides visibility into the time between message production and consumption, allowing for improved cluster performance and scalability. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before you start, ensure you have the following: * Aiven account. * [Aiven for Apache Kafka®](/docs/products/kafka/get-started/create-kafka-service.md) service running. * [Prometheus integration](/docs/platform/howto/integrations/prometheus-metrics.md) set up for your Aiven for Apache Kafka for extracting metrics. * Necessary permissions to modify service configurations. * The consumer lag predictor for Aiven for Apache Kafka® is a **limited availability** feature and requires activation on your Aiven account. Contact the [sales team](https://aiven.io/contact) to request activation. ## Enable and configure the consumer lag predictor[​](#enable-and-configure-the-consumer-lag-predictor "Direct link to Enable and configure the consumer lag predictor") * Aiven Console * Aiven CLI 1. Once the consumer lag predictor is activated for your account, log in to the [Aiven Console](https://console.aiven.io/), select your project, and choose your Aiven for Apache Kafka® service. 2. On the **Overview** page, click **Service settings** from the sidebar. 3. Go to the **Advanced configuration** section, and click **Configure**. 4. In the **Advanced configuration** window, click **Add configuration options**. 5. Set `kafka_lag_predictor.enabled` to **Enabled**. This enables the lag predictor to compute predictions for all consumer groups and topics. 6. Configure the following options: * **Set `kafka_lag_predictor.group_filters`**: Specify the consumer group pattern to include only the desired consumer groups in the lag prediction. By default, the consumer lag predictor calculates the lag for all consumer groups, but you can restrict this by specifying group patterns. Example group patterns: * `consumer_group_*`: Matches any consumer group that starts with `consumer_group_`, such as `consumer_group_1` or `consumer_group_a`. * `important_group`: Matches exactly the consumer group named `important_group`. * `group?-test`: Matches consumer groups like `group1-test` or `groupA-test`, where the `?` represents any single character. * **Set `kafka_lag_predictor.topics`**: Specify which topics to include in the lag prediction. By default, predictions are computed for all topics, but you can restrict this by using topic names or patterns. Example topic patterns: * `important_topic_*`: Matches any topic that starts with `important_topic_`, such as `important_topic_1`, `important_topic_data`. * `secondary_topic`: Matches exactly the topic named `secondary_topic`. * `topic?-logs`: Matches topics like `topic1-logs` or `topicA-logs`, where the `?` represents any single character. 7. Click **Save configuration** to save your changes and enable consumer lag prediction. To enable the consumer lag predictor for your Aiven for Apache Kafka service using [Aiven CLI](/docs/tools/cli.md): 1. Ensure the consumer lag predictor feature is activated for your account by contacting the [sales team](https://aiven.io/contact). The consumer lag predictor is a limited availability feature and needs to be activated for your account. 2. Get the project information: ``` avn project details ``` If you need details for a specific project, use: ``` avn project details --project PROJECT_NAME ``` 3. Get the name of the Aiven for Apache Kafka service: ``` avn service list ``` Make a note of the `SERVICE_NAME` corresponding to your Aiven for Apache Kafka service. 4. Enable the consumer lag predictor for your service: ``` avn service update SERVICE_NAME -c kafka_lag_predictor.enabled=true ``` Replace `SERVICE_NAME` with your service name. note This enables the lag predictor to compute predictions for all consumer groups across all topics. 5. Configure the consumer groups and topics to be included in the lag prediction: * **For consumer groups**: Set the `kafka_lag_predictor.group_filters` option to specify which consumer groups should be included in the lag prediction. By default, the consumer lag predictor calculates the lag for all consumer groups, but you can restrict this by specifying group patterns. ``` avn service update SERVICE_NAME \ -c kafka_lag_predictor.group_filters='["example_consumer_group_1", "example_consumer_group_2"]' ``` * Replace `SERVICE_NAME` with the actual name or ID of your Aiven for Apache Kafka® service. * Replace `example_consumer_group_1` and `example_consumer_group_2` with your consumer group names. * **For topics**: Set the `kafka_lag_predictor.topics` option to specify which topics should be included in the lag prediction. By default, predictions are computed for all topics, but you can restrict this by using topic names or patterns. ``` avn service update SERVICE_NAME \ -c kafka_lag_predictor.topics='["important_topic_*", "secondary_topic"]' ``` Replace `important_topic_*` and `secondary_topic` with your topic names or patterns. ## Monitor metrics with Prometheus[​](#monitor-metrics-with-prometheus "Direct link to Monitor metrics with Prometheus") After enabling the consumer lag predictor, you can use [Prometheus](/docs/platform/howto/integrations/prometheus-metrics.md) to access and monitor detailed metrics that provide insights into your Apache Kafka cluster's performance: | Metric | Type | Description | | -------------------------------------------------- | ------- | ------------------------------------------------------------------------------------------------------ | | `kafka_lag_predictor_topic_produced_records_total` | Counter | Represents the total count of records produced. | | `kafka_lag_predictor_group_consumed_records_total` | Counter | Represents the total count of records consumed. | | `kafka_lag_predictor_group_lag_predicted_seconds` | Gauge | Represents the estimated time lag, in seconds, for a consumer group to catch up to the latest message. | For example, you can monitor the average estimated time lag in seconds for a consumer group to consume produced messages using the following PromQL query: ``` avg by(topic,group)(kafka_lag_predictor_group_lag_predicted_seconds_gauge) ``` Another useful metric to monitor is the consume/produce ratio. You can monitor this per topic and partition for consumer groups by using the following PromQL query: ``` sum by(group, topic, partition)( kafka_lag_predictor_group_consumed_records_total_counter ) / on(topic, partition) group_left() sum by(topic, partition)( kafka_lag_predictor_topic_produced_records_total_counter ) ``` --- # Use Apache Flink® with Aiven for Apache Kafka® [Apache Flink®](https://flink.apache.org/) is an open-source platform for handling distributed streaming and batch data. It enhances Apache Kafka's® event streaming abilities by offering advanced features for consuming, transforming, aggregating, and enriching data. note To experience the power of streaming SQL transformations with Flink, Aiven provides a managed [Aiven for Apache Flink®](/docs/products/flink.md) with built-in data flow integration with Aiven for Apache Kafka®. The following example demonstrates how to create a simple Java Flink job. This job reads data from a Apache Kafka topic, processes it,and sends it to another Apache Kafka topic. It uses the Java API on a [local installation of Apache Flink 1.16](https://nightlies.apache.org/flink/flink-docs-release-1.16/docs/try-flink/local_installation/). However, the same approach can be applied to use Aiven for Apache Kafka with any self-hosted cluster. ## Prerequisites[​](#kafka-flink-java-prereq "Direct link to Prerequisites") Before you start, make sure you have the following: * An active **Aiven for Apache Kafka** service with two topics: `test-flink-input` and `test-flink-output`. To create topics, see [Create an Apache Kafka topic](/docs/products/kafka/howto/create-topic.md). * Gather the following details about your Aiven for Apache Kafka service: * `APACHE_KAFKA_HOST`: The hostname of your Apache Kafka service. * `APACHE_KAFKA_PORT`: The port number of your Apache Kafka service. * [**Apache Maven™**](https://maven.apache.org/install.html) installed on your machine build the example. ### Setup the truststore and keystore[​](#setup-the-truststore-and-keystore "Direct link to Setup the truststore and keystore") Create a [Java keystore and truststore](/docs/products/kafka/howto/keystore-truststore.md) for the Aiven for Apache Kafka service. For this example, the configuration is as follows: * The keystore is available at `KEYSTORE_PATH/client.keystore.p12` * The truststore is available at `TRUSTSTORE_PATH/client.truststore.jks` * For simplicity, use the same secret (password) for both the keystore and the truststore, referred to as `KEY_TRUST_SECRET` ## Use Apache Flink with Aiven for Apache Kafka[​](#use-apache-flink-with-aiven-for-apache-kafka "Direct link to Use Apache Flink with Aiven for Apache Kafka") The following example shows how to customise the `DataStreamJob` generated from the [Quickstart](https://nightlies.apache.org/flink/flink-docs-release-1.19/docs/dev/configuration/overview/) to work with Aiven for Apache Kafka. note Find the full code in the [Aiven examples GitHub repository](https://github.com/aiven/aiven-examples/tree/master/kafka/flink-capitalizer). 1. Generate a Flink job skeleton named `flink-capitalizer` using the Maven archetype: ``` mvn archetype:generate -DinteractiveMode=false \ -DarchetypeGroupId=org.apache.flink \ -DarchetypeArtifactId=flink-quickstart-java \ -DarchetypeVersion=1.16.0 \ -DgroupId=io.aiven.example \ -DartifactId=flink-capitalizer \ -Dpackage=io.aiven.example.flinkcapitalizer \ -Dversion=0.0.1-SNAPSHOT ``` 2. Uncomment the Kafka connector in \`pom.xml\`: ``` org.apache.flink flink-connector-kafka ${flink.version} ``` ### Customize the `DataStreamJob` application[​](#customize-the-datastreamjob-application "Direct link to customize-the-datastreamjob-application") In the generated code, `DataStreamJob` is the main entry point, and has already been configured with all of the context necessary to interact with the cluster for your processing. 1. Create a class called `io.aiven.example.flinkcapitalizer.StringCapitalizer` which performs a `MapFunction` transformation on incoming records, emitting every incoming string in uppercase. ``` package io.aiven.example.flinkcapitalizer; import org.apache.flink.api.common.functions.MapFunction; public class StringCapitalizer implements MapFunction { public String map(String s) { return s.toUpperCase(); } } ``` 2. Import the following classes in the `DataStreamJob` ``` import java.util.Properties; import org.apache.flink.api.common.eventtime.WatermarkStrategy; import org.apache.flink.api.common.serialization.SimpleStringSchema; import org.apache.flink.connector.base.DeliveryGuarantee; import org.apache.flink.connector.kafka.sink.KafkaRecordSerializationSchema; import org.apache.flink.connector.kafka.sink.KafkaSink; import org.apache.flink.connector.kafka.source.KafkaSource; import org.apache.flink.connector.kafka.source.enumerator.initializer.OffsetsInitializer; ``` 3. Modify the `main` method in `DataStreamJob` to read and write from the Apache Kafka topics, replacing the `APACHE_KAFKA_HOST`, `APACHE_KAFKA_PORT`, `KEYSTORE_PATH`, `TRUSTSTORE_PATH` and `KEY_TRUST_SECRET` placeholders with the values from the [prerequisites](/docs/products/kafka/howto/flink-with-aiven-for-kafka.md#kafka-flink-java-prereq). ``` public static void main(String[] args) throws Exception { final StreamExecutionEnvironment env = StreamExecutionEnvironment.getExecutionEnvironment(); Properties props = new Properties(); props.put("security.protocol", "SSL"); props.put("ssl.keystore.type", "PKCS12"); props.put("ssl.keystore.location", "KEYSTORE_PATH/client.keystore.p12"); props.put("ssl.keystore.password", "KEY_TRUST_SECRET"); props.put("ssl.key.password", "KEY_TRUST_SECRET"); props.put("ssl.truststore.type", "JKS"); props.put("ssl.truststore.location", "TRUSTSTORE_PATH/client.truststore.jks"); props.put("ssl.truststore.password", "KEY_TRUST_SECRET"); KafkaSource source = KafkaSource.builder() .setBootstrapServers("APACHE_KAFKA_HOST:APACHE_KAFKA_PORT") .setGroupId("test-flink-input-group") .setTopics("test-flink-input") .setProperties(props) .setStartingOffsets(OffsetsInitializer.earliest()) .setValueOnlyDeserializer(new SimpleStringSchema()) .build(); KafkaSink sink = KafkaSink.builder() .setBootstrapServers("APACHE_KAFKA_HOST:APACHE_KAFKA_PORT") .setKafkaProducerConfig(props) .setRecordSerializer(KafkaRecordSerializationSchema.builder() .setTopic("test-flink-output") .setValueSerializationSchema(new SimpleStringSchema()) .build() ) .setDeliverGuarantee(DeliveryGuarantee.AT_LEAST_ONCE) .build(); // ... processing continues here } ``` 4. Tie the Apache Kafka sources and sinks together with the `StringCapitalizer` in a single processing pipeline. ``` // ... processing continues here env .fromSource(source, WatermarkStrategy.noWatermarks(), "Kafka Source") .map(new StringCapitalizer()) .sinkTo(sink); env.execute("Flink Java capitalizer"); ``` ### Build the application[​](#build-the-application "Direct link to Build the application") From the main `flink-capitalizer` folder, execute the following Maven command to build the application: ``` mvn -DskipTests=true clean package ``` The above command should create a `jar` file named `target/flink-capitalizer-0.0.1-SNAPSHOT.jar`. ### Run the applications[​](#run-the-applications "Direct link to Run the applications") If you have installed a [local cluster installation of Apache Flink 1.16](https://nightlies.apache.org/flink/flink-docs-release-1.19/docs/try-flink/local_installation/), you can launch the job on your local machine. `$FLINK_HOME` is the Flink installation directory. ``` $FLINK_HOME/bin/flink run target/flink-capitalizer-0.0.1-SNAPSHOT.jar ``` You can see that the job is running in the Flink web UI at `http://localhost:8081`. By integrating [Aiven for Apache Flink®](/docs/products/flink.md) with Aiven for Apache Kafka®, you can process string events and transform them to uppercase before forwarding them to the output topic. --- # Generate Java classes from Avro schemas Generate Java classes from Avro schema files (`.avsc`) to use with Apache Kafka® producers and consumers. Use the `avro-tools` JAR to create Java classes that match your schema structure. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An active [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md) * [Karapace Schema Registry](/docs/products/kafka/karapace/howto/enable-karapace.md) enabled * Java (JDK 8 or later) * [Apache Maven](https://maven.apache.org/) installed to manage dependencies * [`avro-tools`](https://avro.apache.org/releases.html) JAR file downloaded (such as `avro-tools-1.12.0.jar`) * An Avro schema file with a `.avsc` extension note You can also use `.json` files if they contain a valid Avro schema. Do not use `.avro` files. They include both schema and data and are not compatible with `avro-tools`. ## Generate the Java classes[​](#generate-the-java-classes "Direct link to Generate the Java classes") Run the following command to generate Java classes from your Avro schema: note Use the latest version of `avro-tools` if available. You can check for newer releases at . ``` java -jar avro-tools-1.12.0.jar compile schema src/main/resources/user-value.avsc src/main/java/ ``` * Replace `user-value.avsc` with your schema file. * Replace `src/main/java/` with your preferred output directory. The generated class is named based on the `name` field in your schema, and it is placed in a subdirectory matching the `namespace`. ## Example schema[​](#example-schema "Direct link to Example schema") ``` { "type": "record", "name": "User", "namespace": "io.aiven.example", "fields": [ { "name": "id", "type": "int" }, { "name": "email", "type": "string" } ] } ``` This schema produces the following file: ``` src/main/java/io/aiven/example/User.java ``` ## Add Maven dependencies[​](#add-maven-dependencies "Direct link to Add Maven dependencies") Add these dependencies to your `pom.xml` to compile and use the generated classes for Avro serialization and deserialization in Kafka producers and consumers. ``` org.apache.kafka kafka-clients 3.8.1 io.confluent kafka-avro-serializer 8.0.0 io.confluent kafka-schema-registry-client 8.0.0 io.confluent kafka-schema-serializer 8.0.0 org.apache.avro avro 1.12.0 ``` note Ensure that the versions of Avro, Kafka, and Confluent dependencies are compatible with each other and with your Kafka setup. ### Optional dependencies[​](#optional-dependencies "Direct link to Optional dependencies") You can use additional libraries depending on your schema or Avro usage: ``` com.fasterxml.jackson.datatype jackson-datatype-jdk8 2.19.2 com.google.guava guava 33.4.8-jre ``` These dependencies are optional. Include them only if your generated classes use features like `Optional` fields (Jackson) or `ImmutableList` and `ImmutableSet` types (Guava). Related pages * [Use schema registry with Java producers and consumers](/docs/products/kafka/howto/use-schema-registry-in-java.md) * [Official Avro Java Getting Started Guide](https://avro.apache.org/docs/1.12.0/getting-started-java/) --- # Generate Java classes from JSON Schema Generate Java classes from JSON Schema (`.json`) files for use in Apache Kafka® applications. Use the `jsonschema2pojo` CLI tool to generate Java classes that match your schema structure. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An active [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md) * [Karapace Schema Registry](/docs/products/kafka/karapace/howto/enable-karapace.md) enabled * Java (JDK 8 or later) * [Apache Maven](https://maven.apache.org/) installed to manage dependencies * [`jsonschema2pojo`](https://github.com/joelittlejohn/jsonschema2pojo) installed * A JSON Schema file (`.json`) ## Generate the Java classes[​](#generate-the-java-classes "Direct link to Generate the Java classes") Generate Java classes from your JSON schema file using the `jsonschema2pojo` CLI: ``` jsonschema2pojo --source src/main/resources/schema.json --target src/main/java/ ``` * Replace `schema.json` with your schema file. * Replace `src/main/java/` with your preferred output directory. note Use the `--package` option to specify the Java package during code generation. Kafka serialization depends on the fully qualified class name. For example: ``` jsonschema2pojo \ --source src/main/resources/users.json \ --target src/main/java/io/aiven/example \ --package io.aiven.example ``` ## Example schema[​](#example-schema "Direct link to Example schema") ``` { "type": "object", "title": "User", "javaType": "io.aiven.example.User", "properties": { "id": { "type": "integer" }, "email": { "type": "string" } }, "required": ["id", "email"] } ``` This schema generates the following Java file: ``` src/main/java/io/aiven/example/User.java ``` ## Optional: Add Confluent schema annotation[​](#optional-add-confluent-schema-annotation "Direct link to Optional: Add Confluent schema annotation") To enable additional features in Confluent’s deserializers (which are compatible with schema registries like [Karapace](/docs/products/kafka/karapace.md)), you can add this annotation to your JSON Schema: ``` "@io.confluent.kafka.schemaregistry.annotations.Schema": { "value": "{...your schema as a string...}", "refs": [] } ``` This annotation makes runtime methods such as `schema()` and `refs()` available in the generated Java classes. Use this only for advanced use cases that require schema introspection at runtime. ## Add Maven dependencies[​](#add-maven-dependencies "Direct link to Add Maven dependencies") Add these dependencies to your `pom.xml` to compile and use the generated classes for JSON Schema serialization and deserialization in Kafka producers and consumers. ``` org.apache.kafka kafka-clients 3.8.1 io.confluent kafka-json-schema-serializer 8.0.0 com.fasterxml.jackson.core jackson-databind 2.19.2 ``` ### Optional dependencies[​](#optional-dependencies "Direct link to Optional dependencies") Depending on your schema complexity or tooling, you might also need these: ``` com.fasterxml.jackson.datatype jackson-datatype-jdk8 2.19.2 com.fasterxml.jackson.datatype jackson-datatype-jsr310 2.19.2 com.github.erosb everit-json-schema 1.14.6 ``` These optional dependencies support features like Java 8 types, JSR-310 date/time classes, and schema validation. Add them only if your schema or generated class uses these features. Related pages * [Use schema registry with Java producers and consumers](/docs/products/kafka/howto/use-schema-registry-in-java.md) * [jsonschema2pojo CLI reference](https://github.com/joelittlejohn/jsonschema2pojo/wiki/Getting-Started#the-command-line-interface) --- # Generate Java classes from Protobuf schemas Generate Java classes from Protocol Buffers (`.proto`) schema files for use in Apache Kafka® producers and consumers. Use the `protoc` compiler to generate Java classes that match your schema structure. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An active [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md) * [Karapace Schema Registry](/docs/products/kafka/karapace/howto/enable-karapace.md) enabled * Java (JDK 8 or later) * [Apache Maven](https://maven.apache.org/) installed to manage dependencies * [`protoc`](https://grpc.io/docs/protoc-installation/) compiler installed * A Protobuf schema file (`.proto`) ## Generate the Java classes[​](#generate-the-java-classes "Direct link to Generate the Java classes") Generate Java classes from your Protobuf schema file using the `protoc` CLI: ``` protoc -I=. --java_out=src/main/java/ src/main/resources/example.proto ``` * Replace `example.proto` with your Protobuf schema file. * Replace `src/main/java/` with your preferred output directory. The generated class is placed inside a directory that matches the `package` declaration in your `.proto` file. Ensure the `package` in your `.proto` file matches your intended Java package structure. note To control the Java package of the generated classes, add the option `java_package` inside your `.proto` file. This cannot be set through the `protoc` compiler command-line options. Kafka serialization depends on the fully qualified class name, so ensure the `java_package` matches your project structure. ## Example schema[​](#example-schema "Direct link to Example schema") ``` syntax = "proto3"; package io.aiven.example.protobuf; option java_package = "io.aiven.example.protobuf"; message User { int32 id = 1; string email = 2; } ``` This schema produces the following file: ``` src/main/java/io/aiven/example/protobuf/User.java ``` ## Add Maven dependencies[​](#add-maven-dependencies "Direct link to Add Maven dependencies") Add these dependencies to your `pom.xml` to compile and use the generated classes for Protobuf serialization and deserialization in Kafka producers and consumers. ``` org.apache.kafka kafka-clients 3.8.1 io.confluent kafka-protobuf-serializer 8.0.0 io.confluent kafka-schema-registry-client 8.0.0 com.google.protobuf protobuf-java 4.32.0 ``` ### Optional dependencies[​](#optional-dependencies "Direct link to Optional dependencies") You can also use these optional dependencies depending on your use case: ``` io.confluent kafka-protobuf-types 8.0.0 io.confluent kafka-protobuf-provider 8.0.0 ``` These dependencies are optional. Use them if your schema includes well-known Protobuf types or if you use advanced features in Confluent’s Protobuf support. Related pages * [Use schema registry with Java producers and consumers](/docs/products/kafka/howto/use-schema-registry-in-java.md) * [Protobuf Compiler Installation](https://grpc.io/docs/protoc-installation/) --- # Stream sample data from the Aiven Console Use the sample data generator to simulate streaming events and observe how data flows through topics and schemas in your Aiven for Apache Kafka® service. ## About the sample data generator[​](#about-the-sample-data-generator "Direct link to About the sample data generator") The sample data generator helps you explore Aiven for Apache Kafka by producing realistic test messages to a Kafka topic in your service. It's designed for quick onboarding with no client configuration required. note You can use the sample data generator at no additional cost. It does not use service credits. You can choose from the following data scenarios: * **Logistics**: Tracks package events including timestamp, tracking ID, carrier, location, and delivery state such as *received*, *shipped*, or *in transit*. * **User activity**: Captures app or website interactions such as action type, page, user ID, and country code. * **Metrics**: Streams application metrics like percentages, averages, and totals across time windows (instant, hourly, 12-hour). Use the sample data generator to: * Start streaming data in as little as 30 seconds after service creation. * Validate how your Kafka service handles schema-based messages. * Explore how topics, schemas, and consumers interact. ## How it works[​](#how-it-works "Direct link to How it works") When you start a sample data session, the generator: * Enables the Schema Registry and REST Proxy if they are not already active. * Creates a system-generated topic for the selected scenario. * Applies a predefined Avro schema to the topic. * Produces messages at a steady, test-friendly rate. * Streams data for the selected duration, up to 4 hours. ## Limitations[​](#limitations "Direct link to Limitations") * Only one sample data session can run per Aiven for Apache Kafka service at a time. * Each user can run only one data generator scenario at a time. * Only the **Avro** schema format is supported. * Sample data is retained for up to one week. * The stream stops automatically when the selected time ends or the browser session is closed. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to the [Aiven Console](https://console.aiven.io). * Permission to create and manage Aiven for Apache Kafka services in your project. * An Aiven for Apache Kafka® service. To create one, see [Create an Aiven for Apache Kafka® Professional tier service](/docs/products/kafka/get-started/create-kafka-service.md). ## Start a data stream[​](#start-a-data-stream "Direct link to Start a data stream") 1. In the [Aiven Console](https://console.aiven.io), select an existing **Aiven for Apache Kafka** service or [create a service](/docs/products/kafka/get-started/create-kafka-service.md). 2. On the **Overview** page, in the **Start data stream** section, click **Generate sample data**. 3. In the setup wizard: 1. Choose a data scenario and click **Continue**. 2. Click **Enable & continue** to proceed. Aiven ensures the Karapace Schema Registry and REST Proxy are enabled for sample data generation. If they are not already active, Aiven enables them for you. 3. Review the schema your messages will follow. Click **Confirm** to continue. 4. Review the topic and configure stream settings. note On **Standard Kafka** services, you can choose **Classic topic** or **Diskless topic** for sample data. On **Classic Kafka** services, only classic topics are used. For more information, see [Diskless topics overview](/docs/products/kafka/diskless/concepts/diskless-topic-overview.md). 5. Click **Confirm**. 6. Select a stream duration between 15 minutes and 4 hours. Click **Start data stream**. After the stream starts, click **Open service overview** to monitor progress from the **Overview** page. You can view the message rate, remaining time, and a link to review messages. ### Monitor and manage the stream[​](#monitor-and-manage-the-stream "Direct link to Monitor and manage the stream") * To view the streamed data, click **Review messages** on the **Overview** page. The **Messages** page opens for the topic. Messages appear in Avro format within a few seconds. * To stop the stream, click **Stop streaming** in the **Data generator** section on the **Overview** page. This option is only visible in the browser tab where the generator was started. * To view topic settings, message details, or the applied schema, click **Manage stream** > **Topics** in the sidebar, then select the topic created by the generator. Related pages * [View topic details and partitions](/docs/products/kafka/howto/get-topic-partition-details.md) * [Enable Schema Registry and Kafka REST Proxy](/docs/products/kafka/karapace/howto/enable-karapace.md) * [Create an Aiven for Apache Kafka® Professional tier service](/docs/products/kafka/get-started/create-kafka-service.md) --- # Generate sample data with Docker Use a Docker-based producer to generate sample data in Aiven for Apache Kafka®. It creates a customizable stream of messages for testing and development. This example uses [Docker](https://www.docker.com/) images. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Docker](https://www.docker.com/) or [Podman](https://podman.io/) installed * Active [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md) * Access to the [Aiven Console](https://console.aiven.io) * [Kafka REST API](https://aiven.io/docs/products/kafka/karapace/howto/enable-karapace) enabled * [Personal access token](/docs/platform/howto/create_authentication_token.md) ## Set up the fake data producer[​](#set-up-the-fake-data-producer "Direct link to Set up the fake data producer") Use the [Dockerized fake data producer for Aiven for Apache Kafka®](https://github.com/aiven/fake-data-producer-for-apache-kafka-docker) to stream sample messages into a topic. 1. Clone the repository: ``` git clone https://github.com/aiven/fake-data-producer-for-apache-kafka-docker ``` 2. Copy the sample config file to create your own version: ``` cp conf/env.conf.sample conf/env.conf ``` 3. Open the `conf/env.conf` file and update the following values: * `my_project_name`: Aiven project name * `my_kafka_service_name`: Apache Kafka service name * `my_topic_name`: Topic name to receive messages * `my_aiven_email`: Aiven login email * `my_aiven_token`: Personal access token 4. Generate a [personal access token](/docs/platform/howto/create_authentication_token.md) in the [Aiven Console](https://console.aiven.io) or use the [Aiven CLI](/docs/tools/cli.md): ``` avn user access-token create \ --description "Token used by fake data generator" \ --max-age-seconds 3600 \ --json | jq -r '.[].full_token' ``` tip The command uses [`jq`](https://stedolan.github.io/jq/) to extract the token from the Aiven CLI output. If `jq` is not installed, remove the `| jq -r '.[].full_token'` part and copy the token manually from the JSON output. 5. Build the Docker image: ``` docker build -t fake-data-producer-for-apache-kafka-docker . ``` tip Rebuild the Docker image after editing the `conf/env.conf` file. 6. Start the producer: ``` docker run fake-data-producer-for-apache-kafka-docker ``` 7. Once the Docker image is running, verify that the topic is receiving messages: * In the [Aiven Console](https://console.aiven.io), go to your Apache Kafka service and click **Manage stream** > **Topics**. * Or use a command-line tool such as [kcat](/docs/products/kafka/howto/kcat.md) to consume messages from the topic. Related pages [Stream sample data from the Aiven Console](/docs/products/kafka/howto/generate-sample-data.md) --- # Get partition details of an Apache Kafka® topic Learn how to get partition details of an Apache Kafka® topic. * Console * API * CLI 1. Log in to [Aiven Console](https://console.aiven.io/) and select your Aiven for Apache Kafka service. 2. Select **Manage stream** > **Topics** from the left sidebar. 3. Select a specific topic or click the ellipsis (More options). 4. On the **Topics info** screen, select **Partitions** to view detailed information about the partitions. Retrieve topic details with an API call using [the endpoint to get Kafka topic info](https://api.aiven.io/doc/#operation/ServiceKafkaTopicGet). Learn more about API usage in the [Aiven API overview](/docs/tools/api.md). Retrieve topic details by using Aiven CLI commands. Find the full list of commands for `avn service topic` in [the CLI reference](/docs/tools/cli/service/topic.md). For example, this bash script, with a help of `jq` utility, lists topics and their details for a specified Apache Kafka service: ``` #!/bin/bash proj=${1:-YOUR-AIVEN-PROJECT-NAME} serv="${2:-YOUR-KAFKA-SERVICE-NAME}" cloud=$(avn service get --project $proj $serv --json | jq -r '.cloud_name') topics=$(avn service topic-list --project $proj $serv --json | jq -r '.[] | .topic_name') echo "Cloud: $cloud Service: $serv" for topic in $topics do echo "Topic: $topic" avn service topic-get --project $proj $serv $topic done ``` --- # Governance in Aiven for Apache Kafka® [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Governance in Aiven for Apache Kafka® helps manage your Aiven for Apache Kafka clusters securely and efficiently through structured policies, roles, and processes. ## [Governance](/docs/products/kafka/concepts/governance-overview.md) [Governance in Aiven for Apache Kafka® provides a structured and secure way to manage your Aiven for Apache Kafka clusters, ensuring security, compliance, and efficiency.](/docs/products/kafka/concepts/governance-overview.md) ## [Enable governance](/docs/products/kafka/howto/enable-governance.md) [Enable governance in Aiven for Apache Kafka® to create a secure and compliant framework to manage your Aiven for Apache Kafka services efficiently.](/docs/products/kafka/howto/enable-governance.md) ## [Governance with Terraform](/docs/products/kafka/howto/terraform-governance-approvals.md) [Aiven for Apache Kafka® Governance lets you manage approval workflows for Apache Kafka topic changes using Terraform and GitHub Actions.](/docs/products/kafka/howto/terraform-governance-approvals.md) ## [Claim topic ownership](/docs/products/kafka/howto/claim-topic.md) [To take ownership of a topic that your group does not currently own, you can submit a claim request. Once you send the request, the current owner can approve or decline it.](/docs/products/kafka/howto/claim-topic.md) ## [Manage topic requests](/docs/products/kafka/howto/manage-resource-requests.md) [4 items](/docs/products/kafka/howto/manage-resource-requests.md) --- # Manage group requests [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) The **Group requests** page allows you to view and track requests you and other members of your group made for Aiven for Apache Kafka® resources. These requests can include claims for existing topics or requests to create new topics. ## Key elements[​](#key-elements "Direct link to Key elements") The **group requests** page displays the following key elements: * **Topic**: Name of the topic * **Type**: Type of request * **Status**: Current status of the request * **Approved**: Request approved and processed * **Declined**: Request declined. The topic is now available for another group to claim * **Deleted**: Request deleted by the requester * **Failed**: Request cannot complete the required action (for example, due to Apache Kafka cluster issues or connectivity problems) * **Pending**: Request awaiting approval or decline * **Provisioning**: Request in progress. The topic is being created on the Apache Kafka cluster * **Requesting group**: Group that made the request * **Requester**: User who submitted the request * **Requested on**: Date and time the request was made * **Filters**: Filter the list by type, status, requester, and group * **Search function**: Search for specific topics by name ## View and track requests[​](#view-and-track-requests "Direct link to View and track requests") To view and track requests made by you and other members of your group: * Access the [Aiven console](https://console.aiven.io/) and click **Tools** > **Governance** > ****Group requests****. * Use the search bar to find specific requests by topic name. * Filter requests by type, status, requester, and group. * Click the topic name to view detailed information about the request. ## Delete requests[​](#delete-requests "Direct link to Delete requests") You can only delete requests you have created. You can view requests from other group members but cannot delete them. 1. Access the [Aiven console](https://console.aiven.io/) and click **Tool** > **Governance** > ****Group requests****. 2. Click the topic name to delete. 3. In the **Review request** pane, click **Delete**. Related pages * [Governance in Aiven for Apache Kafka](/docs/products/kafka/concepts/governance-overview.md) * [Aiven for Apache Kafka® topic catalog](/docs/products/kafka/concepts/topic-catalog-overview.md) * [Claim topic ownership](/docs/products/kafka/howto/claim-topic.md) * [Approvals](/docs/products/kafka/howto/approvals.md) --- # Integrate an external Apache Kafka® cluster in Aiven You can integrate an external Apache Kafka® cluster with your Aiven for Apache Kafka® service. This setup supports use cases such as replication, data ingestion, or connecting Aiven Kafka to on-premises or third-party clusters. To integrate, define an external Kafka service endpoint in the Aiven Console. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An active Aiven account with access to the [Aiven Console](https://console.aiven.io/) * Connection details for your external Kafka cluster, depending on your setup: * Bootstrap server addresses * Security protocol type * Authentication credentials (if required) * SSL certificates (if using SSL/TLS) ## Define an external Apache Kafka® endpoint[​](#define-an-external-apache-kafka-endpoint "Direct link to Define an external Apache Kafka® endpoint") 1. Log in to the [Aiven Console](https://console.aiven.io/) and select your project. 2. In the left sidebar, click **Integration endpoints**. 3. From the list of available services, select **External Apache Kafka**. 4. Click **Add a new endpoint**. 5. In the **Create new External Apache Kafka endpoint** dialog, enter the following details: * **Endpoint name**: Enter a descriptive name for your Kafka connection. * **Bootstrap servers**: Enter a comma-separated list of bootstrap server addresses (for example, `kafka-broker-1:9092,kafka-broker-2:9092`). The minimum length is 3 characters and the maximum is 256 characters. 6. Select the **Security protocol** from the dropdown menu. The available options and required fields depend on your selection: 7. Select the **Security protocol** from the dropdown menu. The available options and required fields depend on your selection: | Security protocol | Description | Required fields | | ----------------- | -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------- | | PLAINTEXT | No encryption or authentication. | - Bootstrap servers | | SSL | SSL encryption and optional client authentication. | - CA certificate
- Client certificate (optional)
- Client key (optional)
- Endpoint identification algorithm | | SASL\_PLAINTEXT | SASL authentication without encryption. | - Username
- Password
- SASL mechanism | | SASL\_SSL | SASL authentication with SSL encryption. | - Username
- Password
- SASL mechanism
- CA certificate
- Endpoint identification algorithm | 8. Click **Create**. The external Kafka cluster appears in the **Integration endpoints** list under the name specified in **Endpoint name**. You can review, edit, or remove it from this list at any time. Related pages * [Apache Kafka® documentation](https://kafka.apache.org/documentation/) * [MirrorMaker 2 overview](/docs/products/kafka/kafka-mirrormaker.md) * [Permissions and internal topics in MirrorMaker 2](/docs/products/kafka/kafka-mirrormaker/concepts/permissions-internal-topics.md) --- # Integration of logs into Apache Kafka® topic You can send logs from your Aiven services into a specified Apache Kafka® topic. The setup can be done through [Aiven console](https://console.aiven.io). note The integration can be used for both Aiven for Apache Kafka, as well as external Kafka clusters registered in the project's service integrations page. Read more on [how to manage Aiven internal and external integrations](/docs/tools/cli/service/integration.md). In this example, you will learn how to send logs from an Aiven for PostgreSQL® service to a topic in your Aiven for Apache Kafka® service. To proceed with this example, you will need the following: * A running Aiven for PostgreSQL service, which will be referred to as the *source* service. * A running Aiven for Apache Kafka service, which will be referred to as the *destination* service. ## Set up Apache Kafka to receive the logs[​](#set-up-apache-kafka-to-receive-the-logs "Direct link to Set up Apache Kafka to receive the logs") Verify that [Apache Kafka REST API](/docs/products/kafka/concepts/kafka-rest-api.md) is enabled and create a Kafka topic where to receive the logs. ## Add a new integration to the source service[​](#add-a-new-integration-to-the-source-service "Direct link to Add a new integration to the source service") 1. Log in to [Aiven console](https://console.aiven.io) and select your source PostgreSQL service. 2. In the sidebar, click **Connect** > **Integrations**. 3. In **Aiven services**, select **Apache Kafka Logs**. 4. Select the destination Kafka service (or external Kafka integration) and select **Continue**. 5. Enter the desired **Topic name** where you want the logs to be produced. ## Test the integration (with Aiven for Apache Kafka)[​](#test-the-integration-with-aiven-for-apache-kafka "Direct link to Test the integration (with Aiven for Apache Kafka)") 1. Access your destination Apache Kafka service. 2. Select **Manage stream** > **Topics** from the left sidebar and locate the topic you specified to send logs. 3. From the **Topic info** screen, select **Messages**. note Alternatively, you can access the messages for a topic by selecting the ellipsis in the row of the topic and choosing **Topic messages**. 4. In the **Messages** screen, select **Fetch Messages** to view the log entries that were sent from your source service. 5. To see the messages in JSON format, use the **FORMAT** drop-down menu and select *json*. ## Edit or remove the integration[​](#edit-or-remove-the-integration "Direct link to Edit or remove the integration") To edit or remove the integration, use **Manage Integrations** in the source service. The created integration is listed in the **Enabled service integrations** section, from where you can edit or remove it. Related pages [Set up a log integration with an Aiven for OpenSearch® service](/docs/products/opensearch/howto/opensearch-log-integration.md) --- # Enable IPv6 connectivity for Aiven for Apache Kafka® [Early availability](/docs/platform/concepts/service-and-feature-releases.md) Aiven for Apache Kafka® supports dual-stack IPv4 and IPv6 connectivity. Kafka clients can connect using either address type. warning This feature is in early availability. Enable it in a non-production environment first. ## Enable IPv6 connectivity[​](#enable-ipv6-connectivity "Direct link to Enable IPv6 connectivity") To enable dual-stack IPv4 and IPv6 support, set the service user configuration `enable_ipv6` to `true`. * Console * CLI * API 1. In the [Aiven Console](https://console.aiven.io), select the Aiven for Apache Kafka® service. 2. Click **Service settings** in the sidebar. 3. In the **Cloud and network** section, click **Actions** and select **More network configurations**. 4. In the **Network configuration** dialog, click **Add configuration options**. 5. Search for `enable_ipv6`, select the option from the list, and set the value to **Enabled**. 6. Click **Save configuration**. Enable IPv6 connectivity on an existing service using Aiven CLI: ``` avn service update SERVICE_NAME -c enable_ipv6=true ``` Parameters: * `SERVICE_NAME`: Name of the Aiven for Apache Kafka® service. Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) API to enable IPv6 connectivity: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer API_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "enable_ipv6": true } }' ``` Parameters: * `PROJECT_NAME`: Name of the project. * `SERVICE_NAME`: Name of the service. * `API_TOKEN`: API token for authentication. ## How IPv6 connectivity works[​](#how-ipv6-connectivity-works "Direct link to How IPv6 connectivity works") ### Server behavior[​](#server-behavior "Direct link to Server behavior") When IPv6 connectivity is enabled: * Existing DNS `A` records continue to resolve to IPv4 addresses. * The service creates DNS `AAAA` records that resolve to IPv6 addresses. * Kafka brokers listen on the same ports for both IPv4 and IPv6 traffic. ### Client behavior[​](#client-behavior "Direct link to Client behavior") Kafka clients that can resolve IPv6 addresses can connect using either IPv4 or IPv6, depending on the client configuration and network setup. To control the preferred address family: * **Java clients** * Prefer IPv6: `-Djava.net.preferIPv6Addresses=true` * Prefer IPv4: `-Djava.net.preferIPv4Addresses=true` * **librdkafka-based clients (for example, kcat)** * Prefer IPv6: `broker.address.family=v6` * Prefer IPv4: `broker.address.family=v4` note The Java JVM setting also affects HTTP connections, including Schema Registry. The `librdkafka` setting does not. ## Supported configurations[​](#supported-configurations "Direct link to Supported configurations") IPv6 connectivity supports the following configurations: **Authentication methods** * SSL certificate * SASL using project CA * SASL using public CA **Access routes** * Dynamic * Public * Private (excluding VPC peering and PrivateLink) ## Limitations[​](#limitations "Direct link to Limitations") The following limitations apply when IPv6 connectivity is enabled: * You cannot fully disable IPv4. * VPC peering routes support IPv4 only. * PrivateLink routes support IPv4 only. * Static IP addresses support IPv4 only. ### Impact[​](#impact "Direct link to Impact") When IPv6 connectivity is enabled, some clients may resolve IPv6 addresses but fail to connect over access routes that rely on IPv4. If connection failures occur, configure Kafka clients to prefer IPv4 addresses: * **Java clients**: Set `-Djava.net.preferIPv4Addresses=true`. * **librdkafka-based clients (for example, kcat)**: Set `broker.address.family=v4`. If Schema Registry connections fail, configure the client to use the service public endpoint. --- # Use Kafbat UI with Aiven for Apache Kafka® [Kafbat UI](https://github.com/kafbat/kafka-ui) is a popular Open-Source web GUI for Apache Kafka® management that allows you to monitor and manage Apache Kafka® clusters. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") To connect Kafbat UI to Aiven for Apache Kafka®, create a [Java keystore and truststore containing the service SSL certificates](/docs/products/kafka/howto/keystore-truststore.md). Also collect the following information: * `APACHE_KAFKA_HOST`: The Aiven for Apache Kafka hostname * `APACHE_KAFKA_PORT`: The Aiven for Apache Kafka port * `SSL_KEYSTORE_FILE_NAME`: The name of the Java keystore containing the Aiven for Apache Kafka SSL certificates * `SSL_TRUSTSTORE_FILE_NAME`: The name of the Java truststore containing the Aiven for Apache Kafka SSL certificates * `SSL_KEYSTORE_PASSWORD`: The password used to secure the Java keystore * `SSL_TRUSTSTORE_PASSWORD`: The password used to secure the Java truststore * `SSL_STORE_FOLDER`: The absolute path of the folder containing both the truststore and keystore ## Run Kafbat UI on Docker or Podman[​](#run-kafbat-ui-on-docker-or-podman "Direct link to Run Kafbat UI on Docker or Podman") ### Share keystores with non-root user[​](#share-keystores-with-non-root-user "Direct link to Share keystores with non-root user") Since container for Kafbat UI uses non-root user, to avoid permission problems, while keeping the secrets safe: 1. Create separate directory for secrets: ``` mkdir SSL_STORE_FOLDER ``` 2. Restrict the directory to current user: ``` chmod 700 SSL_STORE_FOLDER ``` 3. Copy secrets there (replace the `SSL_KEYSTORE_FILE_NAME` and `SSL_TRUSTSTORE_FILE_NAME` with the keystores and truststores file names): ``` cp SSL_KEYSTORE_FILE_NAME SSL_TRUSTSTORE_FILE_NAME SSL_STORE_FOLDER ``` 4. Give read permissions for secret files for everyone: ``` chmod +r SSL_STORE_FOLDER/* ``` ### Execute Kafbat UI on Docker or Podman[​](#execute-kafbat-ui-on-docker-or-podman "Direct link to Execute Kafbat UI on Docker or Podman") You can run Kafbat UI in a Docker/Podman container with the following command, by replacing the placeholders for: * `APACHE_KAFKA_HOST` * `APACHE_KAFKA_PORT` * `SSL_STORE_FOLDER` * `SSL_KEYSTORE_FILE_NAME` * `SSL_KEYSTORE_PASSWORD` * `SSL_TRUSTSTORE_FILE_NAME` * `SSL_TRUSTSTORE_PASSWORD` ``` docker run -p 8080:8080 \ -v SSL_STORE_FOLDER/SSL_TRUSTSTORE_FILE_NAME:/client.truststore.jks:ro \ -v SSL_STORE_FOLDER/SSL_KEYSTORE_FILE_NAME:/client.keystore.p12:ro \ -e KAFKA_CLUSTERS_0_BOOTSTRAPSERVERS=APACHE_KAFKA_HOST:APACHE_KAFKA_PORT \ -e KAFKA_CLUSTERS_0_PROPERTIES_SECURITY_PROTOCOL=SSL \ -e KAFKA_CLUSTERS_0_PROPERTIES_SSL_TRUSTSTORE_LOCATION=/client.truststore.jks \ -e KAFKA_CLUSTERS_0_PROPERTIES_SSL_TRUSTSTORE_PASSWORD=SSL_TRUSTSTORE_PASSWORD \ -e KAFKA_CLUSTERS_0_PROPERTIES_SSL_KEYSTORE_LOCATION=/client.keystore.p12 \ -e KAFKA_CLUSTERS_0_PROPERTIES_SSL_KEYSTORE_PASSWORD=SSL_KEYSTORE_PASSWORD \ -e KAFKA_CLUSTERS_0_PROPERTIES_SSL_KEYSTORE_TYPE=PKCS12 \ -d ghcr.io/kafbat/kafka-ui:latest ``` ## Use Kafbat UI[​](#use-kafbat-ui "Direct link to Use Kafbat UI") Once Kafbat UI for Apache Kafka® starts, you should be able to access it at `localhost:8080`. ![Kafbat in action](/docs/assets/images/kafbat-ui-37d19f4869e418624b994853ee281342.jpg) --- # Use Kafdrop Web UI with Aiven for Apache Kafka® [Kafdrop](https://github.com/obsidiandynamics/kafdrop) is a web UI for Apache Kafka® to monitor clusters, view topics and consumer groups, and integrate with the Schema Registry. It supports Avro, JSON, and Protobuf. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to the [Aiven Console](https://console.aiven.io/) * Aiven CLI * Docker installed. For Mac M1 chip users, ensure Docker is version 4.26.1 or lower to avoid compatibility issues ## Retrieve SSL certificate files[​](#retrieve-ssl-certificate-files "Direct link to Retrieve SSL certificate files") Aiven for Apache Kafka uses TLS security by default. To retrieve the necessary SSL certificates, do one of the following: * [Download the certificates](/docs/platform/concepts/tls-ssl-certificates.md#download-ca-certificates) manually from the service **Overview** page in the [Aiven Console](https://console.aiven.io/). * Use the [Aiven CLI command](/docs/tools/cli/service/user.md#avn_service_user_creds_download). ## Set up a Kafdrop configuration file[​](#set-up-a-kafdrop-configuration-file "Direct link to Set up a Kafdrop configuration file") Kafdrop supports both [SASL and SSL authentication methods](/docs/products/kafka/concepts/auth-types.md). This example uses SSL, which requires a keystore and truststore. 1. Follow [keystore and truststore](/docs/products/kafka/howto/keystore-truststore.md) to create the necessary files. 2. Create a Kafdrop configuration file named `kafdrop.properties` with the following content. Replace `KEYSTORE_PWD` and `TRUSTSTORE_PWD` with your keystore and truststore passwords: ``` security.protocol=SSL ssl.keystore.password=KEYSTORE_PWD ssl.keystore.type=PKCS12 ssl.truststore.password=TRUSTSTORE_PWD ``` ## Run Kafdrop on Docker[​](#run-kafdrop-on-docker "Direct link to Run Kafdrop on Docker") To run Kafdrop in a Docker or Podman container, use the following command. Replace `KAFKA_SERVICE_URI` with your Aiven for Apache Kafka® service URI from the **Overview** page in the [Aiven Console](https://console.aiven.io/). Replace `client.truststore.jks` and `client.keystore.p12` with your keystore and truststore file names: ``` docker run -d --rm -p 9000:9000 \ -e KAFKA_BROKERCONNECT=KAFKA_SERVICE_URI \ -e KAFKA_PROPERTIES="$(cat kafka/kafdrop.properties | base64)" \ -e KAFKA_TRUSTSTORE="$(cat kafka/client.truststore.jks | base64)" \ -e KAFKA_KEYSTORE="$(cat kafka/client.keystore.p12 | base64)" \ obsidiandynamics/kafdrop ``` For users with a Mac M1 chip, add the `--platform linux/amd64` flag to the command: ``` docker run --platform linux/amd64 -d --rm -p 9000:9000 \ -e KAFKA_BROKERCONNECT=KAFKA_SERVICE_URI \ -e KAFKA_PROPERTIES="$(cat kafka/kafdrop.properties | base64)" \ -e KAFKA_TRUSTSTORE="$(cat kafka/client.truststore.jks | base64)" \ -e KAFKA_KEYSTORE="$(cat kafka/client.keystore.p12 | base64)" \ obsidiandynamics/kafdrop ``` note Docker versions above 4.26.1 have known issues on Mac M1 chips. We recommend using Docker version 4.26.1 or lower. If you need Kafdrop to deserialize Avro messages using the [Karapace](https://github.com/aiven/karapace) schema registry, add the following two lines to the `docker run` command: ``` -e SCHEMAREGISTRY_AUTH="avnadmin:SCHEMA_REGISTRY_PWD" \ -e SCHEMAREGISTRY_CONNECT="https://SCHEMA_REGISTRY_URI" \ ``` Replace `SCHEMA_REGISTRY_PWD` with the schema registry password and `SCHEMA_REGISTRY_URI` with the schema registry URI available on the **Overview** page in the [Aiven Console](https://console.aiven.io/). ## Access Kafdrop[​](#access-kafdrop "Direct link to Access Kafdrop") After Kafdrop starts, you can access it at `localhost:9000`: [](/docs/images/content/products/kafka/kafdrop.mp4) With Kafdrop, you can perform the following tasks over an Aiven for Apache Kafka® service: * View and search topics * Create and delete topics * View brokers * View messages --- # Connect to Apache Kafka® with Conduktor [Conduktor](https://www.conduktor.io/) is a friendly user interface for Apache Kafka and it works with Aiven and offers built-in support for setting up the connection. 1. Visit the **Service overview** page for your Aiven for Apache Kafka® service (the [Getting started with Aiven for Apache Kafka®](/docs/products/kafka/get-started/create-kafka-service.md) page is a good place for more information about creating a new service if you don't have one already). 2. Download the **Access Key**, **Access Certificate** and **CA Certificate**. 3. Choose **New Kafka Cluster** on the main pane, and click the **Aiven** icon. Add the following fields: * **Host** and **port**, you can copy these from the **Service overview** page. * The three files downloaded: access key, access certificate and CA certificate. ![Screenshot of the cluster configuration screen](/docs/assets/images/conduktor-config-a502ce09f45adf98b46efd4f953a783e.png) Conduktor will create the keystore and truststore files in the folder that you specified, or you can choose an alternative location. Click **Create** and the helper will create the configuration for Conduktor to connect to your Aiven for Apache Kafka service. 4. Click **Test Kafka Connectivity**. tip If you experience a Java SSL error when testing the connectivity, add the service CA certificate to the list of Conduktor's trusted certificates. * Download the **CA Certificate** file to your computer. * In the Conduktor application, click the settings dropdown in the bottom right hand side and choose **Network**. * On the **Trusted Certificates** tab, select **Import** and supply the CA certificate file you downloaded. Save the settings. Once connected, you can visit the [Conduktor documentation](https://docs.conduktor.io/) to learn more about using this tool. --- # Encrypt client-side with a custom serializer and deserializer With the Aiven platform, there are several deployment models available to meet your security and compliance needs: * VPC peering to securely peer the Aiven services to your cloud VPC * Enhanced Compliance Environments (ECE) to satisfy additional compliance needs such as HIPAA and PCI-DSS * Bring your own cloud (BYOC) which allows deployment of Aiven services directly into your cloud account In addition to the above, all data transmitted to the Aiven services is encrypted in transit and at rest. In some cases however, there are additional compliance and legal needs which require the data to be encrypted prior to entering the Aiven for Kafka® service. This can be achieved using a custom serializer and deserializer (commonly referred to as a "custom serde") . ## AES encryption[​](#aes-encryption "Direct link to AES encryption") In this example, we will use AES encryption. AES provides numerous benefits for streaming systems such as Apache Kafka. * Can configure it to use 128-bit, 192-bit or 256-bit keys giving a high level of security * The algorithm is fast and efficient, making it well suited to real-time data processing use-cases * As the algorithm relies on a symmetric key as it simplifies the key management. Asymmetrical algorithms require managing separate encryption/decryption keys for all the producers/consumers. In this example, we hard code the key. This should never be done in practise. Instead a secure vault service such as HashiVault or native cloud vendor solution should be used to manage the keys. warning You are responsible for the key. If the key is lost, it may be impossible to recover the data. ## Producer[​](#producer "Direct link to Producer") ``` from cryptography.hazmat.primitives.ciphers import Cipher, algorithms, modes from cryptography.hazmat.backends import default_backend from cryptography.hazmat.primitives import padding from kafka import KafkaProducer import base64, os import json from datetime import datetime # 128-bit encryption key # This must be stored securely elsewhere key = b"\x2b\x7e\x15\x16\x28\xae\xd2\xa6\xab\xf7\x15\x88\x09\xcf\x4f\x3c" # Kafka server BOOTSTRAP_SERVER = os.environ.get("BOOTSTRAP_SERVER") def encrypt(plaintext): # Encrypt data using AES-128 in CBC mode cipher = Cipher(algorithms.AES(key), modes.CBC(key), backend=default_backend()) encryptor = cipher.encryptor() padded_plaintext = pad(plaintext) ciphertext = encryptor.update(padded_plaintext) + encryptor.finalize() return base64.b64encode(ciphertext) def pad(data): # Pad data to be encrypted padder = padding.PKCS7(128).padder() padded_data = padder.update(data) + padder.finalize() return padded_data class EncryptedValueSerializer: def __call__(self, msg): return encrypt(bytes(msg, "utf-8")) producer = KafkaProducer( bootstrap_servers=BOOTSTRAP_SERVER, security_protocol="SSL", ssl_cafile="ca.pem", ssl_certfile="service.cert", ssl_keyfile="service.key", value_serializer=EncryptedValueSerializer(), ) for i in range(10): producer.send( "my-secrets", json.dumps({"msg": "this is a test", "time": str(datetime.now())}) ) producer.flush() ``` ## Consumer[​](#consumer "Direct link to Consumer") ``` from cryptography.hazmat.primitives.ciphers import Cipher, algorithms, modes from cryptography.hazmat.backends import default_backend from cryptography.hazmat.primitives import padding from kafka import KafkaConsumer import base64, os # 128-bit encryption key # This must be stored securely elsewhere key = b"\x2b\x7e\x15\x16\x28\xae\xd2\xa6\xab\xf7\x15\x88\x09\xcf\x4f\x3c" # Kafka server BOOTSTRAP_SERVER = os.environ.get("BOOTSTRAP_SERVER") def decrypt(ciphertext): # Decrypt data using AES-128 in CBC mode ciphertext = base64.b64decode(ciphertext) cipher = Cipher(algorithms.AES(key), modes.CBC(key), backend=default_backend()) decryptor = cipher.decryptor() plaintext = decryptor.update(ciphertext) + decryptor.finalize() return unpad(plaintext) def unpad(data): # Unpad data that was encrypted unpadder = padding.PKCS7(128).unpadder() return unpadder.update(data) + unpadder.finalize() class EncryptedValueDeserializer: def __call__(self, value): return decrypt(value) consumer = KafkaConsumer( bootstrap_servers=BOOTSTRAP_SERVER, value_deserializer=EncryptedValueDeserializer(), security_protocol="SSL", ssl_cafile="ca.pem", ssl_certfile="service.cert", ssl_keyfile="service.key", group_id="group_id_1", auto_offset_reset="earliest", ) consumer.subscribe(["my-secrets"]) for message in consumer: print(message.value) ``` --- # Connect Aiven for Apache Kafka® with Klaw [Klaw](https://www.klaw-project.io/) is an open-source, web-based data governance toolkit for managing Apache Kafka® topics, ACLs, schemas, and connectors. It provides a self-service interface where teams can request Kafka configuration changes without administrator intervention. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md) * [Klaw cluster](https://www.klaw-project.io/docs/quickstart) * [Java keystore and truststore containing the service SSL certificates](/docs/products/kafka/howto/keystore-truststore.md) ## Connect Aiven for Apache Kafka to Klaw[​](#connect-aiven-for-apache-kafka-to-klaw "Direct link to Connect Aiven for Apache Kafka to Klaw") To connect your Aiven for Apache Kafka cluster to Klaw: 1. Log in to the **Klaw web interface**. 2. Click **Environments**, then click **Clusters**. 3. On the **Clusters** page, click **Add Cluster**. 4. Enter the following details: * **Cluster type**: Select **Kafka**. * **Cluster name**: Enter a name for the cluster. * **Protocol**: Select the protocol used by your Kafka service. note Depending on the protocol selected, [configure Klaw application.properties file](/docs/products/kafka/howto/kafka-klaw.md#klaw-application-properties-configs) to enable connection between Aiven for Apache Kafka and Klaw clusters. * **Kafka flavor**: Select **Aiven for Apache Kafka®**. * **Project name**: Select your Aiven project. * **Bootstrap server**: Enter the **Service URI** for your Kafka service (available on the **Overview page** in the **Connection information** section in the Aiven Console). * **Service name**: Enter the Kafka service name as defined in the [Aiven Console](https://console.aiven.io/). 5. Click **Save**. 6. Add the cluster to an environment: * Click **Environments**, then select **Environments** from the drop-down menu. * Click **Add Environment** and provide the following: * **Environment Name:** Select environment from the drop-down list. note To learn more, see [Clusters and environments](https://www.klaw-project.io/docs/Concepts/clusters-environments) in Klaw documentation. * **Select Cluster:** Choose the cluster you added. The bootstrap server and protocol fields are populated automatically. * **Default Partitions:** Enter the number of partitions required (default is 2). * **Maximum Partitions:** Enter the maximum number of partitions (default is 2). * **Default Replication factor:** Enter the required replication factor (default is 1). * **Max Replication factor:** Enter the maximum replication factor (default is 1). * **Topic prefix (optional)**: Enter a topic prefix if needed. * **Tenant:** The value is set to default Tenant. note Klaw is multi-tenant by default. Each tenant manages topics with their own teams in isolation. Every tenant has its own set of Apache Kafka® environments, and users of one tenant cannot view/access topics, or ACLS from other tenants. It provides isolation avoiding any security breach. For this topic, I have used the default tenant configuration. For more information, see [Klaw documentation](https://www.klaw-project.io/docs/getstarted#configure-the-cluster-to-sync). 7. Click **Save**. If the connection is successful, the **Environments** page shows a **blue thumbs-up icon**. If it fails, a **red thumbs-down icon** appears. Check the authentication protocol settings. See [Configure Klaw application.properties file](#klaw-application-properties-configs) for troubleshooting. ## Configure the Klaw `application.properties` file[​](#klaw-application-properties-configs "Direct link to klaw-application-properties-configs") [Klaw](https://www.klaw-project.io/) is a Spring Boot application that uses the `application.properties` file to configure application-level settings such as secret keys, authentication protocols, and Kafka cluster connection details. To connect Aiven for Apache Kafka® to Klaw, you must configure additional properties in the `application.properties` file, including the Kafka protocol and bootstrap server. By default, the file is located in the following directories: * `klaw/cluster-api/src/main/resources` * `klaw/core/src/main/resources` If Klaw is running in a Kubernetes environment, the file path inside the pod may vary depending on your deployment method. You can override the file location using the `spring.config.location` property: ``` -Dspring.config.location=/mnt/config/application.properties ``` Review your deployment manifest or use `kubectl exec` to inspect the pod and determine the `application.properties` file location. ### Secret key configuration[​](#secret-key-configuration "Direct link to Secret key configuration") Set the value of `klaw.clusterapi.access.base64.secret` with a secret key in the form of a Base64 encoded string in the `application.properties` file located in the following paths: * `klaw/cluter-api/src/main/resources` * `klaw/core/src/main/resources` ### Configure authentication protocol[​](#configure-authentication-protocol "Direct link to Configure authentication protocol") You can connect Aiven for Apache Kafka® using either of the following authentication protocols: * `PLAINTEXT` * `SSL`, `SASL PLAIN`, `SASL SSL` * `SASL SSL (GSSAPI / Kerberos)`, `SASL_SSL (SCRAM SHA 256/512)` note If you are using `PLAINTEXT`, you do not need to perform any additional configuration. #### Configure SASL authentication[​](#configure-sasl-authentication "Direct link to Configure SASL authentication") To use SSL as the authentication protocol to connect the Apache Kafka® cluster to Klaw: ##### Retrieve SSL certificate files[​](#retrieve-ssl-certificate-files "Direct link to Retrieve SSL certificate files") Aiven for Apache Kafka uses TLS encryption by default. Download the required certificate files from the service overview page in the Aiven Console or from the [Aiven CLI page](/docs/tools/cli/service/user.md#avn_service_user_kafka_java_creds). After downloading the certificates, ensure that you have configured the [Java SSL keystore and truststore](/docs/products/kafka/howto/keystore-truststore.md). Move the following files into a directory accessible to Klaw: * `client.keystore.p12` * `client.truststore.jks` You will reference these files in the `application.properties` configuration. ##### Configure SSL properties[​](#configure-ssl-properties "Direct link to Configure SSL properties") After retrieving the SSL certificate files and configuring the SSL keystore and truststore, update the `application.properties` file to enable SSL for your Kafka cluster. 1. In the **Klaw web interface**, go to **Clusters** and copy the **Cluster ID**. 2. Open the `application.properties` file located in the `klaw/cluster-api/src/main/resources` directory. 3. Configure the SSL properties to connect to Apache Kafka clusters by editing the following lines: ``` klawssl.kafkassl.keystore.location=client.keystore.p12 klawssl.kafkassl.keystore.pwd=klaw1234 klawssl.kafkassl.key.pwd=klaw1234 klawssl.kafkassl.truststore.location=client.truststore.jks klawssl.kafkassl.truststore.pwd=klaw1234 klawssl.kafkassl.keystore.type=pkcs12 klawssl.kafkassl.truststore.type=JKS ``` * Replace every instance of `klawssl` with the **Cluster ID** you used when adding the Kafka cluster in the Klaw web interface. * Replace `client.keystore.p12` with the full path to your keystore file. * Replace `client.truststore.jks` with the full path to your truststore file. * Replace the sample password values (`klaw1234`) with the actual passwords configured for your keystore and truststore. * Save the `application.properties` file. The following is an example of an `application.properties` file configured with Klaw Cluster ID, keystore, and truststore paths and passwords. ``` demo_cluster.kafkassl.keystore.location=/Users/demo.user/Documents/Klaw/demo-certs/client.keystore.p12 demo_cluster.kafkassl.keystore.pwd=Aiventest123! demo_cluster.kafkassl.key.pwd=Aiventest123! demo_cluster.kafkassl.truststore.location=/Users/demo.user/Documents/Klaw/demo-certs/client.truststore.jks demo_cluster.kafkassl.truststore.pwd=Aiventest123! demo_cluster.kafkassl.keystore.type=pkcs12 demo_cluster.kafkassl.truststore.type=JKS ``` note To add multiple SSL configurations, copy the configuration block and update the Cluster ID, file paths, and passwords for each cluster. #### Connect using SASL protocols[​](#connect-using-sasl-protocols "Direct link to Connect using SASL protocols") To use SASL-based authentication methods such as `SASL_PLAIN`, `SASL_SSL/PLAIN`, `SASL_SSL/GSSAPI` or `SASL_SSL/OAUTHBEARER`, update the `application.properties` file: * Locate the lines starting with `acc1.kafkasasl.jaasconfig.` * Uncomment the relevant lines * Enter the required values for your SASL configuration * Save the updated `application.properties` file Related pages * [Klaw documentation](https://www.klaw-project.io/docs). * [Klaw GitHub project repository](https://github.com/aiven/klaw) --- # Set up Kafka OAuth 2.0/OIDC authentication with AWS IAM using Outbound Identity Federation Use AWS IAM Outbound Identity Federation to authenticate Apache Kafka® clients with OAuth 2.0/OIDC. AWS IAM principals can connect to Aiven for Apache Kafka® without managing separate credentials by using short-lived JSON Web Tokens (JWTs) issued by AWS Security Token Service (AWS STS). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before you begin, make sure you have: * An [Aiven for Apache Kafka](/docs/products/kafka/get-started/create-kafka-service.md) service with [SASL authentication](/docs/products/kafka/howto/kafka-sasl-auth.md#enable-sasl-authentication) enabled. * The [Aiven CLI](/docs/tools/cli.md) installed. * The [AWS CLI](https://docs.aws.amazon.com/cli/latest/userguide/getting-started-install.html) installed. * [`jq`](https://jqlang.github.io/jq/) installed. Define the audience value, typically as your Kafka service hostname. Replace `SERVICE_NAME` with your value: ``` AUDIENCE=$(avn service get SERVICE_NAME --json | jq .user_config.kafka.sasl_oauthbearer_expected_audience) ``` Use the same audience value when you configure AWS, Aiven for Apache Kafka, and your Kafka client. ## Enable Outbound Identity Federation in AWS[​](#enable-outbound-identity-federation-in-aws "Direct link to Enable Outbound Identity Federation in AWS") Enable Outbound Identity Federation for your AWS account: ``` aws iam enable-outbound-web-identity-federation ``` ## Verify IAM permissions[​](#verify-iam-permissions "Direct link to Verify IAM permissions") Make sure the AWS principal that requests the token has permission to call `sts:GetWebIdentityToken` with the configured audience. The following example policy grants the minimum required permissions: ``` { "Version": "2012-10-17", "Statement": [{ "Effect": "Allow", "Action": "sts:GetWebIdentityToken", "Resource": "*", "Condition": { "ForAllValues:StringEquals": { "sts:IdentityTokenAudience": "${AUDIENCE}" }, "NumericLessThanEquals": { "sts:DurationSeconds": 3600 } } }] } ``` ## Retrieve the issuer URL[​](#retrieve-the-issuer-url "Direct link to Retrieve the issuer URL") Get the issuer URL from your AWS account configuration: ``` ISSUER_URL=$(aws iam get-outbound-web-identity-federation-info | jq -r .IssuerIdentifier) ``` The issuer URL has the form `https://.tokens.sts.global.api.aws`. ## Configure OIDC for Aiven for Apache Kafka[​](#configure-oidc-for-aiven-for-apache-kafka "Direct link to Configure OIDC for Aiven for Apache Kafka") Configure your Kafka service with the OIDC issuer, JSON Web Key Set (JWKS) endpoint, audience, and subject claim. Replace `PROJECT_NAME` and `SERVICE_NAME` with your values: ``` avn service update --project PROJECT_NAME SERVICE_NAME \ -c kafka.sasl_oauthbearer_jwks_endpoint_url="${ISSUER_URL}/.well-known/jwks.json" \ -c kafka.sasl_oauthbearer_expected_issuer="${ISSUER_URL}" \ -c kafka.sasl_oauthbearer_expected_audience="${AUDIENCE}" \ -c kafka.sasl_oauthbearer_sub_claim_name="sub" ``` note This command triggers a rolling restart of your Apache Kafka brokers. To minimize impact, apply it during a maintenance window. For details about each parameter, see [OIDC parameters](/docs/products/kafka/howto/enable-oidc.md#oidc-parameters). ## Configure Kafka ACLs for your IAM principal[​](#configure-kafka-acls-for-your-iam-principal "Direct link to Configure Kafka ACLs for your IAM principal") The JWT issued by AWS STS contains the IAM principal Amazon Resource Name (ARN) as the `sub` claim. Aiven for Apache Kafka uses this value as the Kafka principal, so configure your access control lists (ACLs) to reference the full ARN prefixed with `User:`. Use the Aiven CLI to add the required ACLs. Replace `PROJECT_NAME`, `SERVICE_NAME`, and the ARN with your values: ``` PRINCIPAL="User:arn:aws:iam::123456789012:user/my-iam-user" # Allow topic describe operations. avn service kafka-acl-add --project PROJECT_NAME SERVICE_NAME \ --permission allow \ --principal "${PRINCIPAL}" \ --operation Describe \ --resource-type topic \ --pattern-type literal \ --resource-name my-topic # Allow produce operations. avn service kafka-acl-add --project PROJECT_NAME SERVICE_NAME \ --permission allow \ --principal "${PRINCIPAL}" \ --operation Write \ --resource-type topic \ --pattern-type literal \ --resource-name my-topic # Allow consume operations from the topic. avn service kafka-acl-add --project PROJECT_NAME SERVICE_NAME \ --permission allow \ --principal "${PRINCIPAL}" \ --operation Read \ --resource-type topic \ --pattern-type literal \ --resource-name my-topic # Allow consume operations with the consumer group. avn service kafka-acl-add --project PROJECT_NAME SERVICE_NAME \ --permission allow \ --principal "${PRINCIPAL}" \ --operation Read \ --resource-type group \ --pattern-type literal \ --resource-name my-consumer-group ``` To verify that the ACLs are in place, run the following command: ``` avn service kafka-acl-list --project PROJECT_NAME SERVICE_NAME ``` The output is similar to the following: ``` ID PERMISSION_TYPE PRINCIPAL OPERATION RESOURCE_TYPE PATTERN_TYPE RESOURCE_NAME HOST ============== =============== ===================================================== ========= ============= ============ ================= ==== acl5baf3a5cada ALLOW User:arn:aws:iam::123456789012:user/my-iam-user Describe Topic LITERAL my-topic * acl5baf3a77c15 ALLOW User:arn:aws:iam::123456789012:user/my-iam-user Write Topic LITERAL my-topic * acl5baf3a60d8f ALLOW User:arn:aws:iam::123456789012:user/my-iam-user Read Topic LITERAL my-topic * acl5baf3ab8316 ALLOW User:arn:aws:iam::123456789012:user/my-iam-user Read Group LITERAL my-consumer-group * ``` ## Configure the Kafka client[​](#configure-the-kafka-client "Direct link to Configure the Kafka client") Your client must request a JWT from AWS STS at runtime and use it to authenticate with the Kafka service by using the `OAUTHBEARER` SASL mechanism. ### Python example[​](#python-example "Direct link to Python example") Install the required libraries: ``` pip install confluent-kafka boto3 PyJWT ``` Download the CA certificate for your service from the Aiven Console. Use the following code: ``` from confluent_kafka import Producer, Consumer import boto3 import jwt # Aiven Kafka config KAFKA_BOOTSTRAP = ":" CA_CERT_PATH = "ca.pem" # Download from the Aiven Console AUDIENCE = "" # Must match the server-side audience setting def fetch_token_from_aws_sts(config): """ Fetch a short-lived JWT from AWS STS. Equivalent to: aws sts get-web-identity-token \ --audience "" \ --signing-algorithm RS256 \ --duration-seconds 3600 """ sts_client = boto3.client("sts") response = sts_client.get_web_identity_token( Audience=[AUDIENCE], DurationSeconds=3600, SigningAlgorithm="RS256", ) token = response["WebIdentityToken"] expiry = response["Expiration"].timestamp() return token, expiry # Verify that token retrieval works before connecting. token, expiry = fetch_token_from_aws_sts(None) decoded = jwt.decode(token, options={"verify_signature": False}) print(decoded) # Client config kafka_config = { "bootstrap.servers": KAFKA_BOOTSTRAP, "security.protocol": "SASL_SSL", "sasl.mechanism": "OAUTHBEARER", "oauth_cb": fetch_token_from_aws_sts, "ssl.ca.location": CA_CERT_PATH, } # Producer example delivery_errors = [] def delivery_callback(err, msg): if err: delivery_errors.append(err) else: print( f"Message produced to {msg.topic()} [{msg.partition()}] at offset {msg.offset()}" ) producer = Producer(kafka_config) producer.produce("my-topic", key="key", value="hello from OIDC", callback=delivery_callback) producer.flush() if delivery_errors: raise RuntimeError(f"Message delivery failed: {delivery_errors[0]}") # Consumer example consumer_config = { **kafka_config, "group.id": "my-consumer-group", "auto.offset.reset": "earliest", } consumer = Consumer(consumer_config) consumer.subscribe(["my-topic"]) try: while True: msg = consumer.poll(timeout=5.0) if msg is None: print("No message received.") break if msg.error(): print(f"Consumer error: {msg.error()}") continue print(f"Received: {msg.value().decode('utf-8')}") finally: consumer.close() ``` ### Java example[​](#java-example "Direct link to Java example") Implement a custom `AuthenticateCallbackHandler` that calls `GetWebIdentityToken` and refreshes the token before it expires. Pass the handler class to your Kafka client configuration: ``` bootstrap.servers=: security.protocol=SASL_SSL sasl.mechanism=OAUTHBEARER sasl.jaas.config=org.apache.kafka.common.security.oauthbearer.OAuthBearerLoginModule required; sasl.login.callback.handler.class= ``` Configure your handler to: 1. Call the AWS SDK `StsClient.getWebIdentityToken()` with the configured audience and a signing algorithm. 2. Return the token and its expiration time. 3. Refresh the token before it expires to avoid authentication failures. ## Optional: Disable other authentication mechanisms[​](#optional-disable-other-authentication-mechanisms "Direct link to Optional: Disable other authentication mechanisms") To use only OAuth 2.0/OIDC authentication, you can now disable all other SASL authentication mechanisms. Replace `SERVICE_NAME` with your value: ``` avn service update SERVICE_NAME \ -c "kafka_sasl_mechanisms.plain=false" \ -c "kafka_sasl_mechanisms.scram_sha_256=false" \ -c "kafka_sasl_mechanisms.scram_sha_512=false" ``` Related pages * [Enable OAuth 2.0/OIDC authentication for Apache Kafka®](/docs/products/kafka/howto/enable-oidc.md) * [Enable and configure SASL authentication](/docs/products/kafka/howto/kafka-sasl-auth.md) --- # Configure Prometheus for Aiven for Apache Kafka® using Privatelink You can integrate Prometheus with your Aiven for Apache Kafka® service using Privatelink for secure monitoring. This setup uses a Privatelink load balancer, which allows for efficient service discovery of Apache Kafka nodes and enables you to connect to your Aiven for Apache Kafka service using a private endpoint in your network or VPCs. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before you start, ensure you have the following: * [Aiven for Apache Kafka®](/docs/products/kafka/get-started/create-kafka-service.md) service running. * [Prometheus integration](/docs/platform/howto/integrations/prometheus-metrics.md) set up for your Aiven for Apache Kafka for extracting metrics. * Necessary permissions to modify service configurations. ## Configuration steps[​](#configuration-steps "Direct link to Configuration steps") ### Basic configuration[​](#basic-configuration "Direct link to Basic configuration") Begin by configuring Prometheus to scrape metrics from your Aiven for Apache Kafka service. This setup involves specifying various parameters for secure data retrieval. Following is an example configuration: ``` scrape_configs: - job_name: aivenmetrics scheme: https tls_config: insecure_skip_verify: true basic_auth: username: password: http_sd_configs: - url: refresh_interval: 120s tls_config: insecure_skip_verify: true basic_auth: username: password: ``` **Configuration details**: * `job_name`: Identifies the set of targets, for example,, `aivenmetrics`. * `scheme`: Specifies the protocol, typically `https`. * `tls_config`: Manages TLS settings. note Setting `insecure_skip_verify: true` is crucial, as it permits Prometheus to disregard TLS certificate validation against host IP addresses, facilitating seamless connectivity. * `basic_auth`: Provides authentication credentials for Apache Kafka service access. * `http_sd_configs`: Configures HTTP Service Discovery. Includes: * `url`: The URI for Prometheus Privatelink service access. * `refresh_interval`: The frequency of target list refresh, for example,, `120s`. note The `basic_auth` and `tls_config` are specified twice - first for scraping the HTTP SD response and to retrieve service metrics. This duplication is necessary because the same authentication and security settings are used to retrieve the service discovery information and scrape the metrics. ### Optional: Metadata and relabeling[​](#optional-metadata-and-relabeling "Direct link to Optional: Metadata and relabeling") If your setup involves multiple Privatelink connections, you can leverage Prometheus's relabeling for better target management. This approach allows you to dynamically modify target label sets before scraping. To manage metrics from different Privatelink connections, include the `__meta_privatelink_connection_id` label in your configuration. This setup helps categorize and filter relevant metrics for each connection. ``` relabel_configs: - source_labels: [__meta_privatelink_connection_id] regex: 1 action: keep ``` The `regex: 1` in the configuration is a placeholder. Make sure to replace `1` with the actual Privatelink connection ID that you wish to monitor. Related pages * [Aiven for Apache Kafka® metrics available via Prometheus](/docs/products/kafka/reference/kafka-metrics-prometheus.md) --- # Connect Aiven for Apache Kafka® with Quix Connect your Aiven for Apache Kafka® service with Quix to consume the data and process it in real-time, and produce it back to Kafka via Quix Cloud. [Quix](https://quix.io?utm_source=aiven) is a complete platform for developing, deploying, and monitoring stream processing pipelines. You use the Quix Streams Python library to develop modular stream processing applications, and deploy them to containers managed in Quix with a single click. You can develop and manage applications on the command line or manage them in Quix Cloud and visualize them as a end-to-end pipeline. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") To connect Aiven for Apache Kafka® and Quix: * A running Aiven for Apache Kafka® service. See [Getting started with Aiven for Apache Kafka](/docs/products/kafka/get-started/create-kafka-service.md) for more information. * A CA certificate file for your Kafka instance. See [Use SASL authentication with Aiven for Apache Kafka®](/docs/products/kafka/howto/kafka-sasl-auth.md). * An account with Quix. If you don't yet have an account, you can [sign up for a free trial](https://portal.platform.quix.io/self-sign-up). ## Connect Aiven for Apache Kafka® to Quix[​](#connect-aiven-for-apache-kafka-to-quix "Direct link to Connect Aiven for Apache Kafka® to Quix") To configure and connect Aiven for Apache Kafka® with Klaw: 1. Log in to the **Quix Portal**. 2. Open an existing **Project** or create a project. For more details, see the **Create a test pipeline section** below. 3. Create an environment. * If you're editing an existing project, open the project settings and click **+ New environment**. ![New environment](/docs/assets/images/1_quix_new_env_menu-f53c8aed597ecc7b79af9c432f580a96.png) * Follow the setup wizard until you get to the broker settings. 4. When you get to the broker settings, select **Aiven** as your broken provider. ![Broker settings](/docs/assets/images/2_quix-aiven-broker-option-597f21e6230a2780f225d79a3ef3da33.png) 5. Configure the required settings: ![Broker setup](/docs/assets/images/3_quix-aiven-broker-settings-39139e427ec9dcb9aa342e7e9fa65d9c.png) * **Service URI**: Enter the Service URI for your Apache Kafka service. Find the service URI in the Connection information page of your service in Aiven Console * **Cluster Size**: This number is used to limit the replication factor of the topics. You'll find it in the Services overview section of your Aiven account * **User name**: Make sure that SASL is enabled in the Aiven Advanced Configuration section and copy the user name into this field. * **Password**: Likewise, copy the password from the Aiven Advanced Configuration section. * **SASL Mechanism**: Use the same SASL mechanism as defined in the Aiven Advanced Configuration section. * **CA Certificate**: Upload the CA file that you downloaded from the Aiven console. ## Create a test pipeline[​](#create-a-test-pipeline "Direct link to Create a test pipeline") To help you get started, the Quix platform includes several pipeline templates that you can deploy in a few clicks. To test your Aiven for Apache Kafka® connection, you can use the [*Hello Quix* template](https://quix.io/templates/hello-quix), which is a three-step pipeline: ![Screenshot of a pipeline](/docs/assets/images/4_helloquix-template-12e80b6528d50ebe96455cce15e027eb.png) 1. Click [**Clone this project**](https://portal.platform.quix.io/signup?projectName=Hello%20Quix\&httpsUrl=https://github.com/quixio/template-hello-quix\&branchName=tutorial). 2. On the **Import Project** screen, click **Quix advanced configuration** to ensure you get the option to configure own broker settings. 3. Follow the project creation wizard and configure your Aiven for Apache Kafka® connection details when prompted. 4. Click **Sync your pipeline**. ## Test the Setup[​](#test-the-setup "Direct link to Test the Setup") In the Quix portal, wait for the services to deploy and their status to become **Running**. ![Screenshot of a pipeline](/docs/assets/images/5_helloquix-pipeline-251fd56f23a8a3cb93666ae185a824db.png) Ensure the `_csv-data_` and `_counted-names_` required topics appear in both Quix and Aiven. In Aiven, topics that originate from Quix have the Quix workspace and project name as a prefix, such as `_quixdemo-helloquix-csv-data_`. Related pages * [Quix documentation](https://quix.io/docs/get-started/welcome) * [Quix guide to creating projects](https://quix.io/blog/how-to-create-a-project-from-a-template#cloning-a-project-template-into-github) * [Quix Streams Python library](https://github.com/quixio/quix-streams) --- # Enable and configure SASL authentication for Apache Kafka® Aiven for Apache Kafka® supports [multiple authentication methods](/docs/products/kafka/concepts/auth-types.md), including Simple Authentication and Security Layer ([SASL](https://en.wikipedia.org/wiki/Simple_Authentication_and_Security_Layer)) over SSL. ## Enable SASL authentication[​](#enable-sasl-authentication "Direct link to Enable SASL authentication") To allow clients to authenticate with SASL, enable `kafka_authentication_methods.sasl` on your Aiven for Apache Kafka service. * Aiven Console * CLI * API * Terraform 1. In the [Aiven Console](https://console.aiven.io), select your Aiven for Apache Kafka service. 2. Click **Service settings**. 3. Scroll to **Advanced configuration** and click **Configure**. 4. Click **Add configuration options**. 5. Select `kafka_authentication_methods.sasl` from the list and set the value to **Enabled**. 6. Click **Save configurations**. The **Connection information** on the **Overview** page now shows connection details for SASL and client certificate authentication. note SASL and client certificate connections use different ports. The host, CA, and user credentials remain the same. Enable SASL authentication for your Aiven for Apache Kafka service using [Aiven CLI](/docs/tools/cli.md): 1. Get the name of the Aiven for Apache Kafka service: ``` avn service list ``` Note the `SERVICE_NAME` corresponding to your Aiven for Apache Kafka service. 2. Enable SASL authentication: ``` avn service update SERVICE_NAME -c kafka_authentication_methods.sasl=true ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `kafka_authentication_methods.sasl`: Set to `true` to enable SASL authentication. Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) API to enable SASL authentication on an existing service: ``` curl -X PUT "https://api.aiven.io/v1/project/{project_name}/service/{service_name}" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "user_config": { "kafka_authentication_methods": { "sasl": true } } }' ``` Parameters: * `project_name`: Name of your Aiven project. * `service_name`: Name of your Aiven for Apache Kafka service. * `API_TOKEN`: Personal Aiven [token](/docs/platform/howto/create_authentication_token.md). * `kafka_authentication_methods.sasl`: Set to `true` to enable SASL authentication. Set the `kafka_authentication_methods.sasl` attribute in [your `aiven_kafka` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/kafka#kafka_authentication_methods) to `true`. ## Configure SASL mechanisms[​](#configure-sasl-mechanisms "Direct link to Configure SASL mechanisms") After [enabling SASL authentication](#enable-sasl-authentication), choose which SASL mechanisms clients can use. ### Supported mechanisms[​](#supported-mechanisms "Direct link to Supported mechanisms") Aiven for Apache Kafka supports the following SASL mechanisms: * **PLAIN**: Enabled by default. Controlled by `kafka_sasl_mechanisms.plain`. * **SCRAM-SHA-256**: Enabled by default. Controlled by `kafka_sasl_mechanisms.scram_sha_256`. * **SCRAM-SHA-512**: Enabled by default. Controlled by `kafka_sasl_mechanisms.scram_sha_512`. * **OAUTHBEARER**: Set `kafka.sasl_oauthbearer_jwks_endpoint_url` to enable [OAuth 2.0/OIDC authentication](/docs/products/kafka/howto/enable-oidc.md). `PLAIN`, `SCRAM-SHA-256`, and `SCRAM-SHA-512` remain enabled by default. Each client selects one SASL mechanism when it connects. To allow only OAuth 2.0/OIDC authentication, disable `kafka_sasl_mechanisms.plain`, `kafka_sasl_mechanisms.scram_sha_256`, and `kafka_sasl_mechanisms.scram_sha_512`. Optional OIDC parameters include `kafka.sasl_oauthbearer_expected_issuer`, `kafka.sasl_oauthbearer_expected_audience`, and `kafka.sasl_oauthbearer_sub_claim_name`. note For **SCRAM-SHA-256** and **SCRAM-SHA-512** with Apache Kafka® 4.x, use `librdkafka` version 2.6.1 or later. Clients built on earlier `librdkafka` versions can fail to connect to Apache Kafka 4.x services using SASL/SCRAM. This affects `librdkafka`-based clients, such as `confluent-kafka`, `confluent-kafka-go`, `Confluent.Kafka`, and `node-rdkafka`. note When SASL authentication is enabled, at least one SASL mechanism must be available. `OAUTHBEARER` satisfies this requirement when `kafka.sasl_oauthbearer_jwks_endpoint_url` is set. If you disable PLAIN, SCRAM-SHA-256, and SCRAM-SHA-512 without setting `kafka.sasl_oauthbearer_jwks_endpoint_url`, the update fails because no SASL mechanism is available. ### Enable or disable PLAIN, SCRAM-SHA-256, and SCRAM-SHA-512[​](#enable-or-disable-plain-scram-sha-256-and-scram-sha-512 "Direct link to Enable or disable PLAIN, SCRAM-SHA-256, and SCRAM-SHA-512") Use `kafka_sasl_mechanisms` to enable or disable these mechanisms using one of the following methods. * Aiven Console * CLI * API * Terraform 1. In the [Aiven Console](https://console.aiven.io), select your Aiven for Apache Kafka service. 2. Click **Service settings**. 3. Scroll to **Advanced configuration** and click **Configure**. 4. In the **Advanced configuration** window, configure **PLAIN**, **SCRAM-SHA-256**, and **SCRAM-SHA-512**: * To enable or disable **PLAIN**, set `kafka_sasl_mechanisms.plain` to **Enabled** or **Disabled**. * To enable or disable **SCRAM-SHA-256**, set `kafka_sasl_mechanisms.scram_sha_256` to **Enabled** or **Disabled**. * To enable or disable **SCRAM-SHA-512**, set `kafka_sasl_mechanisms.scram_sha_512` to **Enabled** or **Disabled**. 5. Click **Save configurations**. Configure SASL mechanisms for your Aiven for Apache Kafka service using [Aiven CLI](/docs/tools/cli.md): 1. Get the name of the Aiven for Apache Kafka service: ``` avn service list ``` Note the `SERVICE_NAME` corresponding to your Aiven for Apache Kafka service. 2. Disable the SASL mechanisms that clients do not use. By default, **PLAIN**, **SCRAM-SHA-256**, and **SCRAM-SHA-512** are enabled. For example, to disable **PLAIN** authentication: ``` avn service update SERVICE_NAME \ -c kafka_sasl_mechanisms.plain=false ``` **SCRAM-SHA-256** and **SCRAM-SHA-512** remain enabled unless you disable them. Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `kafka_sasl_mechanisms.plain`: Set to `true` or `false` to enable or disable **PLAIN**. * `kafka_sasl_mechanisms.scram_sha_256`: Set to `true` or `false` to enable or disable **SCRAM-SHA-256**. * `kafka_sasl_mechanisms.scram_sha_512`: Set to `true` or `false` to enable or disable **SCRAM-SHA-512**. Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) API to configure **PLAIN** and **SCRAM** mechanisms on an existing service: ``` curl -X PUT "https://api.aiven.io/v1/project/{project_name}/service/{service_name}" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "user_config": { "kafka_sasl_mechanisms": { "plain": false, "scram_sha_256": true, "scram_sha_512": true } } }' ``` Parameters: * `project_name`: Name of your Aiven project. * `service_name`: Name of your Aiven for Apache Kafka service. * `API_TOKEN`: Personal Aiven [token](/docs/platform/howto/create_authentication_token.md). * `kafka_sasl_mechanisms.plain`: Set to `true` or `false` to enable or disable **PLAIN**. * `kafka_sasl_mechanisms.scram_sha_256`: Set to `true` or `false` to enable or disable **SCRAM-SHA-256**. * `kafka_sasl_mechanisms.scram_sha_512`: Set to `true` or `false` to enable or disable **SCRAM-SHA-512**. Use the `kafka_sasl_mechanisms` attribute in [your `aiven_kafka` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/kafka#kafka_sasl_mechanisms) to enable or disable **PLAIN**, **SCRAM-SHA-256**, and **SCRAM-SHA-512**. Set each mechanism to `true` or `false`: * **PLAIN**: `kafka_sasl_mechanisms.plain` * **SCRAM-SHA-256**: `kafka_sasl_mechanisms.scram_sha_256` * **SCRAM-SHA-512**: `kafka_sasl_mechanisms.scram_sha_512` ## Enable public CA certificates for SASL authentication[​](#enable-public-ca-certificates-for-sasl-authentication "Direct link to Enable public CA certificates for SASL authentication") After you enable SASL authentication, you can enable public CA certificates for clients that cannot install or trust the default project CA. Enable public CA certificates using one of the following methods. * Aiven Console * CLI * API * Terraform 1. In the [Aiven Console](https://console.aiven.io), select your Aiven for Apache Kafka service. 2. Click **Service settings**. 3. Go to the **Cloud and network** section and click **Actions** > **More network configurations**. 4. In the **Network configuration** dialog: 1. Click **Add configuration options**. 2. Find `letsencrypt_sasl` (or `letsencrypt_sasl_privatelink` for PrivateLink). 3. Select the configuration option. 4. Set the value to **Enabled**. 5. Click **Save configurations**. The **Connection information** on the **Overview** page now supports SASL connections using either Project CA or Public CA. Enable the public CA certificates for SASL authentication using the [Aiven CLI](/docs/tools/cli.md): 1. List the services in your project to find your Aiven for Apache Kafka service name: ``` avn service list ``` Note the `SERVICE_NAME` corresponding to your Aiven for Apache Kafka service. 2. Enable public CA certificates for SASL authentication: ``` avn service update SERVICE_NAME -c letsencrypt_sasl=true ``` For PrivateLink, use `-c letsencrypt_sasl_privatelink=true` instead. Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) API to enable public CA certificates for SASL authentication on an existing service: ``` curl -X PUT "https://api.aiven.io/v1/project/{project_name}/service/{service_name}" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"user_config": {"letsencrypt_sasl": true}}' ``` For PrivateLink, use `letsencrypt_sasl_privatelink` instead of `letsencrypt_sasl`. Parameters: * `project_name`: Name of your Aiven project. * `service_name`: Name of your Aiven for Apache Kafka service. * `API_TOKEN`: Personal Aiven [token](/docs/platform/howto/create_authentication_token.md). * `letsencrypt_sasl`: Set to `true` to enable public CA certificates for SASL authentication. This Terraform example enables SASL authentication and public CA for SASL on a Kafka service. It configures SCRAM-SHA-256 and includes a data source that outputs the SASL port for connections that use the public CA. The complete example is available in the [Aiven Terraform Provider repository](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/kafka/kafka_sasl_authentication) on GitHub. ``` Loading... ``` To apply your Terraform configuration: 1. Initialize Terraform by running: ``` terraform init ``` The output is similar to the following: ``` Initializing the backend... Initializing provider plugins... - Finding aiven/aiven versions matching ">= 4.0.0, < 5.0.0"... - Installing aiven/aiven v4.9.2... - Installed aiven/aiven v4.9.2 ... Terraform has been successfully initialized! ... ``` 2. To create an execution plan and preview the changes, run: ``` terraform plan ``` 3. To deploy your changes, run: ``` terraform apply --auto-approve ``` note * The public certificate is issued and validated by [Let's Encrypt](https://letsencrypt.org), a widely trusted certification authority. For details, see [How it works](https://letsencrypt.org/how-it-works). * When enabling the public CA over a PrivateLink connection, network configuration may take several minutes before clients can connect. A new port must be allocated and the load balancer route table updated before clients can connect. Related pages * [Enable OAuth 2.0/OIDC authentication for Apache Kafka](/docs/products/kafka/howto/enable-oidc.md) * [Authentication types](/docs/products/kafka/concepts/auth-types.md) --- # Use Apache Kafka® Streams with Aiven for Apache Kafka® [Apache Kafka® Streams](https://kafka.apache.org/documentation/streams/) is a client-side library for building real-time applications where input and output data are stored in Kafka clusters. Kafka Streams enables you to build scalable, fault-tolerant applications that process data streams. It reads from one or more input sources (such as Kafka topics) and writes to a sink (such as an output Kafka topic). You write Kafka Streams applications in Java or Scala. The following example shows how to use Kafka Streams with Aiven for Apache Kafka® and [Karapace](https://karapace.io/) schema registry to filter [Apache Avro™](https://avro.apache.org/) messages. The example uses data from the Sample Data Generator for **Logistics**, which writes to the `logistics_data_gen` topic. The example code reads from that topic and writes filtered data to the `logistics_data_delivered` topic: * Writes messages where the `state` is `Delivered`. * Copies the `carrier` and `manifest` fields, and renames `time_utc` to `timeUtc` and `tracking_id` to `trackingId`. Other fields are not copied. note The Avro messages in this example use the [Confluent wire format](https://docs.confluent.io/platform/current/schema-registry/fundamentals/serdes-develop/index.html#wire-format). In this format, a schema ID is inserted before each message value. This format is sometimes referred to as `AvroConfluent`. The input message schema is retrieved from the schema registry. The output schema is defined in [logistics\_delivered.avsc](https://github.com/Aiven-Labs/kafka-streams-example/blob/main/app/src/main/avro/logistics_delivered.avsc), compiled into the Java application, and registered with the schema registry. ## Prerequisites[​](#kafka-streams-prereq "Direct link to Prerequisites") You can run this example with any Apache Kafka service. The steps below use an **Aiven for Apache Kafka®** service. ### Schema registry[​](#schema-registry "Direct link to Schema registry") This example requires a **schema registry**. The producer registers the Avro schema to obtain a schema ID, which is added to each message. The consumer retrieves the schema from the registry to decode the message. Enable the **Karapace schema registry** for the Kafka service. See [Enable Karapace schema registry](/docs/products/kafka/karapace/howto/enable-karapace.md). ### Environment variables[​](#environment-variables "Direct link to Environment variables") Create the following environment variables to connect to the Aiven for Apache Kafka and Karapace services: * `KAFKA_BOOTSTRAP_SERVERS`: Service URL of the Kafka service * `SCHEMA_REGISTRY_URL`: Service URI of the schema registry * `SCHEMA_REGISTRY_USERNAME`: Username for the schema registry * `SCHEMA_REGISTRY_PASSWORD`: Password for the schema registry tip You can find these values in **Connection information** on the **Overview** page in the [Aiven console](https://console.aiven.io/) or by running `avn service get` with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). You can also download the certificate files used in the next step from this section. ### Certificates[​](#certificates "Direct link to Certificates") Create a directory named `certs` and download the following files to this directory: * Access key (`service.key`) * Access certificate (`service.cert`) * CA certificate (`ca.pem`) ### Kafka topic[​](#kafka-topic "Direct link to Kafka topic") Create the output topic [`logistics_data_delivered`](/docs/products/kafka/howto/create-topic.md#create-an-apache-kafka-topic) in the Kafka service. The sample data generator automatically creates the input topic. ### Local tools[​](#local-tools "Direct link to Local tools") Install one of the following to run the example: * [Docker](https://www.docker.com/) to run the application in a container * [Gradle](https://gradle.org/) to build and run the application locally using the `run.sh` script ## Get the example application code[​](#get-the-example-application-code "Direct link to Get the example application code") 1. Clone the `kafka-streams-example` repository from GitHub: ``` git clone https://github.com/Aiven-Labs/kafka-streams-example.git ``` 2. Change into the repository directory: ``` cd kafka-streams-example ``` ## Start the Logistics data stream[​](#start-the-logistics-data-stream "Direct link to Start the Logistics data stream") Follow the instructions in [Stream sample data from the Aiven Console](/docs/products/kafka/howto/generate-sample-data.md) to start the **Logistics** data generator. ## Run the example application[​](#run-the-example-application "Direct link to Run the example application") Set the following environment variables: * `KAFKA_CA_CERT`: Contents of the `ca.pem` file * `SERVICE_CERT_CONTENTS`: Contents of the `service.cert` file * `KAFKA_ACCESS_KEY`: Contents of the `service.key` file Set these variables by sourcing the `prep_cert_env.sh` script in the cloned repository: ``` source prep_cert_env.sh ``` * Run with Docker * Build and run locally 1. Build the container image for the `GenericFilterApp` example: ``` docker build --build-arg APP_NAME=GenericFilterApp -t appimage . ``` 2. Run the container using the environment variables set earlier: ``` docker run -d --name kafka-streams-container -p 3000:3000 \ -e KAFKA_BOOTSTRAP_SERVERS=$KAFKA_BOOTSTRAP_SERVERS \ -e KAFKA_CA_CERT="$KAFKA_CA_CERT" \ -e SERVICE_CERT_CONTENTS="$SERVICE_CERT_CONTENTS" \ -e KAFKA_ACCESS_KEY="$KAFKA_ACCESS_KEY" \ -e SCHEMA_REGISTRY_URL=$SCHEMA_REGISTRY_URL \ -e SCHEMA_REGISTRY_USERNAME=$SCHEMA_REGISTRY_USERNAME \ -e SCHEMA_REGISTRY_PASSWORD=$SCHEMA_REGISTRY_PASSWORD \ appimage ``` Build and run the application locally: 1. Build the application. This command creates a **fat JAR**. ``` gradle GenericFilterAppUberJar ``` 2. Copy the JAR file to the current directory so the `run.sh` script can find it: ``` cp app/build/libs/GenericFilterApp-uber.jar . ``` 3. Run the application using the `run.sh` script: ``` APP_NAME=GenericFilterApp ./run.sh ``` The script uses the environment variables set earlier. tip The Docker image uses the same `run.sh` script. ## Check the produced data[​](#check-the-produced-data "Direct link to Check the produced data") * In the Aiven console * With Python 1. In the [Aiven console](https://console.aiven.io/), open the **Aiven for Apache Kafka®** service. 2. In the sidebar, click **Manage stream** > **Topics**. 3. Select the `logistics_data_delivered` topic. 4. Click **Messages**. 5. In **Format**, select `Avro`. 6. Click **Fetch messages**. The `reporting` directory contains the command-line program `report_messages.py`, which reads messages from the input and output topics and displays them in the terminal. Install [`uv`](https://docs.astral.sh/uv/) to run the script. After installing [`uv`](https://docs.astral.sh/uv/getting-started/installation/) and setting the environment variables from earlier steps, run the script: ``` reporting/report_messages.py ``` ## About the example code[​](#about-the-example-code "Direct link to About the example code") The example repository contains source code for several applications. This example focuses on [`GenericFilterApp.java`](https://github.com/Aiven-Labs/kafka-streams-example/blob/main/app/src/main/java/org/example/GenericFilterApp.java). note See the example repository [README](https://github.com/Aiven-Labs/kafka-streams-example/blob/main/README.md) for additional details and other sample programs. --- # Get started with tiered storage Aiven for Apache Kafka tiered storage optimizes resources by storing recent, frequently accessed data on faster local disks and moving less active data to more economical, slower storage. For a detailed understanding of tiered storage, its workings, and benefits, see [Tiered storage in Aiven for Apache Kafka®](/docs/products/kafka/concepts/kafka-tiered-storage.md). These steps apply only to Classic Kafka clusters. Tiered storage is optional and configurable in Classic Kafka. ## Step 1: Enable tiered storage for the service[​](#step-1-enable-tiered-storage-for-the-service "Direct link to Step 1: Enable tiered storage for the service") Enable tiered storage for your Aiven for Apache Kafka service to set up the necessary infrastructure. 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. Select your project and the Aiven for Apache Kafka service. 3. See [enable tiered storage](/docs/products/kafka/howto/enable-kafka-tiered-storage.md). note * Aiven for Apache Kafka® supports tiered storage starting from Apache Kafka® version 3.6 or later. Upgrade to the latest default version and apply [maintenance updates](/docs/products/kafka/howto/maintenance-updates.md) when using tiered storage for the latest fixes and improvements. * Tiered storage is not available on startup plans. ## Step 2: Configure tiered storage per topic[​](#step-2-configure-tiered-storage-per-topic "Direct link to Step 2: Configure tiered storage per topic") Configure tiered storage for specific topics to control your data storage. * Follow the instructions to [enable and configure tiered storage for topics](/docs/products/kafka/howto/configure-topic-tiered-storage.md). * On the **Manage stream** > **Topics** page, topics using tiered storage display **Active** in the **Tiered storage** column. ## Step 3: Monitor storage usage[​](#step-3-monitor-storage-usage "Direct link to Step 3: Monitor storage usage") Monitor storage usage for your Aiven for Apache Kafka® service. * Open **Observe** > **Tiered storage**. * Review billing, retention settings, and storage usage details. ## Related Pages[​](#related-pages "Direct link to Related Pages") * [Tiered storage in Aiven for Apache Kafka® overview](/docs/products/kafka/concepts/kafka-tiered-storage.md) * [Storage usage and settings in Aiven Console](/docs/products/kafka/howto/view-kafka-storage-in-console.md) --- # Configure properties for Apache Kafka® toolbox The [open source Apache Kafka® code](https://kafka.apache.org/downloads) includes a series of tools under the `bin` directory that can be useful to manage and interact with an Aiven for Apache Kafka® service. Before using the tools, configure a file pointing to a Java keystore and truststore which contain the required certificates for authentication. note There are no restrictions on the file name, but make sure to use the correct name when performing CLI operations. In the examples below we'll name the file `configuration.properties`. ## Define the configuration file[​](#define-the-configuration-file "Direct link to Define the configuration file") 1. Create the Java keystore and truststore for your Aiven for Apache Kafka® service using the [dedicated Aiven CLI command](/docs/tools/cli/service/user.md#avn_service_user_kafka_java_creds). 2. Create a `configuration.properties` file pointing to the keystore and truststore with the following entries: * `security.protocol`: security protocol, SSL for the default TLS security settings * `ssl.keystore.type`: keystore type, `PKCS12` for the keystore generated with the [dedicated Aiven CLI command](/docs/tools/cli/service/user.md#avn_service_user_kafka_java_creds) * `ssl.keystore.location`: keystore location on the file system * `ssl.keystore.password`: keystore password * `ssl.truststore.type`: truststore type * `ssl.truststore.location`: truststore location on the file system * `ssl.truststore.password`: truststore password * `ssl.key.password`: keystore password tip The `avn service user-kafka-java-creds` [Aiven CLI command](/docs/tools/cli/service/user.md#avn_service_user_kafka_java_creds) accepts a `--password` parameter setting the same password for the truststore, keystore and key An example of the `configuration.properties` content is the following: ``` security.protocol=SSL ssl.protocol=TLS ssl.keystore.type=PKCS12 ssl.keystore.location=client.keystore.p12 ssl.keystore.password=changeit ssl.key.password=changeit ssl.truststore.location=client.truststore.jks ssl.truststore.password=changeit ssl.truststore.type=JKS ``` --- # Use kcat with Aiven for Apache Kafka® The `kcat` [tool](https://github.com/edenhill/kcat) (formerly known as `kafkacat`) is a generic non-JVM producer and consumer for Apache Kafka®. It can be used to produce and consume records to Apache Kafka topics as well as to list service configurations. ## Install `kcat`[​](#install-kcat "Direct link to install-kcat") `kcat` is an open source tool available from GitHub at . Installation instructions are provided in the repository README. ## Retrieve Aiven for Apache Kafka® SSL certificate files[​](#retrieve-aiven-for-apache-kafka-ssl-certificate-files "Direct link to Retrieve Aiven for Apache Kafka® SSL certificate files") Aiven for Apache Kafka by default enables TLS security. The certificates can be manually downloaded from the service overview page in the Aiven console, or via the [dedicated Aiven CLI command](/docs/tools/cli/service/user.md#avn_service_user_creds_download). ## Setup a `kcat` configuration file[​](#setup-a-kcat-configuration-file "Direct link to setup-a-kcat-configuration-file") While `kcat` accepts all connection configuration parameters in the command line, using a configuration file helps minimising the code needed for any following calls. A `kcat` configuration file enabling the connection to an Aiven for Apache Kafka® service with TLS security must contain the following entries: * `bootstrap.servers`: Aiven for Apache Kafka® service URI, that can be found in the service overview in Aiven console * `security.protocol`: security protocol, SSL for the default TLS security settings * `ssl.key.location`: location of the `service.key` file downloaded from the service overview in Aiven console * `ssl.certificate.location`: location of the `service.cert` file downloaded from the service overview in Aiven console * `ssl.ca.location`: location of the `ca.pem` file downloaded from the service overview in Aiven console An example of the `kcat` configuration file is provided below: ``` bootstrap.servers=demo-kafka.my-demo-project.aivencloud.com:17072 security.protocol=ssl ssl.key.location=service.key ssl.certificate.location=service.cert ssl.ca.location=ca.pem ``` Once the content is stored in a file named `kcat.config`, this can be referenced using the `-F` flag: ``` kcat -F kcat.config ``` Alternatively, the same settings can be specified directly on the command line with: ``` kcat \ -b demo-kafka.my-demo-project.aivencloud.com:17072 \ -X security.protocol=ssl \ -X ssl.key.location=service.key \ -X ssl.certificate.location=service.cert \ -X ssl.ca.location=ca.pem ``` If [SASL authentication](/docs/products/kafka/howto/kafka-sasl-auth.md) is enabled, then the `kcat` configuration file requires the following entries: ``` bootstrap.servers=demo-kafka.my-demo-project.aivencloud.com:17072 ssl.ca.location=ca.pem security.protocol=SASL_SSL sasl.mechanisms=SCRAM-SHA-256 sasl.username=avnadmin sasl.password=yourpassword ``` tip If you're using Aiven for Apache Kafka®, you can retrieve the `kcat` command parameters using the dedicated [Aiven CLI command](/docs/tools/cli/service/connection-info.md#avn_cli_service_connection_info_kcat). ## Produce data to an Apache Kafka® topic[​](#produce-data-to-an-apache-kafka-topic "Direct link to Produce data to an Apache Kafka® topic") Use the following code to produce a single message into topic named `test-topic`: ``` echo test-message-content | kcat -F kcat.config -P -t test-topic -k test-message-key ``` * `-P`: sets the producer mode * `-t`: specifies the topic * `-k`: sets the message key The output of the above comment is a message sent to Apache Kafka® `test-topic` containing `test-message-key` as key and `test-message-content` as payload. > `kcat` can use a file as input input and specify a delimiter (`-D`) for splitting rows into individual records for bulk loading of data. ## Consume data from an Apache Kafka® topic[​](#consume-data-from-an-apache-kafka-topic "Direct link to Consume data from an Apache Kafka® topic") Use the following code to consume messages coming from a topic named `test-topic`: ``` kcat -F kcat.config -C -t test-topic -o -1 -e ``` * `-C`: sets for consumer mode * `-t`: specifies the topic again * `-o`: defines the topic starting message offset (negative values are considered relative to the latest offset) * `-e`: stops `kcat` once the end of the topic is reached; without it, it will continuously poll Apache Kafka for new messages. The above command retrieves the last message (`-o -1`) from the topic named `test-topic`. Consult the `kcat` helper (by adding the `-h` flag) for the full list of parameters. tip When consuming from a topic, setting the `-f` flag to `%t-%p: %o %S` returns the topic name, partition, offset and size for each message. --- # Configure Java SSL keystore and truststore to access Apache Kafka® Aiven for Apache Kafka® utilises TLS (SSL) to secure the traffic between its services and client applications. This means that clients must be configured with the right tools to be able to communicate with the Aiven services. Keystores and truststores are password-protected files accessible by the client that interacts with the service. To create these files: ## Access service certificates[​](#access-service-certificates "Direct link to Access service certificates") 1. Log in to [Aiven Console](https://console.aiven.io/) and select your Apache Kafka service. 2. On the **Overview** page, go to the **Connect information** panel. 3. Select **Client certificate** as the authentication method. 4. Download the **Access Key**, **Access Certificate**, and **CA Certificate**. The files *service.key*, *service.cert*, and *ca.pem* are necessary. ## Create the keystore[​](#create-the-keystore "Direct link to Create the keystore") * Use the `openssl` utility to create a keystore using the downloaded `service.key` and `service.cert`: ``` openssl pkcs12 -export \ -inkey service.key \ -in service.cert \ -out client.keystore.p12 \ -name service_key ``` note Ensure the keystore format is `PKCS12`, the default since Java 9. * Set a password for the keystore and key when prompted. ## Create the truststore[​](#create-the-truststore "Direct link to Create the truststore") * In the directory containing the certificates, use the `keytool` utility to create a truststore with the `ca.pem` file: ``` keytool -import \ -file ca.pem \ -alias CA \ -keystore client.truststore.jks ``` * When prompted, enter a password for the truststore and confirm trust in the CA certificate. ## Resulting configuration files[​](#resulting-configuration-files "Direct link to Resulting configuration files") The process generates two files: *client.keystore.p12* (keystore) and *client.truststore.jks* (truststore). These files are ready for client configuration. tip Use the [Aiven CLI](/docs/tools/cli.md) command `avn service user-kafka-java-creds` to automate keystore and truststore creation. For more information, see [`avn service user-kafka-java-creds`](/docs/tools/cli/service/user.md#avn_service_user_kafka_java_creds). --- # Use Kpow with Aiven for Apache Kafka® [Kpow by Factor House](https://factorhouse.io/products/kpow) is an enterprise solution for Kafka management and monitoring. Kpow supports **Aiven for Apache Kafka®**. Use Kpow to monitor, manage, and explore your managed Kafka brokers, Karapace Schema Registry, and Standalone Kafka Connect services. ## Cluster authentication[​](#cluster-authentication "Direct link to Cluster authentication") Aiven supports multiple authentication methods. Configure Kpow to connect using the method that matches your cluster security settings. info Aiven secures connections over TLS. Download the CA Certificate, `ca.pem`, from the Aiven Console and provide it to Kpow as your SSL Truststore. Kpow supports raw PEM files, so no `keytool` conversion to Java Keystores, `.jks`, is required. ### SASL/SCRAM[​](#saslscram "Direct link to SASL/SCRAM") Aiven supports SASL/SCRAM authentication. Find your username, password, and the specific Suggest revising to: \`Aiven supports SASL/SCRAM authentication. In the Aiven Console, find your username, password, and SASL port on the **Connection information** page. Set the following connection variables: ``` SECURITY_PROTOCOL=SASL_SSL SASL_MECHANISM=SCRAM-SHA-256 SASL_JAAS_CONFIG=org.apache.kafka.common.security.scram.ScramLoginModule required username="" password=""; SSL_TRUSTSTORE_LOCATION=/path/to/ca.pem SSL_TRUSTSTORE_TYPE=PEM ``` ### Mutual TLS[​](#mutual-tls "Direct link to Mutual TLS") Aiven utilizes mTLS for cluster authentication. Historically, connecting a Java-based Kafka client required using `keytool` to convert your downloaded certificates into a Java Aiven uses mTLS for cluster authentication. Java-based Kafka clients often require using `keytool` to convert downloaded certificates into Java keystore format. Kpow eliminates this hurdle by providing support for raw PEM files. Do not convert your certificates to JKS. Use the files exactly as they are downloaded from the Aiven Kpow supports raw PEM files directly, so you do not need to convert certificates to JKS. Use the files as downloaded from the Aiven Console: `ca.pem`, `service.cert`, and `service.key`. First, combine your access key and certificate into a single keystore file: ``` cat service.key service.cert > keystore.pem ``` Then, configure Kpow with the following properties: ``` SECURITY_PROTOCOL=SSL SSL_TRUSTSTORE_LOCATION=/path/to/ca.pem SSL_TRUSTSTORE_TYPE=PEM SSL_KEYSTORE_LOCATION=/path/to/keystore.pem SSL_KEYSTORE_TYPE=PEM ``` ### OAuth/OIDC[​](#oauthoidc "Direct link to OAuth/OIDC") Aiven supports [OpenID Connect](/docs/products/kafka/howto/enable-oidc.md) for identity federation. If your cluster is configured for OIDC, connect Kpow using standard Kafka Aiven supports [OpenID Connect](/docs/products/kafka/howto/enable-oidc.md) for identity federation. If your cluster uses OIDC, configure Kpow with standard Kafka OAuth properties: ``` SECURITY_PROTOCOL=SASL_SSL SASL_MECHANISM=OAUTHBEARER SASL_LOGIN_CALLBACK_HANDLER_CLASS=org.apache.kafka.common.security.oauthbearer.secured.OAuthBearerLoginCallbackHandler SASL_OAUTHBEARER_TOKEN_ENDPOINT_URL= SASL_JAAS_CONFIG=org.apache.kafka.common.security.oauthbearer.OAuthBearerLoginModule required clientId="" clientSecret=""; SSL_TRUSTSTORE_LOCATION=/path/to/ca.pem SSL_TRUSTSTORE_TYPE=PEM ``` ## Access control[​](#access-control "Direct link to Access control") Aiven uses **Apache Kafka ACLs** for data-plane authorization. After connecting Kpow using your chosen authentication method, configure the connecting user with the appropriate [Aiven ACLs](/docs/products/kafka/concepts/acl.md) to access After connecting Kpow, configure the connecting user with the appropriate [Aiven ACLs](/docs/products/kafka/concepts/acl.md) to access topics and consumer groups. Kpow provides support for managing these ACLs. See the [ACL management documentation](https://docs.factorhouse.io/kpow/management/acls) for details. ## Ecosystem integration[​](#ecosystem-integration "Direct link to Ecosystem integration") If you have provisioned related managed services in Aiven, integrate them into Kpow. ### Aiven Schema Registry[​](#aiven-schema-registry "Direct link to Aiven Schema Registry") Aiven provides a managed Schema Registry, Karapace, that integrates with Kpow. It relies Aiven provides Karapace Schema Registry, which integrates with Kpow. It uses Basic Authentication (`USER_INFO`). Configure Kpow with the following properties: ``` SCHEMA_REGISTRY_NAME=Aiven Schema Registry SCHEMA_REGISTRY_URL= SCHEMA_REGISTRY_AUTH=USER_INFO SCHEMA_REGISTRY_USER= SCHEMA_REGISTRY_PASSWORD= ``` ### Aiven Kafka Connect[​](#aiven-kafka-connect "Direct link to Aiven Kafka Connect") Aiven offers two types of Kafka Connect deployments: **Integrated** and **Standalone**. Because Kpow requires the standard Kafka Connect REST API to monitor and manage connectors, it only supports the **Standalone** Aiven Connect service. Aiven's integrated offering does not expose this public REST URL. Service integration required Creating a Standalone Connect service in Aiven is not enough. Link the Connect service to your Kafka service in the Aiven Console. Click **Manage stream** > **Integrations**. If you skip this step, Aiven returns a 503 Service Unavailable error, and Kpow fails to start. Configure your Standalone Connect cluster with the following properties: ``` CONNECT_NAME=Aiven Connect CONNECT_REST_URL= CONNECT_AUTH=BASIC CONNECT_BASIC_AUTH_USER= CONNECT_BASIC_AUTH_PASS= ``` ## Aiven considerations and limitations[​](#aiven-considerations-and-limitations "Direct link to Aiven considerations and limitations") ### Retention policy fix[​](#retention-policy-fix "Direct link to Retention policy fix") If you use an Aiven Free plan, or a paid plan with strict admin-configured retention limits, Kpow crashes on startup with a `PolicyViolationException`. This happens because Kpow attempts to auto-create its internal audit log topic, `__oprtr_audit_log`, with infinite retention, `retention.ms` set to `-1`, which Aiven rejects to prevent runaway disk usage. To fix this, manually pre-create the `__oprtr_audit_log` topic in your Aiven Console before launching Kpow. Set the Cleanup Policy to `delete` and the `retention_ms` to a finite number allowed by your plan limit, for example `259200000` for 3 days. Kpow detects the existing topic, bypasses auto-creation, and starts successfully. ## Quickstart[​](#quickstart "Direct link to Quickstart") This command starts a Kpow container configured to connect to Aiven using SASL/SCRAM, alongside the Schema Registry and Standalone Connect integrations. Because Kpow supports PEM files, map the downloaded `ca.pem` file directly into the Docker container. License requirements To run Kpow, provide your license details via the `LICENSE_` environment variables shown in the command below. Obtain these values from your welcome email or the Factor House license portal. If you do not have a license, [request a free 30-day trial](https://account.factorhouse.io/cta_action/provision_license_type?code=KPOW_TRIAL). ``` docker run -p 3000:3000 \ -v $(pwd)/ca.pem:/etc/kpow/ca.pem \ --env BOOTSTRAP=":17060" \ --env SECURITY_PROTOCOL="SASL_SSL" \ --env SASL_MECHANISM="SCRAM-SHA-256" \ --env SASL_JAAS_CONFIG='org.apache.kafka.common.security.scram.ScramLoginModule required username="" password="";' \ --env SSL_TRUSTSTORE_LOCATION="/etc/kpow/ca.pem" \ --env SSL_TRUSTSTORE_TYPE="PEM" \ --env SSL_ENDPOINT_IDENTIFICATION_ALGORITHM="" \ --env SCHEMA_REGISTRY_NAME="Aiven Schema Registry" \ --env SCHEMA_REGISTRY_URL="" \ --env SCHEMA_REGISTRY_AUTH="USER_INFO" \ --env SCHEMA_REGISTRY_USER="" \ --env SCHEMA_REGISTRY_PASSWORD="" \ --env CONNECT_NAME="Aiven Connect" \ --env CONNECT_REST_URL="" \ --env CONNECT_AUTH="BASIC" \ --env CONNECT_BASIC_AUTH_USER="" \ --env CONNECT_BASIC_AUTH_PASS="" \ --env LICENSE_ID="" \ --env LICENSE_CODE="" \ --env LICENSEE="" \ --env LICENSE_EXPIRY="" \ --env LICENSE_SIGNATURE="" \ factorhouse/kpow:latest ``` info These steps do not configure Kpow authorization. To restrict user actions in Kpow, see [Simple Access Control](https://factorhouse.io/community) on the Factor House site. After the container starts, open `http://localhost:3000` to access the Kpow UI. ![Kpow - Aiven](/docs/assets/images/kpow-3de36df0e160ae59af69174eda973beb.png) Kpow Community To explore Kafka locally or for non-commercial use, [grab a free Community license](https://factorhouse.io/community) and use the `factorhouse/kpow-ce` Docker image. --- # Use ksqlDB with Aiven for Apache Kafka® Aiven provides a managed Apache Kafka® solution together with a number of auxiliary services like Apache Kafka Connect, Kafka REST and Schema Registry via [Karapace](https://github.com/aiven/karapace). A managed [ksqlDB](https://ksqldb.io/) service in Aiven is, however, not supported. To define streaming data pipelines with SQL, you can: * Use [Aiven for Apache Flink®](/docs/products/flink.md) or, * Run a self-hosted ksqlDB cluster. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") To connect ksqlDB to Aiven for Apache Kafka®, create a [Java keystore and truststore containing the service SSL certificates](/docs/products/kafka/howto/keystore-truststore.md). Also collect the following information: * `APACHE_KAFKA_HOST`: The Aiven for Apache Kafka hostname * `APACHE_KAFKA_PORT`: The Aiven for Apache Kafka port * `SCHEMA_REGISTRY_PORT`: The Aiven for Apache Kafka schema registry port, if enabled * `SCHEMA_REGISTRY_PASSWORD`: The password associated with the `avnadmin` user for Schema Registry * `KEYSTORE_FILE_NAME`: The name of the Java keystore containing the Aiven for Apache Kafka SSL certificates * `TRUSTSTORE_FILE_NAME`: The name of the Java truststore containing the Aiven for Apache Kafka SSL certificates * `SSL_KEYSTORE_PASSWORD`: The password used to secure the Java keystore * `SSL_KEY_PASSWORD`: The password used to secure the Java key * `SSL_TRUSTSTORE_PASSWORD`: The password used to secure the Java truststore * `SSL_STORE_FOLDER`: The absolute path of the folder containing both the truststore and keystore * `TRUSTSTORE_SCHEMA_REGISTRY_FILE_NAME`: The name of the Java truststore containing the schema registry ([Karapace](https://karapace.io/)) certificate * `TRUSTSTORE_SCHEMA_REGISTRY_PASSWORD`: The password used to secure the Java truststore for the schema registry ([Karapace](https://karapace.io/)) certificate ### Create a keystore for schema registry's ca file[​](#create-a-keystore-for-schema-registrys-ca-file "Direct link to Create a keystore for schema registry's ca file") ksqlDB by default uses the `ssl.truststore` settings for the Schema Registry connection. To have ksqlDB working with Aiven's [Karapace](https://karapace.io/) Schema Registry, explicitly define a truststore that contains the commonly trusted root CA of Schema Registry server. To create such a truststore: 1. Obtain the root CA of the server with the following `openssl` command by replacing the `APACHE_KAFKA_HOST` and `SCHEMA_REGISTRY_PORT` placeholders: ``` openssl s_client -connect APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT \ -showcerts < /dev/null 2>/dev/null | \ awk '/BEGIN CERT/{s=1}; s{t=t "\n" $0}; /END CERT/ {last=t; t=""; s=0}; END{print last}' \ > ca_schema_registry.cert ``` 2. Create the truststore with the following `keytool` command by replacing the `TRUSTSTORE_SCHEMA_REGISTRY_FILE_NAME` and `TRUSTSTORE_SCHEMA_REGISTRY_PASSWORD` placeholders: ``` keytool -import -file ca_schema_registry.cert \ -alias CA \ -keystore TRUSTSTORE_SCHEMA_REGISTRY_FILE_NAME \ -storepass TRUSTSTORE_SCHEMA_REGISTRY_PASSWORD \ -noprompt ``` tip The `TRUSTSTORE_SCHEMA_REGISTRY_FILE_NAME` can be any name but the extension should be `.jks` ## Run ksqlDB on Docker[​](#run-ksqldb-on-docker "Direct link to Run ksqlDB on Docker") You can run ksqlDB on Docker with the following command, by replacing the placeholders: * `SSL_STORE_FOLDER` * `APACHE_KAFKA_HOST` * `APACHE_KAFKA_PORT` * `KEYSTORE_FILE_NAME` * `SSL_KEYSTORE_PASSWORD` * `SSL_KEY_PASSWORD` * `TRUSTSTORE_FILE_NAME` * `SSL_TRUSTSTORE_PASSWORD` * `SCHEMA_REGISTRY_PORT` * `SCHEMA_REGISTRY_PASSWORD` * `TRUSTSTORE_SCHEMA_REGISTRY_FILE_NAME` * `TRUSTSTORE_SCHEMA_REGISTRY_PASSWORD` ``` docker run -d --name ksql \ -v SSL_STORE_FOLDER/:/ssl_settings/ \ -p 127.0.0.1:8088:8088 \ -e KSQL_BOOTSTRAP_SERVERS=APACHE_KAFKA_HOST:APACHE_KAFKA_PORT \ -e KSQL_LISTENERS=http://0.0.0.0:8088/ \ -e KSQL_KSQL_SERVICE_ID=ksql_service_1_ \ -e KSQL_OPTS="-Dsecurity.protocol=SSL -Dssl.keystore.type=PKCS12 -Dssl.keystore.location=/ssl_settings/KEYSTORE_FILE_NAME -Dssl.keystore.password=SSL_KEYSTORE_PASSWORD -Dssl.key.password=SSL_KEY_PASSWORD -Dssl.truststore.type=JKS -Dssl.truststore.location=/ssl_settings/TRUSTSTORE_FILE_NAME -Dssl.truststore.password=SSL_TRUSTSTORE_PASSWORD -Dksql.schema.registry.url=APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT -Dksql.schema.registry.basic.auth.credentials.source=USER_INFO -Dksql.schema.registry.basic.auth.user.info=avnadmin:SCHEMA_REGISTRY_PASSWORD -Dksql.schema.registry.ssl.truststore.location=/ssl_settings/TRUSTSTORE_SCHEMA_REGISTRY_FILE_NAME -Dksql.schema.registry.ssl.truststore.password=TRUSTSTORE_SCHEMA_REGISTRY_PASSWORD" \ confluentinc/ksqldb-server:0.23.1 ``` tip `USER_INFO` is **not** a placeholder, but rather a literal that shouldn't be changed warning Some docker setups have issues using the `-v` mounting options. In those cases copying the Keystore and Truststore in the container can be an easier option. This can be achieved with the following: ``` docker container create --name ksql \ -p 127.0.0.1:8088:8088 \ -e KSQL_BOOTSTRAP_SERVERS=APACHE_KAFKA_HOST:APACHE_KAFKA_PORT \ -e KSQL_LISTENERS=http://0.0.0.0:8088/ \ -e KSQL_KSQL_SERVICE_ID=ksql_service_1_ \ -e KSQL_OPTS="-Dsecurity.protocol=SSL -Dssl.keystore.type=PKCS12 -Dssl.keystore.location=/home/appuser/KEYSTORE_FILE_NAME -Dssl.keystore.password=SSL_KEYSTORE_PASSWORD -Dssl.key.password=SSL_KEY_PASSWORD -Dssl.truststore.type=JKS -Dssl.truststore.location=/home/appuser/TRUSTSTORE_FILE_NAME -Dssl.truststore.password=SSL_TRUSTSTORE_PASSWORD -Dksql.schema.registry.url=APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT -Dksql.schema.registry.basic.auth.credentials.source=USER_INFO -Dksql.schema.registry.basic.auth.user.info=avnadmin:SCHEMA_REGISTRY_PASSWORD -Dksql.schema.registry.ssl.truststore.location=/home/appuser/TRUSTSTORE_SCHEMA_REGISTRY_FILE_NAME -Dksql.schema.registry.ssl.truststore.password=TRUSTSTORE_SCHEMA_REGISTRY_PASSWORD" \ confluentinc/ksqldb-server:0.23.1 docker cp KEYSTORE_FILE_NAME ksql:/home/appuser/ docker cp TRUSTSTORE_FILE_NAME ksql:/home/appuser/ docker cp TRUSTSTORE_SCHEMA_REGISTRY_FILE_NAME ksql:/home/appuser/ docker start ksql ``` Once the Docker image is up and running you should be able to access ksqlDB at `localhost:8088` or connect via terminal with the following command: ``` docker exec -it ksql ksql ``` --- ## [Connection methods](/docs/products/kafka/howto/connect-with-python.md) [11 items](/docs/products/kafka/howto/connect-with-python.md) --- # Maintenance and updates for your Aiven for Apache Kafka® service Manage maintenance updates and set the maintenance window for your Aiven for Apache Kafka® service. ## Maintenance updates[​](#maintenance-updates "Direct link to Maintenance updates") Aiven applies some maintenance updates automatically. The following are the types of updates: * **Mandatory updates:** Security updates, quarterly patch releases, and platform updates that affect reliability or stability of the service nodes. * **Optional updates:** All other updates are initially optional. After six months, they become mandatory and are applied in the next week’s maintenance window. * **Periodic infrastructure updates:** Scheduled automatically for services with nodes active for 180 days and more. These updates are mandatory for all services, except those with maintenance turned off. Critical security updates are applied during the next available maintenance window. For other updates, Aiven gives you at least seven days' notice. Maintenance updates are also automatically applied during service upgrades. To view pending updates: * Console * CLI * API 1. In your service, click **Service settings**. 2. Go to the **Service management** section. Use the [`avn service get`](/docs/tools/cli/service-cli.md#avn_service_get) command. Use the [`service`](https://api.aiven.io/doc/#tag/Service/operation/ServiceGet) endpoint. note When Aiven releases a mandatory service update for Apache Kafka®, the [Kafka upgrade procedure](/docs/products/kafka/concepts/upgrade-procedure.md) runs automatically. ## Maintenance window[​](#maintenance-window "Direct link to Maintenance window") The maintenance window is the time period when Aiven can automatically apply maintenance updates to a service. When an update becomes available, Aiven schedules it for the next available maintenance window for each service. The update runs in the first window after it becomes available, and can begin any time after the start time. For example, if a service has a maintenance window of Monday 12:00 UTC, and an update becomes available on Tuesday, the update will be applied on the following Monday. During maintenance, Aiven might restart or replace service nodes. This can cause brief connection interruptions, but services are designed to minimize downtime. Aiven performs maintenance in a rolling-forward style, creating new nodes alongside existing ones and retiring the old nodes after the upgrade completes. Major service upgrades are triggered manually. A manually triggered upgrade starts immediately, regardless of the maintenance window. important You cannot control the order in which services are updated. Each service updates according to its own configured maintenance window, and there is no guaranteed way to control the update sequence. Manual updates and maintenance window adjustments only help for non-critical updates. ## Set the maintenance window[​](#set-the-maintenance-window "Direct link to Set the maintenance window") To set the maintenance window for your service: * Console * Terraform 1. In the Aiven Console, open your service. 2. In the **Maintenance** section, click **Actions** > **Change maintenance window**. 3. Set the day and time. 4. Click **Save changes**. Use the `maintenance_window_dow` and `maintenance_window_time` attributes in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Kafka upgrade procedure](/docs/products/kafka/concepts/upgrade-procedure.md) * [Change the service plan](/docs/products/kafka/howto/change-service-plan.md) --- # Manage access control lists in Aiven for Apache Kafka® Access control lists (ACLs) in Aiven for Apache Kafka® define permissions for topics, schemas, consumer groups, and transactional IDs. ACLs control which authenticated users or applications (principals) can perform specific operations on these resources. ## Types of ACLs[​](#types-of-acls "Direct link to Types of ACLs") Aiven for Apache Kafka supports two types of ACLs: * **Aiven ACLs**: These provide topic-level permissions and support wildcard patterns. * **Kafka-native ACLs**: These offer advanced, resource-level permissions with `ALLOW` and `DENY` rules for operations on multiple resource types, including topics, groups, and clusters. important If both Aiven ACLs and Kafka-native ACLs apply to the same principal and resource, and one allows access while the other denies it, the `DENY` rule takes precedence. This applies even if the `DENY` is defined in only one of the ACL types. note ACL restrictions for Kafka REST are controlled by a user configuration parameter in the service's advanced configuration settings. By default, ACLs do not apply to Kafka REST. To enable ACLs for Kafka REST, set the `kafka_rest_authorization` parameter. For more information, see [Enable Kafka REST Proxy Authorization](https://aiven.io/docs/products/kafka/karapace/howto/enable-kafka-rest-proxy-authorization). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Aiven for Apache Kafka service](/docs/products/kafka/get-started/create-kafka-service.md) running. * Access to the [Aiven Console](https://console.aiven.io/). * Installed and authenticated [Aiven CLI](/docs/tools/cli.md). * An [API token](https://aiven.io/docs/platform/howto/create_authentication_token) for authenticating API requests. * [Aiven Provider for Terraform](/docs/tools/terraform.md). ## Add a Kafka-native ACL entry[​](#add-a-kafka-native-acl-entry "Direct link to Add a Kafka-native ACL entry") * Aiven Console * Aiven CLI * Aiven API * Terraform 1. Log in to [Aiven Console](https://console.aiven.io/) and select your service. 2. Click **Access & Control** > **ACL**. 3. Click **Add entry**. 4. On the **Add access control entry** screen: 1. Select **Kafka-native ACLs** as the ACL type. 2. Fill in the following fields: 1. **Permission type**: Select `ALLOW` or `DENY`. 2. **Principal**: Enter the principal in the format `User:`. 3. **Operation**: Select the operation, such as `Read` or `Write`. 4. **Resource type**: Select the Apache Kafka resource to manage. 5. **Pattern type**: Select `LITERAL` for exact matches or `PREFIXED` for pattern-based matches. 6. **Resource**: Enter the resource name or a prefix for pattern-based matching. 7. **Host**: Enter the allowed host, or use `*` for all. 3. Click **Submit**. To add an Kafka-native ACL entry using the Aiven CLI, run: ``` avn service kafka-acl-add \ --principal \ --operation \ [--topic | --cluster | --group | --transactional-id ] \ --resource-pattern-type \ [--host ] \ [--deny] ``` Parameters: * `service_name`: Enter the name of your Aiven for Apache Kafka service. * `--principal`: Enter the principal in the format `User:`. * `--operation`: Enter the Apache Kafka operation, such as `Read`, `Write`, `Describe`, `Delete`, or any supported operation. * Resource-specific parameters, specify at least one of the following based on your requirements: * `--topic`: Enter the topic resource. * `--group`: Enter the consumer group resource. * `--cluster`: Specify that the ACL applies to the Kafka cluster. * `--transactional-id`: Enter the transactional ID resource. * `--resource-pattern-type`: Specify the resource pattern type. Use `LITERAL` for exact matches or `PREFIXED` for pattern-based matches (default: `LITERAL`). * `--host` Optional: Specify the allowed host. Use `*` to allow all hosts. * `--deny` Optional: Add this flag to create a `DENY` rule. The default is `ALLOW`. **Example:** Allow the `User:analyst` to read from all topics with names starting with `logs-`: ``` avn service kafka-acl-add kafka-service \ --principal User:analyst \ --operation Read \ --topic logs-* \ --resource-pattern-type PREFIXED ``` To add a Kafka-native ACL entry, use the following API request: ``` curl --request POST \ --url https://api.aiven.io/v1/project//service//kafka/acl \ --header 'Authorization: Bearer ' \ --header 'Content-Type: application/json' \ --data '{ "principal": "User:", "host": "", "resource_type": "", "resource_name": "", "pattern_type": "", "operation": "", "permission_type": "" }' ``` Parameter: * `project_name`: Enter the name of your Aiven project. * `service_name`: Enter the name of your Aiven for Apache Kafka service. * `permission_type`: Specify `ALLOW` or `DENY`. * `principal`: Provide the principal in the format `User:`. * `operation`: Specify the Kafka operation, such as `Read`, `Write`, `Describe`, `Delete`. * `resource_type`: Specify the resource type, such as `Cluster`, `Topic`, `Group`, or `TransactionalId`. * `resource_name`: Provide the resource name or prefix for pattern-based matching. * `pattern_type`: Specify `LITERAL` for an exact match or `PREFIXED` for pattern-based matching. * `host`: Specify the allowed host or use `*` to match all hosts. ``` Loading... ``` More information on this resource and its configuration options are available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/kafka_native_acl). ## Add an Aiven ACL entry[​](#add-an-aiven-acl-entry "Direct link to Add an Aiven ACL entry") note When you authenticate with OAuth 2.0/OIDC, an Aiven ACL entry only takes effect when a [matching Aiven service user](/docs/products/kafka/howto/add-manage-service-users.md) also exists. Use the OIDC principal identifier, such as the service principal's GUID, as the service user username. Kafka-native ACLs do not have this requirement. For more information, see [Enable OAuth 2.0/OIDC authentication](/docs/products/kafka/howto/enable-oidc.md). * Aiven Console * Aiven CLI * Aiven API * Terraform 1. Log in to [Aiven Console](https://console.aiven.io/) and select your service. 2. Click **Access & Control** > **ACL**. 3. Click **Add entry**. 4. On the **Add access control entry** screen: 1. Select **Aiven ACLs** as the ACL type. 2. Fill in the following fields: 1. **Resource type**: Select `Topic` or `Schema`. 2. **Permission type**: Select `admin`, `read`, `write`, or `readwrite`. 3. **Username**: Enter the username or pattern to apply the ACL to. Supports wildcards `*` and `?`. 4. **Resource**: Enter the name of the topic or schema, or use `*` to apply to all. 3. Click **Submit**. tip After defining custom ACLs, delete the default `avnadmin` ACL entry by clicking **Delete ACL** under **Actions** to prevent unintended access via wildcard permissions. To add an Aiven ACL entry using the Aiven CLI, run: ``` avn service acl-add \ --username \ --permission \ --topic ``` Parameters: * ``: Enter the name of your Aiven for Apache Kafka service. * `--username`: Enter the username or pattern to apply. Supports wildcards `*` and `?`. * `--permission`: Enter the permission type. Valid values are `read`, `write`, or `readwrite`. * `--topic`: Enter the topic name or pattern to apply. Supports wildcards `*` and `?`. **Example:** Allow the username pattern `developer*` to have `read` permissions for all topics with names starting with `logs-`: ``` avn service acl-add kafka-service \ --username developer* \ --permission read \ --topic logs-* ``` To add an Aiven ACL entry for Apache Kafka topics, use the following API request: ``` curl --request POST \ --url https://api.aiven.io/v1/project//service//acl \ --header 'Authorization: Bearer ' \ --header 'Content-Type: application/json' \ --data '{ "username": "", "resource": "", "permission": "" }' ``` Parameters: * `project_name`: Enter the name of your Aiven project. * `service_name`: Enter the name of your Aiven for Apache Kafka service. * `username`: Enter the username or pattern to apply the ACL to. Use `*` and `?` as wildcards. * `resource`: Enter the name of the topic or schema. Use `*` to apply to all, or include patterns with wildcards. * `permission`: Enter the permission type, such as `read`, `write`, `readwrite`, or `admin`. note Schema-related ACLs control access to schemas in the schema registry. These are configured separately from topic-based ACLs. To configure schema-related ACLs, use the schema registry-specific configuration endpoint: `/service//schema-registry/acl`. For more information, see the [Schema ACL definition](https://aiven.io/docs/products/kafka/karapace/concepts/acl-definition) or the [Schema Registry ACL API documentation](https://api.aiven.io/doc/#tag/Service:_Kafka/operation/ServiceSchemaRegistryAclAdd). ``` Loading... ``` More information on this resource and its configuration options are available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/kafka_native_acl). tip In [the `aiven_kafka` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/kafka), set the `default_acl` attribute to `false` to prevent the creation of an admin user with wildcard permissions. ## View ACL entries[​](#view-acl-entries "Direct link to View ACL entries") * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/) and select your Aiven for Apache Kafka service. 2. Click **Access & Control** > **ACL**. 3. Click the **Kafka-native ACLs** tab to view Kafka-native ACL entries or the **Aiven ACLs** tab to view Aiven ACL entries. 4. Use filters to narrow the list by resource type, operation, or permission type. To view ACL entries using the Aiven CLI: * **Kafka-native ACLs**: ``` avn service kafka-acl-list ``` * **Aiven ACLs**: ``` avn service acl-list ``` Replace `service_name` with the name of your Aiven for Apache Kafka service. To view ACL entries, use the following API request: * **Kafka-native ACLs:** ``` curl -X GET \ https://api.aiven.io/v1/project//service//kafka/acl \ -H 'Authorization: Bearer ' ``` * **Aiven ACLs:** ``` curl -X GET \ https://api.aiven.io/v1/project//service//acl \ -H 'Authorization: Bearer ' ``` Parameters: * `project`: Enter the name of your Aiven project. * `service_name`: Enter the name of your Aiven for Apache Kafka service. ## Delete ACL entries[​](#delete-acl-entries "Direct link to Delete ACL entries") * Aiven Console * Aiven CLI * API 1. Log in to the [Aiven Console](https://console.aiven.io/) and select your Aiven for Kafka service. 2. Click **Access & Control** > **ACL** 3. Click the **Kafka-native ACLs** tab to view Kafka-native ACL entries or the **Aiven ACLs** tab to view Aiven ACL entries. 4. Locate the ACL entry to delete. 5. Click **Delete ACL** under the **Actions** column to remove the entry. 6. Click **Delete**. To delete an ACL entry, use one of the following commands based on the ACL type: * **Kafka-native ACLs**: ``` avn service kafka-acl-delete ``` * **Aiven ACLs**: ``` avn service acl-delete ``` Parameters * ``: Enter the name of the Aiven for Apache Kafka service. * ``: Enter the ID of the ACL entry to delete. You can get the `acl_id` from the output when [viewing ACL entries](#view-acl-entries). To delete ACL entries, use the following API request: * **Kafka-native ACLs:** ``` curl --request DELETE \ --url https://api.aiven.io/v1/project//service//kafka/acl/ \ --header 'Authorization: Bearer ' ``` * **Aiven ACLs:** ``` curl -X DELETE \ https://api.aiven.io/v1/project//service//acl/ \ -H 'Authorization: Bearer ' ``` Parameters: * `project_name`: Enter the name of the project. * `service_name`: Enter the name of the Aiven for Apache Kafka service. * `acl_id` or `kafka_acl_id`: Enter the ID of the ACL entry to delete. Related pages * [Access Control Lists in Aiven for Apache Kafka®](/docs/products/kafka/concepts/acl.md) * [Manage service users in Aiven for Apache Kafka®](/docs/products/kafka/howto/add-manage-service-users.md) * [Apache Kafka documentation](https://kafka.apache.org/42/security/authorization-and-acls/#operations-and-resources-on-protocols) --- # Manage quotas in Aiven for Apache Kafka® Manage quotas in your Aiven for Apache Kafka® service to control network throughput and CPU usage per client. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An [Aiven for Apache Kafka® service](/docs/products/kafka/get-started/create-kafka-service.md). * Access to the [Aiven Console](https://console.aiven.io/) or an authenticated [Aiven CLI](/docs/tools/cli.md). * A [personal access token](/docs/platform/howto/create_authentication_token.md) for API requests. ## Add quota[​](#add-quota "Direct link to Add quota") * Aiven Console * Aiven CLI * Aiven API 1. Log in to [Aiven Console](https://console.aiven.io/) and select your Kafka service. 2. Click **Access & Control** > **Quotas** in the sidebar. 3. Click **Add quota**. 4. Enter the **Client ID** or **User**. 5. Set one or more quota values: * **Consumer throttle**: Maximum data rate for consumers, in bytes per second. * **Producer throttle**: Maximum data rate for producers, in bytes per second. * **CPU throttle**: Maximum CPU usage for the client, as a percentage. note Enter `default` to apply the quota to all clients or users. 6. Click **Add**. Use [`avn service quota create`](/docs/tools/cli/service/quota.md) with at least one of `--client-id` or `--user`, and at least one quota parameter. **Example:** Set a 1 MiB/s producer and consumer throttle for user `alice`: ``` avn service quota create kafka-doc \ --user alice \ --consumer-byte-rate 1048576 \ --producer-byte-rate 1048576 ``` **Example:** Set a default quota for all users: ``` avn service quota create kafka-doc \ --user default \ --consumer-byte-rate 5242880 ``` Send a request to the [ServiceKafkaQuotaCreate](https://api.aiven.io/doc/#tag/Service:_Kafka/operation/ServiceKafkaQuotaCreate) endpoint: ``` curl -X POST https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/quota \ -H "Authorization: Bearer $AIVEN_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "user": "alice", "consumer_byte_rate": 1048576, "producer_byte_rate": 1048576 }' ``` Replace `PROJECT_NAME` and `SERVICE_NAME` with your Aiven project name and Kafka service name. ## View quotas[​](#view-quotas "Direct link to View quotas") * Aiven Console * Aiven CLI * Aiven API 1. Access the [Aiven Console](https://console.aiven.io/) and open your Aiven for Apache Kafka® service. 2. Click **Access & Control** > **Quotas**. The page lists all configured quotas. ``` avn service quota list kafka-doc ``` Send a request to the [ServiceKafkaQuotaList](https://api.aiven.io/doc/#tag/Service:_Kafka/operation/ServiceKafkaQuotaList) endpoint: ``` curl https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/quota \ -H "Authorization: Bearer $AIVEN_TOKEN" ``` Replace `PROJECT_NAME` and `SERVICE_NAME` with your Aiven project name and Kafka service name. ## Update quota[​](#update-quota "Direct link to Update quota") * Aiven Console * Aiven CLI * Aiven API 1. Access the [Aiven Console](https://console.aiven.io/) and open your Aiven for Apache Kafka® service. 2. Click **Access & Control** > **Quotas**. 3. Find the quota to update. 4. Click **Actions** > **Update**. 5. Update the quota values. 6. Click **Save changes**. Run `avn service quota create` again with the same `--user` or `--client-id` to overwrite existing values: ``` avn service quota create kafka-doc \ --user alice \ --consumer-byte-rate 2097152 ``` Send another request to the [ServiceKafkaQuotaCreate](https://api.aiven.io/doc/#tag/Service:_Kafka/operation/ServiceKafkaQuotaCreate) endpoint with the updated values to overwrite the existing quota. ## Delete quota[​](#delete-quota "Direct link to Delete quota") * Aiven Console * Aiven CLI * Aiven API 1. Access the [Aiven Console](https://console.aiven.io/) and open your Aiven for Apache Kafka® service. 2. Click **Access & Control** > **Quotas**. 3. Find the quota to delete. 4. Click **Actions** > **Delete**. 5. Click **Delete quota** to confirm. ``` avn service quota delete kafka-doc --user alice ``` Send a request to the [ServiceKafkaQuotaDelete](https://api.aiven.io/doc/#tag/Service:_Kafka/operation/ServiceKafkaQuotaDelete) endpoint: ``` curl -X DELETE "https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/quota?user=alice" \ -H "Authorization: Bearer $AIVEN_TOKEN" ``` Replace `PROJECT_NAME` and `SERVICE_NAME` with your Aiven project name and Kafka service name. Related pages * [Quotas in Aiven for Apache Kafka®](/docs/products/kafka/concepts/kafka-quotas.md) * [`avn service quota`](/docs/tools/cli/service/quota.md) --- # Manage Aiven for Apache Kafka® resource requests [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Streamline the creation and ownership requests of Apache Kafka resources to enhance governance and ensure efficient management within Aiven for Apache Kafka. ## [Request access](/docs/products/kafka/howto/request-access-topic.md) [Request access to an Apache Kafka topic in Aiven for Apache Kafka Governance to produce or consume messages using access control lists (ACLs).](/docs/products/kafka/howto/request-access-topic.md) ## [Manage approvals](/docs/products/kafka/howto/approvals.md) [The Approvals page allows you to manage requests for Aiven for Apache Kafka® resources owned by your group.](/docs/products/kafka/howto/approvals.md) ## [Manage group requests](/docs/products/kafka/howto/group-requests.md) [The Group requests page allows you to view and track requests you and other members of your group made for Aiven for Apache Kafka® resources.](/docs/products/kafka/howto/group-requests.md) ## [Rotate credentials](/docs/products/kafka/howto/rotate-credentials.md) [Rotate credentials for an Apache Kafka® subscription to replace outdated credentials and maintain secure, approved access.](/docs/products/kafka/howto/rotate-credentials.md) --- # Manage Apache Kafka® topics in detail [Early availability](/docs/platform/concepts/service-and-feature-releases.md) Explore advanced topic management features in the Apache Kafka topic catalog. Navigate through various options and configurations for your Apache Kafka topics. ## Access topic overview page[​](#access-topic-overview-page "Direct link to Access topic overview page") 1. In the Apache Kafka topic catalog, click the name of a topic to open the topic details pane. 2. Click **Open topic page** to open the topic overview page. note If you are not a member of the project associated with the topic, the **Open topic page** option is disabled. On the topic overview page, you can: ### Overview tab[​](#overview-tab "Direct link to Overview tab") View a summary of the key details of your Apache Kafka topic: * **Topic size**: Total size of the topic data * **Partitions**: Number of partitions the topic is divided into * **Replication factor**: Replication factor set for the topic * **Topic tags**: Tags associated with the topic * **Topic owner**: Current owner of the topic ### Configuration tab[​](#configuration-tab "Direct link to Configuration tab") Review and compare the configuration settings of your Apache Kafka topics: * **Configuration details**: Shows all the configuration parameters set for the topic. * **Compare configurations**: Compare the configuration of the current topic with another topic by selecting a topic to compare. ### Users tab[​](#users-tab "Direct link to Users tab") Manage and view user permissions for the topic: * **Search for users**: Find users by username. * **Filter by permission**: Filter users by permissions. * **Compare user permissions**: Compare the permissions of users across different topics. ### Messages tab[​](#messages-tab "Direct link to Messages tab") Fetch and view messages: * Use filters like partition, offset, timeout, max bytes, and format. * Click **Fetch messages**. ### Schemas tab[​](#schemas-tab "Direct link to Schemas tab") Enable and manage schemas for the topic: * **Subject naming strategy**: Choose a naming strategy for subjects. * **Search schemas**: Find schemas by name. * **Set compatibility level**: Configure the compatibility level for your schemas. * **Delete subjects**: Remove schemas that are no longer needed. note Enable the [schema registry authorization](/docs/products/kafka/karapace/concepts/schema-registry-authorization.md) for your Aiven for Apache Kafka service to use the **Schemas** tab. ## Edit topic information[​](#edit-topic-information "Direct link to Edit topic information") 1. Click **Edit topic** to access the topic info screen within the respective Aiven for Apache Kafka service. 2. On the topic info screen, click **Modify** to edit the topic configurations. 3. Update the desired configurations and click **Update** to save the changes. 4. To return to the topic overview page in the Apache Kafka topic catalog, click **Open topic catalog**. Related pages * [Aiven for Apache Kafka® topic catalog overview](/docs/products/kafka/concepts/topic-catalog-overview.md) * [View and manage Apache Kafka topic catalog](/docs/products/kafka/howto/view-kafka-topic-catalog.md) * [Query Kafka topic data in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/query-kafka-topic-data.md) --- # Monitor and alert logs for denied ACL Aiven for Apache Kafka® uses access control lists (ACL) and user definitions in order to establish individual rights to produce or consume a topic. For more information on ACLs permission mapping, see [Access control lists and permission mapping](/docs/products/kafka/concepts/acl.md) section. In cases of ACLs problems, an error `io.aiven.kafka.auth.AivenAclAuthorizer` is generated. You can also use the following log patterns to set up alerts for checking failed authentication and ACL evaluation. ## Failed producer[​](#failed-producer "Direct link to Failed producer") A producer creates the following log in case the client has no privilege to write to a specific topic: ``` HOSTNAME: kafka-pi-3141592-75 SYSTEMD_UNIT: kafka.service MESSAGE: [2020-09-04 06:35:33,509] INFO [DENY] Auth request Write on Topic:nodejs-quickstart-kafka-topic by User test-kuser (io.aiven.kafka.auth.AivenAclAuthorizer) ``` ## Failed consumer[​](#failed-consumer "Direct link to Failed consumer") A consumer creates the following log in case the client has no privilege to describe a specific topic: ``` HOSTNAME: kafka-pi-3141592-74 SYSTEMD_UNIT: kafka.service MESSAGE: [2020-09-04 06:43:09,712] INFO [DENY] Auth request Describe on Topic:nodejs-quickstart-kafka-topic by User test-kuser (io.aiven.kafka.auth.AivenAclAuthorizer) ``` ## Valid certificate with invalid key[​](#valid-cert-with-invalid-key "Direct link to Valid certificate with invalid key") A client creates the following log when using a valid certificate with an invalid key to perform a describe operation over a topic: ``` HOSTNAME: kafka-pi-3141592-75 SYSTEMD_UNIT: kafka.service MESSAGE: [2020-09-04 06:54:10,781] INFO [DENY] Auth request Describe on Topic:nodejs-quickstart-kafka-topic by Invalid CN=delete-user,OU=u6l6y9h1,O=kafka-pi-3141592 (io.aiven.kafka.auth.AivenAclAuthorizer) ``` --- # Optimizing resource usage for Aiven for Apache Kafka® Aiven for Apache Kafka® service plans with CPUs of 2 or less are optimized for lightweight operations, making them suitable for applications that handle fewer messages per second and do not require high throughput. However, you might sometimes encounter an alert showing high resource usage. These alerts typically arise when the Kafka broker memory drops too low, and the CPU idle time is less than 15%. Understanding the reasons behind these alerts and their mitigation ensures optimized Kafka usage and consistent application performance. ## What triggers high resource usage[​](#what-triggers-high-resource-usage "Direct link to What triggers high resource usage") Several factors can lead to high resource usage across Aiven for Apache Kafka plans: * **High Kafka Traffic:** Heavy traffic due to too many producer/consumer requests can cause an overload, leading to increased CPU and memory usage on the Kafka broker. When a Kafka cluster is overloaded, it may struggle to correctly assign leadership for a partition, potentially causing service disruptions. * **Excessive Kafka Partitions:** An excessive number of Kafka partitions for the brokers to manage effectively can lead to increased memory usage and IO load. * **Too many client connections:** When there are too many client connections, the memory usage can dip significantly. When your service's memory is low, it starts to use swap space, adding to the IO load. Regular use of swap space indicates your system may not have enough resources for its workload. ## Additional causes of high resource usage[​](#additional-causes-of-high-resource-usage "Direct link to Additional causes of high resource usage") * **Datadog Integration:** Datadog offers essential monitoring, but its agent is IO-intensive. Especially on plans not designed for high IO operations, Datadog can notably increase the IO load. The load from Datadog also increases with the number of topic partitions in your Kafka service. * **Karapace Integration:** Karapace provides a REST API for Kafka, but this can consume a substantial amount of memory and contribute to a high load when REST API is used in high demand. ## Strategies to minimize resource usage[​](#strategies-to-minimize-resource-usage "Direct link to Strategies to minimize resource usage") * **Reduce Topic Partition Limit:** Decreasing the number of topic partitions reduces the load on the Kafka service. * **Disable Datadog Integration:** If Datadog sends too many metrics, it can affect the reliability of the service and hinder the backup of topic configurations. It is recommended to turn off the integration of the Datadog service. * **Enable Quotas:** Quotas can manage the resources consumed by clients, preventing any single client from using too much of the broker's resources. * **Limit the Number of Integrations:** For smaller plans such as the Startup-2, consider limiting the number of integrations to manage resource consumption effectively. * **Upgrade Your Plan:** If your application demands more resources, upgrading to a larger Kafka plan can ensure stable operation. ## Integration advisory for Kafka Startup-2 plan[​](#integration-advisory-for-kafka-startup-2-plan "Direct link to Integration advisory for Kafka Startup-2 plan") Aiven for Apache Kafka service plans with CPUs of 2 or less operate on relatively small machines. Enabling integrations like Datadog or Karapace might exceed the resources these plans offer, impacting your cluster's performance. If you experience issues with your cluster or require more resources for your integrations, consider upgrading to a higher plan. --- # Power on/off and delete your Aiven for Apache Kafka® service Power off your Aiven for Apache Kafka® service to release resources and save credits, power it back on when you need it, or delete it permanently. ## Power off a service[​](#power-off-a-service "Direct link to Power off a service") When you power off a service: * All virtual machines are removed from the public cloud. * The service configuration is stored on the Aiven Platform. * If there are no backups, all service data is lost. * If the service has time-based or point in time recovery backups, the backups remain on the Aiven Platform. Services powered off for more than 180 days are automatically deleted. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power off service**. To power off a service, run: ``` avn service update SERVICE_NAME --power-off ``` ## Power on a service[​](#power-on-a-service "Direct link to Power on a service") When you power on a service: * New virtual machines are created on the service's public cloud. * The service starts with the stored configuration parameters. * The latest time-based backup is restored. * Maintenance updates are automatically applied. * If a point in time recovery backup is available, the database transaction logs are replayed to recover the service data to a specific point in time. The restoration takes from a few minutes to a few hours, depending on the network bandwidth, the disk IOPS allocated to the service, and the size of the backup. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power on service**. To power on a service, run: ``` avn service update SERVICE_NAME --power-on ``` To see when the service is running, run: ``` avn service wait SERVICE_NAME ``` note When you power on an Aiven for Apache Kafka service, Aiven restores [configuration backups](/docs/products/kafka/concepts/configuration-backup.md) from the most recent backup. Configuration backups do not include classic topic data, consumer groups, or offsets. [Diskless topic](/docs/products/kafka/diskless/concepts/diskless-topic-overview.md) data remains in object storage and is available after you power the service on. important If the service uses [tiered storage](/docs/products/kafka/concepts/kafka-tiered-storage.md), powering off the service permanently deletes all remote data. note Static IP addresses are not removed when a service is powered off or deleted. They continue to generate the usual costs. To avoid these costs, [remove the static IP addresses](/docs/platform/concepts/static-ips.md). ## Delete a service[​](#delete-a-service "Direct link to Delete a service") * Console * CLI 1. In your project, click **Services**. 2. Open the service to delete, and click **Actions** > **Delete service**. To delete a service, run: ``` avn service terminate SERVICE_NAME ``` Related pages * [Configuration backups for Aiven for Apache Kafka®](/docs/products/kafka/concepts/configuration-backup.md) * [Trade-offs and limitations](/docs/products/kafka/concepts/tiered-storage-limitations.md) --- # Prevent full disks Ensure your Aiven for Apache Kafka® services run smoothly by preventing low disk space. The Aiven platform actively monitors disk usage, triggering notifications when it exceeds 90%. If any node in the service surpasses the critical threshold of disk usage (more than 95%), the access control list (ACL) used to authorize API requests by Apache Kafka clients is updated on all nodes. This update prevents operations that can further increase disk usage, including: * The `Write` and `IdempotentWrite` operations are used by clients to produce new messages. * The `CreateTopics` operation creates new topics, each of which carries some overhead on disk. When the disk space is insufficient and the ACL blocks write operations, you encounter an error. For example, when using the Python client for Apache Kafka, you might receive the following error message: ``` TopicAuthorizationFailedError: [Error 29] TopicAuthorizationFailedError: your-topic ``` ## Upgrade to a larger service plan[​](#upgrade-to-a-larger-service-plan "Direct link to Upgrade to a larger service plan") * Console * CLI To resolve disk space issues, you can upgrade to a larger service plan: 1. In the [Aiven Console](https://console.aiven.io/), select your project and choose your Aiven for Apache Kafka® service. 2. On the sidebar, Click **Service settings**. 3. In the **Service plan** section, click **Change plan**. 4. Choose the new service plan or use the slider to adjust disk storage. 5. Click **Change plan**. This deploys new nodes with increased disk space. After data migration from the old nodes to the new ones, disk usage returns to an acceptable level, and write operations are allowed again. To upgrade your service plan using the [Aiven CLI](/docs/tools/cli.md): ``` avn service update --project --service_name --plan --disk-space-gib ``` Parameters: * ``: The name of the project. * ``: The name of the Aiven for Apache service to update. * `--plan `: The Aiven subscription plan name. Refer to [`avn_service_plan`](/docs/tools/cli/service-cli.md#avn-service-plan) for available plans. * `--disk-space-gib `: The total amount of disk space for data storage (in GiB). ## Manage storage usage and settings[​](#manage-storage-usage-and-settings "Direct link to Manage storage usage and settings") See [storage usage and settings](/docs/products/kafka/howto/view-kafka-storage-in-console.md). To add or remove disk on Classic Kafka, see [Scale disk storage](/docs/products/kafka/howto/scale-disk-storage.md). ## Delete one or more topics[​](#delete-one-or-more-topics "Direct link to Delete one or more topics") * Console * CLI * API Free up disk space by deleting topics: 1. In the [Aiven Console](https://console.aiven.io/), select your project and choose your Aiven for Apache Kafka® service. 2. On the sidebar, click **Manage stream** > **Topics**. 3. To delete an existing topic, click the topic to remove. * In the **Topic info** screen, click **Delete** and confirm the deletion. * Alternatively, in the topic row, click the **Actions** > **Delete topic**. To delete topics using the [Aiven CLI](/docs/tools/cli.md): ``` avn service topic-delete --service_name SERVICE_NAME --topic TOPIC_NAME ``` Parameters: * `SERVICE_NAME`: The name of your Aiven for Apache Kafka service. * `TOPIC_NAME`: The name of the topic to delete. To delete topics using the Aiven API: ``` curl -X DELETE \ "https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/topic/TOPIC_NAME" \ -H "Authorization: Bearer YOUR_API_TOKEN" ``` Parameters: * `PROJECT_NAME`: The name of your project. * `SERVICE_NAME`: The name of your Aiven you for Apache Kafka service. * `TOPIC_NAME`: The name of the topic to delete. * `API_TOKEN`: Your [Aiven token](/docs/platform/concepts/authentication-tokens.md). Deleting topics frees up the disk space they used. The log cleaner process can take a few minutes to remove the associated data files from the disk. Once complete, the access control list (ACL) updates to allow write operations. note [Admin](/docs/platform/concepts/permissions.md) access is required to perform this action. ## Decrease retention time/size[​](#decrease-retention-timesize "Direct link to Decrease retention time/size") To free up space without deleting a topic, consider reducing the retention time or size for one or more topics. If you know the age of the oldest messages in a topic, you can lower the retention time to make more space available. For more details, see [how to change the retention period](/docs/products/kafka/howto/change-retention-period.md). --- # Use Provectus® UI for Apache Kafka® with Aiven for Apache Kafka® [Provectus® UI for Apache Kafka®](https://github.com/provectus/kafka-ui) is a popular Open-Source web GUI for Apache Kafka® management that allows you to monitor and manage Apache Kafka® clusters. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") To connect Provectus® UI for Apache Kafka® to Aiven for Apache Kafka®, create a [Java keystore and truststore containing the service SSL certificates](/docs/products/kafka/howto/keystore-truststore.md). Also collect the following information: * `APACHE_KAFKA_HOST`: The Aiven for Apache Kafka hostname * `APACHE_KAFKA_PORT`: The Aiven for Apache Kafka port * `SSL_KEYSTORE_FILE_NAME`: The name of the Java keystore containing the Aiven for Apache Kafka SSL certificates * `SSL_TRUSTSTORE_FILE_NAME`: The name of the Java truststore containing the Aiven for Apache Kafka SSL certificates * `SSL_KEYSTORE_PASSWORD`: The password used to secure the Java keystore * `SSL_TRUSTSTORE_PASSWORD`: The password used to secure the Java truststore * `SSL_STORE_FOLDER`: The absolute path of the folder containing both the truststore and keystore ## Run Provectus® UI for Apache Kafka® on Docker or Podman[​](#run-provectus-ui-for-apache-kafka-on-docker-or-podman "Direct link to Run Provectus® UI for Apache Kafka® on Docker or Podman") ### Share keystores with non-root user[​](#share-keystores-with-non-root-user "Direct link to Share keystores with non-root user") Since container for Provectus® UI for Apache Kafka® uses non-root user, to avoid permission problems, while keeping the secrets safe: 1. Create separate directory for secrets: ``` mkdir SSL_STORE_FOLDER ``` 2. Restrict the directory to current user: ``` chmod 700 SSL_STORE_FOLDER ``` 3. Copy secrets there (replace the `SSL_KEYSTORE_FILE_NAME` and `SSL_TRUSTSTORE_FILE_NAME` with the keystores and truststores file names): ``` cp SSL_KEYSTORE_FILE_NAME SSL_TRUSTSTORE_FILE_NAME SSL_STORE_FOLDER ``` 4. Give read permissions for secret files for everyone: ``` chmod +r SSL_STORE_FOLDER/* ``` ### Execute Provectus® UI for Apache Kafka® on Docker or Podman[​](#execute-provectus-ui-for-apache-kafka-on-docker-or-podman "Direct link to Execute Provectus® UI for Apache Kafka® on Docker or Podman") You can run Provectus® UI for Apache Kafka® in a Docker/Podman container with the following command, by replacing the placeholders for: * `APACHE_KAFKA_HOST` * `APACHE_KAFKA_PORT` * `SSL_STORE_FOLDER` * `SSL_KEYSTORE_FILE_NAME` * `SSL_KEYSTORE_PASSWORD` * `SSL_TRUSTSTORE_FILE_NAME` * `SSL_TRUSTSTORE_PASSWORD` ``` docker run -p 8080:8080 \ -v SSL_STORE_FOLDER/SSL_TRUSTSTORE_FILE_NAME:/client.truststore.jks:ro \ -v SSL_STORE_FOLDER/SSL_KEYSTORE_FILE_NAME:/client.keystore.p12:ro \ -e KAFKA_CLUSTERS_0_BOOTSTRAPSERVERS=APACHE_KAFKA_HOST:APACHE_KAFKA_PORT \ -e KAFKA_CLUSTERS_0_PROPERTIES_SECURITY_PROTOCOL=SSL \ -e KAFKA_CLUSTERS_0_PROPERTIES_SSL_TRUSTSTORE_LOCATION=/client.truststore.jks \ -e KAFKA_CLUSTERS_0_PROPERTIES_SSL_TRUSTSTORE_PASSWORD=SSL_TRUSTSTORE_PASSWORD \ -e KAFKA_CLUSTERS_0_PROPERTIES_SSL_KEYSTORE_LOCATION=/client.keystore.p12 \ -e KAFKA_CLUSTERS_0_PROPERTIES_SSL_KEYSTORE_PASSWORD=SSL_KEYSTORE_PASSWORD \ -e KAFKA_CLUSTERS_0_PROPERTIES_SSL_KEYSTORE_TYPE=PKCS12 \ -d provectuslabs/kafka-ui:latest ``` ## Use Provectus® UI for Apache Kafka®[​](#use-provectus-ui-for-apache-kafka "Direct link to Use Provectus® UI for Apache Kafka®") Once Provectus® UI for Apache Kafka® starts, you should be able to access it at `localhost:8080`. ![Provectus in action](/docs/assets/images/provectus-ui-afe22048ae5e8c0fa596d7cfd1f5dd62.jpg) --- # Renew and acknowledge service user SSL certificates Aiven for Apache Kafka® automatically generates a new SSL certificate for service users about three months before the existing certificate's expiration date. This new certificate includes a renewed private key. ## SSL certificate renewal schedule[​](#ssl-certificate-renewal-schedule "Direct link to SSL certificate renewal schedule") SSL certificates for Aiven for Apache Kafka® services are valid for 820 days, approximately two years, and three months. This renewal involves regenerating the SSL certificate and its private key to enhance security. Renewal notifications are sent to project administrators, operators, and technical contacts. The current certificate stays valid until expiration to ensure a smooth transition. ## Download the new SSL certificates[​](#download-the-new-ssl-certificates "Direct link to Download the new SSL certificates") Once renewed, you can download the new SSL certificate from the [Aiven Console](https://console.aiven.io/), [Aiven API](https://api.aiven.io/doc/), or [Aiven CLI](/docs/tools/cli.md). If your Aiven for Apache Kafka service has a certificate about to expire, the [Aiven Console](https://console.aiven.io/) will display a notification on the service page, prompting you to download the new certificate. To download the new certificate, 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka service. 3. Click **Access & Control** > **Users** in the sidebar. 4. Select the required user and click **Show access key** and **Show access cert** to download the new certificate. note You can also use the Aiven CLI command [`avn service user-creds-download`](/docs/tools/cli/service/user.md#avn_service_user_creds_download) to download the renewed SSL certificate and key. ## Acknowledge new SSL certificate usage[​](#acknowledge-new-ssl-certificate-usage "Direct link to Acknowledge new SSL certificate usage") Confirm that the new certificate is in use to stop receiving notifications about certificate expiration. To acknowledge the new SSL certificate with the [Aiven Console](https://console.aiven.io/): * Select `...` next to the certificate. * Select `Acknowledge certificate`. note You can also use the Aiven CLI command [`avn service user-creds-acknowledge`](/docs/tools/cli/service/user.md#avn_service_user_creds_acknowledge) to acknowledge the user credentials. Similarly, the Aiven API provides a way to acknowledge the new SSL certificate through the [Modify service user credentials endpoint](https://api.aiven.io/doc/#operation/ServiceUserCredentialsModify): ``` curl --request PUT \ --url https://api.aiven.io/v1/project//service//user/ \ --header 'Authorization: Bearer ' \ --header 'content-type: application/json' \ --data '{"operation": "acknowledge-renewal"}' ``` ## Turn off certificate expiration notifications for SASL services[​](#turn-off-certificate-expiration-notifications-for-sasl-services "Direct link to Turn off certificate expiration notifications for SASL services") When using SASL authentication in Aiven for Kafka services, you might still receive certificate expiration notifications, even if your service doesn't use certificates for authorization. Aiven updates certificates across all services to maintain security standards, which includes services that combine TLS encryption with SASL authentication. To turn off these notifications: 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka service. 3. Click **Service settings** from the sidebar. 4. Scroll to **Advanced configurations**, and click **Configure**. 5. Click **Add configuration options**. 6. Search for `kafka_authentication_methods.certificate` and disable it. --- # Request access to an Apache Kafka topic Request access to an Apache Kafka topic in Aiven for Apache Kafka Governance to produce or consume messages using access control lists (ACLs). ## How access requests work[​](#how-access-requests-work "Direct link to How access requests work") When you request access to an Apache Kafka topic, the following happens: * A **service user** is created to authenticate and authorize access to the topic. * A [Kafka-native ACL](/docs/products/kafka/concepts/acl.md#kafka-native-acl-capabilities) is created to define the permissions. * The request goes through an approval process before the credentials are available. You can view the service user and ACLs in the following locations in the [Aiven Console](https://console.aiven.io/): * Select your **Aiven for Apache Kafka** service. In the sidebar, click **Access & Control** > **ACL** or **Access & Control** > **Users**. * Click **Tools** > **Apache Kafka governance operations**. In the sidebar, click **Streaming catalog** > **Access**. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Console * Terraform [Governance](/docs/products/kafka/howto/enable-governance.md) enabled for your organization * [Terraform installed](https://www.terraform.io/downloads) * A [personal token](https://docs.aiven.io/docs/platform/howto/create_authentication_token.html) * A GitHub repository with [approval workflows configured](/docs/products/kafka/howto/terraform-governance-approvals.md) * To use beta features of the Aiven Provider for Terraform, set: ``` export PROVIDER_AIVEN_ENABLE_BETA=1 ``` ## Request access to a topic[​](#request-access-to-a-topic "Direct link to Request access to a topic") * Console * Terraform 1. In the [Aiven console](https://console.aiven.io/), click **Tools** > **Apache Kafka governance operations**. 2. In the sidebar, click **Streaming catalog** > **Topics**. 3. Click the topic you need access to. 4. In the Topic details panel, click **Request access**. 5. Fill in the **Request access** form: * **Project and service**: Auto-populated based on the selected topic. * **Service user**: Enter a username. If left blank, a name is generated automatically. * **Purpose description**: Describe the purpose of this service user. * **Access control list (ACL)**: * **Pattern type**: Auto-populated as **Literal**. note Only the **Literal** pattern type is supported. **Prefix** will be available later. * **Topic**: Auto-populated from the selected topic. * **Permission type**: Auto-populated as **Allow**. * **Operation**: Select **Read** or **Write**. * **Host**: Enter an IP address or use `*` to allow access from any host. * Optional: Click **Add another ACL** to define multiple ACLs. * **Approval information**: * **Service user owner**: Select the responsible team. * **Message for approval**: Provide details for review. 6. Click **Submit**. After submitting: * The request is sent for approval. To check the status, go to the [Group requests](/docs/products/kafka/howto/group-requests.md) page under **Governance operations**. * If approved, you can view and download the credentials for authentication in **Streaming catalog** > **Access overview**. Use the [`aiven_governance_access` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/governance_access) to request access to an Apache Kafka topic. The request is reviewed and approved in a GitHub pull request before access is granted. ### How it works[​](#how-it-works "Direct link to How it works") 1. Define the request: * Use Terraform to define the service user, topic, and the required access control lists (ACLs). * Specify the `owner_user_group_id` to indicate the group responsible for approving the request. **Example Terraform configuration:** ``` resource "aiven_governance_access" "example" { organization_id = data.aiven_organization.main.id access_name = "example-topic-access" access_type = "KAFKA" access_data { project = data.aiven_project.main.project service_name = aiven_kafka.main.service_name acls { resource_name = "example-topic" resource_type = "Topic" operation = "Read" permission_type = "ALLOW" host = "*" } } owner_user_group_id = aiven_organization_user_group.example.group_id } ``` * Commit and push the configuration to a GitHub repository with governance approval workflows enabled. 2. Review and approve the request: * The request appears as a pull request in GitHub. * A GitHub Action checks the request: * The requester must belong to the group defined by `owner_user_group_id`. * An approval must come from another member of the same group. note To verify group membership, GitHub user IDs must be mapped to Aiven user IDs using the [`aiven_external_identity`](/docs/products/kafka/howto/terraform-governance-approvals.md#step-2-map-github-users-to-aiven-identities) resource. * If the request meets all governance rules, the workflow applies the configuration using `terraform apply`. * After the request is approved, Aiven creates the service user, applies the ACLs to the specified topic, and generates the credentials. 3. Download the credentials: After access is provisioned, download the credentials from the [Aiven Console](https://console.aiven.io/). For more details, see [View and download service user credentials](#view-and-download-credentials). note Credentials are not available in Terraform or GitHub Actions output. ## View and download credentials[​](#view-and-download-credentials "Direct link to View and download credentials") After the request is approved, you can view and download the credentials for the service user. ### Why credentials can be viewed once[​](#why-credentials-can-be-viewed-once "Direct link to Why credentials can be viewed once") For security reasons, access certificates and access keys are shown only once to limit exposure and prevent unauthorized access. To access credentials later or perform tasks like resetting credentials, go to the **Aiven for Apache Kafka** service page > **Access & Control** > **Users**. For more information, see [Manage service users](/docs/products/kafka/howto/add-manage-service-users.md#manage-users). This approach: * Prevents storing sensitive credentials in plain text, reducing the risk of unauthorized access. * Encourages secure storage, as users must save access certificates and keys immediately after viewing them. * Future updates will further improve credential security. ### Steps to view and download credentials[​](#steps-to-view-and-download-credentials "Direct link to Steps to view and download credentials") 1. Access the [Aiven console](https://console.aiven.io/) and go to **Tools > Apache Kafka governance operations**. 2. In the sidebar, click **Streaming catalog** > **Access overview**. 3. In the **Access overview** page, locate the service user for which you need credentials. 4. Click **Actions** > **View credentials**. 5. On the confirmation window, click **Show credentials**. warning * Credentials can only be viewed once and only by members of the service owner group. * Once credentials are viewed, they cannot be retrieved again from the **Access overview** page. 6. In the **Save service user credentials** window, click **Show** to reveal the password, access certificate, or access key. Click **Download credentials** to save all at once. Related pages * [Manage approvals in the Aiven Console](/docs/products/kafka/howto/approvals.md) --- # Rotate credentials for an Apache Kafka® subscription Rotate credentials for an Apache Kafka® subscription to replace outdated credentials and maintain secure, approved access. The request must be approved by another member of the owner group before new credentials can be downloaded. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") You must be a member of the **owner group** of the access to rotate. ## Request a credential rotation[​](#request-a-credential-rotation "Direct link to Request a credential rotation") 1. In the [Aiven Console](https://console.aiven.io/), go to **Tools > Apache Kafka governance operations**. 2. In the sidebar, click **Streaming catalog** > **Access overview**. 3. In the **Access overview** page, locate the service user for which you need credentials. 4. Click **Actions** > **Request credential rotation**. 5. In the **Request credential rotation** dialog, enter a message for the approvers. 6. Click **Submit**. After you submit a rotation request: * The request is sent to the **owner group** for approval. * The request must be approved by another member of the owner group. * After approval: * New credentials are generated for the service user associated with the access. * The requester is notified when credentials are ready. * The requester can download the credentials from the request details. * Only one active rotation request is allowed per access at a time. ## Approve a credential rotation[​](#approve-a-credential-rotation "Direct link to Approve a credential rotation") 1. Go to **Tools > Apache Kafka governance operations**. 2. In the sidebar, click **Approvals**. 3. Find the request with type **Rotate credentials**. 4. Click the request to view details. 5. Click **Approve** or **Decline**. note Another member of the owner group must approve the request. ## View and download credentials[​](#view-and-download-credentials "Direct link to View and download credentials") After your request is approved, you can download the new credentials. 1. Open the request from the email notification or go to **Tools > Apache Kafka governance operations**. 2. In the sidebar, click **Streaming catalog** > **Access overview**. 3. Locate the service user related to your request. 4. Click **Actions** > **Review credentials**. 5. In the confirmation dialog, click **Show credentials**. 6. In the **Save service user credentials** dialog: * Click **Show** next to each field to reveal the credentials. * Click **Copy** next to each field to copy credentials. * Click **Download credentials** to save them. warning Credentials can only be viewed once. Make sure to save them securely. Related pages [Request access to an Apache Kafka topic](/docs/products/kafka/howto/request-access-topic.md) --- # Scale disk storage for your Classic Kafka service Use dynamic disk sizing (DDS) to add or remove disk storage on a Classic Kafka service. note DDS is available on [Classic Kafka](/docs/products/kafka/classic-kafka-overview.md) services only. /eol-for-major-versions#aiven-for-flinkAdding or removing disk storage does not disrupt the running service. You pay only for extra storage instead of upgrading compute resources. You can add extra storage when you create a service or after it is running. When you add storage to a running service, the Aiven Platform provisions the extra disk and adds it to the running instances. For a clustered service such as Aiven for Apache Kafka®, Aiven divides extra storage equally between the nodes. For a shared service, each node receives the full extra capacity. ## Limitations[​](#limitations "Direct link to Limitations") * Disk added for extra storage is slower than the original disk until the next maintenance update. The slower disk can reduce performance for I/O-intensive workloads. * Maximum storage depends on the plan, service type, and cloud provider. It can be up to five times the plan's base storage size. * Cloud providers limit how many times you can increase storage between maintenance updates. If you reach the limit, run a maintenance update to optimize performance. * You cannot add storage during a maintenance update. * Dynamic disk sizing (DDS) is not supported on custom service plans. Pricing If you add storage when you create a service, **Additional disk storage** shows an estimated monthly cost. The **Service summary** lists plan storage plus additional storage. The estimated monthly price includes the additional storage cost. If you add storage to a running service, the Aiven Console shows the cost of the additional storage and related backups. The same costs appear on your invoices. ## Add or remove storage[​](#add-or-remove-storage "Direct link to Add or remove storage") ### Add storage when you create a service[​](#add-storage-when-you-create-a-service "Direct link to Add storage when you create a service") To add storage while you create a service: 1. In **Additional disk storage**, set the size with the slider or enter a value in GB. 2. Review the estimated monthly cost. 3. In the **Service summary**, click **Create service**. Change additional storage later on the running service, or enable automatic disk scaling with Aiven Autoscaler. ### Change storage on a running service[​](#change-storage-on-a-running-service "Direct link to Change storage on a running service") You cannot add or remove storage when service nodes are in the rebuilding state, for example during a maintenance update or a service upgrade. If you are removing disk storage: * Make sure the data in your service does not exceed the allocated storage. If it does, you cannot remove the additional storage. * Plan for the time it takes to rebuild the service. The time depends on the service. - Console - CLI - Terraform 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Actions** > **Manage additional storage**. 3. Change the disk storage. note * The price shown for the additional storage includes backup costs. * You can only remove storage that you previously added using this feature. To downgrade further, you can change your service plan. 4. Click **Save Changes**. Use [Aiven CLI](/docs/tools/cli.md) to add or remove additional storage using [`avn service update`](/docs/tools/cli/service-cli.md#avn-cli-service-update) with the `--disk-space-gib` flag to specify the total disk space to provide to your service. For example, if your service has a 80-GiB disk and you would like to add an extra 10-GiB disk, use: ``` avn service update --disk-space-gib 90 --project PROJECT_NAME SERVICE_NAME ``` note * When you perform a horizontal service upgrade or downgrade, remember to include all additional disks the service uses. For example, when switching from `Startup-4` to `Business-4` or from `Business-4` to `Startup-4`, include all the additional disks available for this service. * When you fork an existing service, include all additional disks the service uses. Use the `additional_disk_space` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). If you added storage, the additional storage is available immediately. If you removed additional storage, the service nodes go through a rolling restart. Depending on the service type and configuration, there might be a short downtime for services with no HA capabilities. note Storage optimization is performed at the next maintenance update after a change to the storage size. Due to cloud provider limitations, there is a limit on how many times storage can be increased between two maintenance updates. When this limit is reached, perform a maintenance update for performance optimization. Plan increases to avoid reaching this limit. Related pages * [Disk autoscaler](/docs/products/kafka/howto/disk-autoscaler.md) * [Change the service plan](/docs/products/kafka/howto/change-service-plan.md) * [Prevent full disks](/docs/products/kafka/howto/prevent-full-disks.md) * [Scaling options in Apache Kafka®](/docs/products/kafka/concepts/horizontal-vertical-scaling.md) * [Classic Kafka overview](/docs/products/kafka/classic-kafka-overview.md) * [Create an Aiven for Apache Kafka® Professional tier service](/docs/products/kafka/get-started/create-kafka-service.md) --- # Manage Apache Kafka® parameters Every Aiven for Apache Kafka® service comes with a set of configuration properties that define values like partition count, replication factor, retention time, whether an automatic topic creation is enabled or not and others. tip Check a short [video tutorial](https://www.youtube.com/watch?v=pXQZWI0ddLg\&t=25s) for an end-to-end example of how to manage Aiven for Apache Kafka® parameters. You can read and edit these configuration properties in the Advanced configuration section of the Overview page in the [Aiven Console](https://console.aiven.io/) or using the [Aiven CLI service update command](/docs/tools/cli/service-cli.md#avn-cli-service-update). warning Most of the Apache Kafka settings cause the service to restart when changed. Aiven for Apache Kafka restarts nodes one at a time to ensure minimal disruption to service availability. However, it can take a few minutes from the change before the new settings are in use. ## Retrieve the current service parameters with Aiven CLI[​](#retrieve-the-current-service-parameters-with-aiven-cli "Direct link to Retrieve the current service parameters with Aiven CLI") To retrieve the existing Aiven for Apache Kafka configuration use the following command: ``` avn service get SERVICE_NAME --json ``` The output is the JSON representation of the service configuration. ## Retrieve the customizable parameters with Aiven CLI[​](#retrieve-the-customizable-parameters-with-aiven-cli "Direct link to Retrieve the customizable parameters with Aiven CLI") Not all Aiven for Apache Kafka parameters are customizable. To retrieve the list of parameters you can change, use the following command: ``` avn service types -v ``` The output is a set of customizable parameters for all the services, browse to the `kafka` section to check the ones available for Aiven for Apache Kafka. ## Update a service parameter with the Aiven CLI[​](#update-a-service-parameter-with-the-aiven-cli "Direct link to Update a service parameter with the Aiven CLI") To modify a service parameter, use the [Aiven CLI service update command](/docs/tools/cli/service-cli.md#avn-cli-service-update). For example, to modify the `message.max.bytes` parameter, use the following command: ``` avn service update SERVICE_NAME -c "kafka.message_max_bytes=newmaximumbytelimit" ``` note For some changes, like the `message.max.bytes`, client settings need to be amended as well. Otherwise, you may encounter issues with processing Kafka messages. --- # Set up Aiven for Apache Kafka® using Skills Use [Skills to automate Kafka workflows](/docs/products/kafka/dev-tier/kafka-dev-tier.md#automate-workflows-with-skills) to create and configure an Aiven for Apache Kafka® service from the command line. A Skill can create a service and configure topics, access control lists (ACLs), and Karapace Schema Registry. note If you use the [Free tier](/docs/products/kafka/free-tier/kafka-free-tier.md), create services in the [Aiven Console](https://console.aiven.io). Skills are available only on **Developer** and **Professional** tiers. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven account and a project where you can create services. * [Aiven CLI](/docs/tools/cli.md) installed and authenticated. * [Node.js](https://nodejs.org/) installed so `npm` provides `npx`. ## Install the Skills bundle[​](#install-the-skills-bundle "Direct link to Install the Skills bundle") Install the Aiven Skills bundle: ``` npx skills add Aiven-Open/aiven-skills-bundle ``` ## Run the Kafka service Skill[​](#run-the-kafka-service-skill "Direct link to Run the Kafka service Skill") Start the Kafka service creation Skill: ``` npx skills run kafka-create-service ``` Follow the prompts to choose your project, cloud, and region from the options the Skill shows, and your service tier: **Developer** or **Professional**. note If you select **Developer** tier, the **Aiven Console** only asks for a **region** when you create the service there. The Skill may still prompt for **cloud** and **region** so the CLI can create the service—pick the options the Skill offers for Developer tier. When the run finishes, your service is configured and ready to use. For more about what Skills can automate, see [Use Skills to automate workflows](/docs/products/kafka/dev-tier/kafka-dev-tier.md#automate-workflows-with-skills). ## Use an AI agent[​](#use-an-ai-agent "Direct link to Use an AI agent") If your editor or automation exposes Skills, trigger the same shell commands through that integration. Local `npx`, an authenticated Aiven CLI session, or an equivalent documented setup is still required. ## Next steps[​](#next-steps "Direct link to Next steps") * [Connect to Aiven for Apache Kafka®](/docs/products/kafka/howto/list-code-samples.md) * [Generate sample data for Aiven for Apache Kafka®](/docs/products/kafka/howto/generate-sample-data.md) * [Create a Kafka topic](/docs/products/kafka/howto/create-topic.md) Related pages * [Aiven for Apache Kafka® Developer tier](/docs/products/kafka/dev-tier/kafka-dev-tier.md) --- # Switch a classic topic to a diskless topic [Early availability](/docs/platform/concepts/service-and-feature-releases.md) Switch an existing [classic topic](/docs/products/kafka/diskless/concepts/topics-vs-classic.md) in Aiven for Apache Kafka® to a diskless topic without copying data or renaming the topic. The topic remains available during the switch. Records written before the switch remain readable from the classic topic log. Aiven writes new records to the diskless topic. note This feature is in [early availability](/docs/platform/concepts/service-and-feature-releases.md#early-availability-) and is not enabled by default. To request access, contact your account team or [Aiven support](/docs/platform/howto/support.md#create-a-support-ticket). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * The [diskless topics](/docs/products/kafka/diskless/concepts/diskless-topic-overview.md) feature is enabled for the service. To request access, contact your account team or [Aiven support](/docs/platform/howto/support.md#create-a-support-ticket). * The service runs Apache Kafka® 4.1 or later. * You have access to one of the following: * [Aiven CLI](/docs/tools/cli.md) installed and authenticated. * [Aiven API token](/docs/platform/howto/create_authentication_token.md) for API requests. ## Considerations[​](#considerations "Direct link to Considerations") Before switching a topic, review the following: * The switch is one-way. You can't switch a diskless topic back to a classic topic. This might change in a future release. * Diskless topic limitations apply after the switch. Review [Limitations of diskless topics](/docs/products/kafka/diskless/concepts/limitations.md) to confirm that diskless topics support your workload. * Unclean leader election must be turned off for the topic. If unclean leader election is enabled, the switch does not start. * If tiered storage is not enabled for the topic, Aiven enables it automatically as part of the switch. Topics that cannot have tiered storage enabled, such as topics with compaction configured, are not eligible for the diskless switch. * The switch runs in the background. A successful update request doesn't mean that every partition has finished switching. You can't view per-partition switch status or progress in the Aiven CLI, Aiven API, topic configuration, or customer-facing metrics integrations. ## Switch a topic to a diskless topic[​](#switch-a-topic-to-a-diskless-topic "Direct link to Switch a topic to a diskless topic") To switch a classic topic to a diskless topic, enable diskless for the topic using the Aiven CLI or Aiven API. * Aiven CLI * Aiven API Run the following command: ``` avn service topic-update SERVICE_NAME TOPIC_NAME \ --project PROJECT_NAME \ --diskless-enable ``` Replace the following values: * `PROJECT_NAME`: Name of your Aiven project. * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `TOPIC_NAME`: Name of the topic to switch. Send a request to update the topic configuration: ``` API_URL="https://api.aiven.io/v1/project/PROJECT_NAME/service" API_URL="${API_URL}/SERVICE_NAME/topic/TOPIC_NAME" curl --request PUT \ --url "${API_URL}" \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "config": { "diskless_enable": true } }' ``` Replace the following values: * `PROJECT_NAME`: Name of your Aiven project. * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `TOPIC_NAME`: Name of the topic to switch. * `TOKEN`: Your Aiven API token. After the update request succeeds, Aiven starts switching the topic to a diskless topic. ## Verify the switch request[​](#verify-the-switch-request "Direct link to Verify the switch request") Verify the topic configuration to confirm that Aiven accepted the diskless switch request. * Aiven CLI * Aiven API Run the following command: ``` avn service topic-get SERVICE_NAME TOPIC_NAME \ --project PROJECT_NAME ``` Replace the following values: * `PROJECT_NAME`: Name of your Aiven project. * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `TOPIC_NAME`: Name of the topic. In the command output, find `diskless_enable` and `remote_storage_enable` and verify that their `VALUE` is `true`. This confirms that Aiven accepted the diskless switch request. It does not confirm that every partition has finished switching. Per-partition switch status or progress is not exposed in the topic configuration. Send a request to get the topic configuration: ``` API_URL="https://api.aiven.io/v1/project/PROJECT_NAME/service" API_URL="${API_URL}/SERVICE_NAME/topic/TOPIC_NAME" curl --request GET \ --url "${API_URL}" \ --header "Authorization: Bearer TOKEN" ``` Replace the following values: * `PROJECT_NAME`: Name of your Aiven project. * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `TOPIC_NAME`: Name of the topic. * `TOKEN`: Your Aiven API token. In the response, verify that `diskless_enable` and `remote_storage_enable` are both set to `true`. This confirms that Aiven accepted the diskless switch request. Per-partition switch status or progress is not exposed in the topic configuration. ## What to expect during the switch[​](#what-to-expect-during-the-switch "Direct link to What to expect during the switch") During the switch: * The topic remains available. * Producers might briefly receive errors that clients can retry. Most Kafka clients retry these errors by default. If you have not changed the default producer retry settings, no special tuning is usually required. * Consumer applications do not need changes to read records written before or after the switch. * The topic name and retention settings do not change. * Aiven manages the partition-level switch process. You do not need to take action after the update request succeeds. ### Producer settings[​](#producer-settings "Direct link to Producer settings") If you changed any of these Kafka producer settings, verify that the values allow producers to retry records for at least as long as the defaults: | Producer setting | Default | Why it matters during the switch | | --------------------- | ------------ | ----------------------------------------------------------- | | `delivery.timeout.ms` | `120000` | Allows up to 2 minutes for retrying a record. | | `retries` | `2147483647` | Effectively unlimited and bounded by `delivery.timeout.ms`. | | `enable.idempotence` | `true` | Avoids duplicate or reordered records. | | `acks` | `all` | Supports `enable.idempotence` and durable failover. | ## How the switch works[​](#how-the-switch-works "Direct link to How the switch works") When you switch a topic to a diskless topic, Aiven does the following for each partition: 1. Stops writing new records to the classic topic log. 2. Waits until records written to the classic topic partitions are safely replicated. 3. Records the offsets where the diskless topic partitions start. 4. Initializes the diskless topic partitions from those offsets. 5. Routes reads and writes for those partitions through the diskless storage system. Records written before the switch, including records already moved to tiered storage, remain readable until they expire based on the topic retention settings. Related pages * [Diskless topics](/docs/products/kafka/diskless/concepts/diskless-topic-overview.md) * [Diskless vs. classic topics](/docs/products/kafka/diskless/concepts/topics-vs-classic.md) * [Enable and configure tiered storage for topics](/docs/products/kafka/howto/configure-topic-tiered-storage.md) * [Create Apache Kafka topics](/docs/products/kafka/howto/create-topic.md) --- # Tag your Aiven for Apache Kafka® service Add key-value tags to your Aiven for Apache Kafka® service to organize services and track ownership, cost allocation, and governance. Use tags to add metadata to Aiven services to categorize them or run custom logic on them. Typical uses include: * Tagging for governance to deploy services with specific tags only. * Tagging for internal cost reporting, ownership, allocation, and accountability. A tag is a key/value pair: * **Key**: A case-sensitive string that starts with a letter and consists of letters, numbers, dashes, and underscores. The maximum length for a key is 64 characters. * **Value**: A string value limited to 64 UTF-8 characters. Within a service, the tag keys must be unique. * Console * Terraform 1. In the service, click **Service settings**. 2. In the **Service status** section, click **Actions** > **Add service tags**. 3. Enter a key and value for each tag. 4. Click **Save changes**. Use the `tag` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Power on/off and delete your Aiven for Apache Kafka® service](/docs/products/kafka/howto/power-cycle-service.md) --- # Manage approvals in Aiven for Apache Kafka® Governance using Terraform & GitHub Aiven for Apache Kafka® Governance lets you manage approval workflows for Apache Kafka topic changes using Terraform and GitHub Actions. To manage approvals in the Aiven Console, see [Manage approvals in the Aiven Console](/docs/products/kafka/howto/approvals.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Token](/docs/platform/howto/manage-application-users.md) for an [application user](/docs/platform/concepts/application-users.md) created specifically for Terraform, and assigned to a user group with governance permissions * [GitHub repository](https://docs.github.com/en/get-started/quickstart/create-a-repo) with GitHub Actions enabled * [User groups](/docs/platform/howto/manage-groups.md) configured to manage governance workflows * Optional: To access early stage Aiven Terraform Provider features, enable the following environment variable: ``` export PROVIDER_AIVEN_ENABLE_BETA=1 ``` ## How it works[​](#how-it-works "Direct link to How it works") Aiven for Apache Kafka® Governance enables approval workflows for Apache Kafka topic changes using Terraform and GitHub Actions. These workflows follow three key principles: * **Topic ownership**: Apache Kafka topics are assigned to user groups using `owner_user_group_id`. Only members of the assigned group can approve changes to that topic. * **External identity mapping**: The `aiven_external_identity` resource links GitHub usernames to Aiven users, ensuring that requesters and approvers are correctly identified. * **Approval validation**: A GitHub Action automatically checks that pull request creators and approvers meet governance requirements before changes are applied. note When Terraform governance is enabled, operational requests such as access rotation must be performed in the [Aiven Console](https://console.aiven.io/). These actions are not available through Terraform. To learn how to rotate access credentials, see [Rotate credentials](/docs/products/kafka/howto/rotate-credentials.md). ### Workflow steps:[​](#workflow-steps "Direct link to Workflow steps:") 1. Define governance policies using Aiven Terraform Provider: * Assign topic ownership (control who can approve changes) using `owner_user_group_id`. * Configure user groups (`aiven_organization_user_group`) to manage approval permissions. * Enforce team-based approval policies to ensure compliance. * Automatically validate changes against governance policies. 2. Map GitHub users to Aiven identities: * Link GitHub usernames to Aiven users using the `aiven_external_identity` resource. * Ensure that only mapped users can request or approve changes. 3. Submit and validate changes in GitHub: * A user submits a pull request (PR) to modify an Apache Kafka topic. * The GitHub Action automatically runs governance compliance checks. * Without all required approvals, the check fails. 4. Approve and apply changes: * An authorized user approves the PR. * The GitHub Action reruns governance validation. * If all policies are met, the PR passes, and Terraform applies the changes. ### Approval rules[​](#approval-rules "Direct link to Approval rules") Governance enforces approval policies to maintain security and compliance: * If an approver is a member of multiple user groups, their approval applies to all groups they belong to. * If a second approval is required, another member of the same user group must approve. * The pull request creator must be a member of the user group that owns the Apache Kafka topic. ## Set up governance approvals[​](#set-up-governance-approvals "Direct link to Set up governance approvals") Set up approval workflows using Terraform and GitHub Actions. ### Step 1. Define topic ownership and user groups[​](#step-1-define-topic-ownership-and-user-groups "Direct link to Step 1. Define topic ownership and user groups") To restrict modifications to authorized users, you must specify `owner_user_group_id` in your Terraform configuration: ``` Loading... ``` Setting `owner_user_group_id` alone does not enforce approvals. The GitHub Action must be integrated to validate changes and enforce compliance. ### Step 2. Map GitHub users to Aiven identities[​](#step-2-map-github-users-to-aiven-identities "Direct link to Step 2. Map GitHub users to Aiven identities") To verify requesters and approvers, map their GitHub user IDs to their Aiven user IDs using the `aiven_external_identity` resource. ``` Loading... ``` For more information, see the [Aiven Terraform Provider documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/data-sources/external_identity). * This mapping ensures that GitHub users are recognized as Aiven users for governance approvals. * Only mapped users can request or approve changes. * Repeat this for all users who need to request or approve changes. ### Step 3. Set up compliance checks in GitHub[​](#step-3-set-up-compliance-checks-in-github "Direct link to Step 3. Set up compliance checks in GitHub") In your GitHub repository, define a GitHub Actions workflow to enforce governance approvals. **Example GitHub Actions workflow** ``` name: Kafka governance compliance check on: pull_request: types: [opened, synchronize, reopened] pull_request_review: types: [submitted] jobs: compliance: runs-on: ubuntu-latest steps: - name: Checkout code uses: actions/checkout@v4 - name: Get PR approvers id: get_approvers uses: octokit/request-action@v2.x with: route: GET /repos/${{ github.repository }}/pulls/${{ github.event.pull_request.number }}/reviews env: GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} - name: Run compliance check uses: Aiven-Open/Aiven-apache-kafka-governance-compliance-checker-for-terraform@v1.2.0 with: requester: ${{ github.event.pull_request.user.login }} approvers: ${{ steps.get_approvers.outputs.data }} aiven_token: ${{ secrets.AIVEN_API_TOKEN }} ``` ### Step 4. Validate and apply changes using Terraform[​](#step-4-validate-and-apply-changes-using-terraform "Direct link to Step 4. Validate and apply changes using Terraform") Before applying approved changes, check governance compliance with Terraform: ``` terraform plan -out=tfplan terraform show -json tfplan > plan.json ``` Use `plan.json` for governance compliance checks before merging the pull request. If your workflow does not automatically apply Terraform changes, run: ``` terraform apply "tfplan" ``` note Terraform does not enforce governance rules. It applies changes as long as the authentication token is valid. The GitHub Action validates compliance by checking the Terraform plan. Repository administrators must configure GitHub to enforce workflow checks before merging changes. ### Step 5. View compliance check results[​](#step-5-view-compliance-check-results "Direct link to Step 5. View compliance check results") After validation, check the GitHub Actions logs or Terraform output to verify compliance. * Successful check: The request meets governance requirements. ``` { "ok": true } ``` * Failed check: The request does not meet governance rules, and an error message appears in the pull request. ``` { "ok": false, "errors": [ { "error": "Requesting user is not a member of the owner group", "resource": "aiven_kafka_topic.orders" } ] } ``` The GitHub Action reports compliance results but does not block pull requests or fail workflows automatically. To enforce workflow checks, repository administrators must configure GitHub before merging changes. * Ensure that the requester is a member of the owner user group for the Apache Kafka topic. * Rerun the Terraform plan and GitHub compliance check before merging. ## Troubleshoot governance compliance issues[​](#troubleshoot-governance-compliance-issues "Direct link to Troubleshoot governance compliance issues") | Error | Cause | Fix | | ------------------------------------------------------ | ------------------------------------------ | ---------------------------------------------- | | **Invalid plan JSON file** | Missing or incorrectly formatted plan file | Ensure a valid plan JSON file is provided | | **User is not a member of the owner group** | Requester lacks required permissions | Assign requester to the correct user group | | **Approval required from a member of the owner group** | PR lacks required approval | Ensure a valid team member has approved the PR | Related pages [Manage approvals in the Aiven Console](/docs/products/kafka/howto/approvals.md) --- # Use schema registry with Java producers and consumers Aiven for Apache Kafka® provides schema registry functionality through [Karapace](https://github.com/Aiven-Open/karapace). Karapace lets you store, retrieve, and evolve schemas without rebuilding producer or consumer code. The examples use Avro. For Protobuf or JSON Schema, generate the classes first, then apply the same connection and authentication settings. ## Workflow overview[​](#workflow-overview "Direct link to Workflow overview") To produce and consume Avro messages in Java using the schema registry: 1. Define your Avro schema. 2. Generate Java classes from the schema. 3. Add the required Maven dependencies. 4. Optional: Create a keystore, and create a truststore only if you use SASL authentication. 5. Configure your Kafka producer and consumer properties. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running [Aiven for Apache Kafka®](/docs/products/kafka/get-started/create-kafka-service.md) service * [Karapace schema registry enabled](/docs/products/kafka/karapace/howto/enable-karapace.md) * [Keystore and truststore files](/docs/products/kafka/howto/keystore-truststore.md) for SSL authentication ## Get connection details[​](#get-connection-details "Direct link to Get connection details") * Aiven Console * Aiven CLI On the service **Overview** page, open **Connection information**. 1. On the **Apache Kafka** tab, copy the **Service URI** for the bootstrap servers. 2. On the **Schema Registry** tab, copy the **Service URI**, **User**, and **Password**. Generate the keystore and truststore using the Aiven CLI: ``` avn service user-kafka-java-creds \ --project PROJECT_NAME \ --service SERVICE_NAME \ --username USERNAME ``` ## Variables[​](#kafka_schema_registry_variables "Direct link to Variables") Replace the following placeholders in the example configuration: | Variable | Description | | ------------------------ | ------------------------------------------------------------------------------ | | `BOOTSTRAPSERVERS` | Kafka service URI from **Connection information** on the service overview page | | `KEYSTORE` | Path to the keystore file | | `KEYSTOREPASSWORD` | Password for the keystore | | `TRUSTSTORE` | Path to the truststore file | | `TRUSTSTOREPASSWORD` | Password for the truststore | | `SSLKEYPASSWORD` | Password for the private key in the keystore | | `SCHEMAREGISTRYURL` | Schema registry URI from **Connection information** | | `SCHEMAREGISTRYUSER` | Schema registry username from **Connection information** | | `SCHEMAREGISTRYPASSWORD` | Schema registry password from **Connection information** | | `TOPIC_NAME` | Kafka topic name | ## Define an Avro schema[​](#define-an-avro-schema "Direct link to Define an Avro schema") Create an Avro schema file. For example, save the following schema in a file named `ClickRecord.avsc`: ``` { "type": "record", "name": "ClickRecord", "namespace": "io.aiven.avro.example", "fields": [ {"name": "session_id", "type": "string"}, {"name": "browser", "type": ["string", "null"]}, {"name": "campaign", "type": ["string", "null"]}, {"name": "channel", "type": "string"}, {"name": "referrer", "type": ["string", "null"], "default": "None"}, {"name": "ip", "type": ["string", "null"]} ] } ``` This schema defines a record named `ClickRecord` in the namespace `io.aiven.avro.example`. The record has the fields `session_id`, `browser`, `campaign`, `channel`, `referrer`, and `ip`. ## Generate Java classes and add dependencies[​](#generate-java-classes-and-add-dependencies "Direct link to Generate Java classes and add dependencies") Generate Java classes from your schema, then add the required dependencies to your `pom.xml`: * [Generate Java classes from Avro schemas](/docs/products/kafka/howto/generate-avro-java-classes.md) (used in these examples) * [Generate Java classes from Protobuf schemas](/docs/products/kafka/howto/generate-protobuf-java-classes.md) * [Generate Java classes from JSON Schema](/docs/products/kafka/howto/generate-json-java-classes.md) ## Configure producer and consumer properties[​](#configure-producer-and-consumer-properties "Direct link to Configure producer and consumer properties") For complete example code, see the [Aiven examples GitHub repository](https://github.com/aiven/aiven-examples/tree/master/solutions/kafka-schema-registry). ### Producer configuration[​](#producer-configuration "Direct link to Producer configuration") ``` props.put(CommonClientConfigs.BOOTSTRAP_SERVERS_CONFIG, BOOTSTRAPSERVERS); props.put(CommonClientConfigs.SECURITY_PROTOCOL_CONFIG, "SSL"); props.put(SslConfigs.SSL_TRUSTSTORE_LOCATION_CONFIG, TRUSTSTORE); props.put(SslConfigs.SSL_TRUSTSTORE_PASSWORD_CONFIG, TRUSTSTOREPASSWORD); props.put(SslConfigs.SSL_KEYSTORE_TYPE_CONFIG, "PKCS12"); props.put(SslConfigs.SSL_KEYSTORE_LOCATION_CONFIG, KEYSTORE); props.put(SslConfigs.SSL_KEYSTORE_PASSWORD_CONFIG, KEYSTOREPASSWORD); props.put(SslConfigs.SSL_KEY_PASSWORD_CONFIG, SSLKEYPASSWORD); props.put("schema.registry.url", SCHEMAREGISTRYURL); props.put("basic.auth.credentials.source", "USER_INFO"); props.put("basic.auth.user.info", SCHEMAREGISTRYUSER + ":" + SCHEMAREGISTRYPASSWORD); props.put(ProducerConfig.KEY_SERIALIZER_CLASS_CONFIG, StringSerializer.class.getName()); props.put(ProducerConfig.VALUE_SERIALIZER_CLASS_CONFIG, KafkaAvroSerializer.class.getName()); ``` ### Consumer configuration[​](#consumer-configuration "Direct link to Consumer configuration") ``` props.put(CommonClientConfigs.BOOTSTRAP_SERVERS_CONFIG, BOOTSTRAPSERVERS); props.put(CommonClientConfigs.SECURITY_PROTOCOL_CONFIG, "SSL"); props.put(SslConfigs.SSL_TRUSTSTORE_LOCATION_CONFIG, TRUSTSTORE); props.put(SslConfigs.SSL_TRUSTSTORE_PASSWORD_CONFIG, TRUSTSTOREPASSWORD); props.put(SslConfigs.SSL_KEYSTORE_TYPE_CONFIG, "PKCS12"); props.put(SslConfigs.SSL_KEYSTORE_LOCATION_CONFIG, KEYSTORE); props.put(SslConfigs.SSL_KEYSTORE_PASSWORD_CONFIG, KEYSTOREPASSWORD); props.put(SslConfigs.SSL_KEY_PASSWORD_CONFIG, SSLKEYPASSWORD); props.put("schema.registry.url", SCHEMAREGISTRYURL); props.put("basic.auth.credentials.source", "USER_INFO"); props.put("basic.auth.user.info", SCHEMAREGISTRYUSER + ":" + SCHEMAREGISTRYPASSWORD); props.put(ConsumerConfig.KEY_DESERIALIZER_CLASS_CONFIG, StringDeserializer.class.getName()); props.put(ConsumerConfig.VALUE_DESERIALIZER_CLASS_CONFIG, KafkaAvroDeserializer.class.getName()); props.put(KafkaAvroDeserializerConfig.SPECIFIC_AVRO_READER_CONFIG, true); props.put(ConsumerConfig.GROUP_ID_CONFIG, "clickrecord-example-group"); ``` Replace the placeholders with the values from the [variables section](#kafka_schema_registry_variables). Related pages * [Generate Java classes from Avro schemas](/docs/products/kafka/howto/generate-avro-java-classes.md) * [Generate Java classes from Protobuf schemas](/docs/products/kafka/howto/generate-protobuf-java-classes.md) * [Generate Java classes from JSON Schema](/docs/products/kafka/howto/generate-json-java-classes.md) * [Enable schema registry and REST proxy](/docs/products/kafka/karapace/howto/enable-karapace.md) * [Karapace](/docs/products/kafka/karapace.md) --- # Storage usage and settings in Aiven Console Use the storage overview page to review storage usage, billing, and retention settings for your Aiven for Apache Kafka® service. In the Aiven Console, storage is labeled differently based on the service type. Classic Kafka services show **Tiered storage**, which can be enabled and configured. Standard Kafka services show **Storage**, reflecting that object storage is always part of the data path. ## Prerequisite[​](#prerequisite "Direct link to Prerequisite") * An Aiven for Apache Kafka® service. * For Classic Kafka clusters, tiered storage enabled. For details, see [Enable tiered storage](/docs/products/kafka/howto/enable-kafka-tiered-storage.md). ## Access the storage overview page[​](#access-the-storage-overview-page "Direct link to Access the storage overview page") 1. In the [Aiven Console](https://console.aiven.io/), select your project and your Aiven for Apache Kafka service. 2. In the sidebar, click: * For Classic Kafka services: **Observe** > **Tiered storage** * For Standard Kafka services: **Observe** > **Storage** For Classic Kafka clusters: * If tiered storage is not enabled for the service, the option to enable it is shown. * If tiered storage is enabled but not configured for any topics, the option to enable it for topics is shown. For details, see [Enable and configure tiered storage for topics](/docs/products/kafka/howto/configure-topic-tiered-storage.md). 3. Once configured, you can view storage usage and settings. ## Key insights[​](#key-insights "Direct link to Key insights") View essential metrics and details related to tiered storage: * **Current billing expenses in USD:** Your tiered storage costs, calculated at hourly rates * **Forecasted month cost in USD:** Your upcoming monthly costs based on current usage * **Remote tier usage in bytes:** The volume of data that has been tiered ## Storage settings and retention[​](#storage-settings-and-retention "Direct link to Storage settings and retention") View the current local cache details and retention policy configurations for tiered storage: * **Default local retention time (ms):** The current local data retention set in milliseconds * **Default local retention bytes:** The configured volume of data, in bytes, for local retention ### Modify retention policies[​](#modify-retention-polices "Direct link to Modify retention policies") note You can update tiered storage retention settings only in Classic Kafka services. In Standard Kafka services, storage settings for classic topics are managed automatically and cannot be changed. 1. In the **Tiered storage settings** section, click **Actions** > **Update tiered storage settings**. 2. In the **Update tiered storage settings** page, adjust the values for: * **Default local retention time (ms)** * **Default local retention bytes** 3. Click **Save changes**. ## Remote storage overview[​](#remote-storage-overview "Direct link to Remote storage overview") View storage usage and configurations: * **Remote storage usage by topics:** The amount of tiered storage each topic uses * **Filter by topic:** Narrow down your view to specific topics --- # Manage Apache Kafka® topics with topic catalog [Early availability](/docs/platform/concepts/service-and-feature-releases.md) The [Aiven for Apache Kafka® topic catalog](/docs/products/kafka/concepts/topic-catalog-overview.md) offers a user-friendly interface to manage your Apache Kafka topics within your Aiven for Apache Kafka services. ## Access the topic catalog[​](#access-the-topic-catalog "Direct link to Access the topic catalog") 1. Log in to the [Aiven console](https://console.aiven.io/). 2. Click **Tools**. 3. Select **Apache Kafka governance operations**. ## Browse the topic catalog[​](#browse-the-topic-catalog "Direct link to Browse the topic catalog") On the Apache Kafka topic catalog page, you can: * **View topics in a table**: This default view lists all topics in a table format, showing columns for the topic name, service, project, and owner. * **Switch to card view**: Click **Tiles** to switch to the card view and browse topics displayed as cards. * **Search for topics**: Enter the topic name in the search bar at the top to find specific topics. ## Request a new topic[​](#request-a-new-topic "Direct link to Request a new topic") You can request the creation of a new topic from the Apache Kafka topic catalog page, which goes through an approval process to ensure proper [governance](/docs/products/kafka/concepts/governance-overview.md). To request a new topic: 1. On the Apache Kafka topic catalog page, click **Request new topic**. 2. On the Request topic form: * Select the project. * Select the Aiven for Apache Kafka service. * Select the group for topic ownership. * Enter a unique topic name. * Set the replication factor. * Set the number of partitions. * Provide a brief description of the topic. * If needed, enable advanced configuration and complete the additional fields. 3. Click **Submit**. After submitting a new topic creation request, you can view and track its status on the **Group requests** page in **Governance**. You receive a notification when your request is approved or declined. ## View and manage topic details[​](#view-and-manage-topic-details "Direct link to View and manage topic details") Use the topic details pane to view detailed information about a topic and its advanced configurations. 1. Click the name of a topic to open the topic details pane. 2. In the topic details pane, you can: * View topic details. * View advanced configurations. * Claim the topic. 3. To perform advanced operations, click **Open topic page**. note Only users who are added to the project associated with the topic will see the **Open topic page** link. For more detailed information and advanced operations available on the topic overview page, see [Manage Apache Kafka topics in detail](/docs/products/kafka/howto/manage-topics-details.md). note * The **Claim topic** option is available only if [governance is enabled](/docs/products/kafka/howto/enable-governance.md) for your organization. * If a group has claimed this topic, you can view the details on the **Approvals** page under **Governance**. Related pages * [Aiven for Apache Kafka® topic catalog](/docs/products/kafka/concepts/topic-catalog-overview.md) * [Manage Apache Kafka topics in detail](/docs/products/kafka/howto/manage-topics-details.md). * [Aiven for Apache Kafka® governance](/docs/products/kafka/concepts/governance-overview.md) --- # View and reset consumer group offsets The [open source Apache Kafka® code](https://kafka.apache.org/downloads) includes a `kafka-consumer-groups.sh` utility enabling you to view and manipulate the state of consumer groups. note Before using the `kafka-consumer-groups.sh`, configure a `consumer.properties` file pointing to a Java keystore and truststore which contain the required certificates for authentication. See how to do it in the [dedicated page](/docs/products/kafka/howto/kafka-tools-config-file.md). ## Managing consumer group offsets with `kafka-consumer-groups.sh`[​](#managing-consumer-group-offsets-with-kafka-consumer-groupssh "Direct link to managing-consumer-group-offsets-with-kafka-consumer-groupssh") The `kafka-consumer-groups.sh` tool enables to manage consumer group offsets, the following commands are available. ### List active consumer groups[​](#list-active-consumer-groups "Direct link to List active consumer groups") To list the currently active consumer groups use the following command replacing the `demo-kafka.my-project.aivencloud.com:17072` with your service URI: ``` kafka-consumer-groups.sh \ --bootstrap-server demo-kafka.my-project.aivencloud.com:17072 \ --command-config consumer.properties \ --list ``` ### Retrieve the details of a consumer group[​](#retrieve-the-details-of-a-consumer-group "Direct link to Retrieve the details of a consumer group") To retrieve the details of a consumer group use the following command replacing the `demo-kafka.my-project.aivencloud.com:17072` with the Aiven for Apache Kafka service URI and the `my-group` with the required consumer group name: ``` kafka-consumer-groups.sh \ --bootstrap-server demo-kafka.my-project.aivencloud.com:17072 \ --command-config consumer.properties \ --group my-group \ --describe ``` The details of the consumer group `my-group` are printed out in the following output: ``` GROUP TOPIC PARTITION CURRENT-OFFSET LOG-END-OFFSET LAG CONSUMER-ID HOST CLIENT-ID my-group test-topic 0 5509 5515 6 rdkafka-39404560-f8f2-4b0b /151.62.82.140 rdkafka ``` ### List the current members of a consumer group[​](#list-the-current-members-of-a-consumer-group "Direct link to List the current members of a consumer group") To retrieve the current members of a consumer group use the following command replacing the `demo-kafka.my-project.aivencloud.com:17072` with the Aiven for Apache Kafka service URI and the `my-group` with the required consumer group name: ``` kafka-consumer-groups.sh \ --bootstrap-server demo-kafka.my-project.aivencloud.com:17072 \ --command-config consumer.properties \ --group my-group \ --describe \ --members ``` The members of the `my-group` consumer group are printed out in the following output: ``` GROUP CONSUMER-ID HOST CLIENT-ID #PARTITIONS my-group rdkafka-a4c0a09c-8c6e-457e-bf9e-354a8e2f4bb8 /151.62.82.140 rdkafka 0 my-group rdkafka-39404560-f8f2-4b0b-9518-811e2eb20074 /151.62.82.140 rdkafka 1 ``` ### Reset the offset of a consumer group[​](#reset-the-offset-of-a-consumer-group "Direct link to Reset the offset of a consumer group") You might want to reset the consumer group offset when the topic parsing needs to start at a specific (non default) offset. To reset the offset use the following command replacing: * `demo-kafka.my-project.aivencloud.com:17072` with the Aiven for Apache Kafka service URI * `my-group` with the required consumer group name * `test-topic` with the required topic name warning The consumer group must be inactive when you make offset changes. ``` kafka-consumer-groups.sh \ --bootstrap-server demo-kafka.my-project.aivencloud.com:17072 \ --command-config consumer.properties \ --group test-group \ --topic test-topic \ --reset-offsets \ --to-earliest \ --execute ``` The `--reset-offsets` command has the following additional options: * `--to-earliest` : resets the offset to the beginning of the topic. * `--to-lastest` : resets the offset to the end of the topic. * `--to-offset` : resets the offset to a known, fixed offset number. * `--shift-by` : performs a relative shift from the current offset position using the given integer. tip Use positive values to skip forward or negative values to move backwards. * `--to-datetime ` : resets to the given timestamp. * `--topic :`: Applies the change to a specific partition, for example `--topic test-topic:0`. By default, the `--topic` argument applies to all partitions. --- # Aiven for Apache Kafka® Connect Aiven for Apache Kafka® Connect is a fully managed **distributed Apache Kafka® integration component**, deployable in the cloud of your choice. Apache Kafka Connect lets you integrate your existing data sources and sinks with Apache Kafka. With an Apache Kafka Connect connector, you can source data from an existing technology into a topic or sink data from a topic to a target technology by defining the endpoints. ## Source connectors[​](#source-connectors "Direct link to Source connectors") ## [Get started](/docs/products/kafka/kafka-connect/get-started.md) [Get started with Aiven for Apache Kafka® Connect and integrate it with an Aiven for Apache Kafka® service.](/docs/products/kafka/kafka-connect/get-started.md) ## [Get the best from Apache Kafka® Connect](/docs/products/kafka/kafka-connect/howto/best-practices.md) [We recommend to follow these best practices to ensure that your Apache](/docs/products/kafka/kafka-connect/howto/best-practices.md) ## [Setup and configuration](/docs/products/kafka/kafka-connect/howto/enable-connect.md) [4 items](/docs/products/kafka/kafka-connect/howto/enable-connect.md) ## [Manage connectors](/docs/products/kafka/kafka-connect/concepts/list-of-connector-plugins.md) [6 items](/docs/products/kafka/kafka-connect/concepts/list-of-connector-plugins.md) ## [Source connectors](/docs/products/kafka/kafka-connect/howto/amqp-source-connector.md) [13 items](/docs/products/kafka/kafka-connect/howto/amqp-source-connector.md) ## [Sink connectors](/docs/products/kafka/kafka-connect/howto/s3-sink.md) [23 items](/docs/products/kafka/kafka-connect/howto/s3-sink.md) ## [Observability](/docs/products/kafka/kafka-connect/reference/connect-metrics-prometheus.md) [1 item](/docs/products/kafka/kafka-connect/reference/connect-metrics-prometheus.md) ## Apache Kafka® Connect resources[​](#apache-kafka-connect-resources "Direct link to Apache Kafka® Connect resources") If you are new to Apache Kafka Connect, try these resources to learn more: * The main Apache Kafka project page: * The Karapace schema registry that Aiven maintains and makes available for every Aiven for Apache Kafka service: * Our code samples repository, to get you started: * [Aiven MCP](/docs/tools/mcp-server.md), to manage your Apache Kafka Connect services and connectors from **AI assistants** such as Cursor and Claude Code *Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries.* *Couchbase is a trademark of Couchbase, Inc.* --- # Troubleshoot connector list unavailable in Apache Kafka® Connect When you try to view connectors in Aiven for Apache Kafka® Connect, you might see the message `connector list not currently available`. This means the Kafka Connect service failed to return the list of installed connectors. note [Aiven Terraform Provider](/docs/tools/terraform.md) also displays the `connector list not currently available` message, for example when running `terraform plan`, because it uses the same backend API as the Aiven Console. ## Common causes and solutions[​](#common-causes-and-solutions "Direct link to Common causes and solutions") Kafka Connect service is starting If you recently created the service, wait 2 to 5 minutes for all nodes to become fully operational. The connector list loads automatically after initialization. Kafka Connect was recently enabled After you enable Kafka Connect on an existing Kafka service, the service takes 30 to 60 seconds to initialize. Refresh the page after that. Connector creation or update in progress During connector creation or updates, the list might be unavailable for 10 to 30 seconds. It refreshes automatically after the operation completes. Kafka Connect service is low on memory If the service is running out of memory, the connector list might continue to be unavailable. Upgrade to a larger service plan to resolve the issue. One of the Kafka Connect nodes is unavailable The connector list is retrieved from a randomly selected node. If a node is unavailable, the request might fail intermittently. The node usually recovers automatically. If the issue persists, contact [Aiven support](/docs/platform/howto/support.md). --- # JDBC source connector modes JDBC source connector extracts data from a relational database, such as PostgreSQL® or MySQL, and pushes it to Apache Kafka® where can be transformed and read by multiple consumers. The details of the connector are covered in the [Aiven JDBC source connector GitHub documentation](https://github.com/aiven/jdbc-connector-for-apache-kafka/blob/master/docs/source-connector.md). This connector type periodically queries the tables to extract the data, and can be configured in four **modes**. ## Bulk mode[​](#bulk-mode "Direct link to Bulk mode") In `bulk` mode the connector will periodically query the full table retrieving all the rows and publishing them into the Apache Kafka topic. As a result, if the source table contains `100.000` rows, the connector will insert `100.000` new messages in the Apache Kafka topic for every poll, no matter how many rows in the database are new or stale. tip Since the bulk mode replicates the whole table content into the Apache Kafka topic at every poll, it's a suitable option only for tables with limited amount of data which don't have any incremental or timestamp column. ## Incrementing mode[​](#incrementing-mode "Direct link to Incrementing mode") Using the `incrementing` mode, the connector will query the table and append a `WHERE` condition based on an **incrementing column** in order to fetch new rows. The incrementing mode requires that a column containing an always growing number (like a series) is present in the source table. The incrementing column is used to check which rows have been added since last query. note The column name is passed via the `incrementing.column.name` parameter If for example the database `students` table contains the following entries: | `student_id` | `student_name` | | ------------ | -------------- | | 1 | `Jon Doe` | | 2 | `Mary English` | | 3 | `Carol Tunder` | The column `student_id` can be used as an incremental column. On the first poll, the Apache Kafka connector will select all rows from the table and record the maximum `student_id` value in the table (`3` in the above example). The following polls will append a `WHERE` condition to the query selecting only rows with `student_id` greater than the previously recorded maximum value. In the example below, the condition will be `WHERE student_id > 3`. If the new records are available in the table, then the highest value for the incremental column is stored, and used as filter for the following polls. | `student_id` | `student_name` | | ------------ | -------------- | | 1 | `Jon Doe` | | 2 | `Mary English` | | 3 | `Carol Tunder` | | **6** | `Sam Cricket` | In the case above, where a new row for `Sam Cricket` is added, a new record will be sent to the Apache Kafka topic, and the maximum `student_id` value will be updated to `6` and used in `WHERE` condition in the next polls. warning With the incremental mode, any change which doesn't generate rows with an id higher than the maximum recorded in the previous poll will **not** be detected. for example, updating the `student_name` without changing the `student_id` will not generate any new records in Apache Kafka. ## Timestamp mode[​](#timestamp-mode "Direct link to Timestamp mode") Using the `timestamp` mode, the connector will query the table appending a `WHERE` condition based on one or more **timestamp columns**. This requires that timestamps columns (like creation date and modification date) are present for every row. In cases of two columns (for example, `creation_date` and `modification_date`) the polling query will apply the `COALESCENCE` function, parsing the value of the second column only when the first column is null. tip The timestamp columns are passed via the `timestamp.column.name parameter`. If, for example, the database `students` table contains the following entries: | `student_id` | `student_name` | `created_date` | `modified_date` | | ------------ | -------------- | -------------- | --------------- | | 1 | `Jon Doe` | 2021-01-01 | | | 2 | `Mary English` | 2021-03-01 | 2021-04-05 | | 3 | `Carol Tunder` | 2021-03-02 | 2021-04-06 | The columns `created_date` and `modified_date` can be used as timestamp columns. On the first poll, the Kafka connector will select all rows from the table and record the value in the `modified_date` and `created_date` columns (`2021-04-06` in the above example). The following polls will append a `WHERE` condition to the query selecting only rows with `modified_date` or `created_date` greater than the previously recorded maximum value using the `COALESCENCE` function. In the example below, the condition will be: ``` WHERE COALESCENCE(modified_date, created_date) > '2021-04-06' ``` If new records with more recent `modified_date` or `created_date` are available in the table, then the highest value for the timestamp columns is stored, and used as filter for the following polls. warning With the timestamp mode, any change which doesn't generate a more recent timestamp than the maximum recorded in the previous poll will **not** be detected. for example, updating the `Jon Doe`'s `modified_date` to `2021-04-03` will not be captured since a more recent date (`2021-04-06`) was already recorded in the previous poll. ## Timestamp and incrementing mode[​](#timestamp-and-incrementing-mode "Direct link to Timestamp and incrementing mode") Using the `timestamp+incrementing` mode, Kafka connect implements both the incrementing and timestamp functionalities described above. tip The incremental column name is passed via the `incrementing.column.name` and timestamp columns are passed via the `timestamp.column.name parameter`. See the [Aiven JDBC source connector GitHub documentation](https://github.com/aiven/jdbc-connector-for-apache-kafka/blob/master/docs/source-connector.md) for more information. --- # Available Apache Kafka® Connect connectors Discover a variety of connectors available for use with any Aiven for Apache Kafka® service with [Apache Kafka® Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md). ## Source connectors[​](#source-connectors "Direct link to Source connectors") Source connectors enable the integration of data from an existing technology into an Apache Kafka topic. The available source connectors include: * [Amazon S3 source connector](https://github.com/Aiven-Open/cloud-storage-connectors-for-apache-kafka) * [AMQP source connector](/docs/products/kafka/kafka-connect/howto/amqp-source-connector.md) * [Azure Blob Storage source connector](/docs/products/kafka/kafka-connect/howto/azure-blob-source.md) * [Couchbase](https://github.com/couchbase/kafka-connect-couchbase) * [Debezium for MongoDB®](https://debezium.io/docs/connectors/mongodb/) * [Debezium for MySQL](https://debezium.io/docs/connectors/mysql/) * [Debezium for Oracle](/docs/products/kafka/kafka-connect/howto/debezium-source-connector-oracle.md) * [Debezium for PostgreSQL®](/docs/products/kafka/kafka-connect/howto/debezium-source-connector-pg.md) * [Debezium for SQL Server](https://debezium.io/docs/connectors/sqlserver/) * [Debezium for PostgreSQL with TLS support](/docs/products/kafka/kafka-connect/howto/kafka-connect-debezium-tls-pg.md) * [Debezium for PostgreSQL with node replacement](/docs/products/kafka/kafka-connect/howto/debezium-source-connector-pg-node-replacement.md) * [Google Cloud Pub/Sub](https://github.com/googleapis/java-pubsub-group-kafka-connector/) * [Google Cloud Pub/Sub Lite](https://github.com/googleapis/java-pubsub-group-kafka-connector/) * [JDBC source for MySQL](/docs/products/kafka/kafka-connect/howto/jdbc-source-connector-mysql.md) * [JDBC source for PostgreSQL](/docs/products/kafka/kafka-connect/howto/jdbc-source-connector-pg.md) * [JDBC source for SQL Server](/docs/products/kafka/kafka-connect/howto/jdbc-source-connector-sql-server.md) * [MongoDB Kafka Connector (Official)](https://www.mongodb.com/docs/kafka-connector/current/) (supports both source and sink functionality) * [Salesforce source connector](/docs/products/kafka/kafka-connect/howto/salesforce-source-connector.md) * [Stream Reactor Cassandra®](https://docs.lenses.io/5.1/connectors/sources/cassandrasourceconnector/) * [Stream Reactor MQTT](https://docs.lenses.io/5.1/connectors/sources/mqttsourceconnector/) ## Sink connectors[​](#sink-connectors "Direct link to Sink connectors") Sink connectors enable the integration of data from an existing Apache Kafka topic to a target technology. The available sink connectors include: * [Amazon S3 sink connector](/docs/products/kafka/kafka-connect/howto/s3-sink-connector-aiven.md) * [Azure Blob Storage sink connector](/docs/products/kafka/kafka-connect/howto/azure-blob-sink.md) * [ClickHouse sink connector](https://github.com/ClickHouse/clickhouse-kafka-connect) * [Confluent Amazon S3 sink](/docs/products/kafka/kafka-connect/howto/s3-sink-connector-confluent.md) * [Couchbase®](https://github.com/couchbase/kafka-connect-couchbase) * [Elasticsearch](/docs/products/kafka/kafka-connect/howto/elasticsearch-sink.md) * [Google BigQuery](https://github.com/confluentinc/kafka-connect-bigquery) * [Google Cloud Pub/Sub](https://github.com/googleapis/java-pubsub-group-kafka-connector/) * [Google Cloud Pub/Sub Lite](https://github.com/googleapis/java-pubsub-group-kafka-connector/) * [Google Cloud Storage](/docs/products/kafka/kafka-connect/howto/gcs-sink.md) * [HTTP](https://github.com/aiven/http-connector-for-apache-kafka) * [IBM MQ sink connector](/docs/products/kafka/kafka-connect/howto/ibm-mq-sink-connector.md) * [Iceberg sink connector](/docs/products/kafka/kafka-connect/howto/iceberg-sink-connector.md) * [InfluxDB sink connector](/docs/products/kafka/kafka-connect/howto/influx-sink.md) * [JDBC sink](https://github.com/aiven/jdbc-connector-for-apache-kafka/blob/master/docs/sink-connector.md) * [MongoDB sink (Lenses)](/docs/products/kafka/kafka-connect/howto/mongodb-sink-lenses.md) * [OpenSearch®](/docs/products/kafka/kafka-connect/howto/opensearch-sink.md) * [Salesforce sink connector](/docs/products/kafka/kafka-connect/howto/salesforce-sink-connector.md) * [Snowflake](https://docs.snowflake.com/en/user-guide/kafka-connector) * [Splunk](https://github.com/splunk/kafka-connect-splunk) * [Stream Reactor Cassandra®](https://docs.lenses.io/5.1/connectors/sinks/cassandrasinkconnector/) * [Stream Reactor InfluxDB®](https://docs.lenses.io/5.1/connectors/sinks/influxsinkconnector/) * [Stream Reactor MongoDB®](https://docs.lenses.io/5.1/connectors/sinks/mongosinkconnector/) * [Stream Reactor MQTT](https://docs.lenses.io/5.1/connectors/sinks/mqttsinkconnector/) * [Stream Reactor Redis®\*](https://docs.lenses.io/5.1/connectors/sinks/redissinkconnector/) * [S3 IAM Assume Role](/docs/products/kafka/kafka-connect/howto/s3-iam-assume-role.md) ## Request new connectors[​](#request-new-connectors "Direct link to Request new connectors") To request a new connector, [submit an idea through the Aiven Ideas portal](https://ideas.aiven.io/). Aiven regularly reviews new ideas to help prioritize future updates. Aiven evaluates new Apache Kafka Connect connectors based on the following criteria: * License compatibility * Technical implementation * Active repository maintenance tip If the connector is not on the pre-approved list, include the name of the Aiven for Apache Kafka service you plan to use it with. This helps us better understand your use case. *** *Couchbase is a trademark of Couchbase, Inc.* --- # Get started with Aiven for Apache Kafka® Connect Get started with Aiven for Apache Kafka® Connect and integrate it with an Aiven for Apache Kafka® service. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Your project must include at least one Aiven for Apache Kafka® service. If it does not, [create one](/docs/products/kafka/get-started/create-kafka-service.md). * Console * Terraform - Access to the [Aiven Console](https://console.aiven.io) * [Terraform installed](https://www.terraform.io/downloads) * A [personal token](/docs/platform/howto/create_authentication_token.md) ## Create a dedicated Aiven for Apache Kafka® Connect service[​](#apache_kafka_connect_dedicated_cluster "Direct link to Create a dedicated Aiven for Apache Kafka® Connect service") Create a Kafka Connect service and integrate it with an Aiven for Apache Kafka® service. note For Standard Kafka services, Kafka Connect must run as a dedicated service. If no Kafka Connect service exists, create a Kafka Connect service and enable the integration from the Kafka service. * Console * Terraform 1. Log in to the [Aiven Console](https://console.aiven.io) and open the **Aiven for Apache Kafka** service. 2. In the left sidebar, click **Manage stream** > **Connectors**. * For Classic Kafka services, click **Integrate standalone service**. * For Standard Kafka services, click **Create Kafka Connect**. 3. In the Kafka Connect integration dialog, do one of the following: * If a compatible Kafka Connect service exists, select **Existing service** and choose the service. * If no compatible service exists, select **New service** to create one. 4. If you create a new Kafka Connect service, configure the service: 1. Select a cloud provider and region. 2. Select a service plan. 3. In **Service basics**, enter a service name. You cannot change the service name after creation. 5. Click **Create service**. 6. After the service is created, return to the integration dialog and click **Enable** to enable the Kafka Connect integration. After the integration is enabled, Kafka Connect appears as an integration on the Kafka service. Wait until the service status changes from **Rebuilding** to **Running** before using it. The following example files are also available in the [Aiven Terraform Provider repository](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/kafka_connect) on GitHub. 1. Create a file named `provider.tf` and add the following: ``` Loading... ``` 2. Create a file named `service.tf` and add the following: ``` Loading... ``` 3. Create a file named `variables.tf` and add the following: ``` Loading... ``` 4. Create the `terraform.tfvars` file and add the values for your token and project name. To apply your Terraform configuration: 1. Initialize Terraform by running: ``` terraform init ``` The output is similar to the following: ``` Initializing the backend... Initializing provider plugins... - Finding aiven/aiven versions matching ">= 4.0.0, < 5.0.0"... - Installing aiven/aiven v4.9.2... - Installed aiven/aiven v4.9.2 ... Terraform has been successfully initialized! ... ``` 2. To create an execution plan and preview the changes, run: ``` terraform plan ``` 3. To deploy your changes, run: ``` terraform apply --auto-approve ``` ## Next steps[​](#next-steps "Direct link to Next steps") * Review the [Aiven examples repository](https://github.com/aiven/aiven-examples) for Kafka Connect configuration examples. * Generate test data using the [sample data generator](https://github.com/aiven/python-fake-data-producer-for-apache-kafka). * Find available connectors in the [list of supported Kafka Connect plugins](/docs/products/kafka/kafka-connect/concepts/list-of-connector-plugins.md). --- # Create an AMQP source connector for Aiven for Apache Kafka® [Early availability](/docs/platform/concepts/service-and-feature-releases.md) The AMQP source connector retrieves messages from an AMQP-compatible queue and writes them to Apache Kafka® topics. For architecture, JSON field reference, and the full list of configuration properties, see the open-source [AMQP Source Connector Architecture Notes](https://aiven-open.github.io/amqp-connector-for-apache-kafka/source/architecture.html) and [Source Connector Configuration](https://aiven-open.github.io/amqp-connector-for-apache-kafka/source/configuration.html). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An [Aiven for Apache Kafka® service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with Aiven for Kafka Connect enabled, or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * Access to an AMQP-compatible broker, such as RabbitMQ. * A configured AMQP queue with available messages. * The following AMQP connection details: * `amqp.host`: Hostname or IP address of the AMQP broker. * `amqp.port`: Port for the AMQP broker. Use the port your broker documents for AMQP connections. * `amqp.address`: Queue address to consume from. For many brokers, including RabbitMQ, this value is the queue name. * `amqp.user`: Username for authentication. * `amqp.password`: Password for authentication. Use the hostname, port, and credentials from your AMQP broker documentation. If the connection fails, confirm that the port matches your broker's AMQP endpoint. ## Connector behavior[​](#connector-behavior "Direct link to Connector behavior") The connector connects to the configured AMQP queue and reads queued messages. For each message: * One Apache Kafka record is created. * The Kafka record key is a ULID (Universally Unique Lexicographically Sortable Identifier). AMQP messages are not guaranteed to expose a unique identifier, so the connector generates one for each record. * The AMQP message body and supported metadata are serialized as JSON. * The record is written to the configured Apache Kafka topic. caution The connector does not guarantee that every message on the AMQP queue is written to Apache Kafka. Messages available while the connector is not running can be missed. If a source message includes a stable `messageId`, use that field for downstream deduplication. ## Serialized record fields[​](#serialized-record-fields "Direct link to Serialized record fields") The connector serializes each AMQP message into a JSON record. Properties that are not set on the source AMQP message are omitted from the JSON. For each JSON field and serialization details, see **Data Mapping** and **Data Serialization** in the [AMQP Source Connector Architecture Notes](https://aiven-open.github.io/amqp-connector-for-apache-kafka/source/architecture.html). ## Limitations[​](#limitations "Direct link to Limitations") * Delivery is not guaranteed. Messages available while the connector is not running can be missed. * Queues only. The connector reads AMQP queues, not streams. * AMQP broker connection settings are limited to `amqp.host`, `amqp.port`, `amqp.user`, and `amqp.password`. * For a single queue, keep `tasks.max` at `1`. Additional tasks create separate consumers. Depending on the broker, this can cause competing consumers, duplicate or reordered processing, or no added throughput. Increase `tasks.max` only when you understand how your broker handles multiple receivers. * `key.converter` must use `org.apache.kafka.connect.storage.StringConverter`. * `value.converter` must use one of the following: * `org.apache.kafka.connect.storage.StringConverter` * `org.apache.kafka.connect.json.JsonConverter` * Performance depends on queue throughput and Kafka Connect task configuration. ## Output record format[​](#output-record-format "Direct link to Output record format") Each AMQP message becomes one record in the target Apache Kafka topic. The connector builds JSON from each AMQP message using Jackson and connector logic that maps AMQP-specific data to ordinary JSON values. Set `key.converter` and `value.converter` to the classes listed in [Limitations](#limitations). For more detail on serialization, see **Data Serialization** in the [AMQP Source Connector Architecture Notes](https://aiven-open.github.io/amqp-connector-for-apache-kafka/source/architecture.html). * **Key**: ULID string. * **Value**: JSON representation of the AMQP payload. What consumers read depends on `value.converter`. With `StringConverter`, the record value is **plain UTF-8 text**, a single string that contains the JSON. Consumers parse that string as JSON. `JsonConverter` serializes the connector output as JSON without generating a schema for the AMQP message structure. Example value payload inside the JSON string when using `StringConverter`: ``` { "messageId": "12345", "contentType": "application/json", "body": { "orderId": 1001, "status": "created" } } ``` If the AMQP body contains binary data, the connector encodes it using Base64. ## Create an AMQP source connector configuration file[​](#create-an-amqp-source-connector-configuration-file "Direct link to Create an AMQP source connector configuration file") Create a file named `amqp_source_connector.json` and add the following configuration: ``` { "name": "amqp-source", "connector.class": "io.aiven.kafka.connect.amqp.source.AmqpSourceConnector", "tasks.max": 1, "key.converter": "org.apache.kafka.connect.storage.StringConverter", "value.converter": "org.apache.kafka.connect.storage.StringConverter", "topic": "amqp.messages", "amqp.host": "AMQP_HOST", "amqp.port": AMQP_PORT, "amqp.address": "QUEUE_NAME", "amqp.user": "AMQP_USERNAME", "amqp.password": "AMQP_PASSWORD" } ``` Parameters: * `name`: Unique name for the connector. * `connector.class`: Connector class: `io.aiven.kafka.connect.amqp.source.AmqpSourceConnector`. * `tasks.max`: Maximum number of Kafka Connect tasks. Use `1` when consuming from a single queue unless you intentionally run multiple competing consumers. See [Limitations](#limitations) for guidance on multiple tasks per queue. * `key.converter`: Must use `org.apache.kafka.connect.storage.StringConverter`. * `value.converter`: Must use one of the following: * `org.apache.kafka.connect.storage.StringConverter` * `org.apache.kafka.connect.json.JsonConverter` * `topic`: Apache Kafka topic that receives AMQP messages. * `amqp.host`: Hostname or IP address of the AMQP broker. * `amqp.port`: AMQP broker port. Set this as a JSON number that matches your broker configuration. * `amqp.address`: AMQP queue address to consume from—for RabbitMQ, usually the queue name or receiver address for that queue. * `amqp.user`: Username for authentication. * `amqp.password`: Password for authentication. ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Aiven Console * Aiven CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is enabled on the service. If not, click **Enable connector on this service**. Alternatively: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 5. In the source connectors list, select **AMQP Source Connector**. 6. Click **Get started**. 7. On the **AMQP Source Connector** page, click the **Common** tab. 8. In the **Connector configuration** text box, click **Edit**. 9. Paste the contents of your `amqp_source_connector.json` file. 10. Click **Create connector**. 11. Confirm the connector status on the **Manage stream** > **Connectors** page. To create the AMQP source connector, run: ``` avn service connector create SERVICE_NAME @amqp_source_connector.json ``` To check the connector status, run: ``` avn service connector status SERVICE_NAME CONNECTOR_NAME ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `@amqp_source_connector.json`: Path to your connector configuration file. * `CONNECTOR_NAME`: Value of the `name` field in the JSON file. ## Verify the connector[​](#verify-the-connector "Direct link to Verify the connector") After you create the connector: 1. Confirm that the connector status is `RUNNING`. 2. Verify that messages are written to the configured Apache Kafka topic. 3. Consume from that Apache Kafka topic. Confirm that keys and values match [Output record format](#output-record-format). If not, compare your `key.converter` and `value.converter` settings with [Limitations](#limitations). Related pages * [Source Connector Configuration](https://aiven-open.github.io/amqp-connector-for-apache-kafka/source/configuration.html) * [AMQP Source Connector Architecture Notes](https://aiven-open.github.io/amqp-connector-for-apache-kafka/source/architecture.html) * [AMQP connector for Apache Kafka on GitHub](https://github.com/Aiven-Open/amqp-connector-for-apache-kafka) * [AMQP connector project documentation](https://aiven-open.github.io/amqp-connector-for-apache-kafka) * [Apache Qpid ProtonJ2 documentation](https://qpid.apache.org/proton/) --- # Configure the Iceberg sink connector with AWS Glue catalog The AWS Glue catalog directly manages Iceberg metadata within AWS Glue. It supports automatic table creation and schema evolution. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An [Aiven for Apache Kafka® service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with Apache Kafka Connect enabled, or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * Apache Kafka client settings: The Iceberg sink connector requires these settings to connect to the Iceberg control topic. For a full list of supported configurations, see [Iceberg configuration](https://iceberg.apache.org/docs/latest/kafka-connect/#kafka-configuration). * AWS-specific setup: * Create an **S3 bucket** to store data. * Configure **AWS IAM roles** with the appropriate permissions. See [Configure AWS IAM permissions](#configure-aws-iam-permissions). * Create an **AWS Glue database and tables**, and specify the S3 bucket as the storage location. For more details, see the [AWS Glue data catalog documentation](https://docs.aws.amazon.com/glue/latest/dg/start-data-catalog.html). ## Configure AWS IAM permissions[​](#configure-aws-iam-permissions "Direct link to Configure AWS IAM permissions") The Iceberg sink connector requires an IAM user with permissions to access Amazon S3 and AWS Glue. These permissions allow the connector to write data to an S3 bucket and manage metadata in the AWS Glue catalog. To set up the required permissions: 1. Create an IAM user in [AWS Identity and Access Management (IAM)](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_create.html) with permissions for Amazon S3 and AWS Glue. 2. Attach the following policy to the IAM user: ``` { "Version": "2012-10-17", "Statement": [ { "Sid": "S3Access", "Effect": "Allow", "Action": [ "s3:GetObject", "s3:PutObject", "s3:DeleteObject", "s3:ListBucket", "s3:GetBucketLocation", "s3:AbortMultipartUpload", "s3:ListMultipartUploadParts" ], "Resource": [ "arn:aws:s3:::/*" ] }, { "Sid": "S3ListBucket", "Effect": "Allow", "Action": "s3:ListBucket", "Resource": [ "arn:aws:s3:::" ] }, { "Sid": "GlueAccess", "Effect": "Allow", "Action": [ "glue:CreateDatabase", "glue:GetDatabase", "glue:GetTables", "glue:SearchTables", "glue:CreateTable", "glue:UpdateTable", "glue:GetTable", "glue:BatchCreatePartition", "glue:CreatePartition", "glue:UpdatePartition", "glue:GetPartition", "glue:GetPartitions" ], "Resource": [ "arn:aws:glue:::catalog", "arn:aws:glue:::database/*", "arn:aws:glue:::table/*" ] } ] } ``` Replace the placeholder values in the policy: * ``: Your AWS Glue catalog’s region * ``: Your AWS account ID * ``: The name of your Amazon S3 bucket 3. Obtain the access key ID and secret access key for the IAM user. 4. Add these credentials to the Iceberg sink connector configuration. For more information on creating and managing AWS IAM users and policies, see the [AWS IAM documentation](https://docs.aws.amazon.com/IAM/latest/UserGuide/introduction.html). ### AWS Glue naming conventions[​](#aws-glue-naming-conventions "Direct link to AWS Glue naming conventions") When creating databases and tables in AWS Glue for the Iceberg sink connector, follow these naming conventions to ensure compatibility: * **Database names**: * Use only lowercase letters (a-z), numbers (0-9), and underscores (\_). * Must be between **1 and 252 characters** long. * Examples: * Valid: `sales_data`, `customer_orders_2024` * Invalid: `SalesData`, `customer orders` * **Table names**: * Use only lowercase letters, numbers, and underscores. * Must be between **1 and 255 characters** long. * Examples: * Valid: `product_catalog`, `order_history_2023` * Invalid: `ProductCatalog` , `order-history` * **Column names**: AWS Glue has minimal restrictions on column names. Using only letters, numbers, and underscores is recommended for best compatibility. For more details, see the [AWS Athena naming conventions](https://docs.aws.amazon.com/athena/latest/ug/tables-databases-columns-names.html). ## Create an Iceberg sink connector configuration[​](#create-an-iceberg-sink-connector-configuration "Direct link to Create an Iceberg sink connector configuration") To configure the Iceberg sink connector, define a JSON configuration file based on your catalog type. note Loading worker properties is not supported yet. Use `iceberg.kafka.*` properties instead. 1. Create AWS resources, including an S3 bucket, Glue database, and tables. 2. Add the following configurations to the Iceberg sink connector: ``` { "name": "CONNECTOR_NAME", "connector.class": "org.apache.iceberg.connect.IcebergSinkConnector", "tasks.max": "2", "topics": "KAFKA_TOPICS", "key.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "key.converter.schemas.enable": "false", "value.converter.schemas.enable": "false", "consumer.override.auto.offset.reset": "earliest", "iceberg.kafka.auto.offset.reset": "earliest", "iceberg.tables": "DATABASE_NAME.TABLE_NAME", "iceberg.tables.auto-create-enabled": "true", "iceberg.control.topic": "ICEBERG_CONTROL_TOPIC_NAME", "iceberg.control.commit.interval-ms": "1000", "iceberg.control.commit.timeout-ms": "60000", "iceberg.catalog.type": "glue", "iceberg.catalog.glue_catalog.glue.id": "AWS_ACCOUNT_ID", "iceberg.catalog.warehouse": "s3://BUCKET_NAME", "iceberg.catalog.client.region": "AWS_REGION", "iceberg.catalog.client.credentials-provider": "org.apache.iceberg.aws.StaticCredentialsProvider", "iceberg.catalog.client.credentials-provider.access-key-id": "AWS_ACCESS_KEY_ID", "iceberg.catalog.client.credentials-provider.secret-access-key": "AWS_SECRET_ACCESS_KEY", "iceberg.catalog.io-impl": "org.apache.iceberg.aws.s3.S3FileIO", "iceberg.catalog.s3.access-key-id": "AWS_ACCESS_KEY_ID", "iceberg.catalog.s3.secret-access-key": "AWS_SECRET_ACCESS_KEY", "iceberg.catalog.s3.path-style-access": "true", "iceberg.kafka.bootstrap.servers": "KAFKA_HOST:KAFKA_PORT", "iceberg.kafka.security.protocol": "SSL", "iceberg.kafka.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "iceberg.kafka.ssl.keystore.password": "KEYSTORE_PASSWORD", "iceberg.kafka.ssl.keystore.type": "PKCS12", "iceberg.kafka.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "iceberg.kafka.ssl.truststore.password": "TRUSTSTORE_PASSWORD", "iceberg.kafka.ssl.key.password": "KEY_PASSWORD" } ``` Parameters: Most connector parameters are shared with the [AWS Glue REST catalog parameters](/docs/products/kafka/kafka-connect/howto/aws-glue-rest-catalog.md#parameters) configuration. The key differences for the AWS Glue catalog are: * `iceberg.tables.auto-create-enabled`: Set to `true` to enable automatic table creation for AWS Glue catalog * `iceberg.catalog.type`: Specify `glue` for AWS Glue catalog * `iceberg.catalog.glue_catalog.glue.id`: Enter the AWS account ID for AWS Glue catalog * `iceberg.catalog.client.credentials-provider`: Specify the credentials provider for AWS Glue catalog note Apache Kafka security settings are the same for both AWS Glue REST and AWS Glue catalog configurations. ## Create the Iceberg sink connector[​](#create-the-iceberg-sink-connector "Direct link to Create the Iceberg sink connector") * Aiven Console * Aiven CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 3. In the sink connectors list, select **Iceberg Sink Connector**, and click **Get started**. 4. On the **Iceberg Sink Connector** page, go to the **Common** tab. 5. Locate the **Connector configuration** text box and click **Edit**. 6. Paste the configuration from your `iceberg_sink_connector.json` file into the text box. 7. Click **Create connector**. 8. Verify the connector status on the **Connectors** page. To create the Iceberg sink connector using the [Aiven CLI](/docs/tools/cli.md), run: ``` avn service connector create SERVICE_NAME @iceberg_sink_connector.json ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka® service. * `@iceberg_sink_connector.json`: Path to the JSON configuration file. ## Example[​](#example "Direct link to Example") This example shows how to create an Iceberg sink connector using AWS Glue Catalog with the following properties: * Connector name: `iceberg_sink_glue` * Apache Kafka topic: `test-topic` * AWS Account ID: `your-aws-account-id` * AWS Glue region: `us-west-1` * AWS S3 bucket: `your-s3-bucket` * AWS IAM access key ID: `your-access-key-id` * AWS IAM secret access key: `your-secret-access-key` * Target table: `mydatabase.mytable` * Commit interval: `1000 ms` * Tasks: `2` ``` { "name": "iceberg_sink_glue", "connector.class": "org.apache.iceberg.connect.IcebergSinkConnector", "tasks.max": "2", "topics": "test-topic", "iceberg.catalog.type": "glue", "iceberg.catalog.glue_catalog.glue.id": "your-aws-account-id", "iceberg.catalog.client.region": "us-west-1", "iceberg.catalog.client.credentials-provider": "org.apache.iceberg.aws.StaticCredentialsProvider", "iceberg.catalog.client.credentials-provider.access-key-id": "your-access-key-id", "iceberg.catalog.client.credentials-provider.secret-access-key": "your-secret-access-key", "iceberg.catalog.io-impl": "org.apache.iceberg.aws.s3.S3FileIO", "iceberg.catalog.s3.access-key-id": "your-access-key-id", "iceberg.catalog.s3.secret-access-key": "your-secret-access-key", "iceberg.catalog.warehouse": "s3://your-s3-bucket", "iceberg.tables": "mydatabase.mytable", "iceberg.tables.auto-create-enabled": "true", "iceberg.control.commit.interval-ms": "1000", "iceberg.control.commit.timeout-ms": "60000", "key.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "iceberg.kafka.bootstrap.servers": "kafka.example.com:9092", "iceberg.kafka.security.protocol": "SSL", "iceberg.kafka.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "iceberg.kafka.ssl.keystore.password": "password", "iceberg.kafka.ssl.keystore.type": "PKCS12", "iceberg.kafka.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "iceberg.kafka.ssl.truststore.password": "password", "iceberg.kafka.ssl.key.password": "password" } ``` Related pages * [Iceberg sink connector overview](/docs/products/kafka/kafka-connect/howto/iceberg-sink-connector.md) * [AWS Glue REST catalog](/docs/products/kafka/kafka-connect/howto/aws-glue-rest-catalog.md) * [AWS Glue documentation](https://docs.aws.amazon.com/glue/latest/dg/start-data-catalog.html) * [Iceberg connector configuration](https://iceberg.apache.org/docs/latest/kafka-connect/) --- # Configure the Iceberg sink connector with AWS Glue REST catalog The AWS Glue REST catalog stores metadata using the Iceberg REST API. It integrates Apache Kafka with AWS Glue using REST-based communication. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An [Aiven for Apache Kafka® service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with Apache Kafka Connect enabled, or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * Apache Kafka client settings: The Iceberg sink connector requires these settings to connect to the Iceberg control topic. For a full list of supported configurations, see [Iceberg configuration](https://iceberg.apache.org/docs/latest/kafka-connect/#kafka-configuration). * AWS-specific setup: * Create an **S3 bucket** to store data. * Configure **AWS IAM roles** with the appropriate permissions. See [Configure AWS IAM permissions](#configure-aws-iam-permissions). * Create an **AWS Glue database and tables**: * If you use the **AWS Glue REST catalog**, manually create tables. Follow the [naming conventions](/docs/products/kafka/kafka-connect/howto/aws-glue-rest-catalog.md#aws-glue-naming-conventions), select **Apache Iceberg table** as the type, make sure the table definition matches the schema of Apache Kafka record schema. * Specify the S3 bucket as the storage location. For more details, see the [AWS Glue data catalog documentation](https://docs.aws.amazon.com/glue/latest/dg/start-data-catalog.html). ## Configure AWS IAM permissions[​](#configure-aws-iam-permissions "Direct link to Configure AWS IAM permissions") The Iceberg sink connector requires an IAM user with permissions to access Amazon S3 and AWS Glue. These permissions allow the connector to write data to an S3 bucket and manage metadata in the AWS Glue catalog. To set up the required permissions: 1. Create an IAM user in [AWS Identity and Access Management (IAM)](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_create.html) with permissions for Amazon S3 and AWS Glue. 2. Attach the following policy to the IAM user: ``` { "Version": "2012-10-17", "Statement": [ { "Sid": "S3Access", "Effect": "Allow", "Action": [ "s3:GetObject", "s3:PutObject", "s3:DeleteObject", "s3:ListBucket", "s3:GetBucketLocation", "s3:AbortMultipartUpload", "s3:ListMultipartUploadParts" ], "Resource": [ "arn:aws:s3:::/*" ] }, { "Sid": "S3ListBucket", "Effect": "Allow", "Action": "s3:ListBucket", "Resource": [ "arn:aws:s3:::" ] }, { "Sid": "GlueAccess", "Effect": "Allow", "Action": [ "glue:CreateDatabase", "glue:GetDatabase", "glue:GetTables", "glue:SearchTables", "glue:CreateTable", "glue:UpdateTable", "glue:GetTable", "glue:BatchCreatePartition", "glue:CreatePartition", "glue:UpdatePartition", "glue:GetPartition", "glue:GetPartitions" ], "Resource": [ "arn:aws:glue:::catalog", "arn:aws:glue:::database/*", "arn:aws:glue:::table/*" ] } ] } ``` Replace the placeholder values in the policy: * ``: Your AWS Glue catalog’s region * ``: Your AWS account ID * ``: The name of your Amazon S3 bucket 3. Obtain the access key ID and secret access key for the IAM user. 4. Add these credentials to the Iceberg sink connector configuration. For more information on creating and managing AWS IAM users and policies, see the [AWS IAM documentation](https://docs.aws.amazon.com/IAM/latest/UserGuide/introduction.html). ### AWS Glue naming conventions[​](#aws-glue-naming-conventions "Direct link to AWS Glue naming conventions") When creating databases and tables in AWS Glue for the Iceberg sink connector, follow these naming conventions to ensure compatibility: * **Database names**: * Use only lowercase letters (a-z), numbers (0-9), and underscores (\_). * Must be between **1 and 252 characters** long. * Examples: * Valid: `sales_data`, `customer_orders_2024` * Invalid: `SalesData`, `customer orders` * **Table names**: * Use only lowercase letters, numbers, and underscores. * Must be between **1 and 255 characters** long. * Examples: * Valid: `product_catalog`, `order_history_2023` * Invalid: `ProductCatalog` , `order-history` * **Column names**: AWS Glue has minimal restrictions on column names. Using only letters, numbers, and underscores is recommended for best compatibility. For more details, see the [AWS Athena naming conventions](https://docs.aws.amazon.com/athena/latest/ug/tables-databases-columns-names.html). ## Create an Iceberg sink connector configuration[​](#create-an-iceberg-sink-connector-configuration "Direct link to Create an Iceberg sink connector configuration") To configure the Iceberg sink connector, define a JSON configuration file based on your catalog type. note Loading worker properties is not supported yet. Use `iceberg.kafka.*` properties instead. 1. Create AWS resources, including an S3 bucket, Glue database, and tables. note The AWS Glue REST catalog does not support automatic table creation. Manually create tables in AWS Glue and ensure the schema matches the Apache Kafka data. 2. Add the following configurations to the Iceberg sink connector: ``` { "name": "CONNECTOR_NAME", "connector.class": "org.apache.iceberg.connect.IcebergSinkConnector", "tasks.max": "2", "topics": "KAFKA_TOPICS", "key.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "key.converter.schemas.enable": "false", "value.converter.schemas.enable": "false", "consumer.override.auto.offset.reset": "earliest", "iceberg.kafka.auto.offset.reset": "earliest", "iceberg.tables": "DATABASE_NAME.TABLE_NAME", "iceberg.tables.auto-create-enabled": "false", "iceberg.control.topic": "ICEBERG_CONTROL_TOPIC_NAME", "iceberg.control.commit.interval-ms": "1000", "iceberg.control.commit.timeout-ms": "60000", "iceberg.catalog.type": "rest", "iceberg.catalog.uri": "https://glue.AWS_REGION.amazonaws.com/iceberg", "iceberg.catalog.warehouse": "AWS_ACCOUNT_ID", "iceberg.catalog.client.region": "AWS_REGION", "iceberg.catalog.io-impl": "org.apache.iceberg.aws.s3.S3FileIO", "iceberg.catalog.rest.signing-name": "glue", "iceberg.catalog.rest.signing-region": "AWS_REGION", "iceberg.catalog.rest.sigv4-enabled": "true", "iceberg.catalog.rest.access-key-id": "AWS_ACCESS_KEY_ID", "iceberg.catalog.rest.secret-access-key": "AWS_SECRET_ACCESS_KEY", "iceberg.catalog.rest-metrics-reporting-enabled": "false", "iceberg.catalog.s3.access-key-id": "AWS_ACCESS_KEY_ID", "iceberg.catalog.s3.secret-access-key": "AWS_SECRET_ACCESS_KEY", "iceberg.catalog.s3.path-style-access": "true", "iceberg.kafka.bootstrap.servers": "KAFKA_HOST:KAFKA_PORT", "iceberg.kafka.security.protocol": "SSL", "iceberg.kafka.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "iceberg.kafka.ssl.keystore.password": "KEYSTORE_PASSWORD", "iceberg.kafka.ssl.keystore.type": "PKCS12", "iceberg.kafka.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "iceberg.kafka.ssl.truststore.password": "TRUSTSTORE_PASSWORD", "iceberg.kafka.ssl.key.password": "KEY_PASSWORD" } ``` ### Parameters[​](#parameters "Direct link to Parameters") * `name`: Specify the connector name * `connector.class`: Defines the connector class Use `org.apache.iceberg.connect.IcebergSinkConnector` * `tasks.max`: Define the maximum number of tasks the connector can run * `topics`: List the Apache Kafka topics containing data for Iceberg tables * `key.converter`: Set the key converter class. Use `org.apache.kafka.connect.json.JsonConverter` for JSON data * `value.converter`: Set the value converter class. Use `org.apache.kafka.connect.json.JsonConverter` for JSON data * `key.converter.schemas.enable`: Enable (`true`) or disable (`false`) schema support for the key converter * `value.converter.schemas.enable`: Enable (`true`) or disable (`false`) schema support for the value converter * `consumer.override.auto.offset.reset`: Set the Kafka consumer offset reset policy Options: `earliest` (consume from the beginning) or `latest` (consume new messages) * `iceberg.kafka.auto.offset.reset`: Set the offset reset policy for Iceberg’s Apache Kafka consumer * `iceberg.tables`: Define the target Iceberg table in `DATABASE_NAME.TABLE_NAME` format * `iceberg.tables.auto-create-enabled`: Enable (`true`) or disable (`false`) automatic table creation * `iceberg.control.topic`: Set the Kafka topic for Iceberg control operations. Defaults to `control-iceberg` if not set * `iceberg.control.commit.interval-ms`: Define how often (in milliseconds) the connector commits data to Iceberg tables. Default: `1000` (1 second) * `iceberg.control.commit.timeout-ms`: Set the maximum wait time (in milliseconds) for a commit before timing out. Default: `30000` (30 seconds) * `iceberg.catalog.type`: Specify the Iceberg catalog type. Use `rest` for AWS Glue REST catalog * `iceberg.catalog.uri`: Set the URI of the Iceberg REST catalog * `iceberg.catalog.warehouse`: Set the AWS account ID when using the REST catalog * `iceberg.catalog.client.region`: Set the AWS region for Iceberg catalog operations * `iceberg.catalog.io-impl`: Specify the file I/O implementation. Use `org.apache.iceberg.aws.s3.S3FileIO` for AWS S3 * `iceberg.catalog.rest.signing-name`: Specify the AWS service name for signing requests (for example, `glue`) * `iceberg.catalog.rest.signing-region`: Set the AWS region used for request signing. * `iceberg.catalog.rest.sigv4-enabled`: Enable (`true`) or disable (`false`) AWS SigV4 authentication for REST requests. Deprecated in version 1.8 * `iceberg.catalog.rest.auth.type`: Sets the authentication method for REST requests to `basic` (HTTP credentials), `sigv4` (AWS access key), or `oauth2` (OIDC token). Introduced in version 1.8 * `iceberg.catalog.rest.access-key-id`: Set the AWS access key ID for REST catalog authentication * `iceberg.catalog.rest.secret-access-key`: Set the AWS secret access key for REST catalog authentication * `iceberg.catalog.rest-metrics-reporting-enabled`: Enable (`true`) or disable (`false`) metrics reporting for the Iceberg REST catalog * `iceberg.catalog.s3.access-key-id`: Set the AWS access key ID for S3 authentication * `iceberg.catalog.s3.secret-access-key`: Set the AWS secret access key for S3 authentication * `iceberg.catalog.s3.path-style-access`: Enable (`true`) or disable (`false`) path-style access for S3 buckets * `iceberg.kafka.bootstrap.servers`: Define the Kafka broker connection details in `KAFKA_HOST:KAFKA_PORT` format * `iceberg.kafka.security.protocol`: Defines the security protocol. Use `SSL` for encrypted communication * `iceberg.kafka.ssl.keystore.location`: Specify the file path to the keystore containing the SSL certificate * `iceberg.kafka.ssl.keystore.password`: Set the password for the keystore * `iceberg.kafka.ssl.keystore.type`: Set the keystore type (for example, `PKCS12`). * `iceberg.kafka.ssl.truststore.location`: Specify the file path to the truststore containing trusted SSL certificates * `iceberg.kafka.ssl.truststore.password`: Set the password for the truststore * `iceberg.kafka.ssl.key.password`: Set the password to access the private key stored in the keystore note Apache Kafka security settings are the same for both AWS Glue REST and AWS Glue catalog configurations. ## Create the Iceberg sink connector[​](#create-the-iceberg-sink-connector "Direct link to Create the Iceberg sink connector") * Aiven Console * Aiven CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 3. In the sink connectors list, select **Iceberg Sink Connector**, and click **Get started**. 4. On the **Iceberg Sink Connector** page, go to the **Common** tab. 5. Locate the **Connector configuration** text box and click **Edit**. 6. Paste the configuration from your `iceberg_sink_connector.json` file into the text box. 7. Click **Create connector**. 8. Verify the connector status on the **Connectors** page. To create the Iceberg sink connector using the [Aiven CLI](/docs/tools/cli.md), run: ``` avn service connector create SERVICE_NAME @iceberg_sink_connector.json ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka® service. * `@iceberg_sink_connector.json`: Path to the JSON configuration file. ## Example[​](#example "Direct link to Example") This example shows how to create an Iceberg sink connector using AWS Glue as REST Catalog with the following properties: * Connector name: `iceberg_sink_rest` * Apache Kafka topic: `test-topic` * AWS Account ID: `your-aws-account-id` * AWS Glue region: `us-west-1` * AWS IAM access key ID: `your-access-key-id` * AWS IAM secret access key: `your-secret-access-key` * Target table: `mydatabase.mytable` * Commit interval: `1000 ms` * Tasks: `2` ``` { "name": "iceberg_sink_rest", "connector.class": "org.apache.iceberg.connect.IcebergSinkConnector", "tasks.max": "2", "topics": "test-topic", "iceberg.catalog.type": "rest", "iceberg.catalog.uri": "https://glue.us-west-1.amazonaws.com/iceberg", "iceberg.catalog.rest.signing-name": "glue", "iceberg.catalog.rest.signing-region": "us-west-1", "iceberg.catalog.rest.sigv4-enabled": "true", "iceberg.catalog.rest.access-key-id": "your-access-key-id", "iceberg.catalog.rest.secret-access-key": "your-secret-access-key", "iceberg.catalog.io-impl": "org.apache.iceberg.aws.s3.S3FileIO", "iceberg.catalog.s3.access-key-id": "your-access-key-id", "iceberg.catalog.s3.secret-access-key": "your-secret-access-key", "iceberg.catalog.warehouse": "your-aws-account-id", "iceberg.tables": "mydatabase.mytable", "iceberg.tables.auto-create-enabled": "false", "iceberg.control.commit.interval-ms": "1000", "iceberg.control.commit.timeout-ms": "60000", "key.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "iceberg.kafka.bootstrap.servers": "kafka.example.com:9092", "iceberg.kafka.security.protocol": "SSL", "iceberg.kafka.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "iceberg.kafka.ssl.keystore.password": "password", "iceberg.kafka.ssl.keystore.type": "PKCS12", "iceberg.kafka.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "iceberg.kafka.ssl.truststore.password": "password", "iceberg.kafka.ssl.key.password": "password" } ``` Related pages * [Iceberg sink connector overview](/docs/products/kafka/kafka-connect/howto/iceberg-sink-connector.md) * [AWS Glue catalog](/docs/products/kafka/kafka-connect/howto/aws-glue-catalog.md) * [AWS Glue documentation](https://docs.aws.amazon.com/glue/latest/dg/start-data-catalog.html) * [Iceberg connector configuration](https://iceberg.apache.org/docs/latest/kafka-connect/) --- # Create an Azure Blob Storage sink connector for Aiven for Apache Kafka® The Azure Blob Storage sink connector moves data from Apache Kafka® topics to Azure Blob Storage containers for long-term storage, such as archiving or creating backups. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before you begin, make sure you have: * An [Aiven for Apache Kafka® service](https://docs.aiven.io/docs/products/kafka/kafka-connect/howto/enable-connect) with Kafka Connect enabled, or a [dedicated Aiven for Apache Kafka Connect® service](https://docs.aiven.io/docs/products/kafka/kafka-connect/get-started#apache_kafka_connect_dedicated_cluster). * Access to an Azure Storage account and a container with the following: * Azure Storage connection string: Required to authenticate and connect to your Azure Storage account. * Azure Storage container name: Name of the Azure Blob Storage container where data is saved. ## Create an Azure Blob Storage sink connector configuration file[​](#create-an-azure-blob-storage-sink-connector-configuration-file "Direct link to Create an Azure Blob Storage sink connector configuration file") Create a file named `azure_blob_sink_connector.json` with the following configuration: ``` { "name": "azure_blob_sink", "connector.class": "io.aiven.kafka.connect.azure.sink.AzureBlobSinkConnector", "tasks.max": "1", "topics": "test-topic", "azure.storage.connection.string": "DefaultEndpointsProtocol=https;AccountName=myaccount;AccountKey=mykey;EndpointSuffix=core.windows.net", "azure.storage.container.name": "my-container", "key.converter": "org.apache.kafka.connect.storage.StringConverter", "value.converter": "org.apache.kafka.connect.storage.StringConverter", "header.converter": "org.apache.kafka.connect.storage.StringConverter", "file.name.prefix": "connect-azure-blob-sink/test-run/", "file.compression.type": "gzip", "format.output.fields": "key,value,offset,timestamp", "reload.action": "restart" } ``` Parameters: * `name`: Name of the connector. * `topics`: Apache Kafka topics to sink data from. * `azure.storage.connection.string`: Azure Storage connection string. * `azure.storage.container.name`: Azure Blob Storage container name. * `key.converter`: Class used to convert the Apache Kafka record key. * `value.converter`: Class used to convert the Apache Kafka record value. * `header.converter`: Class used to convert message headers. * `file.name.prefix`: Prefix for the files created in Azure Blob Storage. * `file.compression.type`: Compression type for the files, such as `gzip`. * `reload.action`: Action to take when reloading the connector, set to `restart`. You can view the full set of available parameters and advanced configuration options in the [Aiven Azure Blob Storage sink connector GitHub repository](https://github.com/Aiven-Open/cloud-storage-connectors-for-apache-kafka/blob/main/azure-sink-connector/README.md). ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Aiven Console * Aiven CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Kafka Connect is already enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 5. In the sink connectors, find **Azure Blob Storage sink**, and click **Get started**. 6. On the **Azure Blob Storage sink** connector page, go to the **Common** tab. 7. Locate the **Connector configuration** text box and click **Edit**. 8. Paste the configuration from your `azure_blob_sink_connector.json` file into the text box. 9. Click **Create connector**. 10. Verify the connector status on the **Manage stream** > **Connectors** page. To create the Azure Blob Storage sink connector using the Aiven CLI, run: ``` avn service connector create SERVICE_NAME @azure_blob_sink_connector.json ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka® service. * `@azure_blob_sink_connector.json`: Path to your JSON configuration file. ## Example: Define and create an Azure Blob Storage sink connector[​](#example-define-and-create-an-azure-blob-storage-sink-connector "Direct link to Example: Define and create an Azure Blob Storage sink connector") This example shows how to create an Azure Blob Storage sink connector with the following properties: * Connector name: `azure_blob_sink` * Apache Kafka topic: `test-topic` * Azure Storage connection string: `DefaultEndpointsProtocol=https;AccountName=myaccount;AccountKey=mykey;EndpointSuffix=core.windows.net` * Azure container: `my-container` * Output fields: `key, value, offset, timestamp` * File name prefix: `connect-azure-blob-sink/test-run/` * Compression type: `gzip` ``` { "name": "azure_blob_sink", "connector.class": "io.aiven.kafka.connect.azure.sink.AzureBlobSinkConnector", "tasks.max": "1", "topics": "test-topic", "azure.storage.connection.string": "DefaultEndpointsProtocol=https;AccountName=myaccount;AccountKey=mykey;EndpointSuffix=core.windows.net", "azure.storage.container.name": "my-container", "format.output.fields": "key,value,offset,timestamp", "file.name.prefix": "connect-azure-blob-sink/test-run/", "file.compression.type": "gzip" } ``` Once this configuration is saved in the `azure_blob_sink_connector.json` file, you can create the connector using the Aiven Console or CLI, and verify that data from the Apache Kafka topic `test-topic` is successfully delivered to your Azure Blob Storage container. --- # Create an Azure Blob Storage source connector for Aiven for Apache Kafka® Use the Azure Blob source connector to stream data from Blob Storage into Apache Kafka® for real-time processing, analytics, or recovery. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An [Aiven for Apache Kafka® service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with Apache Kafka Connect enabled, or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster) * An Azure Storage container with data to stream into an Apache Kafka topic * A storage account connection string (`azure.storage.connection.string`) for authentication * Azure Blob Storage details: * `azure.storage.container.name`: Name of the container to read from * Optional: `azure.blob.prefix` or `file.name.prefix` to filter files * An existing Apache Kafka topic where the connector sends the data ## Input formats[​](#input-formats "Direct link to Input formats") Use the `input.format` parameter to configure how the connector interprets the contents of source files. Supported formats: * `bytes` (default) * `jsonl` * `avro` * `parquet` ## Create an Azure Blob Storage source connector configuration file[​](#create-an-azure-blob-storage-source-connector-configuration-file "Direct link to Create an Azure Blob Storage source connector configuration file") Create a file named `azure_blob_source_config.json` with the following configuration: ``` { "name": "azure-blob-source", "connector.class": "io.aiven.kafka.connect.azure.source.AzureBlobSourceConnector", "azure.storage.connection.string": "CONNECTION_STRING", "azure.storage.container.name": "CONTAINER_NAME", "file.name.template": "{{topic}}-{{timestamp}}.gz", "tasks.max": 1, "azure.blob.prefix": "data/logs/", "file.compression.type": "gzip", "input.format": "jsonl", "poll.interval.ms": 10000 } ``` Parameters: * `name`: Name of the connector * `connector.class`: Class name of the Azure Blob Storage source connector * `azure.storage.connection.string`: Connection string for the Azure Storage account * `azure.storage.container.name`: Name of the Azure blob container to read from * `file.name.template`: Pattern used to match blob filenames, such as `{{topic}}-{{timestamp}}.gz`. Supported placeholders: * `{{topic}}`: Kafka topic name * `{{partition}}`: Partition number * `{{start_offset}}`: Starting Kafka offset * `{{timestamp}}`: Timestamp when the file is written Example: `{{topic}}-{{partition:padding=true}}-{{start_offset:padding=true}}.gz` * `tasks.max`: Maximum number of parallel ingestion tasks * `azure.blob.prefix`: Optional. Prefix path in the container to filter files * `file.compression.type`: Optional. Compression type used in the files. Valid values are `none`, `gzip`, `snappy`, or `zstd` * `input.format`: Optional. Format of the input files. Valid values are `bytes` (default), `avro`, `json`, or `parquet` * `poll.interval.ms`: Optional. How often the connector checks for new blobs, in milliseconds. The default is `5000` **Advanced options** For advanced use cases, such as Avro or Parquet formats, byte buffering, or topic overrides, you can customize the following settings: * `schema.registry.url`: URL of the schema registry. Required when `input.format` is set to `avro` or `parquet` If the schema registry requires authentication, provide the following properties: ``` "basic.auth.credentials.source": "USER_INFO", "basic.auth.user.info": "username:password" ``` The `basic.auth.user.info` value should contain your Schema Registry credentials in the format `username:password`. * `value.serializer`: Serializer used for values with Avro input format * `transformer.max.buffer.size`: Maximum size in bytes of each blob read when using the `bytes` input format with byte distribution * `distribution.type`: File distribution strategy. Valid values are `hash` (default) or `partition` * `errors.tolerance`: Whether to skip records with decoding or formatting errors. Set to `all` to prevent connector failure * `topic`: Kafka topic to use if not specified in the file name template, or to override the topic defined in the template For a complete list of configuration options, see the [Azure Blob source connector configuration reference](https://aiven-open.github.io/cloud-storage-connectors-for-apache-kafka/azure-source-connector/AzureBlobSourceConfig.html). caution This file is auto-generated and may change when the connector is updated. ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Console * CLI * Terraform 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 5. In the source connectors list, select **Azure Blob source connector**, and click **Get started**. 6. On the **Azure Blob Source Connector** page, go to the **Common** tab. 7. Locate the **Connector configuration** text box and click **Edit**. 8. Paste the configuration from your `azure_blob_source_config.json` file into the text box. 9. Click **Create connector**. 10. Verify the connector status on the **Manage stream** > **Connectors** page. To create the Azure Blob Storage source connector using the Aiven CLI, run: ``` avn service connector create SERVICE_NAME @azure_blob_source_config.json ``` Replace: * `SERVICE_NAME`: Name of your Apache Kafka or Apache Kafka Connect service. * `@azure_blob_source_config.json`: Path to your JSON configuration file. You can configure this connector using the [`aiven_kafka_connector`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/kafka_connector#example-usage) resource in the Aiven Terraform Provider. Use the `config` block to define the connector settings as key-value pairs. ## Example: Define and create an Azure Blob Storage source connector[​](#example-define-and-create-an-azure-blob-storage-source-connector "Direct link to Example: Define and create an Azure Blob Storage source connector") This example creates an Azure Blob Storage source connector with the following settings: * Connector name: `azure_blob_source` * Apache Kafka topic: `blob-ingest-topic` * Azure Storage connection string: `CONNECTION_STRING` * Azure container: `CONTAINER_NAME` * File prefix: `data/logs/` * File name template: `{{topic}}-{{timestamp}}.gz` * Input format: `json` * Compression type: `gzip` * Poll interval: `10 seconds` ``` { "name": "azure_blob_source", "connector.class": "io.aiven.kafka.connect.azure.source.AzureBlobSourceConnector", "tasks.max": 1, "kafka.topic": "blob-ingest-topic", "azure.storage.connection.string": "CONNECTION_STRING", "azure.storage.container.name": "CONTAINER_NAME", "azure.blob.prefix": "data/logs/", "file.name.template": "{{topic}}-{{timestamp}}.gz", "file.compression.type": "gzip", "input.format": "json", "poll.interval.ms": 10000 } ``` ## Disaster recovery[​](#disaster-recovery "Direct link to Disaster recovery") To use the Azure Blob Storage source connector for disaster recovery: 1. On your primary Aiven for Apache Kafka service (or a dedicated Aiven for Apache Kafka Connect service), configure an [Azure Blob sink connector](/docs/products/kafka/kafka-connect/howto/azure-blob-sink.md) to write data to Azure Blob Storage. 2. On your secondary Apache Kafka service, configure an Azure Blob source connector to restore data from the same Azure container. ### Configuration requirements[​](#configuration-requirements "Direct link to Configuration requirements") The source and sink connectors should use the same settings to ensure successful recovery: * **File naming**: Use the same `file.name.template` in both sink and source connectors for compatibility. * For source connectors, a recommended format is `{{topic}}-{{timestamp:padding=true}}.gz`. * For sink connectors, formats like `{{topic}}-{{partition:padding=true}}-{{start_offset:padding=true}}.gz` are more common. This format may not suit source connectors, as it can make file matching harder. * **Format**: Use the same `input.format`. If using Avro or Parquet, make sure schema-related settings match. * **Compression**: Use the same compression method. Using different settings can cause the source connector to skip files, fail to decode content, or ingest records with incorrect topic, partition, or offset information. note The source connector uses Apache Kafka Connect offset tracking. If it restarts before committing offsets, it may reprocess previously ingested files. Design downstream systems to tolerate duplicates or implement idempotent processing. --- # Get the best from Apache Kafka® Connect We recommend to follow these best practices to ensure that your Apache Kafka® Connect service is fast and reliable. ## Pay attention to `tasks.max` for connector configurations[​](#pay-attention-to-tasksmax-for-connector-configurations "Direct link to pay-attention-to-tasksmax-for-connector-configurations") By default, connectors run a maximum of **1** task, which usually leads to under-utilization for large Apache Kafka Connect services (unless you have many connectors with one task each). In general, it is good to keep the cluster CPUs occupied with connector tasks without overloading them. If your connector is under-performing, you can try increasing `tasks.max` to match the number of partitions. ## Consider a standalone Apache Kafka Connect service[​](#consider-a-standalone-apache-kafka-connect-service "Direct link to Consider a standalone Apache Kafka Connect service") You can run Apache Kafka Connect as part of your existing Aiven for Apache Kafka service (for *business-4* and higher service plans). While this allows you to try out Apache Kafka Connect by enabling the feature within your existing service, for heavy usage we recommend that you enable a standalone Apache Kafka Connect service to run your connectors. This allows you to scale the Apache Kafka service and the connector service independently and offers more CPU time and memory for the Apache Kafka Connect service. Having a dedicated cluster that moves the data and does not add load to the main nodes goes according to the "separation of concerns" paradigm. --- # Bring your own Apache Kafka® Connect cluster Aiven provides Apache Kafka® Connect as a managed service in combination with the Aiven for Apache Kafka® managed service. However, there are circumstances where you may want to roll your own Kafka Connect cluster. Integrate your own Apache Kafka Connect cluster with Aiven for Apache Kafka and use the schema registry offered by [Karapace](/docs/products/kafka/karapace.md). The example below shows how to create a JDBC sink connector to a PostgreSQL® database. ## Prerequisites[​](#bring_your_own_kafka_connect_prereq "Direct link to Prerequisites") To bring your own Apache Kafka Connector, you need an Aiven for Apache Kafka service up and running. For the JDBC sink connector database example, collect the following information about the Aiven for Apache Kafka service and the target database upfront: * `APACHE_KAFKA_HOST`: The hostname of the Apache Kafka service * `APACHE_KAFKA_PORT`: The port of the Apache Kafka service * `REST_API_PORT`: The Apache Kafka's REST API port, only needed when testing data flow with REST APIs * `REST_API_USERNAME`: The Apache Kafka's REST API username, only needed when testing data flow with REST APIs * `REST_API_PASSWORD`: The Apache Kafka's REST API password, only needed when testing data flow with REST APIs * `SCHEMA_REGISTRY_PORT`: The Apache Kafka's schema registry port, only needed when using Avro as data format * `SCHEMA_REGISTRY_USER`: The Apache Kafka's schema registry username, only needed when using Avro as data format * `SCHEMA_REGISTRY_PASSWORD`: The Apache Kafka's schema registry user password, only needed when using Avro as data format * `PG_HOST`: The PostgreSQL service hostname * `PG_PORT`: The PostgreSQL service port * `PG_USERNAME`: The PostgreSQL service username * `PG_PASSWORD`: The PostgreSQL service password * `PG_DATABASE_NAME`: The PostgreSQL service database name note If you're using Aiven for PostgreSQL and Aiven for Apache Kafka the above details are available in the [Aiven console](https://console.aiven.io/) service Overview tab or via the dedicated `avn service get` command with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). ## Attach your own Apache Kafka Connect cluster to Aiven for Apache Kafka®[​](#attach-your-own-apache-kafka-connect-cluster-to-aiven-for-apache-kafka "Direct link to Attach your own Apache Kafka Connect cluster to Aiven for Apache Kafka®") The following example demonstrates how to setup a local Apache Kafka Connect cluster with a working JDBC sink connector and attach it to an Aiven for Apache Kafka service. ### Set up the truststore and keystore[​](#setup_trustore_keystore_bring_your_own_connect "Direct link to Set up the truststore and keystore") Create a [Java keystore and truststore](/docs/products/kafka/howto/keystore-truststore.md) for the Aiven for Apache Kafka service. For the following example we assume: * The keystore is available at `KEYSTORE_PATH/client.keystore.p12` * The truststore is available at `TRUSTSTORE_PATH/client.truststore.jks` * For simplicity, the same secret (password) is used for both the keystore and the truststore, and is shown as `KEY_TRUST_SECRET` ### Configure the Aiven for Apache Kafka service[​](#configure-the-aiven-for-apache-kafka-service "Direct link to Configure the Aiven for Apache Kafka service") Enable the schema registry features offered by [Karapace](/docs/products/kafka/karapace.md). You can do it in the [Aiven Console](https://console.aiven.io/) in the Aiven for Apache Kafka service Overview tab. 1. Enable the **Schema Registry (Karapace)** and **Apache Kafka REST API (Karapace)** 2. In the **Topic** tab, create a topic called `jdbc_sink`, the topic will be used by the Apache Kafka Connect connector ### Download the required binaries[​](#download-the-required-binaries "Direct link to Download the required binaries") The following binaries are needed to setup a Apache Kafka Connect cluster locally: * [Apache Kafka](https://kafka.apache.org/quickstart) * [Aiven for Kafka connect JDBC connector](https://github.com/aiven/jdbc-connector-for-apache-kafka/releases) * If you are going to use Avro as the data format, [Avro Value Converter](https://www.confluent.io/hub/confluentinc/kafka-connect-avro-converter). The examples below show how to do this. ### Set up the local Apache Kafka Connect cluster[​](#set-up-the-local-apache-kafka-connect-cluster "Direct link to Set up the local Apache Kafka Connect cluster") The following process defines the setup required to create a local Apache Kafka Connect cluster with Apache Kafka `3.1.0`, Avro converter `7.1.0` and JDBC connector `6.7.0`: 1. Extract the Apache Kafka binaries ``` tar -xzf kafka_2.13-3.1.0.tgz ``` 2. Within the newly created `kafka_2.13-3.1.0` folder, create a `plugins` folder containing a `lib` sub-folder ``` cd kafka_2.13-3.1.0 mkdir -p plugins/lib ``` 3. Unzip the JDBC and Avro binaries and copy the `jar` files in the `plugins/lib` folder ``` # extract aiven connect jdbc unzip jdbc-connector-for-apache-kafka-6.7.0.zip # extract confluent kafka connect avro converter unzip confluentinc-kafka-connect-avro-converter-7.1.0.zip # copying plugins in the plugins/lib folder cp jdbc-connector-for-apache-kafka-6.7.0/*.jar plugins/lib/ cp confluentinc-kafka-connect-avro-converter-7.1.0/*.jar plugins/lib/ ``` 4. Create a properties file, `my-connect-distributed.properties`, under the main `kafka_2.13-3.1.0` folder, for the Apache Kafka Connect settings. Change the following placeholders: * `PATH_TO_KAFKA_HOME` to the path to the `kafka_2.13-3.1.0` folder * `APACHE_KAFKA_HOST`, `APACHE_KAFKA_PORT`, `SCHEMA_REGISTRY_PORT`, `SCHEMA_REGISTRY_USER`, `SCHEMA_REGISTRY_PASSWORD`, to the related parameters fetched in the [prerequisite step](/docs/products/kafka/kafka-connect/howto/bring-your-own-kafka-connect-cluster.md#bring_your_own_kafka_connect_prereq) * `KEYSTORE_PATH`, `TRUSTSTORE_PATH` and `KEY_TRUST_SECRET` to the keystore, truststore location and related secret as defined in the [related step](/docs/products/kafka/kafka-connect/howto/bring-your-own-kafka-connect-cluster.md#setup_trustore_keystore_bring_your_own_connect) ``` # Define the folders for plugins, including the JDBC and Avro plugin.path=PATH_TO_KAFKA_HOME/kafka_2.13-3.1.0/plugins # Defines the location of the Apache Kafka bootstrap servers bootstrap.servers=APACHE_KAFKA_HOST:APACHE_KAFKA_PORT # Defines the group.id used by the connection cluster group.id=connect-cluster # Defines the input data format for key and value: JSON without schema key.converter=org.apache.kafka.connect.json.JsonConverter value.converter=org.apache.kafka.connect.json.JsonConverter key.converter.schemas.enable=false value.converter.schemas.enable=false # Defines the internal data format for key and value: JSON without schema internal.key.converter=org.apache.kafka.connect.json.JsonConverter internal.value.converter=org.apache.kafka.connect.json.JsonConverter internal.key.converter.schemas.enable=false internal.value.converter.schemas.enable=false # Connect clusters create three topics to manage offsets, configs, and status # information. Note that these contribute towards the total partition limit quota. offset.storage.topic=connect-offsets offset.storage.replication.factor=3 offset.storage.partitions=3 config.storage.topic=connect-configs config.storage.replication.factor=3 status.storage.topic=connect-status status.storage.replication.factor=3 # Defines the flush interval for the offset comunication offset.flush.interval.ms=10000 # Defines the SSL endpoint ssl.endpoint.identification.algorithm=https request.timeout.ms=20000 retry.backoff.ms=500 security.protocol=SSL ssl.protocol=TLS ssl.truststore.location=TRUSTSTORE_PATH/client.truststore.jks ssl.truststore.password=KEY_TRUST_SECRET ssl.keystore.location=KEYSTORE_PATH/client.keystore.p12 ssl.keystore.password=KEY_TRUST_SECRET ssl.key.password=KEY_TRUST_SECRET ssl.keystore.type=PKCS12 # Defines the consumer SSL endpoint consumer.ssl.endpoint.identification.algorithm=https consumer.request.timeout.ms=20000 consumer.retry.backoff.ms=500 consumer.security.protocol=SSL consumer.ssl.protocol=TLS consumer.ssl.truststore.location=TRUSTSTORE_PATH/client.truststore.jks consumer.ssl.truststore.password=KEY_TRUST_SECRET consumer.ssl.keystore.location=KEYSTORE_PATH/client.keystore.p12 consumer.ssl.keystore.password=KEY_TRUST_SECRET consumer.ssl.key.password=KEY_TRUST_SECRET consumer.ssl.keystore.type=PKCS12 # Defines the producer SSL endpoint producer.ssl.endpoint.identification.algorithm=https producer.request.timeout.ms=20000 producer.retry.backoff.ms=500 producer.security.protocol=SSL producer.ssl.protocol=TLS producer.ssl.truststore.location=TRUSTSTORE_PATH/client.truststore.jks producer.ssl.truststore.password=KEY_TRUST_SECRET producer.ssl.keystore.location=KEYSTORE_PATH/client.keystore.p12 producer.ssl.keystore.password=KEY_TRUST_SECRET producer.ssl.key.password=KEY_TRUST_SECRET producer.ssl.keystore.type=PKCS12 ``` 5. Start the local Apache Kafka Connect cluster, executing the following from the `kafka_2.13-3.1.0` folder: ``` ./bin/connect-distributed.sh ./my-connect-distributed.properties ``` ### Add the JDBC sink connector[​](#add-the-jdbc-sink-connector "Direct link to Add the JDBC sink connector") To add a JDBC connector to the local Apache Kafka Connect cluster: 1. Create the JDBC sink connector JSON configuration file named `jdbc-sink-pg.json` with the following content, replacing the placeholders `PG_HOST`, `PG_PORT`, `PG_USERNAME`, `PG_PASSWORD`, `PG_DATABASE_NAME`, `APACHE_KAFKA_HOST`, `SCHEMA_REGISTRY_PORT`, `SCHEMA_REGISTRY_USER`, `SCHEMA_REGISTRY_PASSWORD`. ``` { "name": "jdbc-sink-pg", "config": { "connector.class": "io.aiven.connect.jdbc.JdbcSinkConnector", "connection.url": "jdbc:postgresql://PG_HOST:PG_PORT/PG_DATABASE_NAME?user=PG_USERNAME&password=PG_PASSWORD&ssl=required", "tasks.max": "1", "topics": "jdbc_sink", "auto.create": "true", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } } ``` 2. Create the JDBC sink connector instance using Kafka Connect REST APIs ``` curl -s -H "Content-Type: application/json" -X POST \ -d @jdbc-sink-pg.json \ http://localhost:8083/connectors/ ``` 3. Check the status of the JDBC sink connector instance, `jq` is used to beautify the output ``` curl localhost:8083/connectors/jdbc-sink-pg/status | jq ``` The result should be similar to the following ``` { "name": "jdbc-sink-pg", "connector": { "state": "RUNNING", "worker_id": "10.128.0.12:8083" }, "tasks": [ { "id": 0, "state": "RUNNING", "worker_id": "10.128.0.12:8083" } ], "type": "sink" } ``` tip Check the [dedicated blog post](https://aiven.io/blog/connecting-twitter-to-aiven-for-apache-kafka) for an end-to-end example of how to setup a Kafka Connect cluster to host a custom connector. ### Verify the JDBC connector using Karapace REST APIs[​](#verify-the-jdbc-connector-using-karapace-rest-apis "Direct link to Verify the JDBC connector using Karapace REST APIs") To verify that the connector is working, you can write messages to the `jdbc_sink` topic in Avro format using [Karapace REST APIs](https://github.com/aiven/karapace): 1. Create a **Avro schema** using the `/subjects/` endpoint, after changing the placeholders for `REST_API_USER`, `REST_API_PASSWORD`, `APACHE_KAFKA_HOST`, `REST_API_PORT` ``` curl -X POST -H "Content-Type: application/vnd.schemaregistry.v1+json" \ --data ''' {"schema": "{\"type\": \"record\",\"name\": \"jdbcsinkexample\",\"namespace\": \"example\",\"doc\": \"example\",\"fields\": [{ \"type\": \"string\", \"name\": \"name\", \"doc\": \"person name\", \"namespace\": \"example\", \"default\": \"mario\"},{ \"type\": \"int\", \"name\": \"age\", \"doc\": \"persons age\", \"namespace\": \"example\", \"default\": 5}]}" }''' \ https://REST_API_USER:REST_API_PASSWORD@APACHE_KAFKA_HOST:REST_API_PORT/subjects/jdbcsinkexample/versions/ ``` The above call creates a new schema called `jdbcsinkexample` with a schema containing two fields (`name` and `age`). 2. Create a **message** in the `jdbc_sink` topic using the `jdbcsinkexample` schema, after changing the placeholders for `REST_API_USER`, `REST_API_PASSWORD`, `APACHE_KAFKA_HOST`, `REST_API_PORT` ``` curl -H "Content-Type: application/vnd.kafka.avro.v2+json" -X POST \ -d ''' {"value_schema": "{\"namespace\": \"test\", \"type\": \"record\", \"name\": \"example\", \"fields\": [{\"name\": \"name\", \"type\": \"string\"},{\"name\": \"age\", \"type\": \"int\"}]}", "records": [{"value": {"name": "Eric","age":77}}]}''' \ https://REST_API_USER:REST_API_PASSWORD@APACHE_KAFKA_HOST:REST_API_PORT/topics/jdbc_sink ``` 3. Verify the presence of a table called `jdbc_sink` in PostgreSQL containing the row with name `Eric` and age `77` --- # Create a Stream Reactor sink connector from Apache Kafka® to Apache Cassandra® The Apache Cassandra® Stream Reactor sink connector enables you to move data from an Aiven for Apache Kafka® cluster to a Apache Cassandra® database. It uses [KCQL transformations](https://docs.lenses.io/5.0/integrations/connectors/stream-reactor/sinks/cassandrasinkconnector/) to filter and map topic data before sending it to Cassandra. note See the full set of available parameters and configuration options in the [connector's documentation](https://docs.lenses.io/connectors/sink/cassandra). caution **Version compatibility** Stream Reactor version 9.0.2 includes class and package name updates introduced in version 6.0.0 by Lenses to standardize connector and converter names. Version 9.x is not compatible with version 4.2.0. To continue using version 4.2.0, [set the connector version](/docs/products/kafka/kafka-connect/howto/manage-connector-versions.md#set-version) before you upgrade. If you upgrade from version 4.2.0, recreate the connector using the updated class name. For example: ``` "connector.class": "io.lenses.streamreactor.connect.." ``` For details about these changes, see the [Stream Reactor release notes](https://docs.lenses.io/stream-reactor/docs/releases/). ## Prerequisites[​](#connect_cassandra_lenses_sink_prereq "Direct link to Prerequisites") * An Aiven for Apache Kafka service [with Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * Gather the following information for the target Cassandra database: * `CASSANDRA_HOSTNAME`: The Cassandra hostname. * `CASSANDRA_PORT`: The Cassandra port. * `CASSANDRA_USERNAME`: The Cassandra username. * `CASSANDRA_PASSWORD`: The Cassandra password. * `CASSANDRA_SSL`: Set to `true`, `false`, or `default`, depending on your SSL setup. * `CASSANDRA_KEYSTORE`: The path to the keystore containing the CA certificate, used for SSL connections. * `CASSANDRA_KEYSTORE_PASSWORD`: The password for the keystore. note If you are using Aiven for Apache Cassandra, use the following values: * `CASSANDRA_TRUSTSTORE`: `/run/aiven/keys/public.truststore.jks` * `CASSANDRA_TRUSTSTORE_PASSWORD`: `password` * `CASSANDRA_KEYSPACE`: The Cassandra keyspace to use to sink the data warning The Cassandra keyspace and destination table need to be created before starting the connector, otherwise the connector task will fail. * `TOPIC_LIST`: A comma-separated list of Kafka topics to sink. * `KCQL_TRANSFORMATION`: A KCQL statement to map topic fields to table columns. Use the following format: ``` INSERT INTO CASSANDRA_TABLE SELECT LIST_OF_FIELDS FROM APACHE_KAFKA_TOPIC ``` warning Create the Cassandra keyspace and destination table (`CASSANDRA_TABLE`) before starting the connector. The connector fails to start if they do not exist. * `APACHE_KAFKA_HOST`: The Apache Kafka host. Required only when using Avro as the data format. * `SCHEMA_REGISTRY_PORT`: The schema registry port. Required only when using Avro. * `SCHEMA_REGISTRY_USER`: The schema registry username. Required only when using Avro. * `SCHEMA_REGISTRY_PASSWORD`: The schema registry password. Required only when using Avro. note If you are using Aiven for Cassandra and Aiven for Apache Kafka, get all required connection details, including schema registry information, from the **Connection information** section on the **Overview** page. As of version 3.0, Aiven for Apache Kafka uses Karapace as the schema registry and no longer supports the Confluent Schema Registry. For a complete list of supported parameters and configuration options, see the [connector's documentation](https://docs.lenses.io/connectors/sink/cassandra). ## Create a connector configuration file[​](#create-a-connector-configuration-file "Direct link to Create a connector configuration file") Create a file named `cassandra_sink.json` and add the following configuration: ``` { "name":"CONNECTOR_NAME", "connector.class": "com.datamountaineer.streamreactor.connect.cassandra.sink.CassandraSinkConnector", "topics": "TOPIC_LIST", "connect.cassandra.host": "CASSANDRA_HOSTNAME", "connect.cassandra.port": "CASSANDRA_PORT", "connect.cassandra.username": "CASSANDRA_USERNAME", "connect.cassandra.password": "CASSANDRA_PASSWORD", "connect.cassandra.ssl.enabled": "CASSANDRA_SSL", "connect.cassandra.trust.store.path": "CASSANDRA_TRUSTSTORE", "connect.cassandra.trust.store.password": "CASSANDRA_TRUSTSTORE_PASSWORD", "connect.cassandra.key.space": "CASSANDRA_KEYSPACE", "connect.cassandra.kcql": "KCQL_TRANSFORMATION", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } ``` Parameters: * `name`: The connector name. Replace `CONNECTOR_NAME` with your desired name. * `connect.cassandra.*`: Cassandra connection parameters collected in the [prerequisite step](/docs/products/kafka/kafka-connect/howto/cassandra-streamreactor-sink.md#connect_cassandra_lenses_sink_prereq). * `key.converter` and `value.converter`: d Define the message data format in the Kafka topic. This example uses `io.confluent.connect.avro.AvroConverter` to translate messages in Avro format. The schema is retrieved from Aiven's [Karapace schema registry](https://github.com/aiven/karapace) using the `schema.registry.url` and related credentials. note The `key.converter` and `value.converter` fields define how Kafka messages are parsed and must be included in the configuration. When using Avro as the source format, set the following: * `value.converter.schema.registry.url`: Use the Aiven for Apache Kafka schema registry URL in the format `https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT`. * `value.converter.basic.auth.credentials.source`: Set to `USER_INFO`, which means authentication is done using a username and password. * `value.converter.schema.registry.basic.auth.user.info`: Provide the schema registry credentials in the format `SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD`. You can retrieve these values from the [prerequisite step](/docs/products/kafka/kafka-connect/howto/cassandra-streamreactor-sink.md#connect_cassandra_lenses_sink_prereq). ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Console * CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 5. In the sink connectors list, select **Amazon S3 source connector**, and click **Get started**. 6. On the **Stream Reactor Cassandra Sink** page, go to the **Common** tab. 7. Locate the **Connector configuration** text box and click **Edit**. 8. Paste the configuration from your `cassandra_sink.json` file into the text box. 9. Click **Create connector**. 10. Verify the connector status on the **Manage stream** > **Connectors** page. 11. Confirm that data appears in the Cassandra target table. To create the connector using the [Aiven CLI](/docs/tools/cli/service/connector.md#avn_service_connector_create), run: ``` avn service connector create SERVICE_NAME @cassandra_sink.json ``` Replace: * `SERVICE_NAME`: Your Kafka or Kafka Connect service name. * `@cassandra_sink.json`: Path to your configuration file. ## Sink topic data to Cassandra[​](#sink-topic-data-to-cassandra "Direct link to Sink topic data to Cassandra") The following example shows how to sink data from a Kafka topic to a Cassandra table. If your Kafka topic `students` contains the following data: ``` {"id":1, "name":"carlo", "age": 77} {"id":2, "name":"lucy", "age": 55} {"id":3, "name":"carlo", "age": 33} {"id":2, "name":"lucy", "age": 21} ``` To write this data to the `students_tbl` table in the `students_keyspace` keyspace, use the following connector configuration: ``` { "name": "my-cassandra-sink", "connector.class": "com.datamountaineer.streamreactor.connect.cassandra.sink.CassandraSinkConnector", "topics": "students", "connect.cassandra.host": "CASSANDRA_HOSTNAME", "connect.cassandra.port": "CASSANDRA_PORT", "connect.cassandra.username": "CASSANDRA_USERNAME", "connect.cassandra.password": "CASSANDRA_PASSWORD", "connect.cassandra.ssl.enabled": "CASSANDRA_SSL", "connect.cassandra.trust.store.path": "CASSANDRA_TRUSTSTORE", "connect.cassandra.trust.store.password": "CASSANDRA_TRUSTSTORE_PASSWORD", "connect.cassandra.key.space": "students_keyspace", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter.schemas.enable": "false", "connect.cassandra.kcql": "INSERT INTO students_tbl SELECT id, name, age FROM students" } ``` Replace all placeholder values (such as `CASSANDRA_HOSTNAME`, `CASSANDRA_PORT`, and `CASSANDRA_USERNAME`) with your actual Cassandra connection details. This configuration does the following: * `"topics": "students"`: Specifies the Kafka topic to sink. * Connection settings (`connect.cassandra.*`)\*\*: Provide the Cassandra host, port, credentials, SSL settings, and truststore paths. * `"value.converter"` and `"value.converter.schemas.enable"`: Set the message format. The topic uses raw JSON without a schema. * `"connect.cassandra.kcql"`: Defines the insert logic. Each Kafka message is written as a new row in the `students_tbl` Cassandra table. After creating the connector, check the Cassandra database to verify that the data has been written. --- # Create a Stream Reactor source connector from Apache Cassandra® to Apache Kafka® The Apache Cassandra® Stream Reactor source connector enables you to move data from an Apache Cassandra® database to an Aiven for Apache Kafka® cluster. It supports [KCQL transformations](https://docs.lenses.io/5.0/integrations/connectors/stream-reactor/sources/cassandrasourceconnector/) to parse and filter Cassandra table data before sending it to Kafka. caution **Version compatibility** Stream Reactor version 9.0.2 includes class and package name updates introduced in version 6.0.0 by Lenses to standardize connector and converter names. Version 9.x is not compatible with version 4.2.0. To continue using version 4.2.0, [set the connector version](/docs/products/kafka/kafka-connect/howto/manage-connector-versions.md#set-version) before you upgrade. If you upgrade from version 4.2.0, recreate the connector using the updated class name. For example: ``` "connector.class": "io.lenses.streamreactor.connect.." ``` For details about these changes, see the [Stream Reactor release notes](https://docs.lenses.io/stream-reactor/docs/releases/). ## Prerequisites[​](#connect_cassandra_lenses_source_prereq "Direct link to Prerequisites") * An Aiven for Apache Kafka service [with Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * Gather the following information for the source Cassandra database: * `CASSANDRA_HOSTNAME`: The Cassandra hostname. * `CASSANDRA_PORT`: The Cassandra port. * `CASSANDRA_USERNAME`: The Cassandra username. * `CASSANDRA_PASSWORD`: The Cassandra password. * `CASSANDRA_SSL`: Set to `true`, `false`, or `default`, depending on your SSL setup. * `CASSANDRA_KEYSTORE`: The path to the keystore containing the CA certificate, used for SSL connections. * `CASSANDRA_KEYSTORE_PASSWORD`: The password for the keystore. note If you are using Aiven for Apache Cassandra, use the following values: * `CASSANDRA_TRUSTSTORE`: `/run/aiven/keys/public.truststore.jks` * `CASSANDRA_TRUSTSTORE_PASSWORD`: `password` * `CASSANDRA_KEYSPACE`: The Cassandra keyspace to source the data from. * `KCQL_TRANSFORMATION`: A KCQL statement to map table fields to Kafka topics. Use the following format: ``` INSERT INTO APACHE_KAFKA_TOPIC SELECT LIST_OF_FIELDS FROM CASSANDRA_TABLE [PK CASSANDRA_TABLE_COLUMN] [INCREMENTALMODE=MODE] ``` warning By default, the connector runs in **bulk mode** and polls all rows in the table periodically. To source data incrementally, use the `PK` and `INCREMENTALMODE` parameters. See [KCQL syntax for Cassandra Source](https://docs.lenses.io/5.0/integrations/connectors/stream-reactor/sources/cassandrasourceconnector/) for details. * `APACHE_KAFKA_HOST`: The Apache Kafka host. Required only when using Avro as the data format. * `SCHEMA_REGISTRY_PORT`: The schema registry port. Required only when using Avro. * `SCHEMA_REGISTRY_USER`: The schema registry username. Required only when using Avro. * `SCHEMA_REGISTRY_PASSWORD`: The schema registry password. Required only when using Avro. note If you are using Aiven for Cassandra and Aiven for Apache Kafka, get all required connection details, including schema registry information, from the **Connection information** section on the **Overview** page. As of version 3.0, Aiven for Apache Kafka uses Karapace as the schema registry and no longer supports the Confluent Schema Registry. For a complete list of supported parameters and configuration options, see the [connector's documentation](https://docs.lenses.io/connectors/source/cassandra). ## Create a connector configuration file[​](#create-a-connector-configuration-file "Direct link to Create a connector configuration file") Create a file named `cassandra_source.json` and add the following configuration: ``` { "name":"CONNECTOR_NAME", "connector.class": "com.datamountaineer.streamreactor.connect.cassandra.source.CassandraSourceConnector", "connect.cassandra.host": "CASSANDRA_HOSTNAME", "connect.cassandra.port": "CASSANDRA_PORT", "connect.cassandra.username": "CASSANDRA_USERNAME", "connect.cassandra.password": "CASSANDRA_PASSWORD", "connect.cassandra.ssl.enabled": "CASSANDRA_SSL", "connect.cassandra.trust.store.path": "CASSANDRA_TRUSTSTORE", "connect.cassandra.trust.store.password": "CASSANDRA_TRUSTSTORE_PASSWORD", "connect.cassandra.key.space": "CASSANDRA_KEYSPACE", "connect.cassandra.kcql": "KCQL_TRANSFORMATION", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } ``` Parameters: * `name`: The connector name. Replace `CONNECTOR_NAME` with your desired name. * `connect.cassandra.*`: Cassandra connection parameters collected in the [prerequisite step](#connect_cassandra_lenses_source_prereq). * `key.converter` and `value.converter`: Define the message data format. This example uses `AvroConverter`. The schema is retrieved from Aiven’s [Karapace schema registry](https://github.com/aiven/karapace). note The `key.converter` and `value.converter` fields define how Kafka messages are parsed and must be included in the configuration. When using Avro as the output format, set the following fields: * `value.converter.schema.registry.url`: Use the Aiven for Apache Kafka schema registry URL in the format `https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT`. Retrieve these values from the [prerequisite step](#connect_cassandra_lenses_source_prereq). * `value.converter.basic.auth.credentials.source`: Set to `USER_INFO` to use username and password authentication. * `value.converter.schema.registry.basic.auth.user.info`: Provide the credentials in the format `SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD`. Retrieve these values from the [prerequisite step](#connect_cassandra_lenses_source_prereq). ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Console * CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 5. In the source connectors list, select **Stream Reactor Cassandra Source**, and click **Get started**. 6. On the **Stream Reactor Cassandra Source** page, go to the **Common** tab. 7. Locate the **Connector configuration** text box and click **Edit**. 8. Paste the configuration from your `cassandra_source.json` file into the text box. 9. Click **Create connector**. 10. Verify the connector status on the **Manage stream** > **Connectors** page. 11. Confirm that data appears in the target Kafka topic. To create the connector using the [Aiven CLI](/docs/tools/cli/service/connector.md#avn_service_connector_create), run: ``` avn service connector create SERVICE_NAME @cassandra_source.json ``` Replace: * `SERVICE_NAME`: Your Kafka or Kafka Connect service name. * `@cassandra_source.json`: Path to your configuration file. ## Source Cassandra data to a Kafka topic[​](#source-cassandra-data-to-a-kafka-topic "Direct link to Source Cassandra data to a Kafka topic") The following example shows how to source data from a Cassandra table to a Kafka topic. If your Cassandra table `students` in the `students_keyspace` keyspace contains the following data: | id | name | age | timestamp\_added | | -- | ----- | --- | ---------------- | | 1 | carlo | 77 | 1719838880 | | 2 | lucy | 55 | 1719839999 | To write this data incrementally to a Kafka topic named `students_topic`, use the following connector configuration: ``` { "name": "my-cassandra-source", "connector.class": "com.datamountaineer.streamreactor.connect.cassandra.source.CassandraSourceConnector", "connect.cassandra.host": "CASSANDRA_HOSTNAME", "connect.cassandra.port": "CASSANDRA_PORT", "connect.cassandra.username": "CASSANDRA_USERNAME", "connect.cassandra.password": "CASSANDRA_PASSWORD", "connect.cassandra.ssl.enabled": "CASSANDRA_SSL", "connect.cassandra.trust.store.path": "CASSANDRA_TRUSTSTORE", "connect.cassandra.trust.store.password": "CASSANDRA_TRUSTSTORE_PASSWORD", "connect.cassandra.key.space": "students_keyspace", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter.schemas.enable": "false", "connect.cassandra.kcql": "INSERT INTO students_topic SELECT id, name, age, timestamp_added FROM students PK timestamp_added INCREMENTALMODE=TIMESTAMP" } ``` Replace all placeholder values (such as `CASSANDRA_HOSTNAME`, `CASSANDRA_PORT`, and `CASSANDRA_USERNAME`) with your actual Cassandra connection details. This configuration does the following: * `connect.cassandra.kcql`: Defines how Cassandra data is mapped to the Kafka topic. This example uses the `timestamp_added` column for incremental polling. * `value.converter` and `value.converter.schemas.enable`: Set the message format. This example uses raw JSON without a schema. * Connection settings (`connect.cassandra.*`): Provide the Cassandra host, port, credentials, SSL settings, and truststore paths. After creating the connector, check the Kafka topic to verify that the data has been written. tip If your Aiven for Apache Kafka instance does not have [automatic topic creation enabled](/docs/products/kafka/howto/create-topics-automatically.md), create the `students_topic` manually before starting the connector. --- # Create a ClickHouse sink connector for Aiven for Apache Kafka® The ClickHouse sink connector delivers data from Apache Kafka® topics to a ClickHouse database for efficient querying and analysis. tip You can also: * [Query Apache Kafka® topic data in Aiven for ClickHouse®](/docs/products/clickhouse/concepts/query-kafka-topic-data.md) using a managed setup that connects a Kafka topic to a ClickHouse table. * [Connect Aiven for ClickHouse® with Apache Kafka® or Aiven for Apache Kafka® using ClickHouse Kafka Engine](/docs/products/clickhouse/howto/integrate-kafka.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before you begin, ensure that you have the following: * An [Aiven for Apache Kafka® service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with Apache Kafka Connect enabled or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * Access to a ClickHouse service (either Aiven for ClickHouse or an external instance), including: * Hostname, port, and credentials for the ClickHouse service. * A pre-created target database and table. ClickHouse version compatibility ClickHouse 25.10 and later enable HTTP response compression by default. The ClickHouse sink connector can't decode this compression, so queries that return data fail even though the connector reports a successful connection. To use the sink connector with a ClickHouse service running version 25.10 or later, turn off HTTP compression in your connector configuration: ``` "jdbcConnectionProperties": "custom_http_params=enable_http_compression=0" ``` ## Limitations[​](#limitations "Direct link to Limitations") The ClickHouse sink connector has the following limitations related to data consistency and exactly-once delivery: 1. **No exactly-once delivery after restore**: The connector does not guarantee exactly-once delivery after the ClickHouse service is restored from a backup, powered off, or forked. However, at least once delivery is guaranteed, which may result in duplicate records in ClickHouse. 2. **Manual removal of duplicate records**: If duplicates occur, manually remove them to maintain data consistency in ClickHouse. For detailed instructions, see [Remove duplicate records](#remove-duplicate-records). ### Remove duplicate records[​](#remove-duplicate-records "Direct link to Remove duplicate records") Ensure all potential duplicates are processed before removing them: 1. Verify the committed offset in Aiven for Apache Kafka: 1. Access the [Aiven Console](https://console.aiven.io/) and select your Aiven for Apache Kafka service. 2. Click **Manage stream** > **Topics** and select the topic used by the connector. 3. Go to the **Consumer Group** tab and check the **Offset** column for the committed offset. 2. Verify the committed offset in ClickHouse: 1. In the ClickHouse service, access the query editor. 2. If you are using Aiven for ClickHouse®, go to the service's **Overview** page, and click **Query editor**. 3. Run the following query to get offset details: ``` SELECT key, minOffset, maxOffset, state FROM connect_state; ``` 3. Confirm the following conditions from the query result: * The `state` column is set to `AFTER_PROCESSING`. * The `minOffset` and `maxOffset` columns have the same value. * The committed offset from Apache Kafka is **equal to or greater than** the `minOffset` value in ClickHouse. 4. Remove duplicate records: After confirming these conditions, remove any duplicate records by running the following SQL command in ClickHouse: ``` OPTIMIZE TABLE table_name DEDUPLICATE; ``` ## Create a ClickHouse sink connector configuration file[​](#create-a-clickhouse-sink-connector-configuration-file "Direct link to Create a ClickHouse sink connector configuration file") Create a file named `clickhouse_sink_connector.json` with the following configuration: ``` { "name": "clickhouse_sink_connector", "connector.class": "com.clickhouse.kafka.connect.ClickHouseSinkConnector", "tasks.max": "1", "topics": "test_topic", "hostname": "my-clickhouse-hostname", "port": "12345", "database": "default", "username": "avnadmin", "password": "mypassword", "ssl": "true", "jdbcConnectionProperties": "custom_http_params=enable_http_compression=0", "key.converter": "org.apache.kafka.connect.storage.StringConverter", "value.converter": "org.apache.kafka.connect.storage.StringConverter" } ``` ### Parameters[​](#parameters "Direct link to Parameters") * `name`: Name of the connector. * `topics`: Apache Kafka topics from which to pull data. * `hostname`: Hostname of the ClickHouse service * `port`: Port of the ClickHouse service. * `database`: Target database in ClickHouse. * `username`: Username for authentication in the ClickHouse service. * `password`: Password for authentication in the ClickHouse service. * `ssl`: Set to `true` to enable SSL encryption. * `jdbcConnectionProperties`: Turns off HTTP response compression for ClickHouse 25.10 or later. Set `custom_http_params=enable_http_compression=0`. For more configuration options, see the [ClickHouse sink connector GitHub repository](https://github.com/ClickHouse/clickhouse-kafka-connect). ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Aiven Console * Aiven CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is already enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 5. In the sink connectors list, select **ClickHouse**, and click **Get started**. 6. On the **ClickHouse** connector page, go to the **Common** tab. 7. Locate the **Connector configuration** text box and click **Edit**. 8. Paste the configuration from your `clickhouse_sink_connector.json` file into the text box. 9. Click **Create connector**. 10. Verify the connector status on the **Manage stream** > **Connectors** page. To create the ClickHouse sink connector using the [Aiven CLI](/docs/tools/cli.md), run: ``` avn service connector create SERVICE_NAME @clickhouse_sink_connector.json ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka® service. * `@clickhouse_sink_connector.json`: Path to the JSON configuration file. ## Example: Define and create a ClickHouse sink connector[​](#example-define-and-create-a-clickhouse-sink-connector "Direct link to Example: Define and create a ClickHouse sink connector") This example shows how to create a ClickHouse sink connector with the following properties: * Connector name: `clickhouse_sink_connector` * Apache Kafka topic: `test-topic` * ClickHouse hostname: `clickhouse-31d766f9-systest-project.avns.net` * ClickHouse port: `14420` * Target database: `default` * Username: `avnadmin` * Password: `mypassword` * SSL: `true` ``` { "name": "clickhouse_sink_connector", "connector.class": "com.clickhouse.kafka.connect.ClickHouseSinkConnector", "tasks.max": "1", "topics": "test-topic", "hostname": "clickhouse-31d766f9-systest-project.avns.net", "port": "14420", "database": "default", "username": "avnadmin", "password": "mypassword", "ssl": "true", "jdbcConnectionProperties": "custom_http_params=enable_http_compression=0" } ``` Once this configuration is saved in the `clickhouse_sink_connector.json` file, you can create the connector using the Aiven Console or CLI, and verify that data from the Apache Kafka topic `test-topic` is successfully delivered to your ClickHouse instance. --- # Configure AWS Secrets Manager Configure and use [AWS Secrets Manager](https://docs.aws.amazon.com/secretsmanager/latest/userguide/intro.html) as a secret provider in Aiven for Apache Kafka® Connect services. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Aiven for Apache Kafka service with Apache Kafka Connect](/docs/products/kafka/kafka-connect/get-started.md) set up and running. * [Aiven CLI](/docs/tools/cli.md). * [Aiven Terraform Provider](https://registry.terraform.io/providers/aiven/aiven/latest/docs) installed. * [AWS Secrets Manager access key and secret key](https://docs.aws.amazon.com/secretsmanager/latest/userguide/auth-and-access.html). note The integration with AWS Secrets Manager is not yet available on the Aiven Console. ## Configure secret providers[​](#configure-secret-providers "Direct link to Configure secret providers") Set up AWS Secrets Manager in your Aiven for Apache Kafka Connect service to manage and access sensitive information. * API * Terraform * CLI Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) API to update your service configuration. Add the AWS Secrets Manager configuration to the `user_config` using the following API request: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/{PROJECT_NAME}/service/{SERVICE_NAME} \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'Content-Type: application/json' \ --data '{ "user_config": { "secret_providers": [ { "name": "aws", "aws": { "auth_method": "credentials", "region": "your-aws-region", "access_key": "your-aws-access-key", "secret_key": "your-aws-secret-key" } } ] } }' ``` Parameters: * `url`: API endpoint for updating the service configuration. Replace `PROJECT_NAME` and `SERVICE_NAME` with your project and service names. * `Authorization`: Token used for authentication. Replace `YOUR_BEARER_TOKEN` with your [Aiven API token](/docs/platform/howto/create_authentication_token.md). * `Content-Type`: Specifies that the request body is in JSON format. * `auth_method`: Authentication method used by AWS Secrets Manager. In this case, it is `credentials`. * `region`: AWS region where your secrets are stored. * `access_key`: Your AWS access key. * `secret_key`: Your AWS secret key. Configure AWS Secrets Manager using Terraform. Add this configuration to your `main.tf` file, or optionally create a `secrets.tf` file to manage secret providers: 1. Create the `main.tf` file for your resources: ``` terraform { required_providers { aiven = { source = "aiven/aiven" version = ">=4.0.0, < 5.0.0" } } } provider "aiven" { api_token = var.aiven_api_token } resource "aiven_kafka_connect" "kafka_connect" { project = var.project_name cloud_name = "your-cloud-region" plan = "startup-4" service_name = "kafka-connect" kafka_connect_user_config { secret_providers { name = "aws" aws { auth_method = "credentials" region = var.aws_region access_key = var.aws_access_key secret_key = var.aws_secret_key } } } } ``` 2. Declare variables in a `variables.tf` file: ``` variable "aiven_api_token" { description = "Aiven API token" type = string } variable "project_name" { description = "Aiven project name" type = string } variable "aws_region" { description = "AWS region" type = string } variable "aws_access_key" { description = "AWS access key" type = string } variable "aws_secret_key" { description = "AWS secret key" type = string sensitive = true } ``` 3. Provide variable values in a `terraform.tfvars` file: ``` aiven_api_token = "YOUR_AIVEN_API_TOKEN" project_name = "YOUR_PROJECT_NAME" aws_region = "your-aws-region" aws_access_key = "your-aws-access-key" aws_secret_key = "your-aws-secret-key" ``` Parameters: * `project_name`: Name of your Aiven project. * `cloud_name`: Cloud provider and region for hosting Aiven for Apache Kafka Connect service. Replace with the appropriate region, for example `google-europe-west3`. * `plan`: Service plan for Aiven for Apache Kafka Connect service. * `service_name`: Name of your Aiven for Kafka Connect service in Aiven. * `auth_method`: Authentication method used by AWS Secrets Manager. Set to `credentials`. * `region`: AWS region where your secrets are stored. * `access_key`: AWS access key for authentication. * `secret_key`: AWS secret key for authentication. * `aiven_api_token`:API token for authentication. Replace with your Aiven API token to manage resources. Add AWS Secrets Manager using the [Aiven CLI](/docs/tools/cli.md): ``` avn service update SERVICE_NAME \ -c secret_providers='[ { "name": "aws", "aws": { "auth_method": "credentials", "region": "your-aws-region", "access_key": "your-aws-access-key", "secret_key": "your-aws-secret-key" } } ]' ``` Parameters: * `PROJECT_NAME`: Name of your Aiven project. * `SERVICE_NAME`: Name of your Aiven Kafka service. * `name`: Name of the secret provider. In this case, `aws`. * `auth_method`: Authentication method used by AWS Secrets Manager. In this case, it is credentials. * `region`: AWS region where your secrets are stored. * `access_key`: Your AWS access key. * `secret_key`: Your AWS secret key. ## Reference secrets in connector configurations[​](#reference-secrets-in-connector-configurations "Direct link to Reference secrets in connector configurations") You can use secrets stored in AWS Secrets Manager with any connector. The examples below show how to configure secrets for JDBC connectors, but you can follow the same steps for other connectors. ### JDBC sink connector[​](#jdbc-sink-connector "Direct link to JDBC sink connector") * API * Terraform * CLI Configure a JDBC sink connector using the API with secrets referenced from AWS Secrets Manager. ``` curl --request POST \ --url https://api.aiven.io/v1/project/{PROJECT_NAME}/service/{SERVICE_NAME}/connectors \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'Content-Type: application/json' \ --data '{ "name": "YOUR_CONNECTOR_NAME", "connector.class": "io.aiven.connect.jdbc.JdbcSinkConnector", "connection.url": "jdbc:{DATABASE_TYPE}://{HOST}:{PORT}/{DATABASE_NAME}?user=${aws:PATH/TO/SECRET:USERNAME}&password=${aws:PATH/TO/SECRET:PASSWORD}&ssl=require", "topics": "YOUR_TOPIC", "auto.create": true }' ``` Parameters: * `PROJECT_NAME`: Name of your Aiven project. * `SERVICE_NAME`: Name of your Aiven Kafka service. * `name`: Name of the connector. * `connector.class`: Specifies the connector class to use, in this case, `io.aiven.connect.jdbc.JdbcSinkConnector`. * `connection.url`: JDBC connection URL with placeholders for `DATABASE_TYPE`, `HOST`, `PORT`, `DATABASE_NAME`, and the username and password retrieved from AWS Secrets Manager. * `topics`: Apache Kafka topic where the data can be sent. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. Configure a JDBC sink connector using the Aiven Terraform Provider with secrets referenced from AWS Secrets Manager. Add this configuration to your `main.tf` file, or optionally create a dedicated `connectors.tf` file for managing Apache Kafka connectors: ``` resource "aiven_kafka_connector" "jdbc_sink_connector" { project = var.project_name service_name = aiven_kafka_connect.kafka_connect.service_name connector_name = "jdbc-sink-connector" config = { "connector.class" = "io.aiven.connect.jdbc.JdbcSinkConnector" "connection.url" = "jdbc:postgresql://{HOST}:{PORT}/{DATABASE_NAME}?user=${aws:PATH/TO/SECRET:USERNAME}&password=${aws:PATH/TO/SECRET:PASSWORD}&ssl=require" "topics" = "your-topic" "auto.create" = "true" } } ``` Parameters: * `project`: Name of your Aiven project. * `service_name`: Name of the Aiven for Apache Kafka service where the connector is to be created. * `connector_name`: Name for your JDBC sink connector. * `connector.class`: Java class that implements the connector. For JDBC sink connectors, use `"io.aiven.connect.jdbc.JdbcSinkConnector"`. * `connection.url`: JDBC URL for your database. Replace `{HOST}`, `{PORT}`, and `{DATABASE_NAME}` with your actual database details. The username and password are retrieved from AWS Secrets Manager using `${aws:PATH/TO/SECRET:USERNAME}` and `${aws:PATH/TO/SECRET:PASSWORD}`. * `topics`: Apache Kafka topics that the connector consumes data from. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. Configure a JDBC sink connector using the Aiven CLI with secrets referenced from AWS Secrets Manager. ``` avn service connector create SERVICE_NAME '{ "name": "jdbc-sink-connector", "connector.class": "io.aiven.connect.jdbc.JdbcSinkConnector", "connection.url": "jdbc:{DATABASE_TYPE}://{HOST}:{PORT}/{DATABASE_NAME}?user=${aws:PATH/TO/SECRET:USERNAME}&password=${aws:PATH/TO/SECRET:PASSWORD}&ssl=require", "topics": "your-topic", "auto.create": true }' ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `name`: Name of the connector. * `connector.class`: Specifies the connector class to use, in this case, `io.aiven.connect.jdbc.JdbcSinkConnector`. * `connection.url`: JDBC connection URL with placeholders for `DATABASE_TYPE`, `HOST`, `PORT`, `DATABASE_NAME`, and the username and password retrieved from AWS Secrets Manager. * `topics`: Apache Kafka topic where the data can be sent. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. ### JDBC source connector[​](#jdbc-source-connector "Direct link to JDBC source connector") * API * Terraform * CLI Configure a JDBC source connector using the API with secrets referenced from AWS Secrets Manager. ``` curl -X POST https://api.aiven.io/v1/project/{PROJECT_NAME}/service/{SERVICE_NAME}/connectors \ -H "Authorization: Bearer YOUR_BEARER_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "your-source-connector-name", "connector.class": "io.aiven.connect.jdbc.JdbcSourceConnector", "connection.url": "jdbc:{DATABASE_TYPE}://{HOST}:{PORT}/{DATABASE_NAME}?ssl=require", "connection.user": "${aws:PATH/TO/SECRET:USERNAME}", "connection.password": "${aws:PATH/TO/SECRET:PASSWORD}", "incrementing.column.name": "id", "mode": "incrementing", "table.whitelist": "your-table", "topic.prefix": "your-prefix_", "auto.create": "true" }' ``` Parameters: * `PROJECT_NAME`: Name of your Aiven project. * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `name`: Name of the connector. * `connector.class`: Specifies the connector class to use, in this case, `io.aiven.connect.jdbc.JdbcSinkConnector`. * `connection.url`: JDBC connection URL with placeholders for `DATABASE_TYPE`, `HOST`, `PORT`, `DATABASE_NAME`, and the username and password retrieved from AWS Secrets Manager. * `connection.user`: Database username retrieved from AWS Secrets Manager * `connection.password`: Database password retrieved from AWS Secrets Manager. * `incrementing.column.name`: Column used for incrementing mode. * `mode`: Mode of operation, in this case, `incrementing`. * `table.whitelist`: Tables to include. * `topic.prefix`: Prefix for Apache Kafka topics. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. Configure a JDBC source connector using the Aiven Terraform Provider with secrets referenced from AWS Secrets Manager. Add this configuration to your `main.tf` file, or optionally create a dedicated `connectors.tf` file for managing Apache Kafka connectors: ``` resource "aiven_kafka_connector" "jdbc_source_connector" { project = var.project_name service_name = aiven_kafka_connect.kafka_connect.service_name connector_name = "jdbc-source-connector" config = { "connector.class" = "io.aiven.connect.jdbc.JdbcSourceConnector" "connection.url" = "jdbc:postgresql://{HOST}:{PORT}/{DATABASE_NAME}?ssl=require" "connection.user" = "${aws:PATH/TO/SECRET:USERNAME}" "connection.password" = "${aws:PATH/TO/SECRET:PASSWORD}" "incrementing.column.name" = "id" "mode" = "incrementing" "table.whitelist" = "your-table" "topic.prefix" = "your-prefix_" "auto.create" = "true" } } ``` Parameters: * `project`: Name of your Aiven project. * `service_name`: Name of the Aiven for Apache Kafka service where the connector is to be created. * `connector_name`: Name for your JDBC source connector. * `connector.class`: Java class that implements the connector. For JDBC source connectors, use `"io.aiven.connect.jdbc.JdbcSourceConnector"`. * `connection.url`: JDBC URL for your database. Replace `{HOST}`, `{PORT}`, and `{DATABASE_NAME}` with your actual database details. The username and password are retrieved from AWS Secrets Manager using `${aws:PATH/TO/SECRET:USERNAME}` and `${aws:PATH/TO/SECRET:PASSWORD}`. * `connection.user`: Username for connecting to the database, retrieved from AWS Secrets Manager. * `connection.password`: Password for connecting to the database, retrieved from AWS Secrets Manager. * `incrementing.column.name`: Name of the column used for incrementing mode, typically a primary key column. * `mode`: Mode of operation for the connector. `incrementing` mode reads data incrementally based on the specified column. * `table.whitelist`: List of tables that the connector includes when reading data from the database. * `topic.prefix`: Prefix that the connector adds to the Kafka topic names it produces data to. * `auto.create`: If set to `true`, the connector automatically creates the target table in the database if it does not exist. Configure a JDBC source connector using the Aiven CLI with secrets referenced from AWS Secrets Manager. ``` avn service connector create SERVICE_NAME '{ "name": "jdbc-source-connector", "connector.class": "io.aiven.connect.jdbc.JdbcSourceConnector", "connection.url": "jdbc:{DATABASE_TYPE}://{HOST}:{PORT}/{DATABASE_NAME}?ssl=require", "connection.user": "${aws:PATH/TO/SECRET:USERNAME}", "connection.password": "${aws:PATH/TO/SECRET:PASSWORD}", "incrementing.column.name": "id", "mode": "incrementing", "table.whitelist": "your-table", "topic.prefix": "your-prefix_", "auto.create": true }' ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `name`: Name of the connector. * `connector.class`: Specifies the connector class to use, in this case, `io.aiven.connect.jdbc.JdbcSinkConnector`. * `connection.url`:JDBC connection URL with placeholders for `DATABASE_TYPE`, `HOST`, `PORT`, `DATABASE_NAME`, and the username and password retrieved from AWS Secrets Manager. * `connection.user`: Database username retrieved from AWS Secrets Manager * `connection.password`: Database password retrieved from AWS Secrets Manager. * `incrementing.column.name`: Column used for incrementing mode. * `mode`: Mode of operation, in this case, `incrementing`. * `table.whitelist`: Tables to include. * `topic.prefix`: Prefix for Apache Kafka topics. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. --- # Configure Azure Key Vault Configure and use [Azure Key Vault](https://learn.microsoft.com/en-us/azure/key-vault/general/overview) as a secret provider in Aiven for Apache Kafka® Connect services. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Aiven for Apache Kafka service with Apache Kafka Connect](/docs/products/kafka/kafka-connect/get-started.md) set up and running. * [Aiven CLI](/docs/tools/cli.md). * [Aiven Terraform Provider](https://registry.terraform.io/providers/aiven/aiven/latest/docs) installed. * Azure Key Vault with the following: * A [service principal](https://learn.microsoft.com/en-us/azure/active-directory/develop/howto-create-service-principal-portal) with access to the Key Vault. * The service principal's client ID, tenant ID, and client secret. note The integration with Azure Key Vault is not yet available on the Aiven Console. important Secrets stored in Azure Key Vault must have an expiration date set. Secrets without an expiration date cause the connector to restart continuously. ## Configure secret providers[​](#configure-secret-providers "Direct link to Configure secret providers") Set up Azure Key Vault in your Aiven for Apache Kafka Connect service to manage and access sensitive information. * API * Terraform * CLI Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) API to update your service configuration. Add the Azure Key Vault configuration to the `user_config` using the following API request: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/{PROJECT_NAME}/service/{SERVICE_NAME} \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'Content-Type: application/json' \ --data '{ "user_config": { "secret_providers": [ { "name": "azure", "azure": { "auth_method": "credentials", "client_id": "your-azure-client-id", "tenant_id": "your-azure-tenant-id", "secret": "your-azure-client-secret" } } ] } }' ``` Parameters: * `url`: API endpoint for updating the service configuration. Replace `PROJECT_NAME` and `SERVICE_NAME` with your project and service names. * `Authorization`: Token used for authentication. Replace `YOUR_BEARER_TOKEN` with your [Aiven API token](/docs/platform/howto/create_authentication_token.md). * `Content-Type`: Specifies that the request body is in JSON format. * `auth_method`: Authentication method used by Azure Key Vault. Set to `credentials`. * `client_id`: Azure service principal client ID. * `tenant_id`: Azure service principal tenant ID. * `secret`: Azure service principal client secret. Configure Azure Key Vault using Terraform. Add this configuration to your `main.tf` file, or optionally create a `secrets.tf` file to manage secret providers: 1. Create the `main.tf` file for your resources: ``` terraform { required_providers { aiven = { source = "aiven/aiven" version = ">=4.0.0, < 5.0.0" } } } provider "aiven" { api_token = var.aiven_api_token } resource "aiven_kafka_connect" "kafka_connect" { project = var.project_name cloud_name = "your-cloud-region" plan = "startup-4" service_name = "kafka-connect" kafka_connect_user_config { secret_providers { name = "azure" azure { auth_method = "credentials" client_id = var.azure_client_id tenant_id = var.azure_tenant_id secret = var.azure_client_secret } } } } ``` 2. Declare variables in a `variables.tf` file: ``` variable "aiven_api_token" { description = "Aiven API token" type = string } variable "project_name" { description = "Aiven project name" type = string } variable "azure_client_id" { description = "Azure service principal client ID" type = string } variable "azure_tenant_id" { description = "Azure service principal tenant ID" type = string } variable "azure_client_secret" { description = "Azure service principal client secret" type = string sensitive = true } ``` 3. Provide variable values in a `terraform.tfvars` file: ``` aiven_api_token = "YOUR_AIVEN_API_TOKEN" project_name = "YOUR_PROJECT_NAME" azure_client_id = "your-azure-client-id" azure_tenant_id = "your-azure-tenant-id" azure_client_secret = "your-azure-client-secret" ``` Parameters: * `project_name`: Name of your Aiven project. * `cloud_name`: Cloud provider and region for hosting Aiven for Apache Kafka Connect service. Replace with the appropriate region, for example `azure-westeurope`. * `plan`: Service plan for Aiven for Apache Kafka Connect service. * `service_name`: Name of your Aiven for Kafka Connect service in Aiven. * `auth_method`: Authentication method used by Azure Key Vault. Set to `credentials`. * `client_id`: Azure service principal client ID. * `tenant_id`: Azure service principal tenant ID. * `secret`: Azure service principal client secret. * `aiven_api_token`: API token for authentication. Replace with your Aiven API token to manage resources. Add Azure Key Vault using the [Aiven CLI](/docs/tools/cli.md): ``` avn service update SERVICE_NAME \ -c secret_providers='[ { "name": "azure", "azure": { "auth_method": "credentials", "client_id": "your-azure-client-id", "tenant_id": "your-azure-tenant-id", "secret": "your-azure-client-secret" } } ]' ``` Parameters: * `SERVICE_NAME`: Name of your Aiven Kafka service. * `name`: Name of the secret provider. In this case, `azure`. * `auth_method`: Authentication method used by Azure Key Vault. Set to `credentials`. * `client_id`: Azure service principal client ID. * `tenant_id`: Azure service principal tenant ID. * `secret`: Azure service principal client secret. ## Reference secrets in connector configurations[​](#reference-secrets-in-connector-configurations "Direct link to Reference secrets in connector configurations") You can use secrets stored in Azure Key Vault with any connector. The format for referencing secrets is: ``` ${PROVIDER_NAME:VAULT_NAME.vault.azure.net:SECRET_NAME} ``` Parameters: * `PROVIDER_NAME`: Name of your secret provider configuration, such as `azure`. * `VAULT_NAME.vault.azure.net`: Your Azure Key Vault hostname without `https://`. * `SECRET_NAME`: Name of the secret in Azure Key Vault. The examples below show how to configure secrets for JDBC connectors, but you can follow the same steps for other connectors. ### JDBC sink connector[​](#jdbc-sink-connector "Direct link to JDBC sink connector") * API * Terraform * CLI Configure a JDBC sink connector using the API with secrets referenced from Azure Key Vault. ``` curl --request POST \ --url https://api.aiven.io/v1/project/{PROJECT_NAME}/service/{SERVICE_NAME}/connectors \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'Content-Type: application/json' \ --data '{ "name": "YOUR_CONNECTOR_NAME", "connector.class": "io.aiven.connect.jdbc.JdbcSinkConnector", "connection.url": "jdbc:{DATABASE_TYPE}://{HOST}:{PORT}/{DATABASE_NAME}?user=${azure:your-vault.vault.azure.net:db-username}&password=${azure:your-vault.vault.azure.net:db-password}&ssl=require", "topics": "YOUR_TOPIC", "auto.create": true }' ``` Parameters: * `PROJECT_NAME`: Name of your Aiven project. * `SERVICE_NAME`: Name of your Aiven Kafka service. * `name`: Name of the connector. * `connector.class`: Specifies the connector class to use, in this case, `io.aiven.connect.jdbc.JdbcSinkConnector`. * `connection.url`: JDBC connection URL with placeholders for `DATABASE_TYPE`, `HOST`, `PORT`, `DATABASE_NAME`, and the username and password retrieved from Azure Key Vault. * `topics`: Apache Kafka topic where the data can be sent. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. Configure a JDBC sink connector using the Aiven Terraform Provider with secrets referenced from Azure Key Vault. Add this configuration to your `main.tf` file, or optionally create a dedicated `connectors.tf` file for managing Apache Kafka connectors: ``` resource "aiven_kafka_connector" "jdbc_sink_connector" { project = var.project_name service_name = aiven_kafka_connect.kafka_connect.service_name connector_name = "jdbc-sink-connector" config = { "connector.class" = "io.aiven.connect.jdbc.JdbcSinkConnector" "connection.url" = "jdbc:postgresql://{HOST}:{PORT}/{DATABASE_NAME}?user=$${azure:your-vault.vault.azure.net:db-username}&password=$${azure:your-vault.vault.azure.net:db-password}&ssl=require" "topics" = "your-topic" "auto.create" = "true" } } ``` Parameters: * `project`: Name of your Aiven project. * `service_name`: Name of the Aiven for Apache Kafka service where the connector is to be created. * `connector_name`: Name for your JDBC sink connector. * `connector.class`: Java class that implements the connector. For JDBC sink connectors, use `"io.aiven.connect.jdbc.JdbcSinkConnector"`. * `connection.url`: JDBC URL for your database. Replace `{HOST}`, `{PORT}`, and `{DATABASE_NAME}` with your actual database details. The username and password are retrieved from Azure Key Vault using `${azure:your-vault.vault.azure.net:db-username}` and `${azure:your-vault.vault.azure.net:db-password}`. * `topics`: Apache Kafka topics that the connector consumes data from. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. Configure a JDBC sink connector using the Aiven CLI with secrets referenced from Azure Key Vault. ``` avn service connector create SERVICE_NAME '{ "name": "jdbc-sink-connector", "connector.class": "io.aiven.connect.jdbc.JdbcSinkConnector", "connection.url": "jdbc:{DATABASE_TYPE}://{HOST}:{PORT}/{DATABASE_NAME}?user=${azure:your-vault.vault.azure.net:db-username}&password=${azure:your-vault.vault.azure.net:db-password}&ssl=require", "topics": "your-topic", "auto.create": true }' ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `name`: Name of the connector. * `connector.class`: Specifies the connector class to use, in this case, `io.aiven.connect.jdbc.JdbcSinkConnector`. * `connection.url`: JDBC connection URL with placeholders for `DATABASE_TYPE`, `HOST`, `PORT`, `DATABASE_NAME`, and the username and password retrieved from Azure Key Vault. * `topics`: Apache Kafka topic where the data can be sent. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. ### JDBC source connector[​](#jdbc-source-connector "Direct link to JDBC source connector") * API * Terraform * CLI Configure a JDBC source connector using the API with secrets referenced from Azure Key Vault. ``` curl -X POST https://api.aiven.io/v1/project/{PROJECT_NAME}/service/{SERVICE_NAME}/connectors \ -H "Authorization: Bearer YOUR_BEARER_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "your-source-connector-name", "connector.class": "io.aiven.connect.jdbc.JdbcSourceConnector", "connection.url": "jdbc:{DATABASE_TYPE}://{HOST}:{PORT}/{DATABASE_NAME}?ssl=require", "connection.user": "${azure:your-vault.vault.azure.net:db-username}", "connection.password": "${azure:your-vault.vault.azure.net:db-password}", "incrementing.column.name": "id", "mode": "incrementing", "table.whitelist": "your-table", "topic.prefix": "your-prefix_", "auto.create": "true" }' ``` Parameters: * `PROJECT_NAME`: Name of your Aiven project. * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `name`: Name of the connector. * `connector.class`: Specifies the connector class to use, in this case, `io.aiven.connect.jdbc.JdbcSourceConnector`. * `connection.url`: JDBC connection URL with placeholders for `DATABASE_TYPE`, `HOST`, `PORT`, and `DATABASE_NAME`. * `connection.user`: Database username retrieved from Azure Key Vault. * `connection.password`: Database password retrieved from Azure Key Vault. * `incrementing.column.name`: Column used for incrementing mode. * `mode`: Mode of operation, in this case, `incrementing`. * `table.whitelist`: Tables to include. * `topic.prefix`: Prefix for Apache Kafka topics. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. Configure a JDBC source connector using the Aiven Terraform Provider with secrets referenced from Azure Key Vault. Add this configuration to your `main.tf` file, or optionally create a dedicated `connectors.tf` file for managing Apache Kafka connectors: ``` resource "aiven_kafka_connector" "jdbc_source_connector" { project = var.project_name service_name = aiven_kafka_connect.kafka_connect.service_name connector_name = "jdbc-source-connector" config = { "connector.class" = "io.aiven.connect.jdbc.JdbcSourceConnector" "connection.url" = "jdbc:postgresql://{HOST}:{PORT}/{DATABASE_NAME}?ssl=require" "connection.user" = "$${azure:your-vault.vault.azure.net:db-username}" "connection.password" = "$${azure:your-vault.vault.azure.net:db-password}" "incrementing.column.name" = "id" "mode" = "incrementing" "table.whitelist" = "your-table" "topic.prefix" = "your-prefix_" "auto.create" = "true" } } ``` Parameters: * `project`: Name of your Aiven project. * `service_name`: Name of the Aiven for Apache Kafka service where the connector is to be created. * `connector_name`: Name for your JDBC source connector. * `connector.class`: Java class that implements the connector. For JDBC source connectors, use `"io.aiven.connect.jdbc.JdbcSourceConnector"`. * `connection.url`: JDBC URL for your database. Replace `{HOST}`, `{PORT}`, and `{DATABASE_NAME}` with your actual database details. * `connection.user`: Username for connecting to the database, retrieved from Azure Key Vault. * `connection.password`: Password for connecting to the database, retrieved from Azure Key Vault. * `incrementing.column.name`: Name of the column used for incrementing mode, typically a primary key column. * `mode`: Mode of operation for the connector. `incrementing` mode reads data incrementally based on the specified column. * `table.whitelist`: List of tables that the connector includes when reading data from the database. * `topic.prefix`: Prefix that the connector adds to the Kafka topic names it produces data to. * `auto.create`: If set to `true`, the connector automatically creates the target table in the database if it does not exist. Configure a JDBC source connector using the Aiven CLI with secrets referenced from Azure Key Vault. ``` avn service connector create SERVICE_NAME '{ "name": "jdbc-source-connector", "connector.class": "io.aiven.connect.jdbc.JdbcSourceConnector", "connection.url": "jdbc:{DATABASE_TYPE}://{HOST}:{PORT}/{DATABASE_NAME}?ssl=require", "connection.user": "${azure:your-vault.vault.azure.net:db-username}", "connection.password": "${azure:your-vault.vault.azure.net:db-password}", "incrementing.column.name": "id", "mode": "incrementing", "table.whitelist": "your-table", "topic.prefix": "your-prefix_", "auto.create": true }' ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `name`: Name of the connector. * `connector.class`: Specifies the connector class to use, in this case, `io.aiven.connect.jdbc.JdbcSourceConnector`. * `connection.url`: JDBC connection URL with placeholders for `DATABASE_TYPE`, `HOST`, `PORT`, and `DATABASE_NAME`. * `connection.user`: Database username retrieved from Azure Key Vault. * `connection.password`: Database password retrieved from Azure Key Vault. * `incrementing.column.name`: Column used for incrementing mode. * `mode`: Mode of operation, in this case, `incrementing`. * `table.whitelist`: Tables to include. * `topic.prefix`: Prefix for Apache Kafka topics. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. --- # Configure the ENV secret provider Configure and use the ENV secret provider in Aiven for Apache Kafka® Connect services. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Aiven for Apache Kafka service with Apache Kafka Connect](/docs/products/kafka/kafka-connect/get-started.md) set up and running. * [Aiven CLI](/docs/tools/cli.md). * [Aiven Terraform Provider](https://registry.terraform.io/providers/aiven/aiven/latest/docs) installed. note The ENV secret provider is not yet available in the Aiven Console. ## Configure the secret provider[​](#configure-the-secret-provider "Direct link to Configure the secret provider") Configure the ENV secret provider in your Aiven for Apache Kafka Connect service to store and reference secrets in `user_config`. * API * Terraform * CLI Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) API to update your service configuration. Add the ENV secret provider configuration to `user_config` with the following API request: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer AIVEN_API_TOKEN' \ --header 'Content-Type: application/json' \ --data '{ "user_config": { "secret_providers": [ { "name": "db_credentials", "env": { "secrets": { "db_password": "DB_PASSWORD_VALUE" } } } ] } }' ``` Parameters: * `name`: Name of the secret provider, for example `db_credentials`. * `env.secrets`: Map of secret keys and values stored in `user_config`. * `db_password`: Secret key that you use later in connector configuration. Configure the ENV secret provider using Terraform. Add this configuration to your `main.tf` file, or create a dedicated file for secret providers: ``` resource "aiven_kafka_connect" "kafka_connect" { project = var.project_name cloud_name = var.cloud_name plan = var.plan service_name = var.service_name kafka_connect_user_config { secret_providers { name = "db_credentials" env { secrets = { db_password = var.db_password } } } } } ``` Parameters: * `name`: Name of the secret provider, for example `db_credentials`. * `env.secrets`: Map of secret keys and values. * `db_password`: Terraform variable containing the secret value. Add the ENV secret provider using the [Aiven CLI](/docs/tools/cli.md): ``` avn service update SERVICE_NAME \ -c secret_providers='[ { "name": "db_credentials", "env": { "secrets": { "db_password": "DB_PASSWORD_VALUE" } } } ]' ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `name`: Name of the secret provider, for example `db_credentials`. * `env.secrets`: Map of secret keys and values. ## Reference secrets in connector configurations[​](#reference-secrets-in-connector-configurations "Direct link to Reference secrets in connector configurations") Reference secrets in connector configuration values using the provider name and secret key. Use the syntax `${PROVIDER_NAME:SECRET_KEY}`. Example values: * **Provider name**: `db_credentials` * **Secret key**: `db_password` * **Secret reference**: `${db_credentials:db_password}` ### JDBC sink connector[​](#jdbc-sink-connector "Direct link to JDBC sink connector") Example JDBC sink connector configuration that references a secret from the ENV secret provider. ``` { "name": "jdbc-sink-connector", "connector.class": "io.aiven.connect.jdbc.JdbcSinkConnector", "connection.url": "jdbc:postgresql://DB_HOST:5432/DB_NAME?user=DB_USER&password=${db_credentials:db_password}&ssl=require", "topics": "YOUR_TOPIC", "auto.create": true } ``` ### JDBC source connector[​](#jdbc-source-connector "Direct link to JDBC source connector") Example JDBC source connector configuration that references a secret from the ENV secret provider. ``` { "name": "jdbc-source-connector", "connector.class": "io.aiven.connect.jdbc.JdbcSourceConnector", "connection.url": "jdbc:postgresql://DB_HOST:5432/DB_NAME?ssl=require", "connection.user": "DB_USER", "connection.password": "${db_credentials:db_password}", "mode": "incrementing", "incrementing.column.name": "id", "table.whitelist": "YOUR_TABLE", "topic.prefix": "jdbc_" } ``` ## Security behavior[​](#security-behavior "Direct link to Security behavior") The ENV secret provider stores secrets in encrypted form at rest. The service decrypts secrets in memory only when a connector resolves them at runtime. ## Base64 encoding for complex secret values[​](#base64-encoding-for-complex-secret-values "Direct link to Base64 encoding for complex secret values") If your secret value contains complex strings such as JSON, use base64 encoding to avoid escaping issues. Use the format `ENV-base64:BASE64_ENCODED_VALUE`. The secret provider automatically decodes base64-encoded values at runtime. ### Example: JSON credential[​](#example-json-credential "Direct link to Example: JSON credential") If you need to store a JSON credential as a secret: 1. Create your JSON value: ``` { "username": "USER_NAME", "password": "PASSWORD", "api_key": "API_KEY_VALUE" } ``` 2. Encode it with base64: ``` echo '{"username":"user","password":"p@ssw0rd","api_key":"sk-1234567890"}' | base64 ``` Output example: ``` eyJ1c2VybmFtZSI6InVzZXIiLCJwYXNzd29yZCI6InBAc3N3MHJkIiwiYXBpX2tleSI6InNrLTEyMzQ1Njc4OTAifQ== ``` 3. Add the encoded value to your secret provider configuration with the `ENV-base64:` prefix: * API * Terraform * CLI ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer AIVEN_API_TOKEN' \ --header 'Content-Type: application/json' \ --data '{ "user_config": { "secret_providers": [ { "name": "api_credentials", "env": { "secrets": { "api_config": "ENV-base64:BASE64_ENCODED_VALUE" } } } ] } }' ``` ``` resource "aiven_kafka_connect" "kafka_connect" { project = var.project_name cloud_name = var.cloud_name plan = var.plan service_name = var.service_name kafka_connect_user_config { secret_providers { name = "api_credentials" env { secrets = { api_config = "ENV-base64:BASE64_ENCODED_VALUE" } } } } } ``` ``` avn service update SERVICE_NAME \ -c secret_providers='[ { "name": "api_credentials", "env": { "secrets": { "api_config": "ENV-base64:BASE64_ENCODED_VALUE" } } } ]' ``` The secret provider automatically decodes the base64 value and resolves it to the original value when a connector references the secret. --- # Configure HashiCorp Vault Configure and use [HashiCorp Vault](https://developer.hashicorp.com/vault/docs) as a secret provider in Aiven for Apache Kafka® Connect services. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to the [Aiven Console](https://console.aiven.io/). * [Aiven for Apache Kafka service with Apache Kafka Connect](/docs/products/kafka/kafka-connect/get-started.md) set up and running. * [Aiven CLI](/docs/tools/cli.md). * [Aiven Terraform Provider](https://registry.terraform.io/providers/aiven/aiven/latest/docs) installed. * [HashiCorp Vault address and token](https://developer.hashicorp.com/vault/docs/concepts/tokens). note The integration with HashiCorp Vault is not yet available on the Aiven Console. ## Configure secret providers[​](#configure-secret-providers "Direct link to Configure secret providers") Set up HashiCorp Vault in your Aiven for Apache Kafka Connect service to manage and access sensitive information. * API * Terraform * CLI Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) API to update your service configuration. Add the HashiCorp Vault configuration to the `user_config` using the following API request: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/{PROJECT_NAME}/service/{SERVICE_NAME} \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'Content-Type: application/json' \ --data '{ "user_config": { "secret_providers": [ { "name": "vault", "vault": { "auth_method": "token", "address": "https://vault.aiven.fi:8200/ "token": "YOUR_VAULT_TOKEN" } } ] } }' ``` Parameters: * `PROJECT_NAME`: Name of your Aiven project. * `SERVICE_NAME`: Name of your Aiven Kafka service. * `url`: API endpoint for updating service configuration. Replace `{project_name}` and `{service_name}` with your project and service names. * `Authorization`: Token used for authentication. Replace `YOUR_BEARER_TOKEN` with your [Aiven API token](/docs/platform/howto/create_authentication_token.md). * `Content-Type`: Specifies that the request body is in JSON format. * `auth_method`: Authentication method used by HashiCorp Vault. In this case, it is token. * `address`: Address of the HashiCorp Vault server. * `token`: Your HashiCorp Vault token. Configure HashiCorp Vault using the Aiven Terraform Provider. Add this configuration to your `main.tf` file, or optionally create a dedicated `vault.tf` file for managing secret providers: 1. Create the `main.tf` file for your resources: ``` terraform { required_providers { aiven = { source = "aiven/aiven" version = ">=4.0.0, < 5.0.0" } } } provider "aiven" { api_token = var.aiven_api_token } resource "aiven_kafka_connect" "kafka_connect" { project = var.project_name cloud_name = "your-cloud-region" plan = "startup-4" service_name = "kafka-connect" kafka_connect_user_config { secret_providers { name = "vault" vault { auth_method = "token" address = "https://vault.aiven.fi:8200/" token = var.vault_token } } } } ``` 2. Declare variables in a `variables.tf` file: ``` variable "aiven_api_token" { description = "Aiven API token" type = string } variable "project_name" { description = "Aiven project name" type = string } variable "vault_token" { description = "HashiCorp Vault token" type = string sensitive = true } ``` 3. Provide variable values in a `terraform.tfvars` file: ``` aiven_api_token = "YOUR_AIVEN_API_TOKEN" project_name = "YOUR_PROJECT_NAME" vault_token = "YOUR_VAULT_TOKEN" ``` Parameters: * `project_name`: Name of your Aiven project. * `cloud_name`: Cloud provider and region for hosting the Aiven for Apache Kafka service. Replace with the appropriate region. For example, `google-europe-west3`. * `plan`: Service plan for the Aiven for Apache Kafka Connect service. * `service_name`: Name of your Aiven for Apache Kafka Connect service. * `auth_method`: Authentication method used by HashiCorp Vault. Set to `token`. * `address`: URL address of the HashiCorp Vault server. * `token`: Token used for authenticating with HashiCorp Vault. * `aiven_api_token`: API token for authentication. Replace with your Aiven API token to manage resources. Configure HashiCorp Vault as a secret provider using [Aiven CLI](/docs/tools/cli.md): ``` avn service update SERVICE_NAME \ -c secret_providers='[ { "vault": { "auth_method": "token", "address": "https://vault.aiven.fi:8200/", "token": "YOUR_VAULT_TOKEN" }, "name": "vault" } ]' ``` Parameters: * `SERVICE_NAME`: Name of your Aiven Kafka service.. * `name`: Name of the secret provider. In this case, `vault`. * `auth_method`: Authentication method used by HashiCorp Vault. In this case, it is `token`. * `address`: Address of the HashiCorp Vault server. * `token`: Your HashiCorp Vault token. ## Reference secrets in connector configurations[​](#reference-secrets-in-connector-configurations "Direct link to Reference secrets in connector configurations") You can use secrets stored in HashiCorp Vault with any connector. The examples below show how to configure secrets for JDBC connectors, but you can follow the same steps for other connectors. ### JDBC sink connector[​](#jdbc-sink-connector "Direct link to JDBC sink connector") * API * Terraform * CLI Configure a JDBC sink connector using the API with secrets referenced from HashiCorp Vault: ``` curl -X POST https://api.aiven.io/v1/project/{PROJECT_NAME}/service/{SERVICE_NAME}/connectors \ -H "Authorization: Bearer YOUR_BEARER_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "your-connector-name", "connector.class": "io.aiven.connect.jdbc.JdbcSinkConnector", "connection.url": "jdbc:{DATABASE_TYPE}://{HOST}:{PORT}/{DATABASE_NAME}?user=${vault:PATH/TO/SECRET:USERNAME}&password=${vault:PATH/TO/SECRET:PASSWORD}&ssl=require", "topics": "your-topic", "auto.create": true }' ``` Parameters: * `PROJECT_NAME`: Name of your Aiven project. * `SERVICE_NAME`: Name of your Aiven Kafka service. * `name`: Name of the connector. * `connector.class`: Specifies the connector class to use, in this case, `io.aiven.connect.jdbc.JdbcSinkConnector`. * `connection.url`: JDBC connection URL with placeholders for `DATABASE_TYPE`, `HOST`, `PORT`, `DATABASE_NAME`, and the username and password retrieved from HashiCorp Vault. * `topics`: Apache Kafka topic where the data can be sent. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. Configure a JDBC sink connector using the Aiven Terraform Provider with secrets referenced from HashiCorp Vault. Add this configuration to your `main.tf` file, or optionally create a dedicated `vault.tf` file for managing secret providers: ``` resource "aiven_kafka_connector" "jdbc_sink_connector" { project = var.project_name service_name = aiven_kafka_connect.kafka_connect.service_name connector_name = "jdbc-sink-connector" config = { "connector.class" = "io.aiven.connect.jdbc.JdbcSinkConnector" "connection.url" = "jdbc:postgresql://{HOST}:{PORT}/{DATABASE_NAME}?user=${vault:PATH/TO/SECRET:USERNAME}&password=${vault:PATH/TO/SECRET:PASSWORD}&ssl=require" "topics" = "your-topic" "auto.create" = "true" } } ``` Parameters: * `project`: Name of your Aiven project where the Kafka Connect service is located. * `service_name`: Name of the Aiven for Kafka Connect service where the connector is to be created. * `connector_name`: Name for your JDBC sink connector. * `connector.class`: The Java class that implements the connector. For JDBC sink connectors, use `"io.aiven.connect.jdbc.JdbcSinkConnector"`. * `connection.url`: The JDBC URL for your database. Replace `{HOST}`, `{PORT}`, and `{DATABASE_NAME}` with your actual database details. The username and password are retrieved from HashiCorp Vault using the `${vault:PATH/TO/SECRET:USERNAME}` and `${vault:PATH/TO/SECRET:PASSWORD}` syntax. * `topics`: Apache Kafka topics that the connector consumes data from. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. Configure a JDBC sink connector using the Aiven CLI with secrets referenced from HashiCorp Vault. ``` avn service connector create SERVICE_NAME '{ "name": "jdbc-sink-connector", "connector.class": "io.aiven.connect.jdbc.JdbcSinkConnector", "connection.url": "jdbc:{DATABASE_TYPE}://{HOST}:{PORT}/{DATABASE_NAME}?user=${vault:PATH/TO/SECRET:USERNAME}&password=${vault:PATH/TO/SECRET:PASSWORD}&ssl=require", "topics": "your-topic", "auto.create": true ``` Parameters: * `SERVICE_NAME`: Name of your Aiven Kafka service. * `name`: Name of the connector. * `connector.class`: Specifies the connector class to use, in this case, `io.aiven.connect.jdbc.JdbcSinkConnector`. * `connection.url`: JDBC connection URL with placeholders for `DATABASE_TYPE`, `HOST`, `PORT`, `DATABASE_NAME`, and the username and password retrieved from HashiCorp Vault. * `topics`: Apache Kafka topic where the data can be sent. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. ### JDBC source connector[​](#jdbc-source-connector "Direct link to JDBC source connector") * API * Terraform * CLI Configure a JDBC source connector using the API with secrets referenced from HashiCorp Vault: ``` curl -X POST https://api.aiven.io/v1/project/{PROJECT_NAME}/service/{SERVICE_NAME}/connectors \ -H "Authorization: Bearer YOUR_BEARER_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "name": "your-connector-name", "connector.class": "io.aiven.connect.jdbc.JdbcSourceConnector", "connection.url": "jdbc:{DATABASE_TYPE}://{HOST}:{PORT}/{DATABASE_NAME}?ssl=require", "connection.user": "${vault:PATH/TO/SECRET:USERNAME}", "connection.password": "${vault:PATH/TO/SECRET:PASSWORD}", "incrementing.column.name": "id", "mode": "incrementing", "table.whitelist": "your-table", "topic.prefix": "your-prefix_", "auto.create": true }' ``` Parameters: * `PROJECT_NAME`: Name of your Aiven project. * `SERVICE_NAME`: Name of your Aiven Kafka service. * `name`: Name of the connector. * `connector.class`: Specifies the connector class to use, in this case, `io.aiven.connect.jdbc.JdbcSinkConnector`. * `connection.url`: JDBC connection URL with placeholders for `DATABASE_TYPE`, `HOST`, `PORT`, `DATABASE_NAME`, and the username and password retrieved from HashiCorp Vault. * `connection.user`: Database username retrieved from HashiCorp Vault. * `connection.password`: Database password retrieved from HashiCorp Vault. * `incrementing.column.name`: Column used for incrementing mode. * `mode`: Mode of operation, in this case, `incrementing`. * `table.whitelist`: Tables to include. * `topic.prefix`: Prefix for Apache Kafka topics. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. Configure a JDBC source connector using the Aiven Terraform Provider with secrets referenced from HashiCorp Vault. Add this configuration to your `main.tf` file, or optionally create a dedicated `connectors.tf` file for managing Apache Kafka connectors: ``` resource "aiven_kafka_connector" "jdbc_source_connector" { project = var.project_name service_name = aiven_kafka_connect.kafka_connect.service_name connector_name = "jdbc-source-connector" config = { "connector.class" = "io.aiven.connect.jdbc.JdbcSourceConnector" "connection.url" = "jdbc:postgresql://{HOST}:{PORT}/{DATABASE_NAME}?ssl=require" "connection.user" = "${vault:PATH/TO/SECRET:USERNAME}" "connection.password" = "${vault:PATH/TO/SECRET:PASSWORD}" "incrementing.column.name" = "id" "mode" = "incrementing" "table.whitelist" = "your-table" "topic.prefix" = "your-prefix_" "auto.create" = "true" } } ``` Parameters: * `project`: Name of your Aiven project where the Kafka Connect service is located. * `service_name`: Name of the Kafka Connect service where the connector is to be created. * `connector_name`: Name for your JDBC source connector. * `connector.class`: The Java class that implements the connector. For JDBC source connectors, use `"io.aiven.connect.jdbc.JdbcSourceConnector"`. * `connection.url`: The JDBC URL for your database. Replace `{HOST}`, `{PORT}`, and `{DATABASE_NAME}` with your actual database details. The username and password are retrieved from HashiCorp Vault using the `${vault:PATH/TO/SECRET:USERNAME}` and `${vault:PATH/TO/SECRET:PASSWORD}` syntax. * `connection.user`: The username for connecting to the database, retrieved from HashiCorp Vault. * `connection.password`: The password for connecting to the database, retrieved from HashiCorp Vault. * `incrementing.column.name`: The name of the column that the connector will use for incrementing mode, typically a primary key column. * `mode`: The mode of operation for the connector. `incrementing` mode reads data incrementally based on the specified column. * `table.whitelist`: A list of tables that the connector includes when reading data from the database. * `topic.prefix`: Prefix for Apache Kafka topics. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. Configure a JDBC source connector using the Aiven CLI with secrets referenced from HashiCorp Vault: ``` avn service connector create SERVICE_NAME '{ "name": "jdbc-source-connector", "connector.class": "io.aiven.connect.jdbc.JdbcSourceConnector", "connection.url": "jdbc:{DATABASE_TYPE}://{HOST}:{PORT}/{DATABASE_NAME}?ssl=require", "connection.user": "${vault:PATH/TO/SECRET:USERNAME}", "connection.password": "${vault:PATH/TO/SECRET:PASSWORD}", "incrementing.column.name": "id", "mode": "incrementing", "table.whitelist": "your-table", "topic.prefix": "your-prefix_", "auto.create": true }' ``` Parameters: * `SERVICE_NAME`: Name of your Aiven Kafka service. * `name`: Name of the connector. * `connector.class`: Specifies the connector class to use, in this case, `io.aiven.connect.jdbc.JdbcSinkConnector`. * `connection.url`: JDBC connection URL with placeholders for `DATABASE_TYPE`, `HOST`, `PORT`, `DATABASE_NAME`, and the username and password retrieved from HashiCorp Vault. * `connection.user`: Database username retrieved from HashiCorp Vault. * `connection.password`: Database password retrieved from HashiCorp Vault. * `incrementing.column.name`: Column used for incrementing mode. * `mode`: Mode of operation, in this case, `incrementing`. * `table.whitelist`: Tables to include. * `topic.prefix`: Prefix for Apache Kafka topics. * `auto.create`: If `true`, the connector automatically creates the table in the target database if it does not exist. --- # Aiven for Apache Kafka® Connect secret providers Configure and use secret providers in Apache Kafka Connect services on Aiven for Apache Kafka. Securely reference secrets stored in external secret managers within your connector configurations, ensuring that sensitive information is not stored in plain text. ## What are secret providers?[​](#what-are-secret-providers "Direct link to What are secret providers?") Secret providers are tools that manage sensitive information, such as passwords and API keys, in a secure manner. Instead of directly including these secrets in your configuration files, secret providers allow you to store them securely in external secret managers like AWS Secrets Manager and HashiCorp Vault, or in encrypted service configuration with the ENV secret provider. Aiven for Apache Kafka Connect dynamically retrieves these secrets when needed, enhancing the security of your setup. ## Supported secret managers[​](#supported-secret-managers "Direct link to Supported secret managers") * [AWS Secrets Manager](/docs/products/kafka/kafka-connect/howto/configure-aws-secrets-manager.md) * **Auth method**: `credentials` * **Required parameters**: `access key`, `secret key` * [Azure Key Vault](/docs/products/kafka/kafka-connect/howto/configure-azure-key-vault.md) * **Auth method**: `credentials` * **Required parameters**: `client_id`, `tenant_id`, `secret` * [HashiCorp Vault](/docs/products/kafka/kafka-connect/howto/configure-hashicorp-vault.md) * **Auth method**: `token` * **Required parameters**: `token`, `address` * [ENV secret provider](/docs/products/kafka/kafka-connect/howto/configure-env-secret-provider.md) * **Auth method**: not applicable * **Required parameters**: `name`, `env.secrets` --- # Create a sink connector from Apache Kafka® to Couchbase The [Couchbase](https://www.couchbase.com/) sink connector pushes Apache Kafka® data to the NoSQL database. note See the full set of available parameters and configuration options in the [connector's documentation](https://github.com/couchbase/kafka-connect-couchbase). ## Prerequisites[​](#connect_couchbase_sink_prereq "Direct link to Prerequisites") To setup a [Couchbase](https://www.couchbase.com/) sink connector, you need an Aiven for Apache Kafka service [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). Also collect the following information about the sink Couchbase database upfront: * `COUCHBASE_SEED_NODES`: The database seed nodes * `COUCHBASE_USER`: The database user to connect * `COUCHBASE_PASSWORD`: The database password for the `COUCHBASE_USER` * `COUCHBASE_BUCKET`: The bucket where to land the data * `APACHE_KAFKA_HOST`: The hostname of the Apache Kafka service, only needed when using Avro as data format * `SCHEMA_REGISTRY_PORT`: The Apache Kafka's schema registry port, only needed when using Avro as data format * `SCHEMA_REGISTRY_USER`: The Apache Kafka's schema registry username, only needed when using Avro as data format * `SCHEMA_REGISTRY_PASSWORD`: The Apache Kafka's schema registry user password, only needed when using Avro as data format note If you're using Aiven for Apache Kafka®, the Kafka related details are available in the [Aiven console](https://console.aiven.io/) service Overview tab or via the dedicated `avn service get` command with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). ## Setup a Couchbase sink connector with Aiven Console[​](#setup-a-couchbase-sink-connector-with-aiven-console "Direct link to Setup a Couchbase sink connector with Aiven Console") The following example demonstrates how to setup a Couchbase sink connector for Apache Kafka using the [Aiven Console](https://console.aiven.io/). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Define the connector configurations in a file (we'll refer to it with the name `couchbase-sink.json`) with the following content. Creating a file is not strictly necessary but allows to have all the information in one place before copy/pasting them in the [Aiven Console](https://console.aiven.io/): ``` { "name":"CONNECTOR_NAME", "connector.class": "com.couchbase.connect.kafka.CouchbaseSourceConnector", "couchbase.seed.nodes": "COUCHBASE_SEED_NODES", "couchbase.username": "COUCHBASE_USER", "couchbase.password": "COUCHBASE_PASSWORD", "couchbase.bucket": "COUCHBASE_BUCKET", "topics": "TOPIC_LIST" } ``` The configuration file contains the following entries: * `name`: the connector name, replace CONNECTOR\_NAME with the name you want to use for the connector * `COUCHBASE_SEED_NODES`, `COUCHBASE_BUCKET`, `COUCHBASE_USER`, `COUCHBASE_PASSWORD`: sink database parameters collected in the [prerequisite](/docs/products/kafka/kafka-connect/howto/couchbase-sink.md#connect_couchbase_sink_prereq) phase. ### Create a Kafka Connect connector with the Aiven Console[​](#create-a-kafka-connect-connector-with-the-aiven-console "Direct link to Create a Kafka Connect connector with the Aiven Console") To create an Apache Kafka Connect connector: 1. Log in to the [Aiven Console](https://console.aiven.io/) and select the Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service where the connector needs to be defined. 2. Select **Manage stream** > **Connectors** from the left sidebar. 3. Select **Create New Connector** note Is is enabled only for services [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md). 4. Select **Couchbase Sink**. 5. In the **Common** tab, locate the **Connector configuration** text box and select on **Edit**. 6. Paste the connector configuration (stored in the `couchbase-sink.json` file) in the form. 7. Select **Apply**. note The Aiven Console parses the configuration file and fills the relevant UI fields. You can review the UI fields across the various tabs and change them if necessary. The changes will be reflected in JSON format in the **Connector configuration** text box. 8. After all the settings are correctly configured, select **Create connector**. 9. Verify the connector status under **Manage stream** > **Connectors**. 10. Verify the presence of the data in the target Couchbase bucket. note You can also create connectors using the [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create). ## Example: define a Couchbase sink connector[​](#example-define-a-couchbase-sink-connector "Direct link to Example: define a Couchbase sink connector") The example creates an Couchbase sink connector with the following properties: * connector name: `couchbase_sink` * Couchbase seeds: `test.cloud.couchbase.com` * Couchbase username: `testuser` * Couchbase password: `Test123!` * Couchbase bucket: `travel-sample` * topic to sink: `inventory` The connector configuration is the following: ``` { "name": "couchbase_sink", "connector.class": "com.couchbase.connect.kafka.CouchbaseSinkConnector", "couchbase.seed.nodes": "test.cloud.couchbase.com", "couchbase.username": "testuser", "couchbase.password": "Test123!", "couchbase.bucket": "travel-sample", "topics": "inventory" } ``` With the above configuration stored in a `couchbase-sink.json` file, you can create the connector in the `demo-kafka` instance and you should see the data landing in an Couchbase bucket topic named `travel-sample`. *Couchbase is a trademark of Couchbase, Inc.* --- # Create a source connector from Couchbase to Apache Kafka® The [Couchbase](https://www.couchbase.com/) source connector pushes data from the NoSQL database, to Apache Kafka® where it can be transformed and read by multiple consumers. tip Sourcing data from a database into Apache Kafka decouples the database from the set of consumers. Once the data is in Apache Kafka, multiple applications can access it without adding any additional query overhead to the source database. note See the full set of available parameters and configuration options in the [connector's documentation](https://github.com/couchbase/kafka-connect-couchbase). ## Prerequisites[​](#connect_couchbase_source_prereq "Direct link to Prerequisites") To setup a [Couchbase](https://www.couchbase.com/) source connector, you need an Aiven for Apache Kafka service [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). Also collect the following information about the source Couchbase database upfront: * `COUCHBASE_SEED_NODES`: The database seed nodes * `COUCHBASE_USER`: The database user to connect * `COUCHBASE_PASSWORD`: The database password for the `COUCHBASE_USER` * `COUCHBASE_BUCKET`: The bucket from which extract the data * `COUCHBASE_SCOPE`: The scope from which extract the data (if none all the scopes within the bucket will be sourced) * `COUCHBASE_COLLECTIONS`: The list of collections from which extract the data (if none all the collections within the bucket and scope will be sourced) * `APACHE_KAFKA_HOST`: The hostname of the Apache Kafka service, only needed when using Avro as data format * `SCHEMA_REGISTRY_PORT`: The Apache Kafka's schema registry port, only needed when using Avro as data format * `SCHEMA_REGISTRY_USER`: The Apache Kafka's schema registry username, only needed when using Avro as data format * `SCHEMA_REGISTRY_PASSWORD`: The Apache Kafka's schema registry user password, only needed when using Avro as data format note If you're using Aiven for Apache Kafka®, the Kafka related details are available in the [Aiven console](https://console.aiven.io/) service Overview tab or via the dedicated `avn service get` command with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). ## Setup a Couchbase source connector with Aiven Console[​](#setup-a-couchbase-source-connector-with-aiven-console "Direct link to Setup a Couchbase source connector with Aiven Console") The following example demonstrates how to setup a Couchbase source connector for Apache Kafka using the [Aiven Console](https://console.aiven.io/). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Define the connector configurations in a file (we'll refer to it with the name `couchbase-source.json`) with the following content. Creating a file is not strictly necessary but allows to have all the information in one place before copy/pasting them in the [Aiven Console](https://console.aiven.io/): ``` { "name":"CONNECTOR_NAME", "connector.class": "com.couchbase.connect.kafka.CouchbaseSourceConnector", "couchbase.seed.nodes": "COUCHBASE_SEED_NODES", "couchbase.username": "COUCHBASE_USER", "couchbase.password": "COUCHBASE_PASSWORD", "couchbase.bucket": "COUCHBASE_BUCKET", "couchbase.scope": "COUCHBASE_SCOPE", "couchbase.collections": "COUCHBASE_COLLECTIONS", "couchbase.source.handler": "com.couchbase.connect.kafka.handler.source.RawJsonSourceHandler", "value.converter": "org.apache.kafka.connect.converters.ByteArrayConverter", } ``` The configuration file contains the following entries: * `name`: the connector name, replace CONNECTOR\_NAME with the name you want to use for the connector * `COUCHBASE_SEED_NODES`, `COUCHBASE_BUCKET`, `COUCHBASE_SCOPE`, `COUCHBASE_COLLECTIONS`, `COUCHBASE_USER`, `COUCHBASE_PASSWORD`: source database parameters collected in the [prerequisite](/docs/products/kafka/kafka-connect/howto/couchbase-source.md#connect_couchbase_source_prereq) phase. * `couchbase.source.handler` and `value.converter`: defines the messages data format in the Apache Kafka topic. The combination of `com.couchbase.connect.kafka.handler.source.RawJsonSourceHandler` and `org.apache.kafka.connect.converters.ByteArrayConverter` pushes the Couchbase documents in the Kafka topic in JSON format. ### Create a Kafka Connect connector with the Aiven Console[​](#create-a-kafka-connect-connector-with-the-aiven-console "Direct link to Create a Kafka Connect connector with the Aiven Console") To create a Kafka Connect connector: 1. Log in to the [Aiven Console](https://console.aiven.io/) and select the Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service where the connector needs to be defined. 2. Select **Manage stream** > **Connectors** from the left sidebar. 3. Select **Create New Connector**, it is enabled only for services [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md). 4. Select **Couchbase Source**. 5. In the **Common** tab, locate the **Connector configuration** text box and select on **Edit**. 6. Paste the connector configuration (stored in the `couchbase-source.json` file) in the form. 7. Select **Apply**. note The Aiven Console parses the configuration file and fills the relevant UI fields. You can review the UI fields across the various tabs and change them if necessary. The changes will be reflected in JSON format in the **Connector configuration** text box. 8. After all the settings are correctly configured, select **Create connector**. tip If you're using Aiven for Apache Kafka, topics will not be created automatically. Either create them manually following the `database.server.name.schema_name.table_name` naming pattern or enable the `kafka.auto_create_topics_enable` advanced parameter. 9. Verify the connector status under **Manage stream** > **Connectors** 10. Verify the presence of the data in the target Apache Kafka topic coming from the MongoDB dataset. The topic name is equal to the concatenation of the database and collection name. To change the target table name, you can do so using the Kafka Connect `RegexRouter` transformation. note You can also create connectors using the [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create). ## Example: define a Couchbase source connector[​](#example-define-a-couchbase-source-connector "Direct link to Example: define a Couchbase source connector") The example creates an Couchbase source connector with the following properties: * connector name: `couchbase_source` * Couchbase seeds: `test.cloud.couchbase.com` * Couchbase username: `testuser` * Couchbase password: `Test123!` * Couchbase bucket: `travel-sample` * Couchbase scope: `inventory` * Couchbase collections: `airline` The connector configuration is the following: ``` { "name": "couchbase_source", "connector.class": "com.couchbase.connect.kafka.CouchbaseSourceConnector", "couchbase.seed.nodes": "test.cloud.couchbase.com", "couchbase.username": "testuser", "couchbase.password": "Test123!", "couchbase.bucket": "travel-sample", "couchbase.scope": "inventory", "couchbase.collections": "airline", "couchbase.source.handler": "com.couchbase.connect.kafka.handler.source.RawJsonSourceHandler", "value.converter": "org.apache.kafka.connect.converters.ByteArrayConverter" } ``` With the above configuration stored in a `couchbase-source.json` file, you can create the connector in the `demo-kafka` instance and you should see the data landing in an Apache Kafka topic named `${bucket}.${scope}.${collection}` by default, you can change the landing topic logic by modifying the `couchbase.topic` parameter definition. *Couchbase is a trademark of Couchbase, Inc.* --- # Create a Debezium source connector from MongoDB to Apache Kafka® Track and write MongoDB database changes to an Apache Kafka® topic in a standard format with the Debezium source connector, enabling transformation and access by multiple consumers using a MongoDB replica set or sharded cluster. note Aiven supports multiple Debezium versions through multi-version support, including versions 1.9.7, 2.5.0, 2.7.4, and 3.1.0. Debezium 2.5 introduced changes to connector configuration and behavior. To prevent unintentional upgrades during maintenance updates, pin the connector version using the `plugin_versions` configuration property. For details, see [Manage connector versions](/docs/products/kafka/kafka-connect/howto/manage-connector-versions.md). If you use Debezium for PostgreSQL version 1.9.7 with the `wal2json` replication format, do not upgrade to version 2.0 or later until you migrate to a supported format such as `pgoutput`. To upgrade from version 1.9.7, use multi-version support to test your configuration before applying changes in production. For further assistance, contact [Aiven support](mailto:support@aiven.io). ## Prerequisites[​](#connect_debezium_mongodb_source_prereq "Direct link to Prerequisites") To configure a Debezium source connector for MongoDB, you need either an Aiven for Apache Kafka service with [Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). note You can view the full set of available parameters and configuration options in the [connector's documentation](https://debezium.io/docs/connectors/mongodb/). Before you begin, gather the necessary information about your source MongoDB database: * `MONGODB_HOST`: The database hostname * `MONGODB_PORT`: The database port * `MONGODB_USER`: The database user to connect * `MONGODB_PASSWORD`: The database password for the `MONGODB_USER` * `MONGODB_DATABASE_NAME`: The database name to include in the replica * `MONGODB_REPLICA_SET_NAME`: The name of MongoDB's replica set * `APACHE_KAFKA_HOST`: The hostname of the Apache Kafka service, only needed when using Avro as data format * `SCHEMA_REGISTRY_PORT`: The Apache Kafka's schema registry port, only needed when using Avro as data format * `SCHEMA_REGISTRY_USER`: The Apache Kafka's schema registry username, only needed when using Avro as data format * `SCHEMA_REGISTRY_PASSWORD`: The Apache Kafka's schema registry user password, only needed when using Avro as data format note With Aiven for Apache Kafka, you can gather the necessary Apache Kafka details from the service's Overview page on the [Aiven console](https://console.aiven.io/) or by using the `avn service get` command with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). ## Setup a MongoDB Debezium source connector with Aiven Console[​](#setup-a-mongodb-debezium-source-connector-with-aiven-console "Direct link to Setup a MongoDB Debezium source connector with Aiven Console") The following example demonstrates how to setup a Debezium source connector for Apache Kafka to a MongoDB database using the [Aiven Console](https://console.aiven.io/). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Create a configuration file named `debezium_source_mongodb.json` with the following connector configurations. While optional, creating this file helps you organize your settings in one place and copy/paste them into the [Aiven Console](https://console.aiven.io/) later. ``` { "name":"CONNECTOR_NAME", "connector.class": "io.debezium.connector.mongodb.MongoDbConnector", "mongodb.hosts": "MONGODB_REPLICA_SET_NAME/MONGODB_HOST:MONGODB_PORT", "mongodb.name" : "MONGODB_DATABASE_NAME", "mongodb.user": "MONGODB_USER", "mongodb.password": "MONGODB_PASSWORD", "tasks.max":"NR_TASKS", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } ``` The configuration file contains the following entries: * `name`: The connector name, replace CONNECTOR\_NAME with the name you want to use for the connector * `MONGODB_HOST`, `MONGODB_PORT`, `MONGODB_DATABASE_NAME`, `MONGODB_USER`, `MONGODB_PASSWORD` and `MONGODB_REPLICA_SET_NAME`: Source database parameters collected in the [prerequisite](/docs/products/kafka/kafka-connect/howto/debezium-source-connector-mongodb.md#connect_debezium_mongodb_source_prereq) phase. * `tasks.max`: Maximum number of tasks to execute in parallel. By default this is 1, the connector can use at most 1 task for each collection defined. Replace `NR_TASKS` with the amount of parallel task based on the number of input collections. * `key.converter` and `value.converter`: Defines the messages data format in the Apache Kafka topic. The `io.confluent.connect.avro.AvroConverter` converter pushes messages in Avro format. To store the message schemas, Aiven's [Karapace schema registry](https://github.com/Aiven-Open/karapace) is used, specified by the `schema.registry.url` parameter and related credentials. note The `key.converter` and `value.converter` sections are only needed when pushing data in Avro format. Otherwise, messages default to JSON format. The `USER_INFO` is not a placeholder and does not require any parameter substitution. ### Create a Kafka Connect connector with the Aiven Console[​](#create-a-kafka-connect-connector-with-the-aiven-console "Direct link to Create a Kafka Connect connector with the Aiven Console") To create a Kafka Connect connector: 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. Select the Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service to define the connector. 3. Select **Manage stream** > **Connectors** from the sidebar. 4. Select **Create New Connector**, which is available only for services [that have Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md). 5. Select **Debezium - MongoDB**. 6. In the **Common** tab, locate the **Connector configuration** text box and select on **Edit**. 7. Paste the connector configuration (stored in the `debezium_source_mongodb.json` file) in the form. 8. Select **Apply**. note The Aiven Console reads through the configuration file and automatically populates the relevant UI fields. You can view and modify these fields across different tabs. Any change you make is reflected in JSON format within the **Connector configuration** text box. 9. After all the settings are correctly configured, select **Create new connector**. tip With Aiven for Apache Kafka, topics are not created automatically. You have two options: * Manually create topics using the naming pattern: `database.server.name.schema_name.table_name`. * Enable the `Kafka topic auto-creation` feature. See [Enable automatic topic creation with Aiven CLI](/docs/products/kafka/howto/create-topics-automatically.md). 10. Verify the connector status under **Manage stream** > **Connectors**. 11. Verify the presence of the data in the target Apache Kafka topic coming from the MySQL dataset. The topic name is equal to concatenation of the database and table name. To change the target table name, you can use Apache Kafka Connect `RegexRouter` transformation. You can also create connectors using the [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create). --- # Create a Debezium source connector from MySQL to Apache Kafka® The MySQL Debezium source connector extracts the changes committed to the database binary log (binlog), and writes them to an Apache Kafka® topic in a standard format where they can be transformed and read by multiple consumers. note Aiven supports multiple Debezium versions through multi-version support, including versions 1.9.7, 2.5.0, 2.7.4, and 3.1.0. Debezium 2.5 introduced changes to connector configuration and behavior. To prevent unintentional upgrades during maintenance updates, pin the connector version using the `plugin_versions` configuration property. For details, see [Manage connector versions](/docs/products/kafka/kafka-connect/howto/manage-connector-versions.md). If you use Debezium for PostgreSQL version 1.9.7 with the `wal2json` replication format, do not upgrade to version 2.0 or later until you migrate to a supported format such as `pgoutput`. To upgrade from version 1.9.7, use multi-version support to test your configuration before applying changes in production. For further assistance, contact [Aiven support](mailto:support@aiven.io). ## Schema versioning[​](#connect_debezium_mysql_schema_versioning "Direct link to Schema versioning") Database table schemas can evolve over time by adding, modifying, or removing columns. The MySQL Debezium source connector keeps track of schema changes by storing them in a separate "history" topic that you can set up with dedicated `history.*` configuration parameters. warning The MySQL Debezium source connector `history.*` parameters are not visible in the list of options available in the [Aiven Console](https://console.aiven.io/). However, you can insert or modify them by editing the JSON configuration in the **Connector configuration** section. ## Prerequisites[​](#connect_debezium_mysql_source_prereq "Direct link to Prerequisites") To configure a Debezium source connector for MySQL, you need either an Aiven for Apache Kafka service with [Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). Before you begin, gather the necessary information about your source MySQL database: note You can view the full set of available parameters and configuration options in the [connector's documentation](https://debezium.io/docs/connectors/mysql/). * `MYSQL_HOST`: The database hostname * `MYSQL_PORT`: The database port * `MYSQL_USER`: The database user to connect * `MYSQL_PASSWORD`: The database password for the `MYSQL_USER` * `MYSQL_DATABASE_NAME`: The database name * `SSL_MODE`: The [SSL mode](https://dev.mysql.com/doc/refman/5.7/en/connection-options.html) * `MYSQL_TABLES`: The list of database tables to be included in Apache Kafka. Format the list as `schema_name1.table_name1,schema_name2.table_name2`. * `APACHE_KAFKA_HOST`: The hostname of the Apache Kafka service, needed when storing the [schema definition changes](/docs/products/kafka/kafka-connect/howto/debezium-source-connector-mysql.md#connect_debezium_mysql_schema_versioning) * `APACHE_KAFKA_PORT`: The port of the Apache Kafka service, needed when storing the [schema definition changes](/docs/products/kafka/kafka-connect/howto/debezium-source-connector-mysql.md#connect_debezium_mysql_schema_versioning) * `SCHEMA_REGISTRY_PORT`: The Apache Kafka's schema registry port is only needed when using Avro as a data format * `SCHEMA_REGISTRY_USER`: The Apache Kafka's schema registry username is only needed when using Avro as a data format * `SCHEMA_REGISTRY_PASSWORD`: The Apache Kafka's schema registry user password is only needed when using Avro as a data format * `TOPIC_PREFIX`: The namespace for the MySQL database server from which Debezium captures changes. It should be unique across all connectors. note When using Aiven for MySQL and Aiven for Apache Kafka, you can gather the necessary details from the service's Overview page on [Aiven console](https://console.aiven.io/) or by using the `avn service get` command with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). ## Setup a MySQL Debezium source connector with Aiven Console[​](#setup-a-mysql-debezium-source-connector-with-aiven-console "Direct link to Setup a MySQL Debezium source connector with Aiven Console") The following example demonstrates how to set up a Debezium source Connector for Apache Kafka to a MySQL database using the [Aiven CLI dedicated command](/docs/tools/cli/service/connector.md). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Create a configuration file named `debezium_source_mysql.json` with the following connector configurations. While optional, creating this file helps you organize your settings in one place and copy/paste them into the [Aiven Console](https://console.aiven.io/) later. * 2.5 config * 1.9 config ``` { "name": "CONNECTOR_NAME", "connector.class": "io.debezium.connector.mysql.MySqlConnector", "database.hostname": "MYSQL_HOST", "database.port": "MYSQL_PORT", "database.user": "MYSQL_USER", "database.password": "MYSQL_PASSWORD", "database.dbname": "MYSQL_DATABASE_NAME", "database.sslmode": "SSL_MODE", "database.server.id": "MYSQL_SERVER_ID", "topic.prefix": "TOPIC_PREFIX", "table.include.list": "MYSQL_TABLES", "tasks.max": "NR_TASKS", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "schema.history.internal.kafka.topic": "HISTORY_TOPIC_NAME", "schema.history.internal.kafka.bootstrap.servers": "APACHE_KAFKA_HOST:APACHE_KAFKA_PORT", "schema.history.internal.producer.security.protocol": "SSL", "schema.history.internal.producer.ssl.keystore.type": "PKCS12", "schema.history.internal.producer.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "schema.history.internal.producer.ssl.keystore.password": "password", "schema.history.internal.producer.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "schema.history.internal.producer.ssl.truststore.password": "password", "schema.history.internal.producer.ssl.key.password": "password", "schema.history.internal.consumer.security.protocol": "SSL", "schema.history.internal.consumer.ssl.keystore.type": "PKCS12", "schema.history.internal.consumer.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "schema.history.internal.consumer.ssl.keystore.password": "password", "schema.history.internal.consumer.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "schema.history.internal.consumer.ssl.truststore.password": "password", "schema.history.internal.consumer.ssl.key.password": "password", "include.schema.changes" : "true" } ``` ``` { "name": "CONNECTOR_NAME", "connector.class": "io.debezium.connector.mysql.MySqlConnector", "database.hostname": "MYSQL_HOST", "database.port": "MYSQL_PORT", "database.user": "MYSQL_USER", "database.password": "MYSQL_PASSWORD", "database.dbname": "MYSQL_DATABASE_NAME", "database.sslmode": "SSL_MODE", "database.server.name": "TOPIC_PREFIX", "table.include.list": "MYSQL_TABLES", "tasks.max": "NR_TASKS", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "database.history.consumer.security.protocol": "SSL", "database.history.consumer.ssl.key.password": "password", "database.history.consumer.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "database.history.consumer.ssl.keystore.password": "password", "database.history.consumer.ssl.keystore.type": "PKCS12", "database.history.consumer.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "database.history.consumer.ssl.truststore.password": "password", "database.history.kafka.bootstrap.servers": "APACHE_KAFKA_HOST:APACHE_KAFKA_PORT", "database.history.kafka.topic": "HISTORY_TOPIC_NAME", "database.history.producer.security.protocol": "SSL", "database.history.producer.ssl.key.password": "password", "database.history.producer.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "database.history.producer.ssl.keystore.password": "password", "database.history.producer.ssl.keystore.type": "PKCS12", "database.history.producer.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "database.history.producer.ssl.truststore.password": "password", "include.schema.changes": "true" } ``` The configuration file contains the following entries: * `name`: The connector name, replace CONNECTOR\_NAME with the name you want to use for the connector. * `MYSQL_HOST`, `MYSQL_PORT`, `MYSQL_DATABASE_NAME`, `SSL_MODE`, `MYSQL_USER`, `MYSQL_PASSWORD`, `MYSQL_TABLES`: Source database parameters collected in the [prerequisite](/docs/products/kafka/kafka-connect/howto/debezium-source-connector-mysql.md#connect_debezium_mysql_source_prereq) phase. * `database.server.name` (Debezium v1): The logical name of the database, which also determines the prefix used in combination with table names for Apache Kafka topic names. * `database.server.id` (Debezium v2): The numeric ID of the MySQL client. It must be unique across all currently-running database processes in the MySQL cluster. * `topic.prefix` (Debezium v2): The prefix used in combination with table names for Apache Kafka topic names. * `tasks.max`: Maximum number of tasks to execute in parallel. By default this is 1, the connector can use at most 1 task for each source table defined. Replace `NR_TASKS` with the amount of parallel task based on the number of tables. * `database.history.kafka.topic`: The name of the Apache Kafka topic that contains the history of schema changes. * `database.history.kafka.bootstrap.servers`: Directs to the Aiven for Apache Kafka service that runs the connector and stores [schema definition changes](/docs/products/kafka/kafka-connect/howto/debezium-source-connector-mysql.md#connect_debezium_mysql_schema_versioning). * `database.history.producer` and `database.history.consumer`: Refers to truststores and keystores pre-created on the Aiven for Apache Kafka node to handle SSL authentication warning The values defined for each `database.history.producer` and `database.history.consumer` parameters are already set to work with the predefined truststore and keystore created in the Aiven for Apache Kafka nodes. Modifying these values is not recommended. * `key.converter` and `value.converter`: Defines the messages data format in the Apache Kafka topic. The `io.confluent.connect.avro.AvroConverter` converter pushes messages in Avro format. To store the message schemas, Aiven's [Karapace schema registry](https://github.com/Aiven-Open/karapace) is used, specified by the `schema.registry.url` parameter and related credentials. note The `key.converter` and `value.converter` sections are only needed when pushing data in Avro format. Otherwise, messages default to JSON format. The `USER_INFO` is not a placeholder and does not require any parameter substitution. ### Create a Kafka Connect connector with the Aiven Console[​](#create-a-kafka-connect-connector-with-the-aiven-console "Direct link to Create a Kafka Connect connector with the Aiven Console") To create a Kafka Connect connector: 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. Select the Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service to define the connector. 3. Select **Manage stream** > **Connectors** from the sidebar. 4. Select **Create New Connector**, which is available only for services [that have Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md). 5. Select the **Debezium - MySQL**. 6. In the **Common** tab, locate the **Connector configuration** text box and select on **Edit**. 7. Paste the connector configuration (stored in the `debezium_source_mysql.json` file) in the form. 8. Select **Apply**. note The Aiven Console reads through the configuration file and automatically populates the relevant UI fields. You can view and modify these fields across different tabs. Any change you make is reflected in JSON format within the **Connector configuration** text box. 9. After all the settings are correctly configured, select **Create connector** tip With Aiven for Apache Kafka, topics are not created automatically. You have two options: * Manually create topics using the naming pattern: `database.server.name.schema_name.table_name`. * Enable the `Kafka topic auto-creation` feature. See [Enable automatic topic creation with Aiven CLI](/docs/products/kafka/howto/create-topics-automatically.md). 10. Verify the connector status under **Manage stream** > **Connectors**. 11. Verify the presence of the data in the target Apache Kafka topic coming from the MySQL dataset. The topic name is equal to concatenation of the database and table name. To change the target table name, you can use Apache Kafka Connect `RegexRouter` transformation. note You can also create connectors using the [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create). --- # Create an Oracle Debezium source connector for Aiven for Apache Kafka® [Early availability](/docs/platform/concepts/service-and-feature-releases.md) The Oracle Debezium source connector streams change data from an Oracle database to Apache Kafka® topics. The connector uses Oracle LogMiner to read redo logs and capture row-level inserts, updates, and deletes from selected tables. Each change event is written to an Apache Kafka topic based on the configured topic prefix and table selection. Use this connector to capture near real-time changes from Oracle transactional tables into Apache Kafka. The connector does not support bulk ingestion or loading existing data. important This connector does not perform large initial snapshots. It captures only changes that occur after the connector starts. ## Enable CDC in Oracle[​](#enable-cdc-in-oracle "Direct link to Enable CDC in Oracle") To use the Oracle Debezium source connector, the source database must be configured to support LogMiner-based change data capture (CDC). ### Required Oracle configuration[​](#required-oracle-configuration "Direct link to Required Oracle configuration") Ensure that the following settings are enabled on the Oracle database: * Supplemental logging * ARCHIVELOG mode * Redo log retention long enough for the connector to process changes note On Amazon RDS for Oracle, enable supplemental logging using the RDS administrative procedures instead of standard `ALTER DATABASE` commands. For detailed instructions, see: * [Creating and connecting to an Oracle DB instance on Amazon RDS](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/CHAP_GettingStarted.CreatingConnecting.Oracle.html) * [Enabling supplemental logging on Oracle RDS](https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/Appendix.Oracle.CommonDBATasks.Log.html#Appendix.Oracle.CommonDBATasks.AddingSupplementalLogging) * [Debezium Oracle connector documentation](https://debezium.io/documentation/reference/stable/connectors/oracle.html) ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An [Aiven for Apache Kafka® service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with Apache Kafka Connect enabled, or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * The Kafka Connect service can connect to the Oracle database host and port. * * A [secret provider configured for Kafka Connect](/docs/products/kafka/kafka-connect/howto/configure-secret-providers.md), such as [AWS Secrets Manager](/docs/products/kafka/kafka-connect/howto/configure-aws-secrets-manager.md). The configuration example below uses a secret provider as a best practice for managing database credentials. * [Schema Registry enabled](/docs/products/kafka/karapace/howto/enable-karapace.md) when using Avro serialization. * An Oracle Database instance with: * Oracle LogMiner enabled * Supplemental logging enabled * ARCHIVELOG mode enabled * A database user with sufficient privileges for LogMiner-based change data capture (CDC) note The Oracle user must have permissions to access redo logs and LogMiner views. For the required permissions, see the [Debezium Oracle connector documentation](https://debezium.io/documentation/reference/stable/connectors/oracle.html). ### Collect Oracle and connector requirements for the connector[​](#collect-oracle-and-connector-requirements-for-the-connector "Direct link to Collect Oracle and connector requirements for the connector") Before creating the connector, gather the following details: * `ORACLE_HOST`: Oracle database hostname * `ORACLE_PORT`: Oracle database port (default: 1521) * `ORACLE_USER`: Database user * `ORACLE_PASSWORD`: Database password * `ORACLE_DATABASE_NAME`: Database name (SID or service name) * `ORACLE_TABLES`: Tables to capture, formatted as `SCHEMA.TABLE` * `TOPIC_PREFIX`: Prefix for Apache Kafka topic names * Schema Registry endpoint and credentials, if using Avro ## Create an Oracle Debezium source connector configuration file[​](#create-an-oracle-debezium-source-connector-configuration-file "Direct link to Create an Oracle Debezium source connector configuration file") Create a file named `oracle_debezium_source_connector.json` with the following configuration. This example streams only changes that occur after the connector starts. ``` { "name": "oracle-debezium-source", "connector.class": "io.debezium.connector.oracle.OracleConnector", "tasks.max": 1, "database.hostname": "${aws:oracle/secrets:database.hostname}", "database.port": "1521", "database.user": "admin", "database.password": "${aws:oracle/secrets:database.password}", "database.dbname": "ORCL", "topic.prefix": "oracle.cdc", "table.include.list": "ADMIN.PERSON", "snapshot.mode": "no_data", "include.schema.changes": "true", "key.converter": "io.confluent.connect.avro.AvroConverter", "value.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "", "value.converter.schema.registry.url": "", "schema.name.adjustment.mode": "avro", "schema.history.internal.kafka.topic": "oracle.cdc.history", "schema.history.internal.kafka.bootstrap.servers": ":", "schema.history.internal.producer.security.protocol": "SSL", "schema.history.internal.producer.ssl.keystore.type": "PKCS12", "schema.history.internal.producer.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "schema.history.internal.producer.ssl.keystore.password": "password", "schema.history.internal.producer.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "schema.history.internal.producer.ssl.truststore.password": "password", "schema.history.internal.producer.ssl.key.password": "password", "schema.history.internal.consumer.security.protocol": "SSL", "schema.history.internal.consumer.ssl.keystore.type": "PKCS12", "schema.history.internal.consumer.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "schema.history.internal.consumer.ssl.keystore.password": "password", "schema.history.internal.consumer.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "schema.history.internal.consumer.ssl.truststore.password": "password", "schema.history.internal.consumer.ssl.key.password": "password" } ``` Parameters: * `name`: A unique name for the connector. * `connector.class`: The Java class for the connector. Set to `io.debezium.connector.oracle.OracleConnector`. * `tasks.max`: The maximum number of tasks that run in parallel. For Oracle CDC, set this value to `1`. * `database.hostname`: Hostname or IP address of the Oracle database. * `database.port`: Oracle database port. Default is `1521`. * `database.user`: Database user used by the connector. * `database.password`: Password for the database user. * `database.dbname`: Oracle database name, specified as a SID or service name. * `topic.prefix`: Prefix for Apache Kafka topic names. Topics are created using the pattern `..`. * `table.include.list`: Comma-separated list of tables to capture, formatted as `SCHEMA.TABLE`. * `snapshot.mode`: Controls initial snapshot behavior. Set to `no_data` to stream only changes that occur after the connector starts. * `include.schema.changes`: Specifies whether schema change events are published to Apache Kafka. * `key.converter`: Converter used to serialize record keys. * `value.converter`: Converter used to serialize record values. * `key.converter.schema.registry.url`: Schema Registry endpoint used for key schemas when using Avro. * `value.converter.schema.registry.url`: Schema Registry endpoint used for value schemas when using Avro. * `schema.name.adjustment.mode`: Adjusts schema names for Avro compatibility. Set to `avro` when using Avro serialization. * `schema.history.internal.kafka.topic`: Kafka topic Debezium uses to store schema history, including DDL and table structure changes. * `schema.history.internal.kafka.bootstrap.servers`: Kafka bootstrap servers Debezium uses to connect to Kafka for schema history storage. * `schema.history.internal.producer.*`: SSL settings Debezium uses to write schema history to Kafka. * `schema.history.internal.consumer.*`: SSL settings Debezium uses to read schema history from Kafka. note On Aiven, Debezium requires SSL to access Kafka for schema history. Use the keystore and truststore provided by the Kafka Connect service at `/run/aiven/keys/`. Keep the keystore and truststore passwords set to `password`. ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Console * CLI 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector**. 5. In the source connectors list, select **Debezium - Oracle**, and click **Get started**. 6. On the **Oracle Debezium Source Connector** page, open the **Common** tab. 7. Locate the **Connector configuration** field and click **Edit**. 8. Paste the contents of `oracle_debezium_source_connector.json`. 9. Click **Create connector**. 10. Verify the connector status on the **Manage stream** > **Connectors** page. To create the connector using the [Aiven CLI](/docs/tools/cli/service/connector.md#avn_service_connector_create), run: ``` avn service connector create SERVICE_NAME @oracle_debezium_source_connector.json ``` To check the connector status: ``` avn service connector status SERVICE_NAME oracle-debezium-source ``` ## Verify data capture[​](#verify-data-capture "Direct link to Verify data capture") After creating the connector: 1. Confirm that the connector state is `RUNNING`. 2. Verify that Apache Kafka topics exist for the selected tables, using the configured topic prefix. 3. Make an insert, update, or delete in the source Oracle table. 4. Confirm that a corresponding record appears in the Apache Kafka topic. Once running, the connector continuously streams row-level changes from the selected Oracle tables to Apache Kafka topics until it is stopped or reconfigured. ## Understand Kafka topic names[​](#understand-kafka-topic-names "Direct link to Understand Kafka topic names") The connector writes change events to Apache Kafka topics using the following pattern: ``` ..
``` Use this pattern to identify the topic that contains change events for a specific Oracle table. For example, if `topic.prefix` is set to `oracle.cdc`, changes from the `ADMIN.PERSON` table are written to: ``` oracle.cdc.ADMIN.PERSON ``` Consume records from this topic to read change events for the `ADMIN.PERSON` table. ## Limitations[​](#limitations "Direct link to Limitations") * Large initial snapshots are not supported. * The connector streams ongoing changes only. * Streaming performance depends on redo log volume and retention. * High change rates can increase load on the source database. * Exactly-once delivery is not supported. * This connector is available as an early availability feature. ## Best practices[​](#best-practices "Direct link to Best practices") * Use the Oracle Debezium source connector for transactional tables that require change data capture. * Avoid using this connector for append-only or bulk ingestion workloads. For those cases, use a JDBC source connector instead. * Limit the number of tables per connector to reduce load on the source database. * Start with streaming ongoing changes only before expanding to additional tables. * Monitor the source database for increased load during initial deployment, especially in early access environments. ## Example: Oracle Debezium source connector[​](#example-oracle-debezium-source-connector "Direct link to Example: Oracle Debezium source connector") The following example shows an Oracle Debezium source connector that streams ongoing changes from two tables using Avro serialization. ``` { "name": "oracle-debezium-example", "connector.class": "io.debezium.connector.oracle.OracleConnector", "tasks.max": 1, "database.hostname": "${aws:oracle/secrets:database.hostname}", "database.port": "1521", "database.user": "admin", "database.password": "${aws:oracle/secrets:database.password}", "database.dbname": "ORCL", "topic.prefix": "oracle.cdc", "table.include.list": "SALES.ORDERS,SALES.CUSTOMERS", "snapshot.mode": "no_data", "include.schema.changes": "true", "key.converter": "io.confluent.connect.avro.AvroConverter", "value.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "", "value.converter.schema.registry.url": "", "schema.name.adjustment.mode": "avro" } ``` This connector creates Apache Kafka topics using the pattern `..
`, for example: ``` oracle.cdc.SALES.ORDERS oracle.cdc.SALES.CUSTOMERS ``` Related pages * [Debezium Oracle connector documentation](https://debezium.io/documentation/reference/stable/connectors/oracle.html) * [Enable Apache Kafka Connect on Aiven](/docs/products/kafka/kafka-connect/howto/enable-connect.md) --- # Create a Debezium source connector from PostgreSQL® to Apache Kafka® The Debezium source connector extracts the changes committed to the transaction log in a relational database, such as PostgreSQL®, and writes them to an Apache Kafka® topic in a standard format where they can be transformed and read by multiple consumers. note Aiven supports multiple Debezium versions through multi-version support, including versions 1.9.7, 2.5.0, 2.7.4, and 3.1.0. Debezium 2.5 introduced changes to connector configuration and behavior. To prevent unintentional upgrades during maintenance updates, pin the connector version using the `plugin_versions` configuration property. For details, see [Manage connector versions](/docs/products/kafka/kafka-connect/howto/manage-connector-versions.md). If you use Debezium for PostgreSQL version 1.9.7 with the `wal2json` replication format, do not upgrade to version 2.0 or later until you migrate to a supported format such as `pgoutput`. To upgrade from version 1.9.7, use multi-version support to test your configuration before applying changes in production. For further assistance, contact [Aiven support](mailto:support@aiven.io). ## Prerequisites[​](#connect_debezium_pg_source_prereq "Direct link to Prerequisites") To configure a Debezium source connector for PostgreSQL, you need either an Aiven for Apache Kafka service with [Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). Before you begin, gather the necessary information about your source PostgreSQL database: * `PG_HOST`: The database hostname * `PG_PORT`: The database port * `PG_USER`: The database user to connect * `PG_PASSWORD`: The database password for the `PG_USER` * `PG_DATABASE_NAME`: The database name * `SSL_MODE`: The [SSL mode](https://www.postgresql.org/docs/current/libpq-ssl.html) * `PLUGIN_NAME`: The [logical decoding plugin](https://debezium.io/documentation/reference/stable/connectors/postgresql.html), possible values are `decoderbufs` and `pgoutput`. note Starting with Debezium version 2.5, the `wal2json` plugin is deprecated. * `PG_TABLES`: The list of database tables to be included in Apache Kafka; the list must be in the form of `schema_name1.table_name1,schema_name2.table_name2` * `PG_PUBLICATION_NAME`: The name of the [PostgreSQL logical replication publication](https://www.postgresql.org/docs/current/logical-replication-publication.html), if left empty, `debezium` is used as default * `PG_SLOT_NAME`: name of the [PostgreSQL replication slot](/docs/products/postgresql/howto/setup-logical-replication.md), if left empty, `debezium` is be used as default * `APACHE_KAFKA_HOST`: The hostname of the Apache Kafka service, only needed when using Avro as data format * `SCHEMA_REGISTRY_PORT`: The Apache Kafka's schema registry port, only needed when using Avro as data format * `SCHEMA_REGISTRY_USER`: The Apache Kafka's schema registry username, only needed when using Avro as data format * `SCHEMA_REGISTRY_PASSWORD`: The Apache Kafka's schema registry user password, only needed when using Avro as data format If you are using Aiven for PostgreSQL and Aiven for Apache Kafka the above details are available in the [Aiven console](https://console.aiven.io/) service Overview page or via the dedicated `avn service get` command with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). For a complete list of all available parameters and configuration options, see Debezium [connector's documentation](https://debezium.io/documentation/reference/stable/connectors/postgresql.html). important PostgreSQL® node replacements and major version upgrades can interrupt Debezium change data capture (CDC). For recovery steps and recommended practices, see [Handle PostgreSQL node replacements with Debezium](/docs/products/kafka/kafka-connect/howto/debezium-source-connector-pg-node-replacement.md). warning Debezium updates the LSN positions of the PostgreSQL replication slot only when changes occur in the connected database. If any replication slots have not acknowledged receiving old WAL segments, PostgreSQL cannot delete them. This means that if your system is idle (even though Aiven for PostgreSQL still generates 16 MiB of WAL every 5 minutes) or if changes are only happening in databases not connected to Debezium, PostgreSQL is unable to clean up the WAL, eventually leading to the service running out of disk space. Therefore, ensure you regularly update any database connected to Debezium. ## Set up a PostgreSQL Debezium source connector using Aiven CLI[​](#set-up-a-postgresql-debezium-source-connector-using-aiven-cli "Direct link to Set up a PostgreSQL Debezium source connector using Aiven CLI") The following example demonstrates how to set up a Debezium source Connector for Apache Kafka to a PostgreSQL database using the [Aiven CLI dedicated command](/docs/tools/cli/service/connector.md). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Create a configuration file named `debezium_source_pg.json` with the following connector configurations: * Debezium 2.5 config * Debezium 1.9 config ``` { "connector.class": "io.debezium.connector.postgresql.PostgresConnector", "database.dbname": "PG_DATABASE_NAME", "database.hostname": "PG_HOST", "database.password": "PG_PASSWORD", "database.port": "PG_PORT", "database.names": "testing", "database.server.name": "KAFKA_TOPIC_PREFIX", "database.sslmode": "SSL_MODE", "database.trustServerCertificate": "true", "database.user": "sqlserver", "include.schema.changes": "true", "key.converter.basic.auth.credentials.source": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "key.converter.schema.registry.basic.auth.user.info": "USER:PASS", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter": "io.confluent.connect.avro.AvroConverter", "name": "CONNECTOR_NAME", "plugin.name": "PLUGIN_NAME", "poll.interval.ms": "500", "publication.name": "PG_PUBLICATION_NAME", "schema.history.internal.consumer.security.protocol": "SSL", "schema.history.internal.consumer.ssl.key.password": "password", "schema.history.internal.consumer.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "schema.history.internal.consumer.ssl.keystore.password": "password", "schema.history.internal.consumer.ssl.keystore.type": "PKCS12", "schema.history.internal.consumer.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "schema.history.internal.consumer.ssl.truststore.password": "password", "schema.history.internal.kafka.bootstrap.servers": "URL.com:10934", "schema.history.internal.kafka.topic": "sql-testing-history", "schema.history.internal.producer.security.protocol": "SSL", "schema.history.internal.producer.ssl.key.password": "password", "schema.history.internal.producer.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "schema.history.internal.producer.ssl.keystore.password": "password", "schema.history.internal.producer.ssl.keystore.type": "PKCS12", "schema.history.internal.producer.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "schema.history.internal.producer.ssl.truststore.password": "password", "slot.name": "PG_SLOT_NAME", "table.include.list": "PG_TABLES", "tasks.max":"NR_TASKS", "topic.prefix": "sql_topic", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter": "io.confluent.connect.avro.AvroConverter" } ``` ``` { "connector.class": "io.debezium.connector.postgresql.PostgresConnector", "database.dbname": "PG_DATABASE_NAME", "database.hostname": "PG_HOST", "database.password": "PG_PASSWORD", "database.port": "PG_PORT", "database.server.name": "KAFKA_TOPIC_PREFIX", "database.sslmode": "SSL_MODE", "database.user": "PG_USER", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter": "io.confluent.connect.avro.AvroConverter", "name":"CONNECTOR_NAME", "plugin.name": "PLUGIN_NAME", "publication.name": "PG_PUBLICATION_NAME", "slot.name": "PG_SLOT_NAME", "table.include.list": "PG_TABLES", "tasks.max":"NR_TASKS", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter": "io.confluent.connect.avro.AvroConverter" } ``` The configuration file contains the following entries: * `name`: The name of the connector. * `PG_HOST`, `PG_PORT`, `PG_DATABASE_NAME`, `SSL_MODE`, `PG_USER`, `PG_PASSWORD`, `PG_TABLES`, `PG_PUBLICATION_NAME` and `PG_SLOT_NAME`: Source database parameters collected in the [prerequisite](/docs/products/kafka/kafka-connect/howto/debezium-source-connector-pg.md#connect_debezium_pg_source_prereq) phase. * `database.server.name`: The logical name of the database, which determines the prefix used for Apache Kafka topic names. The resulting topic name is a combination of the `database.server.name` and the table name. * `tasks.max`: The maximum number of tasks to execute in parallel. By default, this is 1, the connector can use at most 1 task for each source table defined. * `plugin.name`: Defines the [PostgreSQL output plugin](https://debezium.io/documentation/reference/connectors/postgresql.html) used to convert changes in the database into events in Apache Kafka. warning The `wal2json` logical decoding plugin has certain limitations with the data types it can support. Apart from the basic data types, it converts all other data types into strings based on their textual representation. If you use complex data types, verify the corresponding `wal2json` string representation. * `key.converter` and `value.converter`: Defines the messages data format in the Apache Kafka topic. The `io.confluent.connect.avro.AvroConverter` converter pushes messages in Avro format. To store the message schemas, Aiven's [Karapace schema registry](https://github.com/Aiven-Open/karapace) is used, specified by the `schema.registry.url` parameter and related credentials. note The `key.converter` and `value.converter` sections are only needed when pushing data in Avro format. Otherwise, messages default to JSON format. tip Check the [dedicated blog post](https://aiven.io/blog/db-technology-migration-with-apache-kafka-and-kafka-connect) for an end-to-end example of the Debezium source connector in action with PostgreSQL. ### Create a Kafka Connect connector with Aiven CLI[​](#create-a-kafka-connect-connector-with-aiven-cli "Direct link to Create a Kafka Connect connector with Aiven CLI") To create the connector, execute the following [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create), replacing the `SERVICE_NAME` with the name of the Aiven service where the connector needs to run: ``` avn service connector create SERVICE_NAME @debezium_source_pg.json ``` Check the connector status with the following command, replacing the `SERVICE_NAME` with the Aiven service and the `CONNECTOR_NAME` with the name of the connector defined before: ``` avn service connector status SERVICE_NAME CONNECTOR_NAME ``` Verify the presence of the topic and data in the Apache Kafka target instance. tip With Aiven for Apache Kafka, topics are not created automatically. You have two options: * Manually create topics using the naming pattern: `database.server.name.schema_name.table_name`. * Enable the `Kafka topic auto-creation` feature. See [Enable automatic topic creation with Aiven CLI](/docs/products/kafka/howto/create-topics-automatically.md) ## Solve the error `must be superuser to create FOR ALL TABLES publication`[​](#solve-the-error-must-be-superuser-to-create-for-all-tables-publication "Direct link to solve-the-error-must-be-superuser-to-create-for-all-tables-publication") When creating a Debezium source connector with Aiven for PostgreSQL as the target using `pgoutput` plugin, you might encounter the following error: ``` Caused by: org.postgresql.util.PSQLException: ERROR: must be superuser to create FOR ALL TABLES publication ``` This error occurs when Debezium attempts to create a publication and fails because `avnadmin` is not a superuser. You can address this issue in two ways: * Add the `"publication.autocreate.mode": "filtered"` parameter to the Debezium connector configuration. This enables the creation of publications only for the tables specified in the `table.include.list` parameter. * Create the publication on the source database before configuring the connector, as detailed in the following section. The older versions of Debezium had a bug that prevented the addition of more tables to the filter when using `filtered` mode. As a result, this configuration did not conflict with a `FOR ALL TABLES` publication. Starting with Debezium 1.9.7, these configurations are causing conflicts which result in the following error message: ``` Caused by: org.postgresql.util.PSQLException: ERROR: publication "dbz_publication" is defined as FOR ALL TABLES Detail: Tables cannot be added to or dropped from FOR ALL TABLES publications. ``` This error is triggered when Debezium tries to include more tables in the publication, which is incompatible with `FOR ALL TABLES`. To resolve this error, remove the `publication.autocreate.mode` configuration default to `all_tables`. To maintain `filtered` mode, recreate the publication and replication slot accordingly. ### Create the publication in PostgreSQL[​](#create-the-publication-in-postgresql "Direct link to Create the publication in PostgreSQL") To create the publication in PostgreSQL: * Installing the `aiven-extras` extension: ``` CREATE EXTENSION aiven_extras CASCADE; ``` * Create a publication (with name for example, `my_test_publication`) for all the tables: ``` SELECT * FROM aiven_extras.pg_create_publication_for_all_tables( 'my_test_publication', 'INSERT,UPDATE,DELETE' ); ``` * Make sure to use the correct publication name (for example, `my_test_publication`) in the connector definition and restart the connector --- # Handle PostgreSQL® node replacements when using Debezium for change data capture When you run a [Debezium source connector for PostgreSQL®](/docs/products/kafka/kafka-connect/howto/debezium-source-connector-pg.md) with an Aiven for PostgreSQL® service, some database operations can interrupt change data capture (CDC). For example, if the PostgreSQL service undergoes an operation that replaces nodes (such as maintenance, a plan change, a cloud region change, or a node replacement), the Debezium connector can lose its connection to the database. In many cases, Debezium resumes CDC after the connector tasks restart. In some cases, the PostgreSQL replication slot used by Debezium starts lagging. This can cause WAL files to accumulate and increase disk usage. tip Use the GitHub repository to set up and test a Debezium node replacement scenario. For guidance, see the [Aiven Debezium test repository](https://github.com/aiven/debezium-pg-kafka-connect-test). ## Common Debezium errors during PostgreSQL node replacement[​](#common-debezium-errors-during-postgresql-node-replacement "Direct link to Common Debezium errors during PostgreSQL node replacement") If the Debezium connector cannot recover during or after a PostgreSQL node replacement, the following errors commonly appear in the logs: ``` # ERROR 1 org.apache.kafka.connect.errors.ConnectException: Could not create PostgreSQL connection # ERROR 2 io.debezium.DebeziumException: Could not execute heartbeat action (Error: 57P01) # ERROR 3 org.PostgreSQL.util.PSQLException: ERROR: replication slot "SLOT_NAME" is active for PID xxxx ``` These errors are not recoverable automatically. Restart the connector tasks to resume CDC. You can restart connector tasks in one of the following ways: * **Aiven Console:** Open the service where the connector is running and go to **Connectors**. * **Kafka Connect REST API:** Send the request to the Kafka Connect REST API endpoint. To find the Kafka Connect service URI, open the Aiven Console and go to the Aiven for Apache Kafka Connect service used by the connector. On the **Overview** page, under **Connection information**, copy the value in **Service URI**. ![The Aiven Console page showing the Debezium connector error](/docs/assets/images/pg-debezium-cdc_image-e3ae1c96c6bf5214bf7a0c06b2195125.png) tip To automatically restart failed tasks, set `_aiven.restart.on.failure=true` in the connector configuration. Aiven checks task status every 15 minutes by default. You can change the interval if needed. For details, see [Enable automatic restart](/docs/products/kafka/kafka-connect/howto/enable-automatic-restart.md). ## Handle growing replication lag after Debezium connector restart[​](#handle-growing-replication-lag-after-debezium-connector-restart "Direct link to Handle growing replication lag after Debezium connector restart") As described in the [Debezium documentation](https://debezium.io/documentation/reference/stable/connectors/postgresql.html#postgresql-wal-disk-space), replication lag can increase after the connector tasks restart for two common reasons: 1. **High write volume in the database:** Many updates occur in the database, but only a small portion affects the tables and schemas Debezium monitors. In this case, Debezium might not advance the replication slot fast enough to prevent WAL files from accumulating. To address this issue, enable periodic heartbeat events by setting `heartbeat.interval.ms`. 2. **Multiple databases on the same PostgreSQL instance:** One database generates high write traffic, while Debezium captures changes from a low-traffic database. Replication slots operate per database, so Debezium cannot advance `confirmed_flush_lsn` unless it receives events for the monitored database. Because WAL is shared across databases, WAL files can accumulate until Debezium receives an event from the monitored database. In Aiven testing, this behavior occurred in the following scenarios: 1. The monitored tables had no changes and heartbeat events were not enabled. 2. The monitored tables had no changes, heartbeat events were enabled (`heartbeat.interval.ms` and `heartbeat.action.query`), but the connector did not emit heartbeat events. note This heartbeat issue is tracked in Debezium bug [DBZ-3746](https://issues.redhat.com/browse/DBZ-3746). ### Clear the replication lag[​](#clear-the-replication-lag "Direct link to Clear the replication lag") To clear replication lag, generate activity on monitored tables so Debezium can advance the replication slot: * Resume database traffic (if you paused writes). * Make a change to any monitored table. After Debezium processes the event, it updates the confirmed LSN and PostgreSQL can reclaim WAL space. ## Ensure Debezium survives a PostgreSQL node replacement[​](#ensure-debezium-survives-a-postgresql-node-replacement "Direct link to Ensure Debezium survives a PostgreSQL node replacement") To prevent missing CDC events during node replacement and ensure Debezium recovers correctly, use one of the following approaches. ### Recover safely from Debezium failures[​](#recover-safely-from-debezium-failures "Direct link to Recover safely from Debezium failures") This approach prevents missing events by stopping writes to the source database until Debezium recreates the replication slot on the new primary. 1. Stop write traffic to the database immediately. After node replacement, replication slots are not recreated automatically on the newly promoted primary. When Debezium recovers, it recreates the replication slot. If writes occur before the slot is recreated, Debezium cannot capture those changes. If you must recover missing data, temporarily set `snapshot.mode=always` and restart the connector tasks. This republishes snapshot data to the Kafka topics. After the snapshot completes, restore the configuration to the default to avoid generating a snapshot on every restart. 2. Restart the failed connector tasks. 3. Verify that the replication slot exists and is active. Query `pg_replication_slots` and confirm the slot is present and active. 4. Resume write operations. ### Automate replication slot recreation and verification[​](#automate-replication-slot-recreation-and-verification "Direct link to Automate replication slot recreation and verification") This approach requires an automated process that recreates the replication slot on the new primary after node replacement. Debezium recommends recreating the replication slot before allowing applications to write to the new primary: note There must be a process that re-creates the Debezium replication slot before allowing applications to write to the new primary. This is crucial. Without this process, your application can miss change events. Debezium can recreate the replication slot after it recovers, but this can take time. An automated process that recreates the slot immediately after node replacement reduces downtime and lowers the risk of missed change events. When Debezium restarts, it reads all changes written *after* the replication slot is created. note The example in the [Aiven test repository](https://github.com/aiven/debezium-pg-kafka-connect-test/blob/6f1e6e829ba06bbc396fc0faf28be9e0268ad4f8/bin/python_scripts/debezium_pg_producer.py#L164) blocks inserts unless the Debezium replication slot is active. In most cases, it is sufficient to verify that the replication slot exists, even if it is inactive. Once the connector resumes CDC, it reads all changes written since the slot was created. Debezium also recommends verifying that it has read all changes from the replication slot before the old primary failed. To ensure downstream consumers receive all events, implement a verification method that confirms changes to monitored tables were recorded. tip The [Aiven test repository](https://github.com/aiven/debezium-pg-kafka-connect-test/blob/main/bin/python_scripts/debezium_pg_producer.py) includes an example implementation. ## Handle PostgreSQL major version upgrades[​](#handle-postgresql-major-version-upgrades "Direct link to Handle PostgreSQL major version upgrades") A PostgreSQL major version upgrade replaces database nodes. This can interrupt Debezium change data capture (CDC) in the same way as a PostgreSQL node replacement. During the upgrade, Debezium tasks can fail and the replication slot might not be available immediately on the new primary. If applications write to monitored tables before the slot is recreated, downstream consumers can miss change events, leading to potential data loss. For guidance, see [Steps to protect against data loss when using Debezium and upgrading the PostgreSQL source](https://debezium.io/documentation/reference/stable/connectors/postgresql.html#upgrading-postgresql). Related pages [Upgrade Aiven for PostgreSQL® to a major version](/docs/products/postgresql/howto/upgrade.md#upgrade-to-a-major-version) --- # Create a Debezium source connector from SQL Server to Apache Kafka® with CDC The SQL Server Debezium source connector uses the [change data capture (CDC) ](https://learn.microsoft.com/en-us/sql/relational-databases/track-changes/about-change-data-capture-sql-server?view=sql-server-2017)feature to extract database changes from designated tables and write them to Apache Kafka® topic in a standard format for multiple consumers to read and transform. note Aiven supports multiple Debezium versions through multi-version support, including versions 1.9.7, 2.5.0, 2.7.4, and 3.1.0. Debezium 2.5 introduced changes to connector configuration and behavior. To prevent unintentional upgrades during maintenance updates, pin the connector version using the `plugin_versions` configuration property. For details, see [Manage connector versions](/docs/products/kafka/kafka-connect/howto/manage-connector-versions.md). If you use Debezium for PostgreSQL version 1.9.7 with the `wal2json` replication format, do not upgrade to version 2.0 or later until you migrate to a supported format such as `pgoutput`. To upgrade from version 1.9.7, use multi-version support to test your configuration before applying changes in production. For further assistance, contact [Aiven support](mailto:support@aiven.io). ## Enable CDC in SQL Server[​](#connect_debezium_sql_server_schema_versioning "Direct link to Enable CDC in SQL Server") To use the Debezium source connector for SQL Server, enable the SQL Server Change Data Capture (CDC) at both the database and table levels. This creates the necessary schemas and tables containing a history of all change events in the tables you wish to track with the connector. ### Enable CDC at database level[​](#enable-cdc-at-database-level "Direct link to Enable CDC at database level") To enable the CDC at database level, you can use the following command: ``` USE GO EXEC sys.sp_cdc_enable_db GO ``` note If you're using GCP Cloud SQL for SQL Server, you can enable database CDC with: ``` EXEC msdb.dbo.gcloudsql_cdc_enable_db '' ``` Enabling CDC creates a new schema called `cdc` in the target database, which contains all necessary tables. ### Enable CDC at table level[​](#enable-cdc-at-table-level "Direct link to Enable CDC at table level") To enable CDC for a table you can execute the following command: ``` USE GO EXEC sys.sp_cdc_enable_table @source_schema = N'', @source_name = N'', @role_name = N'', @filegroup_name = N'', @supports_net_changes = 0 GO ``` The command above has the following parameters: * ``, ``, ``: The references to the table where CDC needs to be set. * ``: The database role that will have access to the change tables. Leave it `NULL` to only allow access to members of `sysadmin` or `db_owner` groups. * ``: Specifies the [file group](https://learn.microsoft.com/en-us/sql/relational-databases/databases/database-files-and-filegroups) where the files will be written, needs to be pre-existing. note If you're using GCP Cloud SQL for SQL Server, you can enable database CDC on a table with: ``` EXEC sys.sp_cdc_enable_table @source_schema = N'', @source_name = N'', @role_name = N'' ``` warning When modifying table schemas online, new column information can be lost until CDC is re-enabled for the table. For more details, refer to the [related Debezium documentation](https://debezium.io/documentation/reference/stable/connectors/sqlserver.html#sqlserver-schema-evolution). ## Prerequisites[​](#connect_debezium_sql_server_source_prereq "Direct link to Prerequisites") To configure a Debezium source connector for MongoDB, you need either an Aiven for Apache Kafka service with [Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). Before you begin, gather the necessary information about your source MongoDB database: * `SQLSERVER_HOST`: The database hostname * `SQLSERVER_PORT`: The database port * `SQLSERVER_USER`: The database user to connect * `SQLSERVER_PASSWORD`: The database password for the `SQLSERVER_USER` * `SQLSERVER_DATABASE_NAME`: The database name * `SQLSERVER_TABLES`: The list of database tables to be included in Apache Kafka. Format the list as `schema_name1.table_name1,schema_name2.table_name2` * `APACHE_KAFKA_HOST`: The hostname of the Apache Kafka service, needed when storing the [schema definition changes](/docs/products/kafka/kafka-connect/howto/debezium-source-connector-sql-server.md#connect_debezium_sql_server_schema_versioning) * `APACHE_KAFKA_PORT`: The port of the Apache Kafka service, needed when storing the [schema definition changes](/docs/products/kafka/kafka-connect/howto/debezium-source-connector-sql-server.md#connect_debezium_sql_server_schema_versioning) * `SCHEMA_REGISTRY_PORT`: The Apache Kafka's schema registry port, only needed when using Avro as data format * `SCHEMA_REGISTRY_USER`: The Apache Kafka's schema registry username, only needed when using Avro as data format * `SCHEMA_REGISTRY_PASSWORD`: The Apache Kafka's schema registry user password, only needed when using Avro as data format If you are using Aiven for PostgreSQL and Aiven for Apache Kafka the above details are available in the [Aiven console](https://console.aiven.io/) service Overview page or via the dedicated `avn service get` command with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). For a complete list of all available parameters and configuration options, see Debezium [connector's documentation](https://debezium.io/documentation/reference/stable/connectors/sqlserver.html). ## Setup a SQL Server Debezium source connector with Aiven Console[​](#setup-a-sql-server-debezium-source-connector-with-aiven-console "Direct link to Setup a SQL Server Debezium source connector with Aiven Console") The following example demonstrates how to setup a Debezium source Connector for Apache Kafka to a SQL Server database using the [Aiven Console](https://console.aiven.io/). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Create a configuration file named `debezium_source_mysql.json` with the following connector configurations. While optional, creating this file helps you organize your settings in one place and copy/paste them into the [Aiven Console](https://console.aiven.io/) later. ``` { "name":"CONNECTOR_NAME", "connector.class": "io.debezium.connector.sqlserver.SqlServerConnector", "database.hostname": "SQLSERVER_HOST", "database.port": "SQLSERVER_PORT", "database.user": "SQLSERVER_USER", "database.password": "SQLSERVER_PASSWORD", "database.dbname": "SQLSERVER_DATABASE_NAME", "database.server.name": "KAFKA_TOPIC_PREFIX", "table.include.list": "SQLSERVER_TABLES", "tasks.max":"NR_TASKS", "poll.interval.ms": 500, "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "schema.history.internal.kafka.topic": "HISTORY_TOPIC_NAME", "schema.history.internal.kafka.bootstrap.servers": "APACHE_KAFKA_HOST:APACHE_KAFKA_PORT", "schema.history.internal.producer.security.protocol": "SSL", "schema.history.internal.producer.ssl.keystore.type": "PKCS12", "schema.history.internal.producer.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "schema.history.internal.producer.ssl.keystore.password": "password", "schema.history.internal.producer.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "schema.history.internal.producer.ssl.truststore.password": "password", "schema.history.internal.producer.ssl.key.password": "password", "schema.history.internal.consumer.security.protocol": "SSL", "schema.history.internal.consumer.ssl.keystore.type": "PKCS12", "schema.history.internal.consumer.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "schema.history.internal.consumer.ssl.keystore.password": "password", "schema.history.internal.consumer.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "schema.history.internal.consumer.ssl.truststore.password": "password", "schema.history.internal.consumer.ssl.key.password": "password", "include.schema.changes": "true" } ``` The configuration file contains the following entries: * `name`: The connector name, replace CONNECTOR\_NAME with the name you want to use for the connector. * `SQLSERVER_HOST`, `SQLSERVER_PORT`, `SQLSERVER_DATABASE_NAME`, `SSL_MODE`, `SQLSERVER_USER`, `SQLSERVER_PASSWORD`, `SQLSERVER_TABLES`: Source database parameters collected in the [prerequisite](/docs/products/kafka/kafka-connect/howto/debezium-source-connector-sql-server.md#connect_debezium_sql_server_source_prereq) phase. * `database.server.name`: The logical name of the database, which determines the prefix used for Apache Kafka topic names. The resulting topic name is a combination of the `database.server.name` and the table name. * `tasks.max`: Maximum number of tasks to execute in parallel. By default this is 1, the connector can use at most 1 task for each source table defined. Replace `NR_TASKS` with the amount of parallel task based on the number of input tables. * `poll.interval.ms`: The frequency of the queries to the CDC tables. * `database.history.kafka.topic`: The name of the Apache Kafka topic that contains the history of schema changes. * `database.history.kafka.bootstrap.servers`: Directs to the Aiven for Apache Kafka service where the connector is running and is needed to store [schema definition changes](/docs/products/kafka/kafka-connect/howto/debezium-source-connector-sql-server.md#connect_debezium_sql_server_schema_versioning) * `database.history.producer` and `database.history.consumer`: Refers to truststores and keystores pre-created on the Aiven for Apache Kafka node to handle SSL authentication warning The values defined for each `database.history.producer` and `database.history.consumer` parameters are already set to work with the predefined truststore and keystore created in the Aiven for Apache Kafka nodes. Modifying these values is not recommended. * `key.converter` and `value.converter`: Defines the messages data format in the Apache Kafka topic. The `io.confluent.connect.avro.AvroConverter` converter pushes messages in Avro format. To store the message schemas, Aiven's [Karapace schema registry](https://github.com/Aiven-Open/karapace) is used, specified by the `schema.registry.url` parameter and related credentials. note The `key.converter` and `value.converter` sections are only needed when pushing data in Avro format. Otherwise, messages default to JSON format. The `USER_INFO` is not a placeholder and does not require any parameter substitution. ### Create a Kafka Connect connector with the Aiven Console[​](#create-a-kafka-connect-connector-with-the-aiven-console "Direct link to Create a Kafka Connect connector with the Aiven Console") To create a Kafka Connect connector: 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. Select the Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service to define the connector. 3. Select **Manage stream** > **Connectors** from the sidebar. 4. Select **Create New Connector**, which is available only for services [that have Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md). 5. Select **Debezium - SQL Server**. 6. In the **Common** tab, locate the **Connector configuration** text box and select on **Edit**. 7. Paste the connector configuration (stored in the `debezium_source_sql_server.json` file) in the form. 8. Select **Apply**. note The Aiven Console reads through the configuration file and automatically populates the relevant UI fields. You can view and modify these fields across different tabs. Any change you make is reflected in JSON format within the **Connector configuration** text box. 9. After all the settings are correctly configured, select **Create connector**. tip With Aiven for Apache Kafka, topics are not created automatically. You have two options: * Manually create topics using the naming pattern: `database.server.name.schema_name.table_name`. * Enable the `Kafka topic auto-creation` feature. See [Enable automatic topic creation with Aiven CLI](/docs/products/kafka/howto/create-topics-automatically.md). 10. Verify the connector status under **Manage stream** > **Connectors**. 11. Verify the presence of the data in the target Apache Kafka topic coming from the MySQL dataset. The topic name is equal to concatenation of the database and table name. To change the target table name, you can use Apache Kafka Connect `RegexRouter` transformation. note You can also create connectors using the [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create). --- # Create a sink connector from Apache Kafka® to Elasticsearch The Elasticsearch sink connector enables you to move data from an Aiven for Apache Kafka® cluster to an Elasticsearch instance for further processing and analysis. note * To see the equivalent instructions for OpenSearch®, see [Create a sink connector to OpenSearch®](/docs/products/kafka/kafka-connect/howto/opensearch-sink.md). * See the full set of available parameters and configuration options in the [connector's documentation](https://docs.confluent.io/current/connect/kafka-connect-elasticsearch/index). ## Prerequisites[​](#connect_elasticsearch_sink_prereq "Direct link to Prerequisites") * an Aiven for Apache Kafka service [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * Collect the following information about the target Elasticsearch service: * `ES_CONNECTION_URL`: The Elasticsearch connection URL, in the form of `https://HOST:PORT` * `ES_USERNAME`: The Elasticsearch username to connect * `ES_PASSWORD`: The password for the username selected * `TOPIC_LIST`: The list of topics to sink divided by comma * `APACHE_KAFKA_HOST`: The hostname of the Apache Kafka service, only needed when using Avro as data format * `SCHEMA_REGISTRY_PORT`: The Apache Kafka's schema registry port, only needed when using Avro as data format * `SCHEMA_REGISTRY_USER`: The Apache Kafka's schema registry username, only needed when using Avro as data format * `SCHEMA_REGISTRY_PASSWORD`: The Apache Kafka's schema registry user password, only needed when using Avro as data format note If you're using Aiven for Elasticsearch and Aiven for Apache Kafka® the above details are available in the [Aiven console](https://console.aiven.io/) service Overview tab or via the dedicated `avn service get` command with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). The `SCHEMA_REGISTRY` related parameters are available in the Aiven for Apache Kafka® service page, *Overview* tab, and *Schema Registry* subtab As of version 3.0, Aiven for Apache Kafka no longer supports Confluent Schema Registry. For more information, read [the article describing the replacement, Karapace](https://help.aiven.io/en/articles/5651983) ## Setup an Elasticsearch sink connector with Aiven Console[​](#setup-an-elasticsearch-sink-connector-with-aiven-console "Direct link to Setup an Elasticsearch sink connector with Aiven Console") The following example demonstrates how to setup a Elasticsearch sink connector for Apache Kafka using the [Aiven Console](https://console.aiven.io/). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Define the connector configurations in a file (we'll refer to it with the name `elasticsearch_sink.json`) with the following content: ``` { "name":"CONNECTOR_NAME", "connector.class": "io.aiven.connect.elasticsearch.ElasticsearchSinkConnector", "topics": "TOPIC_LIST", "connection.url": "ES_CONNECTION_URL", "connection.username": "ES_USERNAME", "connection.password": "ES_PASSWORD", "type.name": "TYPE_NAME", "tasks.max":"1", "key.ignore": "true", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } ``` The configuration file contains the following entries: * `name`: the connector name * `connection.url`, `connection.username`, `connection.password`: sink Elasticsearch parameters collected in the [prerequisite](/docs/products/kafka/kafka-connect/howto/elasticsearch-sink.md#connect_elasticsearch_sink_prereq) phase. * `type.name`: the Elasticsearch type name to be used when indexing. * `key.ignore`: boolean flag dictating if to ignore the message key. If set to true, the document ID is generated as message's `topic+partition+offset`, the message key is used as ID otherwise. * `tasks.max`: maximum number of tasks to execute in parallel. By default this is 1. * `key.converter` and `value.converter`: defines the messages data format in the Apache Kafka topic. The `io.confluent.connect.avro.AvroConverter` converter translates messages from the Avro format. To retrieve the messages schema we use Aiven's [Karapace schema registry](https://github.com/aiven/karapace) as specified by the `schema.registry.url` parameter and related credentials. note The `key.converter` and `value.converter` sections are only needed when the source data is in Avro format. If omitted the messages will be read as binary format. When using Avro as source data format, set following parameters: * `value.converter.schema.registry.url`: pointing to the Aiven for Apache Kafka schema registry URL in the form of `https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT` with the `APACHE_KAFKA_HOST` and `SCHEMA_REGISTRY_PORT` parameters [retrieved in the previous step](/docs/products/kafka/kafka-connect/howto/elasticsearch-sink.md#connect_elasticsearch_sink_prereq). * `value.converter.basic.auth.credentials.source`: to the value `USER_INFO`, since you're going to login to the schema registry using username and password. * `value.converter.schema.registry.basic.auth.user.info`: passing the required schema registry credentials in the form of `SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD` with the `SCHEMA_REGISTRY_USER` and `SCHEMA_REGISTRY_PASSWORD` parameters [retrieved in the previous step](/docs/products/kafka/kafka-connect/howto/elasticsearch-sink.md#connect_elasticsearch_sink_prereq). ### Create a Kafka Connect connector with Aiven Console[​](#create-a-kafka-connect-connector-with-aiven-console "Direct link to Create a Kafka Connect connector with Aiven Console") To create a Kafka Connect connector: 1. Log in to the [Aiven Console](https://console.aiven.io/) and select the Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service where the connector needs to be defined. 2. Select **Manage stream** > **Connectors** from the left sidebar. 3. Select **Create New Connector**, it is enabled only for services [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md). 4. Select **Elasticsearch sink**. 5. In the **Common** tab, locate the **Connector configuration** text box and select on **Edit**. 6. Paste the connector configuration (stored in the `elasticsearch_sink.json` file) in the form. 7. Select **Apply**. note The Aiven Console parses the configuration file and fills the relevant UI fields. You can review the UI fields across the various tab and change them if necessary. The changes will be reflected in JSON format in the **Connector configuration** text box. 8. After all the settings are correctly configured, select **Create connector**. 9. Verify the connector status under **Manage stream** > **Connectors**. 10. Verify the presence of the data in the target Elasticsearch service, the index name is equal to the Apache Kafka topic name. note You can also create connectors using the [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create). ## Create daily Elasticsearch indices[​](#create-daily-elasticsearch-indices "Direct link to Create daily Elasticsearch indices") You might need to create an Elasticsearch index on daily basis to store the Apache Kafka messages. Adding the following `TimestampRouter` transformation in the connector properties file provides a way to define the index name as concatenation of the topic name and message date. ``` "transforms": "TimestampRouter", "transforms.TimestampRouter.topic.format": "${topic}-${timestamp}", "transforms.TimestampRouter.timestamp.format": "yyyy-MM-dd", "transforms.TimestampRouter.type": "org.apache.kafka.connect.transforms.TimestampRouter" ``` warning The current version of the Elasticsearch sink connector is not able to automatically create daily indices in Elasticsearch. Therefore you need to create the indices with the correct name before starting the sink connector. You can create Elasticsearch indices in many ways including [CURL commands](/docs/products/opensearch/howto/opensearch-with-curl.md). ## Example: Create an Elasticsearch sink connector on a topic with a JSON schema[​](#example-create-an-elasticsearch-sink-connector-on-a-topic-with-a-json-schema "Direct link to Example: Create an Elasticsearch sink connector on a topic with a JSON schema") If you have a topic named `iot_measurements` containing the following data in JSON format, with a defined JSON schema: ``` { "schema": { "type":"struct", "fields":[{ "type":"int64", "optional": false, "field": "iot_id" },{ "type":"string", "optional": false, "field": "metric" },{ "type":"int32", "optional": false, "field": "measurement" }] }, "payload":{ "iot_id":1, "metric":"Temperature", "measurement":14} } { "schema": { "type":"struct", "fields":[{ "type":"int64", "optional": false, "field": "iot_id" },{ "type":"string", "optional": false, "field": "metric" },{ "type":"int32", "optional": false, "field": "measurement" }] }, "payload":{"iot_id":2, "metric":"Humidity", "measurement":60}} } ``` note Since the JSON schema needs to be defined in every message, there is a big overhead to transmit the information. To achieve a better performance in term of information-message ratio you should use the Avro format together with the [Karapace schema registry](https://karapace.io/) provided by Aiven You can sink the `iot_measurements` topic to Elasticsearch with the following connector configuration, after replacing the placeholders for `ES_CONNECTION_URL`, `ES_USERNAME` and `ES_PASSWORD`: ``` { "name":"sink_iot_json_schema", "connector.class": "io.aiven.connect.elasticsearch.ElasticsearchSinkConnector", "topics": "iot_measurements", "connection.url": "ES_CONNECTION_URL", "connection.username": "ES_USERNAME", "connection.password": "ES_PASSWORD", "type.name": "iot_measurements", "tasks.max":"1", "key.ignore": "true", "value.converter": "org.apache.kafka.connect.json.JsonConverter" } ``` The configuration file contains the following peculiarities: * `"topics": "iot_measurements"`: setting the topic to sink * `"value.converter": "org.apache.kafka.connect.json.JsonConverter"`: the message value is in plain JSON format without a schema * `"key.ignore": "true"`: the connector is ignoring the message key (empty), and generating documents with ID equal to `topic+partition+offset` ## Example: Create an Elasticsearch sink connector on a topic in plain JSON format[​](#example-create-an-elasticsearch-sink-connector-on-a-topic-in-plain-json-format "Direct link to Example: Create an Elasticsearch sink connector on a topic in plain JSON format") If you have a topic named `students` containing the following data in JSON format, without a defined schema: ``` Key: 1 Value: {"student_id":1, "student_name":"Carla"} Key: 2 Value: {"student_id":2, "student_name":"Ugo"} Key: 3 Value: {"student_id":3, "student_name":"Mary"} ``` You can sink the `students` topic to Elasticsearch with the following connector configuration, after replacing the placeholders for `ES_CONNECTION_URL`, `ES_USERNAME` and `ES_PASSWORD`: ``` { "name":"sink_students_json", "connector.class": "io.aiven.connect.elasticsearch.ElasticsearchSinkConnector", "topics": "students", "connection.url": "ES_CONNECTION_URL", "connection.username": "ES_USERNAME", "connection.password": "ES_PASSWORD", "type.name": "students", "tasks.max":"1", "key.converter": "org.apache.kafka.connect.storage.StringConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter.schemas.enable": "false", "schema.ignore": "true" } ``` The configuration file contains the following peculiarities: * `"topics": "students"`: setting the topic to sink * `"key.converter": "org.apache.kafka.connect.storage.StringConverter"`: the message key is a string * `"value.converter": "org.apache.kafka.connect.json.JsonConverter"`: the message value is in plain JSON format without a schema * `"value.converter.schemas.enable": "false"`: since the data in the value doesn't have a schema, the connector shouldn't try to read it and sets it to null * `"schema.ignore": "true"`: since the value schema is null, the connector doesn't infer it before pushing the data to Elasticsearch note The Elasticsearch document ID is set as the message key Related pages * [Create a sink connector to OpenSearch®](/docs/products/kafka/kafka-connect/howto/opensearch-sink.md). --- # Enable automatic restart for Apache Kafka® Connect connectors Automatic restart can help recover a connector task after a rare transient failure, such as an out-of-memory error caused by a sudden data surge. Before you enable automatic restart, investigate the cause of the connector failure. Automatic restart is not recommended for recurring failures because the connector might continue to fail until the underlying issue is fixed. note In some cases, the Debezium source connector for PostgreSQL® can stop after PostgreSQL maintenance if PostgreSQL cannot create replication slots before failover. Automatic restart can help recover the connector in this scenario. ## Enable automatic restart in the Aiven Console[​](#enable-automatic-restart-in-the-aiven-console "Direct link to Enable automatic restart in the Aiven Console") All Aiven-managed Apache Kafka Connect connectors support automatic restart. To enable automatic restart for an existing connector: 1. In the [Aiven Console](https://console.aiven.io/), select the Aiven for Apache Kafka or Aiven for Apache Kafka Connect service where the connector is defined. 2. Select **Manage stream** > **Connectors** from the sidebar. 3. Open the connector to update. 4. Select the **Aiven** tab. 5. For **Automatic restart**, select **True**. After you enable automatic restart, the connector restarts automatically when a failure is detected. warning Automatic restart restarts the entire connector, not only the failed tasks. --- # Enable Apache Kafka® Connect on Aiven for Apache Kafka® For a low-cost way to get started with Aiven for Apache Kafka® Connect, you can run Kafka Connect on the same nodes as your Apache Kafka cluster, sharing the resources. The Kafka service must be running on a business or premium plan. To reduce load on the Kafka nodes and make the cluster more stable, you can [create a standalone Kafka Connect service](/docs/products/kafka/kafka-connect/get-started.md) instead. A standalone service offers more CPU time and memory, and allows you to [scale the service independently](/docs/products/kafka/howto/change-service-plan.md). To enable Apache Kafka Connect on Aiven for Apache Kafka nodes: 1. Select an Aiven for Apache Kafka service. 2. Click **Service settings**. 3. In the **Service management** section, click **Actions** > **Enable Connect**. You can view information about the Apache Kafka Connect service on the **Manage stream** > **Connectors** page. --- # Create a sink connector from Apache Kafka® to Google BigQuery Set up the [BigQuery sink connector](https://github.com/Aiven-Open/bigquery-connector-for-apache-kafka) to move data from Aiven for Apache Kafka® into BigQuery tables for analysis and storage. See the [full list of parameters](https://aiven-open.github.io/bigquery-connector-for-apache-kafka/configuration.html) in the GitHub documentation. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven for Apache Kafka service with [Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md), or a [dedicated Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster) * A Google Cloud project with: * A [service account and JSON service key](https://cloud.google.com/docs/authentication/client-libraries) * BigQuery API enabled * A [BigQuery dataset](https://cloud.google.com/bigquery/docs/datasets) to store data * Dataset access granted to the service account Collect the following details: * `GOOGLE_CLOUD_PROJECT_NAME`: Target Google Cloud project name * `GOOGLE_CLOUD_SERVICE_KEY`: Google Cloud service account key in JSON format warning When adding the service key to the connector configuration, provide it as an escaped string: * Escape all `"` characters as `\"` * Escape any `\n` in the `private_key` field as `\\n` Example: `{\"type\": \"service_account\",\"project_id\": \"XXXXXX\", ...}` * `BIGQUERY_DATASET_NAME`: BigQuery dataset name * `TOPIC_LIST`: Comma-separated list of Kafka topics to sink * Only if using Avro: Schema Registry connection details: * `APACHE_KAFKA_HOST` * `SCHEMA_REGISTRY_PORT` * `SCHEMA_REGISTRY_USER` * `SCHEMA_REGISTRY_PASSWORD` note Schema Registry connection details are available in the **Overview** page of your Kafka service, under **Connection information**, in the **Schema Registry** tab. Aiven uses [Karapace](https://github.com/aiven/karapace) as the Schema Registry. ### Google credential source restrictions[​](#google-credential-source-restrictions "Direct link to Google credential source restrictions") When using Google Cloud external account credentials with Aiven for Apache Kafka® Connect, Aiven applies security restrictions to prevent unauthorized file access and network requests. If the Google Cloud credential JSON file includes a `credential_source` object, the following restrictions apply: * `credential_source.file`: Not allowed * `credential_source.executable.command`: Not allowed * `credential_source.url`: Allowed only when allow-listed in the service configuration (`gcp_auth_allowed_urls`) Example `credential_source` object using a URL-based credential: ``` { "credential_source": { "url": "https://sts.googleapis.com/v1/token", "headers": { "Metadata-Flavor": "Google" } } } ``` important `gcp_auth_allowed_urls` is a **Kafka Connect service-level configuration**, not a connector configuration. Configure it in the Kafka Connect service settings. This setting applies to all connectors in the service. #### Configure allowed authentication URLs[​](#configure-allowed-authentication-urls "Direct link to Configure allowed authentication URLs") To use URL-based credentials (`credential_source.url`), configure the following: 1. Configure the Kafka Connect service with allowed authentication URLs. 2. Configure each connector to use one of the allowed URLs. Set `gcp_auth_allowed_urls` on the Kafka Connect service to define which HTTPS authentication endpoints the service can access. This setting applies to all connectors in the service. * Console * CLI 1. Go to the [Aiven Console](https://console.aiven.io/). 2. Select the Aiven for Apache Kafka Connect service. 3. Click **Service settings**. 4. In **Advanced configuration**, click **Configure**. 5. Set **`gcp_auth_allowed_urls`** to the required HTTPS endpoints. 6. Click **Save configuration**. Run the following command: ``` avn service update KAFKA_CONNECT_SERVICE_NAME \ -c gcp_auth_allowed_urls='["https://sts.googleapis.com","https://iamcredentials.googleapis.com"]' ``` If multiple connectors use URL-based credentials, add all required authentication URLs to `gcp_auth_allowed_urls`. Each unique URL needs to be added only once. If `credential_source.url` is set but the URL is not included in `gcp_auth_allowed_urls`, connector creation fails. In the connector configuration JSON, set `credential_source.url` to match one of the URLs configured in the service. ## Configure Google Cloud[​](#configure-google-cloud "Direct link to Configure Google Cloud") 1. **Create a service account and generate a JSON key:** In the [Google Cloud Console](https://console.cloud.google.com/), create a service account and generate a JSON key. See [Google’s guide](https://cloud.google.com/docs/authentication/client-libraries) for details. You’ll use this key in the connector configuration. 2. **Enable the BigQuery API:** In the [API & Services dashboard](https://console.cloud.google.com/apis), enable the **BigQuery API** if it isn’t already. See [Google’s reference](https://cloud.google.com/bigquery/docs/reference/rest) for details. 3. **Create a BigQuery dataset:** In the [BigQuery Console](https://console.cloud.google.com/bigquery), create a dataset or use an existing one. See [Google’s guide](https://cloud.google.com/bigquery/docs/datasets). Select a region close to your Kafka service to reduce latency. 4. **Grant dataset access to the service account:** In the BigQuery Console, grant your service account the **BigQuery Data Editor** role on the dataset. See [Google’s access control guide](https://cloud.google.com/bigquery/docs/dataset-access-controls). ## Write methods[​](#write-methods "Direct link to Write methods") * **Google Cloud Storage (default):** Uses GCS as an intermediate step. Supports all features, including delete and upsert. Parameters used only with delete or upsert: * `intermediateTableSuffix` * `kafkaKeyFieldName` (required) * `mergeIntervalMs` * **Storage Write API:** Streams data directly into BigQuery. Enable by setting `useStorageWriteApi` to `true`. This method provides lower latency for streaming workloads. Parameters used only with Storage Write API: * `bigQueryPartitionDecorator` * `commitInterval` * `enableBatchMode` warning Do not use the Storage Write API with `deleteEnabled` or `upsertEnabled`. If `useStorageWriteApi` is not set, the connector uses the standard Google Cloud Storage API by default. ## Create a BigQuery sink connector configuration[​](#create-a-bigquery-sink-connector-configuration "Direct link to Create a BigQuery sink connector configuration") Define the connector configuration in a JSON file, for example `bigquery_sink.json`. ``` { "name": "CONNECTOR_NAME", "connector.class": "com.wepay.kafka.connect.bigquery.BigQuerySinkConnector", "topics": "TOPIC_LIST", "project": "GOOGLE_CLOUD_PROJECT_NAME", "defaultDataset": ".*=BIGQUERY_DATASET_NAME", "schemaRetriever": "com.wepay.kafka.connect.bigquery.retrieve.IdentitySchemaRetriever", "schemaRegistryClient.basic.auth.credentials.source": "URL", "schemaRegistryLocation": "https://SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD@APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "autoCreateTables": "true", "allBQFieldsNullable": "true", "keySource": "JSON", "keyfile": "GOOGLE_CLOUD_SERVICE_KEY" } ``` note To use the Storage Write API instead of the default GCS method, add `"useStorageWriteApi": "true"` to the configuration. The configuration file includes: * `name`: Connector name * `connector.class`: Must be `com.wepay.kafka.connect.bigquery.BigQuerySinkConnector` * `topics`: Comma-separated list of Kafka topics to write to BigQuery * `project`: Target Google Cloud project name * `defaultDataset`: BigQuery dataset name, prefixed with `.*=` note By default, table names in BigQuery match the Kafka topic names. Use the Kafka Connect `RegexRouter` transformation to rename tables if needed. If your messages are in Avro format, also set these parameters: * `schemaRegistryLocation`: Karapace schema registry endpoint * `key.converter` and `value.converter`: Set to `io.confluent.connect.avro.AvroConverter` for Avro * `key.converter.schema.registry.url` and `value.converter.schema.registry.url`: Schema registry URL (`https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT`) * `key.converter.schema.registry.basic.auth.user.info` and `value.converter.schema.registry.basic.auth.user.info`: Schema registry credentials in the format `SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD` For table management, you can set: * `autoCreateTables`: Create target BigQuery tables automatically if they do not exist * `allBQFieldsNullable`: Set new BigQuery fields to `NULLABLE` instead of `REQUIRED` For schema evolution, you can also set: * `allowNewBigQueryFields`: Add new fields from Kafka schemas to BigQuery tables * `allowBigQueryRequiredFieldRelaxation`: Relax `REQUIRED` fields back to `NULLABLE` If multiple connector tasks or connector instances write to the same BigQuery table, you can also set: * `mediateConcurrentSchemaUpdates`: Retry failed schema updates and check whether another connector already applied the required schema change. Default: `false` * `concurrentSchemaUpdateRetryWaitMs`: Time to wait before each retry attempt when `mediateConcurrentSchemaUpdates` is enabled, in milliseconds. Allowed range: `0` to `300000` (up to five minutes). Default: `10000` * `concurrentSchemaUpdateMaxRetries`: Maximum number of retry attempts. Default: `3` warning Automatic schema evolution reduces control over table definitions and may cause errors if message schemas change in ways BigQuery cannot support. note Enable `mediateConcurrentSchemaUpdates` when multiple connector tasks or connector instances write to the same BigQuery table and schema evolution is enabled. This setting helps recover from concurrent schema update failures, but it does not remove BigQuery limits, including limits on table metadata updates. For authentication, set: * `keySource`: Format of the Google Cloud service account key, set to `JSON` * `keyfile`: Google Cloud service account key as an escaped string ## Track write attempts for deduplication[​](#track-write-attempts-for-deduplication "Direct link to Track write attempts for deduplication") To help downstream systems distinguish rows written by different Kafka Connect `put()` attempts, enable write-attempt tracking. Set both of the following: ``` { "kafkaDataFieldName": "__kafka", "trackPutAttempts": "true" } ``` When `trackPutAttempts` is enabled, the connector adds a `putAttemptId` field to the Kafka metadata struct in each BigQuery row. The connector generates one ULID value for each Kafka Connect `put()` call. If Kafka Connect retries the same records in a later `put()` call, the retried rows get a different `putAttemptId`. This setting does not prevent duplicate rows. Use `putAttemptId` with Kafka metadata, such as topic, partition, and offset, to support downstream deduplication. important If you enable `trackPutAttempts` on an existing BigQuery table that already has a Kafka metadata field, also set `allowNewBigQueryFields` to `true`. The connector must add `putAttemptId` as a new nullable subfield in the existing metadata `RECORD` schema. ## Create a BigQuery sink connector[​](#create-a-bigquery-sink-connector "Direct link to Create a BigQuery sink connector") * Aiven Console * Aiven CLI 1. Go to the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka® or Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** (enable Kafka Connect if required). 5. Select **Google BigQuery Sink** from the list. 6. On the **Common** tab, click **Edit** in the **Connector configuration** box. 7. Paste the contents of `bigquery_sink.json`. Replace placeholders with actual values. 8. Click **Apply**, then **Create connector**. 9. Verify the connector status on the **Manage stream** > **Connectors** page. 10. Check that data appears in the BigQuery dataset. By default, table names match topic names. Use the Kafka Connect `RegexRouter` transformation to rename tables if required. Run the following command to create the connector: ``` avn service connector create SERVICE_NAME @bigquery_sink.json ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service * `@bigquery_sink.json`: Path to the connector configuration file ## Examples[​](#examples "Direct link to Examples") ### Sink a JSON topic[​](#sink-a-json-topic "Direct link to Sink a JSON topic") The topic `iot_measurements` contains JSON messages with an inline schema: ``` { "schema": { "type": "struct", "fields": [ { "type": "int64", "field": "iot_id" }, { "type": "string", "field": "metric" }, { "type": "int32", "field": "measurement" } ] }, "payload": { "iot_id": 1, "metric": "Temperature", "measurement": 14 } } ``` Connector configuration: ``` { "name": "iot_sink", "connector.class": "com.wepay.kafka.connect.bigquery.BigQuerySinkConnector", "topics": "iot_measurements", "project": "GOOGLE_CLOUD_PROJECT_NAME", "defaultDataset": ".*=BIGQUERY_DATASET_NAME", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "autoCreateTables": "true", "keySource": "JSON", "keyfile": "GOOGLE_CLOUD_SERVICE_KEY" } ``` Parameters: * `topics`: Source topic * `value.converter`: JSON converter without schema note Inline JSON schemas increase message size and add processing overhead. For efficiency, prefer Avro format with Karapace Schema Registry. ### Sink an Avro topic[​](#sink-an-avro-topic "Direct link to Sink an Avro topic") The topic `students` contains messages in Avro format, with schemas stored in Karapace. Connector configuration: ``` { "name": "students_sink", "connector.class": "com.wepay.kafka.connect.bigquery.BigQuerySinkConnector", "topics": "students", "project": "GOOGLE_CLOUD_PROJECT_NAME", "defaultDataset": ".*=BIGQUERY_DATASET_NAME", "schemaRetriever": "com.wepay.kafka.connect.bigquery.retrieve.IdentitySchemaRetriever", "schemaRegistryLocation": "https://SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD@APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "autoCreateTables": "true", "keySource": "JSON", "keyfile": "GOOGLE_CLOUD_SERVICE_KEY" } ``` Parameters: * `topics`: Source topic * `key.converter` and `value.converter`: Enable Avro parsing with Karapace schema registry ### Sink an Avro topic using Storage Write API[​](#sink-an-avro-topic-using-storage-write-api "Direct link to Sink an Avro topic using Storage Write API") To stream Avro messages directly into BigQuery with lower latency. Connector configuration: ``` { "name": "students_sink_write_api", "connector.class": "com.wepay.kafka.connect.bigquery.BigQuerySinkConnector", "topics": "students", "project": "GOOGLE_CLOUD_PROJECT_NAME", "defaultDataset": ".*=BIGQUERY_DATASET_NAME", "schemaRetriever": "com.wepay.kafka.connect.bigquery.retrieve.IdentitySchemaRetriever", "schemaRegistryLocation": "https://SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD@APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "autoCreateTables": "true", "keySource": "JSON", "keyfile": "GOOGLE_CLOUD_SERVICE_KEY", "useStorageWriteApi": "true", "commitInterval": "1000" } ``` Parameters: * `useStorageWriteApi`: Enables direct streaming into BigQuery * `commitInterval`: Flush interval for Storage Write API batches, in milliseconds ### Enable schema update retries[​](#enable-schema-update-retries "Direct link to Enable schema update retries") Use this configuration when multiple connector tasks or connector instances write to the same BigQuery table and schema evolution is enabled. ``` { "allowNewBigQueryFields": "true", "mediateConcurrentSchemaUpdates": "true", "concurrentSchemaUpdateRetryWaitMs": "10000", "concurrentSchemaUpdateMaxRetries": "3" } ``` Parameters: * `allowNewBigQueryFields`: Allows the connector to add new fields to BigQuery tables * `mediateConcurrentSchemaUpdates`: Enables retry-and-reconcile handling for concurrent schema updates * `concurrentSchemaUpdateRetryWaitMs`: Sets the wait time before each retry, in milliseconds (allowed range `0` to `300000`; default `10000`) * `concurrentSchemaUpdateMaxRetries`: Sets the maximum number of retry attempts ### Track write attempts[​](#track-write-attempts "Direct link to Track write attempts") Use this configuration to add write-attempt metadata to BigQuery rows for downstream deduplication. ``` { "kafkaDataFieldName": "__kafka", "trackPutAttempts": "true" } ``` Parameters: * `kafkaDataFieldName`: Adds Kafka metadata to BigQuery rows under the specified field name * `trackPutAttempts`: Adds a ULID-based `putAttemptId` to the Kafka metadata struct for each `put()` call If the BigQuery table already exists and has a Kafka metadata field, also set `allowNewBigQueryFields` to `true`. --- # Create a sink connector from Apache Kafka® to Google Pub/Sub Lite The [Google Pub/Sub Lite sink connector](https://github.com/googleapis/java-pubsub-group-kafka-connector) enables you to push data from an Aiven for Apache Kafka® topic to a Google Pub/Sub Lite topic. note See the full set of available parameters and configuration options in the [connector's documentation](https://github.com/googleapis/java-pubsub-group-kafka-connector). ## Prerequisites[​](#connect_pubsub_lite_sink_prereq "Direct link to Prerequisites") To setup an Google Pub/Sub Lite sink connector, you need an Aiven for Apache Kafka service [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). Also collect the following information about the target Google Pub/Sub Lite upfront: * `GCP_PROJECT_NAME`: The GCP project name where the target Google Pub/Sub Lite is located * `GCP_TOPIC`: the name of the target Google Pub/Sub Lite topic * `GCP_PUBSUB_LOCATION`: the name of the [Google Pub/Sub Lite location](https://cloud.google.com/pubsub/lite/docs/locations) * `KAFKA_TOPIC`: The name of the target topic in Aiven for Apache Kafka * `APACHE_KAFKA_HOST`: The hostname of the Apache Kafka service, only needed when using Avro as data format * `SCHEMA_REGISTRY_PORT`: The Apache Kafka's schema registry port, only needed when using Avro as data format * `SCHEMA_REGISTRY_USER`: The Apache Kafka's schema registry username, only needed when using Avro as data format * `SCHEMA_REGISTRY_PASSWORD`: The Apache Kafka's schema registry user password, only needed when using Avro as data format note The `SCHEMA_REGISTRY` related parameters are available in the Aiven for Apache Kafka® service page, *Overview* tab, and *Schema Registry* subtab Starting with version 3.0, Aiven for Apache Kafka no longer supports Confluent Schema Registry. For more information, read [the article describing the replacement, Karapace](/docs/products/kafka/karapace.md). ## Setup a Google Pub/Sub Lite sink connector with Aiven Console[​](#setup-a-google-pubsub-lite-sink-connector-with-aiven-console "Direct link to Setup a Google Pub/Sub Lite sink connector with Aiven Console") The following example demonstrates how to setup a Google Pub/Sub Lite sink connector for Apache Kafka using the [Aiven Console](https://console.aiven.io/). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Define the connector configurations in a file (we'll refer to it with the name `pubsub_sink.json`) with the following content: ``` { "name":"CONNECTOR_NAME", "connector.class": "com.google.pubsublite.kafka.sink.PubSubLiteSinkConnector", "topics": "KAFKA_TOPIC", "pubsublite.project": "GCP_PROJECT_NAME", "pubsublite.location": "GCP_PUBSUB_LOCATION", "pubsublite.topic": "GCP_TOPIC", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.sink": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.sink": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } ``` The configuration file contains the following entries: * `name`: the connector name * `topics`: the source Apache Kafka topic names, divided by comma * `pubsublite.project`: the GCP project name where the target Google Pub/Sub is located * `pubsublite.location`: the name of the [Google Pub/Sub Lite location](https://cloud.google.com/pubsub/lite/docs/locations) * `pubsublite.topic`: the name of the target Google Pub/Sub topic * `key.converter` and `value.converter`: define the message data format in the Apache Kafka topic. The `io.confluent.connect.avro.AvroConverter` converter translates messages from the Avro format. To retrieve the message schema we use Aiven's [Karapace schema registry](https://github.com/aiven/karapace), as specified by the `schema.registry.url` parameter and related credentials. note The `key.converter` and `value.converter` sections are only needed when the sink data is in Avro format. If omitted the messages will be read as binary format. When using Avro as sink data format, set following parameters: * `value.converter.schema.registry.url`: pointing to the Aiven for Apache Kafka schema registry URL in the form of `https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT` with the `APACHE_KAFKA_HOST` and `SCHEMA_REGISTRY_PORT` parameters [retrieved in the previous step](/docs/products/kafka/kafka-connect/howto/gcp-pubsub-lite-sink.md#connect_pubsub_lite_sink_prereq). * `value.converter.basic.auth.credentials.sink`: to the value `USER_INFO`, since you're going to login to the schema registry using username and password. * `value.converter.schema.registry.basic.auth.user.info`: passing the required schema registry credentials in the form of `SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD` with the `SCHEMA_REGISTRY_USER` and `SCHEMA_REGISTRY_PASSWORD` parameters [retrieved in the previous step](/docs/products/kafka/kafka-connect/howto/gcp-pubsub-lite-sink.md#connect_pubsub_lite_sink_prereq). The full list of parameters is available in the [dedicated GitHub page](https://github.com/googleapis/java-pubsub-group-kafka-connector/). ### Create a Kafka Connect connector with the Aiven Console[​](#create-a-kafka-connect-connector-with-the-aiven-console "Direct link to Create a Kafka Connect connector with the Aiven Console") To create an Apache Kafka Connect connector: 1. Log in to the [Aiven Console](https://console.aiven.io/) and select the Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service where the connector needs to be defined. 2. Select **Manage stream** > **Connectors** from the left sidebar. 3. Select **Create New Connector**, it is enabled only for services [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md). 4. Select **Google Pub/Sub Lite sink**. 5. In the **Common** tab, locate the **Connector configuration** text box and select on **Edit**. 6. Paste the connector configuration (stored in the `pubsub_sink.json` file) in the form. 7. Select **Apply**. .. note: The Aiven Console parses the configuration file and fills the relevant UI fields. You can review the UI fields across the various tabs and change them if necessary. The changes will be reflected in JSON format in the **Connector configuration** text box. 8. After all the settings are correctly configured, select **Create connector**. 9. Verify the connector status under **Manage stream** > **Connectors**. 10. Verify the presence of the data in the target Pub/Sub Lite topic, the table name is equal to the Apache Kafka topic name. note You can also create connectors using the [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create). ## Example: Create a Google Pub/Sub sink connector[​](#example-create-a-google-pubsub-sink-connector "Direct link to Example: Create a Google Pub/Sub sink connector") You have an Apache Kafka topic `iot_metrics` that you want to push to a Google Pub/Sub Lite topic `iot_metrics_pubsub`, you can create a sink connector with the following configuration, after replacing the placeholders for `GCP_PROJECT_NAME` and `GCP_PUBSUB_LOCATION`: ``` { "name":"CONNECTOR_NAME", "connector.class": "com.google.pubsublite.kafka.sink.PubSubLiteSinkConnector", "topics": "iot_metrics", "pubsublite.project": "GCP_PROJECT_NAME", "pubsublite.location": "GCP_PUBSUB_LOCATION", "pubsublite.topic": "iot_metrics_pubsub", "gcp.credentials.json": "GCP_SERVICE_KEY" } ``` --- # Create a Google Pub/Sub Lite source connector to Apache Kafka® The [Google Pub/Sub Lite source connector](https://github.com/googleapis/java-pubsub-group-kafka-connector/) enables you to push from a Google Pub/Sub subscription to an Aiven for Apache Kafka® topic. note See the full set of available parameters and configuration options in the [connector's documentation](https://github.com/googleapis/java-pubsub-group-kafka-connector/). ## Prerequisites[​](#connect_pubsub_lite_source_prereq "Direct link to Prerequisites") To setup an Google Pub/Sub source connector, you need an Aiven for Apache Kafka service [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). Also collect the following information about the target Google Pub/Sub upfront: * `GCP_PROJECT_NAME`: The GCP project name where the target Google Pub/Sub Lite is located * `GCP_SUBSCRIPTION`: the name of the [Google Pub/Sub Lite subscription](https://cloud.google.com/pubsub/docs/create-subscription) * `GCP_PUBSUB_LOCATION`: the name of the [Google Pub/Sub Lite location](https://cloud.google.com/pubsub/lite/docs/locations) * `KAFKA_TOPIC`: The name of the target topic in Aiven for Apache Kafka * `APACHE_KAFKA_HOST`: The hostname of the Apache Kafka service, only needed when using Avro as data format * `SCHEMA_REGISTRY_PORT`: The Apache Kafka's schema registry port, only needed when using Avro as data format * `SCHEMA_REGISTRY_USER`: The Apache Kafka's schema registry username, only needed when using Avro as data format * `SCHEMA_REGISTRY_PASSWORD`: The Apache Kafka's schema registry user password, only needed when using Avro as data format note The `SCHEMA_REGISTRY` related parameters are available in the Aiven for Apache Kafka® service page, *Overview* tab, and *Schema Registry* subtab Starting with version 3.0, Aiven for Apache Kafka no longer supports Confluent Schema Registry. For more information, read [the article describing the replacement, Karapace](/docs/products/kafka/karapace.md). ## Setup a Google Pub/Sub Lite source connector with Aiven Console[​](#setup-a-google-pubsub-lite-source-connector-with-aiven-console "Direct link to Setup a Google Pub/Sub Lite source connector with Aiven Console") The following example demonstrates how to setup a Google Pub/Sub source connector for Apache Kafka using the [Aiven Console](https://console.aiven.io/). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Define the connector configurations in a file (we'll refer to it with the name `pubsub_lite_source.json`) with the following content: ``` { "name":"CONNECTOR_NAME", "connector.class": "com.google.pubsublite.kafka.source.PubSubLiteSourceConnector", "pubsublite.project": "GCP_PROJECT_NAME", "pubsublite.subscription": "GCP_SUBSCRIPTION", "pubsublite.location": "GCP_PUBSUB_LOCATION", "kafka.topic": "KAFKA_TOPIC", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } ``` The configuration file contains the following entries: * `name`: the connector name * `kafka-topic`: the target Apache Kafka topic name * `pubsublite.project`: the GCP project name where the target Google Pub/Sub is located * `pubsublite.subscription`: the name of the [Google Pub/Sub lite subscription](https://cloud.google.com/pubsub/docs/create-subscription) * `pubsublite.location`: the name of the [Google Pub/Sub Lite location](https://cloud.google.com/pubsub/lite/docs/locations) * `key.converter` and `value.converter`: define the message data format in the Apache Kafka topic. The `io.confluent.connect.avro.AvroConverter` converter translates messages from the Avro format. To retrieve the message schema we use Aiven's [Karapace schema registry](https://github.com/aiven/karapace), as specified by the `schema.registry.url` parameter and related credentials. note The `key.converter` and `value.converter` sections are only needed when the source data is in Avro format. If omitted the messages will be read as binary format. When using Avro as source data format, set following parameters: * `value.converter.schema.registry.url`: pointing to the Aiven for Apache Kafka schema registry URL in the form of `https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT` with the `APACHE_KAFKA_HOST` and `SCHEMA_REGISTRY_PORT` parameters [retrieved in the previous step](/docs/products/kafka/kafka-connect/howto/gcp-pubsub-lite-source.md#connect_pubsub_lite_source_prereq). * `value.converter.basic.auth.credentials.source`: to the value `USER_INFO`, since you're going to login to the schema registry using username and password. * `value.converter.schema.registry.basic.auth.user.info`: passing the required schema registry credentials in the form of `SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD` with the `SCHEMA_REGISTRY_USER` and `SCHEMA_REGISTRY_PASSWORD` parameters [retrieved in the previous step](/docs/products/kafka/kafka-connect/howto/gcp-pubsub-lite-source.md#connect_pubsub_lite_source_prereq). The full list of parameters is available in the [dedicated GitHub page](https://github.com/googleapis/java-pubsub-group-kafka-connector/). ### Create a Kafka Connect connector with the Aiven Console[​](#create-a-kafka-connect-connector-with-the-aiven-console "Direct link to Create a Kafka Connect connector with the Aiven Console") To create a Kafka Connect connector: 1. Log in to the [Aiven Console](https://console.aiven.io/) and select the Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service where the connector needs to be defined. 2. Select **Manage stream** > **Connectors** from the left sidebar. 3. Select **Create New Connector**, it is enabled only for services [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md). 4. Select **Google Pub/Sub source**. 5. In the **Common** tab, locate the **Connector configuration** text box and select on **Edit**. 6. Paste the connector configuration (stored in the `pubsub_lite_source.json` file) in the form. 7. Select **Apply**. note The Aiven Console parses the configuration file and fills the relevant UI fields. You can review the UI fields across the various tabs and change them if necessary. The changes will be reflected in JSON format in the **Connector configuration** text box. 8. After all the settings are correctly configured, select **Create connector**. 9. Verify the connector status under **Manage stream** > **Connectors**. 10. Verify the presence of the data in the target Pub/Sub dataset, the table name is equal to the Apache Kafka topic name. To change the target table name, you can do so using the Kafka Connect `RegexRouter` transformation. note You can also create connectors using the [Aiven CLI command](/docs/tools/cli/service/connector.md). ## Example: Create a Google Pub/Sub source connector[​](#example-create-a-google-pubsub-source-connector "Direct link to Example: Create a Google Pub/Sub source connector") You have a Google Pub/Sub Lite subscription `GCP_SUBSCRIPTION` that you want to push to a Aiven for Apache Kafka topic named `measurements` you can create a source connector with the following configuration, after replacing the placeholders for `GCP_PROJECT_NAME`, `GCP_SERVICE_KEY` and `GCP_PUBSUB_LOCATION`: ``` { "name":"CONNECTOR_NAME", "connector.class": "com.google.pubsub.kafka.source.CloudPubSubSourceConnector", "kafka.topic": "measurements", "cps.project": "GCP_PROJECT_NAME", "cps.subscription": "GCP_SUBSCRIPTION", "gcp.credentials.json": "GCP_SERVICE_KEY" } ``` The Apache Kafka topic format will be the default bytes by default, you can use the AVRO schema by including the `value.converter` and `key.converter` properties defined previously. --- # Create a sink connector from Apache Kafka® to Google Pub/Sub The [Google Pub/Sub sink connector](https://github.com/googleapis/java-pubsub-group-kafka-connector) enables you to push data from an Aiven for Apache Kafka® topic to a Google Pub/Sub topic. note See the full set of available parameters and configuration options in the [connector's documentation](https://github.com/googleapis/java-pubsub-group-kafka-connector). ## Prerequisites[​](#connect_pubsub_sink_prereq "Direct link to Prerequisites") To setup an Google Pub/Sub sink connector, you need an Aiven for Apache Kafka service [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). Also collect the following information about the target Google Pub/Sub upfront: * `GCP_PROJECT_NAME`: The GCP project name where the target Google Pub/Sub is located * `GCP_TOPIC`: the name of the target Google Pub/Sub topic * `GCP_SERVICE_KEY`: A valid GCP service account key for the `GCP_PROJECT_NAME`. To create the project key review the [dedicated document](/docs/products/kafka/kafka-connect/howto/gcp-bigquery-sink.md#configure-google-cloud) warning The GCP Pub/Sub sink connector accepts the `GCP_SERVICE_KEY` JSON service key as a string, therefore all `"` symbols within it must be escaped `\"`. The `GCP_SERVICE_KEY` parameter should be in the format `{\"type\": \"service_account\",\"project_id\": \"XXXXXX\", ...}` Additionally, any `\n` symbols contained in the `private_key` field need to be escaped (by substituting with `\\n`) * `KAFKA_TOPIC`: The name of the target topic in Aiven for Apache Kafka * `APACHE_KAFKA_HOST`: The hostname of the Apache Kafka service, only needed when using Avro as data format * `SCHEMA_REGISTRY_PORT`: The Apache Kafka's schema registry port, only needed when using Avro as data format * `SCHEMA_REGISTRY_USER`: The Apache Kafka's schema registry username, only needed when using Avro as data format * `SCHEMA_REGISTRY_PASSWORD`: The Apache Kafka's schema registry user password, only needed when using Avro as data format note The `SCHEMA_REGISTRY` related parameters are available in the Aiven for Apache Kafka® service page, *Overview* tab, and *Schema Registry* subtab As of version 3.0, Aiven for Apache Kafka no longer supports Confluent Schema Registry. For more information, read [the article describing the replacement, Karapace](https://help.aiven.io/en/articles/5651983) ## Setup a Google Pub/Sub sink connector with Aiven Console[​](#setup-a-google-pubsub-sink-connector-with-aiven-console "Direct link to Setup a Google Pub/Sub sink connector with Aiven Console") The following example demonstrates how to setup a Google Pub/Sub sink connector for Apache Kafka using the [Aiven Console](https://console.aiven.io/). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Define the connector configurations in a file (we'll refer to it with the name `pubsub_sink.json`) with the following content: ``` { "name":"CONNECTOR_NAME", "connector.class": "com.google.pubsub.kafka.sink.CloudPubSubSinkConnector", "topics": "KAFKA_TOPIC", "cps.project": "GCP_PROJECT_NAME", "cps.topic": "GCP_TOPIC", "gcp.credentials.json": "GCP_SERVICE_KEY", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.sink": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.sink": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } ``` The configuration file contains the following entries: * `name`: the connector name * `topics`: the source Apache Kafka topic names, divided by comma * `cps.project`: the GCP project name where the target Google Pub/Sub is located * `cps.topic`: the name of the target Google Pub/Sub topic * `gcp.credentials.json`: contains the GCP service account key, correctly escaped as defined in the [prerequisite phase](/docs/products/kafka/kafka-connect/howto/gcp-pubsub-sink.md#connect_pubsub_sink_prereq) * `key.converter` and `value.converter`: define the message data format in the Apache Kafka topic. The `io.confluent.connect.avro.AvroConverter` converter translates messages from the Avro format. To retrieve the message schema we use Aiven's [Karapace schema registry](https://github.com/aiven/karapace), as specified by the `schema.registry.url` parameter and related credentials. note The `key.converter` and `value.converter` sections are only needed when the sink data is in Avro format. If omitted the messages will be read as binary format. When using Avro as sink data format, set following parameters: * `value.converter.schema.registry.url`: pointing to the Aiven for Apache Kafka schema registry URL in the form of `https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT` with the `APACHE_KAFKA_HOST` and `SCHEMA_REGISTRY_PORT` parameters [retrieved in the previous step](/docs/products/kafka/kafka-connect/howto/gcp-pubsub-sink.md#connect_pubsub_sink_prereq). * `value.converter.basic.auth.credentials.sink`: to the value `USER_INFO`, since you're going to login to the schema registry using username and password. * `value.converter.schema.registry.basic.auth.user.info`: passing the required schema registry credentials in the form of `SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD` with the `SCHEMA_REGISTRY_USER` and `SCHEMA_REGISTRY_PASSWORD` parameters [retrieved in the previous step](/docs/products/kafka/kafka-connect/howto/gcp-pubsub-sink.md#connect_pubsub_sink_prereq). The full list of parameters is available in the [dedicated GitHub page](https://github.com/googleapis/java-pubsub-group-kafka-connector/). ### Create a Kafka Connect connector with the Aiven Console[​](#create-a-kafka-connect-connector-with-the-aiven-console "Direct link to Create a Kafka Connect connector with the Aiven Console") To create an Apache Kafka Connect connector: 1. Log in to the [Aiven Console](https://console.aiven.io/) and select the Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service where the connector needs to be defined. 2. Select **Manage stream** > **Connectors** from the left sidebar. 3. Select **Create New Connector**, it is enabled only for services [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md). 4. Select **Google Pub/Sub sink**. 5. In the **Common** tab, locate the **Connector configuration** text box and select on **Edit**. 6. Paste the connector configuration (stored in the `pubsub_sink.json` file) in the form. 7. Select **Apply**. note The Aiven Console parses the configuration file and fills the relevant UI fields. You can review the UI fields across the various tabs and change them if necessary. The changes will be reflected in JSON format in the **Connector configuration** text box. 8. After all the settings are correctly configured, select **Create connector**. 9. Verify the connector status under **Manage stream** > **Connectors**. 10. Verify the presence of the data in the target Pub/Sub dataset, the table name is equal to the Apache Kafka topic name. note You can also create connectors using the [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create). ## Example: Create a Google Pub/Sub sink connector[​](#example-create-a-google-pubsub-sink-connector "Direct link to Example: Create a Google Pub/Sub sink connector") You have an Apache Kafka topic `iot_metrics` that you want to push to a Google Pub/Sub topic `iot_metrics_pubsub`, you can create a sink connector with the following configuration, after replacing the placeholders for `GCP_PROJECT_NAME` and `GCP_SERVICE_KEY`: ``` { "name":"CONNECTOR_NAME", "connector.class": "com.google.pubsub.kafka.sink.CloudPubSubSinkConnector", "topics": "iot_metrics", "cps.project": "GCP_PROJECT_NAME", "cps.topic": "iot_metrics_pubsub", "gcp.credentials.json": "GCP_SERVICE_KEY" } ``` --- # Create a Google Pub/Sub source connector to Apache Kafka® The [Google Pub/Sub source connector](https://github.com/googleapis/java-pubsub-group-kafka-connector) enables you to push from a Google Pub/Sub subscription to an Aiven for Apache Kafka® topic. note See the complete set of available parameters and configuration options in the [connector's documentation](https://github.com/googleapis/java-pubsub-group-kafka-connector). ## Prerequisites[​](#connect_pubsub_source_prereq "Direct link to Prerequisites") To set up a Google Pub/Sub source connector, ensure you have the following: * An Aiven for Apache Kafka service with [Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster) * Google Cloud Pub/Sub details: * `GCP_PROJECT_NAME`: The name of the Google Cloud project where the Pub/Sub subscription is located * `GCP_SUBSCRIPTION`: The name of the [Google Pub/Sub subscription](https://cloud.google.com/pubsub/docs/create-subscription) * `GCP_SERVICE_KEY`: A valid Google Cloud service account key for the `GCP_PROJECT_NAME`: To create this key, see [Configure GCP for a Google BigQuery sink connector](/docs/products/kafka/kafka-connect/howto/gcp-bigquery-sink.md#configure-google-cloud) warning The GCP Pub/Sub source connector accepts the `GCP_SERVICE_KEY` as a JSON string. Escape all `"` symbols as `\"`, and replace any `\n` symbols in the `private_key` field with `\\n`. Format the `GCP_SERVICE_KEY` like this: `{\"type\": \"service_account\",\"project_id\": \"XXXXXX\", ...}` * IAM Roles: The service account used by the Google Pub/Sub source connector must have the following IAM roles: * `roles/pubsub.subscriber`: Allows the connector to subscribe to the Pub/Sub topic * `roles/pubsub.viewer`: Required to verify the subscription and perform `pubsub.subscriptions.get` operations * Alternatively, use `roles/pubsub.editor`, which includes both the `subscriber` and `viewer` permissions note Without these roles, the connector will fail with a `PERMISSION_DENIED` error during the verification process. * Apache Kafka details: The target topic in Aiven for Apache Kafka, `KAFKA_TOPIC`, where data from the Pub/Sub subscription is written * Schema registry details (if using Avro as the data format): * `APACHE_KAFKA_HOST`: The hostname of your Apache Kafka service * `SCHEMA_REGISTRY_PORT`: The schema registry port for your Apache Kafka service * `SCHEMA_REGISTRY_USER`: The schema registry username for your Apache Kafka service * `SCHEMA_REGISTRY_PASSWORD`: The schema registry password for your Apache Kafka service note To find the `SCHEMA_REGISTRY` details, go to the Aiven for Apache Kafka® **Overview** page, click **Service settings**, and scroll to the **Service management** section. Since Apache Kafka version 3.0, Aiven for Apache Kafka® uses Karapace instead of Confluent Schema Registry. For more information, see [Karapace](/docs/products/kafka/karapace.md). ## Setup a Google Pub/Sub source connector[​](#setup-a-google-pubsub-source-connector "Direct link to Setup a Google Pub/Sub source connector") This example shows how to set up a Google Pub/Sub source connector for Aiven for Apache Kafka using the [Aiven Console](https://console.aiven.io/). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Create a configuration file, `pubsub_source.json`, with the following content: ``` { "name":"CONNECTOR_NAME", "connector.class": "com.google.pubsub.kafka.source.CloudPubSubSourceConnector", "kafka.topic": "KAFKA_TOPIC", "cps.project": "GCP_PROJECT_NAME", "cps.subscription": "GCP_SUBSCRIPTION", "gcp.credentials.json": "GCP_SERVICE_KEY", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } ``` This configuration file includes: * `name`: Connector name * `kafka.topic`: Target Apache Kafka topic name * `cps.project`: GCP project name where the target Google Pub/Sub is located * `cps.subscription`: Name of the [Google Pub/Sub subscription](https://cloud.google.com/pubsub/docs/create-subscription) * `gcp.credentials.json`: GCP service account key, correctly escaped as defined in the [prerequisite](/docs/products/kafka/kafka-connect/howto/gcp-pubsub-source.md#connect_pubsub_source_prereq) * `key.converter` and `value.converter`: Define the message format in the Apache Kafka topic. The `io.confluent.connect.avro.AvroConverter` translates messages from the Avro format. To retrieve the message schema, use Aiven's [Karapace schema registry](https://github.com/aiven/karapace), as specified by the `schema.registry.url` parameter and related credentials note Use the `key.converter` and `value.converter` sections only if the source data is in Avro format. If omitted, messages are read as binary. When using Avro as the source data format, set the following parameters: * `value.converter.schema.registry.url`: Points to the Aiven for Apache Kafka schema registry URL (`https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT`) * `value.converter.basic.auth.credentials.source`: Set to `USER_INFO` to log in to the schema registry using a username and password * `value.converter.schema.registry.basic.auth.user.info`: Pass the schema registry credentials in the form `SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD` For a full list of parameters, see the [java-pubsub-group-kafka-connector GitHub repository](https://github.com/googleapis/java-pubsub-group-kafka-connector/). ### Create a Google Pub/Sub source connector[​](#create-a-google-pubsub-source-connector "Direct link to Create a Google Pub/Sub source connector") * Aiven Console * Aiven CLI 1. Log in to the [Aiven Console](https://console.aiven.io/) and select the Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service. 2. Click **Manage stream** > **Connectors** from the left sidebar. 3. Click **Create New Connector**. If connectors are not enabled, click **Enable connector on this service**. 4. Click **See all connectors** to view all available options. 5. Select **Google Pub/Sub source** and click **Get started**. 6. In the **Common** tab, find the **Connector configuration** text box and click **Edit**. 7. Paste the `pubsub_source.json` connector configuration into the form. 8. Click **Apply**. note The Aiven Console parses the configuration file and fills in the relevant UI fields. Review and modify fields as needed. Changes are reflected in the **Connector configuration** text box in JSON format. 9. After verifying the settings, click **Create connector**. 10. Check the connector status on **Manage stream** > **Connectors**. 11. Verify that data appears in the target Pub/Sub dataset. By default, the table name is derived from the Apache Kafka topic name. To change the table name, use the Kafka Connect `RegexRouter` transformation. This transformation allows you to rename topics using regular expressions, ensuring that the data is routed to the desired table. note You can also create connectors using the [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create). To create a Google Pub/Sub source connector using the [Aiven CLI](/docs/tools/cli/service-cli.md), run: ``` avn service connector create SERVICE_NAME @pubsub_source.json ``` Parameters: * `SERVICE_NAME`: Your Aiven for Apache Kafka service name. * `@pubsub_source.json`: The path to your JSON configuration file. To check the connector status, run: ``` avn service connector status SERVICE_NAME CONNECTOR_NAME ``` Verify that topic and data are present in the Apache Kafka target instance. ## Example: Create a Google Pub/Sub source connector[​](#example-create-a-google-pubsub-source-connector "Direct link to Example: Create a Google Pub/Sub source connector") Use the following configuration to push data from a Google Pub/Sub subscription (`GCP_SUBSCRIPTION`) to an Aiven for Apache Kafka topic named `measurements`. Replace the placeholders `GCP_PROJECT_NAME` and `GCP_SERVICE_KEY` with the appropriate values: ``` { "name":"CONNECTOR_NAME", "connector.class": "com.google.pubsub.kafka.source.CloudPubSubSourceConnector", "kafka.topic": "measurements", "cps.project": "GCP_PROJECT_NAME", "cps.subscription": "GCP_SUBSCRIPTION", "gcp.credentials.json": "GCP_SERVICE_KEY" } ``` The default Apache Kafka topic format is bytes. To use the Avro schema, include the `value.converter` and `key.converter` properties defined earlier. --- # Create a Google Cloud Storage sink connector for Apache Kafka® The Google Cloud Storage (GCS) sink connector moves data from Aiven for Apache Kafka® topics to a Google Cloud Storage bucket for long-term storage. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven for Apache Kafka® service with [Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md), or a [dedicated Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster) * Access to a Google Cloud project where you can create the following: * A Google Cloud Storage bucket * A Google service account with a JSON service key * The following values for the connector configuration: * `GCS_NAME`: The name of the target Google Cloud Storage bucket * `GCS_CREDENTIALS`: The Google service account JSON key note For a full list of configuration options, see the [Google Cloud Storage sink connector documentation](https://github.com/aiven/gcs-connector-for-apache-kafka). warning The connector expects `GCS_CREDENTIALS` as a single JSON string. Escape all `"` symbols as `\"`. Example: `{\"type\":\"service_account\",\"project_id\":\"XXXXXX\",...}` If the `private_key` field contains `\n`, escape it as `\\n`. ### Google credential source restrictions[​](#google-credential-source-restrictions "Direct link to Google credential source restrictions") When using Google Cloud external account credentials with Aiven for Apache Kafka® Connect, Aiven applies security restrictions to prevent unauthorized file access and network requests. If the Google Cloud credential JSON file includes a `credential_source` object, the following restrictions apply: * `credential_source.file`: Not allowed * `credential_source.executable.command`: Not allowed * `credential_source.url`: Allowed only when allow-listed in the service configuration (`gcp_auth_allowed_urls`) Example `credential_source` object using a URL-based credential: ``` { "credential_source": { "url": "https://sts.googleapis.com/v1/token", "headers": { "Metadata-Flavor": "Google" } } } ``` important `gcp_auth_allowed_urls` is a **Kafka Connect service-level configuration**, not a connector configuration. Configure it in the Kafka Connect service settings. This setting applies to all connectors in the service. #### Configure allowed authentication URLs[​](#configure-allowed-authentication-urls "Direct link to Configure allowed authentication URLs") To use URL-based credentials (`credential_source.url`), configure the following: 1. Configure the Kafka Connect service with allowed authentication URLs. 2. Configure each connector to use one of the allowed URLs. Set `gcp_auth_allowed_urls` on the Kafka Connect service to define which HTTPS authentication endpoints the service can access. This setting applies to all connectors in the service. * Console * CLI 1. Go to the [Aiven Console](https://console.aiven.io/). 2. Select the Aiven for Apache Kafka Connect service. 3. Click **Service settings**. 4. In **Advanced configuration**, click **Configure**. 5. Set **`gcp_auth_allowed_urls`** to the required HTTPS endpoints. 6. Click **Save configuration**. Run the following command: ``` avn service update KAFKA_CONNECT_SERVICE_NAME \ -c gcp_auth_allowed_urls='["https://sts.googleapis.com","https://iamcredentials.googleapis.com"]' ``` If multiple connectors use URL-based credentials, add all required authentication URLs to `gcp_auth_allowed_urls`. Each unique URL needs to be added only once. If `credential_source.url` is set but the URL is not included in `gcp_auth_allowed_urls`, connector creation fails. In the connector configuration JSON, set `credential_source.url` to match one of the URLs configured in the service. ## Configure Google Cloud for the connector[​](#configure-google-cloud-for-the-connector "Direct link to Configure Google Cloud for the connector") Create a Google Cloud Storage bucket and a Google service account key that the connector can use to write objects. ### Create a Google Cloud Storage bucket[​](#gcs-sink-connector-google-bucket "Direct link to Create a Google Cloud Storage bucket") 1. In the [Google Cloud console](https://console.cloud.google.com/), open **Cloud Storage**. 2. Create a bucket using the [Cloud Storage buckets page](https://console.cloud.google.com/storage/). 3. Specify the bucket name and location. 4. Keep the other settings as default unless your organization requires otherwise. ### Create a Google service account and JSON key[​](#gcs-sink-connector-google-account "Direct link to Create a Google service account and JSON key") 1. Create a Google service account and JSON service key by following [Google authentication instructions](https://cloud.google.com/docs/authentication/client-libraries). 2. Download the JSON service key. You use this key in the connector configuration as `GCS_CREDENTIALS`. ### Grant the service account access to the bucket[​](#gcs-sink-connector-grant-permissions "Direct link to Grant the service account access to the bucket") 1. Open the bucket in the Cloud Storage console. 2. Go to the **Permissions** tab. 3. Grant access to the service account. Grant the following permissions: * `storage.objects.create` * `storage.objects.delete`, required for overwriting during reprocessing Grant these permissions with a custom role or the standard role **Storage Legacy Bucket Writer**. Also ensure the bucket does not have a retention policy that prevents overwriting. ## Naming and data formats[​](#naming-and-data-formats "Direct link to Naming and data formats") ### Filename format[​](#filename-format "Direct link to Filename format") The connector uses the following format for output files: ``` --[.gz] ``` The filename format has the following parts: * ``: Filename prefix, useful for defining subdirectories in the storage bucket * ``: Source Apache Kafka® topic name * ``: Source Apache Kafka® topic partition number * ``: Offset of the first record in the file * `[.gz]`: File suffix added when you enable compression, depending on compression type ### Data format[​](#data-format "Direct link to Data format") Output files are text files with one record per line, separated by `\n`. Two data formats are available: * **Flat structure**: Default format. Commas separate field values in CSV format. Set `format.output.type` to `csv`. * **Complex structure**: JSON Lines format. Each line is a valid JSON object. Set `format.output.type` to `jsonl`. ## Create the connector configuration[​](#create-the-connector-configuration "Direct link to Create the connector configuration") Create a JSON configuration file, for example `gcs_sink.json`: ``` { "name": "my-gcs-connector", "connector.class": "io.aiven.kafka.connect.gcs.GcsSinkConnector", "tasks.max": "1", "topics": "TOPIC_NAME", "key.converter": "org.apache.kafka.connect.storage.StringConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "gcs.credentials.json": "GCS_CREDENTIALS", "gcs.bucket.name": "GCS_NAME", "file.name.prefix": "my-custom-prefix/", "file.compression.type": "gzip", "format.output.type": "jsonl", "format.output.fields": "value,offset" } ``` Parameters: * `name`: The connector name * `topics`: Comma-separated list of Apache Kafka® topics to sink to the bucket * `key.converter` and `value.converter`: Message converters based on your topic format * `gcs.credentials.json`: The Google service account JSON key as a JSON string * `gcs.bucket.name`: The name of the target bucket * `file.name.prefix`: Prefix for files created in the bucket * `file.compression.type`: Compression type for output files * `format.output.type`: Output file format * `format.output.fields`: Message fields to include in output files ## Create a Google Cloud Storage sink connector[​](#create-a-google-cloud-storage-sink-connector "Direct link to Create a Google Cloud Storage sink connector") * Aiven Console * Aiven CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Open your Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service. 3. In the sidebar, click **Manage stream** > **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is already enabled on the service. If not, click **Enable connector on this service**. To enable connectors: 1. In the sidebar, click **Service settings**. 2. In the **Service management** section, click **Actions** > **Enable Kafka Connect**. 5. In the list of sink connectors, click **Get started** under **Google Cloud Storage sink**. 6. On the connector page, open the **Common** tab. 7. In **Connector configuration**, click **Edit**. 8. Paste the configuration from your `gcs_sink.json` file into the text box. Replace placeholders with your actual values. 9. Click **Apply**. note When you paste the JSON configuration, Aiven Console parses it and populates the corresponding fields in the UI. Changes you make in the UI appear in the **Connector configuration** JSON. 10. Click **Create connector**. 11. Verify the connector status on the **Manage stream** > **Connectors** page. 12. Confirm that data from the Apache Kafka topics appears in the target bucket. To create a GCS sink connector using the [Aiven CLI](/docs/tools/cli/service-cli.md), run: ``` avn service connector create SERVICE_NAME @gcs_sink.json ``` Parameters: * `SERVICE_NAME`: The name of your Aiven for Apache Kafka® service * `@gcs_sink.json`: The path to your connector configuration file ## Examples[​](#examples "Direct link to Examples") ### Create a GCS sink connector for a JSON topic[​](#create-a-gcs-sink-connector-for-a-json-topic "Direct link to Create a GCS sink connector for a JSON topic") This example creates a connector with the following settings: * Connector name: `my_gcs_sink` * Source topic: `test` * Bucket name: `my-test-bucket` * Name prefix: `my-custom-prefix/` * Compression: `gzip` * Output format: `jsonl` * Output fields: `value,offset` * Maximum records per file: 1 ``` { "name": "my_gcs_sink", "connector.class": "io.aiven.kafka.connect.gcs.GcsSinkConnector", "topics": "test", "key.converter": "org.apache.kafka.connect.storage.StringConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "gcs.credentials.json": "{\"type\": \"service_account\", \"project_id\": \"XXXXXXXXX\", ...}", "gcs.bucket.name": "my-test-bucket", "file.name.prefix": "my-custom-prefix/", "file.compression.type": "gzip", "file.max.records": "1", "format.output.type": "jsonl", "format.output.fields": "value,offset" } ``` --- # Create a sink connector from Apache Kafka® via HTTP The HTTP sink connector enables you to move data from an Aiven for Apache Kafka® cluster to a remote server via HTTP. The full list of parameters and setup details is available in the [dedicated GitHub repository](https://github.com/aiven/http-connector-for-apache-kafka/). note See the full set of available parameters and configuration options in the [connector's documentation](https://github.com/aiven/aiven-kafka-connect-http). ## Prerequisites[​](#connect_http_sink_prereq "Direct link to Prerequisites") To setup an HTTP sink connector, you need an Aiven for Apache Kafka service [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). Also collect the following information about the target server: * `SERVER_URL`: The remote server URL that will be called via POST method * `SERVER_AUTHORIZATION_TYPE`: The HTTP authorization type, supported types are `none`, `oauth2` and `static` * `TOPIC_LIST`: The list of topics to sink divided by comma and, if you are using Avro as the data format: * `APACHE_KAFKA_HOST`: The hostname of the Apache Kafka service * `SCHEMA_REGISTRY_PORT`: The Apache Kafka's schema registry port * `SCHEMA_REGISTRY_USER`: The Apache Kafka's schema registry username * `SCHEMA_REGISTRY_PASSWORD`: The Apache Kafka's schema registry user password note You can browse the additional parameters available for the `static` and `oauth2` authorization types in the [dedicated documentation](https://github.com/aiven/http-connector-for-apache-kafka/blob/main/docs/sink-connector-config-options.rst). ## Setup an HTTP sink connector with Aiven Console[​](#setup-an-http-sink-connector-with-aiven-console "Direct link to Setup an HTTP sink connector with Aiven Console") The following example demonstrates how to setup an HTTP sink connector for Apache Kafka using the [Aiven Console](https://console.aiven.io/). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Define the connector configurations in a file (we'll refer to it with the name `http_sink.json`) with the following content: ``` { "name":"CONNECTOR_NAME", "connector.class": "io.aiven.kafka.connect.http.HttpSinkConnector", "topics": "TOPIC_LIST", "http.url": "SERVER_URL", "http.authorization.type": "SERVER_AUTHORIZATION_TYPE", "key.converter": "org.apache.kafka.connect.storage.StringConverter", "value.converter": "org.apache.kafka.connect.storage.StringConverter" } ``` The configuration file contains the following entries: * `name`: the connector name * `http.url` and `http.authorization.type`: remote server URL and authorization parameters collected in the [prerequisite](/docs/products/kafka/kafka-connect/howto/http-sink.md#connect_http_sink_prereq) phase. * `key.converter` and `value.converter`: defines the message data format in the Apache Kafka topic. The `io.confluent.connect.avro.AvroConverter` converter translates messages from the Avro format. To retrieve the message schema we use Aiven's [Karapace schema registry](https://github.com/aiven/karapace) as specified by the `schema.registry.url` parameter and related credentials. note The `key.converter` and `value.converter` sections define how the topic messages will be parsed and need to be included in the connector configuration. When using Avro as source data format, set the following parameters: * `value.converter.schema.registry.url`: pointing to the Aiven for Apache Kafka schema registry URL in the form of `https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT` with the `APACHE_KAFKA_HOST` and `SCHEMA_REGISTRY_PORT` parameters [retrieved in the previous step](/docs/products/kafka/kafka-connect/howto/http-sink.md#connect_http_sink_prereq). * `value.converter.basic.auth.credentials.source`: to the value `USER_INFO`, since you're going to login to the schema registry using username and password. * `value.converter.schema.registry.basic.auth.user.info`: passing the required schema registry credentials in the form of `SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD` with the `SCHEMA_REGISTRY_USER` and `SCHEMA_REGISTRY_PASSWORD` parameters [retrieved in the previous step](/docs/products/kafka/kafka-connect/howto/http-sink.md#connect_http_sink_prereq). ### Create a Kafka Connect connector with the Aiven Console[​](#create-a-kafka-connect-connector-with-the-aiven-console "Direct link to Create a Kafka Connect connector with the Aiven Console") To create a Kafka Connect connector: 1. Log in to the [Aiven Console](https://console.aiven.io/) and select the Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service where the connector needs to be defined. 2. Select **Manage stream** > **Connectors** from the left sidebar. 3. Select **Create New Connector**, it is enabled only for services [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md). 4. Select **HTTP sink**. 5. In the **Common** tab, locate the **Connector configuration** text box and select on **Edit**. 6. Paste the connector configuration (stored in the `http_sink.json` file) in the form. 7. Select **Apply**. note The Aiven Console parses the configuration file and fills the relevant UI fields. You can review the UI fields across the various tabs and change them if necessary. The changes will be reflected in JSON format in the **Connector configuration** text box. 8. After all the settings are correctly configured, select **Create connector**. 9. Verify the connector status under **Manage stream** > **Connectors**. 10. Verify the flow of HTTP POST calls in the target server. note You can also create connectors using the [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create). ## Example: Create an HTTP sink connector with a server having no authorization[​](#example-create-an-http-sink-connector-with-a-server-having-no-authorization "Direct link to Example: Create an HTTP sink connector with a server having no authorization") If you have a topic named `iot_measurements` containing the following data in JSON format: ``` Key: 1 Value: {"iot_id":1, "metric":"Temperature", "measurement":14} Key: 2 Value: {"iot_id":2, "metric":"Humidity", "measurement":60} Key: 1 Value: {"iot_id":1, "metric":"Temperature", "measurement":16} ``` You can sink the `iot_measurements` topic to a remote server over HTTP with the following connector configuration, after replacing the placeholders for `SERVER_URL`, and `SERVER_AUTHORIZATION_TYPE`: ``` { "name":"iot_measurements_sink", "connector.class": "io.aiven.kafka.connect.http.HttpSinkConnector", "topics": "iot_measurements", "http.url": "SERVER_URL", "http.authorization.type": "SERVER_AUTHORIZATION_TYPE", "key.converter": "org.apache.kafka.connect.storage.StringConverter", "value.converter": "org.apache.kafka.connect.storage.StringConverter" } ``` The configuration file contains the following things to note: * `"topics": "iot_measurements"`: setting the topic to sink * `"value.converter": "org.apache.kafka.connect.json.StringConverter"`: the message value and key are in plain JSON format without a schema, therefore we can just pass them as plain string via HTTP --- # Create an IBM MQ sink connector in Aiven for Apache Kafka® The IBM MQ sink connector allows you to route messages from Apache Kafka® topics to IBM MQ queues. ## Prerequisites[​](#connect_ibm_mq_sink_prereq "Direct link to Prerequisites") * An [Aiven for Apache Kafka service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with Apache Kafka Connect enabled, or * A [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * Access to an IBM MQ instance. ### Required IBM MQ details[​](#required-ibm-mq-details "Direct link to Required IBM MQ details") Gather the following details about your IBM MQ setup: note You can view the full set of available parameters and configuration options in the [IBM MQ sink connector GitHub repository](https://github.com/ibm-messaging/kafka-connect-mq-sink). * `mq.queue.manager`: The name of the IBM MQ queue manager. * `mq.connection.mode`: The connection mode, typically set to `client`. * `mq.connection.name.list`: The hostname and port of the IBM MQ instance, in the format `host(port)`. * `mq.channel.name`: The name of the server-connection channel in IBM MQ. * `mq.queue`: The name of the IBM MQ target queue. * `mq.user.name`: The username for IBM MQ authentication. * `mq.password`: The password associated with the `mq.user.name`. * `mq.ssl.cipher.suite`: The cipher suite name for TLS (SSL) connections. * `mq.ssl.use.ibm.cipher.mappings`: A setting that controls whether to apply IBM-specific cipher mappings. Typically set to `false` to ensure compatibility with standard TLS configurations. ## Create an IBM MQ sink connector configuration file[​](#create-an-ibm-mq-sink-connector-configuration-file "Direct link to Create an IBM MQ sink connector configuration file") Create a file named `ibm_mq_sink_connector.json` with the following configurations. This file is optional but helps you organize your settings, making it easier to copy and paste the configurations into the [Aiven Console](https://console.aiven.io/) later. ``` { "name": "IBM_MQ_Sink_Connector", "connector.class": "com.ibm.eventstreams.connect.mqsink.MQSinkConnector", "tasks.max": "2", "topics": "kafka_topic_name", "mq.queue.manager": "QueueManagerName", "mq.connection.mode": "client", "mq.connection.name.list": "mq-host(1414)", "mq.channel.name": "CLOUD.APP.SVRCONN", "mq.queue": "TargetQueueName", "mq.user.name": "YourUsername", "mq.password": "YourPassword", "mq.ssl.cipher.suite": "TLS_AES_256_GCM_SHA384", "mq.ssl.use.ibm.cipher.mappings": "false", "mq.message.builder": "com.ibm.eventstreams.connect.mqsink.builders.DefaultMessageBuilder", "mq.message.body.jms": true, "config.action.reload": "restart" } ``` Parameters: * `name`: The connector name. Replace `IBM_MQ_Sink_Connector` with the desired connector name. * `topics`: The Apache Kafka topics from which messages will be sent to IBM MQ. * `mq.queue.manager`, `mq.connection.mode`, `mq.connection.name.list`, `mq.channel.name`, `mq.queue`, `mq.user.name`, `mq.password`: IBM MQ connection details collected in the [prerequisite](#connect_ibm_mq_sink_prereq) phase. * `mq.ssl.cipher.suite`: The cipher suite for the SSL connection. * `mq.ssl.use.ibm.cipher.mappings`: Controls the use of IBM cipher mappings, typically set to `false`. * `mq.message.builder`: Defines how Apache Kafka messages are converted to MQ messages. * `config.action.reload`: Defines the action to take when the configuration changes, typically set to `restart`. ## Create an IBM MQ sink connector[​](#create-an-ibm-mq-sink-connector "Direct link to Create an IBM MQ sink connector") * Aiven Console * Aiven CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is already enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the Service management section, click **Actions** > **Enable Kafka connect**. 5. In the sink connectors, click **Get started** under **IBM MQ Sink**. 6. On the **IBM MQ Sink** connector page, go to the **Common** tab. 7. Locate the **Connector configuration** text box and click **Edit**. 8. Paste the configuration from your `ibm_mq_sink_connector.json` file into the text box. 9. Click **Create connector**. 10. Verify the connector status on the **Manage stream** > **Connectors** page. Ensure that data from the Apache Kafka topics is successfully transferred to the target IBM MQ queue. To create an IBM MQ Sink connector using the [Aiven CLI](/docs/tools/cli/service-cli.md), run the following command: ``` avn service connector create SERVICE_NAME @ibm_mq_sink_connector.json ``` Parameters: * `SERVICE_NAME`: The name of your Aiven for Apache Kafka service. * `@ibm_mq_sink_connector.json`: This denotes the path to your JSON configuration file. ## Example: Define and create an IBM MQ sink connector[​](#example-define-and-create-an-ibm-mq-sink-connector "Direct link to Example: Define and create an IBM MQ sink connector") This example demonstrates how to create an IBM MQ sink connector using the following properties: * Connector name: `ibm_mq_sink` * Queue manager: `kafka_connect_connector_testing_queue_manager` * Connection mode: `client` * Connection name list: `kafka-connect-connector-testing-queue-manager-95e0.qm2.eu-de.mq.appdomain.cloud(32275)` * Channel name: `CLOUD.APP.SVRCONN` * Target queue: `test_topic_mq_1` * Username: `testuser` * Password: `Test123!` * SSL cipher suite: `TLS_AES_256_GCM_SHA384` * IBM cipher mappings: `false` * Topics to sink: `test-topic-1` **Connector configuration example:** ``` { "name": "ibm_mq_sink", "connector.class": "com.ibm.eventstreams.connect.mqsink.MQSinkConnector", "tasks.max": "2", "topics": "test-topic-1", "mq.queue.manager": "kafka_connect_connector_testing_queue_manager", "mq.connection.mode": "client", "mq.connection.name.list": "kafka-connect-connector-testing-queue-manager-95e0.qm2.eu-de.mq.appdomain.cloud(32275)", "mq.channel.name": "CLOUD.APP.SVRCONN", "mq.queue": "test_topic_mq_1", "mq.user.name": "testuser", "mq.password": "Test123!", "mq.ssl.cipher.suite": "TLS_AES_256_GCM_SHA384", "mq.ssl.use.ibm.cipher.mappings": "false", "mq.message.builder": "com.ibm.eventstreams.connect.mqsink.builders.DefaultMessageBuilder", "config.action.reload": "restart" } ``` Once you've saved this configuration in the `ibm_mq_sink_connector.json` file, you can create the connector in your Aiven for Apache Kafka service. You can observe data from the `test-topic-1` Apache Kafka topic being successfully delivered to the IBM MQ queue named `test_topic_mq_1`. --- # Create an Iceberg sink connector for Aiven for Apache Kafka® Use the Iceberg sink connector to write real-time Apache Kafka® data to Iceberg tables for analytics and long-term storage. The connector supports exactly-once delivery, schema evolution, and metadata management. It is optimized for high-throughput, large-scale processing. For more information, see the [official Iceberg sink connector documentation](https://iceberg.apache.org/docs/latest/kafka-connect/#apache-iceberg-sink-connector). ## Catalogs in Iceberg[​](#catalogs-in-iceberg "Direct link to Catalogs in Iceberg") In Apache Iceberg, a catalog stores table metadata and supports key operations such as creating, renaming, and deleting tables. It manages collections of tables organized into namespaces and provides the metadata needed for access. The Iceberg sink connector writes data to a storage backend. The catalog manages metadata so that multiple systems can read and write to the same tables. The connector supports the following catalog types: * [AWS Glue REST catalog](/docs/products/kafka/kafka-connect/howto/aws-glue-rest-catalog.md) * [AWS Glue catalog](/docs/products/kafka/kafka-connect/howto/aws-glue-catalog.md) * [PostgreSQL JDBC catalog](/docs/products/kafka/kafka-connect/howto/jdbc-catalog-postgres.md) * [Snowflake Open Catalog](/docs/products/kafka/kafka-connect/howto/snowflake-open-catalog.md) note The AWS Glue REST catalog does not support automatic table creation. You must manually create tables in AWS Glue and ensure the schema matches the Apache Kafka data. For more details, see the [Iceberg catalogs documentation](https://iceberg.apache.org/terms/#catalog/). ## File I/O and write format[​](#file-io-and-write-format "Direct link to File I/O and write format") The Iceberg sink connector supports the following settings: * **File I/O**: Supports `S3FileIO` for AWS S3 storage. * **Write format**: Supports the Parquet format. ## Future enhancements[​](#future-enhancements "Direct link to Future enhancements") Future updates to the Iceberg sink connector include: * **FileIO implementations:** Support for GCS and Azure FileIO. * **Write formats:** Additional support for Avro and ORC formats. * **Catalogs:** Planned support for Hive and Amazon S3 Tables. Related pages * [Official Apache Iceberg sink connector docs](https://iceberg.apache.org/docs/latest/kafka-connect/) * [Configure secret providers for Apache Kafka Connect](/docs/products/kafka/kafka-connect/howto/configure-secret-providers.md) --- # Create a sink connector from Apache Kafka® to InfluxDB® The InfluxDB® Stream Reactor sink connector enables you to move data from an Aiven for Apache Kafka® cluster to an InfluxDB® instance. It uses [KCQL transformations](https://docs.lenses.io/connectors/sink/influx) to filter and map topic data before writing it to InfluxDB. caution **Version compatibility** Stream Reactor version 9.0.2 includes class and package name updates introduced in version 6.0.0 by Lenses to standardize connector and converter names. Version 9.x is not compatible with version 4.2.0. To continue using version 4.2.0, [set the connector version](/docs/products/kafka/kafka-connect/howto/manage-connector-versions.md#set-version) before you upgrade. If you upgrade from version 4.2.0, recreate the connector using the updated class name. For example: ``` "connector.class": "io.lenses.streamreactor.connect.." ``` For details about these changes, see the [Stream Reactor release notes](https://docs.lenses.io/stream-reactor/docs/releases/). ## Prerequisites[​](#connect_influx_lenses_sink_prereq "Direct link to Prerequisites") * An Aiven for Apache Kafka service [with Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * Gather the following information for the target InfluxDB database: * `INFLUXDB_HOST`: The InfluxDB hostname. * `INFLUXDB_PORT`: The InfluxDB port. * `INFLUXDB_DATABASE_NAME`: The InfluxDB database name. * `INFLUXDB_USERNAME`: The InfluxDB username. * `INFLUXDB_PASSWORD`: The InfluxDB password. * `TOPIC_LIST`: A comma-separated list of Kafka topics to sink. * `KCQL_TRANSFORMATION`: A KCQL statement to map topic fields to InfluxDB measurements. Use the following format: ``` INSERT INTO MEASUREMENT_NAME SELECT LIST_OF_FIELDS FROM APACHE_KAFKA_TOPIC ``` * `APACHE_KAFKA_HOST`: The Apache Kafka host. Required only when using Avro as the data format. * `SCHEMA_REGISTRY_PORT`: The schema registry port. Required only when using Avro. * `SCHEMA_REGISTRY_USER`: The schema registry username. Required only when using Avro. * `SCHEMA_REGISTRY_PASSWORD`: The schema registry password. Required only when using Avro. note If you are using Aiven for InfluxDB and Aiven for Apache Kafka, get all required connection details, including schema registry information, from the **Connection information** section on the **Overview** page. As of version 3.0, Aiven for Apache Kafka uses Karapace as the schema registry and no longer supports the Confluent Schema Registry. For a complete list of supported parameters and configuration options, see the [connector's documentation](https://docs.lenses.io/latest/connectors/kafka-connectors/sinks/influxdb). ## Create a connector configuration file[​](#create-a-connector-configuration-file "Direct link to Create a connector configuration file") Create a file named `influxdb_sink.json` and add the following configuration: ``` { "name": "CONNECTOR_NAME", "connector.class": "com.datamountaineer.streamreactor.connect.influx.InfluxSinkConnector", "topics": "TOPIC_LIST", "connect.influx.url": "https://INFLUXDB_HOST:INFLUXDB_PORT", "connect.influx.db": "INFLUXDB_DATABASE_NAME", "connect.influx.username": "INFLUXDB_USERNAME", "connect.influx.password": "INFLUXDB_PASSWORD", "connect.influx.kcql": "KCQL_TRANSFORMATION", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } ``` Parameters: * `name`: The connector name. Replace `CONNECTOR_NAME` with your desired name. * `connect.influx.*`: InfluxDB connection parameters collected in the [prerequisite step](/docs/products/kafka/kafka-connect/howto/influx-sink.md#connect_influx_lenses_sink_prereq). * `topics`: A comma-separated list of Kafka topics to sink. * `key.converter` and `value.converter`: Define the message data format in the Kafka topic. This example uses `io.confluent.connect.avro.AvroConverter` to translate messages in Avro format. The schema is retrieved from Aiven's [Karapace schema registry](https://github.com/aiven/karapace) using the `schema.registry.url` and related credentials. note The `key.converter` and `value.converter` fields define how Kafka messages are parsed. Include these fields in the connector configuration. If you use Avro as the message format, set the following parameters: * `value.converter.schema.registry.url`: The Aiven for Apache Kafka schema registry URL, in the format `https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT`. * `value.converter.basic.auth.credentials.source`: Set to `USER_INFO` to enable authentication with a username and password. * `value.converter.schema.registry.basic.auth.user.info`: The schema registry credentials, in the format `SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD`. You can get these values from the [prerequisite step](/docs/products/kafka/kafka-connect/howto/influx-sink.md#connect_influx_lenses_sink_prereq). ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Console * CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 5. In the sink connectors list, select **Stream Reactor InfluxDB Sink**, and click **Get started**. 6. On the **Stream Reactor InfluxDB Sink** page, go to the **Common** tab. 7. Locate the **Connector configuration** text box and click **Edit**. 8. Paste the configuration from your `influxdb_sink.json` file into the text box. 9. Click **Create connector**. 10. Verify the connector status on the **Manage stream** > **Connectors** page. 11. Confirm that data is written to the InfluxDB target database. To create the connector using the [Aiven CLI](/docs/tools/cli/service/connector.md#avn_service_connector_create), run: ``` avn service connector create SERVICE_NAME @influxdb_sink.json ``` Replace: * `SERVICE_NAME`: Your Kafka or Kafka Connect service name. * `@influxdb_sink.json`: Path to your configuration file. ## Sink topic data to InfluxDB[​](#sink-topic-data-to-influxdb "Direct link to Sink topic data to InfluxDB") The following example shows how to sink data from a Kafka topic to an InfluxDB measurement. If your Kafka topic `measurements` contains the following data: ``` { "ts": "2022-10-24T13:09:43.406000Z", "device_name": "mydevice1", "measurement": 17 } ``` To write this data to InfluxDB, use the following connector configuration: ``` { "name": "my-influxdb-sink", "connector.class": "com.datamountaineer.streamreactor.connect.influx.InfluxSinkConnector", "topics": "measurements", "connect.influx.url": "https://INFLUXDB_HOST:INFLUXDB_PORT", "connect.influx.db": "INFLUXDB_DATABASE_NAME", "connect.influx.username": "INFLUXDB_USERNAME", "connect.influx.password": "INFLUXDB_PASSWORD", "connect.influx.kcql": "INSERT INTO measurements SELECT ts, device_name, measurement FROM measurements", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter.schemas.enable": "false" } ``` Replace all placeholder values (such as `INFLUXDB_HOST`, `INFLUXDB_PORT`, and `INFLUXDB_PASSWORD`) with your actual InfluxDB connection details. This configuration does the following: * `"topics": "measurements"`: Specifies the Kafka topic to sink. * Connection settings (`connect.influx.*`): Provide the InfluxDB connection details. * `"value.converter"` and `"value.converter.schemas.enable"`: Set the message format. The topic uses raw JSON without a schema. * `"connect.influx.kcql"`: Defines the insert logic. Each Kafka message is written as a row in the `measurements` table. After creating the connector, verify the presence of the data in the target InfluxDB database. The measurement name matches the Kafka topic (`measurements`). --- # Configure the Iceberg sink connector with a PostgreSQL JDBC catalog The JDBC catalog stores Iceberg table metadata in a PostgreSQL® database. Use it to integrate Apache Kafka®, Amazon S3, and a PostgreSQL-based metadata store. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Aiven for Apache Kafka® service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with Apache Kafka Connect enabled, or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster) * Apache Kafka topic that contains the source data * Apache Kafka control topic used to coordinate Iceberg table commits (for example, `control-iceberg`) * [Aiven for PostgreSQL® service](/docs/products/postgresql/get-started.md) or another PostgreSQL instance that is publicly accessible * Amazon S3 bucket for storing Iceberg table data * IAM user or role with permissions to read and write to the S3 bucket * AWS access key ID and secret access key for the IAM credentials ## Create an Iceberg sink connector configuration[​](#create-an-iceberg-sink-connector-configuration "Direct link to Create an Iceberg sink connector configuration") To configure the Iceberg sink connector with the JDBC catalog, define a JSON file that uses PostgreSQL for metadata and Amazon S3 for storage. note Loading worker properties is not supported. Use `iceberg.kafka.*` properties instead. ``` { "name": "CONNECTOR_NAME", "connector.class": "org.apache.iceberg.connect.IcebergSinkConnector", "tasks.max": "2", "topics": "KAFKA_TOPICS", "key.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "key.converter.schemas.enable": "false", "value.converter.schemas.enable": "false", "consumer.override.auto.offset.reset": "earliest", "iceberg.kafka.auto.offset.reset": "earliest", "iceberg.tables": "DATABASE.TABLE", "iceberg.tables.auto-create-enabled": "true", "iceberg.control.commit.interval-ms": "1000", "iceberg.control.commit.timeout-ms": "20000", "iceberg.catalog.type": "jdbc", "iceberg.catalog.uri": "jdbc:postgresql://HOST:PORT/DATABASE?user=USERNAME&password=PASSWORD&ssl=require", "iceberg.catalog.warehouse": "s3://BUCKET_NAME", "iceberg.catalog.client.region": "AWS_REGION", "iceberg.catalog.io-impl": "org.apache.iceberg.aws.s3.S3FileIO", "iceberg.catalog.jdbc.useSSL": "true", "iceberg.catalog.jdbc.verifyServerCertificate": "true", "iceberg.catalog.s3.access-key-id": "AWS_ACCESS_KEY_ID", "iceberg.catalog.s3.secret-access-key": "AWS_SECRET_ACCESS_KEY", "iceberg.catalog.s3.path-style-access": "true", "iceberg.kafka.bootstrap.servers": "KAFKA_HOST:PORT", "iceberg.kafka.security.protocol": "SSL", "iceberg.kafka.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "iceberg.kafka.ssl.keystore.password": "KEYSTORE_PASSWORD", "iceberg.kafka.ssl.keystore.type": "PKCS12", "iceberg.kafka.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "iceberg.kafka.ssl.truststore.password": "TRUSTSTORE_PASSWORD", "iceberg.kafka.ssl.key.password": "KEY_PASSWORD" } ``` Parameters: * `name`: Name of the connector * `connector.class`: Connector implementation class * `topics`: Comma-separated list of Apache Kafka topics with source data * `iceberg.catalog.type`: Catalog type. Use `jdbc` for PostgreSQL * `iceberg.catalog.uri`: JDBC connection string for PostgreSQL * `iceberg.catalog.warehouse`: S3 bucket URI for table storage * `iceberg.catalog.client.region`: AWS region where the S3 bucket is located. Required if no region is set in the environment or system properties * `iceberg.tables`: Target Iceberg table in `DATABASE.TABLE` format. * `iceberg.tables.auto-create-enabled`: Automatically create tables if they do not exist * `iceberg.catalog.io-impl`: File I/O implementation for S3 * `iceberg.catalog.s3.access-key-id` and `secret-access-key`: AWS credentials for S3 * `iceberg.kafka.*`: Apache Kafka security settings For the full list of parameters, see the [Iceberg Kafka Connect configuration](https://iceberg.apache.org/docs/latest/kafka-connect/). ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Aiven Console * Aiven CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 3. In the sink connectors list, select **Iceberg Sink Connector**, and click **Get started**. 4. On the **Iceberg Sink Connector** page, go to the **Common** tab. 5. Locate the **Connector configuration** text box and click **Edit**. 6. Paste the configuration from your `iceberg_sink_connector.json` file into the text box. 7. Click **Create connector**. 8. Verify the connector status on the **Connectors** page. To create the Iceberg sink connector using the [Aiven CLI](/docs/tools/cli.md), run: ``` avn service connector create SERVICE_NAME @iceberg_sink_connector.json ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka® service. * `@iceberg_sink_connector.json`: Path to the JSON configuration file. ## Example[​](#example "Direct link to Example") This example creates an Iceberg sink connector with PostgreSQL as the catalog: ``` { "name": "iceberg_sink_jdbc", "connector.class": "org.apache.iceberg.connect.IcebergSinkConnector", "tasks.max": "2", "topics": "test-topic", "iceberg.catalog.type": "jdbc", "iceberg.catalog.uri": "jdbc:postgresql://postgres.example.com:5432/iceberg_db?user=iceberg_user&password=secret&ssl=require", "iceberg.catalog.warehouse": "s3://my-s3-bucket", "iceberg.catalog.client.region": "us-west-2", "iceberg.catalog.io-impl": "org.apache.iceberg.aws.s3.S3FileIO", "iceberg.catalog.s3.access-key-id": "your-access-key-id", "iceberg.catalog.s3.secret-access-key": "your-secret-access-key", "iceberg.catalog.s3.path-style-access": "true", "iceberg.catalog.jdbc.useSSL": "true", "iceberg.catalog.jdbc.verifyServerCertificate": "true", "iceberg.tables": "mydatabase.mytable", "iceberg.tables.auto-create-enabled": "true", "iceberg.control.commit.interval-ms": "1000", "iceberg.control.commit.timeout-ms": "20000", "key.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "key.converter.schemas.enable": "false", "value.converter.schemas.enable": "false", "iceberg.kafka.bootstrap.servers": "kafka.example.com:9092", "iceberg.kafka.security.protocol": "SSL", "iceberg.kafka.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "iceberg.kafka.ssl.keystore.password": "password", "iceberg.kafka.ssl.keystore.type": "PKCS12", "iceberg.kafka.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "iceberg.kafka.ssl.truststore.password": "password", "iceberg.kafka.ssl.key.password": "password" } ``` Related pages * [Iceberg sink connector overview](/docs/products/kafka/kafka-connect/howto/iceberg-sink-connector.md) * [AWS Glue REST catalog](/docs/products/kafka/kafka-connect/howto/aws-glue-rest-catalog.md) * [PostgreSQL documentation](https://www.postgresql.org/docs/) * [Iceberg Kafka Connect configuration](https://iceberg.apache.org/docs/latest/kafka-connect/) --- # Create a JDBC sink connector from Apache Kafka® to another database The JDBC (Java Database Connectivity) sink connector enables you to move data from an Aiven for Apache Kafka® cluster to any relational database offering JDBC drivers like PostgreSQL® or MySQL. warning The JDBC sink connector requires topics to have a schema to transfer data to relational databases. You can define or manage this schema for each topic through the [Karapace](/docs/products/kafka/karapace.md) schema registry. For a complete list of parameters and configuration options, see the [connector's documentation](https://github.com/aiven/aiven-kafka-connect-jdbc/blob/master/docs/sink-connector.md). ## Prerequisites[​](#connect_jdbc_sink_prereq "Direct link to Prerequisites") * Aiven for Apache Kafka service [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * Gather the following information about your target database service: * `DB_CONNECTION_URL`: The database JDBC connection URL. Examples for PostgreSQL and MySQL: * PostgreSQL: `jdbc:postgresql://HOST:PORT/DB_NAME?sslmode=SSL_MODE`. * MySQL: `jdbc:mysql://HOST:PORT/DB_NAME?ssl-mode=SSL_MODE`. * `DB_USERNAME`: The database username to connect. * `DB_PASSWORD`: The password for the username selected. * `TOPIC_LIST`: The list of topics to sink divided by comma. * `APACHE_KAFKA_HOST`: The hostname of the Apache Kafka service. Required only when using Avro as the data format. * `SCHEMA_REGISTRY_PORT`: The Apache Kafka's schema registry port. Required only when using Avro as data format. * `SCHEMA_REGISTRY_USER`: The Apache Kafka's schema registry username. Required only when using Avro as data format. * `SCHEMA_REGISTRY_PASSWORD`: The Apache Kafka's schema registry user password. Required only when using Avro as data format. For Aiven for PostgreSQL® and Aiven for MySQL®, access the connection details (URL, username, and password) on the **Overview** page of your service in the [Aiven Console](https://console.aiven.io/), or retrieve them using the `avn service get` command in the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). Locate the `SCHEMA_REGISTRY` parameters in the **Schema Registry** tab under **Connection information** on the service **Overview** page. note As of Apache Kafka version 3.0, Aiven for Apache Kafka no longer supports Confluent Schema Registry. Consider using [Karapace](/docs/products/kafka/karapace.md) instead. ## Setup a JDBC sink connector with Aiven Console[​](#setup-a-jdbc-sink-connector-with-aiven-console "Direct link to Setup a JDBC sink connector with Aiven Console") The following example demonstrates setting up a JDBC sink connector for Apache Kafka using the [Aiven Console](https://console.aiven.io/). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Define a Kafka Connect configuration file, such as `jdbc_sink.json`, with the following connector settings: ``` { "name":"CONNECTOR_NAME", "connector.class": "io.aiven.connect.jdbc.JdbcSinkConnector", "topics": "TOPIC_LIST", "connection.url": "DB_CONNECTION_URL", "connection.user": "DB_USERNAME", "connection.password": "DB_PASSWORD", "tasks.max":"1", "auto.create": "true", "auto.evolve": "true", "insert.mode": "upsert", "pk.mode": "record_key", "pk.fields": "field1,field2", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } ``` The configuration file contains the following entries: * `name`: The connector name. * `connector.class`: Specifies the class Kafka Connect will use to create the connector. For JDBC sink connectors, use `io.aiven.connect.jdbc.JdbcSinkConnector`. * `connection.url`, `connection.username`, `connection.password`: JDBC parameters for the sink, collected in the [prerequisite](/docs/products/kafka/kafka-connect/howto/jdbc-sink.md#connect_jdbc_sink_prereq) phase. * `topics`: List the Kafka topics you wish to sink into the database. * `tasks.max`: The maximum number of tasks to execute in parallel. The maximum is 1 per topic and partition. * `auto.create`: Enables automatic creation of the target table in the database if it doesn't exist. * `auto.evolve`: Enables automatic modification of the target table schema to match changes in the Kafka topic messages. * `insert.mode`: Defines how data is inserted into the database: * `insert`: Standard `INSERT` statements. * `upsert`: Upsert semantics supported by the target database. See the [dedicated GitHub repository](https://github.com/aiven/jdbc-connector-for-apache-kafka/blob/master/docs/sink-connector.md). * `update`: Update semantics supported by the target database. See [dedicated GitHub repository](https://github.com/aiven/jdbc-connector-for-apache-kafka/blob/master/docs/sink-connector.md). * `pk.mode`: Defines how the connector identifies rows in the target table (primary key): * `none`: No primary key is used. * `kafka`: Apache Kafka coordinates are used. * `record_key`: Entire or part of the message key is used. * `record_value`: Entire or part of the message value is used. For more information, see the [dedicated GitHub repository](https://github.com/aiven/jdbc-connector-for-apache-kafka/blob/master/docs/sink-connector.md). * `pk.fields`: Defines which fields of the composite key or value to use as record key in the database. * `key.converter` and `value.converter`: Defines the data format of messages within the Apache Kafka topic. The `io.confluent.connect.avro.AvroConverter` translates messages from the Avro format. The message schemas are retrieved from Aiven's [Karapace schema registry](https://github.com/aiven/karapace), as specified by the `schema.registry.url` parameter and related credentials. note The `key.converter` and `value.converter` settings in the connector configuration define how the connector parses messages. When using Avro as source data format, set the following parameters: * `value.converter.schema.registry.url`: Points to the Aiven for Apache Kafka schema registry URL. Use the format `https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT` where `APACHE_KAFKA_HOST` and `SCHEMA_REGISTRY_PORT` are the values [retrieved earlier in the prerequisites](/docs/products/kafka/kafka-connect/howto/jdbc-sink.md#connect_jdbc_sink_prereq). * `value.converter.basic.auth.credentials.source`: Set to `USER_INFO` to enable username and password access to the schema registry. * `value.converter.schema.registry.basic.auth.user.info`: Enter the required schema registry credentials in the `SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD` format, using the `SCHEMA_REGISTRY_USER` and `SCHEMA_REGISTRY_PASSWORD` parameters [retrieved earlier in the prerequisites](/docs/products/kafka/kafka-connect/howto/jdbc-sink.md#connect_jdbc_sink_prereq). ### Create a Kafka Connect connector with the Aiven Console[​](#create-a-kafka-connect-connector-with-the-aiven-console "Direct link to Create a Kafka Connect connector with the Aiven Console") To create a Kafka Connect connector: 1. Log in to the [Aiven Console](https://console.aiven.io/) and select the Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service where the connector needs to be defined. 2. Click **Manage stream** > **Connectors** from the sidebar. 3. Click **Create connector** to start setting up a new connector. This option is visible only if [Kafka Connect](/docs/products/kafka/kafka-connect/howto/enable-connect.md) is enabled for your service. 4. On the **Select connector** page, locate **JDBC Sink** and click **Get started**. 5. In the **Common** tab, find the **Connector configuration** text box. 6. Click **Edit** to modify the connector configuration. 7. Paste the configuration details from your `jdbc_sink.json` file into the text box. 8. Click **Apply**. note The Aiven Console automatically populates the UI fields with the data from the configuration file. You can review and edit these fields across the different tabs. Any modifications you make are updated in the **Connector configuration** text box in JSON format. 9. Once you've entered all the required settings, click **Create connector**. 10. Verify the connector status on the **Manage stream** > **Connectors** page. 11. Confirm that the data has appeared in the target database service. The table name should match the Apache Kafka topic name. You can also create connectors using the [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create). ## Example: Create a JDBC sink connector to PostgreSQL® on a topic with a JSON schema[​](#example-create-a-jdbc-sink-connector-to-postgresql-on-a-topic-with-a-json-schema "Direct link to Example: Create a JDBC sink connector to PostgreSQL® on a topic with a JSON schema") Suppose you have a topic named `iot_measurements` that contains data in JSON format with a defined JSON schema as follows: ``` { "schema": { "type":"struct", "fields":[{ "type":"int64", "optional": false, "field": "iot_id" },{ "type":"string", "optional": false, "field": "metric" },{ "type":"int32", "optional": false, "field": "measurement" }] }, "payload":{ "iot_id":1, "metric":"Temperature", "measurement":14} } { "schema": { "type":"struct", "fields":[{ "type":"int64", "optional": false, "field": "iot_id" },{ "type":"string", "optional": false, "field": "metric" },{ "type":"int32", "optional": false, "field": "measurement" }] }, "payload":{"iot_id":2, "metric":"Humidity", "measurement":60} } ``` note Embedding a JSON schema in every message can increase the data size. For a more efficient size-to-content ratio, consider using the Avro format with the [Karapace schema registry](https://karapace.io/). To sink the `iot_measurements` topic to PostgreSQL, use the following connector configuration. Replace the placeholders for `DB_HOST`, `DB_PORT`, `DB_NAME`, `DB_SSL_MODE`, `DB_USERNAME`, and `DB_PASSWORD`: ``` { "name":"sink_iot_json_schema", "connector.class": "io.aiven.connect.jdbc.JdbcSinkConnector", "topics": "iot_measurements", "connection.url": "jdbc:postgresql://DB_HOST:DB_PORT/DB_NAME?sslmode=DB_SSL_MODE", "connection.user": "DB_USERNAME", "connection.password": "DB_PASSWORD", "tasks.max":"1", "auto.create": "true", "auto.evolve": "true", "insert.mode": "upsert", "pk.mode": "record_value", "pk.fields": "iot_id", "value.converter": "org.apache.kafka.connect.json.JsonConverter" } ``` Key aspects of the configuration: * `"topics": "iot_measurements"`: Identifies `iot_measurements` topic as the data source for the sink operation. * `"value.converter": "org.apache.kafka.connect.json.JsonConverter"`: Indicates that the message value is in plain JSON format without a schema. Since the key is empty, no converter is defined for it. * `"pk.mode": "record_value"`: Indicates the connector is using the message value to set the target database key. * `"pk.fields": "iot_id"`: Indicates the connector is using the field `iot_id` on the message value to set the target database key. ## Example: Create a JDBC sink connector to MySQL on a topic using Avro and schema registry[​](#example-create-a-jdbc-sink-connector-to-mysql-on-a-topic-using-avro-and-schema-registry "Direct link to Example: Create a JDBC sink connector to MySQL on a topic using Avro and schema registry") Suppose you have a topic named `students` that contains data in Avro format. The schema is stored in the schema registry provided by [Karapace](/docs/products/kafka/karapace.md) and has the following structure: ``` key: {"student_id": 1234} value: {"student_name": "Mary", "exam": "Math", "exam_result":"A"} ``` To sink the `students` topic to MySQL, use the following connector configuration. Make sure to replace the placeholders for `DB_HOST`, `DB_PORT`, `DB_NAME`, `DB_SSL_MODE`, DB\_USERNAME, DB\_PASSWORD, APACHE\_KAFKA\_HOST, SCHEMA\_REGISTRY\_PORT, `SCHEMA_REGISTRY_USER`, and `SCHEMA_REGISTRY_PASSWORD`: ``` { "name": "sink_students_avro_schema", "connector.class": "io.aiven.connect.jdbc.JdbcSinkConnector", "topics": "students", "connection.url": "jdbc:mysql://DB_HOST:DB_PORT/DB_NAME?ssl-mode=DB_SSL_MODE", "connection.user": "DB_USERNAME", "connection.password": "DB_PASSWORD", "insert.mode": "upsert", "table.name.format": "students", "pk.mode": "record_key", "pk.fields": "student_id", "auto.create": "true", "auto.evolve": "true", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } ``` Key aspects of the configuration: * `"topics": "students"`: Identifies `students` topic as the data source for the sink operation. * `"pk.mode": "record_key"`: Uses the message key as the database key. * `"pk.fields": "student_id"`: Sets the database key using the `student_id` field from the message key. * `key.converter` and `value.converter`: Defines the Avro data format with `io.confluent.connect.avro.AvroConverter` and provides the URL and credentials for the [Karapace](/docs/products/kafka/karapace.md) schema registry. With `"auto.create": "true"`, the connector automatically creates a `students` table in the MySQL database. This table is populated with data from the `students` Apache Kafka topic and includes the `student_id`, `student_name`, `exam`, and `exam_result` columns. Related pages * View the [Database migration with Apache Kafka® and Apache Kafka® Connect](https://aiven.io/blog/db-technology-migration-with-apache-kafka-and-kafka-connect) blog post --- # Create a JDBC source connector from MySQL to Apache Kafka® The JDBC source connector pushes data from a relational database, such as MySQL, to Apache Kafka® where can be transformed and read by multiple consumers. tip Sourcing data from a database into Apache Kafka decouples the database from the set of consumers. Once the data is in Apache Kafka, multiple applications can access it without adding any additional query overhead to the source database. note See the full set of available parameters and configuration options in the [connector's documentation](https://github.com/aiven/aiven-kafka-connect-jdbc/blob/master/docs/source-connector.md). ## Prerequisites[​](#connect_jdbc_mysql_source_prereq "Direct link to Prerequisites") To setup a JDBC source connector pointing to MySQL, you need an Aiven for Apache Kafka service [with Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). Also collect the following information about the source MySQL database upfront: * `MYSQL_HOST`: The database hostname * `MYSQL_PORT`: The database port * `MYSQL_USER`: The database user to connect * `MYSQL_PASSWORD`: The database password for the `MYSQL_USER` * `MYSQL_DATABASE_NAME`: The database name * `MYSQL_TABLES`: The list of database tables to be included in Apache Kafka; the list must be in the form of `table_name1,table_name2` note If you're using Aiven for MySQL the above details are available in the [Aiven console](https://console.aiven.io/) service Overview tab or via the dedicated `avn service get` command with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). ## Setup a MySQL JDBC source connector with Aiven CLI[​](#setup-a-mysql-jdbc-source-connector-with-aiven-cli "Direct link to Setup a MySQL JDBC source connector with Aiven CLI") The following example demonstrates how to setup an Apache Kafka JDBC source connector to a MySQL database using the [Aiven CLI dedicated command](/docs/tools/cli/service/connector.md#avn_service_connector_create). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Define the connector configurations in a file (we'll refer to it with the name `jdbc_source_mysql.json`) with the following content: ``` { "name":"CONNECTOR_NAME", "connector.class":"io.aiven.connect.jdbc.JdbcSourceConnector", "connection.url":"jdbc:mysql://MYSQL_HOST:MYSQL_PORT/MYSQL_DATABASE_NAME?&verifyServerCertificate=false&useSSL=true&requireSSL=true", "connection.user":"MYSQL_USER", "connection.password":"MYSQL_PASSWORD", "table.whitelist":"MYSQL_TABLES", "mode":"JDBC_MODE", "topic.prefix":"KAFKA_TOPIC_PREFIX", "tasks.max":"NR_TASKS", "poll.interval.ms":"POLL_INTERVAL" } ``` The configuration file contains the following entries: * `name`: the connector name * `MYSQL_HOST`, `MYSQL_PORT`, `MYSQL_DATABASE_NAME`, `MYSQL_USER`, `MYSQL_PASSWORD` and `MYSQL_TABLES`: source database parameters collected in the [prerequisite](/docs/products/kafka/kafka-connect/howto/jdbc-source-connector-mysql.md#connect_jdbc_mysql_source_prereq) phase. * `mode`: the query mode, more information in the [dedicated page](/docs/products/kafka/kafka-connect/concepts/jdbc-source-modes.md); depending on the selected mode, additional configuration entries might be required. * `topic.prefix`: the prefix that will be used for topic names. The resulting topic name will be the concatenation of the `topic.prefix` and the schema and table name. * `tasks.max`: maximum number of tasks to execute in parallel. By default is 1, the connector can use at max 1 task for each source table defined. * `poll.interval.ms`: query frequency, default 5000 milliseconds See the [dedicated documentation](https://github.com/aiven/jdbc-connector-for-apache-kafka/blob/master/docs/source-connector-config-options.rst) for the full list of parameters. ### Create a Kafka Connect connector with Aiven CLI[​](#create-a-kafka-connect-connector-with-aiven-cli "Direct link to Create a Kafka Connect connector with Aiven CLI") To create the connector, execute the following [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create), replacing the `SERVICE_NAME` with the name of the Aiven service where the connector needs to run: ``` avn service connector create SERVICE_NAME @jdbc_source_mysql.json ``` Check the connector status with the following command, replacing the `SERVICE_NAME` with the Aiven service and the `CONNECTOR_NAME` with the name of the connector defined before: ``` avn service connector status SERVICE_NAME CONNECTOR_NAME ``` Verify in the Apache Kafka target instance, the presence of the topic and the data tip If you're using Aiven for Apache Kafka, topics will not be created automatically. Either create them manually following the `topic.prefix.schema_name.table_name` naming pattern or enable the `kafka.auto_create_topics_enable` advanced parameter. ## Example: define a JDBC incremental connector[​](#example-define-a-jdbc-incremental-connector "Direct link to Example: define a JDBC incremental connector") The example creates an [incremental](/docs/products/kafka/kafka-connect/concepts/jdbc-source-modes.md) JDBC connector with the following properties: * connector name: `jdbc_source_mysql_increment` * source tables: `students` and `exams`, available in an Aiven for MySQL database * [incremental column name](/docs/products/kafka/kafka-connect/concepts/jdbc-source-modes.md): `id` * topic prefix: `jdbc_source_mysql_increment.` * maximum number of concurrent tasks: 1 * time interval between queries: 5 seconds The connector configuration is the following: ``` { "name":"jdbc_source_mysql_increment", "connector.class":"io.aiven.connect.jdbc.JdbcSourceConnector", "connection.url":"jdbc:mysql://demo-mysql-myproject.aivencloud.com:13039/defaultdb?sslmode=require", "connection.user":"avnadmin", "connection.password":"mypassword123", "table.whitelist":"students,exams", "mode":"incrementing", "incrementing.column.name":"id", "topic.prefix":"jdbc_source_mysql_increment.", "tasks.max":"1", "poll.interval.ms":"5000" } ``` With the above configuration stored in a `jdbc_incremental_source_mysql.json` file, you can create the connector in the `demo-kafka` instance with: ``` avn service connector create demo-kafka @jdbc_incremental_source_mysql.json ``` --- # Create a JDBC source connector from PostgreSQL® to Apache Kafka® The JDBC source connector pushes data from a relational database, such as PostgreSQL®, to Apache Kafka® where can be transformed and read by multiple consumers. tip Sourcing data from a database into Apache Kafka decouples the database from the set of consumers. Once the data is in Apache Kafka, multiple applications can access it without adding any additional query overhead to the source database. note See the full set of available parameters and configuration options in the [connector's documentation](https://github.com/aiven/aiven-kafka-connect-jdbc/blob/master/docs/source-connector.md). ## Prerequisites[​](#connect_jdbc_pg_source_prereq "Direct link to Prerequisites") To setup a JDBC source connector pointing to PostgreSQL, you need an Aiven for Apache Kafka service [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). Also collect the following information about the source PostgreSQL database upfront: * `PG_HOST`: The database hostname * `PG_PORT`: The database port * `PG_USER`: The database user to connect * `PG_PASSWORD`: The database password for the `PG_USER` * `PG_DATABASE_NAME`: The database name * `SSL_MODE`: The [SSL mode](https://www.postgresql.org/docs/current/libpq-ssl.html) * `PG_TABLES`: The list of database tables to be included in Apache Kafka; the list must be in the form of `schema_name1.table_name1,schema_name2.table_name2` note If you're using Aiven for PostgreSQL the above details are available in the [Aiven console](https://console.aiven.io/) service Overview tab or via the dedicated `avn service get` command with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). ## Setup a PostgreSQL JDBC source connector with Aiven CLI[​](#setup-a-postgresql-jdbc-source-connector-with-aiven-cli "Direct link to Setup a PostgreSQL JDBC source connector with Aiven CLI") The following example demonstrates how to setup an Apache Kafka JDBC source connector to a PostgreSQL database using the [Aiven CLI dedicated command](/docs/tools/cli/service/connector.md#avn_service_connector_create). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Define the connector configurations in a file (we'll refer to it with the name `jdbc_source_pg.json`) with the following content: ``` { "name":"CONNECTOR_NAME", "connector.class":"io.aiven.connect.jdbc.JdbcSourceConnector", "connection.url":"jdbc:postgresql://PG_HOST:PG_PORT/PG_DATABASE_NAME?sslmode=SSL_MODE", "connection.user":"PG_USER", "connection.password":"PG_PASSWORD", "table.whitelist":"PG_TABLES", "mode":"JDBC_MODE", "topic.prefix":"KAFKA_TOPIC_PREFIX", "tasks.max":"NR_TASKS", "poll.interval.ms":"POLL_INTERVAL" } ``` The configuration file contains the following entries: * `name`: the connector name * `PG_HOST`, `PG_PORT`, `PG_DATABASE_NAME`, `SSL_MODE`, `PG_USER`, `PG_PASSWORD` and `PG_TABLES`: source database parameters collected in the [prerequisite](/docs/products/kafka/kafka-connect/howto/jdbc-source-connector-pg.md#connect_jdbc_pg_source_prereq) phase. * `mode`: the query mode, more information in the [dedicated page](/docs/products/kafka/kafka-connect/concepts/jdbc-source-modes.md); depending on the selected mode, additional configuration entries might be required. * `topic.prefix`: the prefix that will be used for topic names. The resulting topic name will be the concatenation of the `topic.prefix` and the table name. * `tasks.max`: maximum number of tasks to execute in parallel. By default is 1, the connector can use at max 1 task for each source table defined. * `poll.interval.ms`: query frequency, default 5000 milliseconds See the [dedicated documentation](https://github.com/aiven/jdbc-connector-for-apache-kafka/blob/master/docs/source-connector-config-options.rst) for the full list of parameters. tip Check the [dedicated blog post](https://aiven.io/blog/using-kafka-connect-jdbc-source-a-postgresql-example) for an end-to-end example of the JDBC source connector in action with PostgreSQL®. ### Create a Kafka Connect connector with Aiven CLI[​](#create-a-kafka-connect-connector-with-aiven-cli "Direct link to Create a Kafka Connect connector with Aiven CLI") To create the connector, execute the following [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create), replacing the `SERVICE_NAME` with the name of the Aiven service where the connector needs to run: ``` avn service connector create SERVICE_NAME @jdbc_source_pg.json ``` Check the connector status with the following command, replacing the `SERVICE_NAME` with the Aiven service and the `CONNECTOR_NAME` with the name of the connector defined before: ``` avn service connector status SERVICE_NAME CONNECTOR_NAME ``` Verify in the Apache Kafka target instance, the presence of the topic and the data tip If you're using Aiven for Apache Kafka, topics will not be created automatically. Either create them manually following the `topic.prefix.schema_name.table_name` naming pattern or enable the `kafka.auto_create_topics_enable` advanced parameter. ## Example: define a JDBC incremental connector[​](#example-define-a-jdbc-incremental-connector "Direct link to Example: define a JDBC incremental connector") The example creates an [incremental](/docs/products/kafka/kafka-connect/concepts/jdbc-source-modes.md) JDBC connector with the following properties: * connector name: `jdbc_source_pg_increment` * source tables: `students` and `exams` from the `public` schema, available in an Aiven for PostgreSQL database * [incremental column name](/docs/products/kafka/kafka-connect/concepts/jdbc-source-modes.md): `id` * topic prefix: `jdbc_source_pg_increment.` * maximum number of concurrent tasks: `1` * time interval between queries: 5 seconds The connector configuration is the following: ``` { "name":"jdbc_source_pg_increment", "connector.class":"io.aiven.connect.jdbc.JdbcSourceConnector", "connection.url":"jdbc:postgresql://demo-pg-myproject.aivencloud.com:13039/defaultdb?sslmode=require", "connection.user":"avnadmin", "connection.password":"mypassword123", "table.whitelist":"public.students,public.exams", "mode":"incrementing", "incrementing.column.name":"id", "topic.prefix":"jdbc_source_pg_increment.", "tasks.max":"1", "poll.interval.ms":"5000" } ``` With the above configuration stored in a `jdbc_incremental_source_pg.json` file, you can create the connector in the `demo-kafka` instance with: ``` avn service connector create demo-kafka @jdbc_incremental_source_pg.json ``` --- # Create a JDBC source connector from SQL Server to Apache Kafka® The JDBC source connector pushes data from a relational database, such as SQL Server, to Apache Kafka® where can be transformed and read by multiple consumers. tip Sourcing data from a database into Apache Kafka decouples the database from the set of consumers. Once the data is in Apache Kafka, multiple applications can access it without adding any additional query overhead to the source database. note See the full set of available parameters and configuration options in the [connector's documentation](https://github.com/aiven/aiven-kafka-connect-jdbc/blob/master/docs/source-connector.md). ## Prerequisites[​](#connect_jdbc_sqlserver_source_prereq "Direct link to Prerequisites") To setup a JDBC source connector pointing to SQL Server, you need an Aiven for Apache Kafka service [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). Also collect the following information about the source SQL Server database upfront: * `SQLSERVER_HOST`: The database hostname * `SQLSERVER_PORT`: The database port * `SQLSERVER_USER`: The database user to connect * `SQLSERVER_PASSWORD`: The database password for the `SQLSERVER_USER` * `SQLSERVER_DATABASE_NAME`: The database name * `SQLSERVER_TABLES`: The list of database tables to be included in Apache Kafka; the list must be in the form of `table_name1,table_name2` note If you're using Aiven for SQL Server the above details are available in the [Aiven console](https://console.aiven.io/) service Overview tab or via the dedicated `avn service get` command with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). ## Setup a SQL Server JDBC source connector with Aiven CLI[​](#setup-a-sql-server-jdbc-source-connector-with-aiven-cli "Direct link to Setup a SQL Server JDBC source connector with Aiven CLI") The following example demonstrates how to setup an Apache Kafka JDBC source connector to a SQL Server database using the [Aiven CLI dedicated command](/docs/tools/cli/service/connector.md#avn_service_connector_create). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Define the connector configurations in a file (we'll refer to it with the name `jdbc_source_sqlserver.json`) with the following content: ``` { "name":"CONNECTOR_NAME", "connector.class":"io.aiven.connect.jdbc.JdbcSourceConnector", "connection.url":"jdbc:sqlserver://SQLSERVER_HOST:SQLSERVER_PORT;databaseName=SQLSERVER_DATABASE_NAME;", "connection.user":"SQLSERVER_USER", "connection.password":"SQLSERVER_PASSWORD", "table.whitelist":"SQLSERVER_TABLES", "mode":"JDBC_MODE", "topic.prefix":"KAFKA_TOPIC_PREFIX", "tasks.max":"NR_TASKS", "poll.interval.ms":"POLL_INTERVAL" } ``` The configuration file contains the following entries: * `name`: the connector name. * `SQLSERVER_HOST`, `SQLSERVER_PORT`, `SQLSERVER_DATABASE_NAME`, `SQLSERVER_USER`, `SQLSERVER_PASSWORD` and `SQLSERVER_TABLES`: source database parameters collected in the [prerequisite](/docs/products/kafka/kafka-connect/howto/jdbc-source-connector-sql-server.md#connect_jdbc_sqlserver_source_prereq) phase. * `mode`: the query mode, more information in the [dedicated page](/docs/products/kafka/kafka-connect/concepts/jdbc-source-modes.md); depending on the selected mode, additional configuration entries might be required. * `topic.prefix`: the prefix that will be used for topic names. The resulting topic name will be the concatenation of the `topic.prefix` and the schema and table name. * `tasks.max`: maximum number of tasks to execute in parallel. By default is 1, the connector can use at max 1 task for each source table defined. * `poll.interval.ms`: query frequency, default 5000 milliseconds. See the [dedicated documentation](https://github.com/aiven/jdbc-connector-for-apache-kafka/blob/master/docs/source-connector-config-options.rst) for the full list of parameters. ### Create a Kafka Connect connector with Aiven CLI[​](#create-a-kafka-connect-connector-with-aiven-cli "Direct link to Create a Kafka Connect connector with Aiven CLI") To create the connector, execute the following [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create), replacing the `SERVICE_NAME` with the name of the Aiven service where the connector needs to run: ``` avn service connector create SERVICE_NAME @jdbc_source_sqlserver.json ``` Check the connector status with the following command, replacing the `SERVICE_NAME` with the Aiven service and the `CONNECTOR_NAME` with the name of the connector defined before: ``` avn service connector status SERVICE_NAME CONNECTOR_NAME ``` Verify in the Apache Kafka target instance, the presence of the topic and the data. tip If you're using Aiven for Apache Kafka, topics will not be created automatically. Either create them manually following the `topic.prefix.schema_name.table_name` naming pattern or enable the `kafka.auto_create_topics_enable` advanced parameter. ## Example: define a JDBC incremental connector[​](#example-define-a-jdbc-incremental-connector "Direct link to Example: define a JDBC incremental connector") The example creates an [incremental](/docs/products/kafka/kafka-connect/concepts/jdbc-source-modes.md) JDBC connector with the following properties: * connector name: `jdbc_source_sqlserver_increment` * source tables: `students` and `exams`, available in an Aiven for Server database * [incremental column name](/docs/products/kafka/kafka-connect/concepts/jdbc-source-modes.md): `id` * topic prefix: `jdbc_source_sqlserver_increment` * maximum number of concurrent tasks: `1` * time interval between queries: 5 seconds The connector configuration is the following: ``` { "name":"jdbc_source_sqlserver_increment", "connector.class":"io.aiven.connect.jdbc.JdbcSourceConnector", "connection.url":"jdbc:sqlserver://demo-sqlserver-myproject.aivencloud.com:13039;databaseName=defaultdb;", "connection.user":"avnadmin", "connection.password":"mypassword123", "table.whitelist":"students,exams", "mode":"incrementing", "incrementing.column.name":"id", "topic.prefix":"jdbc_source_sqlserver_increment.", "tasks.max":"1", "poll.interval.ms":"5000" } ``` With the above configuration stored in the `jdbc_incremental_source_sqlserver.json` file, you can create the connector in the `demo-kafka` instance with: ``` avn service connector create demo-kafka @jdbc_incremental_source_sqlserver.json ``` --- # Integrate Aiven for Apache Kafka® Connect with PostgreSQL using Debezium with mutual TLS Integrate Aiven for Apache Kafka® Connect with PostgreSQL using Debezium with mutual TLS (mTLS) to enhance security with mutual authentication. This configuration establishes a secure and efficient data synchronization channel between Aiven for Apache Kafka® Connect and a PostgreSQL database. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to an Aiven for Apache Kafka® and Aiven for Apache Kafka® Connect service. * Administrative access to a PostgreSQL database with SSL enabled. * The following SSL certificates and keys from your PostgreSQL database: * Client certificate: Public certificate for client authentication. * Root certificate: Certificate authority (CA) certificate used to sign the server and client certificates. * Client key: Private key for encrypting data sent to the server. For additional details, see [Certificate requirements](/docs/platform/concepts/tls-ssl-certificates.md#certificate-requirements). ### For CloudSQL PostgreSQL databases[​](#for-cloudsql-postgresql-databases "Direct link to For CloudSQL PostgreSQL databases") If you are using CloudSQL, complete these additional steps if they are not already configured: * **Allow Aiven IP addresses**: Whitelist Aiven IP addresses to enable connections from Apache Kafka Connect to your CloudSQL instance. * **Enable WAL and logical decoding**: * Set `cloudsql.logical_decoding` to `on` to capture database change events. * Set `cloudsql.enable_pglogical` to `on` to configure the Write-Ahead Log (WAL) for logical replication. * Restart the CloudSQL instance to apply the changes. * **Enable the `pgoutput` extension** In your CloudSQL instance, enable the extension that Debezium uses for logical decoding: ``` CREATE EXTENSION pgoutput; ``` * **Verify SSL configuration**: After whitelisting IP addresses, test the SSL connection to ensure your certificates are configured correctly: ``` psql "sslmode=verify-ca \ sslrootcert=server-ca.pem \ sslcert=client-cert.pem \ sslkey=client-key.pem \ hostaddr= \ port=5432 \ user= \ dbname=" ``` * **Set the Debezium output plugin**: Set the `plugin.name` field in the connector configuration to either `pgoutput` or `decoderbufs`. note Starting with Debezium 2.5, the `wal2json` plugin is deprecated. Use `pgoutput` or `decoderbufs` as the recommended replacement. ## Limitations[​](#limitations "Direct link to Limitations") * The integration is not yet available in the Aiven Console. Use the Aiven CLI or Aiven API instead. * A dedicated Aiven for Apache Kafka® Connect service is required. Enabling Kafka Connect within the same Aiven for Apache Kafka® service is not supported for this integration. * The `kafka_connect_postgresql` integration type is not yet available in the Aiven Console. * Delivering SSL keys to Aiven for Apache Kafka Connect is asynchronous. There may be a delay of up to five minutes between creating the integration and using the keys. * Mutual authentication in `verify-full` SSL mode is not supported in the Apache Kafka setup. The `verify-ca` mode remains secure because the Certificate Authority (CA) is unique to the instance. * For CloudSQL PostgreSQL databases, enable logical decoding and the `pgoutput` extension to ensure replication compatibility. ## Variables[​](#variables "Direct link to Variables") Replace the placeholders in the following table with values from your environment. Use these values in the configuration steps and code snippets. | Variable | Description | | ----------------------------- | -------------------------------------------------- | | `` | Name of your Aiven for Apache Kafka service | | `` | Name of your Apache Kafka Connect service | | `` | Your cloud provider and region identifier | | `` | Service plan for Apache Kafka/Apache Kafka Connect | | `` | Name of your project in the Aiven platform | | `` | JSON configuration for PostgreSQL endpoint | | `` | ID of the created PostgreSQL endpoint | | `` | JSON configuration for Debezium connector | ## Configure the integration[​](#configure-the-integration "Direct link to Configure the integration") 1. Verify that your Aiven for Apache Kafka® service is active and accessible. If you do not have a Kafka service, create one: ``` avn service create \ --service-type kafka \ --cloud \ --plan \ --project ``` note Ensure [topic auto-creation](/docs/products/kafka/howto/create-topics-automatically.md) is enabled to automatically generate required Apache Kafka topics. If disabled, manually create topics before starting the connector. 2. Verify that your Aiven for Apache Kafka® Connect service is active and accessible. If you do not have one, use the following command to create it and connect it to your Kafka service: ``` avn service create \ --service-type kafka_connect \ --cloud \ --plan \ --project $PROJECT ``` 3. Create an external PostgreSQL integration endpoint. This endpoint represents your PostgreSQL database. Replace the placeholders in the following JSON configuration: ``` INTEGRATION_CONFIG=$(cat <<-END { "ssl_mode": "verify-ca", "host": "", "port": , "username": "", "ssl_client_certificate": "$(cat /path/to/your/client-cert.pem)", "ssl_root_cert": "$(cat /path/to/your/ca.pem)", "ssl_client_key": "$(cat /path/to/your/client-key.pem)", "password": my_password } END ) avn service integration-endpoint-create \ --project $PROJECT \ --endpoint-name external_postgresql \ --user-config "$INTEGRATION_CONFIG" ``` 4. Retrieve the ID of the external PostgreSQL integration endpoint: ``` INTEGRATION_ENDPOINT_ID=$( avn service integration-endpoint-list --project $PROJECT \ | grep external_postgresql \ | awk '{print $1}' ) ``` 5. Integrate the Apache Kafka® Connect service with the external PostgreSQL endpoint: ``` avn service integration-create \ --project $PROJECT \ --integration-type kafka_connect_postgresql \ --source-endpoint-id $INTEGRATION_ENDPOINT_ID \ --dest-service ``` 6. Create the Debezium connector configuration to monitor your PostgreSQL database. Replace the placeholders with your PostgreSQL and Apache Kafka® Connect information: ``` CONNECTOR_CONFIG=$(cat <<-END { "name": "debezium-postgres-connector", "connector.class": "io.debezium.connector.postgresql.PostgresConnector", // Omitted as provided by external PostgreSQL integration: // "database.hostname", "database.port", "database.user", "database.password", // "database.dbname", "database.sslcert", "database.sslkey", "database.sslmode", // "database.sslrootcert" "database.server.name": "", "plugin.name": "pgoutput", "publication.name": "debezium_publication", "publication.autocreate.mode": "all_tables", "endpoint_id": "$INTEGRATION_ENDPOINT_ID", "topic.prefix": "my_prefix", // Topics will be named as "{prefix}.{database_name}.{table_name}" "database.tcpKeepAlive": "true", // Optional Transforms (uncomment if needed) //"transforms": "unwrap", //"transforms.unwrap.type": "io.debezium.transforms.ExtractNewRecordState" } END ) avn service connector create \ --project $PROJECT \ --service \ --config "$CONNECTOR_CONFIG" ``` Parameters: * `name`: The connector instance name. For example, `debezium-postgres-connector`. * `connector.class`: The class of the PostgreSQL connector. For example, `io.debezium.connector.postgresql.PostgresConnector`. * `database.server.name`: The identifier for the database server within Apache Kafka Connect. * `plugin.name`: The PostgreSQL logical decoding plugin. * `publication.name`: The PostgreSQL publication for tracking database changes. * `publication.autocreate.mode`: If set to `all_tables`, it captures changes for all tables automatically. * `endpoint_id`: The unique ID for the PostgreSQL endpoint managed externally. * `topic.prefix`: The prefix for Kafka topics receiving database change events. * `database.tcpKeepAlive`: If set to `true`, it prevents database connection timeouts during inactivity. * `transforms` and `transforms.unwrap.type`: The data transformations. `ExtractNewRecordState` extracts the latest data state. These transformations are optional and should be used based on your specific requirements. ## Verification[​](#verification "Direct link to Verification") After completing the setup, confirm the following: 1. Apache Kafka Connect establishes a secure SSL connection with the PostgreSQL database. 2. Apache Kafka topics exist as expected, whether created automatically or manually. 3. Data is streaming into topics using the naming pattern `{connector_name}.{database_name}.{table_name}`. --- # Manage connector versions in Aiven for Apache Kafka Connect® Aiven for Apache Kafka Connect® lets you control which connector version is used in your service. ## Benefits[​](#benefits "Direct link to Benefits") It prevents compatibility issues from automatic updates, avoids breaking changes when multiple dependencies rely on a specific connector version, and allows testing, upgrading, or reverting versions to maintain production pipeline stability. You can specify the version of a connector by setting the `plugin_versions` property in the service configuration. ## Key considerations[​](#key-considerations "Direct link to Key considerations") * Deprecated connector versions may be removed during maintenance updates. If you [set](#set-version) a deprecated version, the system alerts you and recommends an upgrade. The connector will continue to run, but upgrading to a supported version is recommended to avoid compatibility issues. * Support is limited to the latest connector version and the most recent previous version. Breaking changes, if known, are mentioned in [maintenance update notifications](/docs/products/kafka/howto/maintenance-updates.md). * Setting a connector version applies to the entire plugin, ensuring that all connectors provided by the plugin (such as source and sink connectors) use the same version. * Downgrading a connector to an older version is supported if that version is available. On older service nodes, the selected version takes effect only after a maintenance update. * Multi-version support is available for all connectors where Aiven has published and supports more than one version. Support will continue to expand as new versions are released. * Set a specific version to prevent automatic upgrades. If not set, the latest published version is used. * See [Check available connector versions](#check-available-connector-versions) to confirm which versions are supported before setting a version. * Setting `plugin_version` in the connector configuration (for example, in the `config` block in Terraform) is not supported and has no effect. To control connector versions, use the [`plugin_versions`](#set-version) property in the **Kafka Connect service** configuration. * You can set connector versions even if the service is running on older nodes. However, the selected version only takes effect **after** a maintenance update. Until then, the connector runs the latest available version, and no warning is shown. warning Breaking changes may exist between different connector versions. These changes often involve updates to configuration parameters, such as the removal of deprecated parameters or the introduction of new ones. Before upgrading: * Review the connector release notes for details on changes. * Always test a connector version change in a non-production environment before applying it in production. note Aiven supports multiple Debezium versions through multi-version support, including versions 1.9.7, 2.5.0, 2.7.4, and 3.1.0. To prevent automatic upgrades during maintenance, pin the connector version using the `plugin_versions` property. If you use Debezium for PostgreSQL with version 1.9.7 and the `wal2json` format, do not upgrade to version 2.0 or later until you migrate to `pgoutput`. For upgrade steps, see [Set a connector version](#set-version). ## Limitations[​](#limitations "Direct link to Limitations") Connector version selection is available in the Aiven Console only for [dedicated Apache Kafka Connect services](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). If you enabled [Apache Kafka Connect](/docs/products/kafka/kafka-connect/howto/enable-connect.md) as part of an Aiven for Apache Kafka service, use the [Aiven API](https://api.aiven.io/doc/), [Aiven CLI](/docs/tools/cli.md), or [Aiven Provider for Terraform](https://registry.terraform.io/providers/aiven/aiven/latest/docs) to set the connector version. note The version selector in the Aiven Console appears on individual connector configuration pages, but changing the version affects all instances of that connector within the service, not just the selected instance. Apache Kafka Connect does not support loading multiple versions of the same plugin within a service. This limitation is expected to be addressed in Apache Kafka 4.1.0 through [KIP-891](https://cwiki.apache.org/confluence/display/KAFKA/KIP-891%3A+Running+multiple+versions+of+Connector+plugins). Until then, if you run multiple instances of a connector in a service, they must all use the same plugin version. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Aiven for Apache Kafka® service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster) enabled * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) ## Check available connector versions[​](#check-available-connector-versions "Direct link to Check available connector versions") Before selecting a connector version, check which versions are available for your Aiven for Apache Kafka Connect service. This ensures that the desired version is supported and can be set if needed. Use one of the following methods: * Aiven Console * Aiven API * Aiven CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka Connect service. 3. Click **Manage stream** > **Connectors**. Connectors that support multiple versions display **2 versions** next to their names on the **Manage stream** > **Connectors** page. note When setting up a new connector, a default version is selected. To change it, go to the connector details page and click **Change version**. 1. Run the following command: ``` curl -X GET "https://api.aiven.io/v1/project//service//available-connectors" \ -H "Authorization: Bearer " ``` 2. Review the response to see the available versions. If multiple versions are listed, the connector supports multi-versioning. Example output: ``` { "plugin_name": "aiven-kafka-connect-jdbc", "available_versions": [ { "version": "6.10.0", "deprecated": false }, { "version": "6.9.0", "deprecated": true } ] } ``` 1) Run the following command: ``` avn service connector available ``` 2) Review the output to see the available versions. If multiple versions are listed, the connector supports multi-versioning. Example output: ``` { "plugin_name": "aiven-kafka-connect-jdbc", "available_versions": [ { "version": "6.10.0", "deprecated": false }, { "version": "6.9.0", "deprecated": true } ] } ``` ## Set a connector version[​](#set-version "Direct link to Set a connector version") To set a specific connector version, update the `plugin_versions` property in the service configuration using the API, CLI, or Terraform. In the Aiven Console, you can select the version through the UI. The selected version applies to all instances of the connector, including both source and sink connectors. For example, setting the `aiven-kafka-connect-jdbc` plugin to version `6.9.0` affects both the JDBC source and sink connectors. note Changing the connector version restarts the Apache Kafka Connect service and reloads all plugins. This process can take several minutes. * Aiven Console * Aiven API * Aiven CLI * Terraform note Connector version selection in the Aiven Console is available only for [dedicated Apache Kafka Connect services](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). 1. In your Aiven for Apache Kafka Connect service, click **Connectors**. 2. In the **Enabled connectors** section, locate the connector to update. 3. Click **Actions** > **Change connector version**. 4. In the **Version setup** window: * Select the version to use. * Optional: If you select the latest version, you can turn on **Enable version updates** to automatically update the connector to newer versions during maintenance updates. note **Enable version updates** is available only for the latest (default) version. This option is unavailable for older versions because automatic updates apply only to the latest version. 5. Depending on your selection, click: * **Install version and restart service** to apply the selected version. * **Confirm version** to keep the current version. note If you change the version, the connector installs the new package and restarts. The selected version applies to all instances of the connector, including both source and sink connectors. 1. Run the following command: ``` curl -X PUT "https://api.aiven.io/v1/project//service/" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "user_config": { "plugin_versions": [{ "plugin_name": "", "version": "" }] } }' ``` Parameters: * ``: Aiven project name. * ``: Apache Kafka Connect service name. * ``: Aiven API token. * ``: Plugin name. For example, `aiven-kafka-connect-jdbc`. * ``: Desired version of the plugin. For example, `6.9.0`. 2. Verify the update in the service configuration: ``` curl -X GET "https://api.aiven.io/v1/project//service/" \ -H "Authorization: Bearer " ``` 1) Run the following command: ``` avn service update --project \ -c plugin_versions='[{"plugin_name":"","version":""}]' ``` Parameters: * ``: Apache Kafka Connect service name. * ``: Aiven project name. * ``: Plugin name. For example, `aiven-kafka-connect-jdbc`. * ``: Desired version of the plugin. For example, `6.9.0`. 2) Verify the update: ``` avn service get ``` Use [the `plugin_versions` attribute](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/kafka_connect#nestedblock--kafka_connect_user_config--plugin_versions) in your `aiven_kafka_connect` resource. ## Verify the connector version[​](#verify-the-connector-version "Direct link to Verify the connector version") After setting a version, confirm that the correct version is in use. * Aiven API 1. Run the following command: ``` curl -X GET "https://api.aiven.io/v1/project//service/" \ -H "Authorization: Bearer " ``` 2. Review the `plugin_versions` property in the response to verify the set version. --- # Manage Apache Kafka® Connect logging level During the operation of an Aiven for Apache Kafka® Connect cluster, you may encounter errors from one or more running connectors. Sometimes the stack trace printed in the logs is useful in determining the root cause of an issue, while other times the information provided just isn't enough to work with. It is possible to get access to more detailed logs to debug an issue. This can be done for a specific logger or connector by setting the logging level of an Apache Kafka® Connect cluster using the [Kafka Connect REST APIs](https://kafka.apache.org/documentation.html). warning The REST API only changes the logging level on the node that's accessed, not across an entire distributed Connect cluster. Therefore, for a multi-node cluster, you would have to change the logging level in all of the nodes that run the connector's tasks you wish to debug. For a dedicated Aiven for Apache Kafka® Connect cluster using a Startup plan, there is 1 node, while Business and Premium plans have 3 and 6 nodes respectively. ## Modify the Kafka Connect logging level[​](#modify-the-kafka-connect-logging-level "Direct link to Modify the Kafka Connect logging level") The following procedure allows you to update the logging level in an Aiven for Apache Kafka Connect service. ### Get the Kafka Connect nodes connection URI[​](#get-the-kafka-connect-nodes-connection-uri "Direct link to Get the Kafka Connect nodes connection URI") To update the logging level in all the Kafka Connect nodes, get their connection URI using the [Aiven CLI service get command](/docs/tools/cli/service-cli.md#avn_service_get) ``` avn service get SERVICE_NAME --format '{connection_info}' ``` note The above command will show the connection URI for each node in the form `https://avnadmin:PASSWORD@IP_ADDRESS:PORT` ### Retrieve the list of loggers and connectors[​](#retrieve-the-list-of-loggers-and-connectors "Direct link to Retrieve the list of loggers and connectors") You can retrieve the list of loggers, connectors and their current logging level on each worker using the dedicated `/admin/loggers` Kafka Connect API ``` curl https://avnadmin:PASSWORD@IP_ADDRESS:PORT/admin/loggers --insecure ``` tip The `--insecure` parameter avoids a `curl` non matching certificates error due to using the IP instead of the hostname The output should be similar to the following ``` { "org.apache.zookeeper": { "level": "ERROR" }, "org.reflections": { "level": "ERROR" }, "root": { "level": "INFO" } } ``` warning The previous command shows the standard list of loggers (`org.apache.zookeeper`, `org.reflections` and `root`) and any loggers for which the logging level has been already modified. This means that if you have not previously set a custom logging level for a connector's logger class, the related logger level information will not be visible in the list, even if that connector is currently running in the Kafka Connect cluster. The next section describes how to change the logging level of a particular logger. ### Change the logging level for a particular logger[​](#change-the-logging-level-for-a-particular-logger "Direct link to Change the logging level for a particular logger") To change the logging level for a particular logger you can use the same `admin/loggers` endpoint, specifying the logger name (`LOGGER_NAME` in the following command) ``` curl -X PUT -H "Content-Type:application/json" \ -d '{"level": "TRACE"}' \ https://192.168.0.1:443/admin/loggers/LOGGER_NAME \ --insecure ``` warning When the node is restarted, logging reverts back to using the logging properties defined in the `log4j` configuration file. In an Aiven for Apache Kafka® Connect cluster the default logging level is `INFO`. For example, if you set the custom log level for the logger `io.debezium.connector.postgresql.connection` to be `TRACE`, then this is what you will see upon listing the logger levels: ``` { "io.debezium.connector.postgresql.connection.AbstractMessageDecoder": { "level": "TRACE" }, "io.debezium.connector.postgresql.connection.PostgresConnection": { "level": "TRACE" }, "io.debezium.connector.postgresql.connection.PostgresDefaultValueConverter": { "level": "TRACE" }, "io.debezium.connector.postgresql.connection.PostgresReplicationConnection": { "level": "TRACE" }, "io.debezium.connector.postgresql.connection.pgproto.PgProtoMessageDecoder": { "level": "TRACE" } } ``` ### Get the connector class name[​](#get-the-connector-class-name "Direct link to Get the connector class name") Loggers are Java objects which trigger log events, and each log message produced by the application is sent to a specific logger. Loggers are arranged in hierarchies, for example the logger `io.debezium.connector.postgresql.PostgresConnector` is a child of the logger `io.debezium.connector.postgresql`. When you define the logging level of a logger using the commands above, the logging level will be set for that logger and all of its children in the logger hierarchy. By convention, loggers have the same name as the corresponding Java class. To get the name of the logger for a particular connector, use the connector's class name. The class name is usually the first field of the connector configuration when you select a connector for creation in the Aiven Console. For example, the logger for the Debezium PostgreSQL® source connector is also its class name `io.debezium.connector.postgresql`: ``` { "connector.class": "io.debezium.connector.postgresql" } ``` --- # Create a MongoDB source connector for Aiven for Apache Kafka® Use the MongoDB source connector to stream data from MongoDB collections into Apache Kafka® topics for processing and analytics. tip The MongoDB source connector uses **change streams** to capture and emit changes at defined intervals. Instead of continuously polling the collection, the connector queries the change stream at a configurable interval to detect updates. For a log-based change data capture (CDC) approach, use the [Debezium source connector for MongoDB](https://debezium.io/documentation/reference/stable/connectors/mongodb.html). ## Prerequisites[​](#connect_mongodb_pull_source_prereq "Direct link to Prerequisites") * An [Aiven for Apache Kafka® service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with Kafka Connect enabled, or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster) * A MongoDB database and collection with accessible credentials * The following MongoDB connection details: * `connection.uri`: Connection URI in the format `mongodb://USERNAME:PASSWORD@HOST:PORT` * `database`: Name of the MongoDB database * `collection`: Name of the MongoDB collection * A target Apache Kafka topic where the connector writes the data * Access to one of the following setup methods: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md#get-started) * Authentication configured for your project (for example, set the `AIVEN_API_TOKEN` environment variable if using the CLI or Terraform) tip The connector writes to a topic named `DATABASE.COLLECTION`. Create the topic in advance or enable the `auto_create_topic` parameter in your Kafka service. ## Create a MongoDB source connector configuration file[​](#create-a-mongodb-source-connector-configuration-file "Direct link to Create a MongoDB source connector configuration file") Create a file named `mongodb_source_config.json` with the following configuration: ``` { "name": "mongodb-source", "connector.class": "com.mongodb.kafka.connect.MongoSourceConnector", "connection.uri": "mongodb://USERNAME:PASSWORD@HOST:PORT", "database": "DATABASE_NAME", "collection": "COLLECTION_NAME", "poll.await.time.ms": "5000", "output.format.key": "json", "output.format.value": "json", "publish.full.document.only": "true" } ``` Parameters: * `name`: Name of the connector * `connector.class`: Class name of the MongoDB source connector * `connection.uri`: MongoDB connection URI with authentication * `database`: Name of the MongoDB database to read from * `collection`: Name of the MongoDB collection to stream from * `poll.await.time.ms`: Interval in milliseconds for polling new changes. Default is `5000` * `output.format.key` and `output.format.value`: Format for the key and value of each Kafka record. Supported values: `json`, `bson`, `schema` * `publish.full.document.only`: When `true`, only the changed document is published instead of the full change event ### Advanced options[​](#advanced-options "Direct link to Advanced options") For advanced use cases, such as schema inference, document filtering, or topic overrides, you can customize additional parameters. See the [MongoDB Kafka connector documentation](https://www.mongodb.com/docs/kafka-connector/current/) for the full list of available options. ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Console * CLI * Terraform 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Kafka Connect is enabled on the service. If not, enable Kafka Connect under **Service settings** > **Actions** > **Enable Kafka Connect**. 5. In the source connectors list, select **MongoDB source connector**, and click **Get started**. 6. In the **Common** tab, locate the **Connector configuration** text box and click **Edit**. 7. Paste the configuration from your `mongodb_source_config.json` file into the text box. 8. Click **Create connector**. 9. Verify the connector status on the **Manage stream** > **Connectors** page. 10. Verify that data appears in the target Kafka topic. By default, the connector writes to a topic named after the MongoDB database and collection, for example, `districtA.students`. To create the MongoDB source connector using the Aiven CLI, run: ``` avn service connector create SERVICE_NAME @mongodb_source_config.json ``` Replace: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka or Kafka Connect service * `@mongodb_source_config.json`: Path to your JSON configuration file You can configure this connector using the [`aiven_kafka_connector`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/kafka_connector) resource in the Aiven Provider for Terraform. Example: ``` resource "aiven_kafka_connector" "mongodb_source_connector" { project = var.project_name service_name = aiven_kafka.example_kafka.service_name connector_name = "mongodb-source-connector" config = { "name" = "mongodb-source-connector" "connector.class" = "com.mongodb.kafka.connect.MongoSourceConnector" "connection.uri" = var.mongodb_connection_uri "database" = "sample_airbnb" "collection" = "listingsAndReviews" "copy.existing" = "true" "poll.await.time.ms" = "1000" "output.format.value" = "json" "output.format.key" = "json" "publish.full.document.only" = "true" } } ``` Define variables such as `project_name` and `mongodb_connection_uri` in your Terraform configuration. ## Example: Create a MongoDB source connector[​](#example-create-a-mongodb-source-connector "Direct link to Example: Create a MongoDB source connector") The following example shows how to create a MongoDB source connector that reads data from the `students` collection in the `districtA` database and writes it to a Kafka topic named `districtA.students`. **MongoDB collection (`students`):** ``` {"name": "carlo", "age": 77} {"name": "lucy", "age": 55} {"name": "carlo", "age": 33} ``` **Connector configuration:** ``` { "name": "mongodb-source-students", "connector.class": "com.mongodb.kafka.connect.MongoSourceConnector", "connection.uri": "mongodb://USERNAME:PASSWORD@HOST:PORT", "database": "districtA", "collection": "students", "output.format.key": "json", "output.format.value": "json", "output.schema.infer.value": "true", "poll.await.time.ms": "1000" } ``` This configuration streams data from the `students` collection to the Kafka topic `districtA.students` every second, based on the polling interval (`poll.await.time.ms`). ### Verify data flow[​](#verify-data-flow "Direct link to Verify data flow") After you create the connector: 1. Check the connector status on the **Manage stream** > **Connectors** page in the Aiven Console. 2. Confirm that the Kafka topic `districtA.students` exists in your service. 3. Consume messages from the topic to verify that data from MongoDB is streaming correctly. Related pages * [MongoDB sink connector for Aiven for Apache Kafka®](/docs/products/kafka/kafka-connect/howto/mongodb-sink-mongo.md) * [MongoDB sink connector (Lenses.io) for Aiven for Apache Kafka®](/docs/products/kafka/kafka-connect/howto/mongodb-sink-lenses.md) --- # Create a Lenses.io MongoDB sink connector for Aiven for Apache Kafka® Use the MongoDB sink connector by Lenses.io to write data from Apache Kafka® topics into a MongoDB database. This connector supports [KCQL transformations](https://docs.lenses.io/5.0/integrations/connectors/stream-reactor/sinks/mongosinkconnector/) to filter and map topic data before inserting it into MongoDB. Aiven supports two MongoDB sink connectors with different capabilities: * [MongoDB sink connector by MongoDB](https://www.mongodb.com/docs/kafka-connector/current/) * [MongoDB sink connector by Lenses.io](https://docs.lenses.io/connectors/sink/mongo) caution **Version compatibility** Stream Reactor version 9.0.2 includes class and package name updates introduced in version 6.0.0 by Lenses to standardize connector and converter names. Version 9.x is not compatible with version 4.2.0. To continue using version 4.2.0, [set the connector version](/docs/products/kafka/kafka-connect/howto/manage-connector-versions.md#set-version) before you upgrade. If you upgrade from version 4.2.0, recreate the connector using the updated class name. For example: ``` "connector.class": "io.lenses.streamreactor.connect.." ``` For details about these changes, see the [Stream Reactor release notes](https://docs.lenses.io/stream-reactor/docs/releases/). ## Prerequisites[​](#connect_mongodb_lenses_sink_prereq "Direct link to Prerequisites") * An [Aiven for Apache Kafka® service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with Kafka Connect enabled, or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster) * A MongoDB database with accessible connection credentials * The following MongoDB and Kafka connection details: * `MONGODB_USERNAME`: MongoDB username * `MONGODB_PASSWORD`: MongoDB password * `MONGODB_HOST`: MongoDB hostname * `MONGODB_PORT`: MongoDB port * `MONGODB_DATABASE_NAME`: MongoDB database name * `TOPIC_LIST`: Comma-separated list of Kafka topics to sink * `KCQL_TRANSFORMATION`: KCQL mapping statement. For example: ``` INSERT INTO MONGODB_COLLECTION SELECT * FROM APACHE_KAFKA_TOPIC ``` * Schema Registry details (required only for Avro format): * `APACHE_KAFKA_HOST`: Kafka host * `SCHEMA_REGISTRY_PORT`: Schema Registry port * `SCHEMA_REGISTRY_USER`: Schema Registry username * `SCHEMA_REGISTRY_PASSWORD`: Schema Registry password * Access to one of the following setup methods: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md#get-started) * Authentication configured for your Aiven project (for example, set the `AIVEN_API_TOKEN` environment variable if using CLI or Terraform) tip Aiven for Apache Kafka® uses [Karapace](https://github.com/aiven/karapace) as the built-in Schema Registry. Find connection details in the Aiven Console under **Overview > Connection information > Schema Registry**, or using the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get) command `avn service get`. ## Create a MongoDB sink connector configuration file[​](#create-a-mongodb-sink-connector-configuration-file "Direct link to Create a MongoDB sink connector configuration file") Create a file named `mongodb_sink.json` with the following configuration: ``` { "name": "mongodb-sink", "connector.class": "com.datamountaineer.streamreactor.connect.mongodb.sink.MongoSinkConnector", "topics": "TOPIC_LIST", "connect.mongo.connection": "mongodb://MONGODB_USERNAME:MONGODB_PASSWORD@MONGODB_HOST:MONGODB_PORT", "connect.mongo.db": "MONGODB_DATABASE_NAME", "connect.mongo.kcql": "KCQL_TRANSFORMATION", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } ``` Parameters: * `name`: Connector name * `connector.class`: Class name of the MongoDB sink connector * `connect.mongo.connection`: MongoDB connection URI with credentials * `connect.mongo.db`: MongoDB database name * `connect.mongo.kcql`: KCQL statement defining how Kafka topic data maps to MongoDB * `key.converter` and `value.converter`: Configure the data format and schema registry details * `topics`: Kafka topics to sink data from note The `key.converter` and `value.converter` fields define how Kafka messages are parsed and must be included in the configuration. When using Avro as the source format, configure the following parameters: * `value.converter.schema.registry.url`: Schema Registry URL in the format `https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT`. Retrieve these values from the [prerequisites](#connect_mongodb_lenses_sink_prereq). * `value.converter.basic.auth.credentials.source`: Set to `USER_INFO` to authenticate with a username and password. * `value.converter.schema.registry.basic.auth.user.info`: Schema Registry credentials in the format `SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD`. Retrieve these values from the [prerequisites](#connect_mongodb_lenses_sink_prereq). For advanced configurations such as batch size, write strategy, or custom transformations, see the [Lenses.io MongoDB sink connector reference](https://docs.lenses.io/5.0/integrations/connectors/stream-reactor/sinks/mongosinkconnector/). ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Console * CLI 1. Go to the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Kafka Connect is enabled on the service. If not, enable it under **Service settings** > **Actions** > **Enable Kafka Connect**. 5. From the list of sink connectors, select **Stream Reactor MongoDB Sink**, and click **Get started**. 6. In the **Common** tab, find the **Connector configuration** text box and click **Edit**. 7. Paste the configuration from your `mongodb_sink.json` file. 8. Click **Create connector**. 9. Verify the connector status on the **Manage stream** > **Connectors** page. 10. Confirm that data appears in the target MongoDB collection. The collection name corresponds to the KCQL mapping. To create the connector using the Aiven CLI, run: ``` avn service connector create SERVICE_NAME @mongodb_sink.json ``` Replace: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka or Kafka Connect service * `@mongodb_sink.json`: Path to your configuration file ## Example: Sink data to MongoDB[​](#example-sink-data-to-mongodb "Direct link to Example: Sink data to MongoDB") The following examples show how to write Kafka topic data to MongoDB collections using KCQL transformations. ### Insert mode[​](#insert-mode "Direct link to Insert mode") If the Kafka topic `students` contains: ``` {"name": "carlo", "age": 77} {"name": "lucy", "age": 55} {"name": "carlo", "age": 33} ``` Use this configuration to insert all records into a MongoDB collection named `studentscol`: ``` { "name": "mongodb-sink-insert", "connector.class": "com.datamountaineer.streamreactor.connect.mongodb.sink.MongoSinkConnector", "topics": "students", "connect.mongo.connection": "mongodb://MONGODB_USERNAME:MONGODB_PASSWORD@MONGODB_HOST:MONGODB_PORT", "connect.mongo.db": "MONGODB_DATABASE_NAME", "connect.mongo.kcql": "INSERT INTO studentscol SELECT * FROM students", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter.schemas.enable": "false" } ``` ### Upsert mode[​](#upsert-mode "Direct link to Upsert mode") To ensure only one document per unique name is stored, use upsert mode: ``` { "name": "mongodb-sink-upsert", "connector.class": "com.datamountaineer.streamreactor.connect.mongodb.sink.MongoSinkConnector", "topics": "students", "connect.mongo.connection": "mongodb://MONGODB_USERNAME:MONGODB_PASSWORD@MONGODB_HOST:MONGODB_PORT", "connect.mongo.db": "MONGODB_DATABASE_NAME", "connect.mongo.kcql": "UPSERT INTO studentscol SELECT * FROM students PK name", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter.schemas.enable": "false" } ``` This configuration updates existing records based on the `name` field instead of inserting duplicates. After the connector runs, the MongoDB collection `studentscol` contains: ``` {"name": "lucy", "age": 55} {"name": "carlo", "age": 33} ``` Related pages * [MongoDB sink connector for Aiven for Apache Kafka®](/docs/products/kafka/kafka-connect/howto/mongodb-sink-mongo.md) * [MongoDB source connector for Aiven for Apache Kafka®](/docs/products/kafka/kafka-connect/howto/mongodb-poll-source-connector.md) --- # Create a MongoDB sink connector for Aiven for Apache Kafka® Use the MongoDB sink connector to move data from an Aiven for Apache Kafka® service to a MongoDB database. Aiven provides two MongoDB sink connectors with different capabilities: * **MongoDB Kafka Sink Connector (by MongoDB):** The standard connector maintained by MongoDB. * **MongoDB Sink Connector (by Lenses.io):** Supports KCQL transformations for topic data before writing to MongoDB. For information about the Lenses.io connector, see [MongoDB sink connector (Lenses.io)](/docs/products/kafka/kafka-connect/howto/mongodb-sink-lenses.md). For detailed configuration parameters, refer to the [MongoDB Kafka connector documentation](https://www.mongodb.com/docs/kafka-connector/current/). ## Prerequisites[​](#connect_mongodb_sink_prereq "Direct link to Prerequisites") * An [Aiven for Apache Kafka® service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with Kafka Connect enabled, or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster) * A running MongoDB database with valid credentials and network access from the Kafka Connect service * The following MongoDB connection details: * `MONGODB_USERNAME`: Database username * `MONGODB_PASSWORD`: Database password * `MONGODB_HOST`: MongoDB host name * `MONGODB_PORT`: MongoDB port * `MONGODB_DATABASE_NAME`: Target database name * A source Kafka topic with data to write to MongoDB * Access to one of the following setup methods: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md#get-started) * Authentication configured for your project (for example, set the `AIVEN_API_TOKEN` environment variable if using the CLI or Terraform) ### Additional details for Avro data format[​](#additional-details-for-avro-data-format "Direct link to Additional details for Avro data format") If you use Avro serialization, collect the following Schema Registry details: * `SCHEMA_REGISTRY_URL`: Schema Registry URL, for example `https://HOST:PORT` * `SCHEMA_REGISTRY_PORT`: Schema Registry port * `SCHEMA_REGISTRY_USER`: Schema Registry username * `SCHEMA_REGISTRY_PASSWORD`: Schema Registry password tip Aiven for Apache Kafka® includes [Karapace](https://github.com/aiven/karapace) as the built-in Schema Registry. Find connection details in the Aiven Console under ****Overview** > Connection information > Schema Registry**, or using the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get) command `avn service get`. ## Create a MongoDB sink connector configuration file[​](#create-a-mongodb-sink-connector-configuration-file "Direct link to Create a MongoDB sink connector configuration file") Create a file named `mongodb_sink_config.json` with the following configuration: ``` { "name": "mongodb-sink", "connector.class": "com.mongodb.kafka.connect.MongoSinkConnector", "topics": "students", "connection.uri": "mongodb://USERNAME:PASSWORD@HOST:PORT", "database": "school", "collection": "students", "tasks.max": "1", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://SCHEMA_REGISTRY_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://SCHEMA_REGISTRY_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } ``` Parameters: * `name`: Name of the connector * `connector.class`: Class name of the MongoDB sink connector * `topics`: Comma-separated list of Kafka topics to sink * `connection.uri`: MongoDB connection URI * `database`: Target MongoDB database name * `collection`: Target MongoDB collection name. If not specified, the connector uses the topic name by default * `tasks.max`: Maximum number of parallel tasks for writing data * `key.converter` and `value.converter`: Define the serialization format for records (Avro, JSON, or others supported by your Kafka service) * `key.converter.schema.registry.url` and `value.converter.schema.registry.url`: URL of the Schema Registry * `key.converter.basic.auth.credentials.source` and `value.converter.basic.auth.credentials.source`: Method used to supply Schema Registry credentials * `key.converter.schema.registry.basic.auth.user.info` and `value.converter.schema.registry.basic.auth.user.info`: Schema Registry credentials in the format `username:password` For advanced configuration options, including batch size, document mapping, and topic management, refer to the [MongoDB Kafka connector documentation](https://www.mongodb.com/docs/kafka-connector/current/). ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Console * CLI * Terraform 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Kafka Connect is enabled on the service. If not, enable it under **Service settings** > **Actions** > **Enable Kafka Connect**. 5. From the list of sink connectors, select **MongoDB sink connector**, and click **Get started**. 6. In the **Common** tab, locate the **Connector configuration** text box and click **Edit**. 7. Paste the configuration from your `mongodb_sink_config.json` file into the text box. 8. Click **Create connector**. 9. Verify the connector status on the **Manage stream** > **Connectors** page. 10. Verify that data from the Kafka topic appears in MongoDB. To create the MongoDB sink connector using the Aiven CLI, run: ``` avn service connector create SERVICE_NAME @mongodb_sink_config.json ``` Replace: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka or Kafka Connect service * `@mongodb_sink_config.json`: Path to your JSON configuration file You can configure this connector using the [`aiven_kafka_connector`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/kafka_connector) resource in the Aiven Provider for Terraform. For configuration examples, see the [MongoDB sink connector Terraform example](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/kafka/kafka_connectors/mongo_sink). ## Example: Create a MongoDB sink connector[​](#example-create-a-mongodb-sink-connector "Direct link to Example: Create a MongoDB sink connector") The following example creates a MongoDB sink connector that writes data from the Kafka topic `students` to a MongoDB database named `school`. **Kafka topic (`students`):** ``` key: 1 value: {"name": "carlo"} key: 2 value: {"name": "lucy"} key: 3 value: {"name": "mary"} ``` **Connector configuration:** ``` { "name": "mongodb-sink", "connector.class": "com.mongodb.kafka.connect.MongoSinkConnector", "topics": "students", "connection.uri": "mongodb://USERNAME:PASSWORD@HOST:PORT", "database": "school", "tasks.max": "1" } ``` This configuration writes records from the Kafka topic `students` to a collection named `students` in the MongoDB database `school`. ### Verify data flow[​](#verify-data-flow "Direct link to Verify data flow") After creating the connector: 1. Check the connector status on the **Manage stream** > **Connectors** page in the Aiven Console. 2. Verify that a `students` collection exists in the MongoDB database. 3. Confirm that records from the Kafka topic are written to the collection. Related pages * [MongoDB sink connector (Lenses.io) for Aiven for Apache Kafka®](/docs/products/kafka/kafka-connect/howto/mongodb-sink-lenses.md) * [MongoDB source connector for Aiven for Apache Kafka®](/docs/products/kafka/kafka-connect/howto/mongodb-poll-source-connector.md) --- # Create an MQTT sink connector The [MQTT sink connector](https://docs.lenses.io/connectors/kafka-connectors/sources/mqtt) copies messages from an Apache Kafka® topic to an MQTT queue. tip You can use this connector to send messages to RabbitMQ® when the [RabbitMQ MQTT plugin](https://www.rabbitmq.com/mqtt.html) is enabled. caution **Version compatibility** Stream Reactor version 9.0.2 includes class and package name updates introduced in version 6.0.0 by Lenses to standardize connector and converter names. Version 9.x is not compatible with version 4.2.0. To continue using version 4.2.0, [set the connector version](/docs/products/kafka/kafka-connect/howto/manage-connector-versions.md#set-version) before you upgrade. If you upgrade from version 4.2.0, recreate the connector using the updated class name. For example: ``` "connector.class": "io.lenses.streamreactor.connect.." ``` For details about these changes, see the [Stream Reactor release notes](https://docs.lenses.io/stream-reactor/docs/releases/). ## Prerequisites[​](#connect_mqtt_rbmq_sink_prereq "Direct link to Prerequisites") * An Aiven for Apache Kafka service [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * Gather the following information for the target MQTT server: * `USERNAME`: The MQTT username. * `PASSWORD`: The MQTT password. * `HOST`: The MQTT hostname. * `PORT`: The MQTT port (typically `1883`). * `KCQL_STATEMENT`: A KCQL statement that maps topic data to the MQTT topic. Use the following format: ``` INSERT INTO MQTT_TOPIC SELECT LIST_OF_FIELDS FROM APACHE_KAFKA_TOPIC ``` * `APACHE_KAFKA_HOST`: The Apache Kafka host. Required only when using Avro. * `SCHEMA_REGISTRY_PORT`: The schema registry port. Required only when using Avro. * `SCHEMA_REGISTRY_USER`: The schema registry username. Required only when using Avro. * `SCHEMA_REGISTRY_PASSWORD`: The schema registry password. Required only when using Avro. note The connector writes to the Kafka topic defined in the `connect.mqtt.kcql` parameter. Either create the topic manually or enable the `auto_create_topic` parameter to allow automatic topic creation. For a complete list of parameters and configuration options, see the [connector documentation](https://docs.lenses.io/connectors/kafka-connectors/sources/mqtt). ## Create the connector configuration file[​](#create-the-connector-configuration-file "Direct link to Create the connector configuration file") Create a file named `mqtt_sink.json` and add the following configuration: ``` { "name": "CONNECTOR_NAME", "connector.class": "com.datamountaineer.streamreactor.connect.mqtt.sink.MqttSinkConnector", "connect.mqtt.hosts": "tcp://HOST:PORT", "connect.mqtt.kcql": "KCQL_STATEMENT", "connect.mqtt.username": "USERNAME", "connect.mqtt.password": "PASSWORD", "key.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter" } ``` Parameters: * `name`: The connector name. Replace `CONNECTOR_NAME` with your desired name. * `connect.mqtt.*`: MQTT server connection parameters collected in the [prerequisite step](/docs/products/kafka/kafka-connect/howto/mqtt-sink-connector.md#connect_mqtt_rbmq_sink_prereq). * `key.converter` and `value.converter`: Define the message data format in the Kafka topic. This example uses `JsonConverter` for both key and value. For a full list of supported parameters, see the [Stream Reactor MQTT sink documentation](https://docs.lenses.io/connectors/kafka-connectors/sources/mqtt#storage-to-output-matrix). ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Console * CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 5. In the sink connectors list, select **Stream Reactor MQTT Sink Connector**, and click **Get started**. 6. On the **Stream Reactor MQTT Sink** page, go to the **Common** tab. 7. Locate the **Connector configuration** text box and click **Edit**. 8. Paste the configuration from your `mqtt_sink.json` file into the text box. 9. Click **Create connector**. 10. Verify the connector status on the **Manage stream** > **Connectors** page. 11. Confirm that data is delivered to the MQTT topic defined in the `KCQL_STATEMENT`. To create the connector using the [Aiven CLI](/docs/tools/cli/service/connector.md#avn_service_connector_create), run: ``` avn service connector create SERVICE_NAME @mqtt_sink.json ``` Replace: * `SERVICE_NAME`: Your Kafka or Kafka Connect service name. * `@mqtt_sink.json`: Path to your connector configuration file. ## Sink topic data to an MQTT topic[​](#sink-topic-data-to-an-mqtt-topic "Direct link to Sink topic data to an MQTT topic") The following example shows how to sink data from a Kafka topic to an MQTT topic. If your Kafka topic `sensor_data` contains the following messages: ``` {"device":"sensor-1", "temperature": 22.5} {"device":"sensor-2", "temperature": 19.0} {"device":"sensor-1", "temperature": 23.1} ``` To write this data to an MQTT topic named `iot/devices/temperature`, use the following connector configuration: ``` { "name": "my-mqtt-sink", "connector.class": "com.datamountaineer.streamreactor.connect.mqtt.sink.MqttSinkConnector", "topics": "sensor_data", "connect.mqtt.hosts": "tcp://MQTT_HOST:MQTT_PORT", "connect.mqtt.username": "MQTT_USERNAME", "connect.mqtt.password": "MQTT_PASSWORD", "connect.mqtt.kcql": "INSERT INTO iot/devices/temperature SELECT * FROM sensor_data", "key.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter.schemas.enable": "false" } ``` Replace all placeholder values (such as `MQTT_HOST`, `MQTT_PORT`, and `MQTT_USERNAME`) with your actual MQTT broker connection details. This configuration does the following: * `"topics": "sensor_data"`: Specifies the Kafka topic to sink. * `connect.mqtt.*`: Sets the MQTT broker hostname, port, credentials, and KCQL rules. * `"value.converter"` and `"value.converter.schemas.enable"`: Set the message format. This example uses raw JSON without a schema. * `"connect.mqtt.kcql"`: Defines the KCQL transformation. Each Kafka message is published to the MQTT topic `iot/devices/temperature`. After creating the connector, check your MQTT broker to verify that messages are published to the `iot/devices/temperature` topic. --- # Create a source connector from MQTT to Apache Kafka® The [Stream Reactor MQTT source connector](https://docs.lenses.io/5.0/integrations/connectors/stream-reactor/sources/mqttsourceconnector/) transfers messages from an MQTT topic to an Aiven for Apache Kafka® topic, where they can be processed and consumed by multiple applications. It creates a queue and binds it to the `amq.topic` exchange specified in the KCQL statement, then forwards the messages to Kafka. tip You can use this connector to source messages from RabbitMQ® if the [RabbitMQ MQTT plugin](https://www.rabbitmq.com/mqtt.html) is enabled. caution **Version compatibility** Stream Reactor version 9.0.2 includes class and package name updates introduced in version 6.0.0 by Lenses to standardize connector and converter names. Version 9.x is not compatible with version 4.2.0. To continue using version 4.2.0, [set the connector version](/docs/products/kafka/kafka-connect/howto/manage-connector-versions.md#set-version) before you upgrade. If you upgrade from version 4.2.0, recreate the connector using the updated class name. For example: ``` "connector.class": "io.lenses.streamreactor.connect.." ``` For details about these changes, see the [Stream Reactor release notes](https://docs.lenses.io/stream-reactor/docs/releases/). ## Prerequisites[​](#connect_mqtt_rbmq_source_prereq "Direct link to Prerequisites") * An Aiven for Apache Kafka service [with Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * Collect the following details for your MQTT source: * `HOST`: The MQTT hostname. * `PORT`: The MQTT port (usually `1883`). * `USERNAME`: The MQTT username. * `PASSWORD`: The MQTT password. * `KCQL_STATEMENT`: A KCQL mapping from MQTT to Kafka in the format: `INSERT INTO SELECT * FROM ` * `APACHE_KAFKA_HOST`: The Kafka host (required only when using Avro). * `SCHEMA_REGISTRY_PORT`: The schema registry port (required only when using Avro). * `SCHEMA_REGISTRY_USER`: The schema registry username (required only when using Avro). * `SCHEMA_REGISTRY_PASSWORD`: The schema registry password (required only when using Avro). tip The connector writes to a Kafka topic defined in the `connect.mqtt.kcql` parameter. Make sure the topic exists, or enable [automatic topic creation](/docs/products/kafka/howto/create-topics-automatically.md). ## Create a connector configuration file[​](#create-a-connector-configuration-file "Direct link to Create a connector configuration file") Create a file named `mqtt_source.json` and add the following configuration: ``` { "name": "CONNECTOR_NAME", "connector.class": "com.datamountaineer.streamreactor.connect.mqtt.source.MqttSourceConnector", "connect.mqtt.hosts": "tcp://HOST:PORT", "connect.mqtt.kcql": "KCQL_STATEMENT", "connect.mqtt.username": "USERNAME", "connect.mqtt.password": "PASSWORD", "connect.mqtt.service.quality": "1", "key.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter" } ``` Parameters: * `name`: The connector name. Replace `CONNECTOR_NAME` with your desired name. * `connect.mqtt.*`: MQTT connection details collected in the [prerequisite step](#connect_mqtt_rbmq_source_prereq). * `connect.mqtt.kcql`: A KCQL statement that maps MQTT topics to Kafka topics. * `connect.mqtt.service.quality`: Sets the MQTT Quality of Service (QoS). Common values are `0`, `1`, or `2`. * `key.converter` and `value.converter`: Set the message format. This example uses raw JSON. ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Console * CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector**. If Kafka Connect is not yet enabled, click **Enable connector on this service**. Alternatively: * Go to **Service settings**. * In the **Service management** section, click **Actions** > **Enable Kafka connect**. 5. From the source connectors list, select **Stream Reactor MQTT Source**, then click **Get started**. 6. On the **Common** tab, locate the **Connector configuration** text box and click **Edit**. 7. Paste the contents of your `mqtt_source.json` file. 8. Click **Create connector**. 9. Verify the connector status on the **Manage stream** > **Connectors** page. To create the connector using the [Aiven CLI](/docs/tools/cli/service/connector.md#avn_service_connector_create), run: ``` avn service connector create SERVICE_NAME @mqtt_source.json ``` Replace: * `SERVICE_NAME`: Your Kafka or Kafka Connect service name. * `@mqtt_source.json`: Path to your configuration file. ## Source MQTT data to a Kafka topic[​](#source-mqtt-data-to-a-kafka-topic "Direct link to Source MQTT data to a Kafka topic") The following example shows how to forward messages from an MQTT topic to a Kafka topic. Suppose your MQTT broker publishes JSON messages to a topic named `devices/status` like this: ``` {"device": "alpha", "status": "online"} {"device": "beta", "status": "offline"} ``` To send this data to a Kafka topic named `device_status`, use the following connector configuration: ``` { "name": "mqtt-source-example", "connector.class": "com.datamountaineer.streamreactor.connect.mqtt.source.MqttSourceConnector", "connect.mqtt.hosts": "tcp://mqtt-broker.example.com:1883", "connect.mqtt.kcql": "INSERT INTO device_status SELECT * FROM devices/status", "connect.mqtt.username": "mqtt_user", "connect.mqtt.password": "mqtt_password", "connect.mqtt.service.quality": "1", "key.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter" } ``` Replace all placeholder values (such as `mqtt-broker.example.com`, `mqtt_user`, and `mqtt_password`) with your actual connection details. This configuration does the following: * `connect.mqtt.kcql`: Routes messages from the MQTT topic `devices/status` to the Kafka topic `device_status`. * `connect.mqtt.service.quality`: Uses QoS level `1` for at-least-once delivery. * `key.converter` and `value.converter`: Set the message format to JSON. * `connect.mqtt.*`: Supplies the MQTT connection details. After the connector is running, check the `device_status` Kafka topic to confirm that messages are being delivered. --- # Create a sink connector from Apache Kafka® to OpenSearch® The OpenSearch sink connector writes data from Aiven for Apache Kafka® to OpenSearch®. ## Prerequisites[​](#connect_opensearch_sink_prereq "Direct link to Prerequisites") To set up an OpenSearch sink connector, you need an Aiven for Apache Kafka service [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). Collect the following information about the target OpenSearch service: * `OS_CONNECTION_URL`: The OpenSearch connection URL, in the form of `https://HOST:PORT` * `OS_USERNAME`: The OpenSearch username to connect * `OS_PASSWORD`: The password for the username selected * `TOPIC_LIST`: The comma-separated list of topics to sink If the source data is in Avro format, also collect the following information: * `SCHEMA_REGISTRY_HOST`: Host from the **Schema Registry** tab. This is the same hostname as the Kafka service. * `SCHEMA_REGISTRY_PORT`: Port from the **Schema Registry** tab. * `SCHEMA_REGISTRY_USER`: User from the **Schema Registry** tab. * `SCHEMA_REGISTRY_PASSWORD`: Password from the **Schema Registry** tab. note For Aiven for OpenSearch® and Aiven for Apache Kafka®, find these values on the service **Overview** page in the [Aiven Console](https://console.aiven.io/). On Aiven for Apache Kafka®, Schema Registry connection details are in the **Schema Registry** tab. You can also run `avn service get` with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). As of version 3.0, Aiven for Apache Kafka no longer supports Confluent Schema Registry. For more information, see [Karapace](/docs/products/kafka/karapace.md). ## Setup an OpenSearch sink connector with Aiven Console[​](#setup-an-opensearch-sink-connector-with-aiven-console "Direct link to Setup an OpenSearch sink connector with Aiven Console") The following example demonstrates how to setup a OpenSearch sink connector for Apache Kafka using the [Aiven Console](https://console.aiven.io/). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Define the connector configuration in a file named `opensearch_sink.json` with the following content: ``` { "name":"CONNECTOR_NAME", "connector.class": "io.aiven.kafka.connect.opensearch.OpensearchSinkConnector", "topics": "TOPIC_LIST", "connection.url": "OS_CONNECTION_URL", "connection.username": "OS_USERNAME", "connection.password": "OS_PASSWORD", "type.name": "TYPE_NAME", "tasks.max":"1", "key.ignore": "true", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://SCHEMA_REGISTRY_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://SCHEMA_REGISTRY_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } ``` The configuration file contains the following entries: * `name`: Name of the connector. * `connection.url`, `connection.username`, `connection.password`: OpenSearch connection values from the [prerequisites](#connect_opensearch_sink_prereq). * `type.name`: OpenSearch type name the connector uses when indexing. * `key.ignore`: If `true`, the connector ignores the message key and sets the document ID to `topic+partition+offset`. Otherwise it uses the message key. * `tasks.max`: Maximum number of tasks to run in parallel. The default is `1`. * `key.converter` and `value.converter`: Define the message data format in the Apache Kafka topic. For Avro, use `io.confluent.connect.avro.AvroConverter`. * `existing.resource.type` and `topic.to.existing.resource.mapping`: Optional properties that send records to an existing OpenSearch resource, such as an index alias, instead of an index named after the topic. Add them before you create the connector. For details, see [Write to an existing OpenSearch resource](#connect_opensearch_sink_existing_resource). note Include the `key.converter` and `value.converter` sections only when the source data is in Avro format. If you omit them, Kafka Connect reads the messages as binary. For Avro, the converter retrieves the schema from [Karapace](https://github.com/aiven/karapace). When the source data is Avro, set the following parameters: * `value.converter.schema.registry.url`: Schema Registry URL in the form `https://SCHEMA_REGISTRY_HOST:SCHEMA_REGISTRY_PORT`. Use the values from the [prerequisites](#connect_opensearch_sink_prereq). * `value.converter.basic.auth.credentials.source`: Set to `USER_INFO` to log in with a username and password. * `value.converter.schema.registry.basic.auth.user.info`: Schema Registry credentials in the form `SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD`. Use the values from the [prerequisites](#connect_opensearch_sink_prereq). note For the full set of connector parameters, see the [OpenSearch sink connector configuration options](https://github.com/aiven/opensearch-connector-for-apache-kafka/blob/main/docs/opensearch-sink-connector-config-options.rst). ### Create a Kafka Connect connector with the Aiven Console[​](#create-a-kafka-connect-connector-with-the-aiven-console "Direct link to Create a Kafka Connect connector with the Aiven Console") To create a Kafka Connect connector: 1. Log in to the [Aiven Console](https://console.aiven.io/) and select the Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service where the connector needs to be defined. 2. Click **Manage stream** > **Connectors** from the left sidebar. 3. Select **Create New Connector**, it is enabled only for services [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md). 4. Select **OpenSearch sink**. 5. In the **Common** tab, locate the **Connector configuration** text box and click on **Edit**. 6. Paste the connector configuration (stored in the `opensearch_sink.json` file) in the form. 7. Click **Apply**. note The Aiven Console parses the configuration file and fills the relevant UI fields. You can review the UI fields across the various tab and change them if necessary. The changes will be reflected in JSON format in the **Connector configuration** text box. 8. After all the settings are correctly configured, click **Create connector**. 9. Verify the connector status under **Manage stream** > **Connectors**. 10. Verify that the data is available in the target OpenSearch resource. By default, the connector writes to an index based on the Apache Kafka topic name. In your Aiven for OpenSearch service, select **Indexes** to view the index. If the configuration maps the topic to an existing resource, verify that resource instead. See [Write to an existing OpenSearch resource](#connect_opensearch_sink_existing_resource). note You can also create connectors using the [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create). ## Write to an existing OpenSearch resource[​](#connect_opensearch_sink_existing_resource "Direct link to Write to an existing OpenSearch resource") From connector version 3.2.0, you can write to an existing OpenSearch resource. The connector does not create an index from the Kafka topic name. Writing to an existing resource is useful when you manage index rotation with a write alias. You rotate daily or weekly indices behind a stable alias. Create the target resource in OpenSearch before you start the connector. Then set the following properties in the connector configuration: * `existing.resource.type`: The type of existing OpenSearch resource. To write to an index alias, set this to `index_alias`. * `topic.to.existing.resource.mapping`: Maps Kafka topics to existing OpenSearch resources. Use the format `topic_name:resource_name`. Separate multiple mappings with commas. ``` { "connection.url": "OS_CONNECTION_URL", "connection.username": "OS_USERNAME", "connection.password": "OS_PASSWORD", "existing.resource.type": "index_alias", "topic.to.existing.resource.mapping": "orders:orders_write_alias", "key.ignore": "true", "schema.ignore": "true" } ``` This example maps the `orders` topic to the `orders_write_alias` alias. For other `existing.resource.type` values, see the [connector configuration options](https://github.com/aiven/opensearch-connector-for-apache-kafka/blob/main/docs/opensearch-sink-connector-config-options.rst). ## Create daily OpenSearch indices[​](#create-daily-opensearch-indices "Direct link to Create daily OpenSearch indices") To write through a stable write alias, use an existing resource mapping. To include the message date in the index name, use the `TimestampRouter` transformation and create the indices first. To store the Apache Kafka messages in a daily OpenSearch index, add the following `TimestampRouter` transformation to the connector properties file. The transformation defines the index name as the topic name followed by the message date. ``` "transforms": "TimestampRouter", "transforms.TimestampRouter.topic.format": "${topic}-${timestamp}", "transforms.TimestampRouter.timestamp.format": "yyyy-MM-dd", "transforms.TimestampRouter.type": "org.apache.kafka.connect.transforms.TimestampRouter" ``` warning The current version of the OpenSearch sink connector is not able to automatically create daily indices in OpenSearch. Therefore create the indices with the correct name before starting the sink connector. You can create OpenSearch indices in many ways including [CURL commands](/docs/products/opensearch/howto/opensearch-with-curl.md). ## Create a sink connector for JSON with a schema[​](#create-a-sink-connector-for-json-with-a-schema "Direct link to Create a sink connector for JSON with a schema") If you have a topic named `iot_measurements` that contains the following JSON, including an embedded schema: ``` { "schema": { "type":"struct", "fields":[{ "type":"int64", "optional": false, "field": "iot_id" },{ "type":"string", "optional": false, "field": "metric" },{ "type":"int32", "optional": false, "field": "measurement" }] }, "payload":{ "iot_id":1, "metric":"Temperature", "measurement":14} } { "schema": { "type":"struct", "fields":[{ "type":"int64", "optional": false, "field": "iot_id" },{ "type":"string", "optional": false, "field": "metric" },{ "type":"int32", "optional": false, "field": "measurement" }] }, "payload":{"iot_id":2, "metric":"Humidity", "measurement":60}} } ``` note Each message includes the JSON schema, which increases payload size. For a smaller payload, use Avro with [Karapace](/docs/products/kafka/karapace.md). You can sink the `iot_measurements` topic to OpenSearch with the following connector configuration, after replacing `OS_CONNECTION_URL`, `OS_USERNAME`, and `OS_PASSWORD`: ``` { "name":"sink_iot_json_schema", "connector.class": "io.aiven.kafka.connect.opensearch.OpensearchSinkConnector", "topics": "iot_measurements", "connection.url": "OS_CONNECTION_URL", "connection.username": "OS_USERNAME", "connection.password": "OS_PASSWORD", "type.name": "iot_measurements", "tasks.max":"1", "key.ignore": "true", "value.converter": "org.apache.kafka.connect.json.JsonConverter" } ``` The configuration file contains the following entries: * `topics`: Set to `iot_measurements`. * `value.converter`: JSON converter. The sample messages include an embedded schema, so leave schema support enabled. * `key.ignore`: If `true`, the connector ignores the empty message key and sets the document ID to `topic+partition+offset`. ## Create a sink connector for schemaless JSON[​](#create-a-sink-connector-for-schemaless-json "Direct link to Create a sink connector for schemaless JSON") If you have a topic named `students` that contains the following schemaless JSON: ``` Key: 1 Value: {"student_id":1, "student_name":"Carla"} Key: 2 Value: {"student_id":2, "student_name":"Ugo"} Key: 3 Value: {"student_id":3, "student_name":"Mary"} ``` You can sink the `students` topic to OpenSearch with the following connector configuration, after replacing `OS_CONNECTION_URL`, `OS_USERNAME`, and `OS_PASSWORD`: ``` { "name":"sink_students_json", "connector.class": "io.aiven.kafka.connect.opensearch.OpensearchSinkConnector", "topics": "students", "connection.url": "OS_CONNECTION_URL", "connection.username": "OS_USERNAME", "connection.password": "OS_PASSWORD", "type.name": "students", "tasks.max":"1", "key.converter": "org.apache.kafka.connect.storage.StringConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter.schemas.enable": "false", "schema.ignore": "true" } ``` The configuration file contains the following entries: * `topics`: Set to `students`. * `key.converter`: String converter for the message key. * `value.converter`: JSON converter for the message value. * `value.converter.schemas.enable`: Set to `false` when the value has no schema so the connector does not read a schema. * `schema.ignore`: Set to `true` so the connector does not infer a schema before it writes to OpenSearch. note The connector sets the OpenSearch document ID to the message key. --- # Create a Stream Reactor sink connector from Apache Kafka® to Redis®\* The Redis®\* Stream Reactor sink connector enables you to move data from an Aiven for Apache Kafka® cluster to a Redis®\* database. It uses [KCQL transformations](https://docs.lenses.io/connectors/sink/redis) to filter and map topic data before sending it to Redis. important In version 4.2.0 of the Redis Stream Reactor sink connector, a known issue with the `GEOADD` command may cause exceptions during initialization under specific configurations. For more information, see the [GitHub issue](https://github.com/lensesio/stream-reactor/issues/990). caution **Version compatibility** Stream Reactor version 9.0.2 includes class and package name updates introduced in version 6.0.0 by Lenses to standardize connector and converter names. Version 9.x is not compatible with version 4.2.0. To continue using version 4.2.0, [set the connector version](/docs/products/kafka/kafka-connect/howto/manage-connector-versions.md#set-version) before you upgrade. If you upgrade from version 4.2.0, recreate the connector using the updated class name. For example: ``` "connector.class": "io.lenses.streamreactor.connect.." ``` For details about these changes, see the [Stream Reactor release notes](https://docs.lenses.io/stream-reactor/docs/releases/). ## Prerequisites[​](#connect_redis_lenses_sink_prereq "Direct link to Prerequisites") * An Aiven for Apache Kafka service [with Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * Gather the following information for the target Redis database: * `REDIS_HOSTNAME`: The Redis hostname. * `REDIS_PORT`: The Redis port. * `REDIS_PASSWORD`: The Redis password. * `REDIS_SSL`: Set to `true` or `false`, depending on your SSL setup. * `TOPIC_LIST`: A comma-separated list of Kafka topics to sink. * `KCQL_TRANSFORMATION`: A KCQL statement to map topic fields to Redis cache entries. Use the following format: ``` INSERT INTO REDIS_CACHE SELECT LIST_OF_FIELDS FROM APACHE_KAFKA_TOPIC ``` * `APACHE_KAFKA_HOST`: The Apache Kafka host. Required only when using Avro as the data format. * `SCHEMA_REGISTRY_PORT`: The schema registry port. Required only when using Avro. * `SCHEMA_REGISTRY_USER`: The schema registry username. Required only when using Avro. * `SCHEMA_REGISTRY_PASSWORD`: The schema registry password. Required only when using Avro. note If you are using Aiven for Caching and Aiven for Apache Kafka, get all required connection details, including schema registry information, from the **Connection information** section on the **Overview** page. As of version 3.0, Aiven for Apache Kafka uses Karapace as the schema registry and no longer supports the Confluent Schema Registry. ## Create a connector configuration file[​](#create-a-connector-configuration-file "Direct link to Create a connector configuration file") Create a file named `redis_sink.json` and add the following configuration: ``` { "name": "CONNECTOR_NAME", "connector.class": "com.datamountaineer.streamreactor.connect.redis.sink.RedisSinkConnector", "topics": "TOPIC_LIST", "connect.redis.host": "REDIS_HOSTNAME", "connect.redis.port": "REDIS_PORT", "connect.redis.password": "REDIS_PASSWORD", "connect.redis.ssl.enabled": "REDIS_SSL", "connect.redis.kcql": "KCQL_TRANSFORMATION", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD" } ``` Parameters: * `name`: The connector name. Replace `CONNECTOR_NAME` with your desired name. * `connect.redis.*`: Redis connection parameters collected in the [prerequisite step](#connect_redis_lenses_sink_prereq). * `key.converter` and `value.converter`: Define the message data format in the Kafka topic. This example uses `io.confluent.connect.avro.AvroConverter` to translate messages in Avro format. The schema is retrieved from Aiven's [Karapace schema registry](https://github.com/aiven/karapace) using the `schema.registry.url` and related credentials. note The `key.converter` and `value.converter` fields define how Kafka messages are parsed and must be included in the configuration. When using Avro as the source format, set the following: * `value.converter.schema.registry.url`: Use the Aiven for Apache Kafka schema registry URL in the format `https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT`. * `value.converter.basic.auth.credentials.source`: Set to `USER_INFO`, which means authentication is done using a username and password. * `value.converter.schema.registry.basic.auth.user.info`: Provide the schema registry credentials in the format `SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD`. You can retrieve these values from the [prerequisite step](#connect_redis_lenses_sink_prereq). ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Console * Aiven CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 5. In the sink connectors list, select **Redis Sink Connector**, and click **Get started**. 6. On the **Stream Reactor Redis Sink** page, go to the **Common** tab. 7. Locate the **Connector configuration** text box and click **Edit**. 8. Paste the configuration from your `redis_sink.json` file into the text box. 9. Click **Create connector**. 10. Verify the connector status on the **Manage stream** > **Connectors** page. 11. Confirm that data is written to the Redis target database. To create the connector using the [Aiven CLI](/docs/tools/cli/service/connector.md#avn_service_connector_create), run: ``` avn service connector create SERVICE_NAME @redis_sink.json ``` Replace: * `SERVICE_NAME`: Your Kafka or Kafka Connect service name. * `@redis_sink.json`: Path to your configuration file. ## Sink topic data to Redis[​](#sink-topic-data-to-redis "Direct link to Sink topic data to Redis") The following example shows how to sink data from a Kafka topic to a Redis database. If your Kafka topic `students` contains the following data: ``` {"id":1, "name":"carlo", "age": 77} {"id":2, "name":"lucy", "age": 55} {"id":3, "name":"carlo", "age": 33} {"id":2, "name":"lucy", "age": 21} ``` To write this data to Redis, use the following connector configuration: ``` { "name": "my-redis-sink", "connector.class": "com.datamountaineer.streamreactor.connect.redis.sink.RedisSinkConnector", "connect.redis.host": "REDIS_HOSTNAME", "connect.redis.port": "REDIS_PORT", "connect.redis.password": "REDIS_PASSWORD", "connect.redis.ssl.enabled": "REDIS_SSL", "topics": "students", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter.schemas.enable": "false", "connect.redis.kcql": "INSERT INTO students- SELECT * FROM students PK id" } ``` Replace all placeholder values (such as `REDIS_HOSTNAME`, `REDIS_PORT`, and `REDIS_PASSWORD`) with your actual Redis connection details. This configuration does the following: * `"topics": "students"`: Specifies the Kafka topic to sink. * Connection settings (`connect.redis.*`): Provide the Redis host, port, password, and SSL setting. * `"value.converter"` and `"value.converter.schemas.enable"`: Set the message format. The topic uses raw JSON without a schema. * `"connect.redis.kcql"`: Defines the insert logic. Each Kafka message is written as a key-value pair in Redis. The key is built from the `id` field and prefixed with `students-`. After creating the connector, check your Redis database. You should see the following entries: ``` 1. "students-1" containing "{\"name\":\"carlo\",\"id\":1,\"age\":77}" 2. "students-2" containing "{\"name\":\"lucy\",\"id\":2,\"age\":21}" 3. "students-3" containing "{\"name\":\"carlo\",\"id\":3,\"age\":33}" ``` Only three keys are present because two Kafka messages shared the same `"id": 2`. Redis overwrites entries that use the same key. --- # Request a new connector If you know about new and interesting Apache Kafka® connectors you'd like us to support, email us at . We will evaluate the requested connector and might add support for it. Aiven's evaluation process for new Apache Kafka Connect connectors checks: * license compatibility * technical implementation * active repository maintenance tip When requesting connectors that are not on the pre-approved list through a support ticket, specify the target Aiven for Apache Kafka service you'd like to have it installed to. --- # Use AWS IAM assume role credentials provider The [Aiven for Apache Kafka® S3 sink connector](/docs/products/kafka/kafka-connect/howto/s3-sink-connector-aiven.md) moves data from an Aiven for Apache Kafka cluster to Amazon S3 for long-term storage. You can connect the S3 sink connector to Amazon S3 using either: * Long-term AWS credentials (`ACCESS_KEY_ID` and `SECRET_ACCESS_KEY`) * [AWS IAM assume role credentials](https://docs.aws.amazon.com/sdkref/latest/guide/feature-assume-role-credentials) (recommended) When you use IAM assume role credentials, the connector requests short-term credentials each time it writes data to the S3 bucket. To use IAM assume role credentials: * Request a unique IAM user from Aiven support. * Create an AWS cross-account access role. * Create a Kafka Connect S3 sink connector. ## Request a unique IAM user from Aiven support[​](#request-a-unique-iam-user-from-aiven-support "Direct link to Request a unique IAM user from Aiven support") Each Aiven project has a dedicated IAM user. Aiven does not share IAM users or roles across customers. Contact Aiven support at `support@aiven.io` to request: * An IAM user ARN * An [External ID](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_roles_create_for-user_externalid) (used to identify your role) Example: * IAM user: `arn:aws:iam::012345678901:user/sample-project-user` * External ID: `2f401145-06a0-4938-8e05-2d67196a0695` The [cross-account role](https://docs.aws.amazon.com/IAM/latest/UserGuide/tutorial_cross-account-with-roles) you create provides access to one Aiven project only. ## Create an AWS cross-account access role[​](#create-an-aws-cross-account-access-role "Direct link to Create an AWS cross-account access role") 1. Sign in to the AWS console. 2. Go to **IAM** > **Roles** > **Create role**. 3. Select **Another AWS account** as the trusted entity type. 4. Enter the **Account ID**. This is the numeric string in the IAM user ARN between `aws:iam::` and `:user/`. Example: `012345678901` 5. Select **Require external ID** and enter the External ID provided by Aiven support. 6. Add permissions to allow writing to an S3 bucket. The following permissions are required: * `s3:GetObject` * `s3:PutObject` * `s3:AbortMultipartUpload` * `s3:ListMultipartUploadParts` * `s3:ListBucketMultipartUploads` 7. Optional: Add tags. 8. Enter a name for the role. Example: `AivenKafkaConnectSink` 9. To restrict access, edit the trust relationship for the new role: * Go to **Trust relationships** > **Edit trust relationship** * Set the IAM user as the `Principal`. 10. Copy the new **IAM role ARN**. You will need it in the connector configuration. ## Create a Kafka Connect S3 sink connector[​](#create-a-kafka-connect-s3-sink-connector "Direct link to Create a Kafka Connect S3 sink connector") Create the connector as described in the [S3 sink connector documentation](/docs/products/kafka/kafka-connect/howto/s3-sink-connector-aiven.md). To use IAM assume role credentials, remove these parameters from the connector configuration: * `aws.access.key.id` * `aws.secret.access.key` Add these parameters: * `aws.sts.role.arn`: ARN of the IAM role created in AWS. * `aws.sts.role.external.id`: External ID provided by Aiven support. Optional parameters: * `aws.sts.role.session.name`: Session identifier for the task. Appears in AWS CloudTrail logs and helps distinguish tasks within the same project. * `aws.sts.config.endpoint`: Security Token Service (STS) endpoint. Use the endpoint in the same region as the S3 bucket for better performance. Example: For region `eu-north-1`, set `https://sts.eu-north-1.amazonaws.com`. For the list of STS endpoints, see the [AWS documentation](https://docs.aws.amazon.com/IAM/latest/UserGuide/id_credentials_temp_enable-regions). ### Example configuration file[​](#example-configuration-file "Direct link to Example configuration file") Save the following as `s3_sink.json`: ``` { "name": "", "connector.class": "io.aiven.kafka.connect.s3.AivenKafkaConnectS3SinkConnector", "key.converter": "org.apache.kafka.connect.converters.ByteArrayConverter", "value.converter": "org.apache.kafka.connect.converters.ByteArrayConverter", "topics": "", "aws.sts.role.arn": "", "aws.sts.role.external.id": "", "aws.sts.role.session.name": "", "aws.sts.config.endpoint": "", "aws.s3.bucket.name": "", "aws.s3.region": "" } ``` For the full list of S3 sink connector settings and examples, see the [S3 sink connector documentation](/docs/products/kafka/kafka-connect/howto/s3-sink-connector-aiven.md). --- # Amazon S3 sink connector Use the Amazon S3 sink connectors to move data from Apache Kafka® topics to Amazon S3 for storage and analytics. Before you create a connector, complete the setup steps in [Prepare AWS for S3 sink](/docs/products/kafka/kafka-connect/howto/s3-sink-prepare.md). Choose one of the following connectors: * [Amazon S3 sink connector (Aiven)](/docs/products/kafka/kafka-connect/howto/s3-sink-connector-aiven.md) * [Amazon S3 sink connector (Confluent)](/docs/products/kafka/kafka-connect/howto/s3-sink-connector-confluent.md) --- # S3 sink connector by Aiven naming and data formats The Apache Kafka Connect® S3 sink connector by Aiven moves data from an Aiven for Apache Kafka® cluster to Amazon S3. You can configure object naming and output data formats. Aiven provides two S3 sink connectors: * An Aiven-developed connector * A Confluent-developed connector This content applies to the Aiven version. For the Confluent version, see [S3 sink connector additional parameters (Confluent)](/docs/products/kafka/kafka-connect/howto/s3-sink-additional-parameters-confluent.md). ## Object naming in S3[​](#object-naming-in-s3 "Direct link to Object naming in S3") The connector writes data as objects in the configured S3 bucket. By default, object names follow this pattern: ``` --. ``` The following placeholders define the pattern: * `AWS_S3_PREFIX`: An optional string prefix. Use placeholders such as `{{ utc_date }}` and `{{ local_date }}` to create date-based object paths. * `TOPIC_NAME`: The Kafka topic name written to Amazon S3. * `PARTITION_NUMBER`: The Kafka topic partition number. * `START_OFFSET`: The starting offset of the records in the file. * `FILE_EXTENSION`: Depends on the value of `file.compression.type`. For example, `gzip` compression produces files with a `.gz` extension. You can configure object naming and how records are grouped into files. See the file naming format section in the [Aiven S3 sink connector GitHub repository](https://github.com/aiven/aiven-kafka-connect-s3). The connector creates one file per partition for each interval defined by `offset.flush.interval.ms`. The connector creates a file only if the partition receives at least one record during the interval. The default interval is 60 seconds. ## S3 data format[​](#s3-data-format "Direct link to S3 data format") By default, the connector writes one record per line in CSV format. To change the output format to JSON or Parquet, set `format.output.type`. The `format.output.fields` configuration controls the output fields. If included, the connector Base64-encodes the record key and value. For example, setting `format.output.fields` to `value,key,timestamp` produces output similar to the following: ``` bWVzc2FnZV9jb250ZW50,cGFydGl0aW9uX2tleQ==,1511801218777 ``` To disable Base64 encoding for values, set `format.output.fields.value.encoding` to `none`. Related pages * [File naming format in the Aiven S3 sink connector GitHub repository](https://github.com/aiven/aiven-kafka-connect-s3) --- # S3 sink connector by Confluent naming and data formats The Apache Kafka Connect® S3 sink connector moves data from Aiven for Apache Kafka® to Amazon S3 for long-term storage. Aiven provides two S3 sink connectors: * An Aiven-developed connector * A Confluent-developed connector This content applies to the Confluent version. For the Aiven-developed connector, see [S3 sink connector additional parameters](/docs/products/kafka/kafka-connect/howto/s3-sink-additional-parameters.md). ## S3 naming format[​](#s3-naming-format "Direct link to S3 naming format") The connector stores data as objects in the configured S3 bucket. By default, object names follow this pattern: ``` topics//partition=/++. ``` The following placeholders define the pattern: * `TOPIC_NAME`: The Kafka topic name written to Amazon S3. * `PARTITION_NUMBER`: The Kafka topic partition number. * `START_OFFSET`: The starting offset of the records in the file. * `FILE_EXTENSION`: Depends on the configured serialization format. For example, binary serialization produces files with a `.bin` extension. For example, a topic with 3 partitions initially generates the following files in the destination S3 bucket: ``` topics//partition=0/+0+0000000000.bin topics//partition=1/+1+0000000000.bin topics//partition=2/+2+0000000000.bin ``` ## S3 data format[​](#s3-data-format "Direct link to S3 data format") By default, the connector stores data in binary format, with one message per line. The connector creates a file after a fixed number of messages. The `flush.size` parameter defines this number. Setting `flush.size` to `1` creates one file per message. For example, for a topic with three partitions and ten messages, setting `flush.size` to `1` produces the following files in the destination S3 bucket (one file per message): ``` topics//partition=0/+0+0000000000.bin topics//partition=0/+0+0000000001.bin topics//partition=0/+0+0000000002.bin topics//partition=0/+0+0000000003.bin topics//partition=1/+1+0000000000.bin topics//partition=1/+1+0000000001.bin topics//partition=1/+1+0000000002.bin topics//partition=2/+2+0000000000.bin topics//partition=2/+2+0000000001.bin topics//partition=2/+2+0000000002.bin ``` For more information, see the [Confluent S3 sink connector documentation](https://docs.confluent.io/5.0.0/connect/kafka-connect-s3/index). --- # Create an Amazon S3 sink connector for Aiven from Apache Kafka® The Amazon S3 sink connector sends data from Aiven for Apache Kafka® to Amazon S3 for long-term storage. View the full configuration reference in the [S3 sink connector documentation](https://github.com/Aiven-Open/cloud-storage-connectors-for-apache-kafka/blob/main/s3-sink-connector/README.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven for Apache Kafka® service [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md), or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * AWS access to an S3 bucket. See [Prepare AWS for S3 sink](/docs/products/kafka/kafka-connect/howto/s3-sink-prepare.md) to complete bucket and IAM setup. After setup, you will have the required S3 bucket details and access credentials. To use AWS IAM assume role credentials, see [Use AWS IAM assume role credentials provider](/docs/products/kafka/kafka-connect/howto/s3-iam-assume-role.md). ## Create an S3 sink connector configuration file[​](#create-an-s3-sink-connector-configuration-file "Direct link to Create an S3 sink connector configuration file") Create a file named `s3_sink_connector.json` with the following configuration: ``` { "name": "s3_sink", "connector.class": "io.aiven.kafka.connect.s3.AivenKafkaConnectS3SinkConnector", "tasks.max": "1", "topics": "my-topic", "aws.access.key.id": "AKIAXXXXXXXXXX", "aws.secret.access.key": "XXXXXXXXXXXXXXXXXXXXXXXXXXXX", "aws.s3.bucket.name": "my-bucket", "aws.s3.region": "eu-central-1", "key.converter": "org.apache.kafka.connect.converters.ByteArrayConverter", "value.converter": "org.apache.kafka.connect.converters.ByteArrayConverter" } ``` Parameters: * `name`: Connector name * `topics`: Apache Kafka topics to sink data from * `aws.access.key.id`: AWS access key ID * `aws.secret.access.key`: AWS secret access key * `aws.s3.bucket.name`: S3 bucket name * `aws.s3.region`: AWS region * `key.converter`, `value.converter`: Converters for key and value formats * `file.compression.type` optional: Compression type for output files. Supported values: `gzip`, `snappy`, `zstd`, `none`. Default is `gzip`. If you use the [Amazon S3 source connector](/docs/products/kafka/kafka-connect/howto/s3-source-connector.md), set the same `file.compression.type` in the source configuration to ensure the source can read data written by the sink. The S3 source connector defaults to `none`. For additional parameters (for example, file naming or output format), see [S3 sink connector reference](https://github.com/Aiven-Open/cloud-storage-connectors-for-apache-kafka/blob/main/s3-sink-connector/README.md). ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Aiven Console * Aiven CLI 1. Open the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka® or Apache Kafka Connect® service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector**. If connectors are not enabled, click **Enable connectors on this service**. 5. Search for **Amazon S3 sink** and click **Get started**. 6. Go to the **Common** tab and click **Edit** in the **Connector configuration** section. 7. Paste your `s3_sink_connector.json` content. 8. Click **Create connector**. 9. Confirm the connector is running in the **Connectors** view. Create the connector: ``` avn service connector create SERVICE_NAME @s3_sink_connector.json ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka® service * `@s3_sink_connector.json`: Path to the connector configuration file Check the connector status: ``` avn service connector status SERVICE_NAME CONNECTOR_NAME ``` ## Example: create an Amazon S3 sink connector[​](#example-create-an-amazon-s3-sink-connector "Direct link to Example: create an Amazon S3 sink connector") This example creates an S3 sink connector that writes data from the `students` topic to the `my-test-bucket` bucket in the `eu-central-1` region. ``` { "name": "my_s3_sink", "connector.class": "io.aiven.kafka.connect.s3.AivenKafkaConnectS3SinkConnector", "tasks.max": "1", "topics": "students", "aws.access.key.id": "AKIAXXXXXXXXXX", "aws.secret.access.key": "hELuXXXXXXXXXXXXXXXXXXXXXXXXXX", "aws.s3.bucket.name": "my-test-bucket", "aws.s3.region": "eu-central-1", "key.converter": "org.apache.kafka.connect.converters.ByteArrayConverter", "value.converter": "org.apache.kafka.connect.converters.ByteArrayConverter" } ``` Save this as `s3_sink_connector.json`, then create the connector: ``` avn service connector create demo-kafka @s3_sink_connector.json ``` After creation, confirm that data from the `students` topic appears in the `my-test-bucket` S3 bucket. The S3 sink connector compresses files using `gzip` by default. To read these files using the S3 source connector, set `file.compression.type` in the source configuration to `gzip` or another matching value. Related pages [Amazon S3 sink connector by Confluent](/docs/products/kafka/kafka-connect/howto/s3-sink-connector-confluent.md) --- # Create an Amazon S3 sink connector by Confluent from Apache Kafka® The Amazon S3 sink connector by Confluent moves data from Apache Kafka® topics to Amazon S3 buckets for long-term storage. For the full list of configuration options, see the [Confluent S3 sink connector documentation](https://docs.confluent.io/current/connect/kafka-connect-s3/). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven for Apache Kafka® service [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md), or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * AWS access to an S3 bucket. See [Prepare AWS for S3 sink](/docs/products/kafka/kafka-connect/howto/s3-sink-prepare.md) to complete bucket and IAM setup. After setup, you will have the required S3 bucket details and access credentials. ## Create the connector configuration file[​](#create-the-connector-configuration-file "Direct link to Create the connector configuration file") Create a file named `s3_sink_confluent.json` with the following configuration: ``` { "name": "s3_sink_confluent", "connector.class": "io.confluent.connect.s3.S3SinkConnector", "tasks.max": "1", "topics": "my-topic", "key.converter": "org.apache.kafka.connect.converters.ByteArrayConverter", "value.converter": "org.apache.kafka.connect.converters.ByteArrayConverter", "format.class": "io.confluent.connect.s3.format.bytearray.ByteArrayFormat", "flush.size": "1", "s3.bucket.name": "my-bucket", "s3.region": "eu-central-1", "s3.credentials.provider.class": "io.aiven.kafka.connect.util.AivenAWSCredentialsProvider", "storage.class": "io.confluent.connect.s3.storage.S3Storage", "s3.credentials.provider.access_key_id": "", "s3.credentials.provider.secret_access_key": "" } ``` Parameters: * `connector.class`: Uses the Confluent S3 sink connector * `topics`: Kafka topics to sink * `format.class`: Output format (byte array in this example) * `flush.size`: Number of messages written per file in S3 * `s3.bucket.name` / `s3.region`: Destination S3 bucket * `s3.credentials.provider.class`: Use Aiven provider to supply AWS credentials For all options, see the [Confluent connector documentation](https://docs.confluent.io/current/connect/kafka-connect-s3/). ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Aiven Console * Aiven CLI 1. Open the [Aiven Console](https://console.aiven.io/). 2. Select your Kafka or Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector**. 5. Search for **Amazon S3 sink (Confluent)** and click **Get started**. 6. On the connector page, go to the **Common** tab. 7. Click **Edit** in **Connector configuration**. 8. Paste your `s3_sink_confluent.json` contents. 9. Click **Create connector**. 10. Verify the connector status on the **Manage stream** > **Connectors** page. Create the connector: ``` avn service connector create SERVICE_NAME @s3_sink_confluent.json ``` Check the connector status: ``` avn service connector status SERVICE_NAME s3_sink_confluent ``` ## Example: create an Amazon S3 sink connector (Confluent)[​](#example-create-an-amazon-s3-sink-connector-confluent "Direct link to Example: create an Amazon S3 sink connector (Confluent)") This example creates an S3 sink connector that writes data from the `students` topic to the `my-test-bucket` bucket in the `eu-central-1` region, writing a file every 10 messages. ``` { "name": "my_s3_sink", "connector.class": "io.confluent.connect.s3.S3SinkConnector", "tasks.max": "1", "topics": "students", "key.converter": "org.apache.kafka.connect.converters.ByteArrayConverter", "value.converter": "org.apache.kafka.connect.converters.ByteArrayConverter", "format.class": "io.confluent.connect.s3.format.bytearray.ByteArrayFormat", "flush.size": "10", "s3.bucket.name": "my-test-bucket", "s3.region": "eu-central-1", "storage.class": "io.confluent.connect.s3.storage.S3Storage", "s3.credentials.provider.access_key_id": "AKIAXXXXXXXXXX", "s3.credentials.provider.secret_access_key": "hELuXXXXXXXXXXXXXXXXXXXXXXXXXX" } ``` Save this as `s3_sink_confluent.json`, then create the connector: ``` avn service connector create demo-kafka @s3_sink_confluent.json ``` After creation, confirm that data from the `students` topic appears in the `my-test-bucket` S3 bucket. Related pages [Amazon S3 sink connector by Aiven](/docs/products/kafka/kafka-connect/howto/s3-sink-connector-aiven.md) --- # Prepare AWS for Amazon S3 sink Set up AWS to allow the S3 sink connector to write data from Apache Kafka® to Amazon S3. ## Create the S3 bucket[​](#create-the-s3-bucket "Direct link to Create the S3 bucket") 1. Open the [AWS S3 console](https://s3.console.aws.amazon.com/). 2. Create a bucket. 3. Enter a bucket name and choose a region. Keep the remaining settings as default. note Keep **Block all public access** enabled. The connector uses IAM permissions to access the bucket. ## Create an IAM policy[​](#create-an-iam-policy "Direct link to Create an IAM policy") The Apache Kafka Connect S3 sink connector requires these permissions: * `s3:GetObject` * `s3:PutObject` * `s3:AbortMultipartUpload` * `s3:ListMultipartUploadParts` * `s3:ListBucketMultipartUploads` Create an inline policy in AWS IAM and replace `` with your bucket name: ``` { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Action": [ "s3:GetObject", "s3:PutObject", "s3:AbortMultipartUpload", "s3:ListMultipartUploadParts", "s3:ListBucketMultipartUploads" ], "Resource": [ "arn:aws:s3:::", "arn:aws:s3:::/*" ] } ] } ``` ## Create the IAM user[​](#create-the-iam-user "Direct link to Create the IAM user") 1. Open the [IAM console](https://console.aws.amazon.com/iamv2/home). 2. Create a user. 3. In **Select AWS credential type**, select **Access key - Programmatic access**. Copy the **Access key ID** and **Secret access key**. You use these values in the connector configuration. 4. In **Permissions**, attach the policy created in the previous section. note If you see `Access Denied` errors when starting the connector, review the AWS guidance for [S3 access issues](https://docs.aws.amazon.com/AmazonS3/latest/userguide/troubleshoot-403-errors). --- # Create an Amazon S3 source connector for Aiven for Apache Kafka® The Amazon S3 source connector allows you to ingest data from S3 buckets into Apache Kafka® topics for real-time processing and analytics. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An [Aiven for Apache Kafka® service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with Aiven for Kafka Connect enabled, or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * An **Amazon S3 bucket** containing the data to stream into Aiven for Apache Kafka. * Required S3 bucket details: * `AWS_S3_BUCKET_NAME`: The bucket name. * `AWS_S3_REGION`: The AWS region where the bucket is located. For example, us-east-1. * `AWS_S3_PREFIX`: Optional. A prefix path if the data is in a specific folder. * `AWS_ACCESS_KEY_ID`: The AWS access key ID. * `AWS_SECRET_ACCESS_KEY`: The AWS secret access key. * `TARGET_KAFKA_TOPIC`: The Apache Kafka topic where the data is published. * **AWS IAM credentials** with the following permissions: * `s3:GetObject` * `s3:ListBucket` For additional details on how the connector works, see the [S3 source connector documentation](https://github.com/aiven-open/cloud-storage-connectors-for-apache-kafka/blob/main/s3-source-connector/README.md). ## S3 object key name format[​](#s3-object-key-name-format "Direct link to S3 object key name format") The `file.name.template` setting defines how the connector extracts metadata from S3 object keys. This setting is required. If it is not set, the connector does not process any objects. The configuration depends on how your files are named in S3: **`object_hash` distribution type** In version 3.4.0 and later, you can match any object key: * `file.name.template=".*"` to process all object keys, or * A regular expression to process only keys that match the pattern **`partition` distribution type** Set `file.name.template` using the following placeholders: * `{{topic}}`: The Kafka topic name * `{{partition}}`: The Kafka partition number * `{{start_offset}}`: The offset of the first record in the file * `{{timestamp}}`: Optional. The timestamp when the connector processed the file ### Backfill and prefix behavior[​](#backfill-and-prefix-behavior "Direct link to Backfill and prefix behavior") The connector uses the Amazon S3 `ListObjectsV2` API to list objects. * `aws.s3.prefix` limits processing to keys that begin with the specified prefix * The connector processes only objects under that prefix * The connector does not automatically continue to the next prefix * Objects are processed in the lexicographical order returned by Amazon S3 For time-partitioned folder structures, use a lexicographically sortable pattern such as: ``` YYYY/MM/DD/HH/ ``` note Use prefix filtering to backfill a specific time window from S3. ### Example templates and extracted values[​](#example-templates-and-extracted-values "Direct link to Example templates and extracted values") The following table shows how different templates extract metadata from S3 object keys: | Template | Example S3 object key | Extracted values | | ---------------------------------------------------------------------- | ---------------------------------------------------------------- | -------------------------------------------------------------- | | `{{topic}}-{{partition}}-{{start_offset}}` | `customer-topic-1-1734445664111.txt` | topic=customer-topic, partition=1, start\_offset=1734445664111 | | `{{topic}}-{{partition}}-{{start_offset}}` | `22-10-12/customer-topic-1-1734445664111.txt` | topic=22, partition=10, start\_offset=112 | | `{{topic}}/{{partition}}/{{start_offset}}` | `customer-topic/1/1734445664111.txt` | topic=customer-topic, partition=1, start\_offset=1734445664111 | | `topic/{{topic}}/partition/{{partition}}/startOffset/{{start_offset}}` | `topic/customer-topic/partition/1/startOffset/1734445664111.txt` | topic=customer-topic, partition=1, start\_offset=1734445664111 | ## Supported S3 object formats[​](#supported-s3-object-formats "Direct link to Supported S3 object formats") The connector supports four S3 object formats. Choose the one that best fits your data and processing needs: | Format | Description | Configuration | Example | | ------------------------- | -------------------------------------------------------------------- | ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | | JSON Lines (`jsonl`) | Each line is a valid JSON object, commonly used for event streaming. | `input.format=jsonl` | `{ "key": "k1", "value": "v0", "offset": 1232155, "timestamp": "2020-01-01T00:00:01Z" }` | | Avro (`avro`) | A compact, schema-based binary format for efficient serialization. | `input.format=avro` | `{ "type": "record", "fields": [ { "name": "key", "type": "string" }, { "name": "value", "type": "string" }, { "name": "timestamp", "type": "long" } ] }` | | Parquet (`parquet`) | A columnar format optimized for analytics. | `input.format=parquet` | Uses a schema similar to Avro but optimized for analytics. | | Bytes (`bytes`) (default) | A raw byte stream for unstructured data. | `input.format=bytes` | No predefined structure | ## Acknowledged records and offset tracking[​](#acknowledged-records-and-offset-tracking "Direct link to Acknowledged records and offset tracking") When a record is acknowledged, Apache Kafka confirms receipt but may not immediately write it to the offset topic. If the connector restarts before Apache Kafka updates the offset topic, some records may be duplicated. For details on offset tracking and retry handling, see the [S3 source connector documentation](https://github.com/aiven-open/cloud-storage-connectors-for-apache-kafka/blob/main/s3-source-connector/README.md). ## Create an Amazon S3 source connector configuration file[​](#create-an-amazon-s3-source-connector-configuration-file "Direct link to Create an Amazon S3 source connector configuration file") Create a file named `s3_source_connector.json` and add the following configuration: ``` { "name": "aiven-s3-source-connector", "connector.class": "io.aiven.kafka.connect.s3.source.S3SourceConnector", "aws.access.key.id": "YOUR_AWS_ACCESS_KEY_ID", "aws.secret.access.key": "YOUR_AWS_SECRET_ACCESS_KEY", "aws.s3.bucket.name": "your-s3-bucket-name", "aws.s3.region": "your-s3-region", "aws.s3.prefix": "optional/prefix/", "aws.credentials.provider": "software.amazon.awssdk.auth.credentials.AwsCredentialsProviderChain", "topic": "your-target-kafka-topic", "file.name.template": "{{topic}}-{{partition}}-{{start_offset}}", "tasks.max": 1, "poll.interval.ms": 10000, "error.tolerance": "all", "input.format": "jsonl", "file.compression.type": "none", "timestamp.timezone": "UTC", "timestamp.source": "wallclock" } ``` Parameters: * `name`: A unique name for the connector. * `connector.class`: The Java class for the connector. * `aws.access.key.id` and `aws.secret.access.key`: AWS IAM credentials for authentication. * `aws.s3.bucket.name`: The name of the S3 bucket containing the source data. * `aws.s3.region`: The AWS region where the bucket is located. * `aws.s3.prefix` optional: Limits processing to keys that begin with this prefix. Only objects under this prefix are processed. The connector does not move to the next prefix. * `aws.credentials.provider`: Specifies the AWS credentials provider. * `topic` optional: The connector publishes ingested data to the specified Apache Kafka topic. If not set, it derives the topic from the `file.name.template` configuration. * `file.name.template`: Defines how to parse S3 object keys to extract the topic, partition, and starting offset. For example, `{{topic}}-{{partition}}-{{start_offset}}` matches a file name like `test-topic-1-1734445664111.txt`. For the `object_hash` distribution type (version 3.4.0+), set `file.name.template` to `.*` to match all keys or use a regular expression to match specific keys. * `tasks.max`: The maximum number of tasks that run in parallel. * `poll.interval.ms`: The polling interval (in milliseconds) for checking the S3 bucket for new files. * `error.tolerance`: Specifies the error handling mode. * `all`: Logs and ignores errors. * `none`: Fails on errors. * `input.format`: Specifies the S3 object format. Supported values: * `jsonl` (JSON Lines) * `avro` (Avro) * `parquet` (Parquet) * `bytes` (default) * `timestamp.timezone` optional: Time zone for timestamps. Default is `UTC`. * `timestamp.source` optional: Source of timestamps. Supports `wallclock`. Default is `wallclock`. * `file.compression.type` optional: Compression type for input files. Supported values: `gzip`, `snappy`, `zstd`, `none`. Default is `none`. ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Aiven Console * Aiven CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 5. In the source connectors list, select **Amazon S3 source connector**, and click **Get started**. 6. On the **Amazon S3 Source Connector** page, go to the **Common** tab. 7. Locate the **Connector configuration** text box and click **Edit**. 8. Paste the configuration from your `s3_source_connector.json` file into the text box. 9. Click **Create connector**. 10. Verify the connector status on the **Manage stream** > **Connectors** page. To create the S3 source connector using the [Aiven CLI](/docs/tools/cli/service/connector.md#avn_service_connector_create), run: ``` avn service connector create SERVICE_NAME @s3_source_connector.json ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `@s3_source_connector.json`: Path to your JSON configuration file. To check the connector status, run: ``` avn service connector status SERVICE_NAME CONNECTOR_NAME ``` Verify that data flows to the target Apache Kafka topic. ## Example: Define and create an Amazon S3 source connector[​](#example-define-and-create-an-amazon-s3-source-connector "Direct link to Example: Define and create an Amazon S3 source connector") This example shows how to create an Amazon S3 source connector with the following properties: * Connector name: `aiven-s3-source-connector` * Apache Kafka topic: `test-topic` * AWS region: `us-west-1` * AWS S3 bucket: `my-s3-bucket` * File name template: `{{topic}}-{{partition}}-{{start_offset}}` * Poll interval: `10 seconds` ``` { "name": "aiven-s3-source-connector", "connector.class": "io.aiven.kafka.connect.s3.source.S3SourceConnector", "tasks.max": 1, "topic": "test-topic", "aws.access.key.id": "your-access-key-id", "aws.secret.access.key": "your-secret-access-key", "aws.s3.bucket.name": "my-s3-bucket", "aws.s3.region": "us-west-1", "aws.s3.prefix": "data-uploads/", "file.name.template": "{{topic}}-{{partition}}-{{start_offset}}", "poll.interval.ms": 10000, "file.compression.type": "none", "timestamp.timezone": "UTC", "timestamp.source": "wallclock" } ``` Once this configuration is saved in the `s3_source_connector.json` file, you can create the connector using the Aiven Console or CLI, and verify that data from the Apache Kafka topic `test-topic` is successfully ingested from the Amazon S3 bucket. Related pages * [Amazon S3 sink connector by Aiven](/docs/products/kafka/kafka-connect/howto/s3-sink-connector-aiven.md) * [Amazon S3 sink connector by Confluent](/docs/products/kafka/kafka-connect/howto/s3-sink-connector-confluent.md) --- # Create a Salesforce sink connector for Aiven for Apache Kafka® [Early availability](/docs/platform/concepts/service-and-feature-releases.md) The Salesforce sink connector writes records from Apache Kafka® topics to Salesforce objects, such as `Account` or `Contact`. It consumes Kafka records and submits them as Salesforce Bulk API 2.0 insert jobs. note The connector provides at-least-once delivery. If records are retried before Kafka offsets are committed, Salesforce can receive duplicate records. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An [Aiven for Apache Kafka® service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with Apache Kafka Connect enabled, or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * A Salesforce account with the following: * A connected app configured for the OAuth 2.0 client credentials flow, with a client ID and client secret. For setup details, see [Configure the client credentials flow for external client apps](https://help.salesforce.com/s/articleView?id=xcloud.configure_client_credentials_flow_for_external_client_apps.htm\&type=5) in the Salesforce documentation. * Bulk API 2.0 access enabled. * The target Salesforce object already exists, for example `Account` or `Contact`. * A Kafka topic that contains the records to send to Salesforce. * Kafka records that use one of the following value formats: * `Struct` values with a schema * `Map` values with a schema * Schemaless `Map` values, such as JSON records produced when `value.converter.schemas.enable` is set to `false` * Field names in `Struct` values or keys in `Map` values that match Salesforce field API names on the target object. ## Create a Salesforce sink connector configuration file[​](#create-a-salesforce-sink-connector-configuration-file "Direct link to Create a Salesforce sink connector configuration file") Create a file named `salesforce_sink_connector.json` with the following configuration: ``` { "name": "salesforce_sink_connector", "connector.class": "io.aiven.kafka.connect.salesforce.sink.SalesforceSinkConnector", "tasks.max": "1", "topics": "test_topic", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter.schemas.enable": "true", "salesforce.bulk.api.sink.object": "Contact", "salesforce.client.id": "YOUR_OAUTH_CLIENT_ID", "salesforce.client.secret": "YOUR_OAUTH_CLIENT_SECRET", "salesforce.oauth.uri": "https://YOUR_INSTANCE.salesforce.com/services/oauth2/token", "salesforce.uri": "https://YOUR_INSTANCE.salesforce.com", "salesforce.api.version": "v65.0", "salesforce.max.retries": "3", "offset.flush.interval.ms": "60000", "offset.flush.timeout.ms": "30000" } ``` The example uses schema-based JSON values. To use schemaless JSON values, set `value.converter.schemas.enable` to `false`. Parameters: * `name`: Name of the connector. * `connector.class`: Use `io.aiven.kafka.connect.salesforce.sink.SalesforceSinkConnector`. * `tasks.max`: Maximum number of connector tasks. * `topics`: Kafka topics to read from. * `value.converter`: Converter for Kafka record values. Use `org.apache.kafka.connect.json.JsonConverter` for JSON. For Avro, use the appropriate converter and set `schema.registry.url`. * `value.converter.schemas.enable`: Set to `true` to use schema-based JSON values. Set to `false` to use schemaless JSON values. The connector supports both schema-based `Struct` values and schemaless `Map` values. * `salesforce.bulk.api.sink.object`: Salesforce object to write records to, such as `Account` or `Contact`. * `salesforce.client.id`: OAuth 2.0 client ID from your Salesforce connected app. * `salesforce.client.secret`: OAuth 2.0 client secret from your Salesforce connected app. * `salesforce.oauth.uri`: OAuth token endpoint for your Salesforce org. Set together with `salesforce.uri` so both match your deployment. * `salesforce.uri`: Base URL of your Salesforce instance. * `salesforce.api.version`: Salesforce REST API version string supported by your org. Examples use `v65.0`. * `salesforce.max.retries`: Maximum retries for Salesforce API and authentication requests. See the [configuration reference](https://aiven-open.github.io/salesforce-connector-for-apache-kafka/sink/configuration.html). * `offset.flush.interval.ms`: How often, in milliseconds, Kafka Connect flushes records to Salesforce. Default: `60000`. Each flush is submitted as one Bulk API 2.0 insert job. A larger interval reduces Salesforce API calls but produces larger batches that take longer to process and can require a higher `offset.flush.timeout.ms`. * `offset.flush.timeout.ms`: Maximum time in milliseconds for a flush before Kafka Connect marks it failed. Default: `5000`. The configuration examples use `30000` when each flush includes many records or when Salesforce Bulk API responses are slower than the default allows. Raise this value if flush operations exceed `5000` ms or fail with timeout errors. For all available configuration options, see the [Salesforce sink connector configuration reference](https://aiven-open.github.io/salesforce-connector-for-apache-kafka/sink/configuration.html). ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Aiven Console * Aiven CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is already enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 5. In the sink connectors list, select **Salesforce**, and click **Get started**. 6. On the **Salesforce** connector page, go to the **Common** tab. 7. Locate the **Connector configuration** text box and click **Edit**. 8. Paste the configuration from your `salesforce_sink_connector.json` file into the text box. 9. Click **Create connector**. 10. Verify the connector status on the **Manage stream** > **Connectors** page. To create the Salesforce sink connector using the [Aiven CLI](/docs/tools/cli/service/connector.md#avn_service_connector_create), run: ``` avn service connector create SERVICE_NAME @salesforce_sink_connector.json ``` To check the connector status, run: ``` avn service connector status SERVICE_NAME CONNECTOR_NAME ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service. * `@salesforce_sink_connector.json`: Path to your JSON configuration file. * `CONNECTOR_NAME`: Value of the `name` field in the connector configuration. Verify that records are written to the expected Salesforce object. ## Verify the connector[​](#verify-the-connector "Direct link to Verify the connector") After you create the connector, confirm that records are flowing to Salesforce: 1. Open the target Salesforce object, for example `Contact`. 2. Confirm that new records appear. 3. Optionally, verify that field names and values match the records in the Kafka topic. For a sample record shape when writing to `Contact`, see [Example record for Salesforce](#example-record-for-salesforce). ## Example record for Salesforce[​](#example-record-for-salesforce "Direct link to Example record for Salesforce") If the connector writes to the Salesforce `Contact` object, a Kafka record value can use schema-based or schemaless JSON. For schemaless JSON, the value can look like this: ``` { "FirstName": "Alice", "LastName": "Example", "Email": "alice@example.com" } ``` The connector uses the JSON keys as Salesforce field API names: `FirstName`, `LastName`, and `Email`. ## Data format[​](#data-format "Direct link to Data format") The connector accepts Kafka record values as `Struct` or `Map` values. For `Struct` values, use field names that match Salesforce object field API names. For `Map` values, use keys that match Salesforce object field API names. For example, the following Kafka record value: ``` { "Name": "Alice", "Email": "alice@example.com", "ExternalId__c": "EXT001" } ``` is written to Salesforce using the following field names: ``` Name,Email,ExternalId__c ``` For custom Salesforce fields, use the Salesforce API field name, such as `ExternalId__c`. ## Schema behavior[​](#schema-behavior "Direct link to Schema behavior") The connector detects Salesforce field names dynamically from the records buffered during each flush. For `Struct` values, it uses schema field names. For `Map` values, it uses map keys. If records in the same flush contain different fields or keys, the connector creates one CSV header that includes all detected Salesforce field names. For example, if one record contains `Name` and `Email`, and another record contains `Name` and `ExternalId__c`, the generated CSV header includes all three fields: ``` Email,ExternalId__c,Name ``` Records that do not contain a detected field or key are sent with an empty value in that column. To reduce failures, keep records in a topic consistent with the target Salesforce object schema. ## How it works[​](#how-it-works "Direct link to How it works") When Kafka Connect flushes records, the connector: 1. Collects buffered records for that flush interval. 2. Detects field names from the buffered records. 3. Uses schema field names for `Struct` values and map keys for `Map` values. 4. Creates CSV data with a header row that contains the detected field names. 5. Sends the CSV data to Salesforce as a Bulk API 2.0 multipart insert job. 6. Polls Salesforce until the job reaches a final state. 7. Commits Kafka offsets after the Salesforce job completes successfully. ## Limitations[​](#limitations "Direct link to Limitations") The Salesforce sink connector has the following limitations: * **Insert only:** Only insert operations are supported. Update and upsert operations are not supported. * **At-least-once delivery:** The connector can write the same record more than once if a task restarts or Kafka Connect replays records before offsets are committed. * **Supported value types:** The connector processes `Struct` and `Map` values. Records with other value types are skipped. If the errant record reporter is configured, skipped records are reported as errant records. If a flush contains no supported values (`Struct` or `Map`), the batch can fail. * **No automatic data mapping:** The connector does not transform field names, keys, or values. Use Kafka record field names or keys that match Salesforce field API names. * **Dynamic fields:** Field names and keys are discovered dynamically from buffered records. Inconsistent fields or keys across records in the same flush interval produce sparse rows with empty values where a field is missing. * **Single batch per flush:** All records buffered within one flush interval are submitted as a single Salesforce Bulk API 2.0 job. Large batches may approach Salesforce API limits or exceed the Kafka Connect flush timeout. Related pages * [Salesforce connector for Apache Kafka on GitHub](https://github.com/Aiven-Open/salesforce-connector-for-apache-kafka) * [Salesforce Bulk API 2.0 Developer Guide](https://developer.salesforce.com/docs/atlas.en-us.api_asynch.meta/api_asynch/bulk_api_2_0.htm) --- # Create a Salesforce source connector for Aiven for Apache Kafka® [Early availability](/docs/platform/concepts/service-and-feature-releases.md) The Salesforce source connector retrieves data from Salesforce and writes it to Apache Kafka® topics. It retrieves data in batches and tracks progress using the `LastModifiedDate` field. The connector uses Salesforce Bulk API 2.0 to run SOQL (Salesforce Object Query Language) queries on a polling schedule. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An [Aiven for Apache Kafka® service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with Aiven for Kafka Connect enabled, or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). * A Salesforce connected app configured for the client credentials flow with these OAuth scopes: * `api`: Grants access to Salesforce data using REST and Bulk APIs. * `refresh_token` / `offline_access`: Allows the connector to maintain session access without user interaction. * Complete client credentials setup per [Configure the client credentials flow for external client apps](https://help.salesforce.com/s/articleView?id=xcloud.configure_client_credentials_flow_for_external_client_apps.htm\&type=5) in the Salesforce documentation. From the connected app, copy: * `salesforce.client.id`: Consumer Key from your connected app. * `salesforce.client.secret`: Consumer Secret from your connected app. * `salesforce.oauth.uri`: The OAuth token endpoint. * `salesforce.uri`: Base URL of your Salesforce instance. * SOQL queries for the Salesforce objects and fields to ingest. * The `topic.prefix` or `topic` setting. * Use `topic.prefix` to write records to object-specific Apache Kafka topics. For more information, see [Topic naming](#topic-naming). * Use `topic` to write all records to a single topic. If you set `topic`, it overrides `topic.prefix`. For more information about how the connector works, see the [Salesforce connector for Apache Kafka documentation](https://github.com/aiven-open/salesforce-connector-for-apache-kafka/blob/main/README.md). ## Define SOQL queries[​](#define-soql-queries "Direct link to Define SOQL queries") SOQL (Salesforce Object Query Language) is a SQL-like query language used to retrieve data from Salesforce objects. Set the `salesforce.bulk.api.queries` parameter to one or more SOQL queries, separated by semicolons (`;`). Each query must meet the following requirements: * Include `LastModifiedDate` in the `SELECT` clause. This field is required for offset tracking and incremental reads. * Do not include `LastModifiedDate` in the `WHERE` clause. The connector applies this filter internally. * Do not include `ORDER BY` in SOQL queries. The connector applies ordering internally. These requirements apply to both standard and custom objects. Example: ``` SELECT Id, Name, LastModifiedDate FROM Account; SELECT Id, FirstName, LastModifiedDate FROM Contact; ``` ## Topic naming[​](#topic-naming "Direct link to Topic naming") By default, the connector writes records to topics using this format: ``` .bulkApi. ``` The `bulkApi` segment indicates that records are retrieved using the Salesforce Bulk API. For example, if `topic.prefix` is `salesforce.test` and the query targets the `Account` object, records are written to: ``` salesforce.test.bulkApi.Account ``` If you set `topic`, the connector writes all records to that topic, regardless of the number of SOQL queries or objects. When `topic` is set, it overrides `topic.prefix` and the per-object naming pattern. ## Delivery semantics and duplicates[​](#delivery-semantics-and-duplicates "Direct link to Delivery semantics and duplicates") caution The connector provides at-least-once delivery. Duplicate records can occur. If the connector restarts due to uncommitted offsets, an error, or a configuration change, it reads again from the last `LastModifiedDate` it processed and includes all records with that timestamp. This ensures that if a bulk update in Salesforce assigns the same `LastModifiedDate` to many records, none are skipped. As a result, some records can be written to Kafka more than once if they were already sent before the restart. Duplicates can also occur when: * A record is updated and its `LastModifiedDate` changes. * A field not included in the SOQL query is updated, which still updates `LastModifiedDate`. To remove duplicates downstream, use the Salesforce `Id` field as the unique record key. ## Limitations[​](#limitations "Direct link to Limitations") * Single task only: `tasks.max` must be `1`. If you set a value greater than `1`, configuration validation fails. More than one task is not supported due to timing issues and offset conflicts. * Polling-based ingestion: The connector does not stream changes in real time. * At-least-once delivery: Duplicate records can occur. For more information, see [Delivery semantics and duplicates](#delivery-semantics-and-duplicates). * Salesforce Bulk API quota: Each polling cycle consumes API quota. Use `salesforce.soql.query.wait` to control how long the connector waits before running queries again. The default is 5 minutes (300 seconds) and the maximum is 1 week (604800 seconds). Use longer intervals when data changes are infrequent to reduce API usage. ## Create a Salesforce source connector configuration file[​](#create-a-salesforce-source-connector-configuration-file "Direct link to Create a Salesforce source connector configuration file") Create a file named `salesforce_source_connector.json` and add the following configuration: ``` { "name": "salesforce-source", "connector.class": "io.aiven.kafka.connect.salesforce.source.SalesforceSourceConnector", "tasks.max": 1, "key.converter": "org.apache.kafka.connect.storage.StringConverter", "value.converter": "org.apache.kafka.connect.storage.StringConverter", "topic.prefix": "salesforce.test", "max.retries": 3, "salesforce.api.version": "v65.0", "salesforce.client.id": "YOUR_CLIENT_ID", "salesforce.client.secret": "YOUR_CLIENT_SECRET", "salesforce.oauth.uri": "https://login.salesforce.com/services/oauth2/token", "salesforce.uri": "https://YOUR_INSTANCE.my.salesforce.com", "salesforce.bulk.api.queries": "SELECT Id, Name, LastModifiedDate FROM Account;", "salesforce.soql.query.wait": 3600, "salesforce.status.check.wait": 120 } ``` Parameters: * `name`: A unique name for the connector. * `connector.class`: The Java class for the connector, `io.aiven.kafka.connect.salesforce.source.SalesforceSourceConnector`. * `tasks.max`: Must be `1`. Values greater than `1` fail validation. See [Limitations](#limitations). * `key.converter` and `value.converter`: Set both to `org.apache.kafka.connect.storage.StringConverter`. See [Output record format](#output-record-format). * `topic.prefix`: Prefix for Apache Kafka topic names when you do not set `topic`. For more information, see [Topic naming](#topic-naming). * Optional: `topic`. Apache Kafka topic that receives all records from every SOQL query. This setting overrides `topic.prefix`. For more information, see [Topic naming](#topic-naming). * `salesforce.api.version`: Salesforce API version, for example `v65.0`. * `salesforce.client.id`: Consumer Key from your Salesforce connected app. * `salesforce.client.secret`: Consumer Secret from your Salesforce connected app. * `salesforce.oauth.uri`: The OAuth token endpoint. * `salesforce.uri`: Base URL of your Salesforce instance, for example `https://YOUR_INSTANCE.my.salesforce.com`. * `salesforce.bulk.api.queries`: One or more SOQL queries separated by semicolons. See [Define SOQL queries](#define-soql-queries). * Optional: `salesforce.soql.query.wait`. Time in seconds between query executions. Longer intervals reduce Salesforce Bulk API usage. Default is `300`. * Optional: `salesforce.status.check.wait`. Time in seconds between Bulk API job status checks. Default is `120`. * Optional: `max.retries`. Number of retry attempts on failure. When retries are exhausted, the connector task fails and must be restarted manually. Default is `3`. ### Output record format[​](#output-record-format "Direct link to Output record format") The connector retrieves data from Salesforce and converts each record into a Kafka Connect record. In the example configuration, `value.converter` is set to `org.apache.kafka.connect.storage.StringConverter`. Record values are written as strings that contain JSON objects. The JSON keys match the fields in your SOQL query. Example value for `SELECT Id, Name, LastModifiedDate FROM Account`: ``` { "Id": "0015g00000XyZAbAAL", "Name": "Acme Corp", "LastModifiedDate": "2024-11-01T10:23:45.000Z" } ``` If your Apache Kafka cluster uses a schema registry, set `value.converter` to match your registry format. For example, for Avro use `io.confluent.connect.avro.AvroConverter` and set `value.converter.schema.registry.url` to your registry URL. Aiven for Apache Kafka® services use [Karapace](/docs/products/kafka/karapace.md) as the schema registry. ## Create the connector[​](#create-the-connector "Direct link to Create the connector") * Aiven Console * Aiven CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 5. In the source connectors list, select **Salesforce source connector**, and click **Get started**. 6. On the **Salesforce Source Connector** page, go to the **Common** tab. 7. Locate the **Connector configuration** text box and click **Edit**. 8. Paste the configuration from your `salesforce_source_connector.json` file into the text box. 9. Click **Create connector**. 10. Verify the connector status on the **Manage stream** > **Connectors** page. To create the Salesforce source connector using the [Aiven CLI](/docs/tools/cli/service/connector.md#avn_service_connector_create), run: ``` avn service connector create SERVICE_NAME @salesforce_source_connector.json ``` To check the connector status, run: ``` avn service connector status SERVICE_NAME CONNECTOR_NAME ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka service. * `@salesforce_source_connector.json`: Path to your JSON configuration file. * `CONNECTOR_NAME`: Value of the `name` field in the JSON file—the same connector name the `connector status` command expects. Verify that data flows to the expected Apache Kafka topics. ## Verify the connector[​](#verify-the-connector "Direct link to Verify the connector") After you create the connector, confirm that records reach the topics you expect. In the example configuration in [Create a Salesforce source connector configuration file](#create-a-salesforce-source-connector-configuration-file), `Account` records appear in `salesforce.test.bulkApi.Account`. For more information, see [Topic naming](#topic-naming). Related pages * [Salesforce connector for Apache Kafka on GitHub](https://github.com/Aiven-Open/salesforce-connector-for-apache-kafka) * [Salesforce Bulk API 2.0 Developer Guide](https://developer.salesforce.com/docs/atlas.en-us.api_asynch.meta/api_asynch/bulk_api_2_0.htm) --- # Configure the Iceberg sink connector with Snowflake Open Catalog [Snowflake Open Catalog](https://other-docs.snowflake.com/en/opencatalog/overview) is a managed [Apache Polaris™](https://polaris.apache.org/) service that supports the [Apache Iceberg™](https://iceberg.apache.org/) REST catalog API and uses Amazon S3 for storage. Use the Iceberg sink connector in Aiven for Apache Kafka® Connect to write data to Iceberg tables managed by Snowflake Open Catalog. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A [Snowflake Open Catalog account](https://other-docs.snowflake.com/en/opencatalog/create-open-catalog-account) * A Polaris catalog created using [Amazon S3 as the storage backend](https://other-docs.snowflake.com/en/opencatalog/create-catalog#create-a-catalog-using-amazon-simple-storage-service-amazon-s3) * An [Aiven for Apache Kafka® service](/docs/products/kafka/kafka-connect/howto/enable-connect.md) with Apache Kafka Connect enabled, or a [dedicated Aiven for Apache Kafka Connect® service](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster) * An Apache Kafka topic that contains data to sink * An Apache Kafka control topic (default: `control-iceberg`) * AWS S3 credentials and a configured S3 bucket * Polaris authentication credentials * `iceberg.catalog.scope`: A role scope defined in the Open Catalog * `iceberg.catalog.credential`: A token or hash for authenticating with the catalog note Snowflake Open Catalog does not support automatic catalog creation. [Create the catalog](https://other-docs.snowflake.com/en/opencatalog/create-catalog) manually in the Snowflake console. ## Create an Iceberg sink connector configuration[​](#create-an-iceberg-sink-connector-configuration "Direct link to Create an Iceberg sink connector configuration") To configure the Iceberg sink connector, define a JSON configuration with the required catalog, S3, and Kafka settings. note Loading worker properties is not supported. Use `iceberg.kafka.*` properties instead. ``` { "name": "CONNECTOR_NAME", "connector.class": "org.apache.iceberg.connect.IcebergSinkConnector", "tasks.max": "2", "topics": "KAFKA_TOPICS", "key.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "key.converter.schemas.enable": "false", "value.converter.schemas.enable": "false", "consumer.override.auto.offset.reset": "earliest", "iceberg.kafka.auto.offset.reset": "earliest", "iceberg.kafka.bootstrap.servers": "KAFKA_HOST:KAFKA_PORT", "iceberg.kafka.security.protocol": "SSL", "iceberg.kafka.ssl.key.password": "KEY_PASSWORD", "iceberg.kafka.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "iceberg.kafka.ssl.keystore.password": "KEYSTORE_PASSWORD", "iceberg.kafka.ssl.keystore.type": "PKCS12", "iceberg.kafka.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "iceberg.kafka.ssl.truststore.password": "TRUSTSTORE_PASSWORD", "iceberg.tables": "DATABASE.TABLE", "iceberg.tables.auto-create-enabled": "true", "iceberg.control.commit.interval-ms": "1000", "iceberg.control.commit.timeout-ms": "20000", "iceberg.control.topic": "CONTROL_TOPIC", "iceberg.catalog.type": "rest", "iceberg.catalog.uri": "https://SNOWFLAKE_ACCOUNT_ID.AWS_REGION.snowflakecomputing.com/polaris/api/catalog", "iceberg.catalog.scope": "PRINCIPAL_ROLE:ROLE_NAME", "iceberg.catalog.credential": "CATALOG_CREDENTIAL", "iceberg.catalog.warehouse": "CATALOG_NAME", "iceberg.catalog.client.region": "AWS_REGION", "iceberg.catalog.io-impl": "org.apache.iceberg.aws.s3.S3FileIO", "iceberg.catalog.s3.access-key-id": "AWS_ACCESS_KEY_ID", "iceberg.catalog.s3.secret-access-key": "AWS_SECRET_ACCESS_KEY", "iceberg.catalog.s3.path-style-access": "true" } ``` Parameters: * `name`: Name of the connector * `connector.class`: Set to `org.apache.iceberg.connect.IcebergSinkConnector` * `tasks.max`: Maximum number of tasks the connector can run in parallel * `topics`: Apache Kafka topics to read data from * `key.converter`: Use `org.apache.kafka.connect.json.JsonConverter` * `value.converter`: Use `org.apache.kafka.connect.json.JsonConverter` * `key.converter.schemas.enable`: Enable (`true`) or disable (`false`) schema support in the key converter * `value.converter.schemas.enable`: Enable (`true`) or disable (`false`) schema support in the value converter * `consumer.override.auto.offset.reset`: Kafka consumer offset reset policy (recommended: `earliest`) * `iceberg.kafka.auto.offset.reset`: Offset reset policy for the Iceberg internal Apache Kafka consumer * `iceberg.kafka.bootstrap.servers`: Apache Kafka broker address in `KAFKA_HOST:KAFKA_PORT` format * `iceberg.kafka.security.protocol`: Use `SSL` for secure communication * `iceberg.kafka.ssl.keystore.location`: File path to the SSL keystore * `iceberg.kafka.ssl.keystore.password`: Keystore password * `iceberg.kafka.ssl.keystore.type`: Use `PKCS12` * `iceberg.kafka.ssl.truststore.location`: File path to the truststore * `iceberg.kafka.ssl.truststore.password`: Truststore password * `iceberg.kafka.ssl.key.password`: Password for the SSL private key * `iceberg.tables`: Iceberg table name in `DATABASE.TABLE` format * `iceberg.tables.auto-create-enabled`: Enable (`true`) or disable (`false`) automatic table creation * `iceberg.control.commit.interval-ms`: Frequency (in ms) to commit data to Iceberg * `iceberg.control.commit.timeout-ms`: Max time (in ms) to wait for a commit * `iceberg.control.topic`: Control topic used by Iceberg (default: `control-iceberg`) * `iceberg.catalog.type`: Set to `rest` for Snowflake Open Catalog * `iceberg.catalog.uri`: Polaris REST catalog endpoint (from the Snowflake Open Catalog console) * `iceberg.catalog.scope`: Role scope in the format `PRINCIPAL_ROLE:ROLE_NAME` * `iceberg.catalog.credential`: Authentication credential for Snowflake Open Catalog Use the format `CLIENT_ID:CLIENT_SECRET` from the configured service connection that uses a Principal role * `iceberg.catalog.warehouse`: Name of the catalog created in Snowflake Open Catalog. This is not the S3 bucket name * `iceberg.catalog.client.region`: AWS region of the S3 bucket * `iceberg.catalog.io-impl`: Set to `org.apache.iceberg.aws.s3.S3FileIO` * `iceberg.catalog.s3.access-key-id`: AWS access key ID with write permissions * `iceberg.catalog.s3.secret-access-key`: AWS secret access key * `iceberg.catalog.s3.path-style-access`: Enable (`true`) or disable (`false`) path-style access ## Create the Iceberg sink connector[​](#create-the-iceberg-sink-connector "Direct link to Create the Iceberg sink connector") * Aiven Console * Aiven CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is enabled on the service. If not, click **Enable connector on this service**. Alternatively, to enable connectors: 1. Click **Service settings** in the sidebar. 2. In the **Service management** section, click **Actions** > **Enable Kafka connect**. 3. In the sink connectors list, select **Iceberg Sink Connector**, and click **Get started**. 4. On the **Iceberg Sink Connector** page, go to the **Common** tab. 5. Locate the **Connector configuration** text box and click **Edit**. 6. Paste the configuration from your `iceberg_sink_connector.json` file into the text box. 7. Click **Create connector**. 8. Verify the connector status on the **Connectors** page. To create the Iceberg sink connector using the [Aiven CLI](/docs/tools/cli.md), run: ``` avn service connector create SERVICE_NAME @iceberg_sink_connector.json ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka® service. * `@iceberg_sink_connector.json`: Path to the JSON configuration file. ## Example[​](#example "Direct link to Example") This example shows a complete connector configuration that uses Snowflake Open Catalog (Polaris) as the catalog and Amazon S3 for storage. ``` { "name": "iceberg_sink_polaris", "connector.class": "org.apache.iceberg.connect.IcebergSinkConnector", "tasks.max": "2", "topics": "test-topic", "iceberg.catalog.type": "rest", "iceberg.catalog.uri": "https://1234567890.us-east-1.snowflakecomputing.com/polaris/api/catalog", "iceberg.catalog.scope": "PRINCIPAL_ROLE:my-role", "iceberg.catalog.credential": "my-token", "iceberg.catalog.io-impl": "org.apache.iceberg.aws.s3.S3FileIO", "iceberg.catalog.s3.access-key-id": "your-access-key-id", "iceberg.catalog.s3.secret-access-key": "your-secret-access-key", "iceberg.catalog.warehouse": "my-bucket", "iceberg.tables": "mydatabase.mytable", "iceberg.tables.auto-create-enabled": "true", "iceberg.control.commit.interval-ms": "1000", "iceberg.control.commit.timeout-ms": "20000", "key.converter": "org.apache.kafka.connect.json.JsonConverter", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "iceberg.kafka.bootstrap.servers": "kafka.example.com:9092", "iceberg.kafka.security.protocol": "SSL", "iceberg.kafka.ssl.keystore.location": "/run/aiven/keys/public.keystore.p12", "iceberg.kafka.ssl.keystore.password": "password", "iceberg.kafka.ssl.keystore.type": "PKCS12", "iceberg.kafka.ssl.truststore.location": "/run/aiven/keys/public.truststore.jks", "iceberg.kafka.ssl.truststore.password": "password", "iceberg.kafka.ssl.key.password": "password" } ``` Related pages * [Iceberg sink connector overview](/docs/products/kafka/kafka-connect/howto/iceberg-sink-connector.md) * [Snowflake Open Catalog documentation](https://other-docs.snowflake.com/en/opencatalog) * [Iceberg REST catalog specification](https://iceberg.apache.org/spec/#rest-catalog) --- # Create and configure a Snowflake sink connector for Apache Kafka® The Apache Kafka Connect® Snowflake sink connector moves data from Aiven for Apache Kafka® topics to a Snowflake database. It requires configuration in both Snowflake and Aiven for Apache Kafka. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven for Apache Kafka service with [Apache Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md), or a [dedicated Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster) * Access to the target Snowflake account with privileges to create users, roles, and schemas * OpenSSL installed locally to generate key pairs * Collect the following connection details: * `SNOWFLAKE_URL`: In the format `ACCOUNT_LOCATOR.REGION_ID.snowflakecomputing.com` tip To retrieve your account locator and region ID, run the following in the Snowflake worksheet: ``` SELECT CURRENT_ACCOUNT(), CURRENT_REGION(); ``` * `SNOWFLAKE_USERNAME`: The user created for the connector * `SNOWFLAKE_PRIVATE_KEY`: The private key for the user * `SNOWFLAKE_PRIVATE_KEY_PASSPHRASE`: The key passphrase * `SNOWFLAKE_DATABASE`: The target database * `SNOWFLAKE_SCHEMA`: The target schema * `TOPIC_LIST`: Comma-separated list of Kafka topics to sink If using Avro format: * `APACHE_KAFKA_HOST` * `SCHEMA_REGISTRY_PORT` * `SCHEMA_REGISTRY_USER` * `SCHEMA_REGISTRY_PASSWORD` note For a full list of configuration options, see the [Snowflake Kafka Connector documentation](https://docs.snowflake.com/en/user-guide/kafka-connector). ## Configure Snowflake[​](#configure-snowflake "Direct link to Configure Snowflake") Set up Snowflake to authenticate the connector using key pair authentication. ### Generate a key pair[​](#generate-a-key-pair "Direct link to Generate a key pair") Use OpenSSL to generate a 2048-bit RSA key pair: ``` openssl genrsa 2048 | openssl pkcs8 -topk8 -inform PEM -out rsa_key.p8 openssl rsa -in rsa_key.p8 -pubout -out rsa_key.pub ``` * `rsa_key.p8` contains the private key, secured with a passphrase. * `rsa_key.pub` contains the public key. ### Create a user[​](#create-a-user "Direct link to Create a user") 1. Open the Snowflake UI and go to the **Worksheets** tab. 2. Use a role with `SECURITYADMIN` or `ACCOUNTADMIN` privileges. 3. Create a Snowflake user for the Aiven connector. Run the following SQL command: ``` CREATE USER aiven; ``` 4. Copy the contents of `rsa_key.pub`, excluding the `-----BEGIN` and `-----END` lines. Remove any line breaks. note When copying the public key, exclude the `-----BEGIN PUBLIC KEY-----` and `-----END PUBLIC KEY-----` lines. Remove all line breaks so the key is on a single line. 5. Set the public key for the user: ``` ALTER USER aiven SET RSA_PUBLIC_KEY='PASTE_PUBLIC_KEY_HERE'; ``` ### Create a role and assign it to the user[​](#create-a-role-and-assign-it-to-the-user "Direct link to Create a role and assign it to the user") 1. Create a role for the connector: ``` CREATE ROLE aiven_snowflake_sink_connector_role; ``` 2. Grant the role to the `aiven` user: ``` GRANT ROLE aiven_snowflake_sink_connector_role TO USER aiven; ``` 3. Set the role as the user's default: ``` ALTER USER aiven SET DEFAULT_ROLE = aiven_snowflake_sink_connector_role; ``` ### Grant privileges on the target database and schema[​](#grant-privileges-on-the-target-database-and-schema "Direct link to Grant privileges on the target database and schema") The connector writes data to tables in a specific schema within a Snowflake database. Grant the necessary privileges to the role you created. 1. In the Snowflake UI, open the **Worksheets** tab. 2. Use a role with `SECURITYADMIN` or `ACCOUNTADMIN` privileges. 3. Replace `TESTDATABASE` and `TESTSCHEMA` with your database and schema names, then run the following SQL commands: ``` GRANT USAGE ON DATABASE TESTDATABASE TO ROLE aiven_snowflake_sink_connector_role; GRANT USAGE ON SCHEMA TESTDATABASE.TESTSCHEMA TO ROLE aiven_snowflake_sink_connector_role; GRANT CREATE TABLE ON SCHEMA TESTDATABASE.TESTSCHEMA TO ROLE aiven_snowflake_sink_connector_role; GRANT CREATE STAGE ON SCHEMA TESTDATABASE.TESTSCHEMA TO ROLE aiven_snowflake_sink_connector_role; GRANT CREATE PIPE ON SCHEMA TESTDATABASE.TESTSCHEMA TO ROLE aiven_snowflake_sink_connector_role; ``` These privileges allow the connector to access the database, write to the schema, and manage tables, stages, and pipes. ## Create a Snowflake sink connector configuration[​](#create-a-snowflake-sink-connector-configuration "Direct link to Create a Snowflake sink connector configuration") To configure the Snowflake sink connector, define a JSON file (for example, `snowflake_sink.json`) with the required connection and converter settings. This helps organize your configuration and makes it easier to use in the Aiven Console or CLI. ``` { "name": "my-test-snowflake", "connector.class": "com.snowflake.kafka.connector.SnowflakeSinkConnector", "topics": "TOPIC_LIST", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "snowflake.url.name": "SNOWFLAKE_URL", "snowflake.user.name": "SNOWFLAKE_USERNAME", "snowflake.private.key": "SNOWFLAKE_PRIVATE_KEY", "snowflake.private.key.passphrase": "SNOWFLAKE_PRIVATE_KEY_PASSPHRASE", "snowflake.database.name": "SNOWFLAKE_DATABASE", "snowflake.schema.name": "SNOWFLAKE_SCHEMA" } ``` The configuration file includes the following parameters: * `name`: Specifies the connector name. * `topics`: Lists the Apache Kafka topics to write to the Snowflake database. * `key.converter` and `value.converter`: Specify the format of messages in the Kafka topic. To use Avro, set these to `io.confluent.connect.avro.AvroConverter`. To retrieve the message schema, use Aiven's [Karapace schema registry](https://github.com/aiven/karapace). Set the following parameters when using Avro as the message format: * `key.converter.schema.registry.url` and `value.converter.schema.registry.url`: Specify the schema registry URL in the format `https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT`. Replace `APACHE_KAFKA_HOST` and `SCHEMA_REGISTRY_PORT` with the values retrieved in the [prerequisites section](#prerequisites). * `key.converter.basic.auth.credentials.source` and `value.converter.basic.auth.credentials.source`: Set to `USER_INFO` to enable basic authentication using a username and password. * `key.converter.schema.registry.basic.auth.user.info` and `value.converter.schema.registry.basic.auth.user.info`: Specify the schema registry credentials in the format `SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD`. Replace these values with the credentials retrieved in the [prerequisites section](#prerequisites). These converter parameters are required for parsing Avro messages from the Kafka topics. Snowflake connection parameters: * `snowflake.url.name`: Specifies the URL of the Snowflake service. * `snowflake.user.name`: Specifies the Snowflake username. * `snowflake.private.key`: Specifies the private key used to authenticate the Snowflake user. * `snowflake.private.key.passphrase`: Specifies the passphrase for the private key. * `snowflake.database.name`: Specifies the target Snowflake database. * `snowflake.schema.name`: Specifies the target schema within the database. ## Create a Snowflake sink connector[​](#create-a-snowflake-sink-connector "Direct link to Create a Snowflake sink connector") * Aiven Console * Aiven CLI 1. Access the [Aiven Console](https://console.aiven.io/). 2. Select your Aiven for Apache Kafka or Aiven for Apache Kafka Connect service. 3. Click **Manage stream** > **Connectors**. 4. Click **Create connector** if Apache Kafka Connect is already enabled on the service. If not, click **Enable connector on this service**. To enable connectors: 1. Click **Service settings** in the sidebar. 2. In the Service management section, click **Actions** > **Enable Kafka Connect**. 5. In the list of sink connectors, click **Get started** under **Snowflake Sink**. 6. On the **Snowflake Sink** connector page, go to the **Common** tab. 7. Locate the **Connector configuration** text box and click **Edit**. 8. Paste the configuration from your `snowflake_sink.json` file into the text box. Replace all placeholders with actual values. 9. Click **Apply**. 10. Click **Create connector**. 11. Verify the connector status on the **Manage stream** > **Connectors** page. 12. Confirm that data from the Apache Kafka topics appears in the target Snowflake database. To create a Snowflake sink connector using the [Aiven CLI](/docs/tools/cli/service-cli.md), run: ``` avn service connector create SERVICE_NAME @snowflake_sink.json ``` Parameters: * `SERVICE_NAME`: The name of your Aiven for Apache Kafka service. * `@snowflake_sink.json`: The path to your connector configuration file. ## Examples[​](#examples "Direct link to Examples") ### Create a Snowflake sink connector for an Avro topic[​](#create-a-snowflake-sink-connector-for-an-avro-topic "Direct link to Create a Snowflake sink connector for an Avro topic") This example creates a Snowflake sink connector with the following settings: * Connector name: `my_snowflake_sink` * Source topic: `test` * Snowflake database: `testdb` * Snowflake schema: `testschema` * Snowflake URL: `XX0000.eu-central-1.snowflakecomputing.com` * Snowflake user: `testuser` * Private key: ``` XXXXXXXYYY ZZZZZZZZZZ KKKKKKKKKK YY ``` * Private key passphrase: `password123` ``` { "name": "my_snowflake_sink", "connector.class": "com.snowflake.kafka.connector.SnowflakeSinkConnector", "key.converter": "io.confluent.connect.avro.AvroConverter", "key.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "key.converter.basic.auth.credentials.source": "USER_INFO", "key.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "value.converter": "io.confluent.connect.avro.AvroConverter", "value.converter.schema.registry.url": "https://APACHE_KAFKA_HOST:SCHEMA_REGISTRY_PORT", "value.converter.basic.auth.credentials.source": "USER_INFO", "value.converter.schema.registry.basic.auth.user.info": "SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD", "topics": "test", "snowflake.url.name": "XX0000.eu-central-1.snowflakecomputing.com", "snowflake.user.name": "testkafka", "snowflake.private.key": "XXXXXXXYYYZZZZZZZZZZKKKKKKKKKKYY", "snowflake.private.key.passphrase": "password123", "snowflake.database.name": "testdb", "snowflake.schema.name": "testschema" } ``` ### Create a Snowflake sink connector for a topic with a JSON schema[​](#create-a-snowflake-sink-connector-for-a-topic-with-a-json-schema "Direct link to Create a Snowflake sink connector for a topic with a JSON schema") This example sinks data from a Kafka topic named `iot_measurements` to Snowflake. The topic contains messages in JSON format, each with an inline schema. Sample messages: ``` { "schema": { "type": "struct", "fields": [ { "type": "int64", "optional": false, "field": "iot_id" }, { "type": "string", "optional": false, "field": "metric" }, { "type": "int32", "optional": false, "field": "measurement" } ] }, "payload": { "iot_id": 1, "metric": "Temperature", "measurement": 14 } } { "schema": { "type": "struct", "fields": [ { "type": "int64", "optional": false, "field": "iot_id" }, { "type": "string", "optional": false, "field": "metric" }, { "type": "int32", "optional": false, "field": "measurement" } ] }, "payload": { "iot_id": 2, "metric": "Humidity", "measurement": 60 } } ``` note JSON messages must include the full schema in each record. This results in additional overhead. For better performance, use the Avro format with the [Karapace schema registry](https://karapace.io/) provided by Aiven. You can configure the Snowflake sink connector using the following JSON configuration. Replace the placeholders with your actual values. ``` { "name": "my-test-snowflake-1", "connector.class": "com.snowflake.kafka.connector.SnowflakeSinkConnector", "value.converter": "org.apache.kafka.connect.json.JsonConverter", "topics": "iot_measurements", "snowflake.url.name": "SNOWFLAKE_URL", "snowflake.user.name": "SNOWFLAKE_USERNAME", "snowflake.private.key": "SNOWFLAKE_PRIVATE_KEY", "snowflake.private.key.passphrase": "SNOWFLAKE_PRIVATE_KEY_PASSPHRASE", "snowflake.database.name": "SNOWFLAKE_DATABASE", "snowflake.schema.name": "SNOWFLAKE_SCHEMA" } ``` Details about the configuration: * `"topics": "iot_measurements"`: Specifies the Kafka topic to write to Snowflake. * `"value.converter": "org.apache.kafka.connect.json.JsonConverter"`: Parses message values from JSON with an inline schema. * No key converter is defined because the message key is empty. --- # Create a sink connector from Apache Kafka® to Splunk The [Splunk](https://www.splunk.com/) sink connector enables you to move data from an Aiven for Apache Kafka® cluster to a remote Splunk server via [HTTP event collector](https://docs.splunk.com/Documentation/Splunk/latest/Data/FormateventsforHTTPEventCollector) (HEC). note See the full set of available parameters and configuration options in the [connector's documentation](https://github.com/splunk/kafka-connect-splunk). ## Prerequisites[​](#connect_splunk_sink_prereq "Direct link to Prerequisites") To setup an Splunk sink connector, you need an Aiven for Apache Kafka service [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md) or a [dedicated Aiven for Apache Kafka Connect cluster](/docs/products/kafka/kafka-connect/get-started.md#apache_kafka_connect_dedicated_cluster). Also collect the following information about the target server: * `SPLUNK_HEC_TOKEN`: The [HEC authentication token](https://docs.splunk.com/Documentation/Splunk/latest/Data/FormateventsforHTTPEventCollector) * `SPLUNK_HEC_URI`: The [Splunk endpoint URI](https://docs.splunk.com/Documentation/Splunk/9.0.1/Data/UsetheHTTPEventCollector) * `TOPIC_LIST`: The list of topics to sink divided by comma * `SPLUNK_INDEXES`: The list of Splunk indexes where the data will be landing and, if you are using Avro as the data format: * `APACHE_KAFKA_HOST`: The hostname of the Apache Kafka service * `SCHEMA_REGISTRY_PORT`: The Apache Kafka's schema registry port * `SCHEMA_REGISTRY_USER`: The Apache Kafka's schema registry username * `SCHEMA_REGISTRY_PASSWORD`: The Apache Kafka's schema registry user password note You can browse the additional parameters available for the `static` and `oauth2` authorization types in the [dedicated documentation](https://github.com/aiven/http-connector-for-apache-kafka/blob/main/docs/sink-connector-config-options.rst). ## Setup an Splunk sink connector with Aiven Console[​](#setup-an-splunk-sink-connector-with-aiven-console "Direct link to Setup an Splunk sink connector with Aiven Console") The following example demonstrates how to setup an Splunk sink connector for Apache Kafka using the [Aiven Console](https://console.aiven.io/). ### Define a Kafka Connect configuration file[​](#define-a-kafka-connect-configuration-file "Direct link to Define a Kafka Connect configuration file") Create a file (we'll refer to this one as `splunk_sink.json`) to hold the connector configuration. As an example, see the following configuration for sending JSON payloads to Splunk: ``` { "name":"CONNECTOR_NAME", "connector.class": "com.splunk.kafka.connect.SplunkSinkConnector", "splunk.hec.token": "SPLUNK_HEC_TOKEN", "splunk.hec.uri": "SPLUNK_HEC_URI", "splunk.indexes": "SPLUNK_INDEXES", "topics": "TOPIC_LIST", "splunk.hec.raw" : false, "splunk.hec.ack.enabled" : false, "splunk.hec.ssl.validate.certs": "true", "config.splunk.hec.json.event.formatted": false, "tasks.max":1 } ``` The configuration file contains the following entries: * `name`: the connector name * `splunk.hec.token` and `splunk.hec.uri`: remote Splunk server URI and authorization parameters collected in the [prerequisite](/docs/products/kafka/kafka-connect/howto/splunk-sink.md#connect_splunk_sink_prereq) phase. * `splunk.hec.raw`: if set to `false` defines the data ingestion using the `/raw` HEC endpoint instead of the default `/event` one. * `splunk.hec.ack.enabled`: if set to `true`, Kafka offset is updated only after receiving the ACK for the POST call to Splunk. * `config.splunk.hec.json.event.formatted`: Defines if events are preformatted into the proper [HEC JSON format](https://docs.splunk.com/Documentation/KafkaConnect/2.0.2/User/Parameters). tip When using Splunk with self service SSL certificates it can be useful to set `splunk.hec.ssl.validate.certs` to `false` to disable HTTPS certification validation. ### Create a Kafka Connect connector with the Aiven Console[​](#create-a-kafka-connect-connector-with-the-aiven-console "Direct link to Create a Kafka Connect connector with the Aiven Console") To create an Apache Kafka Connect connector: 1. Log in to the [Aiven Console](https://console.aiven.io/) and select the Aiven for Apache Kafka® or Aiven for Apache Kafka Connect® service where the connector needs to be defined. 2. Select **Manage stream** > **Connectors** from the left sidebar. 3. Select **Create New Connector**, it is enabled only for services [with Kafka Connect enabled](/docs/products/kafka/kafka-connect/howto/enable-connect.md). 4. Select **Splunk sink**. 5. In the **Common** tab, locate the **Connector configuration** text box and select on **Edit**. 6. Paste the connector configuration (stored in the `splunk_sink.json` file) in the form. 7. Select **Apply**. note The Aiven Console parses the configuration file and fills the relevant UI fields. You can review the UI fields across the various tabs and change them if necessary. The changes will be reflected in JSON format in the **Connector configuration** text box. 8. After all the settings are correctly configured, select **Create connector**. 9. Verify the connector status under **Manage stream** > **Connectors**. 10. Verify the data in the target Splunk instance. note You can also create connectors using the [Aiven CLI command](/docs/tools/cli/service/connector.md#avn_service_connector_create). ## Example: Create a simple Splunk sink connector[​](#example-create-a-simple-splunk-sink-connector "Direct link to Example: Create a simple Splunk sink connector") If you have a topic named `data_logs` to sink to a Splunk server in the `kafka_logs` index: ``` { "name":"data_logs_splunk_sink", "connector.class": "com.splunk.kafka.connect.SplunkSinkConnector", "splunk.hec.token": "SPLUNK_HEC_TOKEN", "splunk.hec.uri": "SPLUNK_HEC_URI", "splunk.indexes": "kafka_logs", "topics": "data_logs" } ``` The configuration file contains the following things to note: * `"topics": "data_logs"`: setting the topic to sink --- # Advanced parameters for Apache Kafka® Connect See the configuration options available for Apache Kafka® Connect: | Parameter | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | []()[**ip\_filter**](#ip_filter)`array`- default: `0.0.0.0/0,::/0`IP filterAllow incoming connections from CIDR address block, e.g. '10.20.0.0/16' | | []()[**preferred\_zones**](#preferred_zones)`array,null`List of preferred zone IDs for service node placement. Nodes will be placed in these zones when available. If a specified zone is unavailable (e.g., due to capacity constraints), nodes will be placed in other available zones to maintain the configured number of zones for availability. Invalid zone IDs are rejected at configuration time. Zone IDs are cloud-specific: AWS uses zone IDs like 'euc1-az1', GCP uses zone names like 'europe-west1-a', and Azure uses 'location/zone' format like 'germanywestcentral/1'. If single\_zone is enabled with an availability\_zone, that setting takes precedence over preferred\_zones. Changes take effect on next node recreation (e.g., maintenance or plan change). For eligible plans, nodes outside preferred zones are automatically rebalanced once per day. | | []()[**service\_log**](#service_log)`boolean,null`Service loggingStore logs for the service so that they are available in the HTTP API and console. | | []()[**static\_ips**](#static_ips)`boolean`Use static public IP addresses | | []()[**kafka\_connect**](#kafka_connect)`object`Kafka Connect configuration valueskafka\_connect.prefer\_ipv6\_address\_enable boolean When enabled, connectors will automatically resolve IPv6 addresses from external server names configured with dual-stack. kafka\_connect.connector\_client\_config\_override\_policy string Client config override policy Defines what client configurations can be overridden by the connector. Default is None kafka\_connect.consumer\_auto\_offset\_reset string Consumer auto offset reset What to do when there is no initial offset in Kafka or if the current offset does not exist any more on the server. Default is earliest kafka\_connect.consumer\_fetch\_max\_bytes integer - min: 1048576 - max: 104857600 The maximum amount of data the server should return for a fetch request Records are fetched in batches by the consumer, and if the first record batch in the first non-empty partition of the fetch is larger than this value, the record batch will still be returned to ensure that the consumer can make progress. As such, this is not a absolute maximum. kafka\_connect.consumer\_isolation\_level string Consumer isolation level Transaction read isolation level. read\_uncommitted is the default, but read\_committed can be used if consume-exactly-once behavior is desired. kafka\_connect.consumer\_max\_partition\_fetch\_bytes integer - min: 1048576 - max: 104857600 The maximum amount of data per-partition the server will return. Records are fetched in batches by the consumer.If the first record batch in the first non-empty partition of the fetch is larger than this limit, the batch will still be returned to ensure that the consumer can make progress. kafka\_connect.consumer\_max\_poll\_interval\_ms integer - min: 1 - max: 2147483647 The maximum delay in milliseconds between invocations of poll() when using consumer group management (defaults to 300000). kafka\_connect.consumer\_max\_poll\_records integer - min: 1 - max: 10000 The maximum number of records returned in a single call to poll() (defaults to 500). kafka\_connect.offset\_flush\_interval\_ms integer - min: 1 - max: 100000000 The interval at which to try committing offsets for tasks (defaults to 60000). kafka\_connect.offset\_flush\_timeout\_ms integer - min: 1 - max: 2147483647 Maximum number of milliseconds to wait for records to flush and partition offset data to be committed to offset storage before cancelling the process and restoring the offset data to be committed in a future attempt (defaults to 5000). kafka\_connect.producer\_batch\_size integer - max: 5242880 The batch size in bytes the producer will attempt to collect for the same partition before publishing to broker This setting gives the upper bound of the batch size to be sent. If there are fewer than this many bytes accumulated for this partition, the producer will 'linger' for the linger.ms time waiting for more records to show up. A batch size of zero will disable batching entirely (defaults to 16384). kafka\_connect.producer\_buffer\_memory integer - min: 5242880 - max: 134217728 The total bytes of memory the producer can use to buffer records waiting to be sent to the broker (defaults to 33554432). kafka\_connect.producer\_compression\_type string Specify the default compression type for producers. This configuration accepts the standard compression codecs ('gzip', 'snappy', 'lz4', 'zstd'). It additionally accepts 'none' which is the default and equivalent to no compression. kafka\_connect.producer\_linger\_ms integer - max: 5000 Wait for up to the given delay to allow batching records together This setting gives the upper bound on the delay for batching: once there is batch.size worth of records for a partition it will be sent immediately regardless of this setting, however if there are fewer than this many bytes accumulated for this partition the producer will 'linger' for the specified time waiting for more records to show up. Defaults to 0. kafka\_connect.producer\_max\_request\_size integer - min: 131072 - max: 67108864 The maximum size of a request in bytes This setting will limit the number of record batches the producer will send in a single request to avoid sending huge requests. kafka\_connect.scheduled\_rebalance\_max\_delay\_ms integer - max: 600000 The maximum delay that is scheduled in order to wait for the return of one or more departed workers before rebalancing and reassigning their connectors and tasks to the group. During this period the connectors and tasks of the departed workers remain unassigned. Defaults to 5 minutes. kafka\_connect.session\_timeout\_ms integer - min: 1 - max: 2147483647 The timeout in milliseconds used to detect failures when using Kafka’s group management facilities (defaults to 10000). | | []()[**secret\_providers**](#secret_providers)`array`Kafka Connect secret providersConfigure external secret providers in order to reference external secrets in connector configuration. Currently Hashicorp Vault (provider: vault, auth\_method: token) and AWS Secrets Manager (provider: aws, auth\_method: credentials) are supported. Secrets can be referenced in connector config with ${\:\:\} | | []()[**plugin\_versions**](#plugin_versions)`array`Kafka Connect pluginsThe plugin selected by the user | | []()[**sasl\_oauthbearer\_allowed\_urls**](#sasl_oauthbearer_allowed_urls)`array`List of allowed URLs for SASL OAUTHBEARER authentication. Only HTTPS URLs are allowed for security reasons. | | []()[**gcp\_auth\_allowed\_urls**](#gcp_auth_allowed_urls)`array`Allowed URLs for auth URI validationAllow-list of HTTPS URLs used to validate GCP credential\_source requests for Kafka Connect. | | []()[**private\_access**](#private_access)`object`Allow access to selected service ports from private networksprivate\_access.kafka\_connect boolean Allow clients to connect to kafka\_connect with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.prometheus boolean Allow clients to connect to prometheus with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | []()[**privatelink\_access**](#privatelink_access)`object`Allow access to selected service components through Privatelinkprivatelink\_access.jolokia boolean Enable jolokia privatelink\_access.kafka\_connect boolean Enable kafka\_connect privatelink\_access.prometheus boolean Enable prometheus | | []()[**public\_access**](#public_access)`object`Allow access to selected service ports from the public Internetpublic\_access.kafka\_connect boolean Allow clients to connect to kafka\_connect from the public internet for service nodes that are in a project VPC or another type of private network public\_access.prometheus boolean Allow clients to connect to prometheus from the public internet for service nodes that are in a project VPC or another type of private network | --- # Aiven for Apache Kafka® Connect metrics available via Prometheus Discover metrics offered by Prometheus for the Aiven for Apache Kafka® Connect service. note The metrics only appear if there is activity in the underlying Apache Kafka Connect service. | Metric | Description | | ----------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------- | | `kafka_connect_connect_worker_metrics_connector_count` | Number of connectors run in this worker | | `kafka_connect_connect_worker_metrics_connector_startup_attempts_total` | Total number of connector startups that this worker has attempted | | `kafka_connect_connect_worker_metrics_connector_startup_failure_percentage` | Average percentage of this worker's connector starts that failed | | `kafka_connect_connect_worker_metrics_connector_startup_failure_total` | Total number of connector starts that failed | | `kafka_connect_connect_worker_metrics_connector_startup_success_percentage` | Average percentage of this worker's connector starts that succeeded | | `kafka_connect_connect_worker_metrics_connector_startup_success_total` | Total number of connector starts that succeeded | | `kafka_connect_connect_worker_metrics_task_count` | Number of tasks run in this worker | | `kafka_connect_connect_worker_metrics_task_startup_attempts_total` | Total number of task startups that this worker has attempted | | `kafka_connect_connect_worker_metrics_task_startup_failure_percentage` | Average percentage of this worker's task starts that failed | | `kafka_connect_connect_worker_metrics_task_startup_failure_total` | Total number of task starts that failed | | `kafka_connect_connect_worker_metrics_task_startup_success_percentage` | Average percentage of this worker's task starts that succeeded | | `kafka_connect_connect_worker_metrics_task_startup_success_total` | Total number of task starts that succeeded | | `kafka_connect_connect_worker_rebalance_metrics_completed_rebalances_total` | Total number of rebalances completed by this worker | | `kafka_connect_connect_worker_rebalance_metrics_epoch` | Epoch or generation number of this worker | | `kafka_connect_connect_worker_rebalance_metrics_rebalance_avg_time_ms` | Average time in milliseconds spent by this worker to rebalance | | `kafka_connect_connect_worker_rebalance_metrics_rebalancing` | Whether this worker is currently rebalancing | | `kafka_connect_connect_worker_rebalance_metrics_time_since_last_rebalance_ms` | Time in milliseconds since this worker completed the most recent rebalance | | `kafka_connect:connector_task_metrics_info` | Aggregated information about the connector tasks, such as their statuses and other metadata. | | `kafka.connect:connector-metrics_info` | Aggregated information about the connectors, including statuses and metadata. | | `kafka_connect_connector_metrics_connector_class` | The class name of the connector | | `kafka_connect_connector_metrics_connector_version` | The version of the connector | | `kafka_connect_connector_metrics_connector_type` | The type of the connector, for example, source or sink | | `kafka_connect_connector_metrics_status` | The status of the connector, for example, running, paused | | `kafka_connect_connect_worker_rebalance_metrics_failed_rebalances_total` | Total number of rebalances that failed | | `kafka_connect_connect_worker_rebalance_metrics_rebalance_max_time_ms` | Maximum time in milliseconds taken by this worker to rebalance | | `kafka_connect_connect_worker_rebalance_metrics_rebalance_min_time_ms` | Minimum time in milliseconds taken by this worker to rebalance | Related pages * [Aiven for Apache Kafka® metrics available via Prometheus](/docs/products/kafka/reference/kafka-metrics-prometheus.md) * [Connect Monitoring section of the Apache Kafka® documentation](https://kafka.apache.org/42/operations/monitoring/#connect-monitoring) --- # Aiven for Apache Kafka® MirrorMaker 2 Aiven for Apache Kafka® MirrorMaker 2 is a fully managed **distributed Apache Kafka® data replication utility**, deployable in the cloud of your choice. Apache Kafka® MirrorMaker 2 lets you sync your topic data across Apache Kafka® cluster deployed anywhere in the world. With an Apache Kafka® MirrorMaker 2, you can define replication flows to keep a set of topics in sync between multiple Apache Kafka® clusters. ## Why Apache Kafka® MirrorMaker 2[​](#why-apache-kafka-mirrormaker-2 "Direct link to Why Apache Kafka® MirrorMaker 2") Apache Kafka® represents the best in class data streaming solution. Apache Kafka® MirrorMaker 2 allows cross-cluster replication providing ways to connect Apache Kafka® clusters deployed in different parts of the world and keep their topic data in sync, very useful for disaster recovery or in case of particular geo-distribution requirements. Take your first steps with Aiven for Apache Kafka® MirrorMaker 2 by following our [Getting started](/docs/products/kafka/kafka-mirrormaker/get-started.md). ## Apache Kafka® MirrorMaker 2 resources[​](#apache-kafka-mirrormaker-2-resources "Direct link to Apache Kafka® MirrorMaker 2 resources") If you are new to Apache Kafka® MirrorMaker 2, see these resources: * The main Apache Kafka® project page: * The Karapace schema registry that Aiven maintains and makes available for every Aiven for Apache Kafka® service: * Our code samples repository: --- # Configuration and tuning for Aiven for Apache Kafka® MirrorMaker 2 Learn where Aiven for Apache Kafka® MirrorMaker 2 settings are configured across service, replication-flow, and integration layers, which parameters affect performance, and what restarts when you change them. ## Configuration layers[​](#configuration-layers "Direct link to Configuration layers") Aiven for Apache Kafka® MirrorMaker 2 uses three configuration layers. Each layer controls a different part of the replication process and has a different restart impact. * **Service configurations** * **Replication-flow configurations** * **Integration configurations** ### Service configurations[​](#service-configurations "Direct link to Service configurations") Service configurations control the behavior of nodes and workers in the MirrorMaker 2 cluster. **Example** * **Parameter:** [`kafka_mirrormaker.emit_checkpoints_enabled`](/docs/products/kafka/kafka-mirrormaker/reference/advanced-params.md#kafka_mirrormaker_emit_checkpoints_enabled) * **Description:** Enables or disables periodic emission of consumer group offset checkpoints to the target cluster. * **Impact:** * Restarts workers * Restarts all connectors and tasks ### Replication-flow configurations[​](#replication-flow-configurations "Direct link to Replication-flow configurations") Replication-flow configurations control the behavior of connectors such as Source, Sink, Checkpoint, and Heartbeat connectors. **Example** * **Parameter:** [`topics`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/mirrormaker_replication_flow) * **Description:** Specifies a list of topics or regular expressions to replicate. For more information, see [Topics included in a replication flow](/docs/products/kafka/kafka-mirrormaker/concepts/replication-flow-topics-regex.md). * **Impact:** * Restarts the affected connectors * Restarts their tasks ### Integration configurations[​](#integration-configurations "Direct link to Integration configurations") Integration configurations refine how producers and consumers behave within connectors. **Example** * **Parameter:** [`consumer_fetch_min_bytes`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration#nested-schema-for-kafka_mirrormaker_user_configkafka_mirrormaker) * **Description:** Sets the minimum amount of data the server returns for a fetch request. * **Impact:** * Restarts workers * Restarts all connectors and tasks note Many configuration parameters originate from [KIP-382: MirrorMaker 2.0 configuration properties](https://cwiki.apache.org/confluence/pages/viewpage.action?pageId=95650722#KIP382:MirrorMaker2.0-ConnectorConfigurationProperties). ## Common performance-related parameters[​](#common-performance-related-parameters "Direct link to Common performance-related parameters") Some configuration parameters are commonly adjusted to improve replication throughput, consistency, or topic selection. The configuration layer determines where the parameter is set and what restarts when the value changes. ### Task allocation[​](#task-allocation "Direct link to Task allocation") Increasing the value of [`kafka_mirrormaker.tasks_max_per_cpu`](/docs/products/kafka/kafka-mirrormaker/reference/advanced-params.md#kafka_mirrormaker_tasks_max_per_cpu) in the advanced configuration can improve throughput. Set this value close to the number of partitions when you need more parallelism. ### Interval settings[​](#interval-settings "Direct link to Interval settings") Aligning interval-based settings keeps replication activity consistent. * **Advanced configurations:** * [`kafka_mirrormaker.emit_checkpoints_interval_seconds`](/docs/products/kafka/kafka-mirrormaker/reference/advanced-params.md#kafka_mirrormaker_emit_checkpoints_interval_seconds) * [`kafka_mirrormaker.sync_group_offsets_interval_seconds`](/docs/products/kafka/kafka-mirrormaker/reference/advanced-params.md#kafka_mirrormaker_sync_group_offsets_interval_seconds) * **Replication flow:** * [`sync_group_offsets_interval_seconds`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/mirrormaker_replication_flow#sync_group_offsets_interval_seconds-1) ### Topic exclusion[​](#topic-exclusion "Direct link to Topic exclusion") Adding these patterns to the topic exclusion list prevents internal and system topics from being replicated: * `.*[\-\.]internal` * `.*\.replica` * `__.*` * `connect.*` ### Producer and consumer settings[​](#producer-and-consumer-settings "Direct link to Producer and consumer settings") These [integration configuration](#integration-configurations) parameters control how MirrorMaker 2 producers and consumers interact with the source and target Kafka clusters. Set them on the service integration resource. To update these settings, see [Update integration configurations](/docs/products/kafka/kafka-mirrormaker/howto/update-integration-configurations.md). important * If you do not set a parameter, Kafka applies its built-in default. * The settings apply to MirrorMaker 2 integrations with both Aiven for Apache Kafka services and external Kafka clusters. * Exercise changes incrementally and cautiously, depending on the resources available in your service plan. | Parameter | Description | Kafka default | Maximum value | | ------------------------------------ | ---------------------------------------------------------------------------------------------------- | ------------- | ----------------------- | | `consumer_fetch_min_bytes` | Minimum amount of data the server returns for a fetch request. Higher values reduce fetch frequency. | 1 | — | | `consumer_fetch_max_bytes` | Maximum amount of data the server returns for a fetch request. | 52428800 | 100 MiB | | `consumer_fetch_max_wait_ms` | Maximum amount of time the server waits for a fetch request. | 500 | 600,000 ms (10 minutes) | | `consumer_max_partition_fetch_bytes` | Maximum amount of data per partition the server returns in a single fetch response. | 1048576 | 100 MiB | | `consumer_max_poll_records` | Maximum number of records returned in a single poll request. | 500 | — | | `consumer_receive_buffer_bytes` | Size of the TCP receive buffer for the consumer. A value of `-1` uses the OS default. | 65536 | 100 MiB | | `consumer_request_timeout_ms` | Timeout for consumer requests to the broker, in milliseconds. | 30000 | 600,000 ms (10 minutes) | | `producer_batch_size` | Maximum size of a record batch sent to a single partition, in bytes. | 16384 | — | | `producer_buffer_memory` | Total memory available to the producer for buffering records, in bytes. | 33554432 | — | | `producer_linger_ms` | Time the producer waits for additional records before sending a batch, in milliseconds. | 0 | — | | `producer_max_request_size` | Maximum size of a single producer request, in bytes. | 1048576 | — | | `producer_request_timeout_ms` | Timeout for producer requests to the broker, in milliseconds. | 30000 | 600,000 ms (10 minutes) | | `producer_send_buffer_bytes` | Size of the TCP send buffer for the producer, in bytes. A value of `-1` uses the OS default. | 131072 | 100 MiB | Related pages * [Update integration configurations](/docs/products/kafka/kafka-mirrormaker/howto/update-integration-configurations.md) * [Advanced parameters for Aiven for Apache Kafka® MirrorMaker 2](/docs/products/kafka/kafka-mirrormaker/reference/advanced-params.md) * [Set up an Apache Kafka® MirrorMaker 2 replication flow](/docs/products/kafka/kafka-mirrormaker/howto/setup-replication-flow.md) * [Topics included in a replication flow](/docs/products/kafka/kafka-mirrormaker/concepts/replication-flow-topics-regex.md) * [Known issues](/docs/products/kafka/kafka-mirrormaker/reference/known-issues.md) --- # Disaster recovery and migration MirrorMaker 2 is the standard replication tool packaged with the Apache Kafka® and can be run as a managed service on the Aiven platform. MirrorMaker 2 uses Kafka Connect to consume from one Kafka cluster and produce simultaneously to a second cluster. You can use MirrorMaker 2 to enable Active-Passive disaster recovery architecture or active-active high availability architecture. ## Disaster recovery[​](#disaster-recovery "Direct link to Disaster recovery") Disaster recovery is one of the primary use cases for MirrorMaker 2. MirrorMaker 2 replicates data between Apache Kafka® clusters, including clusters located in different data centres. It's possible to minimise the period of unavailability of Apache Kafka. However, it's important to remember that **MirrorMaker 2 does not offer synchronous replication**, which means that the latest records accepted into a topic in the source cluster may not replicate to the target cluster before the source cluster fails. As MirrorMaker 2 is based on Kafka Connect, there are tradeoffs between throughput and out-of-order or duplicated records. It's important to understand that disaster recovery has no one-size-fits-all solution. Different requirements should be considered, such as the nature of the data, volume, acceptable ops overhead, and costs. ## Migration[​](#migration "Direct link to Migration") Migrating an Apache Kafka® cluster to Aiven will utilize the same pattern as disaster recovery. In this case however, you will need to use an external Kafka integration via the Aiven platform. You can use this registered integration in the replication flow to migrate data from an external Kafka cluster to an Aiven for Apache Kafka cluster. note MirrorMaker 2 can be set up using any Aiven standard provisioning method Console, API, CLI, and TF. ## Consumer requirements[​](#consumer-requirements "Direct link to Consumer requirements") * Ability to process out-of-order records * Tolerate duplicate records --- # MirrorMaker 2 active-active setup An active-active setup of MM2 allows data to be replicated between two clusters simultaneously. It implies a bidirectional mirroring between the two clusters. To do this, MM2 uses topic prefixes in the form of `.`. In this setup, there are two actively used Apache Kafka® clusters: Cluster K1 and cluster K2. Topic exists in both clusters. ![Active-active setup](/docs/assets/images/mirrormaker-active-active-ead6bb2addcd7f8cf97eeb3595ac352b.png) Each cluster has its producers and consumers. Producers produce to a topic, while consumers consume from the same topic using the same group.id. MirrorMaker 2 replicates data between clusters in both directions, so remote topics K1.topic and K2.topic exist in K2 and K1, respectively. In case a disaster happens to the K1 cluster, and it becomes inaccessible for a long time for all the clients as well as MirrorMaker 2, the replication will stop, leaving some data un-replicated in both clusters. The clients (Producers 1 and Consumers 1) switch from K1 to K2. Consumers 1 continue consuming from the remote/replicated topic K1.topic, Producers 1 start producing into a topic. When Consumers 1 finish consuming K1.topic, they switch to a new topic. All consumers act as one group. When K1 is recovered, its clients can switch back. Data produced by Producers 1 into the topic in K2 will be processed by Consumers 2. ## Disaster recovery and high availability[​](#disaster-recovery-and-high-availability "Direct link to Disaster recovery and high availability") This mode has bidirectional data mirroring between two clusters. It allows clients to produce and consume from both clusters simultaneously since data can be produced and consumed from either cluster. This setup is usually utilized for high availability, but it can also be beneficial for disaster recovery as it simplifies recovery. Since data is actively replicated between the clusters at all times, the data in both clusters is identical, making failover to either cluster easy. **Implementation details** * Consumers need to be aware of the prefixed topics and can do this using wildcards or a priority knowledge of the topics to consume from, see [replication flow](/docs/products/kafka/kafka-mirrormaker/concepts/replication-flow-topics-regex.md). * [Karapace](/docs/products/kafka/karapace.md) schemas and topic configurations are not synced and must be created in both clusters. --- # Active-passive setup In this setup, there are two Apache Kafka® clusters, the primary and secondary clusters. The primary cluster contains the *topic* topic. The "active" cluster serves all produce and consume requests, while the "passive" cluster serves as a replica of the "active" cluster without running any applications against it. ![Active-Passive Setup](/docs/assets/images/mirrormaker-active-passive-4ff0893b74e58958c24d1f572afaf96d.png) * All the clients (producers and consumers) work with primary. MirrorMaker 2 replicates the topic from primary to secondary (the remote topic name is `primary_alias.topic_name` where `primary_alias` is remote topic name). * In case a disaster happens to the primary cluster and it becomes inaccessible for a long time for all the clients as well as MirrorMaker 2, the replication stops and some data may remain un-replicated in the primary cluster. It may not reach the secondary cluster before the disaster. * The clients switch to the secondary cluster. Consumers continue consuming from the replicated topic `primary.topic`, and the producers start producing into this topic. * Alternatively, identical topics can be created in both primary and secondary clusters by setting replication flow to use [IdentityReplicationPolicy replication policy](/docs/products/kafka/kafka-mirrormaker/howto/remove-mirrormaker-prefix.md). The producers start producing into it directly, and consumers switch to it when they are done processing remote primary.topic. This approach might be more convenient when a future fallback to the primary cluster is needed. ## Disaster Recovery[​](#disaster-recovery "Direct link to Disaster Recovery") To enable a DR scenario, the backup cluster (Cluster B) must have topics with the same name as the primary cluster (Cluster A). By default, this is not the case, so all replication flows must use the [IdentityReplicationPolicy replication policy](/docs/products/kafka/kafka-mirrormaker/howto/remove-mirrormaker-prefix.md) to guarantee this. When a replication flow is created, it will mirror all topics based upon the allow list or deny list configuration of the replication flow. The allow list should be set to `.*` to guarantee that all internal topics such as consumer offsets for Apache Kafka® Connectors and consumer groups, and schemas for [Karapace](/docs/products/kafka/karapace.md). ## Failover[​](#failover "Direct link to Failover") When a cluster is replicated using MM2, it will have a different service URI and certificates that need to be considered when transitioning applications to the backup cluster instead of the primary one. Given that Cluster B is designed for DR failover, promoting Cluster A as the primary cluster after a failover scenario requires recreating the data in Cluster A from Cluster B using MM2, but in reverse. Existing topic data in Cluster A must be deleted to avoid duplicates since all data from Cluster B would be replicated. --- # Configure permissions for MirrorMaker 2 with external Kafka clusters By default, Apache Kafka® MirrorMaker 2 creates internal topics to store metadata, offsets, and status information. When connected to an external Kafka cluster through an [integration endpoint](/docs/products/kafka/howto/integrate-external-kafka-cluster.md), MirrorMaker 2 uses the credentials defined in that endpoint to authenticate. For MirrorMaker 2 to create these internal topics and write to them, the service account or credentials must have the necessary permissions. In external Kafka clusters, these permissions are managed through Access Control Lists (ACLs). ## Required permissions[​](#required-permissions "Direct link to Required permissions") To replicate topics between clusters, configure the following permissions: * On **source clusters**, the MirrorMaker 2 service account must have **READ** access to replicated topics. * On **target clusters**, the service account must have **WRITE** access to replicated topics. * On **both clusters**, the service account must have **READ** and **WRITE** access to internal topics, plus **CREATE** and **DESCRIBE** permissions to create and manage these topics. ## Source cluster internal topics[​](#source-cluster-internal-topics "Direct link to Source cluster internal topics") MirrorMaker 2 creates the following internal topics on the source cluster: | Topic | Replication factor | Partitions (default) | Cleanup policy | | ------------------------------------- | ------------------ | -------------------- | -------------- | | `heartbeats` | 3 | 1 | compact | | `mm2-configs..internal` | 3 | 1 | compact | | `mm2-offsets..internal` | 3 | 1 | compact | | `mm2-status..internal` | 3 | 5 | compact | In the table, `` represents the cluster alias you configured for the target cluster. note The `__consumer_offsets` topic is a standard Kafka topic used for consumer group management. While MirrorMaker 2 uses this topic, Kafka creates it automatically, not MirrorMaker 2. ## Target cluster internal topics[​](#target-cluster-internal-topics "Direct link to Target cluster internal topics") MirrorMaker 2 creates the following internal topics on the target cluster: | Topic | Replication factor | Partitions (default) | Cleanup policy | | ------------------------------------- | ------------------ | -------------------- | -------------- | | `.checkpoints.internal` | 3 | 1 | compact | | `.heartbeats` | 3 | 1 | compact | | `mm2-configs..internal` | 3 | 1 | compact | | `mm2-offsets..internal` | 3 | 25 | compact | | `mm2-status..internal` | 3 | 5 | compact | In the table, `` represents the cluster alias you configured for the source cluster. ## Configuration-based topics[​](#configuration-based-topics "Direct link to Configuration-based topics") Some MirrorMaker 2 settings, such as heartbeat emission, create additional internal topics that require specific permissions. If you enable these options, ensure the service account has **READ** and **WRITE** access to the corresponding topics: * **`emit_backward_heartbeats_enabled`:** When set to `true`, MirrorMaker 2 requires access to `mm2-offsets..internal` on the source cluster. If disabled, this topic and its permissions are not required on the source cluster. * **`emit_heartbeats_enabled`:** When set to `true`, MirrorMaker 2 requires access to `mm2-offsets..internal` on the target cluster. If disabled, this topic and its permissions are not required on the target cluster. ## Offset sync topic location[​](#offset-sync-topic-location "Direct link to Offset sync topic location") The `offset-syncs.topic.location` setting determines where MirrorMaker 2 creates the offset synchronization topic: * **`source`** (default): Creates `.offset-syncs` on the source cluster * **`target`**: Creates `.offset-syncs` on the target cluster The service account must have **READ** and **WRITE** access to this topic on whichever cluster you configure. note You must configure ACLs on both source and target clusters to ensure that MirrorMaker 2 can describe and create topics, as well as produce to and consume from all required topics. ## Related pages[​](#related-pages "Direct link to Related pages") * [Integrate an external Kafka cluster](/docs/products/kafka/howto/integrate-external-kafka-cluster.md) * [Set up a replication flow](/docs/products/kafka/kafka-mirrormaker/howto/setup-replication-flow.md) * [Get started with MirrorMaker 2](/docs/products/kafka/kafka-mirrormaker/get-started.md) --- # Include or exclude topics in a replication flow When you [define a replication flow](/docs/products/kafka/kafka-mirrormaker/howto/setup-replication-flow.md), specify which topics in the source Apache Kafka® cluster to include or exclude from the cross-cluster replica. Use the **Topics** field to define the topics to include. You can enter a [list of regular expressions in Java format](https://docs.oracle.com/javase/7/docs/api/java/util/regex/Pattern). If you leave **Topics** empty, MirrorMaker 2 uses `.*` and includes all topics. Use the **Topic blacklist** field to define the topics to exclude. To exclude internal topics, click **Internal topics**, which adds these patterns: `.*[\-\.]internal`, `.*\.replica`, and `__.*`. ## Example: Include and exclude topics[​](#example-include-and-exclude-topics "Direct link to Example: Include and exclude topics") To replicate the `warehouse.operations` topic and any topic that starts with `customer.`, but exclude the `customer.support` topic, use these values: * **Topics**: `customer\..*` and `warehouse\.operations` * **Topic blacklist**: `customer\.support` --- # Get started with Apache Kafka® MirrorMaker 2 Create an Apache Kafka® MirrorMaker 2 service and connect it to your Aiven for Apache Kafka® service. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to the [Aiven Console](https://console.aiven.io). * An Aiven project with at least one running Aiven for Apache Kafka® service. To create one, see [Create an Aiven for Apache Kafka service](/docs/products/kafka/get-started/create-kafka-service.md). ## Create an Apache Kafka® MirrorMaker 2 service[​](#create-an-apache-kafka-mirrormaker-2-service "Direct link to Create an Apache Kafka® MirrorMaker 2 service") You can create the service from the Kafka service's integrations page. 1. In the [Aiven Console](https://console.aiven.io), open your project. 2. Open the Aiven for Apache Kafka® service to replicate. 3. Click **Manage stream** > **Integrations**. 4. Under **Aiven services**, select **Apache Kafka MirrorMaker**. 5. Choose one of the following: * **Existing service** to connect to a MirrorMaker 2 service that is already running * **New service** to create a dedicated MirrorMaker 2 service 6. If you select **New service**, click **Create service** to open the service creation page in a new browser tab. ### Create new MirrorMaker 2 service[​](#create-new-mirrormaker-2-service "Direct link to Create new MirrorMaker 2 service") 1. On the **Create service** page, select a **Cloud** and region. 2. Select a **Plan**. 3. In **Service basics**, enter a **service name**. important You cannot change the service name after creation. 4. Review the **Service summary** for the region, plan, and estimated price. 5. Click **Create service**. The service status changes to **Building**. When it changes to **Running**, the service is ready. ### Connect the Kafka service to MirrorMaker 2[​](#connect-the-kafka-service-to-mirrormaker-2 "Direct link to Connect the Kafka service to MirrorMaker 2") After creating the MirrorMaker 2 service: 1. Return to the browser tab with the integration screen. 2. Select **Existing service**. 3. Choose the MirrorMaker 2 service you created. 4. Click **Continue**. 5. Enter a **Cluster alias**. The alias identifies the Kafka cluster within MirrorMaker 2. You cannot modify it after the integration is created. 6. Click **Enable**. You are redirected to the MirrorMaker 2 service overview page. Use the **Replication flows** view to define source and target clusters. Related pages * Browse the [Aiven examples repository](https://github.com/aiven/aiven-examples) for sample code. * Generate sample events with the [Python fake data producer](https://github.com/aiven/python-fake-data-producer-for-apache-kafka). --- # Configure Apache Kafka® MirrorMaker 2 metrics sent to Datadog When creating a [Datadog service integration](https://docs.datadoghq.com/integrations/kafka/?tab=host#kafka-consumer-integration), customize which metrics are sent to the Datadog endpoint using the [Aiven CLI](/docs/tools/cli.md). ## Supported metrics[​](#supported-metrics "Direct link to Supported metrics") The following metric is supported for each replication flow in Apache Kafka® MirrorMaker 2: * `kafka_mirrormaker_summary.replication_lag` The metric is tagged with `replication-flow`, enabling independent monitoring of each one of them. note The `kafka_mirrormaker_summary.replication_lag` metric is available as a custom metric in our Datadog integration. As a custom metric, it may be subject to separate billing by Datadog. ## Variables[​](#variables "Direct link to Variables") Replace the following placeholders in the code samples: | Variable | Description | | ---------------- | --------------------------------------------------------------------------------------- | | `SERVICE_NAME` | Aiven for Apache Kafka® MirrorMaker 2 service name | | `INTEGRATION_ID` | ID of the integration between Aiven for Apache Kafka® MirrorMaker 2 service and Datadog | To find the `INTEGRATION_ID` parameter, run: ``` avn service integration-list SERVICE_NAME ``` ## Customize metrics for Datadog[​](#customize-metrics-for-datadog "Direct link to Customize metrics for Datadog") Before customizing metrics, ensure a Datadog endpoint is configured and enabled in your Aiven for Apache Kafka service. For setup instructions, see [Send metrics to Datadog](/docs/integrations/datadog/datadog-metrics.md). Format any listed parameters as a comma-separated list: `['value0', 'value1', 'value2', ...]`. To customize the metrics sent to Datadog, use the `service integration-update` command with the following customized parameter: * `mirrormaker_custom_metrics`: Define the comma-separated list of custom metrics to include (currently, only `kafka_mirrormaker_summary.replication_lag` is supported). For example, to send the `kafka_mirrormaker_summary.replication_lag` metric, execute the following command: ``` avn service integration-update \ -c 'mirrormaker_custom_metrics=["kafka_mirrormaker_summary.replication_lag"]' \ INTEGRATION_ID ``` After updating settings, view the collected metrics in your Datadog explorer. Related pages * [Datadog and Aiven](/docs/integrations/datadog.md) --- # Exactly-once delivery in Aiven for Apache Kafka MirrorMaker 2 Exactly-once delivery in Aiven for Apache Kafka MirrorMaker 2 replicates each message exactly once between clusters, preventing duplicates or data loss. ## Exactly-once delivery semantics[​](#exactly-once-delivery-semantics "Direct link to Exactly-once delivery semantics") Exactly-once delivery semantics provide a transactional guarantee for message replication, ensuring that all messages in a batch are either fully committed to the target cluster or not replicated at all. This maintains data consistency across clusters. note Exactly-once delivery does not require ACL modifications by default. However, in external Apache Kafka setups where ACLs are applied to the `TransactionalId` resource, review and adjust these ACLs as needed to enable exactly-once delivery. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven for Apache Kafka service with [Aiven for Apache Kafka MirrorMaker 2 integration enabled](/docs/products/kafka/kafka-mirrormaker/get-started.md) * Access to [create or modify replication flows](/docs/products/kafka/kafka-mirrormaker/howto/setup-replication-flow.md) * Access to one of the following, depending on your preferred method: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven Terraform Provider](https://registry.terraform.io/providers/aiven/aiven/latest/docs) * [Aiven API](https://api.aiven.io/) ## Enable or disable exactly-once delivery[​](#enable-or-disable-exactly-once-delivery "Direct link to Enable or disable exactly-once delivery") * Aiven Console * Aiven CLI * Aiven API * Terraform 1. In the [Aiven Console](https://console.aiven.io/), open your Aiven for Apache Kafka service with the Aiven for Apache Kafka MirrorMaker 2 integration. 2. Click **Manage stream** > **Integrations**. 3. Click the Aiven for Apache Kafka MirrorMaker 2 integration. 4. Click **Replication flows** in the sidebar. * To create a new flow, click **Create replication flow** . * To modify an existing flow, select **Edit**. 5. Set **Exactly-once message delivery enabled** to **Enabled** to turn it on or **Disabled** to turn it off. 6. Click **Create** or **Save**. * Create a new replication flow with exactly-once delivery enabled: ``` avn mirrormaker replication-flow create \ --source-cluster \ --target-cluster \ --replication_flow_config '{ "exactly_once_delivery_enabled": true }' ``` * Enable exactly-once delivery on an existing replication flow: ``` avn mirrormaker replication-flow update \ --source-cluster \ --target-cluster \ --replication_flow_config '{ "exactly_once_delivery_enabled": true }' ``` * Disable exactly-once delivery on an existing replication flow: ``` avn mirrormaker replication-flow update \ --source-cluster \ --target-cluster \ --replication_flow_config '{ "exactly_once_delivery_enabled": false }' ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka MirrorMaker 2 service. * `SOURCE_CLUSTER`: Alias of the source Aiven for Apache Kafka cluster. * `TARGET_CLUSTER`: Alias of the target Aiven for Apache Kafka cluster. * `EXACTLY_ONCE_DELIVERY_ENABLED`: Set to `true` to enable or `false` to disable exactly-once delivery. Use the [ServiceKafkaMirrorMakerCreateReplicationFlow](https://api.aiven.io/doc/#tag/Service:_Kafka_MirrorMaker/operation/ServiceKafkaMirrorMakerCreateReplicationFlow) API to create or update a replication flow with `exactly_once_delivery_enabled` in the configuration. * Create a new replication flow with exactly-once delivery enabled: ``` curl -X POST "https://console.aiven.io/v1/project//service//mirrormaker/replication-flows" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "source_cluster": "", "target_cluster": "", "enabled": true, "exactly_once_delivery_enabled": true }' ``` * Update an existing replication flow to enable exactly-once delivery: ``` curl -X PUT "https://console.aiven.io/v1/project//service//mirrormaker/replication-flows//" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "exactly_once_delivery_enabled": true }' ``` * Update an existing replication flow to disable exactly-once delivery: ``` curl -X PUT "https://console.aiven.io/v1/project//service//mirrormaker/replication-flows//" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "exactly_once_delivery_enabled": false }' ``` Parameters: * `PROJECT_NAME`: Name of your project. * `SERVICE_NAME`: Name of your Aiven for Apache Kafka MirrorMaker 2 service. * `API_TOKEN`: Your personal Aiven API [token](/docs/platform/howto/create_authentication_token.md). * `SOURCE_CLUSTER`: Alias of the source Aiven for Apache Kafka cluster. * `TARGET_CLUSTER`: Alias of the target Aiven for Apache Kafka cluster. * `EXACTLY_ONCE_DELIVERY_ENABLED`: Set to `true` to enable or `false` to disable exactly-once delivery. In your `aiven_mirrormaker_replication_flow` resource, set [the `exactly_once_delivery_enabled` attribute](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/mirrormaker_replication_flow#exactly_once_delivery_enabled-1) to true. --- # Offset sync status analysis for Apache Kafka® MirrorMaker 2 Analyze offset synchronization status between source and target clusters with Aiven’s offset sync inspection tool for Aiven for Apache Kafka MirrorMaker 2. This tool monitors sync progress, provides insights into partition health, and flags any needed interventions. Run it locally after configuring the environment to review offset syncing, especially useful during large data migrations. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before you download and run the tool, ensure you have the following: * Basic understanding of Apache Kafka shell tools * A valid Aiven for Apache Kafka MirrorMaker 2 configuration file with the necessary source and target cluster connectivity fields * Java 17 runtime environment (OpenJDK 17) * [Aiven CLI](/docs/tools/cli.md) access * [Aiven personal token](/docs/platform/howto/create_authentication_token.md) * [Aiven for Apache Kafka SSL](/docs/products/kafka/howto/keystore-truststore.md) integration certificates, including: * Client truststore (JKS) * Client keystore (P12) * SSL keystore and truststore password ## Download the offset sync inspection tool[​](#download-the-offset-sync-inspection-tool "Direct link to Download the offset sync inspection tool") [Download the offset sync inspection tool](https://github.com/aiven/kafka/releases/tag/mm2-offset-sync-inspector-0.1) to analyze offset syncing in Aiven for Apache Kafka MirrorMaker 2. ## Run the tool[​](#run-the-tool "Direct link to Run the tool") The offset sync inspection tool outputs data in CSV format. Use this output to check the sync status and detect any issues with the offset syncing process. note The offset sync inspection tool reports only on topics that Aiven for Apache Kafka MirrorMaker 2 actively synchronizes. It does not capture information for all topics in a Kafka cluster. Each replication flow has its own offset sync topic, which contains only offset synchronization data for that flow. If your setup includes multiple MirrorMaker 2 replication flows, run the tool separately for each flow to analyze its synchronization status. The tool does not aggregate data across flows, you must manually combine the results if you need a complete view. Run the tool using the following command: ``` /bin/kafka-mirrormaker-offset-sync-inspector.sh --mm2-config ``` Parameter: * `--mm2-config `: Specifies the path to the Aiven for Apache Kafka MirrorMaker 2 configuration file. This file contains the connection details for the source and target Apache Kafka clusters. To view all available options, run the help command: ``` ./bin/kafka-mirrormaker-offset-sync-inspector.sh --help ``` ## Log messages and actions[​](#log-messages-and-actions "Direct link to Log messages and actions") Below are common log messages generated by the tool, what they mean, and recommended actions. If any issues arise, [create a support ticket](/docs/platform/howto/support.md) or email [Aiven Support](mailto:support@aiven.io). | Sync status message | Description | Action required | Inspect | Analyze | | -------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Successfully synced, operating normally. | Data and offsets are successfully migrated, allowing consumers to read from synced offsets with minimal redelivery. | No action required. | None. | None. | | Intentionally not synced because the source partition is empty. Other partitions in this group are synced. | The source partition has no data, so Apache Kafka MirrorMaker 2 skips syncing it. Consumers start from the beginning when new data arrives. | No action required. | None. | None. | | Intentionally not synced because the source partition is empty. Other partitions in this group are not synced. | The source partition has no data, so Apache Kafka MirrorMaker 2 skips syncing it. Consumers start from the beginning when new data arrives. | No action required. | None. | None. | | Intentionally not synced because the source group is too old. Other partitions in this group are synced. | The source partition has newer data than the committed offset for the consumer group. Consumers start from the beginning when new data arrives. | No action required. | None. | None. | | Erroneously not synced. Source partition has data, and other partitions for this group are synced. | There is an offset syncing issue. Consumers start from the beginning when new data arrives. | [Create a support ticket](/docs/platform/howto/support.md). | - Use `kafka-get-offsets` to compare the earliest and latest offsets in source and target topics.
- Check the source partition's throughput on the Apache Kafka dashboard.
- Inspect offset syncs in `kafka-console-consumer` to locate the topic-partition key.
- Enable TRACE logging for OffsetSyncStore in Apache Kafka MirrorMaker 2. | - The partition has little or no data, causing dropped syncs.
- Zero throughput prevents syncs from being generated.
- Verify if MirrorSourceTask isn’t emitting syncs or if MirrorCheckpointTask is expiring them. | | Erroneously not synced. Source partition has data, and other partitions for this group are not synced. | There is a systematic failure affecting all partitions in this topic. Consumers start from the beginning when new data arrives. | [Create a support ticket](/docs/platform/howto/support.md). | - Run per-partition diagnostics and compare offsets.
- Use `kafka-get-offsets` and inspect sync activity logs. | Investigate sync failures across partitions. | --- # Configure rack awareness in Aiven for Apache Kafka® MirrorMaker 2 Configure rack awareness in Aiven for Apache Kafka® MirrorMaker 2 to reduce cross-availability zone (AZ) network traffic by directing MirrorMaker to read from local follower replicas instead of remote partition leaders. ## About rack awareness in MirrorMaker 2[​](#about-rack-awareness-in-mirrormaker-2 "Direct link to About rack awareness in MirrorMaker 2") Rack awareness in MirrorMaker 2 requires [follower fetching](/docs/products/kafka/howto/enable-follower-fetching.md) to be enabled on the source Aiven for Apache Kafka® service and on the replication flow. When follower fetching is enabled for a replication flow, MirrorMaker automatically sets a rack value based on the availability zone where the MirrorMaker node runs and prefers reading from in-sync follower replicas in the same availability zone. This can reduce cross-availability zone network traffic and associated costs. If the source Kafka cluster does not support follower fetching or uses different rack identifiers, Kafka ignores the rack value and MirrorMaker reads from partition leaders. Follower fetching is enabled by default for new replication flows. If follower fetching is disabled for a replication flow, MirrorMaker reads only from partition leaders for that flow and rack awareness has no effect. For details on follower fetching, see [Follower fetching in Aiven for Apache Kafka®](/docs/products/kafka/concepts/follower-fetching.md). ## When to use rack awareness[​](#when-to-use-rack-awareness "Direct link to When to use rack awareness") Use rack awareness when: * MirrorMaker 2 and the source Kafka service run across multiple availability zones. * Cross-AZ network costs or latency are a concern. * The source Kafka service is hosted on Aiven. Rack awareness does not provide benefits when the Kafka cluster runs in a single availability zone. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running Aiven for Apache Kafka® MirrorMaker 2 service. * Follower fetching enabled on the source Aiven for Apache Kafka® service. ## Enable or disable rack awareness for a replication flow[​](#enable-or-disable-rack-awareness-for-a-replication-flow "Direct link to Enable or disable rack awareness for a replication flow") Rack awareness is controlled by [follower fetching](/docs/products/kafka/howto/enable-follower-fetching.md) and can be enabled or disabled per replication flow. When enabled, MirrorMaker assigns a rack value based on the availability zone where the node runs and prefers reading from in-sync follower replicas in the same availability zone. When disabled, MirrorMaker reads from partition leaders for that replication flow. 1. In the [Aiven Console](https://console.aiven.io), open the MirrorMaker 2 service. 2. Click **Replication flows**. 3. Create a replication flow or edit an existing one. 4. Set **Follower fetching enabled** to on or off. 5. Click **Create** or **Save**. Related pages * [Follower fetching in Aiven for Apache Kafka®](/docs/products/kafka/concepts/follower-fetching.md) * [Enable follower fetching in Aiven for Apache Kafka®](/docs/products/kafka/howto/enable-follower-fetching.md) * [Aiven for Apache Kafka® MirrorMaker 2](/docs/products/kafka/kafka-mirrormaker.md) --- # Monitor replication execution Apache Kafka® MirrorMaker 2 uses Kafka® Connect for monitoring and state management, helping you track replication flows and address issues. ## Monitoring tips[​](#monitoring-tips "Direct link to Monitoring tips") Follow these tips to ensure that replication is up-to-date with message processing: 1. **Monitor consumer lag:** Use the `kafka.consumer_lag` metric to track replication progress and identify delays. 2. **Track dashboard metrics:** Check the `jmx.kafka.connect.mirror.record_count` metric. If MirrorMaker 2 stops adding records to a topic, this metric will show a flat line, indicating no new records are being replicated. 3. **Retrieve the latest messages using `kt`:** Use the [**kt**](https://github.com/fgeller/kt) tool to fetch the latest messages from all partitions. Run the following command: ``` kt consume -auth ./mykafka.conf \ -brokers SERVICE-PROJECT.aivencloud.com:PORT \ -topic topicname -offsets all=newest:newest | \ jq -c -s 'sort_by(.partition) | .[] | \ {partition: .partition, value: .value, timestamp: .timestamp}' ``` --- # Remove topic prefix when replicating with Apache Kafka® MirrorMaker 2 When you use Apache Kafka® MirrorMaker 2 to replicate topics across Apache Kafka® clusters, the default target topic name uses this format: `.`. For example, if the source Apache Kafka cluster alias is `src-kafka`, replicating the `orders` source topic creates the `src-kafka.orders` target topic. For most use cases, the extra prefix is not an issue. If you use a backup Apache Kafka cluster for disaster recovery, you might need consumers and producers to switch clusters with minimal downtime and without changing topic names in their configuration. * Console * API 1. Log in to the [Aiven Console](https://console.aiven.io/) and click the Aiven for Apache Kafka MirrorMaker 2 service. 2. Click **Replication flows**. 3. Open the replication flow to update. 4. In **Replication policy class**, replace the default `org.apache.kafka.connect.mirror.DefaultReplicationPolicy` value with `org.apache.kafka.connect.mirror.IdentityReplicationPolicy`. Replace the ``, ``, and `` placeholders: ``` avn MirrorMaker replication-flow update \ --source-cluster \ --target-cluster \ "{\"replication_policy_class\": \"org.apache.kafka.connect.mirror.IdentityReplicationPolicy\"}" ``` To revert the policy and include the source cluster alias as the topic prefix, run the command with the `org.apache.kafka.connect.mirror.DefaultReplicationPolicy` value. warning The `org.apache.kafka.connect.mirror.IdentityReplicationPolicy` replication policy **does not support active-active replication** because the topics keep the same name and offsets cannot be accurately tracked. Do not create identical replication flows between a source and a destination with the `org.apache.kafka.connect.mirror.IdentityReplicationPolicy` policy. This configuration creates an infinite loop. For active-active, use the `org.apache.kafka.connect.mirror.DefaultReplicationPolicy`. --- # Set up an Apache Kafka® MirrorMaker 2 replication flow Apache Kafka® MirrorMaker 2 replication flows sync topics from a source Apache Kafka® cluster to a target Apache Kafka® cluster. You can define replication flows for Aiven for Apache Kafka services or external [Apache Kafka clusters](/docs/products/kafka/howto/integrate-external-kafka-cluster.md). To define a replication flow between a source Apache Kafka cluster and a target cluster: 1. Log in to the [Aiven Console](https://console.aiven.io/) and click the **Aiven for Apache Kafka MirrorMaker 2** service to define the replication flow. note If needed, [create an Aiven for Apache Kafka MirrorMaker 2 service](/docs/products/kafka/kafka-mirrormaker/get-started.md). 2. Click **Integrations**. 3. Add the source Apache Kafka cluster: * On the **Integrations** page, click **Cluster for replication**. * Click **Existing service** or **New service**. * Click the Apache Kafka service to use as the source cluster. * Click **Continue**. * Add or confirm the cluster alias. 4. Repeat the integration steps for the target Apache Kafka cluster. The cluster alias identifies the Kafka integration in MirrorMaker 2. Use the same aliases when you click **Source Cluster** and **Target Cluster** in the replication flow. If you do not set an alias, Aiven generates one from the project and service name. 5. Click **Replication flows**. 6. Click **Create replication flow**. 7. On the **Create new replication flow** screen, configure the flow: * In **Source Cluster**, click the source Kafka integration alias. * In **Target Cluster**, click the target Kafka integration alias. * In **Topics**, enter the [topics to replicate](/docs/products/kafka/kafka-mirrormaker/concepts/replication-flow-topics-regex.md). To replicate all topics, click **All topics**. * In **Topic blacklist**, enter the topics to exclude. To exclude internal topics, click **Internal topics**. * Configure offset sync, delivery, and replication policy options as needed. For new replications, enable message delivery exactly once if it is available for your setup. * Use **Enable** to create the flow as active or inactive. 8. Click **Create**. 9. In the target Apache Kafka cluster, verify the replicated topics. With the default replication policy, replicated topic names use the `source-cluster-alias.source-topic-name` format. If you use the identity replication policy, topic names stay unchanged. --- # Update integration configurations for Aiven for Apache Kafka® MirrorMaker 2 Update the producer and consumer settings on a MirrorMaker 2 service integration to control how MirrorMaker 2 communicates with the source and target Kafka clusters. Integration configurations are set on the service integration resource. For more information about available parameters and restart impact, see [Configuration and tuning for Aiven for Apache Kafka® MirrorMaker 2](/docs/products/kafka/kafka-mirrormaker/concepts/configuration-layers.md). note Integration configurations are not available in the Aiven Console. Use the Aiven CLI, Aiven API, or Aiven Provider for Terraform to update them. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven for Apache Kafka® MirrorMaker 2 service. * An existing integration between the MirrorMaker 2 service and a source or target Kafka cluster. * Access to one of the following: * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](https://api.aiven.io/) * [Aiven Provider for Terraform](https://registry.terraform.io/providers/aiven/aiven/latest/docs) ## Update integration configurations[​](#update-integration-configurations "Direct link to Update integration configurations") Producer and consumer settings are stored on each `kafka_mirrormaker` service integration. This applies whether the integration connects MirrorMaker 2 to an Aiven for Apache Kafka® service or to an external Kafka cluster through an integration endpoint. * Aiven CLI * Terraform * Aiven API Use [`avn service integration-update`](/docs/tools/cli/service/integration.md#avn_service_integration_update) to update producer and consumer settings on a MirrorMaker 2 integration. 1. List the service integrations for the MirrorMaker 2 service: ``` avn service integration-list MIRRORMAKER_SERVICE_NAME \ --project PROJECT_NAME ``` 2. Copy the service integration ID for the `kafka_mirrormaker` integration that connects MirrorMaker 2 to the source or target cluster you want to configure. 3. Update the integration configuration. To update settings, use `-c KEY=VALUE` for individual parameters or `--user-config-json` for a JSON payload. Pass multiple `-c` options to update several settings at once. Do not use `-c` and `--user-config-json` in the same command. ``` avn service integration-update SERVICE_INTEGRATION_ID \ --project PROJECT_NAME \ -c kafka_mirrormaker.consumer_fetch_min_bytes=1024 \ -c kafka_mirrormaker.producer_linger_ms=100 ``` Parameters: * `PROJECT_NAME`: Your Aiven project name * `MIRRORMAKER_SERVICE_NAME`: Your Aiven for Apache Kafka MirrorMaker 2 service name * `SERVICE_INTEGRATION_ID`: The `kafka_mirrormaker` service integration ID 1. Update the `kafka_mirrormaker_user_config` block in your [`aiven_service_integration`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration) resource: ``` resource "aiven_service_integration" "mirrormaker" { project = "PROJECT_NAME" integration_type = "kafka_mirrormaker" source_service_name = "KAFKA_SERVICE_NAME" destination_service_name = "MIRRORMAKER_SERVICE_NAME" kafka_mirrormaker_user_config { kafka_mirrormaker { consumer_fetch_min_bytes = 1024 producer_linger_ms = 100 } } } ``` Parameters: * `PROJECT_NAME`: Your Aiven project name * `KAFKA_SERVICE_NAME`: The Aiven for Apache Kafka® service connected to MirrorMaker 2. For an external Kafka cluster, use `source_endpoint_id` instead of `source_service_name`. * `MIRRORMAKER_SERVICE_NAME`: Your Aiven for Apache Kafka MirrorMaker 2 service name 2. Run `terraform plan` and `terraform apply` to apply the changes. 1) List the service integrations for the MirrorMaker 2 service using the [ServiceIntegrationList](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationList) endpoint and copy the `service_integration_id` for the `kafka_mirrormaker` integration. 2) Call the [ServiceIntegrationUpdate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationUpdate) endpoint to update the integration configuration: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration/INTEGRATION_ID \ --header "Authorization: Bearer API_TOKEN" \ --header "Content-Type: application/json" \ --data-raw '{ "user_config": { "kafka_mirrormaker": { "consumer_fetch_min_bytes": 1024, "producer_linger_ms": 100 } } }' ``` Parameters: * `PROJECT_NAME`: Your Aiven project name * `INTEGRATION_ID`: The service integration ID * `API_TOKEN`: Your [personal token](/docs/platform/howto/create_authentication_token.md) For the full request schema and supported parameters, see the [Aiven API reference](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationUpdate) and the [`kafka_mirrormaker_user_config` Terraform schema](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration#nested-schema-for-kafka_mirrormaker_user_configkafka_mirrormaker). ## Verify the configuration[​](#verify-the-configuration "Direct link to Verify the configuration") Check that the updated values are applied. * Aiven CLI * Terraform * Aiven API 1. List each `kafka_mirrormaker` integration and its `user_config`: ``` avn service get MIRRORMAKER_SERVICE_NAME \ --project PROJECT_NAME --json | \ jq '[.service_integrations[] | select(.integration_type == "kafka_mirrormaker") | {service_integration_id, user_config}]' ``` Example output: ``` [ { "service_integration_id": "1c2c30f8-413b-4c7c-b393-97165d875952", "user_config": { "cluster_alias": "my-kafka-endpoint-sasl" } }, { "service_integration_id": "57042a2e-aae3-4be3-bcfc-e2c1294b1af3", "user_config": { "cluster_alias": "my-kafka" } }, { "service_integration_id": "a89ca005-e9f1-46a9-9ffb-9f85c517323a", "user_config": { "cluster_alias": "my-kafka-endpoint-ssl", "kafka_mirrormaker": { "consumer_fetch_min_bytes": 1024, "producer_linger_ms": 100 } } } ] ``` 2. Find the `service_integration_id` you updated and confirm that the `kafka_mirrormaker` block under `user_config` contains your updated values. Run `terraform plan` and confirm that Terraform reports no changes: ``` terraform plan ``` Use the [ServiceIntegrationGet](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationGet) endpoint to retrieve the service integration and confirm that the updated values are applied. Related pages * [Configuration and tuning for Aiven for Apache Kafka® MirrorMaker 2](/docs/products/kafka/kafka-mirrormaker/concepts/configuration-layers.md) * [Producer and consumer settings](/docs/products/kafka/kafka-mirrormaker/concepts/configuration-layers.md#producer-and-consumer-settings) * [Aiven CLI service integration commands](/docs/tools/cli/service/integration.md) * [`aiven_service_integration` Terraform resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration) --- # Advanced parameters for Aiven for Apache Kafka® MirrorMaker 2 See the configuration options available for Aiven for Apache Kafka® MirrorMaker 2: | Parameter | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | []()[**ip\_filter**](#ip_filter)`array`- default: `0.0.0.0/0,::/0`IP filterAllow incoming connections from CIDR address block, e.g. '10.20.0.0/16' | | []()[**preferred\_zones**](#preferred_zones)`array,null`List of preferred zone IDs for service node placement. Nodes will be placed in these zones when available. If a specified zone is unavailable (e.g., due to capacity constraints), nodes will be placed in other available zones to maintain the configured number of zones for availability. Invalid zone IDs are rejected at configuration time. Zone IDs are cloud-specific: AWS uses zone IDs like 'euc1-az1', GCP uses zone names like 'europe-west1-a', and Azure uses 'location/zone' format like 'germanywestcentral/1'. If single\_zone is enabled with an availability\_zone, that setting takes precedence over preferred\_zones. Changes take effect on next node recreation (e.g., maintenance or plan change). For eligible plans, nodes outside preferred zones are automatically rebalanced once per day. | | []()[**service\_log**](#service_log)`boolean,null`Service loggingStore logs for the service so that they are available in the HTTP API and console. | | []()[**static\_ips**](#static_ips)`boolean`Use static public IP addresses | | []()[**kafka\_mirrormaker**](#kafka_mirrormaker)`object`Kafka MirrorMaker configuration valueskafka\_mirrormaker.admin\_timeout\_ms integer,null - min: 30000 - max: 1800000 Timeout for administrative tasks, e.g. detecting new topics, loading of consumer group and offsets. Defaults to 60000 milliseconds (1 minute). kafka\_mirrormaker.refresh\_topics\_enabled boolean Refresh topics and partitions Whether to periodically check for new topics and partitions. Defaults to 'true'. kafka\_mirrormaker.refresh\_topics\_interval\_seconds integer - min: 1 - max: 9223372036854776000 Frequency of topic and partitions refresh in seconds. Defaults to 600 seconds (10 minutes). kafka\_mirrormaker.refresh\_groups\_enabled boolean Refresh consumer groups Whether to periodically check for new consumer groups. Defaults to 'true'. kafka\_mirrormaker.refresh\_groups\_interval\_seconds integer - min: 1 - max: 9223372036854776000 Frequency of consumer group refresh in seconds. Defaults to 600 seconds (10 minutes). kafka\_mirrormaker.sync\_group\_offsets\_enabled boolean Whether to periodically write the translated offsets of replicated consumer groups (in the source cluster) to \_\_consumer\_offsets topic in target cluster, as long as no active consumers in that group are connected to the target cluster kafka\_mirrormaker.sync\_group\_offsets\_interval\_seconds integer - min: 1 - max: 9223372036854776000 Frequency of consumer group offset sync Frequency at which consumer group offsets are synced (default: 60, every minute) kafka\_mirrormaker.emit\_checkpoints\_enabled boolean Whether to emit consumer group offset checkpoints to target cluster periodically (default: true) kafka\_mirrormaker.emit\_checkpoints\_interval\_seconds integer - min: 1 - max: 9223372036854776000 Frequency at which consumer group offset checkpoints are emitted (default: 60, every minute) kafka\_mirrormaker.sync\_topic\_configs\_enabled boolean Whether to periodically configure remote topics to match their corresponding upstream topics. kafka\_mirrormaker.tasks\_max\_per\_cpu integer - min: 1 - max: 8 - default: 1 Maximum number of MirrorMaker tasks (of each type) per service CPU 'tasks.max' is set to this multiplied by the number of CPUs in the service. kafka\_mirrormaker.offset\_lag\_max integer - max: 9223372036854776000 Maximum offset lag before it is resynced How out-of-sync a remote partition can be before it is resynced. kafka\_mirrormaker.groups string Comma-separated list of consumer groups to replicate Consumer groups to replicate. Supports comma-separated group IDs and regexes. kafka\_mirrormaker.groups\_exclude string Comma-separated list of group IDs and regexes to exclude from replication Exclude groups. Supports comma-separated group IDs and regexes. Excludes take precedence over includes. | | []()[**sasl\_oauthbearer\_allowed\_urls**](#sasl_oauthbearer_allowed_urls)`array`List of allowed URLs for SASL OAUTHBEARER authentication. Only HTTPS URLs are allowed for security reasons. | --- # Known issues ## MirrorMaker 2 may translate offsets incorrectly if LAG shows negative on the target cluster[​](#mirrormaker-2-may-translate-offsets-incorrectly-if-lag-shows-negative-on-the-target-cluster "Direct link to MirrorMaker 2 may translate offsets incorrectly if LAG shows negative on the target cluster") The following are known issues that are expected to be resolved in version 3.3.0. See [recommended tuning parameters](/docs/products/kafka/kafka-mirrormaker/concepts/configuration-layers.md) to mitigate the impact of these issues. * **MirrorMaker 2 offset sync is incorrect if the target partition is empty**: * `Kafka-12635`: **Don't emit checkpoints for partitions without any offset**: * **MM2 creates invalid checkpoint when offset mapping is not available**: * `Kafka-13452`: **MM2 shouldn't checkpoint when offset mapping is unavailable**: ## MirrorMaker 2 is always set `min.insync.replicas = 1` in destination cluster topics[​](#mirrormaker-2-is-always-set-mininsyncreplicas--1-in-destination-cluster-topics "Direct link to mirrormaker-2-is-always-set-mininsyncreplicas--1-in-destination-cluster-topics") If a topic on the destination cluster is updated or pre-created with `min.insync.replicas > 1`, MirrorMaker 2 would always be overwritten with `min.insync.replicas = 1` in target topics. The replication factor of target topics does not match the replication factor of source topics either when they are newly created by MirrorMaker 2, or when they exist beforehand. * MirrorMaker 2 does not update `replicas`. * MirrorMaker 2 always overwrites `min.insync.replicas` as 1. --- # Terminology for Aiven for Apache Kafka® MirrorMaker 2 * **Cluster alias**: The name alias defined in MirrorMaker 2 for a certain Apache Kafka® source or target cluster. * **Replication flow**: The flow of data between two Apache Kafka® clusters (called source and target) executed by Apache Kafka® MirrorMaker 2. One Apache Kafka® MirrorMaker 2 service can execute multiple replication flows. * **Remote topics**: Topics replicated by MirrorMaker 2 from a source Apache Kafka® cluster to a target Apache Kafka® cluster. There is only one source topic for each remote topic. Remote topics refer to the source cluster by the topic name prefix: `{source_cluster_alias}.{source_topic_name}`. --- # Why topics or partitions not replicated Apache Kafka® MirrorMaker 2 provides reliable message replication across Kafka clusters using configurations, states, and offsets. Various factors can disrupt the replication of topics and partitions, preventing them from progressing as expected. These guidelines assume that you have previously [set up a MirrorMaker replication flow](/docs/products/kafka/kafka-mirrormaker/howto/setup-replication-flow.md) and have identified an issue with the replication of certain topics or partitions. note You can assess Aiven for Apache Kafka MirrorMaker 2 replication issues in several ways: * Basic monitoring * Log search * [Offset sync status analysis](/docs/products/kafka/kafka-mirrormaker/howto/log-analysis-offset-sync-tool.md) ## Excessive message size[​](#excessive-message-size "Direct link to Excessive message size") The error `RecordTooLargeException` appears in service logs when a record exceeds the allowed size. This can occur in two scenarios: 1. **Target Apache Kafka broker rejects a record** * **Cause**: The record is larger than the destination topic's allowed size. * **Solution**: Increase the `max_message_bytes` value for the topic or the broker’s `message_max_bytes` configuration. Restart the worker tasks to apply the new settings. 2. **MirrorMaker connector rejects the record** * **Cause**: The record is larger than the maximum producer request size. * **Solution**: Update the `producer_max_request_size` in the integration configuration. For more details, see the [integration configuration documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration#producer_max_request_size-1). ## Limited replication[​](#limited-replication "Direct link to Limited replication") The following worker configurations can significantly impact topic partition replication: * [`kafka_mirrormaker.offset_lag_max`](/docs/products/kafka/kafka-mirrormaker/reference/advanced-params.md#kafka_mirrormaker_offset_lag_max) * Default value: 100 * Description: This parameter applies globally to all replication topics and defines how far behind a remote partition can be before catching up. Consider its impact on high-throughput and low-throughput topics carefully. * Considerations: * A low value can increase load on Apache Kafka MirrorMaker 2 workers and Kafka brokers due to high-throughput topics. * A high value can prevent low-throughput topics from progressing. * [`kafka_mirrormaker.tasks_max_per_cpu`](/docs/products/kafka/kafka-mirrormaker/reference/advanced-params.md#kafka_mirrormaker_tasks_max_per_cpu) * Default value: 1 * Description: Specifies the maximum number of Apache Kafka MirrorMaker 2 tasks per service CPU. For example, in a typical cluster with 3 nodes, each having 2 CPUs, the `tasks.max` is automatically set to 6, allowing 6 tasks to execute simultaneously. The optimal performance is achieved when each Kafka consumer is assigned to a single partition. If Apache Kafka MirrorMaker 2 processes more partitions than available tasks, multiple partitions are assigned to a single task. * Considerations: * A low value can introduce replication lag on some partitions. * A high value can result in idle tasks. note Starting with Apache Kafka v3.7, Aiven for Apache Kafka MirrorMaker 2 emits offset sync information periodically, enhancing replication monitoring and troubleshooting. ## Non-coherent offset[​](#non-coherent-offset "Direct link to Non-coherent offset") Apache Kafka MirrorMaker 2 stores offsets in the target cluster’s internal topic `mm2-offsets..internal`. By design, it checks if an offset is already stored for the replication topic and resumes replication from the last stored offset to avoid duplicate consumption. This behavior can cause confusion when: * A replication topic was previously replicated and later removed from the replication flow. * A replication topic was deleted and recreated at the source. To resolve these issues, you can reset the offsets using one of the following methods: ### Resetting all offsets[​](#resetting-all-offsets "Direct link to Resetting all offsets") 1. Disable the replication flow. 2. Delete the internal offsets topic. 3. Re-enable the replication flow to recreate the internal offsets topic automatically. 4. MirrorMaker will replicate from the earliest offset. ### Resetting offsets for a specific topic[​](#resetting-offsets-for-a-specific-topic "Direct link to Resetting offsets for a specific topic") 1. [Configure the Kafka toolbox](/docs/products/kafka/howto/kafka-tools-config-file.md) to access the Kafka cluster hosting the internal offsets topic. 2. Disable the replication flow. 3. Produce a tombstone record to the offset storage for the replication topic using [this script](https://gist.github.com/C0urante/30dba7b9dce567f33df0526d68765860). note The `-o` option is not needed if the default offsets topic applies. 4. Re-enable the replication flow. 5. MirrorMaker replicates from the earliest offset. --- # Karapace [Karapace](https://karapace.io/) is an Aiven-built open-source Schema Registry and REST Proxy for Aiven for Apache Kafka®. Use it to store and manage message schemas, and to produce and consume Kafka messages over HTTP. You can [enable or disable each feature independently](/docs/products/kafka/karapace/howto/enable-karapace.md). ## Schema Registry[​](#schema-registry "Direct link to Schema Registry") Schema Registry stores schemas in a central repository. Producers and consumers use those schemas to serialize and deserialize messages. You can version schemas and check compatibility before you publish changes. Supported formats: | Format | Schema registration | Schema references | | ----------- | ------------------- | ----------------- | | Avro | ✓ | ✓ | | Protobuf | ✓ | ✓ | | JSON Schema | ✓ | - | Schema references let one schema depend on other registered schemas instead of inlining every definition. Karapace supports Avro schema references from version 6.1.0 onward. For details, see [Schema references in Karapace](/docs/products/kafka/karapace/concepts/schema-references.md). ## REST Proxy[​](#rest-proxy "Direct link to REST Proxy") The REST Proxy is a Karapace component that produces and consumes Kafka events over HTTP. It exposes REST APIs for those operations. Use those APIs from tools such as `curl`, scripts, or HTTP-based services. You manage Schema Registry over HTTP with REST APIs that register, update, and delete schemas. Those APIs are separate from the Kafka REST APIs served by the REST Proxy. For more information, see [Apache Kafka REST API](/docs/products/kafka/concepts/kafka-rest-api.md). ## Next steps[​](#next-steps "Direct link to Next steps") 1. [Enable schema registry and REST proxy](/docs/products/kafka/karapace/howto/enable-karapace.md) 2. [Use schema registry with Java producers and consumers](/docs/products/kafka/howto/use-schema-registry-in-java.md) 3. [Connect with Kafka REST](/docs/products/kafka/howto/connect-with-kafka-rest.md) 4. [Enable OAuth 2.0/OIDC authentication for Schema Registry](/docs/products/kafka/karapace/howto/enable-oauth-oidc-schema-registry.md) ## Karapace resources[​](#karapace-resources "Direct link to Karapace resources") * [Karapace project site](https://karapace.io/) * [GitHub repository](https://github.com/Aiven-Open/karapace) Related pages * [Enable schema registry and REST proxy](/docs/products/kafka/karapace/howto/enable-karapace.md) * [Apache Kafka REST API](/docs/products/kafka/concepts/kafka-rest-api.md) * [Schema references in Karapace](/docs/products/kafka/karapace/concepts/schema-references.md) --- # Schema registry ACL definitions A Schema Registry ACL controls who can read or write Schema Registry resources in [Karapace](/docs/products/kafka/karapace.md). These ACLs are separate from [Apache Kafka Access Control Lists](/docs/products/kafka/concepts/acl.md), which control access to topics and other Kafka resources. ## ACL entry fields[​](#acl-entry-fields "Direct link to ACL entry fields") Each ACL entry has three parts: * **Username**: A service user on your Aiven for Apache Kafka® service. * **Operation**: One of: * `schema_registry_read` * `schema_registry_write` (always includes `schema_registry_read`) * **Resource**: One of these formats: * `Config:`: Controls access to global compatibility configuration. Users with this resource can get or set the default schema compatibility mode. Getting the configuration requires `schema_registry_read`. Setting it requires `schema_registry_write`. * `Subject:subject_name`: Controls access to a subject in Schema Registry. tip The username and resource `name` values can use wildcards: * `*` matches any characters * `?` matches a single character ## How access decisions work[​](#how-access-decisions-work "Direct link to How access decisions work") When a user requests a resource, Schema Registry checks whether any ACL entry matches the user and the resource. If a matching entry grants the required operation, access is allowed. Entry order does not affect the decision. If no ACL entry grants access, Schema Registry returns an HTTP `401 Unauthorized` status code. ## Endpoint permissions[​](#endpoint-permissions "Direct link to Endpoint permissions") * Read-only endpoints need `schema_registry_read` for the subject. For endpoints that return data for multiple subjects, the response includes only subjects the user can read. * Write endpoints need `schema_registry_write` for the subject. ### Examples[​](#examples "Direct link to Examples") | Username | Operation | Resource | Effect | | ---------------- | ----------------------- | ------------ | ------------------------------------------------------------------------------------------------------ | | `user_1` | `schema_registry_read` | `Config:` | Read global compatibility configuration. | | `user_1` | `schema_registry_read` | `Subject:s1` | Read data for subject `s1` only. List responses omit other subjects. | | `user_1` | `schema_registry_write` | `Subject:s1` | Add, update, or delete data for subject `s1`. Includes read access. | | `user_readonly*` | `schema_registry_read` | `Subject:s*` | Read access for usernames with prefix `user_readonly` to subjects with prefix `s`. | | `user_write*` | `schema_registry_write` | `Subject:s*` | Write access for usernames with prefix `user_write` to subjects with prefix `s`. Includes read access. | ## Superuser access[​](#superuser-access "Direct link to Superuser access") The user that manages ACLs is a superuser with write access to everything in Schema Registry. In the Aiven Console, that superuser can view and modify all schemas on the **Schemas** tab of a Kafka service. The superuser and its ACL entries are not visible in the Console. Aiven adds them automatically. note Create and manage ACL entries with the [Aiven CLI](/docs/tools/cli/service/schema-registry-acl.md) or see [Manage Karapace schema registry authorization](/docs/products/kafka/karapace/howto/manage-schema-registry-authorization.md). Related pages * [Karapace schema registry authorization](/docs/products/kafka/karapace/concepts/schema-registry-authorization.md) * [Enable Karapace schema registry authorization](/docs/products/kafka/karapace/howto/enable-schema-registry-authorization.md) * [Manage Karapace schema registry authorization](/docs/products/kafka/karapace/howto/manage-schema-registry-authorization.md) * [avn service schema-registry-acl](/docs/tools/cli/service/schema-registry-acl.md) --- # Schema references in Karapace Schema references let you register a schema that depends on other schemas already stored in the Schema Registry. Use them to reuse shared schemas instead of copying definitions into every schema. For example, if a `Country` record is used by `Address`, and `Address` is used by `Person`, register each type once and reference it by subject and version. In the Schema Registry, a subject is the named schema stream (for example, `address`). ## When to use schema references[​](#when-to-use-schema-references "Direct link to When to use schema references") Use schema references to: * Keep one source of truth for a shared record or message shape across subjects * Update a shared type in one place instead of editing many duplicated schemas ## Supported schema formats[​](#supported-schema-formats "Direct link to Supported schema formats") Karapace supports schema references for: * Avro (Karapace 6.1.0 or later) * Protobuf note Schema references are not supported for JSON Schema. If you register a JSON Schema with a `references` array, Karapace returns HTTP `422 Unprocessable Entity`. ## How schema references work[​](#how-schema-references-work "Direct link to How schema references work") When a schema uses types from other subjects, include a `references` array in the registration request. Each entry lists `name`, `subject`, and `version`. Each reference points to a specific schema version. Karapace loads those versions when it validates schemas or checks compatibility. Register every referenced schema before you register a schema that lists it in `references`. ## Schema Registry API structure[​](#schema-registry-api-structure "Direct link to Schema Registry API structure") You register schemas with references over the Schema Registry HTTP API. Include a `references` array in the request body. The following example shows the payload shape: ``` POST /subjects/{subject}/versions { "schemaType": "AVRO", "schema": "", "references": [ { "name": "", "subject": "", "version": } ] } ``` Fields in each `references` object: | Field | Purpose | | --------- | ---------------------------------------------------------------------------------- | | `name` | Label or import path used in the registration payload. Behavior differs by format. | | `subject` | Subject where the referenced schema is registered. | | `version` | Schema version to reference. | For format-specific `name` rules, see [Avro references](#avro-references) and [Protobuf references](#protobuf-references). For complete `curl` examples, see [Register schemas with references](/docs/products/kafka/karapace/howto/register-schemas-with-references.md). ## Avro references[​](#avro-references "Direct link to Avro references") In Avro, `name` is a label only. A file-style value such as `address.avsc` is a convention, not a requirement. Karapace resolves each reference from the Avro type names in your schema, together with `subject` and `version`. For example, an `Address` record that uses a `Country` type needs a `references` entry whose `subject` and `version` point to the `Country` registration. For `curl` examples, see [Register schemas with references](/docs/products/kafka/karapace/howto/register-schemas-with-references.md#example-avro-records-with-references). ## Protobuf references[​](#protobuf-references "Direct link to Protobuf references") Protobuf uses the same `references` array as Avro. Unlike Avro, Karapace uses `name` to resolve the reference. Set `name` to the import path from your `.proto` file, such as `address.proto`. The value must match the `import` statement exactly. For `curl` examples, see [Register schemas with references](/docs/products/kafka/karapace/howto/register-schemas-with-references.md#example-protobuf-messages-with-imports). ## Compatibility checks with references[​](#compatibility-checks-with-references "Direct link to Compatibility checks with references") Compatibility checks work the same for schemas with references as for standalone schemas. Karapace resolves referenced schema versions before it evaluates compatibility. Each reference pins a specific schema version. Update `version` in the matching `references` entry when you adopt a newer version of a referenced schema. Related pages * [Register schemas with references](/docs/products/kafka/karapace/howto/register-schemas-with-references.md) * [Enable schema registry and REST proxy](/docs/products/kafka/karapace/howto/enable-karapace.md) * [Karapace](/docs/products/kafka/karapace.md) --- # Karapace schema registry authorization The schema registry authorization feature when enabled in [Karapace schema registry](/docs/products/kafka/karapace/howto/enable-karapace.md) allows you to authenticate the user, and control read or write access to the individual resources available in the Schema Registry. Authorization in Karapace is achieved by using [Access Control Lists (ACLs)](/docs/products/kafka/karapace/concepts/acl-definition.md). ACLs provide a way to achieve fine-grained access control for the resources in Karapace. note Karapace schema registry authorization is enabled on all Aiven for Apache Kafka® services. The exception is older services created before mid-2022, where the feature needs to be [enabled](/docs/products/kafka/karapace/howto/enable-schema-registry-authorization.md). To authenticate Schema Registry with OAuth 2.0/OIDC bearer tokens and restrict operations by JWT roles, see [Enable OAuth 2.0/OIDC authentication for Aiven for Apache Kafka® Schema Registry](/docs/products/kafka/karapace/howto/enable-oauth-oidc-schema-registry.md). ## Common use cases[​](#common-use-cases "Direct link to Common use cases") Some of the common use cases for Karapace schema registry authorization include: * **Schemas as API contract among teams:** Allow teams to leverage schemas as API contracts with other teams by providing read-only access to schemas in Schema Registry. * **Kafka as a communication broker with third parties:** Allow conditional read-only access to schemas by third parties to establish Kafka topics as communication channels. * **Schema registration automated with CI/CD:** Achieve automation of schema registration using Continuous Integration(CI) tools in such a way that add, update and delete operations on schemas is limited to the CI Tools. At the same time, the different applications are allowed read-only access to the schemas as needed. * **Segregated consumers of Schema Registry and REST Proxy:** Limits the attack surface by segregating and limiting access to only what is required for the consumers of the Schema Registry and REST Proxy. --- # Enable Apache Kafka® REST proxy authorization REST proxy authorization applies [Apache Kafka Access Control Lists (ACLs)](/docs/products/kafka/concepts/acl.md) to requests made through the Karapace REST proxy. When authorization is enabled, Karapace forwards HTTP basic authentication credentials to Apache Kafka®. Apache Kafka authenticates the user and authorizes operations based on the ACLs defined for the service. When authorization is disabled, the REST proxy bypasses Apache Kafka ACLs, so REST API calls are not restricted by those rules. REST proxy authorization is disabled by default. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An [Aiven for Apache Kafka®](/docs/products/kafka/get-started/create-kafka-service.md) service * [REST proxy enabled](/docs/products/kafka/karapace/howto/enable-karapace.md) on the service (`kafka_rest`) ## Enable REST proxy authorization[​](#enable-rest-proxy-authorization "Direct link to Enable REST proxy authorization") warning Enabling REST proxy authorization can disrupt access if Kafka ACLs are not configured to allow the operations your clients need. Configure [Access Control Lists](/docs/products/kafka/concepts/acl.md) before you enable authorization. * Console * CLI 1. In the [Aiven Console](https://console.aiven.io/), select your project and choose your Aiven for Apache Kafka® service. 2. Click **Service settings**. 3. In **Advanced configuration**, click **Configure**. 4. Click **Add configuration options**. 5. Find `kafka_rest_authorization` and set it to **Enabled**. 6. Click **Save configuration**. Enable REST proxy authorization with the [Aiven CLI](/docs/tools/cli.md): ``` avn service update -c kafka_rest_authorization=true SERVICE_NAME ``` To disable it: ``` avn service update -c kafka_rest_authorization=false SERVICE_NAME ``` Replace `SERVICE_NAME` with the name of your Aiven for Apache Kafka® service. Related pages * [Enable schema registry and REST proxy](/docs/products/kafka/karapace/howto/enable-karapace.md) * [Access Control Lists in Aiven for Apache Kafka®](/docs/products/kafka/concepts/acl.md) * [Enable OAuth2/OIDC support for Apache Kafka® REST proxy](/docs/products/kafka/karapace/howto/enable-oauth-oidc-kafka-rest-proxy.md) * [Apache Kafka® REST API](/docs/products/kafka/concepts/kafka-rest-api.md) --- # Enable schema registry and REST proxy You can enable the Karapace schema registry and REST proxy independently on Aiven for Apache Kafka®. ## Enable from Connection information[​](#enable-from-connection-information "Direct link to Enable from Connection information") 1. In the [Aiven Console](https://console.aiven.io/), select your project and choose your Aiven for Apache Kafka® service. 2. On the **Overview** page, open **Connection information**. 3. To enable the REST proxy, open the **Apache Kafka REST** tab. When the feature is off, the tab shows **Enable REST Proxy**. 4. Click **Enable**. This sets `kafka_rest`. 5. To enable Schema Registry, open the **Schema Registry** tab. When the feature is off, the tab shows **Enable schema registry**. 6. Click **Enable**. This sets `schema_registry`. ## Enable from Service settings[​](#enable-from-service-settings "Direct link to Enable from Service settings") 1. In the [Aiven Console](https://console.aiven.io/), select your project and choose your Aiven for Apache Kafka® service. 2. Click **Service settings**. 3. In **Service management**, click **Actions** (**...**). 4. To enable the REST proxy, click **Enable REST API (Karapace)**, then click **Enable**. This sets `kafka_rest`. In Service management, the feature appears as **Apache Kafka REST API (Karapace)**. 5. To enable Schema Registry, click **Enable Schema Registry (Karapace)**, then click **Enable**. This sets `schema_registry`. In Service management, the feature appears as **Schema Registry (Karapace)**. tip For automation, set the `schema_registry` and `kafka_rest` service parameters. Related pages * [Karapace](/docs/products/kafka/karapace.md) * [Karapace on GitHub](https://github.com/Aiven-Open/karapace) * [Aiven Terraform Provider `aiven_kafka` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/kafka) --- # Enable OAuth 2.0/OIDC support for Apache Kafka® REST proxy Secure your Apache Kafka® resources by integrating OAuth 2.0/OpenID Connect (OIDC) with the Karapace REST proxy and enabling REST proxy authorization. This setup ensures that only authorized individuals can manage Apache Kafka resources through both token-based authentication and access control rules. ## OAuth 2.0/OIDC token handling[​](#oauth-20oidc-token-handling "Direct link to OAuth 2.0/OIDC token handling") Karapace processes the JSON Web Token (JWT) obtained from the Authorization HTTP header, specifically when employing the Bearer authentication scheme. This allows OAuth 2.0/OIDC credentials to be supplied directly to the REST proxy, which uses the provided token to authorize requests to Apache Kafka. When a Bearer token is presented, Kafka clients configured by Karapace use the SASL OAUTHBEARER mechanism to send the JWT for validation. ## Schema Registry authentication[​](#schema-registry-authentication "Direct link to Schema Registry authentication") Karapace REST proxy also communicates with Schema Registry when processing schema-based messages. With Karapace 6.2.3 or later, the REST proxy can authenticate to Schema Registry using JWT or basic authentication. The REST proxy supports both authentication methods simultaneously. This authentication is separate from the OAuth 2.0/OIDC authentication that the REST proxy uses to authenticate to Apache Kafka. To configure JWT authentication for Schema Registry, see [Enable OAuth 2.0/OIDC authentication for Schema Registry](/docs/products/kafka/karapace/howto/enable-oauth-oidc-schema-registry.md). ## Authorization enforcement[​](#authorization-enforcement "Direct link to Authorization enforcement") In the underlying Aiven for Apache Kafka® service, the default mechanism for authorization, uses the `sub` claim from the JWT as the username. This username is then verified against the configured Access Control Lists (ACLs) to authorize user operations on Apache Kafka resources. While the `sub` claim is the default identifier, this setting is configurable. You can specify a different JWT claim for authentication by adjusting the `kafka.sasl_oauthbearer_sub_claim_name` parameter. For more information on configuring this, see [Enable OAuth 2.0/OIDC via Aiven Console](/docs/products/kafka/howto/enable-oidc.md). To authenticate and authorize a user in Aiven for Apache Kafka, you need a service user and an ACL entry that describes the permissions. Match the JWT claim value used for authentication to the service user in the system. This service user needs to be associated with an ACL entry that outlines their permissions, ensuring that the identity of the user making the request aligns with both the service user and the ACL entry. ## Managing token expiry[​](#managing-token-expiry "Direct link to Managing token expiry") With OAuth 2.0/OIDC enabled, Karapace manages Apache Kafka client connections for security and performance. It automatically cleans up idle clients and those with tokens nearing expiration, typically on a 5-minute cycle. This cleanup prevents unauthorized access with expired tokens and clears idle connections. note Before your token expires, remove any linked consumers and producers to avoid security issues and service interruptions. After removal, refresh your OAuth 2.0 JWT tokens and reconnect with the new tokens. ## Configure OAuth 2.0/OIDC authentication[​](#configure-oauth-20oidc-authentication "Direct link to Configure OAuth 2.0/OIDC authentication") To establish OAuth 2.0/OIDC authentication for the Karapace REST proxy, complete the following prerequisites and configuration steps: ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Aiven for Apache Kafka®](/docs/products/kafka/get-started/create-kafka-service.md) service running with [OAuth 2.0/OIDC enabled](/docs/products/kafka/howto/enable-oidc.md). * [Karapace schema registry and REST APIs enabled](/docs/products/kafka/karapace/howto/enable-karapace.md). * Ensure access to an OIDC-compliant provider, such as Auth0, Okta, Google Identity Platform, or Azure. ### Configuration steps[​](#configuration-steps "Direct link to Configuration steps") * Aiven Console * Aiven CLI 1. In the [Aiven Console](https://console.aiven.io/), select your project and choose your Aiven for Apache Kafka® service. 2. Click **Service settings**. 3. Go to **Advanced configuration** and click **Configure**. 4. In the **Advanced configuration** window, click **Add configuration options**. 5. Find the `kafka_rest_authorization` parameter and set it to `Enabled`. 6. Click **Save configurations**. To enable REST proxy authorization, use the following command in the [Aiven CLI](/docs/tools/cli.md), replacing `SERVICE_NAME` with your actual service name: ``` avn service update -c kafka_rest_authorization=true SERVICE_NAME ``` Disable REST proxy authorization, use: ``` avn service update -c kafka_rest_authorization=false SERVICE_NAME ``` warning Enabling Apache Kafka REST proxy authorization can disrupt access for users if the Kafka access control rules have not been configured properly. For more information, see [Enable Apache Kafka® REST proxy authorization](/docs/products/kafka/karapace/howto/enable-kafka-rest-proxy-authorization.md). Related pages * [Enable OAUTH2/OIDC authentication for Aiven for Apache Kafka](/docs/products/kafka/howto/enable-oidc.md) * [Enable OAuth 2.0/OIDC authentication for Aiven for Apache Kafka® Schema Registry](/docs/products/kafka/karapace/howto/enable-oauth-oidc-schema-registry.md) --- # Enable OAuth 2.0/OIDC authentication for Aiven for Apache Kafka® Schema Registry Use OAuth 2.0/OpenID Connect (OIDC) to authenticate requests to Karapace Schema Registry with a JSON Web Token (JWT) issued by your identity provider. You can also enable role-based authorization to control which Schema Registry operations clients can perform. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before you begin, make sure you have: * An [Aiven for Apache Kafka®](/docs/products/kafka.md) service with [Schema Registry enabled](/docs/products/kafka/karapace/howto/enable-karapace.md) * Karapace version 6.2.1 or later * Access to an OIDC-compliant identity provider * The following OIDC provider settings configured for your Aiven for Apache Kafka service: * `kafka.sasl_oauthbearer_jwks_endpoint_url` * `kafka.sasl_oauthbearer_expected_issuer` * `kafka.sasl_oauthbearer_expected_audience` Schema Registry uses the same OIDC provider settings as Apache Kafka. The Aiven Console does not require the expected issuer or audience settings when you configure Kafka OIDC, but Schema Registry requires both for authentication. For more information about configuring these settings, see [Enable OAuth 2.0/OIDC authentication for Apache Kafka®](/docs/products/kafka/howto/enable-oidc.md). note If your service runs a Karapace version earlier than 6.2.1, apply the available maintenance update first. For more information, see [Set the Karapace version](/docs/products/kafka/karapace/howto/set-karapace-version.md). ## Enable OIDC authentication[​](#enable-oidc-authentication "Direct link to Enable OIDC authentication") Karapace Schema Registry validates the JWT in the `Authorization` header against the OIDC provider settings configured for the service. This differs from the [Karapace REST proxy](/docs/products/kafka/karapace/howto/enable-oauth-oidc-kafka-rest-proxy.md), where Apache Kafka validates the bearer token. In Karapace 6.2.3 and later, enabling OIDC authentication does not disable basic authentication. Clients can authenticate with a bearer token or with basic authentication. Schemas remain visible in the Aiven Console after you enable OIDC authentication because the Console connects to Schema Registry with basic credentials. You cannot disable Schema Registry basic authentication in the Aiven Console or with the Aiven CLI. * Console * CLI 1. In the [Aiven Console](https://console.aiven.io/), select your project and choose your Aiven for Apache Kafka service. 2. Click **Service settings**. 3. Click **Advanced configuration** > **Configure**. 4. Click **Add configuration options**. 5. Add `schema_registry_config.sasl_oauthbearer_authentication_enabled`. 6. Set the option to **Enabled**. 7. Click **Save configuration**. Run the following command: ``` avn service update SERVICE_NAME \ -c schema_registry_config.sasl_oauthbearer_authentication_enabled=true ``` Replace `SERVICE_NAME` with the name of your Aiven for Apache Kafka service. ## Enable role-based authorization[​](#enable-role-based-authorization "Direct link to Enable role-based authorization") When OIDC authentication is enabled and role-based authorization is disabled, any client with a valid token can access Schema Registry. To restrict access based on roles, enable `schema_registry_config.sasl_oauthbearer_authorization_enabled`. note Enabling role-based authorization also enables OIDC authentication if it is not already enabled. When you enable authorization, add the roles claim path and HTTP method roles options. If you do not change the values, Karapace uses the defaults. * Console * CLI 1. In the Aiven Console, select your project and choose your Aiven for Apache Kafka service. 2. Click **Service settings**. 3. Click **Advanced configuration** > **Configure**. 4. Click **Add configuration options**. 5. Add `schema_registry_config.sasl_oauthbearer_authorization_enabled`. 6. Set the option to **Enabled**. 7. Add `schema_registry_config.sasl_oauthbearer_roles_claim_path`. 8. Add `schema_registry_config.sasl_oauthbearer_method_roles`. 9. Click **Save configuration**. Run the following command: ``` avn service update SERVICE_NAME \ -c schema_registry_config.sasl_oauthbearer_authorization_enabled=true ``` Replace `SERVICE_NAME` with the name of your Aiven for Apache Kafka service. To use a different roles claim path, add `schema_registry_config.sasl_oauthbearer_roles_claim_path`. For example: ``` avn service update SERVICE_NAME \ -c schema_registry_config.sasl_oauthbearer_authorization_enabled=true \ -c schema_registry_config.sasl_oauthbearer_roles_claim_path=realm_access.roles ``` By default: * Karapace reads roles from `resource_access.karapace.roles`. * `GET` requests are allowed for `karapace.schema:read` and `karapace.subject:read`. * `POST`, `PUT`, and `DELETE` requests are blocked. ### How role-based authorization works[​](#how-role-based-authorization-works "Direct link to How role-based authorization works") Karapace extracts roles from the JWT using the configured claim path and checks them against the roles allowed for the requested HTTP method. Karapace does not create or assign roles. You create and assign roles in your identity provider. Role names are strings that you define. For example, you can use names such as `karapace.schema:read`. They are not built-in Karapace roles. In Karapace, you configure which roles can use each HTTP method. For each request, Karapace does the following: 1. Validates the JWT signature, expiration, issuer, and audience. 2. Reads the roles from the configured claim path. The default path is `resource_access.karapace.roles`. 3. Looks up the roles allowed for the requested HTTP method in `schema_registry_config.sasl_oauthbearer_method_roles`. 4. Allows the request if at least one role in the JWT matches an allowed role. Karapace matches exact role strings and does not use a role hierarchy. A role grants access only when the same string appears in both the JWT and `schema_registry_config.sasl_oauthbearer_method_roles`. If your identity provider includes roles at a different path, such as `realm_access.roles`, set `schema_registry_config.sasl_oauthbearer_roles_claim_path` to that path. ### Default HTTP method roles[​](#default-http-method-roles "Direct link to Default HTTP method roles") If you do not set `schema_registry_config.sasl_oauthbearer_method_roles`, Karapace uses this default mapping: | Action | HTTP method | Default roles | | -------------------------- | ------------- | ----------------------------------------------- | | Read schemas | `GET` | `karapace.schema:read`, `karapace.subject:read` | | Register or update schemas | `POST`, `PUT` | None | | Delete schemas | `DELETE` | None | An empty array (`[]`) means no role can use that method, even with a valid token. A client whose token includes `karapace.schema:read` can send `GET` requests. `POST`, `PUT`, and `DELETE` requests remain blocked until you [customize roles for HTTP methods](#customize-roles-for-http-methods). ``` { "GET": [ "karapace.schema:read", "karapace.subject:read" ], "POST": [], "PUT": [], "DELETE": [] } ``` ### Example JWT[​](#example-jwt "Direct link to Example JWT") The identity provider issues a token that includes the roles assigned to the user or client. For example: ``` { "sub": "alex", "resource_access": { "karapace": { "roles": [ "karapace.schema:read", "karapace.schema:write" ] } } } ``` This example omits the issuer, audience, and expiration claims. Karapace validates these claims before it reads roles. The default claim path, `resource_access.karapace.roles`, matches this example. ### Configure roles in your identity provider[​](#configure-roles-in-your-identity-provider "Direct link to Configure roles in your identity provider") How you configure roles varies by identity provider. In your identity provider, do the following: 1. Create the roles and assign them to users or clients. 2. Configure the provider to include those roles in the claim that Karapace reads. 3. Confirm that issued tokens include the roles in that claim. Use the same role strings that you plan to list in `schema_registry_config.sasl_oauthbearer_method_roles`. ### Customize roles for HTTP methods[​](#customize-roles-for-http-methods "Direct link to Customize roles for HTTP methods") Set `schema_registry_config.sasl_oauthbearer_method_roles` to a JSON string that maps each HTTP method to the roles that can use it. Karapace does not infer permissions from role names. To allow a client to use an HTTP method, list the role under that method. The following mapping uses a common role convention: | Role | HTTP methods | | ----------------------- | ------------------------------ | | `karapace.schema:read` | `GET` | | `karapace.schema:write` | `GET`, `POST`, `PUT`, `DELETE` | Each key in the JSON object is an HTTP method. Each value is a list of roles allowed for that method: ``` { "GET": [ "karapace.schema:read", "karapace.schema:write" ], "POST": [ "karapace.schema:write" ], "PUT": [ "karapace.schema:write" ], "DELETE": [ "karapace.schema:write" ] } ``` When you set this option, include `GET`, `POST`, `PUT`, and `DELETE`. To block a method, set its value to `[]`. * Console * CLI 1. In the Aiven Console, select your project and choose your Aiven for Apache Kafka service. 2. Click **Service settings**. 3. Click **Advanced configuration** > **Configure**. 4. Click **Add configuration options**. 5. Add `schema_registry_config.sasl_oauthbearer_method_roles`. 6. Enter the JSON object as a single string. 7. Click **Save configuration**. Run the following command: ``` avn service update SERVICE_NAME \ -c 'schema_registry_config.sasl_oauthbearer_method_roles={"GET":["karapace.schema:read","karapace.schema:write"],"POST":["karapace.schema:write"],"PUT":["karapace.schema:write"],"DELETE":["karapace.schema:write"]}' ``` Replace `SERVICE_NAME` with the name of your Aiven for Apache Kafka service. ## Send a request to Schema Registry[​](#send-a-request-to-schema-registry "Direct link to Send a request to Schema Registry") Send the JWT in the `Authorization` header of each Schema Registry request. You can use `curl` or any HTTP client that supports bearer tokens. On the service **Overview** page, open **Connection information** and copy the Schema Registry URL. The following example lists subjects: ``` curl \ --header "Authorization: Bearer ACCESS_TOKEN" \ "SCHEMA_REGISTRY_URL/subjects" ``` Replace the following: * `ACCESS_TOKEN`: A valid JWT from your identity provider. * `SCHEMA_REGISTRY_URL`: The Schema Registry URL from **Connection information**. This example sends a `GET` request. * If OIDC authentication is enabled and role-based authorization is disabled, any client with a valid token can send the request. * If role-based authorization is also enabled, the token must include a role allowed for `GET`. With the default mapping, the allowed roles are `karapace.schema:read` and `karapace.subject:read`. Before sending `POST`, `PUT`, or `DELETE` requests, configure an allowed role for the corresponding method. The default mapping blocks these methods. ## Disable OIDC authentication and authorization[​](#disable-oidc-authentication-and-authorization "Direct link to Disable OIDC authentication and authorization") To disable OIDC authentication and role-based authorization, set both options to **Disabled**. Disabling OIDC authentication does not affect basic authentication. Basic authentication remains enabled. * Console * CLI 1. In the Aiven Console, select your project and choose your Aiven for Apache Kafka service. 2. Click **Service settings**. 3. Click **Advanced configuration** > **Configure**. 4. Set `schema_registry_config.sasl_oauthbearer_authorization_enabled` to **Disabled**. 5. Set `schema_registry_config.sasl_oauthbearer_authentication_enabled` to **Disabled**. 6. Click **Save configuration**. Run the following command: ``` avn service update SERVICE_NAME \ -c schema_registry_config.sasl_oauthbearer_authorization_enabled=false \ -c schema_registry_config.sasl_oauthbearer_authentication_enabled=false ``` Replace `SERVICE_NAME` with the name of your Aiven for Apache Kafka service. Related pages * [Enable OAuth 2.0/OIDC authentication for Apache Kafka®](/docs/products/kafka/howto/enable-oidc.md) * [Enable OAuth 2.0/OIDC support for Apache Kafka® REST proxy](/docs/products/kafka/karapace/howto/enable-oauth-oidc-kafka-rest-proxy.md) * [Karapace schema registry authorization](/docs/products/kafka/karapace/concepts/schema-registry-authorization.md) * [Manage Karapace schema registry authorization](/docs/products/kafka/karapace/howto/manage-schema-registry-authorization.md) * [Enable schema registry and REST proxy](/docs/products/kafka/karapace/howto/enable-karapace.md) * [Set the Karapace version](/docs/products/kafka/karapace/howto/set-karapace-version.md) --- # Shut down Karapace on invalid schema records By default, Karapace Schema Registry skips invalid records in the `_schemas` topic and continues running. Enable strict mode to shut down Karapace when it detects invalid schema records. ## Why enable strict mode[​](#why-enable-strict-mode "Direct link to Why enable strict mode") Skipping invalid records can hide problems and leave the registry inconsistent. Enable strict mode when data consistency is more important than keeping Schema Registry available after it encounters invalid records. ## Enable strict mode[​](#enable-strict-mode "Direct link to Enable strict mode") * Console * CLI * API 1. Log in to the [Aiven Console](https://console.aiven.io/), select your project, and choose your **Aiven for Apache Kafka®** service. 2. On the **Overview** page, click **Service settings** from the sidebar. 3. Scroll to **Advanced configuration** and click **Configure**. 4. Click **Add configuration options**. 5. Find `schema_registry_config.schema_reader_strict_mode` and set it to **Enabled**. 6. Click **Save configuration**. Enable strict mode with the [Aiven CLI](/docs/tools/cli.md): ``` avn service update SERVICE_NAME \ -c schema_registry_config.schema_reader_strict_mode=true ``` Parameters: * `SERVICE_NAME`: Name of your Aiven for Apache Kafka® service. * `schema_registry_config.schema_reader_strict_mode=true`: Shuts down Karapace when invalid schema records are detected. Enable strict mode with the [Aiven API](/docs/tools/api.md): ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "schema_registry_config": { "schema_reader_strict_mode": true } } }' ``` Parameters: * `PROJECT_NAME`: Name of your project in Aiven. * `SERVICE_NAME`: Name of your Aiven for Apache Kafka® service. * `TOKEN`: Your API authentication [token](/docs/platform/concepts/authentication-tokens.md). * `schema_reader_strict_mode`: When `true`, Karapace shuts down if it detects invalid schema records. ## What to do if Karapace shuts down[​](#what-to-do-if-karapace-shuts-down "Direct link to What to do if Karapace shuts down") If strict mode stops Karapace because of invalid records in `_schemas`: 1. Disable strict mode so Schema Registry can start again and skip invalid records. Use the same Console, CLI, or API steps as in [Enable strict mode](#enable-strict-mode), and set the option to **Disabled** or `false`. 2. [Create a support ticket](/docs/platform/howto/support.md) or email [Aiven support](mailto:support@aiven.io) so the invalid records can be investigated and fixed. Related pages * [Enable schema registry and REST proxy](/docs/products/kafka/karapace/howto/enable-karapace.md) * [Karapace](/docs/products/kafka/karapace.md) --- # Enable Karapace schema registry authorization Most Aiven for Apache Kafka® services will automatically have [schema registry authorization](/docs/products/kafka/karapace/concepts/schema-registry-authorization.md) enabled, and the functionality cannot be disabled or enabled once a service has been created. However, some older services may pre-date this feature. To enable or disable this functionality on older services: 1. To enable schema registry authorization for a service, replace the `SERVICE_NAME` placeholder with the name of the Aiven for Apache Kafka® service in the Aiven CLI: ``` avn service update --enable-schema-registry-authorization SERVICE_NAME ``` 2. You can similarly disable the Karapace schema registry authorization using: ``` avn service update --disable-schema-registry-authorization SERVICE_NAME ``` warning Enabling Karapace schema registry authorization can disrupt access for users if the access control rules have not been configured to allow this. For more information, see [Manage Karapace schema registry authorization](/docs/products/kafka/karapace/howto/manage-schema-registry-authorization.md). --- # Manage Karapace schema registry authorization Karapace schema registry authorization allows you to authenticate the user, to control access to individual [Karapace schema registry REST API endpoints](https://github.com/aiven/karapace), and to filter the content the endpoints return. tip Some older Aiven for Apache Kafka® services may not have this feature enabled by default. In this case, [enable Karapace schema registry authorization](/docs/products/kafka/karapace/howto/enable-schema-registry-authorization.md). Karapace schema registry authorization is configured using [Access Control Lists (ACLs)](/docs/products/kafka/karapace/concepts/acl-definition.md). You can manage the Karapace schema registry authorization ACL entries using the [Aiven CLI](/docs/tools/cli/service/schema-registry-acl.md). Using the Aiven CLI commands, you can * Add ACL * Delete ACL * View ACL list For more information on the ACL commands, the required parameters and examples, see [avn service schema-registry-acl](/docs/tools/cli/service/schema-registry-acl.md). ## Manage resources via Terraform[​](#manage-resources-via-terraform "Direct link to Manage resources via Terraform") Additionally, the [Aiven Terraform Provider](/docs/tools/terraform.md) supports managing Karapace schema registry authorization ACL entries with the `aiven_kafka_schema_registry_acl` resource. For more information, see the [resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/kafka_schema_registry_acl). An example of resource configuration via Terraform is as shown below: ``` Loading... ``` --- # Register schemas with references in Karapace Use the Karapace Schema Registry API and `curl` to register Avro and Protobuf schemas that use schema references. [Schema references in Karapace](/docs/products/kafka/karapace/concepts/schema-references.md) explains which formats support references, what each `references` field means, and how Karapace resolves linked schemas during compatibility checks. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An [Aiven for Apache Kafka®](/docs/products/kafka.md) service with [Schema Registry enabled](/docs/products/kafka/karapace/howto/enable-karapace.md) * Schema Registry connection values from **Connection information** on the service overview page: * `SCHEMA_REGISTRY_URL` * `SCHEMA_REGISTRY_USER` * `SCHEMA_REGISTRY_PASSWORD` * `curl` available in your environment A successful registration returns JSON that includes an `id` for the schema version. ## Example: Avro records with references[​](#example-avro-records-with-references "Direct link to Example: Avro records with references") Avro schema references require Karapace 6.1.0 or later. In each `references` entry, `name` is a label only and can be any value. Karapace resolves the reference from the fully qualified type name in your schema (for example, `com.example.Country`), together with `subject` and `version`. Register a `Country` schema, an `Address` schema that references `Country`, a `Job` schema, and a `Person` schema that references `Address` and `Job`. 1. Register the `Country` schema: ``` curl -X POST "$SCHEMA_REGISTRY_URL/subjects/country/versions" \ -u "$SCHEMA_REGISTRY_USER:$SCHEMA_REGISTRY_PASSWORD" \ -H "Content-Type: application/vnd.schemaregistry.v1+json" \ -d '{ "schemaType": "AVRO", "schema": "{\"type\":\"record\",\"name\":\"Country\",\"namespace\":\"com.example\",\"fields\":[{\"name\":\"name\",\"type\":\"string\"},{\"name\":\"code\",\"type\":\"string\"}]}" }' ``` 2. Register the `Address` schema, referencing `Country`: ``` curl -X POST "$SCHEMA_REGISTRY_URL/subjects/address/versions" \ -u "$SCHEMA_REGISTRY_USER:$SCHEMA_REGISTRY_PASSWORD" \ -H "Content-Type: application/vnd.schemaregistry.v1+json" \ -d '{ "schemaType": "AVRO", "schema": "{\"type\":\"record\",\"name\":\"Address\",\"namespace\":\"com.example\",\"fields\":[{\"name\":\"street\",\"type\":\"string\"},{\"name\":\"city\",\"type\":\"string\"},{\"name\":\"country\",\"type\":\"com.example.Country\"}]}", "references": [ { "name": "country.avsc", "subject": "country", "version": 1 } ] }' ``` In this request, `country.avsc` is only a label. Karapace resolves the `Country` record through the `com.example.Country` type in the schema, together with `subject` and `version`. 3. Register the `Job` schema: ``` curl -X POST "$SCHEMA_REGISTRY_URL/subjects/job/versions" \ -u "$SCHEMA_REGISTRY_USER:$SCHEMA_REGISTRY_PASSWORD" \ -H "Content-Type: application/vnd.schemaregistry.v1+json" \ -d '{ "schemaType": "AVRO", "schema": "{\"type\":\"record\",\"name\":\"Job\",\"namespace\":\"com.example\",\"fields\":[{\"name\":\"title\",\"type\":\"string\"},{\"name\":\"salary\",\"type\":\"double\"}]}" }' ``` 4. Register the `Person` schema, referencing `Address` and `Job`: ``` curl -X POST "$SCHEMA_REGISTRY_URL/subjects/person/versions" \ -u "$SCHEMA_REGISTRY_USER:$SCHEMA_REGISTRY_PASSWORD" \ -H "Content-Type: application/vnd.schemaregistry.v1+json" \ -d '{ "schemaType": "AVRO", "schema": "{\"type\":\"record\",\"name\":\"Person\",\"namespace\":\"com.example\",\"fields\":[{\"name\":\"name\",\"type\":\"string\"},{\"name\":\"age\",\"type\":\"int\"},{\"name\":\"address\",\"type\":\"com.example.Address\"},{\"name\":\"job\",\"type\":\"com.example.Job\"}]}", "references": [ { "name": "address.avsc", "subject": "address", "version": 1 }, { "name": "job.avsc", "subject": "job", "version": 1 } ] }' ``` ## Example: Protobuf messages with imports[​](#example-protobuf-messages-with-imports "Direct link to Example: Protobuf messages with imports") Register an `Address` message and a `Customer` message that imports `Address`. In each `references` entry, set `name` to the import path from your `.proto` file. The value must match the `import` statement exactly (for example, `address.proto`). 1. Register the `Address` schema: ``` curl -X POST "$SCHEMA_REGISTRY_URL/subjects/address-proto/versions" \ -u "$SCHEMA_REGISTRY_USER:$SCHEMA_REGISTRY_PASSWORD" \ -H "Content-Type: application/vnd.schemaregistry.v1+json" \ -d '{ "schemaType": "PROTOBUF", "schema": "syntax = \"proto3\"; package com.example; message Address { string street = 1; string city = 2; }" }' ``` 2. Register the `Customer` schema, referencing `Address`: ``` curl -X POST "$SCHEMA_REGISTRY_URL/subjects/customer-proto/versions" \ -u "$SCHEMA_REGISTRY_USER:$SCHEMA_REGISTRY_PASSWORD" \ -H "Content-Type: application/vnd.schemaregistry.v1+json" \ -d '{ "schemaType": "PROTOBUF", "schema": "syntax = \"proto3\"; package com.example; import \"address.proto\"; message Customer { int32 id = 1; string name = 2; Address address = 3; }", "references": [ { "name": "address.proto", "subject": "address-proto", "version": 1 } ] }' ``` Related pages * [Schema references in Karapace](/docs/products/kafka/karapace/concepts/schema-references.md) * [Enable schema registry and REST proxy](/docs/products/kafka/karapace/howto/enable-karapace.md) * [Karapace](/docs/products/kafka/karapace.md) --- # Set the Karapace version You can set which Karapace version your Aiven for Apache Kafka® service runs. Set a specific version to test compatibility or to control upgrades. If you do not set a version, the service uses the latest Karapace version available through maintenance updates. After you set `karapace_version`, the service keeps that version until you change the setting or set it to `null`. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An [Aiven for Apache Kafka service](/docs/products/kafka/get-started/get-started-kafka.md) with Karapace enabled * A [maintenance update](/docs/products/kafka/howto/maintenance-updates.md#maintenance-updates) that includes Karapace version selection ## Set a Karapace version[​](#set-a-karapace-version "Direct link to Set a Karapace version") Set `karapace_version` to one of the Karapace versions available for selection. Aiven makes the two most recent Karapace versions available for selection. To upgrade or roll back, change the setting to another available version. Older Karapace versions can remain supported for services already running them. They might no longer be available for version selection. * Console * CLI * API 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. Select your project and your Aiven for Apache Kafka service. 3. Click **Service settings**. 4. Click **Advanced configuration** > **Configure**. 5. Click **Add configuration options**. 6. Set **`karapace_version`** to a version available for selection, for example, `6.2.2`. 7. Click **Save configuration**. Run the following command: ``` avn service update SERVICE_NAME --project PROJECT_NAME \ -c karapace_version=VERSION ``` Replace the following: * `SERVICE_NAME`: name of your Aiven for Apache Kafka service * `PROJECT_NAME`: name of your Aiven project * `VERSION`: Karapace version available for selection, for example, `6.2.2` Send the following `PUT` request: ``` curl --request PUT \ --url "https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME" \ --header "Authorization: Bearer API_TOKEN" \ --header "Content-Type: application/json" \ --data '{"user_config": {"karapace_version": "VERSION"}}' ``` Replace the following: * `PROJECT_NAME`: name of your Aiven project * `SERVICE_NAME`: name of your Aiven for Apache Kafka service * `API_TOKEN`: your [Aiven API token](/docs/platform/howto/create_authentication_token.md) * `VERSION`: Karapace version available for selection, for example, `6.2.2` ### Troubleshoot a failed change[​](#troubleshoot-a-failed-change "Direct link to Troubleshoot a failed change") If Aiven returns an error when you set the version, apply a [maintenance update](/docs/products/kafka/howto/maintenance-updates.md#maintenance-updates). The update enables Karapace version selection or makes the selected version available on the service. Then set the version again. ## Stop using a specific version[​](#stop-using-a-specific-version "Direct link to Stop using a specific version") Set `karapace_version` to `null` to stop using a specific version. The service then uses the latest Karapace version available through maintenance updates. * CLI * API Run the following command: ``` avn service update SERVICE_NAME --project PROJECT_NAME \ -c karapace_version=null ``` Replace the following: * `SERVICE_NAME`: name of your Aiven for Apache Kafka service * `PROJECT_NAME`: name of your Aiven project Send the following `PUT` request: ``` curl --request PUT \ --url "https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME" \ --header "Authorization: Bearer API_TOKEN" \ --header "Content-Type: application/json" \ --data '{"user_config": {"karapace_version": null}}' ``` Replace the following: * `PROJECT_NAME`: name of your Aiven project * `SERVICE_NAME`: name of your Aiven for Apache Kafka service * `API_TOKEN`: your [Aiven API token](/docs/platform/howto/create_authentication_token.md) ## Check the Karapace version[​](#check-the-karapace-version "Direct link to Check the Karapace version") Check `karapace_version` to see whether the service is set to a specific version. * Console * CLI * API 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. Select your project and your Aiven for Apache Kafka service. 3. Click **Service settings**. 4. Click **Advanced configuration**. 5. Check the **`karapace_version`** value. Run the following command: ``` avn service get SERVICE_NAME --project PROJECT_NAME --json ``` Replace the following: * `SERVICE_NAME`: name of your Aiven for Apache Kafka service * `PROJECT_NAME`: name of your Aiven project Send the following `GET` request: ``` curl --request GET \ --url "https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME" \ --header "Authorization: Bearer API_TOKEN" ``` Replace the following: * `PROJECT_NAME`: name of your Aiven project * `SERVICE_NAME`: name of your Aiven for Apache Kafka service * `API_TOKEN`: your [Aiven API token](/docs/platform/howto/create_authentication_token.md) In the Aiven Console, if **`karapace_version`** shows a version number, the service is set to that version. If the option is not listed or has no value, the service is not set to a specific version. In the CLI output or API response, check `karapace_version` in `user_config`. If it contains a version number, the service is set to that version. If the value is `null` or the field is not present, the service is not set to a specific version. ## Related pages[​](#related-pages "Direct link to Related pages") * [Apache Kafka maintenance updates](/docs/products/kafka/howto/maintenance-updates.md#maintenance-updates) * [Aiven CLI reference: avn service update](/docs/tools/cli/service-cli.md#avn-cli-service-update) * [Aiven REST API reference](https://api.aiven.io/doc/) --- # Advanced parameters for Aiven for Apache Kafka® Classic View the configuration options for Aiven for Apache Kafka® Classic. | Parameter | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | []()[**backup\_interval\_hours**](#backup_interval_hours)`integer,null`- min: `3`
- max: `24`Interval in hours between automatic backups. Minimum value is 3 hours. Must be a divisor of 24 (3, 4, 6, 8, 12, 24). (Applicable to ACU plans only) | | []()[**backup\_retention\_days**](#backup_retention_days)`integer,null`- min: `1`
- max: `30`Backup retention in daysNumber of days to retain automatic backups. Backups older than this value will be automatically deleted. (Applicable to ACU plans only) | | []()[**custom\_domain**](#custom_domain)`string,null`Serve the web frontend using a custom CNAME pointing to the Aiven DNS name. When you set a custom domain for a service deployed in a VPC, the service certificate is only created for the public-\* hostname and the custom domain. | | []()[**ip\_filter**](#ip_filter)`array`- default: `0.0.0.0/0,::/0`IP filterAllow incoming connections from CIDR address block, e.g. '10.20.0.0/16' | | []()[**service\_log**](#service_log)`boolean,null`Service loggingStore logs for the service so that they are available in the HTTP API and console. | | []()[**static\_ips**](#static_ips)`boolean`Use static public IP addresses | | []()[**single\_zone**](#single_zone)`object`Single-zone configurationsingle\_zone.enabled boolean Enabled Whether to allocate nodes on the same Availability Zone or spread across zones available. By default service nodes are spread across different AZs. The single AZ support is best-effort and may temporarily allocate nodes in different AZs e.g. in case of capacity limitations in one AZ. single\_zone.availability\_zone string The availability zone to use for the service. This is only used when enabled is set to true. If not set the service will be allocated in random AZ.The AZ is not guaranteed, and the service may be allocated in a different AZ if the selected AZ is not available. Zones will not be validated and invalid zones will be ignored, falling back to random AZ selection. Common availability zones include: AWS (euc1-az1, euc1-az2, euc1-az3), GCP (europe-west1-a, europe-west1-b, europe-west1-c), Azure (germanywestcentral/1, germanywestcentral/2, germanywestcentral/3). | | []()[**preferred\_zones**](#preferred_zones)`array,null`List of preferred zone IDs for service node placement. Nodes will be placed in these zones when available. If a specified zone is unavailable (e.g., due to capacity constraints), nodes will be placed in other available zones to maintain the configured number of zones for availability. Invalid zone IDs are rejected at configuration time. Zone IDs are cloud-specific: AWS uses zone IDs like 'euc1-az1', GCP uses zone names like 'europe-west1-a', and Azure uses 'location/zone' format like 'germanywestcentral/1'. If single\_zone is enabled with an availability\_zone, that setting takes precedence over preferred\_zones. Changes take effect on next node recreation (e.g., maintenance or plan change). For eligible plans, nodes outside preferred zones are automatically rebalanced once per day. | | []()[**private\_access**](#private_access)`object`Allow access to selected service ports from private networksprivate\_access.kafka boolean Allow clients to connect to kafka with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.kafka\_connect boolean Allow clients to connect to kafka\_connect with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.kafka\_rest boolean Allow clients to connect to kafka\_rest with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.prometheus boolean Allow clients to connect to prometheus with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.schema\_registry boolean Allow clients to connect to schema\_registry with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | []()[**public\_access**](#public_access)`object`Allow access to selected service ports from the public Internetpublic\_access.kafka boolean Allow clients to connect to kafka from the public internet for service nodes that are in a project VPC or another type of private network public\_access.kafka\_connect boolean Allow clients to connect to kafka\_connect from the public internet for service nodes that are in a project VPC or another type of private network public\_access.kafka\_rest boolean Allow clients to connect to kafka\_rest from the public internet for service nodes that are in a project VPC or another type of private network public\_access.prometheus boolean Allow clients to connect to prometheus from the public internet for service nodes that are in a project VPC or another type of private network public\_access.schema\_registry boolean Allow clients to connect to schema\_registry from the public internet for service nodes that are in a project VPC or another type of private network | | []()[**privatelink\_access**](#privatelink_access)`object`Allow access to selected service components through Privatelinkprivatelink\_access.jolokia boolean Enable jolokia privatelink\_access.kafka boolean Enable kafka privatelink\_access.kafka\_connect boolean Enable kafka\_connect privatelink\_access.kafka\_rest boolean Enable kafka\_rest privatelink\_access.prometheus boolean Enable prometheus privatelink\_access.schema\_registry boolean Enable schema\_registry | | []()[**letsencrypt\_sasl**](#letsencrypt_sasl)`boolean,null`Use a Let's Encrypt certificate authority (CA) for Kafka SASL authentication. (Default: False) | | []()[**letsencrypt\_sasl\_privatelink**](#letsencrypt_sasl_privatelink)`boolean,null`Use a Let's Encrypt certificate authority (CA) for Kafka SASL authentication via Privatelink. (Default: False) | | []()[**kafka**](#kafka)`object`- default: `[object Object]`Kafka broker configuration valueskafka.compression\_type string compression.type Specify the final compression type for a given topic. This configuration accepts the standard compression codecs ('gzip', 'snappy', 'lz4', 'zstd'). It additionally accepts 'uncompressed' which is equivalent to no compression; and 'producer' which means retain the original compression codec set by the producer.(Default: producer) kafka.group\_initial\_rebalance\_delay\_ms integer - max: 300000 group.initial.rebalance.delay.ms The amount of time, in milliseconds, the group coordinator will wait for more consumers to join a new group before performing the first rebalance. A longer delay means potentially fewer rebalances, but increases the time until processing begins. The default value for this is 3 seconds. During development and testing it might be desirable to set this to 0 in order to not delay test execution time. (Default: 3000 ms (3 seconds)) kafka.group\_coordinator\_rebalance\_protocols string group.coordinator.rebalance.protocols The enabled consumer group rebalance protocols. Use consumer, classic, share, streams to enable Kafka share groups. kafka.group\_share\_session\_timeout\_ms integer group.share.session.timeout.ms The timeout used to detect share group member failures. kafka.group\_share\_min\_session\_timeout\_ms integer group.share.min.session.timeout.ms The minimum session timeout allowed for share group members. kafka.group\_share\_max\_session\_timeout\_ms integer group.share.max.session.timeout.ms The maximum session timeout allowed for share group members. kafka.group\_share\_heartbeat\_interval\_ms integer group.share.heartbeat.interval.ms The heartbeat interval used by share group members. kafka.group\_share\_min\_heartbeat\_interval\_ms integer group.share.min.heartbeat.interval.ms The minimum heartbeat interval allowed for share group members. kafka.group\_share\_max\_heartbeat\_interval\_ms integer group.share.max.heartbeat.interval.ms The maximum heartbeat interval allowed for share group members. kafka.group\_share\_record\_lock\_duration\_ms integer - min: 1000 - max: 3600000 group.share.record.lock.duration.ms The duration for which a fetched share-group record is locked. kafka.group\_share\_min\_record\_lock\_duration\_ms integer group.share.min.record.lock.duration.ms The minimum record lock duration allowed for share groups. kafka.group\_share\_max\_record\_lock\_duration\_ms integer group.share.max.record.lock.duration.ms The maximum record lock duration allowed for share groups. kafka.group\_share\_partition\_max\_record\_locks integer - min: 100 - max: 10000 group.share.partition.max.record.locks The maximum number of record locks allowed per share group partition. kafka.group\_share\_delivery\_count\_limit integer - min: 2 - max: 10 group.share.delivery.count.limit The maximum delivery attempts for a share-group record. kafka.group\_share\_max\_size integer - min: 10 - max: 1000 group.share.max.size The maximum number of members allowed in a share group. kafka.group\_share\_max\_groups integer group.share.max.groups The maximum number of share groups allowed on the broker. kafka.group\_min\_session\_timeout\_ms integer - max: 60000 group.min.session.timeout.ms The minimum allowed session timeout for registered consumers. Longer timeouts give consumers more time to process messages in between heartbeats at the cost of a longer time to detect failures. (Default: 6000 ms (6 seconds)) kafka.group\_max\_session\_timeout\_ms integer - max: 1800000 group.max.session.timeout.ms The maximum allowed session timeout for registered consumers. Longer timeouts give consumers more time to process messages in between heartbeats at the cost of a longer time to detect failures. Default: 1800000 ms (30 minutes) kafka.connections\_max\_idle\_ms integer - min: 1000 - max: 3600000 connections.max.idle.ms Idle connections timeout: the server socket processor threads close the connections that idle for longer than this. (Default: 600000 ms (10 minutes)) kafka.max\_incremental\_fetch\_session\_cache\_slots integer - min: 1000 - max: 10000 max.incremental.fetch.session.cache.slots The maximum number of incremental fetch sessions that the broker will maintain. (Default: 1000) kafka.message\_max\_bytes integer - max: 100001200 message.max.bytes The maximum size of message that the server can receive. (Default: 1048588 bytes (1 mebibyte + 12 bytes)) kafka.offsets\_retention\_minutes integer - min: 1 - max: 2147483647 offsets.retention.minutes Log retention window in minutes for offsets topic (Default: 10080 minutes (7 days)) kafka.log\_cleaner\_delete\_retention\_ms integer - max: 315569260000 log.cleaner.delete.retention.ms How long are delete records retained? (Default: 86400000 (1 day)) kafka.log\_cleaner\_min\_cleanable\_ratio number - min: 0.2 - max: 0.9 log.cleaner.min.cleanable.ratio Controls log compactor frequency. Larger value means more frequent compactions but also more space wasted for logs. Consider setting log.cleaner.max.compaction.lag.ms to enforce compactions sooner, instead of setting a very high value for this option. (Default: 0.5) kafka.log\_cleaner\_max\_compaction\_lag\_ms integer - min: 30000 - max: 9223372036854776000 log.cleaner.max.compaction.lag.ms The maximum amount of time message will remain uncompacted. Only applicable for logs that are being compacted. (Default: 9223372036854775807 ms (Long.MAX\_VALUE)) kafka.log\_cleaner\_min\_compaction\_lag\_ms integer - max: 9223372036854776000 log.cleaner.min.compaction.lag.ms The minimum time a message will remain uncompacted in the log. Only applicable for logs that are being compacted. (Default: 0 ms) kafka.log\_cleanup\_policy string log.cleanup.policy The default cleanup policy for segments beyond the retention window (Default: delete) kafka.log\_flush\_interval\_messages integer - min: 1 - max: 9223372036854776000 log.flush.interval.messages The number of messages accumulated on a log partition before messages are flushed to disk (Default: 9223372036854775807 (Long.MAX\_VALUE)) kafka.log\_flush\_interval\_ms integer - max: 9223372036854776000 log.flush.interval.ms The maximum time in ms that a message in any topic is kept in memory (page-cache) before flushed to disk. If not set, the value in log.flush.scheduler.interval.ms is used (Default: null) kafka.log\_index\_interval\_bytes integer - max: 104857600 log.index.interval.bytes The interval with which Kafka adds an entry to the offset index (Default: 4096 bytes (4 kibibytes)) kafka.log\_index\_size\_max\_bytes integer - min: 1048576 - max: 104857600 log.index.size.max.bytes The maximum size in bytes of the offset index (Default: 10485760 (10 mebibytes)) kafka.log\_local\_retention\_ms integer - min: -2 - max: 9223372036854776000 log.local.retention.ms The number of milliseconds to keep the local log segments before it gets eligible for deletion. If set to -2, the value of log.retention.ms is used. The effective value should always be less than or equal to log.retention.ms value. (Default: -2) kafka.log\_local\_retention\_bytes integer - min: -2 - max: 9223372036854776000 log.local.retention.bytes The maximum size of local log segments that can grow for a partition before it gets eligible for deletion. If set to -2, the value of log.retention.bytes is used. The effective value should always be less than or equal to log.retention.bytes value. (Default: -2) kafka.log\_message\_downconversion\_enable boolean log.message.downconversion.enable This configuration controls whether down-conversion of message formats is enabled to satisfy consume requests. (Default: true) kafka.log\_message\_timestamp\_type string log.message.timestamp.type Define whether the timestamp in the message is message create time or log append time. (Default: CreateTime) kafka.log\_message\_timestamp\_difference\_max\_ms integer - max: 9223372036854776000 log.message.timestamp.difference.max.ms The maximum difference allowed between the timestamp when a broker receives a message and the timestamp specified in the message (Default: 9223372036854775807 (Long.MAX\_VALUE)) kafka.log\_message\_timestamp\_before\_max\_ms integer - max: 9223372036854776000 log.message.timestamp.before.max.ms The maximum difference allowed between the timestamp when a broker receives a message and the timestamp specified in the message. If message.timestamp.type=CreateTime, a message will be rejected if the difference in timestamp exceeds this threshold. Applies only for messages with timestamps earlier than the broker's timestamp. (Default: 9223372036854775807 (Long.MAX\_VALUE)) kafka.log\_message\_timestamp\_after\_max\_ms integer - max: 9223372036854776000 log.message.timestamp.after.max.ms The maximum difference allowed between the timestamp when a broker receives a message and the timestamp specified in the message. If message.timestamp.type=CreateTime, a message will be rejected if the difference in timestamp exceeds this threshold. Applies only for messages with timestamps later than the broker's timestamp. (Default: 9223372036854775807 (Long.MAX\_VALUE)) kafka.log\_preallocate boolean log.preallocate Should pre allocate file when create new segment? (Default: false) kafka.log\_retention\_bytes integer - min: -1 - max: 9223372036854776000 log.retention.bytes The maximum size of the log before deleting messages (Default: -1) kafka.log\_retention\_hours integer - min: -1 - max: 2147483647 log.retention.hours The number of hours to keep a log file before deleting it. Use -1 for unlimited retention or 1 or higher. Setting 0 is invalid and prevents Kafka from starting. (Default: 168 hours, or 1 week) kafka.log\_retention\_ms integer - min: -1 - max: 9223372036854776000 log.retention.ms The number of milliseconds to keep a log file before deleting it (in milliseconds), If not set, the value in log.retention.minutes is used. If set to -1, no time limit is applied. (Default: null, log.retention.hours applies) kafka.log\_roll\_jitter\_ms integer - max: 9223372036854776000 log.roll.jitter.ms The maximum jitter to subtract from logRollTimeMillis (in milliseconds). If not set, the value in log.roll.jitter.hours is used (Default: null) kafka.log\_roll\_ms integer - min: 1 - max: 9223372036854776000 log.roll.ms The maximum time before a new log segment is rolled out (in milliseconds). (Default: null, log.roll.hours applies (Default: 168, 7 days)) kafka.log\_segment\_bytes integer - min: 10485760 - max: 1073741824 log.segment.bytes The maximum size of a single log file (Default: 1073741824 bytes (1 gibibyte)) kafka.log\_segment\_delete\_delay\_ms integer - max: 3600000 log.segment.delete.delay.ms The amount of time to wait before deleting a file from the filesystem (Default: 60000 ms (1 minute)) kafka.auto\_create\_topics\_enable boolean auto.create.topics.enable Enable auto-creation of topics. (Default: false) kafka.min\_insync\_replicas integer - min: 1 - max: 7 When a producer sets acks to 'all' (or '-1'), min.insync.replicas specifies the minimum number of replicas that must acknowledge a write for the write to be considered successful. (Default: 1) kafka.num\_partitions integer - min: 1 - max: 1000 num.partitions Number of partitions for auto-created topics (Default: 1) kafka.default\_replication\_factor integer - min: 1 - max: 10 default.replication.factor Replication factor for auto-created topics (Default: 3) kafka.replica\_fetch\_max\_bytes integer - min: 1048576 - max: 104857600 replica.fetch.max.bytes The number of bytes of messages to attempt to fetch for each partition . This is not an absolute maximum, if the first record batch in the first non-empty partition of the fetch is larger than this value, the record batch will still be returned to ensure that progress can be made. (Default: 1048576 bytes (1 mebibytes)) kafka.replica\_fetch\_response\_max\_bytes integer - min: 10485760 - max: 1048576000 replica.fetch.response.max.bytes Maximum bytes expected for the entire fetch response. Records are fetched in batches, and if the first record batch in the first non-empty partition of the fetch is larger than this value, the record batch will still be returned to ensure that progress can be made. As such, this is not an absolute maximum. (Default: 10485760 bytes (10 mebibytes)) kafka.max\_connections\_per\_ip integer - min: 256 - max: 2147483647 max.connections.per.ip The maximum number of connections allowed from each ip address (Default: 2147483647). kafka.producer\_purgatory\_purge\_interval\_requests integer - min: 10 - max: 10000 producer.purgatory.purge.interval.requests The purge interval (in number of requests) of the producer request purgatory (Default: 1000). kafka.sasl\_oauthbearer\_expected\_audience string sasl.oauthbearer.expected.audience The (optional) comma-delimited setting for the broker to use to verify that the JWT was issued for one of the expected audiences. (Default: null) kafka.sasl\_oauthbearer\_expected\_issuer string sasl.oauthbearer.expected.issuer Optional setting for the broker to use to verify that the JWT was created by the expected issuer.(Default: null) kafka.sasl\_oauthbearer\_jwks\_endpoint\_url string sasl.oauthbearer.jwks.endpoint.url OIDC JWKS endpoint URL. By setting this the SASL SSL OAuth2/OIDC authentication is enabled. See also other options for SASL OAuth2/OIDC. (Default: null) kafka.sasl\_oauthbearer\_sub\_claim\_name string sasl.oauthbearer.sub.claim.name Name of the scope from which to extract the subject claim from the JWT.(Default: sub) kafka.socket\_request\_max\_bytes integer - min: 10485760 - max: 209715200 socket.request.max.bytes The maximum number of bytes in a socket request (Default: 104857600 bytes). kafka.transaction\_state\_log\_segment\_bytes integer - min: 1048576 - max: 2147483647 transaction.state.log.segment.bytes The transaction topic segment bytes should be kept relatively small in order to facilitate faster log compaction and cache loads (Default: 104857600 bytes (100 mebibytes)). kafka.transaction\_remove\_expired\_transaction\_cleanup\_interval\_ms integer - min: 600000 - max: 3600000 transaction.remove.expired.transaction.cleanup.interval.ms The interval at which to remove transactions that have expired due to transactional.id.expiration.ms passing (Default: 3600000 ms (1 hour)). kafka.transaction\_partition\_verification\_enable boolean transaction.partition.verification.enable Enable verification that checks that the partition has been added to the transaction before writing transactional records to the partition. (Default: true) kafka.audit\_log object Enable Kafka audit logging by providing this object. Removing it disables the feature. Enabling, updating, or disabling audit logging causes a rolling restart of all Kafka brokers. kafka.audit\_log.record\_type string - default: user\_operations Audit log type user\_operations records individual Kafka API calls (produce, fetch, etc.). user\_activity records higher-level user actions. kafka.audit\_log.aggregation\_period\_sec integer - min: 1 - max: 1800 - default: 300 Aggregation period in seconds over which audit log entries are batched before being emitted. kafka.audit\_log.include\_denials boolean Include log denials Whether to include denied authorization attempts in the audit log. kafka.audit\_log.group\_by string - default: user\_and\_ip Group audit log entries by user or by user and IP address. Only valid when record\_type is user\_operations. | | []()[**kafka\_authentication\_methods**](#kafka_authentication_methods)`object`Kafka authentication methodskafka\_authentication\_methods.certificate boolean - default: true Enable certificate/SSL authentication kafka\_authentication\_methods.sasl boolean Enable SASL authentication | | []()[**kafka\_sasl\_mechanisms**](#kafka_sasl_mechanisms)`object`Kafka SASL mechanismskafka\_sasl\_mechanisms.plain boolean - default: true Enable PLAIN mechanism kafka\_sasl\_mechanisms.scram\_sha\_256 boolean - default: true Enable SCRAM-SHA-256 mechanism kafka\_sasl\_mechanisms.scram\_sha\_512 boolean - default: true Enable SCRAM-SHA-512 mechanism | | []()[**follower\_fetching**](#follower_fetching)`object`Enable follower fetchingfollower\_fetching.enabled boolean Enabled Whether to enable the follower fetching functionality | | []()[**kafka\_connect**](#kafka_connect)`boolean`Enable Kafka Connect service | | []()[**kafka\_connect\_config**](#kafka_connect_config)`object`Kafka Connect configuration valueskafka\_connect\_config.prefer\_ipv6\_address\_enable boolean When enabled, connectors will automatically resolve IPv6 addresses from external server names configured with dual-stack. kafka\_connect\_config.connector\_client\_config\_override\_policy string Client config override policy Defines what client configurations can be overridden by the connector. Default is None kafka\_connect\_config.consumer\_auto\_offset\_reset string Consumer auto offset reset What to do when there is no initial offset in Kafka or if the current offset does not exist any more on the server. Default is earliest kafka\_connect\_config.consumer\_fetch\_max\_bytes integer - min: 1048576 - max: 104857600 The maximum amount of data the server should return for a fetch request Records are fetched in batches by the consumer, and if the first record batch in the first non-empty partition of the fetch is larger than this value, the record batch will still be returned to ensure that the consumer can make progress. As such, this is not a absolute maximum. kafka\_connect\_config.consumer\_isolation\_level string Consumer isolation level Transaction read isolation level. read\_uncommitted is the default, but read\_committed can be used if consume-exactly-once behavior is desired. kafka\_connect\_config.consumer\_max\_partition\_fetch\_bytes integer - min: 1048576 - max: 104857600 The maximum amount of data per-partition the server will return. Records are fetched in batches by the consumer.If the first record batch in the first non-empty partition of the fetch is larger than this limit, the batch will still be returned to ensure that the consumer can make progress. kafka\_connect\_config.consumer\_max\_poll\_interval\_ms integer - min: 1 - max: 2147483647 The maximum delay in milliseconds between invocations of poll() when using consumer group management (defaults to 300000). kafka\_connect\_config.consumer\_max\_poll\_records integer - min: 1 - max: 10000 The maximum number of records returned in a single call to poll() (defaults to 500). kafka\_connect\_config.offset\_flush\_interval\_ms integer - min: 1 - max: 100000000 The interval at which to try committing offsets for tasks (defaults to 60000). kafka\_connect\_config.offset\_flush\_timeout\_ms integer - min: 1 - max: 2147483647 Maximum number of milliseconds to wait for records to flush and partition offset data to be committed to offset storage before cancelling the process and restoring the offset data to be committed in a future attempt (defaults to 5000). kafka\_connect\_config.producer\_batch\_size integer - max: 5242880 The batch size in bytes the producer will attempt to collect for the same partition before publishing to broker This setting gives the upper bound of the batch size to be sent. If there are fewer than this many bytes accumulated for this partition, the producer will 'linger' for the linger.ms time waiting for more records to show up. A batch size of zero will disable batching entirely (defaults to 16384). kafka\_connect\_config.producer\_buffer\_memory integer - min: 5242880 - max: 134217728 The total bytes of memory the producer can use to buffer records waiting to be sent to the broker (defaults to 33554432). kafka\_connect\_config.producer\_compression\_type string Specify the default compression type for producers. This configuration accepts the standard compression codecs ('gzip', 'snappy', 'lz4', 'zstd'). It additionally accepts 'none' which is the default and equivalent to no compression. kafka\_connect\_config.producer\_linger\_ms integer - max: 5000 Wait for up to the given delay to allow batching records together This setting gives the upper bound on the delay for batching: once there is batch.size worth of records for a partition it will be sent immediately regardless of this setting, however if there are fewer than this many bytes accumulated for this partition the producer will 'linger' for the specified time waiting for more records to show up. Defaults to 0. kafka\_connect\_config.producer\_max\_request\_size integer - min: 131072 - max: 67108864 The maximum size of a request in bytes This setting will limit the number of record batches the producer will send in a single request to avoid sending huge requests. kafka\_connect\_config.scheduled\_rebalance\_max\_delay\_ms integer - max: 600000 The maximum delay that is scheduled in order to wait for the return of one or more departed workers before rebalancing and reassigning their connectors and tasks to the group. During this period the connectors and tasks of the departed workers remain unassigned. Defaults to 5 minutes. kafka\_connect\_config.session\_timeout\_ms integer - min: 1 - max: 2147483647 The timeout in milliseconds used to detect failures when using Kafka’s group management facilities (defaults to 10000). | | []()[**kafka\_connect\_plugin\_versions**](#kafka_connect_plugin_versions)`array`Kafka Connect pluginsThe plugin selected by the user | | []()[**kafka\_connect\_secret\_providers**](#kafka_connect_secret_providers)`array`Kafka Connect secret providersConfigure external secret providers in order to reference external secrets in connector configuration. Currently Hashicorp Vault (provider: vault, auth\_method: token) and AWS Secrets Manager (provider: aws, auth\_method: credentials) are supported. Secrets can be referenced in connector config with ${\:\:\} | | []()[**kafka\_diskless**](#kafka_diskless)`object`Kafka Diskless configuration valueskafka\_diskless.enabled boolean Enabled Whether to enable the Diskless functionality kafka\_diskless.auto\_diskless\_topic\_regexes array,null The regexes of topics to auto enable diskless. Topics matching any of the regexes will be created as diskless topics. | | []()[**inkless**](#inkless)`object`Inkless configuration valuesinkless.enabled boolean Enabled Whether to enable the Inkless functionality | | []()[**sasl\_oauthbearer\_allowed\_urls**](#sasl_oauthbearer_allowed_urls)`array`List of allowed URLs for SASL OAUTHBEARER authentication. Only HTTPS URLs are allowed for security reasons. | | []()[**gcp\_auth\_allowed\_urls**](#gcp_auth_allowed_urls)`array`Allowed URLs for auth URI validationAllow-list of HTTPS URLs used to validate GCP credential\_source requests for Kafka Connect. | | []()[**kafka\_rest**](#kafka_rest)`boolean`Enable Kafka-REST service | | []()[**kafka\_version**](#kafka_version)`string,null`Kafka major version | | []()[**karapace\_version**](#karapace_version)`string,null`Select a Karapace version for this service, or select Latest to use the latest available version automatically. New versions become available after installation during a maintenance update. | | []()[**schema\_registry**](#schema_registry)`boolean`Enable Schema-Registry service | | []()[**kafka\_rest\_authorization**](#kafka_rest_authorization)`boolean`Enable authorization in Kafka-REST service | | []()[**kafka\_rest\_config**](#kafka_rest_config)`object`Kafka REST configurationkafka\_rest\_config.producer\_acks string - default: 1 producer.acks The number of acknowledgments the producer requires the leader to have received before considering a request complete. If set to 'all' or '-1', the leader will wait for the full set of in-sync replicas to acknowledge the record. kafka\_rest\_config.producer\_compression\_type string producer.compression.type Specify the default compression type for producers. This configuration accepts the standard compression codecs ('gzip', 'snappy', 'lz4', 'zstd'). It additionally accepts 'none' which is the default and equivalent to no compression. kafka\_rest\_config.producer\_linger\_ms integer - max: 5000 producer.linger.ms Wait for up to the given delay to allow batching records together kafka\_rest\_config.producer\_max\_request\_size integer - max: 2147483647 - default: 1048576 producer.max.request.size The maximum size of a request in bytes. Note that Kafka broker can also cap the record batch size. kafka\_rest\_config.consumer\_enable\_auto\_commit boolean - default: true consumer.enable.auto.commit If true the consumer's offset will be periodically committed to Kafka in the background kafka\_rest\_config.consumer\_idle\_disconnect\_timeout integer - max: 2147483647 consumer.idle.disconnect.timeout Specifies the maximum duration (in seconds) a client can remain idle before it is deleted. If a consumer is inactive, it will exit the consumer group, and its state will be discarded. A value of 0 (default) indicates that the consumer will not be disconnected automatically due to inactivity. kafka\_rest\_config.consumer\_request\_max\_bytes integer - max: 671088640 - default: 67108864 consumer.request.max.bytes Maximum number of bytes in unencoded message keys and values by a single request kafka\_rest\_config.consumer\_request\_timeout\_ms integer - min: 1000 - max: 30000 - default: 1000 consumer.request.timeout.ms The maximum total time to wait for messages for a request if the maximum number of messages has not yet been reached kafka\_rest\_config.name\_strategy string - default: topic\_name name.strategy Name strategy to use when selecting subject for storing schemas kafka\_rest\_config.name\_strategy\_validation boolean - default: true name.strategy.validation If true, validate that given schema is registered under expected subject name by the used name strategy when producing messages. kafka\_rest\_config.simpleconsumer\_pool\_size\_max integer - min: 10 - max: 250 - default: 25 simpleconsumer.pool.size.max Maximum number of SimpleConsumers that can be instantiated per broker | | []()[**tiered\_storage**](#tiered_storage)`object`Tiered storage configurationtiered\_storage.enabled boolean Enabled Whether to enable the tiered storage functionality | | []()[**schema\_registry\_config**](#schema_registry_config)`object`Schema Registry configurationschema\_registry\_config.topic\_name string topic\_name The durable single partition topic that acts as the durable log for the data. This topic must be compacted to avoid losing data due to retention policy. Please note that changing this configuration in an existing Schema Registry / Karapace setup leads to previous schemas being inaccessible, data encoded with them potentially unreadable and schema ID sequence put out of order. It's only possible to do the switch while Schema Registry / Karapace is disabled. Defaults to \_schemas. schema\_registry\_config.leader\_eligibility boolean leader\_eligibility If true, Karapace / Schema Registry on the service nodes can participate in leader election. It might be needed to disable this when the schemas topic is replicated to a secondary cluster and Karapace / Schema Registry there must not participate in leader election. Defaults to true. schema\_registry\_config.schema\_reader\_strict\_mode boolean schema\_reader\_strict\_mode If enabled, causes the Karapace schema-registry service to shutdown when there are invalid schema records in the \_schemas topic. Defaults to false. schema\_registry\_config.retriable\_errors\_silenced boolean retriable\_errors\_silenced If enabled, kafka errors which can be retried or custom errors specified for the service will not be raised, instead, a warning log is emitted. This will denoise issue tracking systems, i.e. sentry. Defaults to true. schema\_registry\_config.sasl\_oauthbearer\_authentication\_enabled boolean Enable OIDC authentication for Schema Registry If enabled, the Schema Registry validates OAuth 2.0/OIDC JWT bearer tokens. Requires sasl\_oauthbearer\_jwks\_endpoint\_url, sasl\_oauthbearer\_expected\_issuer, and sasl\_oauthbearer\_expected\_audience under kafka. Defaults to false. schema\_registry\_config.sasl\_oauthbearer\_authorization\_enabled boolean Enable OIDC authorization for Schema Registry If enabled, the Schema Registry enforces role-based authorization using the JWT roles claim. It also enables sasl\_oauthbearer\_authentication\_enabled if it isn't already enabled. Authorization requires authentication. Defaults to false. schema\_registry\_config.sasl\_oauthbearer\_roles\_claim\_path string sasl.oauthbearer.roles.claim.path The JSON path the Schema Registry uses to find the roles claim in the JWT. Defaults to resource\_access.karapace.roles. schema\_registry\_config.sasl\_oauthbearer\_method\_roles string sasl.oauthbearer.method.roles Maps HTTP methods to allowed roles. Use a JSON object with GET, POST, PUT, and DELETE keys mapped to arrays of roles. Role names use the karapace. prefix. Example: \\{"GET": \["karapace.schema:read"], "POST": \[], "PUT": \[], "DELETE": \[]\\}. | | []()[**aiven\_kafka\_topic\_messages**](#aiven_kafka_topic_messages)`boolean`Allow access to read Kafka topic messages in the Aiven Console and REST API. | | []()[**enable\_ipv6**](#enable_ipv6)`boolean`Enable IPv6Register AAAA DNS records for the service, and allow IPv6 packets to service ports | --- # Advanced parameters for Aiven for Apache Kafka® Developer tier View the configuration options for Aiven for Apache Kafka® Developer tier. ## Broker parameters[​](#broker-parameters "Direct link to Broker parameters") | | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | [**backup\_interval\_hours**](#backup_interval_hours)`integer,null`- min: `3`
- max: `24`
- enum: `3,4,6,8,12,24,null`Interval in hours between automatic backups. Minimum value is 3 hours. Must be a divisor of 24 (3, 4, 6, 8, 12, 24). (Applicable to ACU plans only) | | [**backup\_retention\_days**](#backup_retention_days)`integer,null`- min: `1`
- max: `30`Number of days to retain automatic backups. Backups older than this value will be automatically deleted. (Applicable to ACU plans only) | | [**custom\_domain**](#custom_domain)`string,null`- maxLength: `255`Serve the web frontend using a custom CNAME pointing to the Aiven DNS name. When you set a custom domain for a service deployed in a VPC, the service certificate is only created for the public-\* hostname and the custom domain. | | [**ip\_filter**](#ip_filter)`array[string,object]`- maxItems: `8000`
- default: `0.0.0.0/0,::/0`Allow incoming connections from CIDR address block, e.g. '10.20.0.0/16' | | [**ip\_filter.\[0\].description**](#ip_filter.\[0].description)`string`- maxLength: `1024`Description for IP filter list entry | | [**ip\_filter.\[0\].network**](#ip_filter.\[0].network)`string`- maxLength: `43`CIDR address block | | [**service\_log**](#service_log)`boolean,null`Store logs for the service so that they are available in the HTTP API and console. | | [**static\_ips**](#static_ips)`boolean`Use static public IP addresses | | [**single\_zone**](#single_zone)`object`Single-zone configuration | | [**single\_zone.enabled**](#single_zone.enabled)`boolean`Whether to allocate nodes on the same Availability Zone or spread across zones available. By default service nodes are spread across different AZs. The single AZ support is best-effort and may temporarily allocate nodes in different AZs e.g. in case of capacity limitations in one AZ. | | [**single\_zone.availability\_zone**](#single_zone.availability_zone)`string`- maxLength: `40`The availability zone to use for the service. This is only used when enabled is set to true. If not set the service will be allocated in random AZ.The AZ is not guaranteed, and the service may be allocated in a different AZ if the selected AZ is not available. Zones will not be validated and invalid zones will be ignored, falling back to random AZ selection. Common availability zones include: AWS (euc1-az1, euc1-az2, euc1-az3), GCP (europe-west1-a, europe-west1-b, europe-west1-c), Azure (germanywestcentral/1, germanywestcentral/2, germanywestcentral/3). | | [**preferred\_zones**](#preferred_zones)`array[string]`- maxItems: `10`List of preferred zone IDs for service node placement. Nodes will be placed in these zones when available. If a specified zone is unavailable (e.g., due to capacity constraints), nodes will be placed in other available zones to maintain the configured number of zones for availability. Invalid zone IDs are rejected at configuration time. Zone IDs are cloud-specific: AWS uses zone IDs like 'euc1-az1', GCP uses zone names like 'europe-west1-a', and Azure uses 'location/zone' format like 'germanywestcentral/1'. If single\_zone is enabled with an availability\_zone, that setting takes precedence over preferred\_zones. Changes take effect on next node recreation (e.g., maintenance or plan change). For eligible plans, nodes outside preferred zones are automatically rebalanced once per day. | | [**private\_access**](#private_access)`object`Allow access to selected service ports from private networks | | [**private\_access.kafka**](#private_access.kafka)`boolean`Allow clients to connect to kafka with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | [**private\_access.kafka\_connect**](#private_access.kafka_connect)`boolean`Allow clients to connect to kafka\_connect with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | [**private\_access.kafka\_rest**](#private_access.kafka_rest)`boolean`Allow clients to connect to kafka\_rest with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | [**private\_access.prometheus**](#private_access.prometheus)`boolean`Allow clients to connect to prometheus with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | [**private\_access.schema\_registry**](#private_access.schema_registry)`boolean`Allow clients to connect to schema\_registry with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | [**public\_access**](#public_access)`object`Allow access to selected service ports from the public Internet | | [**public\_access.kafka**](#public_access.kafka)`boolean`Allow clients to connect to kafka from the public internet for service nodes that are in a project VPC or another type of private network | | [**public\_access.kafka\_connect**](#public_access.kafka_connect)`boolean`Allow clients to connect to kafka\_connect from the public internet for service nodes that are in a project VPC or another type of private network | | [**public\_access.kafka\_rest**](#public_access.kafka_rest)`boolean`Allow clients to connect to kafka\_rest from the public internet for service nodes that are in a project VPC or another type of private network | | [**public\_access.prometheus**](#public_access.prometheus)`boolean`Allow clients to connect to prometheus from the public internet for service nodes that are in a project VPC or another type of private network | | [**public\_access.schema\_registry**](#public_access.schema_registry)`boolean`Allow clients to connect to schema\_registry from the public internet for service nodes that are in a project VPC or another type of private network | | [**privatelink\_access**](#privatelink_access)`object`Allow access to selected service components through Privatelink | | [**privatelink\_access.jolokia**](#privatelink_access.jolokia)`boolean`Enable jolokia | | [**privatelink\_access.kafka**](#privatelink_access.kafka)`boolean`Enable kafka | | [**privatelink\_access.kafka\_connect**](#privatelink_access.kafka_connect)`boolean`Enable kafka\_connect | | [**privatelink\_access.kafka\_rest**](#privatelink_access.kafka_rest)`boolean`Enable kafka\_rest | | [**privatelink\_access.prometheus**](#privatelink_access.prometheus)`boolean`Enable prometheus | | [**privatelink\_access.schema\_registry**](#privatelink_access.schema_registry)`boolean`Enable schema\_registry | | [**letsencrypt\_sasl**](#letsencrypt_sasl)`boolean,null`Use a Let's Encrypt certificate authority (CA) for Kafka SASL authentication. (Default: False) | | [**letsencrypt\_sasl\_privatelink**](#letsencrypt_sasl_privatelink)`boolean,null`Use a Let's Encrypt certificate authority (CA) for Kafka SASL authentication via Privatelink. (Default: False) | | [**kafka**](#kafka)`object`- default: `[object Object]`Kafka broker configuration values | | [**kafka.group\_coordinator\_rebalance\_protocols**](#kafka.group_coordinator_rebalance_protocols)`string`- enum: `classic,classic,consumer,classic,streams,classic,consumer,streams`The enabled consumer group rebalance protocols. Use consumer, classic, share, streams to enable Kafka share groups. | | [**kafka.message\_max\_bytes**](#kafka.message_max_bytes)`integer`- min: `0`
- max: `100001200`The maximum size of message that the server can receive. (Default: 1048588 bytes (1 mebibyte + 12 bytes)) | | [**kafka.auto\_create\_topics\_enable**](#kafka.auto_create_topics_enable)`boolean`Enable auto-creation of topics. (Default: false) | | [**kafka\_authentication\_methods**](#kafka_authentication_methods)`object`Kafka authentication methods | | [**kafka\_authentication\_methods.certificate**](#kafka_authentication_methods.certificate)`boolean`- default: `true`Enable certificate/SSL authentication | | [**kafka\_authentication\_methods.sasl**](#kafka_authentication_methods.sasl)`boolean`Enable SASL authentication | | [**kafka\_sasl\_mechanisms**](#kafka_sasl_mechanisms)`object`Kafka SASL mechanisms | | [**kafka\_sasl\_mechanisms.plain**](#kafka_sasl_mechanisms.plain)`boolean`- default: `true`Enable PLAIN mechanism | | [**kafka\_sasl\_mechanisms.scram\_sha\_256**](#kafka_sasl_mechanisms.scram_sha_256)`boolean`- default: `true`Enable SCRAM-SHA-256 mechanism | | [**kafka\_sasl\_mechanisms.scram\_sha\_512**](#kafka_sasl_mechanisms.scram_sha_512)`boolean`- default: `true`Enable SCRAM-SHA-512 mechanism | | [**follower\_fetching**](#follower_fetching)`object`Enable follower fetching | | [**follower\_fetching.enabled**](#follower_fetching.enabled)`boolean`Whether to enable the follower fetching functionality | | [**kafka\_connect**](#kafka_connect)`boolean`Enable Kafka Connect service | | [**kafka\_connect\_config**](#kafka_connect_config)`object`Kafka Connect configuration values | | [**kafka\_connect\_config.prefer\_ipv6\_address\_enable**](#kafka_connect_config.prefer_ipv6_address_enable)`boolean`When enabled, connectors will automatically resolve IPv6 addresses from external server names configured with dual-stack. | | [**kafka\_connect\_config.connector\_client\_config\_override\_policy**](#kafka_connect_config.connector_client_config_override_policy)`string`- enum: `None,All`Defines what client configurations can be overridden by the connector. Default is None | | [**kafka\_connect\_config.consumer\_auto\_offset\_reset**](#kafka_connect_config.consumer_auto_offset_reset)`string`- enum: `earliest,latest`What to do when there is no initial offset in Kafka or if the current offset does not exist any more on the server. Default is earliest | | [**kafka\_connect\_config.consumer\_fetch\_max\_bytes**](#kafka_connect_config.consumer_fetch_max_bytes)`integer`- min: `1048576`
- max: `104857600`Records are fetched in batches by the consumer, and if the first record batch in the first non-empty partition of the fetch is larger than this value, the record batch will still be returned to ensure that the consumer can make progress. As such, this is not a absolute maximum. | | [**kafka\_connect\_config.consumer\_isolation\_level**](#kafka_connect_config.consumer_isolation_level)`string`- enum: `read_uncommitted,read_committed`Transaction read isolation level. read\_uncommitted is the default, but read\_committed can be used if consume-exactly-once behavior is desired. | | [**kafka\_connect\_config.consumer\_max\_partition\_fetch\_bytes**](#kafka_connect_config.consumer_max_partition_fetch_bytes)`integer`- min: `1048576`
- max: `104857600`Records are fetched in batches by the consumer.If the first record batch in the first non-empty partition of the fetch is larger than this limit, the batch will still be returned to ensure that the consumer can make progress. | | [**kafka\_connect\_config.consumer\_max\_poll\_interval\_ms**](#kafka_connect_config.consumer_max_poll_interval_ms)`integer`- min: `1`
- max: `2147483647`The maximum delay in milliseconds between invocations of poll() when using consumer group management (defaults to 300000). | | [**kafka\_connect\_config.consumer\_max\_poll\_records**](#kafka_connect_config.consumer_max_poll_records)`integer`- min: `1`
- max: `10000`The maximum number of records returned in a single call to poll() (defaults to 500). | | [**kafka\_connect\_config.offset\_flush\_interval\_ms**](#kafka_connect_config.offset_flush_interval_ms)`integer`- min: `1`
- max: `100000000`The interval at which to try committing offsets for tasks (defaults to 60000). | | [**kafka\_connect\_config.offset\_flush\_timeout\_ms**](#kafka_connect_config.offset_flush_timeout_ms)`integer`- min: `1`
- max: `2147483647`Maximum number of milliseconds to wait for records to flush and partition offset data to be committed to offset storage before cancelling the process and restoring the offset data to be committed in a future attempt (defaults to 5000). | | [**kafka\_connect\_config.producer\_batch\_size**](#kafka_connect_config.producer_batch_size)`integer`- min: `0`
- max: `5242880`This setting gives the upper bound of the batch size to be sent. If there are fewer than this many bytes accumulated for this partition, the producer will 'linger' for the linger.ms time waiting for more records to show up. A batch size of zero will disable batching entirely (defaults to 16384). | | [**kafka\_connect\_config.producer\_buffer\_memory**](#kafka_connect_config.producer_buffer_memory)`integer`- min: `5242880`
- max: `134217728`The total bytes of memory the producer can use to buffer records waiting to be sent to the broker (defaults to 33554432). | | [**kafka\_connect\_config.producer\_compression\_type**](#kafka_connect_config.producer_compression_type)`string`- enum: `gzip,snappy,lz4,zstd,none`Specify the default compression type for producers. This configuration accepts the standard compression codecs ('gzip', 'snappy', 'lz4', 'zstd'). It additionally accepts 'none' which is the default and equivalent to no compression. | | [**kafka\_connect\_config.producer\_linger\_ms**](#kafka_connect_config.producer_linger_ms)`integer`- min: `0`
- max: `5000`This setting gives the upper bound on the delay for batching: once there is batch.size worth of records for a partition it will be sent immediately regardless of this setting, however if there are fewer than this many bytes accumulated for this partition the producer will 'linger' for the specified time waiting for more records to show up. Defaults to 0. | | [**kafka\_connect\_config.producer\_max\_request\_size**](#kafka_connect_config.producer_max_request_size)`integer`- min: `131072`
- max: `67108864`This setting will limit the number of record batches the producer will send in a single request to avoid sending huge requests. | | [**kafka\_connect\_config.scheduled\_rebalance\_max\_delay\_ms**](#kafka_connect_config.scheduled_rebalance_max_delay_ms)`integer`- min: `0`
- max: `600000`The maximum delay that is scheduled in order to wait for the return of one or more departed workers before rebalancing and reassigning their connectors and tasks to the group. During this period the connectors and tasks of the departed workers remain unassigned. Defaults to 5 minutes. | | [**kafka\_connect\_config.session\_timeout\_ms**](#kafka_connect_config.session_timeout_ms)`integer`- min: `1`
- max: `2147483647`The timeout in milliseconds used to detect failures when using Kafka’s group management facilities (defaults to 10000). | | [**kafka\_connect\_plugin\_versions**](#kafka_connect_plugin_versions)`array[object]`The plugin selected by the user | | [**kafka\_connect\_plugin\_versions.\[0\].plugin\_name**](#kafka_connect_plugin_versions.\[0].plugin_name)`string`- maxLength: `128`The name of the plugin | | [**kafka\_connect\_plugin\_versions.\[0\].version**](#kafka_connect_plugin_versions.\[0].version)`string`- maxLength: `128`The version of the plugin | | [**kafka\_connect\_secret\_providers**](#kafka_connect_secret_providers)`array[object]`Configure external secret providers in order to reference external secrets in connector configuration. Currently Hashicorp Vault (provider: vault, auth\_method: token) and AWS Secrets Manager (provider: aws, auth\_method: credentials) are supported. Secrets can be referenced in connector config with ${\<provider\_name\>:\<secret\_path\>:\<key\_name\>} | | [**kafka\_connect\_secret\_providers.\[0\].name**](#kafka_connect_secret_providers.\[0].name)`string`Name of the secret provider. Used to reference secrets in connector config. | | [**kafka\_connect\_secret\_providers.\[0\].vault**](#kafka_connect_secret_providers.\[0].vault)`object`- required: `auth_method,address`Vault secret provider configuration | | [**kafka\_connect\_secret\_providers.\[0\].vault.auth\_method**](#kafka_connect_secret_providers.\[0].vault.auth_method)`string`- enum: `token`An enumeration. | | [**kafka\_connect\_secret\_providers.\[0\].vault.address**](#kafka_connect_secret_providers.\[0].vault.address)`string`- maxLength: `65536`Address of the Vault server | | [**kafka\_connect\_secret\_providers.\[0\].vault.engine\_version**](#kafka_connect_secret_providers.\[0].vault.engine_version)`integer`- enum: `1,2`An enumeration. | | [**kafka\_connect\_secret\_providers.\[0\].vault.prefix\_path\_depth**](#kafka_connect_secret_providers.\[0].vault.prefix_path_depth)`integer`Prefix path depth of the secrets Engine. Default is 1. If the secrets engine path has more than one segment it has to be increased to the number of segments. | | [**kafka\_connect\_secret\_providers.\[0\].vault.token**](#kafka_connect_secret_providers.\[0].vault.token)`string`- maxLength: `256`Token used to authenticate with vault and auth method \`token\`. | | [**kafka\_connect\_secret\_providers.\[0\].vault.server\_pem**](#kafka_connect_secret_providers.\[0].vault.server_pem)`string`- maxLength: `4096`PEM encoded certificate of the Vault server. Required if the vault server uses a self-signed certificate. | | [**kafka\_connect\_secret\_providers.\[0\].aws**](#kafka_connect_secret_providers.\[0].aws)`object`- required: `auth_method,region`AWS secret provider configuration | | [**kafka\_connect\_secret\_providers.\[0\].aws.auth\_method**](#kafka_connect_secret_providers.\[0].aws.auth_method)`string`- enum: `credentials`An enumeration. | | [**kafka\_connect\_secret\_providers.\[0\].aws.access\_key**](#kafka_connect_secret_providers.\[0].aws.access_key)`string`- maxLength: `128`Access key used to authenticate with aws | | [**kafka\_connect\_secret\_providers.\[0\].aws.secret\_key**](#kafka_connect_secret_providers.\[0].aws.secret_key)`string`- maxLength: `128`Secret key used to authenticate with aws | | [**kafka\_connect\_secret\_providers.\[0\].aws.region**](#kafka_connect_secret_providers.\[0].aws.region)`string`- maxLength: `64`Region used to lookup secrets with AWS SecretManager | | [**kafka\_connect\_secret\_providers.\[0\].env**](#kafka_connect_secret_providers.\[0].env)`object`- required: `secrets`ENV secret provider configuration | | [**kafka\_connect\_secret\_providers.\[0\].env.secrets**](#kafka_connect_secret_providers.\[0].env.secrets)`object`Key/value map of secrets for ENV secret provider | | [**kafka\_connect\_secret\_providers.\[0\].azure**](#kafka_connect_secret_providers.\[0].azure)`object`- required: `auth_method`Azure KeyVault secret provider configuration | | [**kafka\_connect\_secret\_providers.\[0\].azure.auth\_method**](#kafka_connect_secret_providers.\[0].azure.auth_method)`string`- enum: `credentials`An enumeration. | | [**kafka\_connect\_secret\_providers.\[0\].azure.client\_id**](#kafka_connect_secret_providers.\[0].azure.client_id)`string`- maxLength: `128`Azure client ID for the service principal. | | [**kafka\_connect\_secret\_providers.\[0\].azure.tenant\_id**](#kafka_connect_secret_providers.\[0].azure.tenant_id)`string`- maxLength: `128`Azure tenant ID for the service principal. | | [**kafka\_connect\_secret\_providers.\[0\].azure.secret**](#kafka_connect_secret_providers.\[0].azure.secret)`string`- maxLength: `256`Azure client secret for the service principal. | | [**kafka\_diskless**](#kafka_diskless)`object`- required: `enabled`Kafka Diskless configuration values | | [**kafka\_diskless.enabled**](#kafka_diskless.enabled)`boolean`Whether to enable the Diskless functionality | | [**kafka\_diskless.auto\_diskless\_topic\_regexes**](#kafka_diskless.auto_diskless_topic_regexes)`array[string]`- maxItems: `32`The regexes of topics to auto enable diskless. Topics matching any of the regexes will be created as diskless topics. | | [**inkless**](#inkless)`object`- required: `enabled`Inkless configuration values | | [**inkless.enabled**](#inkless.enabled)`boolean`Whether to enable the Inkless functionality | | [**sasl\_oauthbearer\_allowed\_urls**](#sasl_oauthbearer_allowed_urls)`array[string]`List of allowed URLs for SASL OAUTHBEARER authentication. Only HTTPS URLs are allowed for security reasons. | | [**gcp\_auth\_allowed\_urls**](#gcp_auth_allowed_urls)`array[string]`Allow-list of HTTPS URLs used to validate GCP credential\_source requests for Kafka Connect. | | [**kafka\_rest**](#kafka_rest)`boolean`Enable Kafka-REST service | | [**kafka\_version**](#kafka_version)`string,null`- enum: `3.9,4.1,4.2`Kafka major version | | [**karapace\_version**](#karapace_version)`string,null`- enum: `6.2.1,6.2.2,null`Select a Karapace version for this service, or select Latest to use the latest available version automatically. New versions become available after installation during a maintenance update. | | [**schema\_registry**](#schema_registry)`boolean`Enable Schema-Registry service | | [**kafka\_rest\_authorization**](#kafka_rest_authorization)`boolean`Enable authorization in Kafka-REST service | | [**kafka\_rest\_config**](#kafka_rest_config)`object`Kafka REST configuration | | [**kafka\_rest\_config.producer\_acks**](#kafka_rest_config.producer_acks)`string`- enum: `all,-1,0,1`
- default: `1`The number of acknowledgments the producer requires the leader to have received before considering a request complete. If set to 'all' or '-1', the leader will wait for the full set of in-sync replicas to acknowledge the record. | | [**kafka\_rest\_config.producer\_compression\_type**](#kafka_rest_config.producer_compression_type)`string`- enum: `gzip,snappy,lz4,zstd,none`Specify the default compression type for producers. This configuration accepts the standard compression codecs ('gzip', 'snappy', 'lz4', 'zstd'). It additionally accepts 'none' which is the default and equivalent to no compression. | | [**kafka\_rest\_config.producer\_linger\_ms**](#kafka_rest_config.producer_linger_ms)`integer`- min: `0`
- max: `5000`
- default: `0`Wait for up to the given delay to allow batching records together | | [**kafka\_rest\_config.producer\_max\_request\_size**](#kafka_rest_config.producer_max_request_size)`integer`- min: `0`
- max: `2147483647`
- default: `1048576`The maximum size of a request in bytes. Note that Kafka broker can also cap the record batch size. | | [**kafka\_rest\_config.consumer\_enable\_auto\_commit**](#kafka_rest_config.consumer_enable_auto_commit)`boolean`- default: `true`If true the consumer's offset will be periodically committed to Kafka in the background | | [**kafka\_rest\_config.consumer\_idle\_disconnect\_timeout**](#kafka_rest_config.consumer_idle_disconnect_timeout)`integer`- min: `0`
- max: `2147483647`
- default: `0`Specifies the maximum duration (in seconds) a client can remain idle before it is deleted. If a consumer is inactive, it will exit the consumer group, and its state will be discarded. A value of 0 (default) indicates that the consumer will not be disconnected automatically due to inactivity. | | [**kafka\_rest\_config.consumer\_request\_max\_bytes**](#kafka_rest_config.consumer_request_max_bytes)`integer`- min: `0`
- max: `671088640`
- default: `67108864`Maximum number of bytes in unencoded message keys and values by a single request | | [**kafka\_rest\_config.consumer\_request\_timeout\_ms**](#kafka_rest_config.consumer_request_timeout_ms)`integer`- min: `1000`
- max: `30000`
- enum: `1000,15000,30000`
- default: `1000`The maximum total time to wait for messages for a request if the maximum number of messages has not yet been reached | | [**kafka\_rest\_config.name\_strategy**](#kafka_rest_config.name_strategy)`string`- enum: `topic_name,record_name,topic_record_name`
- default: `topic_name`Name strategy to use when selecting subject for storing schemas | | [**kafka\_rest\_config.name\_strategy\_validation**](#kafka_rest_config.name_strategy_validation)`boolean`- default: `true`If true, validate that given schema is registered under expected subject name by the used name strategy when producing messages. | | [**kafka\_rest\_config.simpleconsumer\_pool\_size\_max**](#kafka_rest_config.simpleconsumer_pool_size_max)`integer`- min: `10`
- max: `250`
- default: `25`Maximum number of SimpleConsumers that can be instantiated per broker | | [**tiered\_storage**](#tiered_storage)`object`Tiered storage configuration | | [**tiered\_storage.enabled**](#tiered_storage.enabled)`boolean`Whether to enable the tiered storage functionality | | [**schema\_registry\_config**](#schema_registry_config)`object`Schema Registry configuration | | [**schema\_registry\_config.topic\_name**](#schema_registry_config.topic_name)`string`- maxLength: `249`The durable single partition topic that acts as the durable log for the data. This topic must be compacted to avoid losing data due to retention policy. Please note that changing this configuration in an existing Schema Registry / Karapace setup leads to previous schemas being inaccessible, data encoded with them potentially unreadable and schema ID sequence put out of order. It's only possible to do the switch while Schema Registry / Karapace is disabled. Defaults to \`\_schemas\`. | | [**schema\_registry\_config.leader\_eligibility**](#schema_registry_config.leader_eligibility)`boolean`If true, Karapace / Schema Registry on the service nodes can participate in leader election. It might be needed to disable this when the schemas topic is replicated to a secondary cluster and Karapace / Schema Registry there must not participate in leader election. Defaults to \`true\`. | | [**schema\_registry\_config.schema\_reader\_strict\_mode**](#schema_registry_config.schema_reader_strict_mode)`boolean`If enabled, causes the Karapace schema-registry service to shutdown when there are invalid schema records in the \`\_schemas\` topic. Defaults to \`false\`. | | [**schema\_registry\_config.retriable\_errors\_silenced**](#schema_registry_config.retriable_errors_silenced)`boolean`If enabled, kafka errors which can be retried or custom errors specified for the service will not be raised, instead, a warning log is emitted. This will denoise issue tracking systems, i.e. sentry. Defaults to \`true\`. | | [**schema\_registry\_config.sasl\_oauthbearer\_authentication\_enabled**](#schema_registry_config.sasl_oauthbearer_authentication_enabled)`boolean`If enabled, the Schema Registry validates OAuth2/OIDC JWT bearer tokens on incoming requests. Requires the OIDC provider settings under the \`kafka\` configuration (\`sasl\_oauthbearer\_jwks\_endpoint\_url\` and related). Defaults to \`false\`. | | [**schema\_registry\_config.sasl\_oauthbearer\_authorization\_enabled**](#schema_registry_config.sasl_oauthbearer_authorization_enabled)`boolean`If enabled, the Schema Registry enforces role-based authorization derived from the JWT roles claim. Enabling this automatically enables \`sasl\_oauthbearer\_authentication\_enabled\` when it is not already enabled, since authorization requires authentication. Defaults to \`false\`. | | [**schema\_registry\_config.sasl\_oauthbearer\_roles\_claim\_path**](#schema_registry_config.sasl_oauthbearer_roles_claim_path)`string`- maxLength: `128`JSON path used to extract the roles claim from the JWT for Schema Registry authorization. Defaults to \`resource\_access.karapace.roles\`. | | [**schema\_registry\_config.sasl\_oauthbearer\_method\_roles**](#schema_registry_config.sasl_oauthbearer_method_roles)`string`- maxLength: `4096`JSON object mapping HTTP methods to the list of roles allowed to perform them on the Schema Registry, provided as a JSON-encoded string. Role names use the \`karapace.\` prefix, e.g. \`karapace.schema:read\`. Defaults to \`{"GET": \["karapace.schema:read", "karapace.subject:read"], "POST": \[], "PUT": \[], "DELETE": \[]}\`. | | [**aiven\_kafka\_topic\_messages**](#aiven_kafka_topic_messages)`boolean`Allow access to read Kafka topic messages in the Aiven Console and REST API. | | [**enable\_ipv6**](#enable_ipv6)`boolean`Register AAAA DNS records for the service, and allow IPv6 packets to service ports | --- # Advanced parameters for Aiven for Apache Kafka® Free tier View the configuration options for Aiven for Apache Kafka® Free tier. ## Broker parameters[​](#broker-parameters "Direct link to Broker parameters") | | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | [**backup\_interval\_hours**](#backup_interval_hours)`integer,null`- min: `3`
- max: `24`
- enum: `3,4,6,8,12,24,null`Interval in hours between automatic backups. Minimum value is 3 hours. Must be a divisor of 24 (3, 4, 6, 8, 12, 24). (Applicable to ACU plans only) | | [**backup\_retention\_days**](#backup_retention_days)`integer,null`- min: `1`
- max: `30`Number of days to retain automatic backups. Backups older than this value will be automatically deleted. (Applicable to ACU plans only) | | [**custom\_domain**](#custom_domain)`string,null`- maxLength: `255`Serve the web frontend using a custom CNAME pointing to the Aiven DNS name. When you set a custom domain for a service deployed in a VPC, the service certificate is only created for the public-\* hostname and the custom domain. | | [**ip\_filter**](#ip_filter)`array[string,object]`- maxItems: `8000`
- default: `0.0.0.0/0,::/0`Allow incoming connections from CIDR address block, e.g. '10.20.0.0/16' | | [**ip\_filter.\[0\].description**](#ip_filter.\[0].description)`string`- maxLength: `1024`Description for IP filter list entry | | [**ip\_filter.\[0\].network**](#ip_filter.\[0].network)`string`- maxLength: `43`CIDR address block | | [**service\_log**](#service_log)`boolean,null`Store logs for the service so that they are available in the HTTP API and console. | | [**static\_ips**](#static_ips)`boolean`Use static public IP addresses | | [**single\_zone**](#single_zone)`object`Single-zone configuration | | [**single\_zone.enabled**](#single_zone.enabled)`boolean`Whether to allocate nodes on the same Availability Zone or spread across zones available. By default service nodes are spread across different AZs. The single AZ support is best-effort and may temporarily allocate nodes in different AZs e.g. in case of capacity limitations in one AZ. | | [**single\_zone.availability\_zone**](#single_zone.availability_zone)`string`- maxLength: `40`The availability zone to use for the service. This is only used when enabled is set to true. If not set the service will be allocated in random AZ.The AZ is not guaranteed, and the service may be allocated in a different AZ if the selected AZ is not available. Zones will not be validated and invalid zones will be ignored, falling back to random AZ selection. Common availability zones include: AWS (euc1-az1, euc1-az2, euc1-az3), GCP (europe-west1-a, europe-west1-b, europe-west1-c), Azure (germanywestcentral/1, germanywestcentral/2, germanywestcentral/3). | | [**preferred\_zones**](#preferred_zones)`array[string]`- maxItems: `10`List of preferred zone IDs for service node placement. Nodes will be placed in these zones when available. If a specified zone is unavailable (e.g., due to capacity constraints), nodes will be placed in other available zones to maintain the configured number of zones for availability. Invalid zone IDs are rejected at configuration time. Zone IDs are cloud-specific: AWS uses zone IDs like 'euc1-az1', GCP uses zone names like 'europe-west1-a', and Azure uses 'location/zone' format like 'germanywestcentral/1'. If single\_zone is enabled with an availability\_zone, that setting takes precedence over preferred\_zones. Changes take effect on next node recreation (e.g., maintenance or plan change). For eligible plans, nodes outside preferred zones are automatically rebalanced once per day. | | [**private\_access**](#private_access)`object`Allow access to selected service ports from private networks | | [**private\_access.kafka**](#private_access.kafka)`boolean`Allow clients to connect to kafka with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | [**private\_access.kafka\_connect**](#private_access.kafka_connect)`boolean`Allow clients to connect to kafka\_connect with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | [**private\_access.kafka\_rest**](#private_access.kafka_rest)`boolean`Allow clients to connect to kafka\_rest with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | [**private\_access.prometheus**](#private_access.prometheus)`boolean`Allow clients to connect to prometheus with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | [**private\_access.schema\_registry**](#private_access.schema_registry)`boolean`Allow clients to connect to schema\_registry with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | [**public\_access**](#public_access)`object`Allow access to selected service ports from the public Internet | | [**public\_access.kafka**](#public_access.kafka)`boolean`Allow clients to connect to kafka from the public internet for service nodes that are in a project VPC or another type of private network | | [**public\_access.kafka\_connect**](#public_access.kafka_connect)`boolean`Allow clients to connect to kafka\_connect from the public internet for service nodes that are in a project VPC or another type of private network | | [**public\_access.kafka\_rest**](#public_access.kafka_rest)`boolean`Allow clients to connect to kafka\_rest from the public internet for service nodes that are in a project VPC or another type of private network | | [**public\_access.prometheus**](#public_access.prometheus)`boolean`Allow clients to connect to prometheus from the public internet for service nodes that are in a project VPC or another type of private network | | [**public\_access.schema\_registry**](#public_access.schema_registry)`boolean`Allow clients to connect to schema\_registry from the public internet for service nodes that are in a project VPC or another type of private network | | [**privatelink\_access**](#privatelink_access)`object`Allow access to selected service components through Privatelink | | [**privatelink\_access.jolokia**](#privatelink_access.jolokia)`boolean`Enable jolokia | | [**privatelink\_access.kafka**](#privatelink_access.kafka)`boolean`Enable kafka | | [**privatelink\_access.kafka\_connect**](#privatelink_access.kafka_connect)`boolean`Enable kafka\_connect | | [**privatelink\_access.kafka\_rest**](#privatelink_access.kafka_rest)`boolean`Enable kafka\_rest | | [**privatelink\_access.prometheus**](#privatelink_access.prometheus)`boolean`Enable prometheus | | [**privatelink\_access.schema\_registry**](#privatelink_access.schema_registry)`boolean`Enable schema\_registry | | [**letsencrypt\_sasl**](#letsencrypt_sasl)`boolean,null`Use a Let's Encrypt certificate authority (CA) for Kafka SASL authentication. (Default: False) | | [**letsencrypt\_sasl\_privatelink**](#letsencrypt_sasl_privatelink)`boolean,null`Use a Let's Encrypt certificate authority (CA) for Kafka SASL authentication via Privatelink. (Default: False) | | [**kafka**](#kafka)`object`- default: `[object Object]`Kafka broker configuration values | | [**kafka.compression\_type**](#kafka.compression_type)`string`- enum: `gzip,snappy,lz4,zstd,uncompressed,producer`Specify the final compression type for a given topic. This configuration accepts the standard compression codecs ('gzip', 'snappy', 'lz4', 'zstd'). It additionally accepts 'uncompressed' which is equivalent to no compression; and 'producer' which means retain the original compression codec set by the producer.(Default: producer) | | [**kafka.group\_initial\_rebalance\_delay\_ms**](#kafka.group_initial_rebalance_delay_ms)`integer`- min: `0`
- max: `300000`The amount of time, in milliseconds, the group coordinator will wait for more consumers to join a new group before performing the first rebalance. A longer delay means potentially fewer rebalances, but increases the time until processing begins. The default value for this is 3 seconds. During development and testing it might be desirable to set this to 0 in order to not delay test execution time. (Default: 3000 ms (3 seconds)) | | [**kafka.group\_coordinator\_rebalance\_protocols**](#kafka.group_coordinator_rebalance_protocols)`string`- enum: `classic,classic,consumer,classic,streams,classic,consumer,streams`The enabled consumer group rebalance protocols. Use consumer, classic, share, streams to enable Kafka share groups. | | [**kafka.group\_min\_session\_timeout\_ms**](#kafka.group_min_session_timeout_ms)`integer`- min: `0`
- max: `60000`The minimum allowed session timeout for registered consumers. Longer timeouts give consumers more time to process messages in between heartbeats at the cost of a longer time to detect failures. (Default: 6000 ms (6 seconds)) | | [**kafka.group\_max\_session\_timeout\_ms**](#kafka.group_max_session_timeout_ms)`integer`- min: `0`
- max: `1800000`The maximum allowed session timeout for registered consumers. Longer timeouts give consumers more time to process messages in between heartbeats at the cost of a longer time to detect failures. Default: 1800000 ms (30 minutes) | | [**kafka.connections\_max\_idle\_ms**](#kafka.connections_max_idle_ms)`integer`- min: `1000`
- max: `3600000`Idle connections timeout: the server socket processor threads close the connections that idle for longer than this. (Default: 600000 ms (10 minutes)) | | [**kafka.max\_incremental\_fetch\_session\_cache\_slots**](#kafka.max_incremental_fetch_session_cache_slots)`integer`- min: `1000`
- max: `10000`The maximum number of incremental fetch sessions that the broker will maintain. (Default: 1000) | | [**kafka.message\_max\_bytes**](#kafka.message_max_bytes)`integer`- min: `0`
- max: `100001200`The maximum size of message that the server can receive. (Default: 1048588 bytes (1 mebibyte + 12 bytes)) | | [**kafka.offsets\_retention\_minutes**](#kafka.offsets_retention_minutes)`integer`- min: `1`
- max: `2147483647`Log retention window in minutes for offsets topic (Default: 10080 minutes (7 days)) | | [**kafka.log\_cleaner\_delete\_retention\_ms**](#kafka.log_cleaner_delete_retention_ms)`integer`- min: `0`
- max: `315569260000`How long are delete records retained? (Default: 86400000 (1 day)) | | [**kafka.log\_cleaner\_max\_compaction\_lag\_ms**](#kafka.log_cleaner_max_compaction_lag_ms)`integer`- min: `30000`
- max: `9223372036854775807`The maximum amount of time message will remain uncompacted. Only applicable for logs that are being compacted. (Default: 9223372036854775807 ms (Long.MAX\_VALUE)) | | [**kafka.log\_cleaner\_min\_compaction\_lag\_ms**](#kafka.log_cleaner_min_compaction_lag_ms)`integer`- min: `0`
- max: `9223372036854775807`The minimum time a message will remain uncompacted in the log. Only applicable for logs that are being compacted. (Default: 0 ms) | | [**kafka.log\_cleanup\_policy**](#kafka.log_cleanup_policy)`string`- enum: `delete,compact,compact,delete,delete,compact`The default cleanup policy for segments beyond the retention window (Default: delete) | | [**kafka.log\_flush\_interval\_messages**](#kafka.log_flush_interval_messages)`integer`- min: `1`
- max: `9223372036854775807`The number of messages accumulated on a log partition before messages are flushed to disk (Default: 9223372036854775807 (Long.MAX\_VALUE)) | | [**kafka.log\_flush\_interval\_ms**](#kafka.log_flush_interval_ms)`integer`- min: `0`
- max: `9223372036854775807`The maximum time in ms that a message in any topic is kept in memory (page-cache) before flushed to disk. If not set, the value in log.flush.scheduler.interval.ms is used (Default: null) | | [**kafka.log\_index\_interval\_bytes**](#kafka.log_index_interval_bytes)`integer`- min: `0`
- max: `104857600`The interval with which Kafka adds an entry to the offset index (Default: 4096 bytes (4 kibibytes)) | | [**kafka.log\_index\_size\_max\_bytes**](#kafka.log_index_size_max_bytes)`integer`- min: `1048576`
- max: `104857600`The maximum size in bytes of the offset index (Default: 10485760 (10 mebibytes)) | | [**kafka.log\_local\_retention\_ms**](#kafka.log_local_retention_ms)`integer`- min: `-2`
- max: `9223372036854775807`The number of milliseconds to keep the local log segments before it gets eligible for deletion. If set to -2, the value of log.retention.ms is used. The effective value should always be less than or equal to log.retention.ms value. (Default: -2) | | [**kafka.log\_local\_retention\_bytes**](#kafka.log_local_retention_bytes)`integer`- min: `-2`
- max: `9223372036854775807`The maximum size of local log segments that can grow for a partition before it gets eligible for deletion. If set to -2, the value of log.retention.bytes is used. The effective value should always be less than or equal to log.retention.bytes value. (Default: -2) | | [**kafka.log\_message\_timestamp\_type**](#kafka.log_message_timestamp_type)`string`- enum: `CreateTime,LogAppendTime`Define whether the timestamp in the message is message create time or log append time. (Default: CreateTime) | | [**kafka.log\_message\_timestamp\_before\_max\_ms**](#kafka.log_message_timestamp_before_max_ms)`integer`- min: `0`
- max: `9223372036854775807`The maximum difference allowed between the timestamp when a broker receives a message and the timestamp specified in the message. If message.timestamp.type=CreateTime, a message will be rejected if the difference in timestamp exceeds this threshold. Applies only for messages with timestamps earlier than the broker's timestamp. (Default: 9223372036854775807 (Long.MAX\_VALUE)) | | [**kafka.log\_message\_timestamp\_after\_max\_ms**](#kafka.log_message_timestamp_after_max_ms)`integer`- min: `0`
- max: `9223372036854775807`The maximum difference allowed between the timestamp when a broker receives a message and the timestamp specified in the message. If message.timestamp.type=CreateTime, a message will be rejected if the difference in timestamp exceeds this threshold. Applies only for messages with timestamps later than the broker's timestamp. (Default: 9223372036854775807 (Long.MAX\_VALUE)) | | [**kafka.log\_preallocate**](#kafka.log_preallocate)`boolean`Should pre allocate file when create new segment? (Default: false) | | [**kafka.log\_retention\_bytes**](#kafka.log_retention_bytes)`integer`- min: `-1`
- max: `9223372036854775807`The maximum size of the log before deleting messages (Default: -1) | | [**kafka.log\_retention\_hours**](#kafka.log_retention_hours)`integer`- min: `0`
- max: `72`The number of hours to keep a log file before deleting it. Use -1 for unlimited retention or 1 or higher. Setting 0 is invalid and prevents Kafka from starting. (Default: 168 hours, or 1 week) | | [**kafka.log\_retention\_ms**](#kafka.log_retention_ms)`integer`- min: `0`
- max: `259200000`The number of milliseconds to keep a log file before deleting it (in milliseconds), If not set, the value in log.retention.minutes is used. If set to -1, no time limit is applied. (Default: null, log.retention.hours applies) | | [**kafka.log\_roll\_jitter\_ms**](#kafka.log_roll_jitter_ms)`integer`- min: `0`
- max: `9223372036854775807`The maximum jitter to subtract from logRollTimeMillis (in milliseconds). If not set, the value in log.roll.jitter.hours is used (Default: null) | | [**kafka.log\_roll\_ms**](#kafka.log_roll_ms)`integer`- min: `1`
- max: `9223372036854775807`The maximum time before a new log segment is rolled out (in milliseconds). (Default: null, log.roll.hours applies (Default: 168, 7 days)) | | [**kafka.log\_segment\_bytes**](#kafka.log_segment_bytes)`integer`- min: `10485760`
- max: `1073741824`The maximum size of a single log file (Default: 1073741824 bytes (1 gibibyte)) | | [**kafka.log\_segment\_delete\_delay\_ms**](#kafka.log_segment_delete_delay_ms)`integer`- min: `0`
- max: `3600000`The amount of time to wait before deleting a file from the filesystem (Default: 60000 ms (1 minute)) | | [**kafka.auto\_create\_topics\_enable**](#kafka.auto_create_topics_enable)`boolean`Enable auto-creation of topics. (Default: false) | | [**kafka.min\_insync\_replicas**](#kafka.min_insync_replicas)`integer`- min: `1`
- max: `7`When a producer sets acks to 'all' (or '-1'), min.insync.replicas specifies the minimum number of replicas that must acknowledge a write for the write to be considered successful. (Default: 1) | | [**kafka.num\_partitions**](#kafka.num_partitions)`integer`- min: `1`
- max: `2`Number of partitions for auto-created topics (Default: 1) | | [**kafka.default\_replication\_factor**](#kafka.default_replication_factor)`integer`- min: `1`
- max: `10`Replication factor for auto-created topics (Default: 3) | | [**kafka.replica\_fetch\_max\_bytes**](#kafka.replica_fetch_max_bytes)`integer`- min: `1048576`
- max: `104857600`The number of bytes of messages to attempt to fetch for each partition . This is not an absolute maximum, if the first record batch in the first non-empty partition of the fetch is larger than this value, the record batch will still be returned to ensure that progress can be made. (Default: 1048576 bytes (1 mebibytes)) | | [**kafka.replica\_fetch\_response\_max\_bytes**](#kafka.replica_fetch_response_max_bytes)`integer`- min: `10485760`
- max: `1048576000`Maximum bytes expected for the entire fetch response. Records are fetched in batches, and if the first record batch in the first non-empty partition of the fetch is larger than this value, the record batch will still be returned to ensure that progress can be made. As such, this is not an absolute maximum. (Default: 10485760 bytes (10 mebibytes)) | | [**kafka.max\_connections\_per\_ip**](#kafka.max_connections_per_ip)`integer`- min: `256`
- max: `2147483647`The maximum number of connections allowed from each ip address (Default: 2147483647). | | [**kafka.producer\_purgatory\_purge\_interval\_requests**](#kafka.producer_purgatory_purge_interval_requests)`integer`- min: `10`
- max: `10000`The purge interval (in number of requests) of the producer request purgatory (Default: 1000). | | [**kafka.socket\_request\_max\_bytes**](#kafka.socket_request_max_bytes)`integer`- min: `10485760`
- max: `209715200`The maximum number of bytes in a socket request (Default: 104857600 bytes). | | [**kafka.transaction\_state\_log\_segment\_bytes**](#kafka.transaction_state_log_segment_bytes)`integer`- min: `1048576`
- max: `2147483647`The transaction topic segment bytes should be kept relatively small in order to facilitate faster log compaction and cache loads (Default: 104857600 bytes (100 mebibytes)). | | [**kafka.transaction\_remove\_expired\_transaction\_cleanup\_interval\_ms**](#kafka.transaction_remove_expired_transaction_cleanup_interval_ms)`integer`- min: `600000`
- max: `3600000`The interval at which to remove transactions that have expired due to transactional.id.expiration.ms passing (Default: 3600000 ms (1 hour)). | | [**kafka.transaction\_partition\_verification\_enable**](#kafka.transaction_partition_verification_enable)`boolean`Enable verification that checks that the partition has been added to the transaction before writing transactional records to the partition. (Default: true) | | [**kafka\_authentication\_methods**](#kafka_authentication_methods)`object`Kafka authentication methods | | [**kafka\_authentication\_methods.certificate**](#kafka_authentication_methods.certificate)`boolean`- default: `true`Enable certificate/SSL authentication | | [**kafka\_authentication\_methods.sasl**](#kafka_authentication_methods.sasl)`boolean`Enable SASL authentication | | [**kafka\_sasl\_mechanisms**](#kafka_sasl_mechanisms)`object`Kafka SASL mechanisms | | [**kafka\_sasl\_mechanisms.plain**](#kafka_sasl_mechanisms.plain)`boolean`- default: `true`Enable PLAIN mechanism | | [**kafka\_sasl\_mechanisms.scram\_sha\_256**](#kafka_sasl_mechanisms.scram_sha_256)`boolean`- default: `true`Enable SCRAM-SHA-256 mechanism | | [**kafka\_sasl\_mechanisms.scram\_sha\_512**](#kafka_sasl_mechanisms.scram_sha_512)`boolean`- default: `true`Enable SCRAM-SHA-512 mechanism | | [**follower\_fetching**](#follower_fetching)`object`Enable follower fetching | | [**follower\_fetching.enabled**](#follower_fetching.enabled)`boolean`Whether to enable the follower fetching functionality | | [**kafka\_connect**](#kafka_connect)`boolean`Enable Kafka Connect service | | [**kafka\_connect\_config**](#kafka_connect_config)`object`Kafka Connect configuration values | | [**kafka\_connect\_config.prefer\_ipv6\_address\_enable**](#kafka_connect_config.prefer_ipv6_address_enable)`boolean`When enabled, connectors will automatically resolve IPv6 addresses from external server names configured with dual-stack. | | [**kafka\_connect\_config.connector\_client\_config\_override\_policy**](#kafka_connect_config.connector_client_config_override_policy)`string`- enum: `None,All`Defines what client configurations can be overridden by the connector. Default is None | | [**kafka\_connect\_config.consumer\_auto\_offset\_reset**](#kafka_connect_config.consumer_auto_offset_reset)`string`- enum: `earliest,latest`What to do when there is no initial offset in Kafka or if the current offset does not exist any more on the server. Default is earliest | | [**kafka\_connect\_config.consumer\_fetch\_max\_bytes**](#kafka_connect_config.consumer_fetch_max_bytes)`integer`- min: `1048576`
- max: `104857600`Records are fetched in batches by the consumer, and if the first record batch in the first non-empty partition of the fetch is larger than this value, the record batch will still be returned to ensure that the consumer can make progress. As such, this is not a absolute maximum. | | [**kafka\_connect\_config.consumer\_isolation\_level**](#kafka_connect_config.consumer_isolation_level)`string`- enum: `read_uncommitted,read_committed`Transaction read isolation level. read\_uncommitted is the default, but read\_committed can be used if consume-exactly-once behavior is desired. | | [**kafka\_connect\_config.consumer\_max\_partition\_fetch\_bytes**](#kafka_connect_config.consumer_max_partition_fetch_bytes)`integer`- min: `1048576`
- max: `104857600`Records are fetched in batches by the consumer.If the first record batch in the first non-empty partition of the fetch is larger than this limit, the batch will still be returned to ensure that the consumer can make progress. | | [**kafka\_connect\_config.consumer\_max\_poll\_interval\_ms**](#kafka_connect_config.consumer_max_poll_interval_ms)`integer`- min: `1`
- max: `2147483647`The maximum delay in milliseconds between invocations of poll() when using consumer group management (defaults to 300000). | | [**kafka\_connect\_config.consumer\_max\_poll\_records**](#kafka_connect_config.consumer_max_poll_records)`integer`- min: `1`
- max: `10000`The maximum number of records returned in a single call to poll() (defaults to 500). | | [**kafka\_connect\_config.offset\_flush\_interval\_ms**](#kafka_connect_config.offset_flush_interval_ms)`integer`- min: `1`
- max: `100000000`The interval at which to try committing offsets for tasks (defaults to 60000). | | [**kafka\_connect\_config.offset\_flush\_timeout\_ms**](#kafka_connect_config.offset_flush_timeout_ms)`integer`- min: `1`
- max: `2147483647`Maximum number of milliseconds to wait for records to flush and partition offset data to be committed to offset storage before cancelling the process and restoring the offset data to be committed in a future attempt (defaults to 5000). | | [**kafka\_connect\_config.producer\_batch\_size**](#kafka_connect_config.producer_batch_size)`integer`- min: `0`
- max: `5242880`This setting gives the upper bound of the batch size to be sent. If there are fewer than this many bytes accumulated for this partition, the producer will 'linger' for the linger.ms time waiting for more records to show up. A batch size of zero will disable batching entirely (defaults to 16384). | | [**kafka\_connect\_config.producer\_buffer\_memory**](#kafka_connect_config.producer_buffer_memory)`integer`- min: `5242880`
- max: `134217728`The total bytes of memory the producer can use to buffer records waiting to be sent to the broker (defaults to 33554432). | | [**kafka\_connect\_config.producer\_compression\_type**](#kafka_connect_config.producer_compression_type)`string`- enum: `gzip,snappy,lz4,zstd,none`Specify the default compression type for producers. This configuration accepts the standard compression codecs ('gzip', 'snappy', 'lz4', 'zstd'). It additionally accepts 'none' which is the default and equivalent to no compression. | | [**kafka\_connect\_config.producer\_linger\_ms**](#kafka_connect_config.producer_linger_ms)`integer`- min: `0`
- max: `5000`This setting gives the upper bound on the delay for batching: once there is batch.size worth of records for a partition it will be sent immediately regardless of this setting, however if there are fewer than this many bytes accumulated for this partition the producer will 'linger' for the specified time waiting for more records to show up. Defaults to 0. | | [**kafka\_connect\_config.producer\_max\_request\_size**](#kafka_connect_config.producer_max_request_size)`integer`- min: `131072`
- max: `67108864`This setting will limit the number of record batches the producer will send in a single request to avoid sending huge requests. | | [**kafka\_connect\_config.scheduled\_rebalance\_max\_delay\_ms**](#kafka_connect_config.scheduled_rebalance_max_delay_ms)`integer`- min: `0`
- max: `600000`The maximum delay that is scheduled in order to wait for the return of one or more departed workers before rebalancing and reassigning their connectors and tasks to the group. During this period the connectors and tasks of the departed workers remain unassigned. Defaults to 5 minutes. | | [**kafka\_connect\_config.session\_timeout\_ms**](#kafka_connect_config.session_timeout_ms)`integer`- min: `1`
- max: `2147483647`The timeout in milliseconds used to detect failures when using Kafka’s group management facilities (defaults to 10000). | | [**kafka\_connect\_plugin\_versions**](#kafka_connect_plugin_versions)`array[object]`The plugin selected by the user | | [**kafka\_connect\_plugin\_versions.\[0\].plugin\_name**](#kafka_connect_plugin_versions.\[0].plugin_name)`string`- maxLength: `128`The name of the plugin | | [**kafka\_connect\_plugin\_versions.\[0\].version**](#kafka_connect_plugin_versions.\[0].version)`string`- maxLength: `128`The version of the plugin | | [**kafka\_connect\_secret\_providers**](#kafka_connect_secret_providers)`array[object]`Configure external secret providers in order to reference external secrets in connector configuration. Currently Hashicorp Vault (provider: vault, auth\_method: token) and AWS Secrets Manager (provider: aws, auth\_method: credentials) are supported. Secrets can be referenced in connector config with ${\<provider\_name\>:\<secret\_path\>:\<key\_name\>} | | [**kafka\_connect\_secret\_providers.\[0\].name**](#kafka_connect_secret_providers.\[0].name)`string`Name of the secret provider. Used to reference secrets in connector config. | | [**kafka\_connect\_secret\_providers.\[0\].vault**](#kafka_connect_secret_providers.\[0].vault)`object`- required: `auth_method,address`Vault secret provider configuration | | [**kafka\_connect\_secret\_providers.\[0\].vault.auth\_method**](#kafka_connect_secret_providers.\[0].vault.auth_method)`string`- enum: `token`An enumeration. | | [**kafka\_connect\_secret\_providers.\[0\].vault.address**](#kafka_connect_secret_providers.\[0].vault.address)`string`- maxLength: `65536`Address of the Vault server | | [**kafka\_connect\_secret\_providers.\[0\].vault.engine\_version**](#kafka_connect_secret_providers.\[0].vault.engine_version)`integer`- enum: `1,2`An enumeration. | | [**kafka\_connect\_secret\_providers.\[0\].vault.prefix\_path\_depth**](#kafka_connect_secret_providers.\[0].vault.prefix_path_depth)`integer`Prefix path depth of the secrets Engine. Default is 1. If the secrets engine path has more than one segment it has to be increased to the number of segments. | | [**kafka\_connect\_secret\_providers.\[0\].vault.token**](#kafka_connect_secret_providers.\[0].vault.token)`string`- maxLength: `256`Token used to authenticate with vault and auth method \`token\`. | | [**kafka\_connect\_secret\_providers.\[0\].vault.server\_pem**](#kafka_connect_secret_providers.\[0].vault.server_pem)`string`- maxLength: `4096`PEM encoded certificate of the Vault server. Required if the vault server uses a self-signed certificate. | | [**kafka\_connect\_secret\_providers.\[0\].aws**](#kafka_connect_secret_providers.\[0].aws)`object`- required: `auth_method,region`AWS secret provider configuration | | [**kafka\_connect\_secret\_providers.\[0\].aws.auth\_method**](#kafka_connect_secret_providers.\[0].aws.auth_method)`string`- enum: `credentials`An enumeration. | | [**kafka\_connect\_secret\_providers.\[0\].aws.access\_key**](#kafka_connect_secret_providers.\[0].aws.access_key)`string`- maxLength: `128`Access key used to authenticate with aws | | [**kafka\_connect\_secret\_providers.\[0\].aws.secret\_key**](#kafka_connect_secret_providers.\[0].aws.secret_key)`string`- maxLength: `128`Secret key used to authenticate with aws | | [**kafka\_connect\_secret\_providers.\[0\].aws.region**](#kafka_connect_secret_providers.\[0].aws.region)`string`- maxLength: `64`Region used to lookup secrets with AWS SecretManager | | [**kafka\_connect\_secret\_providers.\[0\].env**](#kafka_connect_secret_providers.\[0].env)`object`- required: `secrets`ENV secret provider configuration | | [**kafka\_connect\_secret\_providers.\[0\].env.secrets**](#kafka_connect_secret_providers.\[0].env.secrets)`object`Key/value map of secrets for ENV secret provider | | [**kafka\_connect\_secret\_providers.\[0\].azure**](#kafka_connect_secret_providers.\[0].azure)`object`- required: `auth_method`Azure KeyVault secret provider configuration | | [**kafka\_connect\_secret\_providers.\[0\].azure.auth\_method**](#kafka_connect_secret_providers.\[0].azure.auth_method)`string`- enum: `credentials`An enumeration. | | [**kafka\_connect\_secret\_providers.\[0\].azure.client\_id**](#kafka_connect_secret_providers.\[0].azure.client_id)`string`- maxLength: `128`Azure client ID for the service principal. | | [**kafka\_connect\_secret\_providers.\[0\].azure.tenant\_id**](#kafka_connect_secret_providers.\[0].azure.tenant_id)`string`- maxLength: `128`Azure tenant ID for the service principal. | | [**kafka\_connect\_secret\_providers.\[0\].azure.secret**](#kafka_connect_secret_providers.\[0].azure.secret)`string`- maxLength: `256`Azure client secret for the service principal. | | [**kafka\_diskless**](#kafka_diskless)`object`- required: `enabled`Kafka Diskless configuration values | | [**kafka\_diskless.enabled**](#kafka_diskless.enabled)`boolean`Whether to enable the Diskless functionality | | [**kafka\_diskless.auto\_diskless\_topic\_regexes**](#kafka_diskless.auto_diskless_topic_regexes)`array[string]`- maxItems: `32`The regexes of topics to auto enable diskless. Topics matching any of the regexes will be created as diskless topics. | | [**inkless**](#inkless)`object`- required: `enabled`Inkless configuration values | | [**inkless.enabled**](#inkless.enabled)`boolean`Whether to enable the Inkless functionality | | [**sasl\_oauthbearer\_allowed\_urls**](#sasl_oauthbearer_allowed_urls)`array[string]`List of allowed URLs for SASL OAUTHBEARER authentication. Only HTTPS URLs are allowed for security reasons. | | [**gcp\_auth\_allowed\_urls**](#gcp_auth_allowed_urls)`array[string]`Allow-list of HTTPS URLs used to validate GCP credential\_source requests for Kafka Connect. | | [**kafka\_rest**](#kafka_rest)`boolean`Enable Kafka-REST service | | [**kafka\_version**](#kafka_version)`string,null`- enum: `3.9,4.1,4.2`Kafka major version | | [**karapace\_version**](#karapace_version)`string,null`- enum: `6.2.1,6.2.2,null`Select a Karapace version for this service, or select Latest to use the latest available version automatically. New versions become available after installation during a maintenance update. | | [**schema\_registry**](#schema_registry)`boolean`Enable Schema-Registry service | | [**kafka\_rest\_authorization**](#kafka_rest_authorization)`boolean`Enable authorization in Kafka-REST service | | [**kafka\_rest\_config**](#kafka_rest_config)`object`Kafka REST configuration | | [**kafka\_rest\_config.producer\_acks**](#kafka_rest_config.producer_acks)`string`- enum: `all,-1,0,1`
- default: `1`The number of acknowledgments the producer requires the leader to have received before considering a request complete. If set to 'all' or '-1', the leader will wait for the full set of in-sync replicas to acknowledge the record. | | [**kafka\_rest\_config.producer\_compression\_type**](#kafka_rest_config.producer_compression_type)`string`- enum: `gzip,snappy,lz4,zstd,none`Specify the default compression type for producers. This configuration accepts the standard compression codecs ('gzip', 'snappy', 'lz4', 'zstd'). It additionally accepts 'none' which is the default and equivalent to no compression. | | [**kafka\_rest\_config.producer\_linger\_ms**](#kafka_rest_config.producer_linger_ms)`integer`- min: `0`
- max: `5000`
- default: `0`Wait for up to the given delay to allow batching records together | | [**kafka\_rest\_config.producer\_max\_request\_size**](#kafka_rest_config.producer_max_request_size)`integer`- min: `0`
- max: `2147483647`
- default: `1048576`The maximum size of a request in bytes. Note that Kafka broker can also cap the record batch size. | | [**kafka\_rest\_config.consumer\_enable\_auto\_commit**](#kafka_rest_config.consumer_enable_auto_commit)`boolean`- default: `true`If true the consumer's offset will be periodically committed to Kafka in the background | | [**kafka\_rest\_config.consumer\_idle\_disconnect\_timeout**](#kafka_rest_config.consumer_idle_disconnect_timeout)`integer`- min: `0`
- max: `2147483647`
- default: `0`Specifies the maximum duration (in seconds) a client can remain idle before it is deleted. If a consumer is inactive, it will exit the consumer group, and its state will be discarded. A value of 0 (default) indicates that the consumer will not be disconnected automatically due to inactivity. | | [**kafka\_rest\_config.consumer\_request\_max\_bytes**](#kafka_rest_config.consumer_request_max_bytes)`integer`- min: `0`
- max: `671088640`
- default: `67108864`Maximum number of bytes in unencoded message keys and values by a single request | | [**kafka\_rest\_config.consumer\_request\_timeout\_ms**](#kafka_rest_config.consumer_request_timeout_ms)`integer`- min: `1000`
- max: `30000`
- enum: `1000,15000,30000`
- default: `1000`The maximum total time to wait for messages for a request if the maximum number of messages has not yet been reached | | [**kafka\_rest\_config.name\_strategy**](#kafka_rest_config.name_strategy)`string`- enum: `topic_name,record_name,topic_record_name`
- default: `topic_name`Name strategy to use when selecting subject for storing schemas | | [**kafka\_rest\_config.name\_strategy\_validation**](#kafka_rest_config.name_strategy_validation)`boolean`- default: `true`If true, validate that given schema is registered under expected subject name by the used name strategy when producing messages. | | [**kafka\_rest\_config.simpleconsumer\_pool\_size\_max**](#kafka_rest_config.simpleconsumer_pool_size_max)`integer`- min: `10`
- max: `250`
- default: `25`Maximum number of SimpleConsumers that can be instantiated per broker | | [**tiered\_storage**](#tiered_storage)`object`Tiered storage configuration | | [**tiered\_storage.enabled**](#tiered_storage.enabled)`boolean`Whether to enable the tiered storage functionality | | [**schema\_registry\_config**](#schema_registry_config)`object`Schema Registry configuration | | [**schema\_registry\_config.topic\_name**](#schema_registry_config.topic_name)`string`- maxLength: `249`The durable single partition topic that acts as the durable log for the data. This topic must be compacted to avoid losing data due to retention policy. Please note that changing this configuration in an existing Schema Registry / Karapace setup leads to previous schemas being inaccessible, data encoded with them potentially unreadable and schema ID sequence put out of order. It's only possible to do the switch while Schema Registry / Karapace is disabled. Defaults to \`\_schemas\`. | | [**schema\_registry\_config.leader\_eligibility**](#schema_registry_config.leader_eligibility)`boolean`If true, Karapace / Schema Registry on the service nodes can participate in leader election. It might be needed to disable this when the schemas topic is replicated to a secondary cluster and Karapace / Schema Registry there must not participate in leader election. Defaults to \`true\`. | | [**schema\_registry\_config.schema\_reader\_strict\_mode**](#schema_registry_config.schema_reader_strict_mode)`boolean`If enabled, causes the Karapace schema-registry service to shutdown when there are invalid schema records in the \`\_schemas\` topic. Defaults to \`false\`. | | [**schema\_registry\_config.retriable\_errors\_silenced**](#schema_registry_config.retriable_errors_silenced)`boolean`If enabled, kafka errors which can be retried or custom errors specified for the service will not be raised, instead, a warning log is emitted. This will denoise issue tracking systems, i.e. sentry. Defaults to \`true\`. | | [**schema\_registry\_config.sasl\_oauthbearer\_authentication\_enabled**](#schema_registry_config.sasl_oauthbearer_authentication_enabled)`boolean`If enabled, the Schema Registry validates OAuth2/OIDC JWT bearer tokens on incoming requests. Requires the OIDC provider settings under the \`kafka\` configuration (\`sasl\_oauthbearer\_jwks\_endpoint\_url\` and related). Defaults to \`false\`. | | [**schema\_registry\_config.sasl\_oauthbearer\_authorization\_enabled**](#schema_registry_config.sasl_oauthbearer_authorization_enabled)`boolean`If enabled, the Schema Registry enforces role-based authorization derived from the JWT roles claim. Enabling this automatically enables \`sasl\_oauthbearer\_authentication\_enabled\` when it is not already enabled, since authorization requires authentication. Defaults to \`false\`. | | [**schema\_registry\_config.sasl\_oauthbearer\_roles\_claim\_path**](#schema_registry_config.sasl_oauthbearer_roles_claim_path)`string`- maxLength: `128`JSON path used to extract the roles claim from the JWT for Schema Registry authorization. Defaults to \`resource\_access.karapace.roles\`. | | [**schema\_registry\_config.sasl\_oauthbearer\_method\_roles**](#schema_registry_config.sasl_oauthbearer_method_roles)`string`- maxLength: `4096`JSON object mapping HTTP methods to the list of roles allowed to perform them on the Schema Registry, provided as a JSON-encoded string. Role names use the \`karapace.\` prefix, e.g. \`karapace.schema:read\`. Defaults to \`{"GET": \["karapace.schema:read", "karapace.subject:read"], "POST": \[], "PUT": \[], "DELETE": \[]}\`. | | [**aiven\_kafka\_topic\_messages**](#aiven_kafka_topic_messages)`boolean`Allow access to read Kafka topic messages in the Aiven Console and REST API. | | [**enable\_ipv6**](#enable_ipv6)`boolean`Register AAAA DNS records for the service, and allow IPv6 packets to service ports | ## Topic parameters[​](#topic-parameters "Direct link to Topic parameters") | | | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | [**compression\_type**](#compression_type)`string`- enum: `gzip,snappy,lz4,zstd,uncompressed,producer`Specify the final compression type for a given topic. This configuration accepts the standard compression codecs ('gzip', 'snappy', 'lz4', 'zstd'). It additionally accepts 'uncompressed' which is equivalent to no compression; and 'producer' which means retain the original compression codec set by the producer. | | [**min\_insync\_replicas**](#min_insync_replicas)`integer`- min: `1`
- max: `7`When a producer sets acks to 'all' (or '-1'), this configuration specifies the minimum number of replicas that must acknowledge a write for the write to be considered successful. If this minimum cannot be met, then the producer will raise an exception (either NotEnoughReplicas or NotEnoughReplicasAfterAppend). When used together, min.insync.replicas and acks allow you to enforce greater durability guarantees. A typical scenario would be to create a topic with a replication factor of 3, set min.insync.replicas to 2, and produce with acks of 'all'. This will ensure that the producer raises an exception if a majority of replicas do not receive a write. | | [**cleanup\_policy**](#cleanup_policy)`string`- enum: `compact,delete`Retention policy for old segments. By default, 'delete' discards old segments when they reach the retention time or size limit. The 'compact' setting enables log compaction for the topic. | | [**delete\_retention\_ms**](#delete_retention_ms)`integer`- min: `0`
- max: `9223372036854775807`The amount of time to retain delete tombstone markers for log compacted topics. This setting also gives a bound on the time in which a consumer must complete a read if they begin from offset 0 to ensure that they get a valid snapshot of the final stage (otherwise delete tombstones may be collected before they complete their scan). | | [**file\_delete\_delay\_ms**](#file_delete_delay_ms)`integer`- min: `0`
- max: `9223372036854775807`The time to wait before deleting a file from the filesystem. | | [**flush\_messages**](#flush_messages)`integer`- min: `0`
- max: `9223372036854775807`This setting allows specifying an interval at which we will force an fsync of data written to the log. For example if this was set to 1 we would fsync after every message; if it were 5 we would fsync after every five messages. In general we recommend you not set this and use replication for durability and allow the operating system's background flush capabilities as it is more efficient. | | [**flush\_ms**](#flush_ms)`integer`- min: `0`
- max: `9223372036854775807`This setting allows specifying a time interval at which we will force an fsync of data written to the log. For example if this was set to 1000 we would fsync after 1000 ms had passed. In general we recommend you not set this and use replication for durability and allow the operating system's background flush capabilities as it is more efficient. | | [**index\_interval\_bytes**](#index_interval_bytes)`integer`- min: `0`
- max: `9223372036854775807`This setting controls how frequently Kafka adds an index entry to its offset index. The default setting ensures that we index a message roughly every 4096 bytes. More indexing allows reads to jump closer to the exact position in the log but makes the index larger. You probably don't need to change this. | | [**local\_retention\_bytes**](#local_retention_bytes)`integer`- min: `-2`
- max: `9223372036854775807`This configuration controls the maximum bytes tiered storage will retain segment files locally before it will discard old log segments to free up space. If set to -2, the limit is equal to overall retention time. If set to -1, no limit is applied but it's possible only if overall retention is also -1. | | [**local\_retention\_ms**](#local_retention_ms)`integer`- min: `-2`
- max: `9223372036854775807`This configuration controls the maximum time tiered storage will retain segment files locally before it will discard old log segments to free up space. If set to -2, the time limit is equal to overall retention time. If set to -1, no time limit is applied but it's possible only if overall retention is also -1. | | [**max\_compaction\_lag\_ms**](#max_compaction_lag_ms)`integer`- min: `0`
- max: `9223372036854775807`The maximum time a message will remain ineligible for compaction in the log. Only applicable for logs that are being compacted. | | [**max\_message\_bytes**](#max_message_bytes)`integer`- min: `0`
- max: `9223372036854775807`The largest record batch size allowed by Kafka (after compression if compression is enabled). If this is increased and there are consumers older than 0.10.2, the consumers' fetch size must also be increased so that the they can fetch record batches this large. In the latest message format version, records are always grouped into batches for efficiency. In previous message format versions, uncompressed records are not grouped into batches and this limit only applies to a single record in that case. | | [**message\_timestamp\_after\_max\_ms**](#message_timestamp_after_max_ms)`integer`- min: `0`
- max: `9223372036854775807`The maximum difference allowed between the timestamp when a broker receives a message and the timestamp specified in the message. If message.timestamp.type=CreateTime, a message will be rejected if the difference in timestamp exceeds this threshold. Applies only for messages with timestamps later than the broker's timestamp. | | [**message\_timestamp\_before\_max\_ms**](#message_timestamp_before_max_ms)`integer`- min: `0`
- max: `9223372036854775807`The maximum difference allowed between the timestamp when a broker receives a message and the timestamp specified in the message. If message.timestamp.type=CreateTime, a message will be rejected if the difference in timestamp exceeds this threshold. Applies only for messages with timestamps earlier than the broker's timestamp. | | [**message\_timestamp\_type**](#message_timestamp_type)`string`- enum: `CreateTime,LogAppendTime`Define whether the timestamp in the message is message create time or log append time. | | [**min\_compaction\_lag\_ms**](#min_compaction_lag_ms)`integer`- min: `0`
- max: `9223372036854775807`The minimum time a message will remain uncompacted in the log. Only applicable for logs that are being compacted. | | [**preallocate**](#preallocate)`boolean`True if we should preallocate the file on disk when creating a new log segment. | | [**retention\_bytes**](#retention_bytes)`integer`- min: `-1`
- max: `9223372036854775807`This configuration controls the maximum size a partition (which consists of log segments) can grow to before we will discard old log segments to free up space if we are using the 'delete' retention policy. By default there is no size limit only a time limit. Since this limit is enforced at the partition level, multiply it by the number of partitions to compute the topic retention in bytes. | | [**retention\_ms**](#retention_ms)`integer`- min: `0`
- max: `259200000`This configuration controls the maximum time we will retain a log before we will discard old log segments to free up space if we are using the 'delete' retention policy. This represents an SLA on how soon consumers must read their data. If set to -1, no time limit is applied. | | [**segment\_bytes**](#segment_bytes)`integer`- min: `14`
- max: `9223372036854775807`This configuration controls the segment file size for the log. Retention and cleaning is always done a file at a time so a larger segment size means fewer files but less granular control over retention. Setting this to a very low value has consequences, and the Aiven management plane ignores values less than 10 megabytes. | | [**segment\_index\_bytes**](#segment_index_bytes)`integer`- min: `0`
- max: `9223372036854775807`This configuration controls the size of the index that maps offsets to file positions. We preallocate this index file and shrink it only after log rolls. You generally should not need to change this setting. | | [**segment\_jitter\_ms**](#segment_jitter_ms)`integer`- min: `0`
- max: `9223372036854775807`The maximum random jitter subtracted from the scheduled segment roll time to avoid thundering herds of segment rolling | | [**segment\_ms**](#segment_ms)`integer`- min: `1`
- max: `9223372036854775807`This configuration controls the period of time after which Kafka will force the log to roll even if the segment file isn't full to ensure that retention can delete or compact old data. Setting this to a very low value has consequences, and the Aiven management plane ignores values less than 10 seconds. | --- # Advanced parameters for Aiven for Apache Kafka® Advanced parameters let you customize your Aiven for Apache Kafka® service beyond the default configuration to tune performance, control resource usage, and adjust service behavior to fit your workload. Select your service type to view the available parameters: * [Free tier](/docs/products/kafka/reference/advanced-params-free-tier.md) * [Developer tier](/docs/products/kafka/reference/advanced-params-dev-tier.md) * [Classic Kafka](/docs/products/kafka/reference/advanced-params.md) * [Standard Kafka](/docs/products/kafka/reference/advanced-params-standard.md) --- # Advanced parameters for Standard Kafka View the configuration options for Standard Kafka. These configurations apply to Standard Kafka on Aiven Cloud. Standard Kafka supports two topic types. Classic topics use `remote.storage.enable=true`, and diskless topics use `diskless.enable=true`. For more information about the differences and use cases, see [Classic topics vs. diskless topics](/docs/products/kafka/diskless/concepts/topics-vs-classic.md#compare-classic-and-diskless-topics). ## Broker parameters[​](#broker-parameters "Direct link to Broker parameters") | | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | [**backup\_interval\_hours**](#backup_interval_hours)`integer,null`- min: `3`
- max: `24`
- enum: `3,4,6,8,12,24,null`Interval in hours between automatic backups. Minimum value is 3 hours. Must be a divisor of 24 (3, 4, 6, 8, 12, 24). (Applicable to ACU plans only) | | [**backup\_retention\_days**](#backup_retention_days)`integer,null`- min: `1`
- max: `30`Number of days to retain automatic backups. Backups older than this value will be automatically deleted. (Applicable to ACU plans only) | | [**custom\_domain**](#custom_domain)`string,null`- maxLength: `255`Serve the web frontend using a custom CNAME pointing to the Aiven DNS name. When you set a custom domain for a service deployed in a VPC, the service certificate is only created for the public-\* hostname and the custom domain. | | [**ip\_filter**](#ip_filter)`array[string,object]`- maxItems: `8000`
- default: `0.0.0.0/0,::/0`Allow incoming connections from CIDR address block, e.g. '10.20.0.0/16' | | [**ip\_filter.\[0\].description**](#ip_filter.\[0].description)`string`- maxLength: `1024`Description for IP filter list entry | | [**ip\_filter.\[0\].network**](#ip_filter.\[0].network)`string`- maxLength: `43`CIDR address block | | [**service\_log**](#service_log)`boolean,null`Store logs for the service so that they are available in the HTTP API and console. | | [**static\_ips**](#static_ips)`boolean`Use static public IP addresses | | [**single\_zone**](#single_zone)`object`Single-zone configuration | | [**single\_zone.enabled**](#single_zone.enabled)`boolean`Whether to allocate nodes on the same Availability Zone or spread across zones available. By default service nodes are spread across different AZs. The single AZ support is best-effort and may temporarily allocate nodes in different AZs e.g. in case of capacity limitations in one AZ. | | [**single\_zone.availability\_zone**](#single_zone.availability_zone)`string`- maxLength: `40`The availability zone to use for the service. This is only used when enabled is set to true. If not set the service will be allocated in random AZ.The AZ is not guaranteed, and the service may be allocated in a different AZ if the selected AZ is not available. Zones will not be validated and invalid zones will be ignored, falling back to random AZ selection. Common availability zones include: AWS (euc1-az1, euc1-az2, euc1-az3), GCP (europe-west1-a, europe-west1-b, europe-west1-c), Azure (germanywestcentral/1, germanywestcentral/2, germanywestcentral/3). | | [**preferred\_zones**](#preferred_zones)`array[string]`- maxItems: `10`List of preferred zone IDs for service node placement. Nodes will be placed in these zones when available. If a specified zone is unavailable (e.g., due to capacity constraints), nodes will be placed in other available zones to maintain the configured number of zones for availability. Invalid zone IDs are rejected at configuration time. Zone IDs are cloud-specific: AWS uses zone IDs like 'euc1-az1', GCP uses zone names like 'europe-west1-a', and Azure uses 'location/zone' format like 'germanywestcentral/1'. If single\_zone is enabled with an availability\_zone, that setting takes precedence over preferred\_zones. Changes take effect on next node recreation (e.g., maintenance or plan change). For eligible plans, nodes outside preferred zones are automatically rebalanced once per day. | | [**private\_access**](#private_access)`object`Allow access to selected service ports from private networks | | [**private\_access.kafka**](#private_access.kafka)`boolean`Allow clients to connect to kafka with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | [**private\_access.kafka\_connect**](#private_access.kafka_connect)`boolean`Allow clients to connect to kafka\_connect with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | [**private\_access.kafka\_rest**](#private_access.kafka_rest)`boolean`Allow clients to connect to kafka\_rest with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | [**private\_access.prometheus**](#private_access.prometheus)`boolean`Allow clients to connect to prometheus with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | [**private\_access.schema\_registry**](#private_access.schema_registry)`boolean`Allow clients to connect to schema\_registry with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | [**public\_access**](#public_access)`object`Allow access to selected service ports from the public Internet | | [**public\_access.kafka**](#public_access.kafka)`boolean`Allow clients to connect to kafka from the public internet for service nodes that are in a project VPC or another type of private network | | [**public\_access.kafka\_connect**](#public_access.kafka_connect)`boolean`Allow clients to connect to kafka\_connect from the public internet for service nodes that are in a project VPC or another type of private network | | [**public\_access.kafka\_rest**](#public_access.kafka_rest)`boolean`Allow clients to connect to kafka\_rest from the public internet for service nodes that are in a project VPC or another type of private network | | [**public\_access.prometheus**](#public_access.prometheus)`boolean`Allow clients to connect to prometheus from the public internet for service nodes that are in a project VPC or another type of private network | | [**public\_access.schema\_registry**](#public_access.schema_registry)`boolean`Allow clients to connect to schema\_registry from the public internet for service nodes that are in a project VPC or another type of private network | | [**privatelink\_access**](#privatelink_access)`object`Allow access to selected service components through Privatelink | | [**privatelink\_access.jolokia**](#privatelink_access.jolokia)`boolean`Enable jolokia | | [**privatelink\_access.kafka**](#privatelink_access.kafka)`boolean`Enable kafka | | [**privatelink\_access.kafka\_connect**](#privatelink_access.kafka_connect)`boolean`Enable kafka\_connect | | [**privatelink\_access.kafka\_rest**](#privatelink_access.kafka_rest)`boolean`Enable kafka\_rest | | [**privatelink\_access.prometheus**](#privatelink_access.prometheus)`boolean`Enable prometheus | | [**privatelink\_access.schema\_registry**](#privatelink_access.schema_registry)`boolean`Enable schema\_registry | | [**letsencrypt\_sasl**](#letsencrypt_sasl)`boolean,null`Use a Let's Encrypt certificate authority (CA) for Kafka SASL authentication. (Default: False) | | [**letsencrypt\_sasl\_privatelink**](#letsencrypt_sasl_privatelink)`boolean,null`Use a Let's Encrypt certificate authority (CA) for Kafka SASL authentication via Privatelink. (Default: False) | | [**kafka**](#kafka)`object`- default: `[object Object]`Kafka broker configuration values | | [**kafka.group\_coordinator\_rebalance\_protocols**](#kafka.group_coordinator_rebalance_protocols)`string`- enum: `classic,classic,consumer,classic,share,classic,streams,classic,consumer,share,classic,consumer,streams,classic,consumer,share,streams`The enabled consumer group rebalance protocols. Use consumer, classic, share, streams to enable Kafka share groups. | | [**kafka.message\_max\_bytes**](#kafka.message_max_bytes)`integer`- min: `0`
- max: `20971520`The maximum size of message that the server can receive. (Default: 1048588 bytes (1 mebibyte + 12 bytes)) | | [**kafka.log\_retention\_bytes**](#kafka.log_retention_bytes)`integer`- min: `-1`
- max: `9223372036854775807`The maximum size of the log before deleting messages (Default: -1) | | [**kafka.log\_retention\_hours**](#kafka.log_retention_hours)`integer`- min: `-1`
- max: `9223372036854775807`The number of hours to keep a log file before deleting it. Use -1 for unlimited retention or 1 or higher. Setting 0 is invalid and prevents Kafka from starting. (Default: 168 hours, or 1 week) | | [**kafka.log\_retention\_ms**](#kafka.log_retention_ms)`integer`- range: `60000` - `9223372036854775807`
- min: `-1`The number of milliseconds to keep a log file before deleting it (in milliseconds), If not set, the value in log.retention.minutes is used. If set to -1, no time limit is applied. (Default: null, log.retention.hours applies) | | [**kafka.auto\_create\_topics\_enable**](#kafka.auto_create_topics_enable)`boolean`Enable auto-creation of topics. (Default: true) | | [**kafka.num\_partitions**](#kafka.num_partitions)`integer`- min: `1`
- max: `2048`Number of partitions for auto-created topics (Default: 1) | | [**kafka.default\_replication\_factor**](#kafka.default_replication_factor)`integer`- min: `1`
- max: `3`Replication factor for created topics. For Classic topics, the default is 3. For Diskless topics, the default is 1. | | [**kafka.sasl\_oauthbearer\_expected\_audience**](#kafka.sasl_oauthbearer_expected_audience)`string`The (optional) comma-delimited setting for the broker to use to verify that the JWT was issued for one of the expected audiences. (Default: null) | | [**kafka.sasl\_oauthbearer\_expected\_issuer**](#kafka.sasl_oauthbearer_expected_issuer)`string`Optional setting for the broker to use to verify that the JWT was created by the expected issuer.(Default: null) | | [**kafka.sasl\_oauthbearer\_jwks\_endpoint\_url**](#kafka.sasl_oauthbearer_jwks_endpoint_url)`string`OIDC JWKS endpoint URL. By setting this the SASL SSL OAuth2/OIDC authentication is enabled. See also other options for SASL OAuth2/OIDC. (Default: null) | | [**kafka\_authentication\_methods**](#kafka_authentication_methods)`object`Kafka authentication methods | | [**kafka\_authentication\_methods.certificate**](#kafka_authentication_methods.certificate)`boolean`- default: `true`Enable certificate/SSL authentication | | [**kafka\_authentication\_methods.sasl**](#kafka_authentication_methods.sasl)`boolean`Enable SASL authentication | | [**kafka\_sasl\_mechanisms**](#kafka_sasl_mechanisms)`object`Kafka SASL mechanisms | | [**kafka\_sasl\_mechanisms.plain**](#kafka_sasl_mechanisms.plain)`boolean`- default: `true`Enable PLAIN mechanism | | [**kafka\_sasl\_mechanisms.scram\_sha\_256**](#kafka_sasl_mechanisms.scram_sha_256)`boolean`- default: `true`Enable SCRAM-SHA-256 mechanism | | [**kafka\_sasl\_mechanisms.scram\_sha\_512**](#kafka_sasl_mechanisms.scram_sha_512)`boolean`- default: `true`Enable SCRAM-SHA-512 mechanism | | [**follower\_fetching**](#follower_fetching)`object`Enable follower fetching | | [**follower\_fetching.enabled**](#follower_fetching.enabled)`boolean`Whether to enable the follower fetching functionality | | [**kafka\_connect**](#kafka_connect)`boolean`Enable Kafka Connect service | | [**kafka\_connect\_config**](#kafka_connect_config)`object`Kafka Connect configuration values | | [**kafka\_connect\_config.prefer\_ipv6\_address\_enable**](#kafka_connect_config.prefer_ipv6_address_enable)`boolean`When enabled, connectors will automatically resolve IPv6 addresses from external server names configured with dual-stack. | | [**kafka\_connect\_config.connector\_client\_config\_override\_policy**](#kafka_connect_config.connector_client_config_override_policy)`string`- enum: `None,All`Defines what client configurations can be overridden by the connector. Default is None | | [**kafka\_connect\_config.consumer\_auto\_offset\_reset**](#kafka_connect_config.consumer_auto_offset_reset)`string`- enum: `earliest,latest`What to do when there is no initial offset in Kafka or if the current offset does not exist any more on the server. Default is earliest | | [**kafka\_connect\_config.consumer\_fetch\_max\_bytes**](#kafka_connect_config.consumer_fetch_max_bytes)`integer`- min: `1048576`
- max: `104857600`Records are fetched in batches by the consumer, and if the first record batch in the first non-empty partition of the fetch is larger than this value, the record batch will still be returned to ensure that the consumer can make progress. As such, this is not a absolute maximum. | | [**kafka\_connect\_config.consumer\_isolation\_level**](#kafka_connect_config.consumer_isolation_level)`string`- enum: `read_uncommitted,read_committed`Transaction read isolation level. read\_uncommitted is the default, but read\_committed can be used if consume-exactly-once behavior is desired. | | [**kafka\_connect\_config.consumer\_max\_partition\_fetch\_bytes**](#kafka_connect_config.consumer_max_partition_fetch_bytes)`integer`- min: `1048576`
- max: `104857600`Records are fetched in batches by the consumer.If the first record batch in the first non-empty partition of the fetch is larger than this limit, the batch will still be returned to ensure that the consumer can make progress. | | [**kafka\_connect\_config.consumer\_max\_poll\_interval\_ms**](#kafka_connect_config.consumer_max_poll_interval_ms)`integer`- min: `1`
- max: `2147483647`The maximum delay in milliseconds between invocations of poll() when using consumer group management (defaults to 300000). | | [**kafka\_connect\_config.consumer\_max\_poll\_records**](#kafka_connect_config.consumer_max_poll_records)`integer`- min: `1`
- max: `10000`The maximum number of records returned in a single call to poll() (defaults to 500). | | [**kafka\_connect\_config.offset\_flush\_interval\_ms**](#kafka_connect_config.offset_flush_interval_ms)`integer`- min: `1`
- max: `100000000`The interval at which to try committing offsets for tasks (defaults to 60000). | | [**kafka\_connect\_config.offset\_flush\_timeout\_ms**](#kafka_connect_config.offset_flush_timeout_ms)`integer`- min: `1`
- max: `2147483647`Maximum number of milliseconds to wait for records to flush and partition offset data to be committed to offset storage before cancelling the process and restoring the offset data to be committed in a future attempt (defaults to 5000). | | [**kafka\_connect\_config.producer\_batch\_size**](#kafka_connect_config.producer_batch_size)`integer`- min: `0`
- max: `5242880`This setting gives the upper bound of the batch size to be sent. If there are fewer than this many bytes accumulated for this partition, the producer will 'linger' for the linger.ms time waiting for more records to show up. A batch size of zero will disable batching entirely (defaults to 16384). | | [**kafka\_connect\_config.producer\_buffer\_memory**](#kafka_connect_config.producer_buffer_memory)`integer`- min: `5242880`
- max: `134217728`The total bytes of memory the producer can use to buffer records waiting to be sent to the broker (defaults to 33554432). | | [**kafka\_connect\_config.producer\_compression\_type**](#kafka_connect_config.producer_compression_type)`string`- enum: `gzip,snappy,lz4,zstd,none`Specify the default compression type for producers. This configuration accepts the standard compression codecs ('gzip', 'snappy', 'lz4', 'zstd'). It additionally accepts 'none' which is the default and equivalent to no compression. | | [**kafka\_connect\_config.producer\_linger\_ms**](#kafka_connect_config.producer_linger_ms)`integer`- min: `0`
- max: `5000`This setting gives the upper bound on the delay for batching: once there is batch.size worth of records for a partition it will be sent immediately regardless of this setting, however if there are fewer than this many bytes accumulated for this partition the producer will 'linger' for the specified time waiting for more records to show up. Defaults to 0. | | [**kafka\_connect\_config.producer\_max\_request\_size**](#kafka_connect_config.producer_max_request_size)`integer`- min: `131072`
- max: `67108864`This setting will limit the number of record batches the producer will send in a single request to avoid sending huge requests. | | [**kafka\_connect\_config.scheduled\_rebalance\_max\_delay\_ms**](#kafka_connect_config.scheduled_rebalance_max_delay_ms)`integer`- min: `0`
- max: `600000`The maximum delay that is scheduled in order to wait for the return of one or more departed workers before rebalancing and reassigning their connectors and tasks to the group. During this period the connectors and tasks of the departed workers remain unassigned. Defaults to 5 minutes. | | [**kafka\_connect\_config.session\_timeout\_ms**](#kafka_connect_config.session_timeout_ms)`integer`- min: `1`
- max: `2147483647`The timeout in milliseconds used to detect failures when using Kafka’s group management facilities (defaults to 10000). | | [**kafka\_connect\_plugin\_versions**](#kafka_connect_plugin_versions)`array[object]`The plugin selected by the user | | [**kafka\_connect\_plugin\_versions.\[0\].plugin\_name**](#kafka_connect_plugin_versions.\[0].plugin_name)`string`- maxLength: `128`The name of the plugin | | [**kafka\_connect\_plugin\_versions.\[0\].version**](#kafka_connect_plugin_versions.\[0].version)`string`- maxLength: `128`The version of the plugin | | [**kafka\_connect\_secret\_providers**](#kafka_connect_secret_providers)`array[object]`Configure external secret providers in order to reference external secrets in connector configuration. Currently Hashicorp Vault (provider: vault, auth\_method: token) and AWS Secrets Manager (provider: aws, auth\_method: credentials) are supported. Secrets can be referenced in connector config with ${\<provider\_name\>:\<secret\_path\>:\<key\_name\>} | | [**kafka\_connect\_secret\_providers.\[0\].name**](#kafka_connect_secret_providers.\[0].name)`string`Name of the secret provider. Used to reference secrets in connector config. | | [**kafka\_connect\_secret\_providers.\[0\].vault**](#kafka_connect_secret_providers.\[0].vault)`object`- required: `auth_method,address`Vault secret provider configuration | | [**kafka\_connect\_secret\_providers.\[0\].vault.auth\_method**](#kafka_connect_secret_providers.\[0].vault.auth_method)`string`- enum: `token`An enumeration. | | [**kafka\_connect\_secret\_providers.\[0\].vault.address**](#kafka_connect_secret_providers.\[0].vault.address)`string`- maxLength: `65536`Address of the Vault server | | [**kafka\_connect\_secret\_providers.\[0\].vault.engine\_version**](#kafka_connect_secret_providers.\[0].vault.engine_version)`integer`- enum: `1,2`An enumeration. | | [**kafka\_connect\_secret\_providers.\[0\].vault.prefix\_path\_depth**](#kafka_connect_secret_providers.\[0].vault.prefix_path_depth)`integer`Prefix path depth of the secrets Engine. Default is 1. If the secrets engine path has more than one segment it has to be increased to the number of segments. | | [**kafka\_connect\_secret\_providers.\[0\].vault.token**](#kafka_connect_secret_providers.\[0].vault.token)`string`- maxLength: `256`Token used to authenticate with vault and auth method \`token\`. | | [**kafka\_connect\_secret\_providers.\[0\].vault.server\_pem**](#kafka_connect_secret_providers.\[0].vault.server_pem)`string`- maxLength: `4096`PEM encoded certificate of the Vault server. Required if the vault server uses a self-signed certificate. | | [**kafka\_connect\_secret\_providers.\[0\].aws**](#kafka_connect_secret_providers.\[0].aws)`object`- required: `auth_method,region`AWS secret provider configuration | | [**kafka\_connect\_secret\_providers.\[0\].aws.auth\_method**](#kafka_connect_secret_providers.\[0].aws.auth_method)`string`- enum: `credentials`An enumeration. | | [**kafka\_connect\_secret\_providers.\[0\].aws.access\_key**](#kafka_connect_secret_providers.\[0].aws.access_key)`string`- maxLength: `128`Access key used to authenticate with aws | | [**kafka\_connect\_secret\_providers.\[0\].aws.secret\_key**](#kafka_connect_secret_providers.\[0].aws.secret_key)`string`- maxLength: `128`Secret key used to authenticate with aws | | [**kafka\_connect\_secret\_providers.\[0\].aws.region**](#kafka_connect_secret_providers.\[0].aws.region)`string`- maxLength: `64`Region used to lookup secrets with AWS SecretManager | | [**kafka\_connect\_secret\_providers.\[0\].env**](#kafka_connect_secret_providers.\[0].env)`object`- required: `secrets`ENV secret provider configuration | | [**kafka\_connect\_secret\_providers.\[0\].env.secrets**](#kafka_connect_secret_providers.\[0].env.secrets)`object`Key/value map of secrets for ENV secret provider | | [**kafka\_connect\_secret\_providers.\[0\].azure**](#kafka_connect_secret_providers.\[0].azure)`object`- required: `auth_method`Azure KeyVault secret provider configuration | | [**kafka\_connect\_secret\_providers.\[0\].azure.auth\_method**](#kafka_connect_secret_providers.\[0].azure.auth_method)`string`- enum: `credentials`An enumeration. | | [**kafka\_connect\_secret\_providers.\[0\].azure.client\_id**](#kafka_connect_secret_providers.\[0].azure.client_id)`string`- maxLength: `128`Azure client ID for the service principal. | | [**kafka\_connect\_secret\_providers.\[0\].azure.tenant\_id**](#kafka_connect_secret_providers.\[0].azure.tenant_id)`string`- maxLength: `128`Azure tenant ID for the service principal. | | [**kafka\_connect\_secret\_providers.\[0\].azure.secret**](#kafka_connect_secret_providers.\[0].azure.secret)`string`- maxLength: `256`Azure client secret for the service principal. | | [**kafka\_diskless**](#kafka_diskless)`object`- required: `enabled`Kafka Diskless configuration values | | [**kafka\_diskless.enabled**](#kafka_diskless.enabled)`boolean`Whether to enable the Diskless functionality | | [**kafka\_diskless.auto\_diskless\_topic\_regexes**](#kafka_diskless.auto_diskless_topic_regexes)`array[string]`- maxItems: `32`The regexes of topics to auto enable diskless. Topics matching any of the regexes will be created as diskless topics. | | [**inkless**](#inkless)`object`- required: `enabled`Inkless configuration values | | [**inkless.enabled**](#inkless.enabled)`boolean`Whether to enable the Inkless functionality | | [**sasl\_oauthbearer\_allowed\_urls**](#sasl_oauthbearer_allowed_urls)`array[string]`List of allowed URLs for SASL OAUTHBEARER authentication. Only HTTPS URLs are allowed for security reasons. | | [**gcp\_auth\_allowed\_urls**](#gcp_auth_allowed_urls)`array[string]`Allow-list of HTTPS URLs used to validate GCP credential\_source requests for Kafka Connect. | | [**kafka\_rest**](#kafka_rest)`boolean`Enable Kafka-REST service | | [**kafka\_version**](#kafka_version)`string,null`- enum: `3.9,4.1,4.2`Kafka major version | | [**karapace\_version**](#karapace_version)`string,null`- enum: `6.2.1,6.2.2,null`Select a Karapace version for this service, or select Latest to use the latest available version automatically. New versions become available after installation during a maintenance update. | | [**schema\_registry**](#schema_registry)`boolean`Enable Schema-Registry service | | [**kafka\_rest\_authorization**](#kafka_rest_authorization)`boolean`Enable authorization in Kafka-REST service | | [**kafka\_rest\_config**](#kafka_rest_config)`object`Kafka REST configuration | | [**kafka\_rest\_config.producer\_acks**](#kafka_rest_config.producer_acks)`string`- enum: `all,-1,0,1`
- default: `1`The number of acknowledgments the producer requires the leader to have received before considering a request complete. If set to 'all' or '-1', the leader will wait for the full set of in-sync replicas to acknowledge the record. | | [**kafka\_rest\_config.producer\_compression\_type**](#kafka_rest_config.producer_compression_type)`string`- enum: `gzip,snappy,lz4,zstd,none`Specify the default compression type for producers. This configuration accepts the standard compression codecs ('gzip', 'snappy', 'lz4', 'zstd'). It additionally accepts 'none' which is the default and equivalent to no compression. | | [**kafka\_rest\_config.producer\_linger\_ms**](#kafka_rest_config.producer_linger_ms)`integer`- min: `0`
- max: `5000`
- default: `0`Wait for up to the given delay to allow batching records together | | [**kafka\_rest\_config.producer\_max\_request\_size**](#kafka_rest_config.producer_max_request_size)`integer`- min: `0`
- max: `2147483647`
- default: `1048576`The maximum size of a request in bytes. Note that Kafka broker can also cap the record batch size. | | [**kafka\_rest\_config.consumer\_enable\_auto\_commit**](#kafka_rest_config.consumer_enable_auto_commit)`boolean`- default: `true`If true the consumer's offset will be periodically committed to Kafka in the background | | [**kafka\_rest\_config.consumer\_idle\_disconnect\_timeout**](#kafka_rest_config.consumer_idle_disconnect_timeout)`integer`- min: `0`
- max: `2147483647`
- default: `0`Specifies the maximum duration (in seconds) a client can remain idle before it is deleted. If a consumer is inactive, it will exit the consumer group, and its state will be discarded. A value of 0 (default) indicates that the consumer will not be disconnected automatically due to inactivity. | | [**kafka\_rest\_config.consumer\_request\_max\_bytes**](#kafka_rest_config.consumer_request_max_bytes)`integer`- min: `0`
- max: `671088640`
- default: `67108864`Maximum number of bytes in unencoded message keys and values by a single request | | [**kafka\_rest\_config.consumer\_request\_timeout\_ms**](#kafka_rest_config.consumer_request_timeout_ms)`integer`- min: `1000`
- max: `30000`
- enum: `1000,15000,30000`
- default: `1000`The maximum total time to wait for messages for a request if the maximum number of messages has not yet been reached | | [**kafka\_rest\_config.name\_strategy**](#kafka_rest_config.name_strategy)`string`- enum: `topic_name,record_name,topic_record_name`
- default: `topic_name`Name strategy to use when selecting subject for storing schemas | | [**kafka\_rest\_config.name\_strategy\_validation**](#kafka_rest_config.name_strategy_validation)`boolean`- default: `true`If true, validate that given schema is registered under expected subject name by the used name strategy when producing messages. | | [**kafka\_rest\_config.simpleconsumer\_pool\_size\_max**](#kafka_rest_config.simpleconsumer_pool_size_max)`integer`- min: `10`
- max: `250`
- default: `25`Maximum number of SimpleConsumers that can be instantiated per broker | | [**tiered\_storage**](#tiered_storage)`object`Tiered storage configuration | | [**tiered\_storage.enabled**](#tiered_storage.enabled)`boolean`Whether to enable the tiered storage functionality | | [**schema\_registry\_config**](#schema_registry_config)`object`Schema Registry configuration | | [**schema\_registry\_config.topic\_name**](#schema_registry_config.topic_name)`string`- maxLength: `249`The durable single partition topic that acts as the durable log for the data. This topic must be compacted to avoid losing data due to retention policy. Please note that changing this configuration in an existing Schema Registry / Karapace setup leads to previous schemas being inaccessible, data encoded with them potentially unreadable and schema ID sequence put out of order. It's only possible to do the switch while Schema Registry / Karapace is disabled. Defaults to \`\_schemas\`. | | [**schema\_registry\_config.leader\_eligibility**](#schema_registry_config.leader_eligibility)`boolean`If true, Karapace / Schema Registry on the service nodes can participate in leader election. It might be needed to disable this when the schemas topic is replicated to a secondary cluster and Karapace / Schema Registry there must not participate in leader election. Defaults to \`true\`. | | [**schema\_registry\_config.schema\_reader\_strict\_mode**](#schema_registry_config.schema_reader_strict_mode)`boolean`If enabled, causes the Karapace schema-registry service to shutdown when there are invalid schema records in the \`\_schemas\` topic. Defaults to \`false\`. | | [**schema\_registry\_config.retriable\_errors\_silenced**](#schema_registry_config.retriable_errors_silenced)`boolean`If enabled, kafka errors which can be retried or custom errors specified for the service will not be raised, instead, a warning log is emitted. This will denoise issue tracking systems, i.e. sentry. Defaults to \`true\`. | | [**schema\_registry\_config.sasl\_oauthbearer\_authentication\_enabled**](#schema_registry_config.sasl_oauthbearer_authentication_enabled)`boolean`If enabled, the Schema Registry validates OAuth2/OIDC JWT bearer tokens on incoming requests. Requires the OIDC provider settings under the \`kafka\` configuration (\`sasl\_oauthbearer\_jwks\_endpoint\_url\` and related). Defaults to \`false\`. | | [**schema\_registry\_config.sasl\_oauthbearer\_authorization\_enabled**](#schema_registry_config.sasl_oauthbearer_authorization_enabled)`boolean`If enabled, the Schema Registry enforces role-based authorization derived from the JWT roles claim. Enabling this automatically enables \`sasl\_oauthbearer\_authentication\_enabled\` when it is not already enabled, since authorization requires authentication. Defaults to \`false\`. | | [**schema\_registry\_config.sasl\_oauthbearer\_roles\_claim\_path**](#schema_registry_config.sasl_oauthbearer_roles_claim_path)`string`- maxLength: `128`JSON path used to extract the roles claim from the JWT for Schema Registry authorization. Defaults to \`resource\_access.karapace.roles\`. | | [**schema\_registry\_config.sasl\_oauthbearer\_method\_roles**](#schema_registry_config.sasl_oauthbearer_method_roles)`string`- maxLength: `4096`JSON object mapping HTTP methods to the list of roles allowed to perform them on the Schema Registry, provided as a JSON-encoded string. Role names use the \`karapace.\` prefix, e.g. \`karapace.schema:read\`. Defaults to \`{"GET": \["karapace.schema:read", "karapace.subject:read"], "POST": \[], "PUT": \[], "DELETE": \[]}\`. | | [**aiven\_kafka\_topic\_messages**](#aiven_kafka_topic_messages)`boolean`Allow access to read Kafka topic messages in the Aiven Console and REST API. | | [**enable\_ipv6**](#enable_ipv6)`boolean`Register AAAA DNS records for the service, and allow IPv6 packets to service ports | ## Topic parameters[​](#topic-parameters "Direct link to Topic parameters") | | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | [**cleanup\_policy**](#cleanup_policy)`string`- enum: `compact,delete`The retention policy to use on old segments. The default policy ('delete') will discard old segments when their retention time or size limit has been reached. The 'compact' setting will enable log compaction on the topic. The 'compact' setting is not supported for diskless topics. | | [**delete\_retention\_ms**](#delete_retention_ms)`integer`- min: `60000`
- max: `604800000`The amount of time to retain delete tombstone markers for log compacted topics. This setting also gives a bound on the time in which a consumer must complete a read if they begin from offset 0 to ensure that they get a valid snapshot of the final stage (otherwise delete tombstones may be collected before they complete their scan). | | [**max\_message\_bytes**](#max_message_bytes)`integer`- min: `0`
- max: `20971520`The largest record batch size allowed by Kafka (after compression if compression is enabled). If this is increased and there are consumers older than 0.10.2, the consumers' fetch size must also be increased so that the they can fetch record batches this large. In the latest message format version, records are always grouped into batches for efficiency. In previous message format versions, uncompressed records are not grouped into batches and this limit only applies to a single record in that case. | | [**message\_timestamp\_after\_max\_ms**](#message_timestamp_after_max_ms)`integer`- min: `0`
- max: `9223372036854775807`The maximum difference allowed between the timestamp when a broker receives a message and the timestamp specified in the message. If message.timestamp.type=CreateTime, a message will be rejected if the difference in timestamp exceeds this threshold. Applies only for messages with timestamps later than the broker's timestamp. | | [**message\_timestamp\_before\_max\_ms**](#message_timestamp_before_max_ms)`integer`- min: `0`
- max: `9223372036854775807`The maximum difference allowed between the timestamp when a broker receives a message and the timestamp specified in the message. If message.timestamp.type=CreateTime, a message will be rejected if the difference in timestamp exceeds this threshold. Applies only for messages with timestamps earlier than the broker's timestamp. | | [**message\_timestamp\_type**](#message_timestamp_type)`string`- enum: `CreateTime,LogAppendTime`Define whether the timestamp in the message is message create time or log append time. | | [**min\_insync\_replicas**](#min_insync_replicas)`integer`- min: `0`
- max: `2`When a producer sets acks to 'all' (or '-1'), min.insync.replicas specifies the minimum number of replicas that must acknowledge a write for the write to be considered successful. This configuration is not supported for Diskless topics. (Default: 1) | | [**retention\_bytes**](#retention_bytes)`integer`- min: `-1`
- max: `9223372036854775807`This configuration controls the maximum size a partition (which consists of log segments) can grow to before we will discard old log segments to free up space if we are using the 'delete' retention policy. By default there is no size limit only a time limit. Since this limit is enforced at the partition level, multiply it by the number of partitions to compute the topic retention in bytes. | | [**unclean\_leader\_election\_enable**](#unclean_leader_election_enable)`boolean`Indicates whether to enable replicas not in the ISR set to be elected as leader as a last resort, even though doing so may result in data loss. | | [**retention\_ms**](#retention_ms)`integer`- range: `60000` - `9223372036854775807`
- min: `-1`This configuration controls the maximum time we will retain a log before we will discard old log segments to free up space if we are using the 'delete' retention policy. This represents an SLA on how soon consumers must read their data. If set to -1, no time limit is applied. | | [**diskless\_enable**](#diskless_enable)`boolean`Indicates whether diskless should be enabled. (Default: false) | | [**remote\_storage\_enable**](#remote_storage_enable)`boolean`Indicates whether tiered storage should be enabled. If neither diskless.enable or remote.storage.enable are specified then this configuration is automatically set to 'true' when topic is created. (Default: false) | --- # Advanced topic parameters for Standard Kafka View the topic-level configuration options for Standard Kafka. These configurations apply to Standard Kafka on Aiven Cloud. | | | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | [**cleanup\_policy**](#cleanup_policy)`string`- enum: `compact,delete`The retention policy to use on old segments. The default policy ('delete') will discard old segments when their retention time or size limit has been reached. The 'compact' setting will enable log compaction on the topic. The 'compact' setting is not supported for diskless topics. | | [**delete\_retention\_ms**](#delete_retention_ms)`integer`- min: `60000`
- max: `604800000`The amount of time to retain delete tombstone markers for log compacted topics. This setting also gives a bound on the time in which a consumer must complete a read if they begin from offset 0 to ensure that they get a valid snapshot of the final stage (otherwise delete tombstones may be collected before they complete their scan). | | [**max\_message\_bytes**](#max_message_bytes)`integer`- min: `0`
- max: `20971520`The largest record batch size allowed by Kafka (after compression if compression is enabled). If this is increased and there are consumers older than 0.10.2, the consumers' fetch size must also be increased so that the they can fetch record batches this large. In the latest message format version, records are always grouped into batches for efficiency. In previous message format versions, uncompressed records are not grouped into batches and this limit only applies to a single record in that case. | | [**message\_timestamp\_after\_max\_ms**](#message_timestamp_after_max_ms)`integer`- min: `0`
- max: `9223372036854775807`The maximum difference allowed between the timestamp when a broker receives a message and the timestamp specified in the message. If message.timestamp.type=CreateTime, a message will be rejected if the difference in timestamp exceeds this threshold. Applies only for messages with timestamps later than the broker's timestamp. | | [**message\_timestamp\_before\_max\_ms**](#message_timestamp_before_max_ms)`integer`- min: `0`
- max: `9223372036854775807`The maximum difference allowed between the timestamp when a broker receives a message and the timestamp specified in the message. If message.timestamp.type=CreateTime, a message will be rejected if the difference in timestamp exceeds this threshold. Applies only for messages with timestamps earlier than the broker's timestamp. | | [**message\_timestamp\_type**](#message_timestamp_type)`string`- enum: `CreateTime,LogAppendTime`Define whether the timestamp in the message is message create time or log append time. | | [**min\_insync\_replicas**](#min_insync_replicas)`integer`- min: `0`
- max: `2`When a producer sets acks to 'all' (or '-1'), min.insync.replicas specifies the minimum number of replicas that must acknowledge a write for the write to be considered successful. This configuration is not supported for Diskless topics. (Default: 1) | | [**retention\_bytes**](#retention_bytes)`integer`- min: `-1`
- max: `9223372036854775807`This configuration controls the maximum size a partition (which consists of log segments) can grow to before we will discard old log segments to free up space if we are using the 'delete' retention policy. By default there is no size limit only a time limit. Since this limit is enforced at the partition level, multiply it by the number of partitions to compute the topic retention in bytes. | | [**unclean\_leader\_election\_enable**](#unclean_leader_election_enable)`boolean`Indicates whether to enable replicas not in the ISR set to be elected as leader as a last resort, even though doing so may result in data loss. | | [**retention\_ms**](#retention_ms)`integer`- range: `60000` - `9223372036854775807`
- min: `-1`This configuration controls the maximum time we will retain a log before we will discard old log segments to free up space if we are using the 'delete' retention policy. This represents an SLA on how soon consumers must read their data. If set to -1, no time limit is applied. | | [**diskless\_enable**](#diskless_enable)`boolean`Indicates whether diskless should be enabled. (Default: false) | | [**remote\_storage\_enable**](#remote_storage_enable)`boolean`Indicates whether tiered storage should be enabled. If neither diskless.enable or remote.storage.enable are specified then this configuration is automatically set to 'true' when topic is created. (Default: false) | --- # Aiven for Apache Kafka® metrics available via Prometheus Explore common metrics available via Prometheus for your Aiven for Apache Kafka® service. The available metrics depend on whether your service runs in KRaft mode or ZooKeeper mode. ## How to retrieve metrics[​](#how-to-retrieve-metrics "Direct link to How to retrieve metrics") You can retrieve a complete list of metrics from your service by querying the Prometheus endpoint. To do this: 1. Gather the necessary details: * Aiven project certificate: `ca.pem`. To download the CA certificate, see [Download CA certificates](/docs/platform/concepts/tls-ssl-certificates.md#download-ca-certificates). * Prometheus credentials: `:` * Aiven for Apache Kafka hostname: `` * Prometheus port: `` 2. Run the following `curl` command to query the Prometheus endpoint: ``` curl --cacert ca.pem \ --user ':' \ 'https://:/metrics' ``` For more information about setting up Prometheus integration, see [Use Prometheus with Aiven](/docs/platform/howto/integrations/prometheus-metrics.md) ## Host metrics[​](#host-metrics "Direct link to Host metrics") Host metrics provide insights into system-level performance, including CPU, memory, disk, and network usage. ### CPU utilization[​](#cpu-utilization "Direct link to CPU utilization") CPU utilization metrics offer insights into CPU usage. These metrics include time spent on different processes, system load, and overall uptime. | Metric | Description | | ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | | `cpu_usage_guest` | CPU time spent running a virtual CPU for guest operating systems | | `cpu_usage_guest_nice` | CPU time running low-priority virtual CPUs for guest operating systems; interrupted by higher-priority tasks and measured in hundredths of a second | | `cpu_usage_idle` | Time the CPU spends doing nothing | | `cpu_usage_iowait` | Time waiting for I/O to complete | | `cpu_usage_irq` | Time servicing interrupts | | `cpu_usage_nice` | Time running user-niced processes | | `cpu_usage_softirq` | Time servicing softirqs | | `cpu_usage_steal` | Time spent in other operating systems when running in a virtualized environment | | `cpu_usage_system` | Time spent running system processes | | `cpu_usage_user` | Time spent running user processes | | `system_load1` | System load average for the last minute | | `system_load15` | System load average for the last 15 minutes | | `system_load5` | System load average for the last 5 minutes | | `system_n_cpus` | Number of CPU cores available | | `system_n_users` | Number of users logged in | | `system_uptime` | Time for which the system has been up and running | ### Disk space utilization[​](#disk-space-utilization "Direct link to Disk space utilization") Disk space utilization metrics provide a snapshot of disk usage. These metrics include information about free and used disk space, as well as `inode` usage and total disk capacity. | Metric | Description | | ------------------- | ----------------------------- | | `disk_free` | Amount of free disk space | | `disk_inodes_free` | Number of free inodes | | `disk_inodes_total` | Total number of inodes | | `disk_inodes_used` | Number of used inodes | | `disk_total` | Total disk space | | `disk_used` | Amount of used disk space | | `disk_used_percent` | Percentage of disk space used | ### Disk input and output[​](#disk-input-and-output "Direct link to Disk input and output") Metrics such as `diskio_io_time` and `diskio_iops_in_progress` provide insights into disk I/O operations. These metrics cover read/write operations, the duration of these operations, and the number of bytes read/written. | Metric | Description | | ------------------------- | ------------------------------------------------------------------------------- | | `diskio_io_time` | Total time spent on I/O operations | | `diskio_iops_in_progress` | Number of I/O operations currently in progress | | `diskio_merged_reads` | Number of read operations that were merged | | `diskio_merged_writes` | Number of write operations that were merged | | `diskio_read_bytes` | Total bytes read from disk | | `diskio_read_time` | Total time spent on read operations | | `diskio_reads` | Total number of read operations | | `diskio_weighted_io_time` | Weighted time spent on I/O operations, considering their duration and intensity | | `diskio_write_bytes` | Total bytes written to disk | | `diskio_write_time` | Total time spent on write operations | | `diskio_writes` | Total number of write operations | ### Generic memory[​](#generic-memory "Direct link to Generic memory") The following metrics, including `mem_active` and `mem_available`, provide insights into your system's memory usage. | Metric | Description | | ----------------------- | ------------------------------------------------------------- | | `mem_active` | Amount of actively used memory | | `mem_available` | Amount of available memory | | `mem_available_percent` | Percentage of available memory | | `mem_buffered` | Amount of memory used for buffering I/O | | `mem_cached` | Amount of memory used for caching | | `mem_commit_limit` | Maximum amount of memory that can be committed | | `mem_committed_as` | Total amount of committed memory | | `mem_dirty` | Amount of memory waiting to be written to disk | | `mem_free` | Amount of free memory | | `mem_high_free` | Amount of free memory in the high memory zone | | `mem_high_total` | Total amount of memory in the high memory zone | | `mem_huge_pages_free` | Number of free huge pages | | `mem_huge_page_size` | Size of huge pages | | `mem_huge_pages_total` | Total number of huge pages | | `mem_inactive` | Amount of inactive memory | | `mem_low_free` | Amount of free memory in the low memory zone | | `mem_low_total` | Total amount of memory in the low memory zone | | `mem_mapped` | Amount of memory mapped into the process's address space | | `mem_page_tables` | Amount of memory used by page tables | | `mem_shared` | Amount of memory shared between processes | | `mem_slab` | Amount of memory used by the kernel for data structure caches | | `mem_swap_cached` | Amount of swap memory cached | | `mem_swap_free` | Amount of free swap memory | | `mem_swap_total` | Total amount of swap memory | | `mem_total` | Total amount of memory | | `mem_used` | Amount of used memory | | `mem_used_percent` | Percentage of used memory | | `mem_vmalloc_chunk` | Largest contiguous block of vmalloc memory available | | `mem_vmalloc_total` | Total amount of vmalloc memory | | `mem_vmalloc_used` | Amount of used vmalloc memory | | `mem_wired` | Amount of wired memory | | `mem_write_back` | Amount of memory being written back to disk | | `mem_write_back_tmp` | Amount of temporary memory being written back to disk | ### Network[​](#network "Direct link to Network") The following metrics, including `net_bytes_recv` and `net_packets_sent`, provide insights into your system's network operations. | Metric | Description | | ----------------------------- | ----------------------------------------------------------------- | | `net_bytes_recv` | Total bytes received on the network interfaces | | `net_bytes_sent` | Total bytes sent on the network interfaces | | `net_drop_in` | Incoming packets dropped | | `net_drop_out` | Outgoing packets dropped | | `net_err_in` | Incoming packets with errors | | `net_err_out` | Outgoing packets with errors | | `net_icmp_inaddrmaskreps` | Number of ICMP address mask replies received | | `net_icmp_inaddrmasks` | Number of ICMP address mask requests received | | `net_icmp_incsumerrors` | Number of ICMP checksum errors | | `net_icmp_indestunreachs` | Number of ICMP destination unreachable messages received | | `net_icmp_inechoreps` | Number of ICMP echo replies received | | `net_icmp_inechos` | Number of ICMP echo requests received | | `net_icmp_inerrors` | Number of ICMP messages received with errors | | `net_icmp_inmsgs` | Total number of ICMP messages received | | `net_icmp_inparmprobs` | Number of ICMP parameter problem messages received | | `net_icmp_inredirects` | Number of ICMP redirect messages received | | `net_icmp_insrcquenchs` | Number of ICMP source quench messages received | | `net_icmp_intimeexcds` | Number of ICMP time exceeded messages received | | `net_icmp_intimestampreps` | Number of ICMP timestamp reply messages received | | `net_icmp_intimestamps` | Number of ICMP timestamp request messages received | | `net_icmpmsg_intype3` | Number of ICMP type 3 (destination unreachable) messages received | | `net_icmpmsg_intype8` | Number of ICMP type 8 (echo request) messages received | | `net_icmpmsg_outtype0` | Number of ICMP type 0 (echo reply) messages sent | | `net_icmpmsg_outtype3` | Number of ICMP type 3 (destination unreachable) messages sent | | `net_icmp_outaddrmaskreps` | Number of ICMP address mask reply messages sent | | `net_icmp_outaddrmasks` | Number of ICMP address mask request messages sent | | `net_icmp_outdestunreachs` | Number of ICMP destination unreachable messages sent | | `net_icmp_outechoreps` | Number of ICMP echo reply messages sent | | `net_icmp_outechos` | Number of ICMP echo request messages sent | | `net_icmp_outerrors` | Number of ICMP messages sent with errors | | `net_icmp_outmsgs` | Total number of ICMP messages sent | | `net_icmp_outparmprobs` | Number of ICMP parameter problem messages sent | | `net_icmp_outredirects` | Number of ICMP redirect messages sent | | `net_icmp_outsrcquenchs` | Number of ICMP source quench messages sent | | `net_icmp_outtimeexcds` | Number of ICMP time exceeded messages sent | | `net_icmp_outtimestampreps` | Number of ICMP timestamp reply messages sent | | `net_icmp_outtimestamps` | Number of ICMP timestamp request messages sent | | `net_icmp_outratelimitglobal` | Number of globally rate-limited ICMP messages sent | | `net_icmp_outratelimithost` | Number of ICMP messages rate-limited per host | | `net_ip_defaultttl` | Default time-to-live for IP packets | | `net_ip_forwarding` | Indicates if IP forwarding is enabled | | `net_ip_forwdatagrams` | Number of forwarded IP datagrams | | `net_ip_fragcreates` | Number of IP fragments created | | `net_ip_fragfails` | Number of failed IP fragmentations | | `net_ip_fragoks` | Number of successful IP fragmentations | | `net_ip_inaddrerrors` | Number of incoming IP packets with address errors | | `net_ip_indelivers` | Number of incoming IP packets delivered to higher layers | | `net_ip_indiscards` | Number of incoming IP packets discarded | | `net_ip_inhdrerrors` | Number of incoming IP packets with header errors | | `net_ip_inreceives` | Total number of incoming IP packets received | | `net_ip_inunknownprotos` | Number of incoming IP packets with unknown protocols | | `net_ip_outdiscards` | Number of outgoing IP packets discarded | | `net_ip_outnoroutes` | Number of outgoing IP packets with no route available | | `net_ip_outrequests` | Total number of outgoing IP packets requested to be sent | | `net_ip_outtransmits` | Number of IP packets transmitted successfully | | `net_ip_reasmfails` | Number of failed IP reassembly attempts | | `net_ip_reasmoks` | Number of successful IP reassembly attempts | | `net_ip_reasmreqds` | Number of IP fragments received needing reassembly | | `net_ip_reasmtimeout` | Number of IP reassembly timeouts | | `net_packets_recv` | Total number of packets received on the network interfaces | | `net_packets_sent` | Total number of packets sent on the network interfaces | | `netstat_tcp_close` | Number of TCP connections in the CLOSE state | | `netstat_tcp_close_wait` | Number of TCP connections in the CLOSE\_WAIT state | | `netstat_tcp_closing` | Number of TCP connections in the CLOSING state | | `netstat_tcp_established` | Number of TCP connections in the ESTABLISHED state | | `netstat_tcp_fin_wait1` | Number of TCP connections in the FIN\_WAIT\_1 state | | `netstat_tcp_fin_wait2` | Number of TCP connections in the FIN\_WAIT\_2 state | | `netstat_tcp_last_ack` | Number of TCP connections in the LAST\_ACK state | | `netstat_tcp_listen` | Number of TCP connections in the LISTEN state | | `netstat_tcp_none` | Number of TCP connections in the NONE state | | `netstat_tcp_syn_recv` | Number of TCP connections in the SYN\_RECV state | | `netstat_tcp_syn_sent` | Number of TCP connections in the SYN\_SENT state | | `netstat_tcp_time_wait` | Number of TCP connections in the TIME\_WAIT state | | `netstat_udp_socket` | Number of UDP sockets | | `net_tcp_activeopens` | Number of active TCP open connections | | `net_tcp_attemptfails` | Number of failed TCP connection attempts | | `net_tcp_currestab` | Number of currently established TCP connections | | `net_tcp_estabresets` | Number of established TCP connections reset | | `net_tcp_incsumerrors` | Number of TCP checksum errors in incoming packets | | `net_tcp_inerrs` | Number of incoming TCP packets with errors | | `net_tcp_insegs` | Number of TCP segments received | | `net_tcp_maxconn` | Maximum number of TCP connections supported | | `net_tcp_outrsts` | Number of TCP reset packets sent | | `net_tcp_outsegs` | Number of TCP segments sent | | `net_tcp_passiveopens` | Number of passive TCP open connections | | `net_tcp_retranssegs` | Number of TCP segments retransmitted | | `net_tcp_rtoalgorithm` | TCP retransmission timeout algorithm | | `net_tcp_rtomax` | Maximum TCP retransmission timeout | | `net_tcp_rtomin` | Minimum TCP retransmission timeout | | `net_udp_ignoredmulti` | Number of UDP multicast packets ignored | | `net_udp_incsumerrors` | Number of UDP checksum errors in incoming packets | | `net_udp_indatagrams` | Number of UDP datagrams received | | `net_udp_inerrors` | Number of incoming UDP packets with errors | | `net_udp_memerrors` | Number of UDP packets dropped due to memory errors | | `net_udplite_ignoredmulti` | Number of UDP-Lite multicast packets ignored | | `net_udplite_incsumerrors` | Number of UDP-Lite checksum errors in incoming packets | | `net_udplite_indatagrams` | Number of UDP-Lite datagrams received | | `net_udplite_inerrors` | Number of incoming UDP-Lite packets with errors | | `net_udplite_memerrors` | Number of UDP-L | ### Kernel[​](#kernel "Direct link to Kernel") The metrics listed below, such as `kernel_boot_time` and `kernel_context_switches`, provide insights into the operations of your system's kernel. | Metric | Description | | ------------------------- | ----------------------------------------------------------- | | `kernel_boot_time` | Time at which the system was last booted | | `kernel_context_switches` | Number of context switches that have occurred in the kernel | | `kernel_entropy_avail` | Amount of available entropy in the kernel's entropy pool | | `kernel_interrupts` | Number of interrupts that have occurred | | `kernel_processes_forked` | Number of processes that have been forked | ### Process[​](#process "Direct link to Process") Metrics such as `processes_running` and `processes_zombies` provide insights into the management of the system's processes. | Metric | Description | | ------------------------- | ------------------------------------------------------------------------ | | `processes_blocked` | Number of processes that are blocked | | `processes_dead` | Number of processes that have terminated | | `processes_idle` | Number of processes that are idle | | `processes_paging` | Number of processes that are paging | | `processes_running` | Number of processes currently running | | `processes_sleeping` | Number of processes that are sleeping | | `processes_stopped` | Number of processes that are stopped | | `processes_total` | Total number of processes | | `processes_total_threads` | Total number of threads across all processes | | `processes_unknown` | Number of processes in an unknown state | | `processes_zombies` | Number of zombie processes (terminated but not reaped by parent process) | ### Swap usage[​](#swap-usage "Direct link to Swap usage") Metrics such as `swap_free` and `swap_used` provide insights into the usage of the system's swap memory. | Metric | Description | | ------------------- | ----------------------------------- | | `swap_free` | Amount of free swap memory | | `swap_in` | Amount of data swapped in from disk | | `swap_out` | Amount of data swapped out to disk | | `swap_total` | Total amount of swap memory | | `swap_used` | Amount of used swap memory | | `swap_used_percent` | Percentage of swap memory used | ## Aiven for Apache Kafka-specific metrics[​](#aiven-for-apache-kafka-specific-metrics "Direct link to Aiven for Apache Kafka-specific metrics") Metrics specific to Apache Kafka provide detailed insights into the health and performance of your Kafka clusters, including broker, controller, and topic-level metrics. ## Garbage collector `MXBean`[​](#garbage-collector-mxbean "Direct link to garbage-collector-mxbean") Metrics associated with the `java_lang_GarbageCollector` provide insights into the JVM's garbage collection process. These metrics include the collection count and the duration of collections. | Metric | Description | | ---------------------------------------------------------------- | --------------------------------------------------------------------------- | | `java_lang_GarbageCollector_G1_Young_Generation_CollectionCount` | Returns the total number of collections that have occurred | | `java_lang_GarbageCollector_G1_Young_Generation_CollectionTime` | Returns the approximate accumulated collection elapsed time in milliseconds | | `java_lang_GarbageCollector_G1_Young_Generation_duration` | Duration of G1 Young Generation garbage collections | ## Memory Usage[​](#memory-usage "Direct link to Memory Usage") Metrics starting with `java_lang_Memory` provide insights into the JVM's memory usage, including committed memory, initial memory, max memory, and used memory. | Metric | Description | | ------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | | `java_lang_Memory_committed` | Returns the amount of memory in bytes that is committed for the Java virtual machine to use | | `java_lang_Memory_init` | Returns the amount of memory in bytes that the Java virtual machine initially requests from the operating system for memory management. | | `java_lang_Memory_max` | Returns the maximum amount of memory in bytes that can be used for memory management | | `java_lang_Memory_used` | Returns the amount of used memory in bytes. | | `java_lang_Memory_ObjectPendingFinalizationCount` | Number of objects pending finalization | ## Apache Kafka Connect[​](#apache-kafka-connect "Direct link to Apache Kafka Connect") For a comprehensive list of Apache Kafka Connect metrics exposed through Prometheus, see [Apache Kafka® Connect available via Prometheus](/docs/products/kafka/kafka-connect/reference/connect-metrics-prometheus.md). ## Schema Registry metrics[​](#schema-registry-metrics "Direct link to Schema Registry metrics") When [Karapace Schema Registry](/docs/products/kafka/karapace/howto/enable-karapace.md) is enabled for an Aiven for Apache Kafka® service, metrics from the Schema Registry instance are exposed through the Prometheus metrics endpoint. Use the same [Prometheus integration](/docs/platform/howto/integrations/prometheus-metrics.md) as for Kafka metrics. Metric names use the `karapace_` prefix and cover request volume, latency, and concurrency. These metrics are for the Schema Registry only, not for the Karapace REST proxy. | Metric | Description | | ----------------------------------------- | ------------------------------------------------------------------------------------------------------------------- | | `karapace_http_requests_total` | Total Schema Registry HTTP requests. Use `rate()` for requests per second. | | `karapace_http_requests_duration_seconds` | Schema Registry request duration. Prometheus exposes `_sum` and `_count`; use both for average duration in seconds. | | `karapace_http_requests_in_progress` | Schema Registry HTTP requests currently in progress. | note Schema Registry metrics appear only when Schema Registry is enabled. They are returned with Kafka metrics from the same Prometheus endpoint. ## Apache Kafka broker metrics[​](#apache-kafka-broker-metrics "Direct link to Apache Kafka broker metrics") Apache Kafka brokers expose metrics that provide insights into the health and performance of the Apache Kafka cluster. Find detailed descriptions of these metrics, see the [monitoring section of the Apache Kafka documentation](https://kafka.apache.org/documentation/#monitoring). ### Metric types[​](#metric-types "Direct link to Metric types") #### Cumulative counters (`_count`)[​](#cumulative-counters-_count "Direct link to cumulative-counters-_count") Metrics with a `_count` suffix are cumulative counters. They track the total number of occurrences for a specific event since the broker started. **Example:** `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_Count`: Total number of leader elections that have occurred in the controller. #### Rate counters (`perSec`)[​](#rate-counters-persec "Direct link to rate-counters-persec") Metrics with a `perSec` suffix in their name are also cumulative counters. They track the total number of events per second, not the current rate. **Example:** `kafka_server_BrokerTopicMetrics_MessagesInPerSec_Count`: Total number of incoming messages received by the broker. note To calculate the rate of change for these `_Count` metrics, you can use functions such as `rate()` in PromQL. ### Apache Kafka controller metrics[​](#apache-kafka-controller-metrics "Direct link to Apache Kafka controller metrics") Apache Kafka offers a range of metrics to help you assess the performance and health of your Apache Kafka controller. * **Percentile Metrics**: Metrics like `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_XthPercentile` (where X can be 50th, 75th, 95th, etc.) show the time taken for leader elections to complete at various percentiles. This helps in understanding the distribution of leader election times. * **Interval Metrics**: Metrics ending with `FifteenMinuteRate`, `FiveMinuteRate`, following `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_`, show the rate of leader elections over different time intervals. * **Statistical Metrics**: Metrics ending with `Max`, `Mean`, `Min`, `StdDev`, following `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_`, provide statistical measures about the leader election times. * **Controller State Metrics**: Metrics starting with `kafka_controller_KafkaController_` give insights into the state of the Kafka controller, such as the number of active brokers, offline partitions, and replicas to delete. #### ZooKeeper mode-only metrics[​](#zookeeper-mode-only-metrics "Direct link to ZooKeeper mode-only metrics") Apache Kafka requires a separate ZooKeeper process that, for example, stores metadata. The following metrics are only available when running Kafka in ZooKeeper mode: | Metric | Description | | -------------------------------------------------------------- | ---------------------------- | | `kafka_controller_KafkaController_ActiveControllerCount_Value` | Number of active controllers | #### KRaft mode and metrics changes[​](#kraft-mode-and-metrics-changes "Direct link to KRaft mode and metrics changes") Aiven for Apache Kafka services running Apache Kafka 3.9 and later use KRaft mode, which replaces ZooKeeper for metadata and controller management. In KRaft mode, controller metrics are emitted from a dedicated controller process instead of the broker. Aiven exposes a limited subset of these metrics. The KRaft controller is fully managed by Aiven. Controller and Raft metrics are not exposed. ##### Controller metrics available in KRaft mode[​](#controller-metrics-available-in-kraft-mode "Direct link to Controller metrics available in KRaft mode") The following controller metrics are available in KRaft mode: | Metric | Description | | --------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- | | `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_50thPercentile` | Time taken for leader elections to complete at the 50th percentile | | `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_75thPercentile` | Time taken for leader elections to complete at the 75th percentile | | `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_95thPercentile` | Time taken for leader elections to complete at the 95th percentile | | `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_98thPercentile` | Time taken for leader elections to complete at the 98th percentile | | `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_99thPercentile` | Time taken for leader elections to complete at the 99th percentile | | `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_999thPercentile` | Time taken for leader elections to complete at the 99.9th percentile | | `kafka_controller_ControllerStats_ElectionFromEligibleLeaderReplicasPerSec_Count` | Rate of leader elections successfully completed from eligible leader replicas | | `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_Count` | The total number of leader elections | | `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_FifteenMinuteRate` | Rate of leader elections over the last 15 minutes | | `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_FiveMinuteRate` | Rate of leader elections over the last 5 minutes | | `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_Max` | Maximum time taken for a leader election | | `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_Mean` | Mean time taken for leader elections | | `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_MeanRate` | Mean rate of leader elections | | `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_Min` | Minimum time taken for a leader election | | `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_OneMinuteRate` | Rate of leader elections over the last minute | | `kafka_controller_ControllerStats_LeaderElectionRateAndTimeMs_StdDev` | Standard deviation of leader election times | | `kafka_controller_ControllerStats_UncleanLeaderElectionsPerSec_Count` | Number of unclean leader elections. Unclean leader elections can lead to data loss | | `kafka_controller_KafkaController_ActiveBrokerCount_Value` | Number of active brokers | | `kafka_controller_KafkaController_GlobalTopicCount_Value` | Number of topics | | `kafka_controller_KafkaController_GlobalPartitionCount_Value` | Number of partitions | | `kafka_controller_KafkaController_TopicsToDeleteCount_Value` | Number of topics to delete | | `kafka_controller_KafkaController_ReplicasToDeleteCount_Value` | Number of replicas to delete | | `kafka_controller_KafkaController_OfflinePartitionsCount_Value` | Number of offline partitions | ##### Controller metrics not available in KRaft mode[​](#controller-metrics-not-available-in-kraft-mode "Direct link to Controller metrics not available in KRaft mode") The following controller metrics are **not** available in KRaft mode: | Metric | Description | | --------------------------------------------------------------------------- | --------------------------------------------------- | | `kafka_controller_KafkaController_PreferredReplicaImbalanceCount_Value` | Number of preferred replica imbalances | | `kafka_controller_KafkaController_TopicsIneligibleToDeleteCount_Value` | Number of topics ineligible for deletion | | `kafka_controller_KafkaController_ReplicasIneligibleToDeleteCount_Value` | Number of replicas ineligible for deletion | | `kafka_controller_KafkaController_FencedBrokerCount_Value` | Number of fenced brokers | | `kafka_controller_KafkaController_NewActiveControllersCount_Value` | Number of new active controllers | | `kafka_controller_KafkaController_TimedOutBrokerHeartbeatCount_Value` | Number of timed-out broker heartbeats | | `kafka_controller_KafkaController_ControllerState_Value` | Controller state | | `kafka_controller_KafkaController_MetadataErrorCount_Value` | Number of metadata errors | | `kafka_controller_KafkaController_EventQueueOperationsStartedCount_Value` | Number of event queue operations started | | `kafka_controller_KafkaController_EventQueueOperationsTimedOutCount_Value` | Number of event queue operations timed out | | `kafka_controller_KafkaController_LastAppliedRecordOffset_Value` | Last applied record offset | | `kafka_controller_KafkaController_LastCommittedRecordOffset_Value` | Last committed record offset | | `kafka_controller_KafkaController_LastAppliedRecordTimestamp_Value` | Last applied record timestamp | | `kafka_controller_KafkaController_LastAppliedRecordLagMs_Value` | Last applied record lag in ms | | `kafka_controller_ControllerEventManager_EventQueueProcessingTimeMs_Value` | Event queue processing time in ms | | `kafka_controller_ControllerEventManager_EventQueueSize_Value` | Event queue size | | `kafka_controller_ControllerEventManager_EventQueueTimeMs_Value` | Event queue time in ms | | `kafka_controller_ControllerChannelManager_TotalQueueSize_Value` | Total queue size | | `kafka_controller_ControllerChannelManager_QueueSize_Value` | Queue size per broker | | `kafka_controller_ControllerChannelManager_RequestRateAndQueueTimeMs_Value` | Request rate and queue time per broker | | `kafka_controller_KafkaController_MigratingZkBrokerCount_Value` | Number of brokers migrating from ZooKeeper to KRaft | | `kafka_controller_KafkaController_ZkMigrationState_Value` | ZooKeeper migration state | ### `Jolokia` collector collect time[​](#jolokia-collector-collect-time "Direct link to jolokia-collector-collect-time") Jolokia is a JMX-HTTP bridge that provides an alternative to native JMX access. The following metric provides insights into the time taken by the Jolokia collector to collect metrics. | Metric | Description | | -------------------------------------- | --------------------------------------------------------------------- | | `kafka_jolokia_collector_collect_time` | Represents the time taken by the Jolokia collector to collect metrics | ### Apache Kafka log[​](#apache-kafka-log "Direct link to Apache Kafka log") Apache Kafka provides a variety of metrics that offer insights into its operation. These metrics are useful for understanding the operation of the log cleaner and log flush operations. #### Log cleaner metrics[​](#log-cleaner-metrics "Direct link to Log cleaner metrics") These metrics provide insights into the log cleaner's operation, which helps in compacting the Apache Kafka logs. | Metric | Description | | ---------------------------------------------------------- | ------------------------------------------------------------- | | `kafka_log_LogCleaner_cleaner_recopy_percent_Value` | Percentage of log segments that were recopied during cleaning | | `kafka_log_LogCleanerManager_time_since_last_run_ms_Value` | Time in milliseconds since the last log cleaner run | | `kafka_log_LogCleaner_max_clean_time_secs_Value` | Maximum time in seconds taken for a log cleaning operation | #### Log flush rate metrics[​](#log-flush-rate-metrics "Direct link to Log flush rate metrics") Metrics like `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_XthPercentile` provide the time taken to flush logs at various percentiles. These metrics offer insights into log flush operations, ensuring that the system writes data from memory to disk. They also indicate the time required to flush logs at different percentiles. | Metric | Description | | ----------------------------------------------------------------- | ----------------------------------------------------- | | `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_50thPercentile` | Time taken to flush logs at the 50th percentile | | `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_75thPercentile` | Time taken to flush logs at the 75th percentile | | `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_95thPercentile` | Time taken to flush logs at the 95th percentile | | `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_98thPercentile` | Time taken to flush logs at the 98th percentile | | `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_99thPercentile` | Time taken to flush logs at the 99th percentile | | `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_999thPercentile` | Time taken to flush logs at the 99.9th percentile | | `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_Count` | Total number of log flush operations | | `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_FifteenMinuteRate` | Rate of log flush operations over the last 15 minutes | | `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_FiveMinuteRate` | Rate of log flush operations over the last 5 minutes | | `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_Max` | Maximum time taken for a log flush operation | | `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_Mean` | Mean time taken for log flush operations | | `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_MeanRate` | Mean rate of log flush operations | | `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_Min` | Minimum time taken for a log flush operation | | `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_OneMinuteRate` | Rate of log flush operations over the last minute | | `kafka_log_LogFlushStats_LogFlushRateAndTimeMs_StdDev` | Standard deviation of log flush times | #### Log metrics[​](#log-metrics "Direct link to Log metrics") These metrics provide general information about log sizes and offsets. | Metric | Description | | ------------------------------------ | ----------------------- | | `kafka_log_Log_LogEndOffset_Value` | End offset of the log | | `kafka_log_Log_LogStartOffset_Value` | Start offset of the log | | `kafka_log_Log_Size_Value` | Size of the log | ### Apache Kafka network[​](#apache-kafka-network "Direct link to Apache Kafka network") Apache Kafka provides several metrics, such as `kafka_network_RequestMetrics_RequestsPerSec_Count` and `kafka_network_RequestMetrics_TotalTimeMs_Mean`, to monitor the performance and health of network requests made to the Apache Kafka brokers. | Metric | Description | | ----------------------------------------------------------------- | ------------------------------------------------------------------------ | | `kafka_network_RequestChannel_RequestQueueSize_Value` | Size of the request queue | | `kafka_network_RequestChannel_ResponseQueueSize_Value` | Size of the response queue | | `kafka_network_RequestMetrics_RequestsPerSec_Count` | Total number of requests per second. | | `kafka_network_RequestMetrics_TotalTimeMs_95thPercentile` | Total time for requests at the 95th percentile | | `kafka_network_RequestMetrics_TotalTimeMs_99thPercentile` | 99th percentile of total time taken to process incoming network requests | | `kafka_network_RequestMetrics_TotalTimeMs_Count` | Total number of requests | | `kafka_network_RequestMetrics_TotalTimeMs_Mean` | Mean total time for requests | | `kafka_network_SocketServer_NetworkProcessorAvgIdlePercent_Value` | Average idle percentage of the network processor | ### Apache Kafka server[​](#apache-kafka-server "Direct link to Apache Kafka server") Apache Kafka provides a range of metrics that help monitor the server's performance and health. * **Topic metrics**: `BrokerTopicMetrics` offer insights into various operations related to topics, such as bytes in/out and failed fetch/produce requests. * **Replica metrics**: `kafka_server_ReplicaManager_LeaderCount_Value` provides insights into the state of replicas within the Apache Kafka cluster. The `topic` tag is crucial in these metrics. If you don't specify it, the system displays a combined rate for all topics, along with the rate for each individual topic. To view rates for specific topics, use the `topic` tag. To exclude the combined rate for all topics and only list metrics for individual topics, filter with `topic!=""`. | Metric | Description | | --------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `kafka_server_BrokerTopicMetrics_BytesInPerSec_Count` | Byte in (from the clients) rate per topic. Omitting `topic=(...)` will yield the all-topic rate | | `kafka_server_BrokerTopicMetrics_BytesOutPerSec_Count` | Byte out (to the clients) rate per topic. Omitting `topic=(...)` will yield the all-topic rate | | `kafka_server_BrokerTopicMetrics_BytesRejectedPerSec_Count` | Rejected byte rate per topic due to the record batch size being greater than `max.message.bytes` configuration. Omitting `topic=(...)` will yield the all-topic rate | | `kafka_server_BrokerTopicMetrics_FailedFetchRequestsPerSec_Count` | Failed fetch request (from clients or followers) rate per topic. Omitting `topic=(...)` will yield the all-topic rate | | `kafka_server_BrokerTopicMetrics_FailedProduceRequestsPerSec_Count` | Failed produce request rate per topic. Omitting `topic=(...)` will yield the all-topic rate | | `kafka_server_BrokerTopicMetrics_FetchMessageConversionsPerSec_Count` | Message format conversion rate for produce or fetch requests per topic. Omitting `topic=(...)` will yield the all-topic rate. | | `kafka_server_BrokerTopicMetrics_MessagesInPerSec_Count` | Incoming message rate per topic. Omitting `topic=(...)` will yield the all-topic rate | | `kafka_server_BrokerTopicMetrics_ProduceMessageConversionsPerSec_Count` | Message format conversion rate for produce or fetch requests per topic. Omitting `topic=(...)` will yield the all-topic rate | | `kafka_server_BrokerTopicMetrics_ReassignmentBytesInPerSec_Count` | Incoming byte rate of reassignment traffic. | | `kafka_server_BrokerTopicMetrics_ReassignmentBytesOutPerSec_Count` | Outgoing byte rate of reassignment traffic. | | `kafka_server_BrokerTopicMetrics_ReplicationBytesInPerSec_Count` | Byte in (from other brokers) rate per topic. Omitting `topic=(...)` will yield the all-topic rate | | `kafka_server_BrokerTopicMetrics_ReplicationBytesOutPerSec_Count` | Byte out (to other brokers) rate per topic. Omitting `topic=(...)` will yield the all-topic rate | | `kafka_server_BrokerTopicMetrics_TotalFetchRequestsPerSec_Count` | Fetch request (from clients or followers) rate per topic. Omitting `topic=(...)` will yield the all-topic rate | | `kafka_server_BrokerTopicMetrics_TotalProduceRequestsPerSec_Count` | Total number of produce requests per second. This metric is collected per host and not per topic | | `kafka_server_DelayedOperationPurgatory_NumDelayedOperations_Value` | Number of delayed operations in purgatory. | | `kafka_server_DelayedOperationPurgatory_PurgatorySize_Value` | Size of the purgatory queue. | | `kafka_server_FetchRequestPurgatory_PurgatorySize_Value` | Current number of fetch requests waiting in purgatory (for example, waiting for enough data to satisfy `fetch.min.bytes`) | | `kafka_server_ProducerRequestPurgatory_PurgatorySize_Value` | Current number of produce requests waiting in purgatory (for example, waiting for acknowledgments from followers) | | `kafka_server_KafkaRequestHandlerPool_RequestHandlerAvgIdlePercent_OneMinuteRate` | Average idle percentage of request handlers over the last minute | | `kafka_server_KafkaServer_BrokerState_Value` | State of the broker | | `kafka_server_ReplicaManager_IsrExpandsPerSec_Count` | Number of ISR expansions per second | | `kafka_server_ReplicaManager_IsrShrinksPerSec_Count` | Number of ISR shrinks per second | | `kafka_server_ReplicaManager_LeaderCount_Value` | Number of leader replicas | | `kafka_server_ReplicaManager_PartitionCount_Value` | Number of partitions | | `kafka_server_ReplicaManager_UnderMinIsrPartitionCount_Value` | Number of partitions under the minimum ISR | | `kafka_server_ReplicaManager_UnderReplicatedPartitions_Value` | Number of under-replicated partitions | | `kafka_server_ReplicaFetcherManager_MaxLag_Value` | Maximum message lag measured across all active replica fetcher threads | | `kafka_server_GroupMetadataManager_NumGroups` | Total number of groups (consumer, streams, and so on) managed by the broker's group metadata manager | | `kafka_server_GroupMetadataManager_NumGroupsPreparingRebalance` | Number of groups currently in the preparing-to-rebalance state | | `kafka_server_GroupMetadataManager_NumOffsets` | Total number of consumer offsets currently tracked by the group metadata manager | | `kafka_server_group_coordinator_metrics_batch_flush_rate` | Rate at which offset commit batches are flushed to the `__consumer_offsets` topic | | `kafka_server_group_coordinator_metrics_batch_flush_time_ms_max` | Maximum time taken to flush offset commit batches | | `kafka_server_group_coordinator_metrics_batch_linger_time_ms_max` | Maximum time offset commits linger in memory before being batched and flushed | | `kafka_server_group_coordinator_metrics_consumer_group_count` | Number of active consumer groups handled by this coordinator | | `kafka_server_group_coordinator_metrics_consumer_group_rebalance_count` | Total count of rebalances executed for consumer groups | | `kafka_server_group_coordinator_metrics_consumer_group_rebalance_rate` | Rate of rebalances occurring across consumer groups | | `kafka_server_group_coordinator_metrics_event_processing_time_ms_max` | Maximum time taken to process an event in the group coordinator | | `kafka_server_group_coordinator_metrics_event_purgatory_time_ms_max` | Maximum time coordinator events spend waiting in purgatory before processing | | `kafka_server_group_coordinator_metrics_event_queue_size` | Current number of events waiting in the group coordinator's event queue | | `kafka_server_group_coordinator_metrics_event_queue_time_ms_max` | Maximum time events spend enqueued before processing begins | | `kafka_server_group_coordinator_metrics_group_count` | Total number of all types of groups handled by this coordinator | | `kafka_server_group_coordinator_metrics_group_completed_rebalance_count` | Number of completed group rebalances | | `kafka_server_group_coordinator_metrics_group_completed_rebalance_rate` | Rate of completed group rebalances | | `kafka_server_group_coordinator_metrics_offset_commit_count` | Number of offset commits | | `kafka_server_group_coordinator_metrics_offset_commit_rate` | Rate of offset commits | | `kafka_server_group_coordinator_metrics_offset_deletion_count` | Number of offset deletions | | `kafka_server_group_coordinator_metrics_offset_deletion_rate` | Rate of offset deletions | | `kafka_server_group_coordinator_metrics_offset_expiration_count` | Number of offset expirations | | `kafka_server_group_coordinator_metrics_offset_expiration_rate` | Rate of offset expirations | | `kafka_server_group_coordinator_metrics_num_partitions` | Number of `__consumer_offsets` partitions managed by this broker | | `kafka_server_group_coordinator_metrics_partition_load_time_avg` | Average time taken to load a group metadata partition from disk into memory | | `kafka_server_group_coordinator_metrics_partition_load_time_max` | Maximum time taken to load a group metadata partition | | `kafka_server_group_coordinator_metrics_streams_group_count` | Number of Kafka Streams application groups handled by this coordinator | | `kafka_server_group_coordinator_metrics_streams_group_rebalance_count` | Total count of rebalances specifically for Kafka Streams groups | | `kafka_server_group_coordinator_metrics_streams_group_rebalance_rate` | Rate of rebalances specifically for Kafka Streams groups | | `kafka_server_group_coordinator_metrics_thread_idle_ratio_avg` | Average idle time ratio for the group coordinator background threads | ### Tiered storage metrics[​](#tiered-storage-metrics "Direct link to Tiered storage metrics") Aiven for Apache Kafka includes several metrics to monitor the performance and health of your Apache Kafka broker's tiered storage operations. Access these metrics through Prometheus to gain insights into various aspects of tiered storage, including data copying, fetching, deleting, and their associated lags and errors. | Metric | Description | | ------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------- | | `kafka_server_BrokerTopicMetrics_RemoteCopyBytesPerSec_Count` | Number of bytes per second being copied to remote storage | | `kafka_server_BrokerTopicMetrics_RemoteCopyRequestsPerSec_Count` | Number of copy requests per second to remote storage | | `kafka_server_BrokerTopicMetrics_RemoteCopyErrorsPerSec_Count` | Number of errors per second encountered during remote copy | | `kafka_server_BrokerTopicMetrics_RemoteCopyLagBytes_Value` | Number of bytes in non-active segments eligible for tiering that are not yet uploaded to remote storage | | `kafka_server_BrokerTopicMetrics_RemoteCopyLagSegments_Value` | Number of non-active segments eligible for tiering that are not yet uploaded to remote storage | | `kafka_server_BrokerTopicMetrics_RemoteFetchBytesPerSec_Count` | Number of bytes per second being fetched from remote storage | | `kafka_server_BrokerTopicMetrics_RemoteFetchRequestsPerSec_Count` | Number of fetch requests per second from remote storage | | `kafka_server_BrokerTopicMetrics_RemoteFetchErrorsPerSec_Count` | Number of errors per second encountered during remote fetch | | `kafka_server_BrokerTopicMetrics_RemoteDeleteRequestsPerSec_Count` | Number of delete requests per second to remote storage | | `kafka_server_BrokerTopicMetrics_RemoteDeleteErrorsPerSec_Count` | Number of errors per second encountered during remote delete | | `kafka_server_BrokerTopicMetrics_RemoteDeleteLagBytes_Value` | Number of bytes in non-active segments marked for deletion but not yet deleted from remote storage | | `kafka_server_BrokerTopicMetrics_RemoteDeleteLagSegments_Value` | Number of non-active segments marked for deletion but not yet deleted from remote storage | | `kafka_server_BrokerTopicMetrics_BuildRemoteLogAuxStateErrorsPerSec_Count` | Rate of errors occurring while building remote log auxiliary state | | `kafka_server_BrokerTopicMetrics_BuildRemoteLogAuxStateRequestsPerSec_Count` | Rate of requests to build auxiliary state for remote tiered storage logs | | `kafka_server_BrokerTopicMetrics_RemoteLogMetadataCount_Value` | Total count of remote log metadata entries | | `kafka_server_BrokerTopicMetrics_RemoteLogSizeBytes_Value` | Total size in bytes of log segments currently stored in remote tiered storage | | `kafka_server_BrokerTopicMetrics_RemoteLogSizeComputationTime_Value` | Time taken to compute the aggregate size of the remote logs | | `kafka_server_DelayedRemoteFetchMetrics_ExpiresPerSec_Count` | Rate at which delayed fetch requests from remote storage expire before being fulfilled | | `kafka_log_remote_RemoteLogManager_RemoteLogManagerTasksAvgIdlePercent_Value` | Average idle percentage of tasks within the Remote Log Manager | | `kafka_log_remote_RemoteLogManager_RemoteLogReaderFetchRateAndTimeMs_Count` | Total count of fetch operations initiated by the remote log reader | | `kafka_log_remote_RemoteLogManager_RemoteLogReaderFetchRateAndTimeMs_Mean` | Mean latency of fetch operations performed by the remote log reader | | `kafka_log_remote_RemoteLogManager_RemoteLogReaderFetchRateAndTimeMs_99thPercentile` | 99th percentile latency of fetch operations performed by the remote log reader | | `org_apache_kafka_storage_internals_log_RemoteStorageThreadPool_RemoteLogReaderAvgIdlePercent_OneMinuteRate` | One-minute moving average of the idle percentage of the remote log reader thread pool | | `org_apache_kafka_storage_internals_log_RemoteStorageThreadPool_RemoteLogReaderTaskQueueSize_Value` | Current number of pending tasks in the remote log reader's execution queue | | `kafka_server_RemoteLogManager_remote_copy_throttle_time_remote_copy_throttle_time_avg` | Average time that remote copy operations were artificially throttled | | `kafka_server_RemoteLogManager_remote_copy_throttle_time_remote_copy_throttle_time_max` | Maximum time that remote copy operations were artificially throttled | | `kafka_server_RemoteLogManager_remote_fetch_throttle_time_remote_fetch_throttle_time_avg` | Average time that remote fetch operations were artificially throttled | | `kafka_server_RemoteLogManager_remote_fetch_throttle_time_remote_fetch_throttle_time_max` | Maximum time that remote fetch operations were artificially throttled | --- # Aiven for Apache Kafka® version lifecycle Learn how Aiven manages Aiven for Apache Kafka® version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. ## Aiven version support and upstream EOL[​](#aiven-version-support-and-upstream-eol "Direct link to Aiven version support and upstream EOL") Aiven aims to follow the EOL schedule set by the original authors and maintainers of the open source software (the upstream projects). Once the upstream project retires a specific version, they do not receive security updates and critical bug fixes anymore by the maintainers. Outdated services don't offer the level of protection you need, so Aiven follows the upstream project's EOL schedule to ensure that Aiven services are always running on supported versions. ## Service version numbering[​](#service-version-numbering "Direct link to Service version numbering") Aiven services inherit the upstream project's software versioning scheme. Depending on the service, a major version can be either a single digit or in the format `major.minor`. The exact version of the service is visible in the [Aiven Console](https://console.aiven.io/) when the service is running. ## Service version EOL policy[​](#service-version-eol-policy "Direct link to Service version EOL policy") Aiven sets an EOL date for each major version of the service. This policy covers both running and powered-off services on affected versions. ## EOL notifications[​](#eol-notifications "Direct link to EOL notifications") When Aiven sets the EOL date for a service major version: * You receive an email notification along with instructions on the next steps. * The [Aiven Console](https://console.aiven.io/) shows an EOL alert for affected services. * You receive email reminders monthly. * In the month of the EOL date, you receive weekly reminders. ## EOL best practices[​](#eol-best-practices "Direct link to EOL best practices") * Use service forking to test the version upgrade before upgrading your production services. * Upgrade to the supported version before the EOL date. This gives you time to test compatibility, resolve any issues, and plan the upgrade on your schedule. After the EOL date: * If the service is powered on, it's automatically upgraded to the latest version when possible, or to another supported version. note If it's not possible to upgrade a powered-on service to a supported version, the service is powered off and ultimately deleted. * If the service is powered off, it's deleted. ## Version EOL dates[​](#version-eol-dates "Direct link to Version EOL dates") Aiven for Apache Kafka versions reach EOL one year after they become available on the Aiven Platform. | Version | Aiven EOL | Service creation supported until | Service creation supported from | | ------- | ---------- | -------------------------------- | ------------------------------- | | 3.8.x | 2026-09-30 | 2026-06-30 | 2024-09-06 | | 3.9.x | 2027-09-30 | 2027-06-30 | 2025-03-20 | | 4.0.x | 2026-09-18 | 2026-06-18 | 2025-09-18 | | 4.1.x | 2027-01-31 | 2026-09-10 | 2025-12-10 | | 4.2.x | 2027-06-15 | 2027-03-15 | 2026-06-15 | note Apache Kafka 3.8 is the last version that supports ZooKeeper. Starting with Apache Kafka 3.9, Aiven for Apache Kafka uses KRaft (Kafka Raft) to manage metadata and controllers instead of ZooKeeper. For details about the migration process and rollout limitations, see: * [KRaft in Aiven for Apache Kafka®](/docs/products/kafka/concepts/kraft-mode.md) * [Transitioning to KRaft](/docs/products/kafka/concepts/upgrade-procedure.md#transitioning-to-kraft) To support the transition to KRaft, Aiven supports Apache Kafka 3.8 until the EOL date shown in the table. The EOL date already includes the extended support period. Related pages * [Apache Kafka® upgrade procedure](/docs/products/kafka/concepts/upgrade-procedure.md) * [Maintenance and updates](/docs/products/kafka/howto/maintenance-updates.md) --- # Standard Kafka overview Standard Kafka is an Aiven for Apache Kafka® service type that stores topic data in object storage through diskless topics. It is available on Aiven Cloud, runs on Kafka 4.x and later, and supports classic and diskless topics in the same service. Standard Kafka is compatible with Apache Kafka APIs and clients. New customers can create Standard Kafka services. Classic Kafka remains available only for existing customers. In the Aiven CLI and advanced configuration, Standard Kafka is still identified with `inkless`. ## Key differences from classic Kafka[​](#key-differences-from-classic-kafka "Direct link to Key differences from classic Kafka") Standard Kafka changes how Kafka services store and manage data: * **Classic topics:** Tiered storage of classic topics is enforced with local retention set at 15 minutes or a 5 GB partition limit. * **Diskless topics:** Opt-in diskless topics can be used and store all retained data in object storage. * **Managed configuration:** Some broker-level settings use managed defaults. * **KRaft-based metadata management:** Standard Kafka supports Apache Kafka 4.x and later, so all Standard Kafka services use [KRaft](/docs/products/kafka/concepts/kraft-mode.md) for metadata and consensus instead of ZooKeeper. * **Kafka Connect deployment:** Kafka Connect is deployed as a separate service. ## Billing and cost[​](#billing-and-cost "Direct link to Billing and cost") Aiven bills Standard Kafka services for compute, storage, and network usage as separate components. Billing depends on your selected service plan and actual usage. For details on how network usage is measured and priced, see [Pricing for Aiven for Apache Kafka®](/docs/products/kafka/concepts/kafka-pricing.md). ## When to use Standard Kafka[​](#when-to-use-standard-kafka "Direct link to When to use Standard Kafka") Use Standard Kafka to: * Scale storage without managing broker disk capacity. * Retain larger volumes of data for extended periods. * Scale and recover clusters faster than fixed-storage deployments. * Combine classic and diskless topics in the same service. ## Existing Classic Kafka services[​](#existing-classic-kafka-services "Direct link to Existing Classic Kafka services") Existing Classic Kafka services continue to run unchanged. You cannot upgrade or migrate a Classic Kafka service to Standard Kafka. You set the service type when you create the service and cannot change it later. To use Standard Kafka, create a Standard Kafka service. Related pages * [Create an Aiven for Apache Kafka® Professional tier service](/docs/products/kafka/get-started/create-kafka-service.md) * [Classic Kafka overview](/docs/products/kafka/classic-kafka-overview.md) * [Pricing for Aiven for Apache Kafka®](/docs/products/kafka/concepts/kafka-pricing.md) * [Diskless topics overview](/docs/products/kafka/diskless/concepts/diskless-topic-overview.md) --- # Apache Kafka® terminology A comprehensive glossary of essential Apache Kafka® terms and their meaning. ## Broker[​](#Broker "Direct link to Broker") A server that operates Apache Kafka, responsible for message storage, processing, and delivery. Typically part of a cluster for enhanced scalability and reliability, each broker functions independently but is integral to Kafka's overall operations, separate from tools like Apache Kafka Connect. ## Consumer[​](#consumer "Direct link to Consumer") An application that reads data from Apache Kafka, often processing or acting upon it. Various tools used with Apache Kafka ultimately function as either a producer or a consumer when communicating with Apache Kafka. ## Consumer groups[​](#consumer-groups "Direct link to Consumer groups") Groups of consumers in Apache Kafka are used to scale beyond a single application instance. Multiple instances of an application coordinate to handle messages, with each group allocated to different partitions for even workload distribution. ## Event-driven architecture[​](#event-driven-architecture "Direct link to Event-driven architecture") Application architecture centered around responding to and processing events. ## Event[​](#Event "Direct link to Event") A single discrete data unit in Apache Kafka, consisting of a `value` (the message body) and often a `key` (for quick identification) and `headers` (metadata about the message). ## KRaft[​](#kraft "Direct link to KRaft") Kafka Raft (KRaft) is the metadata and consensus management system in Apache Kafka, replacing [ZooKeeper](#zookeeper). Aiven for Apache Kafka uses [KRaft](/docs/products/kafka/concepts/kraft-mode.md) starting with Apache Kafka 3.9 to manage metadata, controllers, and consensus, removing the need for a separate ZooKeeper cluster. ## Kafka node[​](#kafka-node "Direct link to Kafka node") See [Broker](#Broker) ## Kafka server[​](#kafka-server "Direct link to Kafka server") See [Broker](#Broker) ## Message[​](#message "Direct link to Message") See [Event](#Event) ## Partitioning[​](#partitioning "Direct link to Partitioning") A method used by Apache Kafka to distribute a topic across multiple servers. Each server acts as the `leader` for a partition, ensuring data sharding and message order within each partition. ## Producer[​](#producer "Direct link to Producer") An application that writes data into Apache Kafka without concern for the data consumers. The data can range from well-structured to simple text, often accompanied by metadata. ## Pub/sub[​](#pubsub "Direct link to Pub/sub") A publish-subscribe messaging architecture where messages are broadcast by publishers and received by any listening subscribers, unlike point-to-point systems. ## Queueing[​](#queueing "Direct link to Queueing") A messaging system where messages are sent and received in the order they are produced. Apache Kafka maintains a watermark for each consumer to track the most recent message read. ## Record[​](#record "Direct link to Record") See [Event](#Event) ## Replication[​](#replication "Direct link to Replication") Apache Kafka's feature for data replication across multiple servers, ensuring data preservation even if a server fails. This is configurable per topic. ## Topic[​](#topic "Direct link to Topic") Logical channels in Apache Kafka through which messages are organized. Topics are named in a human-readable manner, like `sensor-readings` or `kubernetes-logs`. ## ZooKeeper[​](#zookeeper "Direct link to ZooKeeper") A distributed system that managed metadata, controller election, and configuration storage in Apache Kafka. Starting in Apache Kafka version 3.9, Aiven for Apache Kafka uses [KRaft](#kraft) instead. --- # NOT\_LEADER\_FOR\_PARTITION errors Aiven continuously monitors services to ensure they are healthy. If problems arise, nodes can be recycled: new nodes are created to substitute old, malfunctioning ones. During nodes replacement in your Aiven for Apache Kafka® cluster, you may find several `NOT_LEADER_FOR_PARTITION` warnings or errors in the logs of your producer. The exact error message depends on your client library and log formatting, but should be similar to the following: ``` [2021-02-04 09:01:20,118] WARN [Producer clientId=test-producer] Received invalid metadata error in produce request on partition topic1-25 due to org.apache.kafka.common.errors.NotLeaderForPartitionException: This server is not the leader for that topic-partition.. Going to request metadata update now (org.apache.kafka.clients.producer.internals.Sender) ``` This is an expected behavior as part of the failover from old to new nodes. Each producer contains a metadata cache that identifies which broker is the leader of each partition. When your code produces a message, it tries to send it to the broker that is the partition leader according to the producer's metadata cache. When nodes are replaced, the partition leaders are expected to change. Every partition leader changes at least once, but it can be more, depending on the number of nodes and how many nodes are replaced at a time. note If a service has 3000 active partitions, then you can expect at least one of these error messages for each partition in each producer. Most producers libraries update their cluster metadata cache on a regular poll interval, or immediately based on errors. After that, they continue producing messages without issue. However, high load on the broker side or unfortunate timing in parallel requests can sometimes trigger several updates. This might cause a very large number of these warnings in the producer logs, which can look worrying but is in fact harmless. --- # Troubleshoot Apache Kafka® consumer connections Apache Kafka® consumers sometimes experience disconnections from a cluster node. On one hand, isolated occurrences are expected and well mitigated by the Apache Kafka® protocol, for example if a partition leadership moved to a different node so the cluster remains available and balanced. On the other hand, repeating occurrences often raise concerns as they can impact consumer lag, and cluster performance. ## What is observed[​](#what-is-observed "Direct link to What is observed") The client behaviour may vary depending on the library and version (remind to keep it up-to-date) in use. A typical consumer disconnection log is as follows: `[AdminClient] Node 3 disconnected. (org.apache.kafka.clients.NetworkClient:937)`. On the server, the GroupCoordinator undertakes some consumer-group re-balancing activity, as for example: `Member ABC has left group my-app through explicit LeaveGroup; client reason: the consumer unsubscribed from all topics (kafka.coordinator.group.GroupCoordinator)` or `Timed out HeartbeatRequest`. ## What it means[​](#what-it-means "Direct link to What it means") As they are processing messages, consumer instances may break, fail to request more local resources or to complete their current task on time. In this case, they leave their group either explicitly or by missing to send a heartbeat to the cluster. As a consequence, the partition assignments of the consumer leaving the group are automatically re-affected to other consumers of the group, if they exist. ## Why it occurs now and never before[​](#why-it-occurs-now-and-never-before "Direct link to Why it occurs now and never before") * The consumer of a given topic `A` can get dropped because it relies on the same `group.id` as another consumer leaving or joining the group, even if the later consumes a totally different topic `B`. Any rebalance triggered by any one consumer would still affect the other consumers in the group. * Although consumer heartbeat and message processing are to be executed by 2 different client threads, the machine (on which the consumer application runs) can get overloaded and so it cannot send a heartbeat to the group coordinator after `heartbeat.interval.ms` (default: 3 s) and before `session.timeout.ms` (default: 45 s). * If the consumer logic consists in expensive transformations or synchronous API calls, then your average single message processing time can dramatically increase, as in the following example scenarios: * The subsystem it depends on has an issue (ex. the destination server is not available). * The incoming message size increased. * If the consumer lag has grown enough so that the consumer now processes `max.poll.records` (default: 500), then your average batch processing time and resources can dramatically increase, with the following consequences: * The consumer may run out-of-memory, resulting in an application crash. * The `max.poll.interval.ms` is exceeded between polls, so the consumer sends a LeaveGroup-request to the coordinator rather than sending a heartbeat. ## How to solve it[​](#how-to-solve-it "Direct link to How to solve it") * Make sure that your consumers are relying on a unique `group.id`. * It is generally a good idea to reduce `max.poll.records` or `max.partition.fetch.bytes`, and monitor if the situation improves. * If it is required to stick to large batches (such as in windowing and data lake ingestion use-cases), then temporarily increasing `session.timeout.ms` might do the trick. * Remind that `session.timeout.ms` must be in the allowable range as configured in the broker configuration by `group.min.session.timeout.ms` and `group.max.session.timeout.ms`. * It is recommended to configure `heartbeat.interval.ms` to be no more than a third of `session.timeout.ms`. This ensures that even if a heartbeat or two are lost over the transient network, then the consumer is still considered alive. * If the problem arose because of an increase in the volume of incoming messages, then we would advise to scale the consumer up/vertically or out/horizontally. If the above didn't help you resolve your issue and you have further concerns or questions, contact [Aiven support](mailto:support@aiven.io). --- # Aiven for Metrics Aiven for Metrics, powered by Thanos, simplifies the management and analysis of large volumes of metrics data. The service is scalable, reliable, and efficient, suitable for organizations of all sizes. This service simplifies the management of large-scale metrics systems, allowing organizations to focus on deriving insights from their data. ## Key components[​](#key-components "Direct link to Key components") Aiven for Metrics includes several core Thanos components: * **Thanos Metrics Query**: Enables users to query and visualize metrics from real-time and historical data sources, aggregating data from various sources for a unified view. * **Thanos Metrics Receiver**: Handles metrics ingestion into the system, acting as a receiver for Prometheus remote write requests to enable real-time metrics collection. * **Thanos Metrics Store**: Interfaces with object storage to provide access to historical data, ensuring scalable and reliable long-term storage. * **Thanos Metrics Compact**: Enhances storage usage and query efficiency by compacting and downsampling data stored in object storage, improving performance and reducing costs. * **Thanos Query Frontend**: Caches query results and splits large queries into smaller sub-queries for efficient execution across multiple Thanos Query instances. No external Thanos access While Aiven for Metrics uses the Thanos architecture internally, **you cannot connect your Aiven-managed services to any external (non-Aiven) Thanos endpoint**. Connect to Aiven for Metrics instead. ## Unified cluster architecture[​](#unified-cluster-architecture "Direct link to Unified cluster architecture") Aiven for Metrics combines these components into a cohesive cluster architecture. This setup ensures a seamless data flow from ingestion (via Thanos Metrics Receive) to long-term storage (in object storage through Thanos Metrics Store) and efficient querying (through Thanos Metrics Query). The Thanos Compact component optimizes data storage and retrieval processes in the background, keeping the system efficient and cost-effective. Aiven for Metrics combines these components into a cohesive cluster architecture, enhancing metrics management: * **Data collection and storage**: Thanos Metrics Receivers collect metrics in real-time and store them in object storage after the Time Series Database (TSDB) block is full, typically every 2 hours. * **Query processing**: The Thanos Query Frontend receives requests and optimizes load distribution by forwarding requests to Thanos Metrics Query. Depending on the query's time range, this component retrieves real-time data from Thanos Metrics Receivers or historical data from the Thanos Metrics Store, directly querying object storage. To maintain data integrity, the system removes duplicate samples received from multiple Thanos Metric Receivers. After processing, Thanos Metrics Query responds to the Query Frontend, which caches the results to speed up future queries before delivering them to the client. ## Benefits of Aiven for Metrics[​](#benefits-of-aiven-for-metrics "Direct link to Benefits of Aiven for Metrics") * **Centralized monitoring**: Query and analyze metrics from multiple Prometheus servers and clusters through a unified query view, simplifying monitoring across your infrastructure. * **Unlimited retention and scalability**: With scalable object storage solutions, you can store unlimited metric data for any duration. * **Prometheus compatibility**: Aiven for Metrics is compatible with Prometheus, allowing you to seamlessly use familiar tools like Grafana. * **Cost-effective and efficient**: Downsampling and compacting data reduces storage needs and improves query performance, resulting in greater efficiency and cost savings. * **Simplified operations**: Reduce the complexity of your metrics system with a managed service that provides a pre-configured and optimized Thanos setup. ## Limitations[​](#limitations "Direct link to Limitations") * **No direct Thanos access**: While built on Thanos, direct API connections to Thanos components are not supported. All access must go through Aiven service integrations. * Aiven for Metrics is not currently available on Azure or Google Cloud Marketplace. Related pages * [Thanos documentation](https://thanos.io/v0.34/thanos/getting-started.md/) --- # Retention rules in Aiven for Metrics Retention rules in Aiven for Metrics define how long your metrics data is stored. By default, all data is retained indefinitely, ensuring uninterrupted access to historical insights. Optimize storage and access to historical data by tailoring retention periods to align with specific data strategies, compliance needs, and cost management goals. No external Thanos access While Aiven for Metrics uses the Thanos architecture internally, **you cannot connect your Aiven-managed services to any external (non-Aiven) Thanos endpoint**. Connect to Aiven for Metrics instead. ## Define retention rules[​](#define-retention-rules "Direct link to Define retention rules") Aiven for Metrics uses the Thanos Metrics Compactor to simplify retention settings. A single parameter, `compactor.retention.days`, sets the same retention period for all types of data: `raw`, `5-minute downsampled`, and `1-hour downsampled`. To adjust the `compactor.retention.days` parameter: 1. Click **Service settings** from the sidebar. 2. Scroll to **Advanced configuration** and click **Configure**. 3. Click **Add configuration options**. 4. Locate the `compactor.retention.days` parameter. 5. Set the desired retention period by adjusting the parameter value. 6. Click **Save configuration**. ## Downsampling time series data[​](#downsampling-time-series-data "Direct link to Downsampling time series data") Downsampling in Aiven for Metrics is the process of lowering the resolution of time series data while preserving its essential informational value. By aggregating this data into distinct 5-minute and 1-hour intervals, the technique significantly enhances the speed and efficiency of queries over long periods. This approach optimizes storage usage and ensures that the data remains manageable and accessible, retaining crucial details. Related pages * [Enforcing retention of data](https://thanos.io/tip/components/compact.md/#enforcing-retention-of-data) --- # Memory and out-of-memory conditions in Aiven for Metrics Understand the memory limits and out-of-memory conditions that apply to your Aiven for Metrics service. ## Service memory limits[​](#service-memory-limits "Direct link to Service memory limits") The practical memory limit will always be less than the service physical memory limit. **All services are subject to operating overhead:** * A small amount of memory is required by the operating system kernel to manage system resources, including networking functions and disk cache. * Aiven's cloud data platform requires memory to monitor availability, provide metrics, logging and manage backups. A server or node's **usable memory** can be calculated as: `usable memory = RAM - overhead` Where: * `overhead` is 350 MiB (≈ 0.34 GiB). Services may utilize optional components, service integrations, connection pooling, or plug-ins, which are not included in overhead calculations. If a service is overcommitted, the operating system, management layer, backups or availability monitoring, may fail status checks or operations due to resource contention. In severe instances, the node may fail completely with an out-of-memory condition. ## Out of memory conditions[​](#out-of-memory-conditions "Direct link to Out of memory conditions") Many processes request more memory from the kernel than they will ever use or need. In these cases, the kernel overallocates memory. This allows it to satisfy multiple processes requesting more memory than is available, which is not used or is freed by the time any other process actually needs it. However, if enough processes start using all their allocated memory simultaneously there may not be enough physical memory available and an `Out Of Memory` (`OOM`) condition occurs. warning This situation is critical and must be resolved immediately. The solution that the Linux kernel employs is to invoke the `Out of Memory Killer` (or `OOM Killer`). This reviews all running processes and kills one or more of them to free up system memory and keep the system running. The `OOM Killer` selects process to kill based on an `oom_score`; a calculation that balances how much memory the process is using with how long the process has been running. Processes that have been running for a long time are less likely to be killed. Subprocesses are summed with parent processes in terms of memory usage, so a process which forks many subprocesses, but itself does not use a lot of memory, may still be killed. In most instances, the hosted data service, or a child process, will have the highest memory footprint and be a prime candidate for termination when the OOM Killer inspects the running processes. Aiven's cloud data platform leverages kernel namespaces or containers to isolate processes from each other. Isolation has several benefits, including: * A smaller footprint for security‑related concerns * A smaller blast radius for failure * Greater control of system resources Left unchecked, the `OOM Killer` may opt to kill the primary service. This is undesirable as unclean termination of the primary service can lead to data loss, inconsistency, or corrupted backups. Further, if Aiven's management platform detects that the primary service is unavailable for , the service will be marked as down and a failover will occur. To mitigate this scenario, namespaces are used, some with additional memory limits, in combination with an `oom_score_adjust` on the primary process, to coax the `OOM Killer` into selection of less critical processes. This will still result in a service restart, but in a more controlled process, where the database is shut down, rather than killed; exposure to data loss is limited and recovery is faster when the service restarts, often avoiding failover. warning Out of Memory conditions can still lead to unexpected behavior, including data unavailable or data loss conditions. ## Avoid running low on memory[​](#avoid-running-low-on-memory "Direct link to Avoid running low on memory") The OOM killer only runs when the system is critically low on memory. To prevent it from running, either reduce your memory usage or increase the available memory. For most databases, the service memory footprint can often be reduced by: * Reducing concurrency or implementing connection pooling * Tuning queries to limit result sets * Tuning indexes for query load * Dropping unused objects from storage In cases where the working set no longer fits into memory, consider scaling your service. Related pages * [Change the service plan](/docs/products/metrics/howto/change-service-plan.md) * [Prepare for high load](/docs/products/metrics/howto/prepare-for-high-load.md) --- # Optimize storage and resources Aiven for Metrics optimizes storage and compute resources using tiered storage, disk storage, memory, and compute power to balance cost and performance. ## Storage solutions[​](#storage-solutions "Direct link to Storage solutions") Aiven for Metrics optimizes storage by using tiered storage for long-term retention and disk storage to temporarily store or process historical data. This combination ensures scalability, reliability, and cost-efficient metric storage. No external Thanos access While Aiven for Metrics uses the Thanos architecture internally, **you cannot connect your Aiven-managed services to any external (non-Aiven) Thanos endpoint**. Connect to Aiven for Metrics instead. ### Tiered storage[​](#tiered-storage "Direct link to Tiered storage") Tiered storage is the **primary storage** for metrics and metadata in Aiven for Metrics. It is **required and automatically enabled for all plans** to provide scalability and long-term cost efficiency. Aiven manages this storage infrastructure to ensure data security and reliability. * Metrics are uploaded to tiered storage every 2 hours for historical analysis. * Stored data remains accessible for queries at all times, ensuring continuous availability for real-time decision making. ### Disk storage[​](#disk-storage "Direct link to Disk storage") Aiven for Metrics uses disk storage for internal processing tasks. Three main components that rely on disk storage: * **Thanos Metric Receiver**: This component is the initial point of contact for your metrics. It accepts incoming data streams, temporarily stores them in a local cache using disk space, and transfers them to tiered storage at regular intervals. * **Thanos Metric Store**: The store component tracks the location of your metric data shards within the tiered storage. It also maintains a small amount of information about these remote blocks on the local disk, ensuring it stays synchronized with the tiered storage. * **Thanos Metric Compactor**: The compactor component is crucial in optimizing long-term data storage usage. It periodically downloads data chunks from object storage, performs downsampling (reducing data granularity for older data), and uploads the compacted data back to tiered storage. This process requires temporary storage of the downloaded data on the local disk. The amount of disk space needed varies depending on your data volume and complexity. Aiven's service plans are designed to handle the most typical use cases. ## Storage costs and billing[​](#storage-costs-and-billing "Direct link to Storage costs and billing") Aiven for Metrics storage costs consist of two components: * **Local disk storage**: Included in the base service plan to temporarily store or process historical data and caching. * **Data stored in tiered storage**: Billed based on the highest amount of data retained in tiered storage during each billing period. **Example billing scenario** If you use a **Start-16** plan with **640 GB** of local disk storage, your billing includes: * The base cost of the Start-16 plan. * Additional tiered storage costs, determined by the volume of metrics data retained. ### BYOC (Bring Your Own Cloud) billing[​](#byoc-bring-your-own-cloud-billing "Direct link to BYOC (Bring Your Own Cloud) billing") [BYOC](/docs/platform/concepts/byoc.md) billing for tiered storage can vary depending on your specific agreement with Aiven. Possible fees include: * **Customer costs**: In all BYOC setups, you are responsible for the full cost of the underlying cloud storage used by tiered storage. This includes all stored data, regardless of local retention settings. * **Aiven management fee**: In addition to cloud storage costs, an Aiven management fee applies to data stored in tiered storage. This fee is based on the total storage used. ## Resource management[​](#resource-management "Direct link to Resource management") Aiven automatically manages memory and compute resources based on your service plan. However, monitoring these resources can help identify potential bottlenecks. * **Memory usage** refers to the RAM needed for Aiven for Metrics to run effectively. Aiven automatically manages memory allocation based on your service plan. However, monitoring memory usage can help you identify potential issues. Factors affecting memory usage include: * Number of ingested metrics * Complexity of metric queries * Number of concurrent users * **Compute usage** refers to the processing power the Aiven for Metrics service uses. Aiven manages compute resources based on your service plan. Monitoring compute usage can help you identify if your plan offers sufficient resources. Factors affecting compute usage include: * The volume of ingested metric data * Frequency of queries * Complexity of queries ## Using Dynamic Disk Sizing (DDS)[​](#using-dynamic-disk-sizing-dds "Direct link to Using Dynamic Disk Sizing (DDS)") Aiven offers Dynamic Disk Sizing (DDS) as a flexible option if you require additional disk space beyond the standard plan allocation. DDS allows you to: * **Scale up**: Increase disk space on demand for bursts of metrics or high cardinality. * **Scale down**: Reduce disk space when needs change, optimizing costs. Consider enabling DDS when: * A large influx of metrics occurs within a short period, such as during data migrations. * A significant amount of data with many unique metric identifiers can strain the disk space used to temporarily store or process data by the compactor and potentially other components. * While Aiven's plans work for most workloads, consider using DDS if you anticipate a recent large data migration or a significant increase in metric volume. Related pages * [Thanos Receiver](https://thanos.io/tip/components/receive.md/) * [Thanos Store](https://thanos.io/tip/components/store.md/#store) * [Thanos Compactor](https://thanos.io/tip/components/compact.md/#disk) --- # Get started with Aiven for Metrics Get started with Aiven for Metrics by creating your service using the [Aiven Console](https://console.aiven.io/) or [Aiven CLI](https://github.com/aiven/aiven-client). note Aiven for Metrics is not currently available on Azure or Google Cloud Marketplace. ## Create a service[​](#create-a-service "Direct link to Create a service") * Console * CLI * Terraform 1. In your project, click **Services**. 2. Click **Create service**. 3. Select **Thanos Metrics**. 4. Select a **Cloud**. 5. Select a **Plan**. note The plans available can vary between cloud providers and regions for the same service. 6. In the **Service details**, enter a name for your service. 7. Optional: Add service tags. 8. In the **Service summary**, click **Create service**. The status of the service is **Rebuilding** during its creation. When the status is **Running**, you can start using the service. This typically takes a couple of minutes and can vary between cloud providers and regions. The [Aiven CLI](https://github.com/aiven/aiven-client) provides a simple and efficient way to create an Aiven for Metrics service. If you prefer creating a new service from the CLI: 1. Determine the service plan, cloud provider, and region to use for your Aiven for Metrics service. 2. Run the following command to create an Aiven for Metrics service named metrics-demo: ``` avn service create metrics-demo \ --service-type thanos \ --cloud aws-europe-west1 \ --plan startup-4 \ --project PROJECT_NAME ``` Where `PROJECT_NAME` is the name of your Aiven project. To see: * A full list of default flags, run `avn service create -h` * Type-specific options, run `avn service types -v` The following example files are also available in the [Aiven Terraform Provider repository](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/thanos) on GitHub. 1. Create a file named `provider.tf` and add the following: ``` Loading... ``` 2. Create a file named `service.tf` and add the following: ``` Loading... ``` 3. Create a file named `variables.tf` and add the following: ``` Loading... ``` 4. Create a file named `terraform.tfvars` and add values for your token and Aiven project. To apply your Terraform configuration: 1. Initialize Terraform by running: ``` terraform init ``` The output is similar to the following: ``` Initializing the backend... Initializing provider plugins... - Finding aiven/aiven versions matching ">= 4.0.0, < 5.0.0"... - Installing aiven/aiven v4.9.2... - Installed aiven/aiven v4.9.2 ... Terraform has been successfully initialized! ... ``` 2. To create an execution plan and preview the changes, run: ``` terraform plan ``` 3. To deploy your changes, run: ``` terraform apply --auto-approve ``` ## Create service integrations[​](#create-service-integrations "Direct link to Create service integrations") Integrate Aiven for Metrics with other Aiven services, such as OpenSearch for advanced queries or Grafana for visualization, or connect it with another Aiven for Metrics service for comprehensive monitoring. Set up integrations using the [Aiven Console](/docs/platform/howto/create-service-integration.md), [Aiven CLI](/docs/tools/cli/service/integration.md), or [Aiven Terraform Provider](/docs/platform/howto/create-service-integration.md). The [Aiven for Metrics integration example](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/thanos/thanos_pg) in GitHub shows you how to use the Aiven Terraform Provider to create and integrate Thanos Metrics with PostgreSQL and Grafana. --- # Change the cloud or region for your Aiven for Metrics service Move your Aiven for Metrics service to a different cloud provider or region. 1. In your service, click **Service settings** from the sidebar. 2. In the **Cloud and network** section, click **Actions** > **Change cloud**. 3. In the **Cloud** section , select a cloud provider and region, and click **Change**. Your service starts a migration to the new location and remains available during the process. When the migration completes, the service continues running in the new cloud or region. Related pages * [Migrate to another cloud or region](/docs/platform/howto/migrate-services-cloud-region.md) * [Change the service plan](/docs/products/metrics/howto/change-service-plan.md) --- # Change the plan for your Aiven for Metrics service Change the service plan for your Aiven for Metrics service to scale resources up or down and optimize costs. Adjust the plan of your services at any time to scale your services as needed and optimize costs. If you can't find a suitable plan, you can [request a custom plan](/docs/platform/concepts/service-pricing.md). tip If you plan to upgrade your service plan, do it immediately after a full backup. This reduces the amount of incremental changes that need to be applied on top of the base backup, which speeds up the upgrade itself. important * When changing a service plan, reserve an additional 25% of disk space. This requirement applies to upgrades and downgrades. * Downgrading to a plan with fewer VMs is supported for most services, including Aiven for Apache Kafka®, Aiven for PostgreSQL®, Aiven for OpenSearch®, Aiven for ClickHouse®, Aiven for MySQL®, Aiven for Metrics, and Aiven for Valkey™. * Changing a service plan triggers a node recycle, service rebuilding, and any pending maintenance updates. - Console - Terraform - CLI 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Change plan**. 3. Select a plan that provides at least 125% of the current disk size and click **Change plan**. Update the `plan` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). To change a service plan in the Aiven CLI, use the [`avn service update --plan `](/docs/tools/cli/service-cli.md#avn-cli-service-update) command. Your service's state becomes **Rebuilding** and remains accessible. When the state switches to **Running**, your new service plan is active. Related pages * [Prepare for high load](/docs/products/metrics/howto/prepare-for-high-load.md) * [Memory and out-of-memory conditions](/docs/products/metrics/concepts/service-memory.md) --- # Controlled upgrade pipelines for your Aiven for Metrics service [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Link Aiven for Metrics services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. Control when your Aiven managed services receive maintenance updates and test maintenance updates in development or staging environments before they reach production. important Controlled upgrade pipeline is a [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-) feature. [Contact Aiven](https://aiven.io/contact) to request access. Aiven performs automatic service maintenance for security fixes, minor software updates, and other platform changes. The controlled upgrade pipeline feature lets you link services of the same type in an ordered sequence to control when each service receives updates. After a maintenance update upgrades a service at the initial pipeline step, you validate that service version before the update proceeds to the service at the next pipeline step. Validating means approving the new version as safe to roll out to the next service. Validation can be manual or automatic after a configurable delay. ## Why use controlled upgrade pipelines[​](#why-use-controlled-upgrade-pipelines "Direct link to Why use controlled upgrade pipelines") Controlled upgrade pipelines prevent production incidents caused by automatic updates reaching production before teams can test the new version in a lower environment. They give you full oversight of the update process: * **Risk mitigation**: Prevents unexpected maintenance updates from breaking your production environment by ensuring they are tested in a non-production setting first. * **Stability**: Keeps destination services (such as production) on a known-good version until you, or the automatic timer, confirm the new version is safe. * **Process control**: Allows platform teams to standardize their deployment and maintenance lifecycle across environments. ## About controlled upgrade pipelines[​](#about-controlled-upgrade-pipelines "Direct link to About controlled upgrade pipelines") ### Upgrade steps[​](#upgrade-steps "Direct link to Upgrade steps") An upgrade step is a pair of services linked by an upgrade constraint: * **Source service**: The service that receives maintenance updates first * **Destination service**: The service that waits for validation before receiving updates Each destination service can have only one source service. A source service can have multiple destination services. ### Upgrade pipelines[​](#upgrade-pipelines "Direct link to Upgrade pipelines") An upgrade pipeline is a chain of upgrade steps that spans multiple environments. For example: * Single chain: development → staging → production * Multiple destinations: development → production-eu and development → production-na ## How validation works[​](#how-validation-works "Direct link to How validation works") When a maintenance update upgrades your source service: 1. The source service receives the update first. 2. Test the updated source service to verify it works as expected. 3. Validate the update manually using the API or CLI, or wait for automatic validation after the configured delay. The default delay is 7 days. 4. After validation, the destination service becomes eligible for the same maintenance update. 5. The destination service receives the update during its next maintenance window. If one source service has multiple destination services, one validation for the source service applies to all connected destination services. ### Validation and maintenance windows[​](#validation-and-maintenance-windows "Direct link to Validation and maintenance windows") Validation and the maintenance window control different things: * **Validation** controls *what* version the destination service upgrades to. * The maintenance window controls *when* the upgrade happens. After you validate an update, or automatic validation applies, the destination service receives the validated version during its next scheduled maintenance window. Validation does not trigger an immediate upgrade outside the maintenance window. Upgrade pipelines add a constraint on what is installed during a maintenance update; they do not change when maintenance runs. Nodes in the destination service maintain the validated version until a newer version is validated, either when you validate it manually or when automatic validation applies after the configured delay. When a node is recycled, it uses the same validated version, not the latest available version. When you create a step, the destination service keeps the newest version that is already validated at that moment. If the destination service is already applying maintenance during step creation, the in-progress target version becomes the initial validated version. warning A powered-off source service cannot receive maintenance updates, so you cannot validate it. If you power off services earlier in the chain, the destination service upgrades regardless. For example, in a development → staging → production chain, if both development and staging are powered off, production upgrades without testing and validation in the earlier environments. Keep services in the chain powered on to preserve the protection that upgrade pipelines provide. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * **Same service type**: You can only link services of the same type. For example, two Aiven for PostgreSQL services. * **Chain length**: The default maximum chain depth is 3 services, which is 2 steps. If you need a longer chain, [contact Aiven](https://aiven.io/contact). * **No cycles**: You cannot create circular dependencies between services. * **Emergency overrides**: Aiven can apply critical security or stability fixes to a destination service before explicit validation. * **Supported services**: This feature supports all Aiven service types except Aiven for Apache Flink® and Aiven for MySQL. * **Automatic maintenance updates only**: Pipelines apply to automatic maintenance updates, such as minor service version updates and node image updates. Major version upgrades, for example Aiven for PostgreSQL® 15 to 16, require manual action and are not promoted automatically through the pipeline. * **No permanent blocking**: You cannot prevent an update indefinitely. Automatic validation applies after the configured delay, up to the maximum delay. * **No validation rollback**: You cannot undo a validation after it is recorded. ## Use controlled upgrade pipelines[​](#use-controlled-upgrade-pipelines "Direct link to Use controlled upgrade pipelines") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") To use controlled upgrade pipelines, you need the following: * The feature enabled by Aiven ([Limited availability](/docs/platform/concepts/service-and-feature-releases.md)) * Dev tool of your choice: * [Aiven CLI](/docs/tools/cli.md) Install the latest version of the Aiven CLI to access the `upgrade-pipeline` commands. * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) * Set `PROVIDER_AIVEN_ENABLE_BETA=true` before running Terraform. * See the [resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) for full schema, import format, and lifecycle behavior. * [Aiven Operator for Kubernetes](/docs/tools/kubernetes.md) Install the operator and create an Aiven token secret named `aiven-token` that the operator uses to authenticate against the Aiven API. * Write access to the source and destination projects * At least two services of the same type (for example, two Aiven for PostgreSQL® services) * Services can be in different projects in the same organization ### Set up an upgrade pipeline[​](#set-up-an-upgrade-pipeline "Direct link to Set up an upgrade pipeline") Use the Aiven CLI or API to create upgrade steps between your services. note The `upgrade-pipeline` CLI commands require Aiven CLI version 4.x or later. Command names and parameters may change before general availability. #### Create an upgrade step[​](#create-an-upgrade-step "Direct link to Create an upgrade step") Create a step to link a source service and a destination service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ [--source-project SOURCE_PROJECT] SOURCE_SERVICE \ [--destination-project DESTINATION_PROJECT] DESTINATION_SERVICE \ [--auto-validation-delay-days DAYS] ``` **Options** * `--organization-id` is required. * `--source-project` and `--destination-project` are optional. If you omit either project option, Aiven CLI uses the current default project set with `avn project switch`. * `--auto-validation-delay-days` is optional. Defaults to 7 days if not specified. ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "SOURCE_PROJECT_NAME", "source_service_name": "SOURCE_SERVICE_NAME", "destination_project_name": "DESTINATION_PROJECT_NAME", "destination_service_name": "DESTINATION_SERVICE_NAME", "auto_validation_delay_days": 7 }' ``` Use the [`aiven_upgrade_step`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) resource: ``` resource "aiven_upgrade_step" "example" { organization_id = "ORGANIZATION_ID" source_project_name = "SOURCE_PROJECT_NAME" source_service_name = "SOURCE_SERVICE_NAME" destination_project_name = "DESTINATION_PROJECT_NAME" destination_service_name = "DESTINATION_SERVICE_NAME" auto_validation_delay_days = 7 } ``` Apply an `UpgradePipelineStep` manifest with `kubectl`: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: upgrade-step-sample spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: SOURCE_PROJECT_NAME sourceServiceName: SOURCE_SERVICE_NAME destinationProjectName: DESTINATION_PROJECT_NAME destinationServiceName: DESTINATION_SERVICE_NAME autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable after the resource is created. Parameters: * `source_project_name`: Name of the project containing the source service * `source_service_name`: Name of the source service * `destination_project_name`: Name of the project containing the destination service * `destination_service_name`: Name of the destination service * `auto_validation_delay_days`: Optional. Number of days before automatic validation. The value must be at least `1`. The default is 7 days. The maximum delay you can configure is 30 days. #### List upgrade steps[​](#list-upgrade-steps "Direct link to List upgrade steps") View all upgrade steps you have access to: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step list --organization-id ORGANIZATION_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) To list managed upgrade steps, use: ``` terraform state list 'aiven_upgrade_step.*' terraform state show 'aiven_upgrade_step.example' ``` List `UpgradePipelineStep` resources in the current namespace: ``` kubectl get upgradepipelinesteps ``` #### View a specific step[​](#view-a-specific-step "Direct link to View a specific step") Get details about a specific upgrade step, including the last validation: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step get \ --organization-id ORGANIZATION_ID \ STEP_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` terraform state show aiven_upgrade_step.example ``` Show the manifest and full status, including `id`, `conditions`, and `lastValidation`: ``` kubectl describe upgradepipelinestep RESOURCE_NAME kubectl get upgradepipelinestep RESOURCE_NAME -o yaml ``` The step details include `last_validation` values such as `validated_at`, `validated_by_user`, and `comment` when validation exists (available through the API). ### Validate an upgrade[​](#validate-an-upgrade "Direct link to Validate an upgrade") After testing your source service with the new update, validate the version to allow the destination service to receive the same update. #### Manual validation[​](#manual-validation "Direct link to Manual validation") Validate the current version of your source service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step validate-for-service \ --project SOURCE_PROJECT \ SERVICE_NAME \ [--comment "COMMENT"] ``` `--comment` is optional. Use it to record a note about the validation, for example `"Tested and verified in development"`. ``` curl -X POST https://api.aiven.io/v1/project/SOURCE_PROJECT/service/SOURCE_SERVICE/upgrade-validation \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "comment": "Tested and verified in development" }' ``` Terraform manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. The operator manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. #### Automatic validation[​](#automatic-validation "Direct link to Automatic validation") If you do not manually validate an update, the system automatically validates the source service version after the configured delay. Auto-validation starts from when the source service receives the update. ### Manage upgrade steps[​](#manage-upgrade-steps "Direct link to Manage upgrade steps") #### Update a step[​](#update-a-step "Direct link to Update a step") Modify the automatic validation delay for an existing step: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step update \ --organization-id ORGANIZATION_ID \ --auto-validation-delay-days 14 \ STEP_ID ``` ``` curl -X PATCH https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "auto_validation_delay_days": 14 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` resource "aiven_upgrade_step" "example" { # ...required fields... auto_validation_delay_days = 14 # Updated from 7 to 14 } ``` Apply the changes: ``` terraform plan terraform apply ``` Edit `autoValidationDelayDays` in your manifest and re-apply: ``` spec: autoValidationDelayDays: 14 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable. To change them, delete the resource and create a new one. #### Delete a step[​](#delete-a-step "Direct link to Delete a step") Remove an upgrade step to allow the destination service to receive updates independently: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step delete --organization-id ORGANIZATION_ID STEP_ID ``` Find `STEP_ID` from the upgrade step list command. ``` curl -X DELETE https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) Remove the resource from configuration and apply, or destroy it directly: ``` terraform apply terraform destroy -target=aiven_upgrade_step.example ``` Delete the `UpgradePipelineStep` resource: ``` kubectl delete upgradepipelinestep RESOURCE_NAME ``` Deleting a step removes all associated validations. ### Example: Three-environment pipeline[​](#example-three-environment-pipeline "Direct link to Example: Three-environment pipeline") Create a pipeline that promotes updates from development to staging to production: * CLI * API * Terraform * Kubernetes 1. Create a step from development to staging: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project dev-project \ --destination-project staging-project \ --auto-validation-delay-days 3 \ pg-dev pg-staging ``` 2. Create a step from staging to production: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project staging-project \ --destination-project prod-project \ --auto-validation-delay-days 7 \ pg-staging pg-prod ``` 1) Create a step from development to staging: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "dev-project", "source_service_name": "pg-dev", "destination_project_name": "staging-project", "destination_service_name": "pg-staging", "auto_validation_delay_days": 3 }' ``` 2) Create a step from staging to production: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "staging-project", "source_service_name": "pg-staging", "destination_project_name": "prod-project", "destination_service_name": "pg-prod", "auto_validation_delay_days": 7 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` # Step 1: Development → Staging resource "aiven_upgrade_step" "dev_to_staging" { organization_id = "ORGANIZATION_ID" source_project_name = "dev-project" source_service_name = "pg-dev" destination_project_name = "staging-project" destination_service_name = "pg-staging" auto_validation_delay_days = 3 } # Step 2: Staging → Production resource "aiven_upgrade_step" "staging_to_prod" { organization_id = "ORGANIZATION_ID" source_project_name = "staging-project" source_service_name = "pg-staging" destination_project_name = "prod-project" destination_service_name = "pg-prod" auto_validation_delay_days = 7 } ``` Apply the configuration: ``` export PROVIDER_AIVEN_ENABLE_BETA=true terraform init terraform plan terraform apply ``` Define both steps in a single manifest and apply it: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: dev-to-staging spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: dev-project sourceServiceName: pg-dev destinationProjectName: staging-project destinationServiceName: pg-staging autoValidationDelayDays: 3 --- apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: staging-to-prod spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: staging-project sourceServiceName: pg-staging destinationProjectName: prod-project destinationServiceName: pg-prod autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-pipeline.yaml ``` When a maintenance update arrives: 1. The development service receives the update. 2. After testing, validate the development version or wait 3 days for auto-validation. 3. The staging service receives the update during its next maintenance window. 4. After testing, validate the staging version or wait 7 days for auto-validation. 5. The production service receives the update during its next maintenance window. Related pages * [Maintenance and updates for your Aiven for Metrics service](/docs/products/metrics/howto/maintenance-updates.md) * [Change the service plan](/docs/products/metrics/howto/change-service-plan.md) * [Service and feature releases](/docs/platform/concepts/service-and-feature-releases.md) * [Aiven CLI](/docs/tools/cli.md) --- # Maintenance and updates for your Aiven for Metrics service Manage maintenance updates and set the maintenance window for your Aiven for Metrics service. ## Maintenance updates[​](#maintenance-updates "Direct link to Maintenance updates") Aiven applies some maintenance updates automatically. The following are the types of updates: * **Mandatory updates:** Security updates, quarterly patch releases, and platform updates that affect reliability or stability of the service nodes. * **Optional updates:** All other updates are initially optional. After six months, they become mandatory and are applied in the next week’s maintenance window. * **Periodic infrastructure updates:** Scheduled automatically for services with nodes active for 180 days and more. These updates are mandatory for all services, except those with maintenance turned off. Critical security updates are applied during the next available maintenance window. For other updates, Aiven gives you at least seven days' notice. Maintenance updates are also automatically applied during service upgrades. To view pending updates: * Console * CLI * API 1. In your service, click **Service settings**. 2. Go to the **Service management** section. Use the [`avn service get`](/docs/tools/cli/service-cli.md#avn_service_get) command. Use the [`service`](https://api.aiven.io/doc/#tag/Service/operation/ServiceGet) endpoint. ## Maintenance window[​](#maintenance-window "Direct link to Maintenance window") The maintenance window is the time period when Aiven can automatically apply maintenance updates to a service. When an update becomes available, Aiven schedules it for the next available maintenance window for each service. The update runs in the first window after it becomes available, and can begin any time after the start time. For example, if a service has a maintenance window of Monday 12:00 UTC, and an update becomes available on Tuesday, the update will be applied on the following Monday. During maintenance, Aiven might restart or replace service nodes. This can cause brief connection interruptions, but services are designed to minimize downtime. Aiven performs maintenance in a rolling-forward style, creating new nodes alongside existing ones and retiring the old nodes after the upgrade completes. Major service upgrades are triggered manually. A manually triggered upgrade starts immediately, regardless of the maintenance window. important You cannot control the order in which services are updated. Each service updates according to its own configured maintenance window, and there is no guaranteed way to control the update sequence. Manual updates and maintenance window adjustments only help for non-critical updates. ## Set the maintenance window[​](#set-the-maintenance-window "Direct link to Set the maintenance window") To set the maintenance window for your service: * Console * Terraform 1. In the Aiven Console, open your service. 2. In the **Maintenance** section, click **Actions** > **Change maintenance window**. 3. Set the day and time. 4. Click **Save changes**. Use the `maintenance_window_dow` and `maintenance_window_time` attributes in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Change the service plan](/docs/products/metrics/howto/change-service-plan.md) --- # Power on/off and delete your Aiven for Metrics service Power off your Aiven for Metrics service to release resources and save credits, power it back on when you need it, or delete it permanently. ## Power off a service[​](#power-off-a-service "Direct link to Power off a service") When you power off a service: * All virtual machines are removed from the public cloud. * The service configuration is stored on the Aiven Platform. * If there are no backups, all service data is lost. * If the service has time-based or point in time recovery backups, the backups remain on the Aiven Platform. Services powered off for more than 180 days are automatically deleted. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power off service**. To power off a service, run: ``` avn service update SERVICE_NAME --power-off ``` ## Power on a service[​](#power-on-a-service "Direct link to Power on a service") When you power on a service: * New virtual machines are created on the service's public cloud. * The service starts with the stored configuration parameters. * The latest time-based backup is restored. * Maintenance updates are automatically applied. * If a point in time recovery backup is available, the database transaction logs are replayed to recover the service data to a specific point in time. The restoration takes from a few minutes to a few hours, depending on the network bandwidth, the disk IOPS allocated to the service, and the size of the backup. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power on service**. To power on a service, run: ``` avn service update SERVICE_NAME --power-on ``` To see when the service is running, run: ``` avn service wait SERVICE_NAME ``` note Aiven for Metrics stores historical metrics in object storage, so that data is retained when the service is powered off and remains available after you power it back on. note Static IP addresses are not removed when a service is powered off or deleted. They continue to generate the usual costs. To avoid these costs, [remove the static IP addresses](/docs/platform/concepts/static-ips.md). ## Delete a service[​](#delete-a-service "Direct link to Delete a service") * Console * CLI 1. In your project, click **Services**. 2. Open the service to delete, and click **Actions** > **Delete service**. To delete a service, run: ``` avn service terminate SERVICE_NAME ``` Related pages * [Monitor storage usage](/docs/products/metrics/howto/storage-usage.md) * [Aiven for Metrics](/docs/products/metrics.md) --- # Prepare your Aiven for Metrics service for high load Prepare your Aiven for Metrics service for higher than usual traffic to avoid outages and keep performance stable. Prepare your services for higher than usual traffic to avoid service outages by doing the following: * **Subscribe to service notifications:** To receive notifications about service health and warnings when resources are low, you can [set service and project contacts](https://aiven.io/docs/platform/howto/technical-emails). You can also view the status of the Aiven Platform and get updates on incidents on the [status page](https://status.aiven.io/). Follow the RSS feed, subscribe to email or SMS notifications, or use the Slack integration to get notifications about incidents. * **Monitor your services:** [Monitor the health of your services](/docs/platform/howto/list-monitoring.md) using metrics, logs, alerts, and dashboards. * **Scale your services:** If you forecast a load that can't be handled by the service, you can scale up your service. * **Set the backup schedule:** To minimize the impact of the higher load during the backup process, schedule backups outside of peak traffic hours. * **Set the maintenance window:** Schedule maintenance updates outside of your peak traffic hours. * **Run load tests on service forks:** To test the impact of high traffic on a production service, fork the service and run your load test on the fork. Additionally, optimizing a service allows it to perform better under stress therefore avoiding the need of an upgrade. The more optimized a service is for your usage, the better you can weather spikes in traffic. Related pages * [Change the service plan](/docs/products/metrics/howto/change-service-plan.md) * [Maintenance and updates](/docs/products/metrics/howto/maintenance-updates.md) --- # Manage storage Get a comprehensive view of your storage usage with Aiven for Metrics. The Storage page in the [Aiven console](https://console.aiven.io/) provides valuable data to help you effectively manage and understand your object storage consumption and associated costs. Key metrics provided on the Storage page include: * **Current billing expenses**: Shows the current cost incurred for the storage used. * **Forecasted monthly cost**: Provides an estimate of the expected cost for the upcoming month based on usage. * **Object storage used**: Indicates the amount of storage currently in use. * **Retention rule**: Displays the current data retention setting, which by default is set to **Keep data forever**. No external Thanos access While Aiven for Metrics uses the Thanos architecture internally, **you cannot connect your Aiven-managed services to any external (non-Aiven) Thanos endpoint**. Connect to Aiven for Metrics instead. ## Access the Storage page[​](#access-the-storage-page "Direct link to Access the Storage page") 1. Log into the [Aiven console](https://console.aiven.io/), select your project, and select your Aiven for Metrics service. 2. Click **Storage** in the sidebar. ## Manage storage settings[​](#manage-storage-settings "Direct link to Manage storage settings") Adjust your data retention settings and set up alerts to keep your storage in check. ### Edit retention rules[​](#edit-retention-rules "Direct link to Edit retention rules") 1. In **Storage** page, click **Actions** > Edit retention rule. 2. Select **Keep data forever** or **Keep data for (days)** and enter the number of days. 3. Click **Save**. ### Set or edit storage alert threshold[​](#set-or-edit-storage-alert-threshold "Direct link to Set or edit storage alert threshold") 1. In **Storage** page, click **Actions** > **Set storage alert threshold** or **Edit storage alert threshold**. 2. To set a new threshold, click **Set storage alert threshold**. To change an existing one, click **Edit storage alert threshold**. 3. Set your threshold in GiB. 4. Click **Save**. --- # Tag your Aiven for Metrics service Add key-value tags to your Aiven for Metrics service to organize services and track ownership, cost allocation, and governance. Use tags to add metadata to Aiven services to categorize them or run custom logic on them. Typical uses include: * Tagging for governance to deploy services with specific tags only. * Tagging for internal cost reporting, ownership, allocation, and accountability. A tag is a key/value pair: * **Key**: A case-sensitive string that starts with a letter and consists of letters, numbers, dashes, and underscores. The maximum length for a key is 64 characters. * **Value**: A string value limited to 64 UTF-8 characters. Within a service, the tag keys must be unique. * Console * Terraform 1. In the service, click **Service settings**. 2. In the **Service status** section, click **Actions** > **Add service tags**. 3. Enter a key and value for each tag. 4. Click **Save changes**. Use the `tag` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Aiven for Metrics](/docs/products/metrics.md) --- # Track restore progress for your Aiven for Metrics service Track the restore progress of individual nodes in your Aiven for Metrics service during node replacement, forking, or maintenance, using the Aiven API. You can track restore progress for individual nodes during service node replacement by using the Aiven API. For example, use this endpoint to monitor the restore progress of a forked service or when applying maintenance. The service object exposes restore progress under `node_states[].progress_updates`: * `service.node_states[]` contains per-node state entries. * When a node is restoring or catching up, its `state` is typically `syncing_data`. * When the state is `syncing_data`, the node may include `progress_updates` with one or more phase objects. * Other node states don't include restore progress data. note `progress_updates` may be missing or empty even when a node is in `syncing_data`. This can occur when a restore completes before detailed progress is reported or when the service does not emit detailed progress counters. ## API endpoints[​](#api-endpoints "Direct link to API endpoints") Restore progress fields are part of the standard service response payload. * Get a single service (recommended for polling): `GET /project/{project}/service/{service_name}` * List services in a project: `GET /project/{project}/service` - Request - Response ``` curl -H "Authorization: aivenv1 API_TOKEN" https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME ``` Replace the following placeholders: * `API_TOKEN`: Your Aiven API token. * `PROJECT_NAME`: Your Aiven project name. * `SERVICE_NAME`: The name of your service. ``` { "service": { ... "node_states": [ { "node_name": "...", "state": "syncing_data", "progress_updates": [ { "phase": "basebackup", "completed": false, "current": 3410567, "min": 0, "max": 7569280, "unit": "bytes_uncompressed" } ] } ], ... } } ``` ## Node states[​](#node-states "Direct link to Node states") Common values for `node_states[].state` include: * `setting_up_vm`: The virtual machine is being created or initialized. * `syncing_data`: The node is restoring data or catching up. * `running`: The node is operating normally. * `leaving`: The node is leaving the cluster. * `unknown`: A transient or error state. ## `progress_updates` data model[​](#progress_updates-data-model "Direct link to progress_updates-data-model") `progress_updates` is a list of phase objects. When present, phases appear in the following order: 1. `prepare` 2. `basebackup` 3. `stream` 4. `finalize` Each phase object includes the following fields: ``` { "completed": false, "current": 3410567, "max": 7569280, "min": 0, "phase": "basebackup", "unit": "bytes_uncompressed" } ``` ### Field semantics[​](#field-semantics "Direct link to Field semantics") * `phase`: String, required. The restore phase. Possible values: `prepare`, `basebackup`, `stream`, and `finalize`. * `completed`: Boolean, required. Whether the phase is complete. * `current`: Number or null, optional. The current progress value. This field can be missing or null. * `min`: Number or null, optional. The starting value for the phase. This field can be missing or null. * `max`: Number or null, optional. The expected total value for the phase. This value can be missing, null, or change while the restore is in progress. * `unit`: String or null, optional. The unit for `current`, `min`, and `max`. New unit values can appear over time. Important considerations * Treat `unit` as an opaque identifier. Unknown values can appear. * `max` may change while a restore is in progress. * Not all phases report numeric counters. Some services only indicate phase completion. ## Why `max` values can change[​](#why-max-values-can-change "Direct link to why-max-values-can-change") The `current`, `min`, and `max` values are best-effort progress indicators. They can be based on estimates or on system state that changes over time. Treat `max` as the latest known expected total, not as a fixed guarantee. Common reasons `max` can change include: * The restore process discovers additional work after it starts, such as files, segments, or objects that become visible only after metadata is read. * New data is added on the backend while the node is catching up, which moves the completion point forward. This is common during incremental catch-up phases. * Progress is calculated from system state, such as replication lag, rather than from a fixed work queue. As the system state changes, the value is recalculated. * The service switches restore strategies during the operation, for example from snapshot restore to replication catch-up, which changes what the counters represent. As a result: * Phase percentage can decrease even when the restore operates normally. * Remaining-time estimates based on `max` are unreliable. * Sudden changes in `max` are expected unless the node remains in `syncing_data` longer than expected. ## Restore phase meanings[​](#restore-phase-meanings "Direct link to Restore phase meanings") Phase names are standardized, but the underlying work and the meaning of the counters are service-specific. * `prepare`: Prepares the node for restore. * `basebackup`: Restores the full backup. * `stream`: Applies incremental changes, such as replication or log replay. * `finalize`: Completes final steps before serving traffic. Not all restores include every phase. ## Compute phase progress percentages[​](#compute-phase-progress-percentages "Direct link to Compute phase progress percentages") You cannot reliably compute overall restore progress. You can compute a phase-specific progress percentage when `min`, `max`, and `current` are present and `max != min`. ``` pct = round(((current - min) / (max - min)) * 100, 1) ``` When handling progress values: * If any of `min`, `max`, or `current` is null or missing, display `n/a`. * If `max == min`, treat the percentage as undefined. * Expect the percentage to decrease when `max` changes. * Clamp displayed values to the range `[0, 100]`. ## Polling guidance[​](#polling-guidance "Direct link to Polling guidance") Progress updates are best-effort and refresh every 10 seconds while a node is in `syncing_data`. Poll the service state every 10 to 30 seconds. More frequent polling does not provide additional detail. For each `node_states[]` entry: * If `state` is not `syncing_data`, no restore progress is available. * If `state` is `syncing_data`: * If `progress_updates` is missing or empty, the node is restoring without detailed progress data. * Otherwise, the current phase is the last phase where `completed` is `false`. Stop polling when all nodes reach the `running` state or when a stall is detected. ### Stall detection[​](#stall-detection "Direct link to Stall detection") The API does not provide per-phase timestamps. To detect stalls, use a time-based threshold, such as a node remaining in `syncing_data` longer than expected. Do not rely on counters or `max` values to estimate remaining time. Related pages * [Maintenance and updates](/docs/products/metrics/howto/maintenance-updates.md) * [Change cloud or region](/docs/products/metrics/howto/change-cloud-region.md) --- # Maintenance and lifecycle in Aiven for Metrics Manage maintenance updates, maintenance windows, and node restore progress for your Aiven for Metrics service. Related pages * [Maintenance and updates](/docs/products/metrics/howto/maintenance-updates.md) * [Track restore progress](/docs/products/metrics/howto/track-restore-progress.md) --- # Scaling and performance in Aiven for Metrics Scale resources and prepare your Aiven for Metrics service for changing load. Related pages * [Change the service plan](/docs/products/metrics/howto/change-service-plan.md) * [Prepare for high load](/docs/products/metrics/howto/prepare-for-high-load.md) * [Memory and out-of-memory conditions](/docs/products/metrics/concepts/service-memory.md) --- # Aiven for MySQL® Aiven for MySQL® is a fully managed relational database service, deployable in the cloud of your choice. With Aiven for MySQL, you can store, retrieve, or modify your data. This service is scalable and available at a size to suit your needs, from single-node starter plans to highly available production platforms. MySQL has been a key part of the open source database landscape for a long time and it's a popular and reliable database platform. Aiven takes care of the management side and provides a MySQL that you can use. ## [Get started](/docs/products/mysql/get-started.md) [2 items](/docs/products/mysql/get-started.md) ## [Connect to service](/docs/products/mysql/howto/list-code-samples.md) [9 items](/docs/products/mysql/howto/list-code-samples.md) ## [Query and analyze data](/docs/products/mysql/howto/create-database.md) [7 items](/docs/products/mysql/howto/create-database.md) ## [Service management](/docs/products/mysql/howto/power-cycle-service.md) [7 items](/docs/products/mysql/howto/power-cycle-service.md) ## [Scaling and performance](/docs/products/mysql/scaling-performance.md) [9 items](/docs/products/mysql/scaling-performance.md) ## [Maintenance and lifecycle](/docs/products/mysql/maintenance-lifecycle.md) [3 items](/docs/products/mysql/maintenance-lifecycle.md) ## [High availability and disaster recovery](/docs/products/mysql/concepts/high-availability.md) [3 items](/docs/products/mysql/concepts/high-availability.md) ## [Backups and migration](/docs/products/mysql/backups-migration.md) [2 items](/docs/products/mysql/backups-migration.md) Related pages * [Aiven.io](https://aiven.io/mysql) * [MySQL proprietary documentation](https://dev.mysql.com/doc/refman/8.4/en/) (upstream project documentation) * [Blog post about MyHoard](https://aiven.io/blog/introducing-myhoard-your-single-solution-to-mysql-backups-and-restoration), our MySQL backup and restore tool * [GitHub repository for the project](https://github.com/aiven/myhoard) --- # Backups and migration in Aiven for MySQL® Back up, restore, and migrate your Aiven for MySQL® service data. Related pages * [Understand MySQL backups](/docs/products/mysql/concepts/mysql-backups.md) * [Use incremental backups](/docs/products/mysql/howto/use-incremental-backups.md) * [Migrate to Aiven for MySQL](/docs/products/mysql/howto/migrate-db-to-aiven-via-console.md) --- # High availability of Aiven for MySQL® Aiven for MySQL® is available on a variety of plans, offering different levels of service availability. The selected plan defines the features available. | Plan | Service availability features | Backup history | | ------------ | -------------------------------------------------------------------------- | -------------- | | **Hobbyist** | Single node (limited availability) | 2 days | | **Startup** | Single node (limited availability) | 2 days | | **Business** | 1 primary node and 1 standby node (higher availability) | 14 days | | **Premium** | 1 primary node and 2 standby nodes (top high availability characteristics) | 30 days | ## Primary and standby nodes[​](#primary-and-standby-nodes "Direct link to Primary and standby nodes") Aiven's Business and Premium plans offer a primary node and standby nodes. A standby node is useful for multiple reasons: * Provides another physical copy of the data in case of hardware, software, or network failures * Typically reduces the data loss window in disaster scenarios * Provides a quicker database time to restore with a controlled failover in case of failures as the standby is already installed, running, and synchronised with the data * Can be used for read-only queries to reduce the load on the primary server ## Failure handling[​](#failure-handling "Direct link to Failure handling") ### Minor failures[​](#minor-failures "Direct link to Minor failures") Minor failures, such as service process crashes or temporary loss of network access, are handled automatically by Aiven in all plans without any major changes to the service deployment. The service automatically restores normal operation once the crashed process is automatically restarted or when network access is restored. ### Severe failures[​](#severe-failures "Direct link to Severe failures") Severe failures, such as losing a node entirely in case of hardware or severe software problems, require radical recovery measures. The Aiven monitoring infrastructure automatically detects a failing node when the node starts reporting issues in the self-diagnostics or when it stops communicating. In such cases, the monitoring infrastructure automatically schedules a new replacement node to be created. note In the event of database failover, the **Service URI** of your service remains the same; only the IP address changes to point to the new primary node. ## Highly available Business and Premium service plans[​](#highly-available-business-and-premium-service-plans "Direct link to Highly available Business and Premium service plans") When a standby node fails, the primary node keeps running normally and provides a normal service level to the client applications. Once the new replacement standby node is ready and synchronised with the primary node, it starts replicating the primary node in real time as the situation gets back to normal. When a primary node fails, the combined information from the Aiven monitoring infrastructure and the standby node is used to make a failover decision. The standby node is promoted as the new primary and immediately starts serving clients. A new replacement node is automatically scheduled and becomes the new standby node. If the primary node and all the standby nodes fail at the same time, new nodes are automatically scheduled for creation to become the new primary and standby. The primary node is restored from the latest available backup, which can involve some degree of data loss. Any write operations made since the backup of the latest binary log file are lost. Typically, this time window is limited to either five minutes or one binary log file. note The amount of time required to replace a failed node depends mainly on the selected cloud region and the amount of data to be restored. In the case of partial loss of the cluster, the surviving node keeps on serving clients even during the recreation of the other node. This happens automatically and requires no administrator intervention. **Premium** plans operate in a similar way as **Business** plans. The main difference comes when one of the standby nodes or the primary node fails. Premium plans have an additional, redundant standby node available, providing platform availability even in the event of losing two nodes. If the primary node fails, the Aiven monitoring infrastructure determines which of the standby nodes is the furthest along in replication (has the least potential for data loss) and does a controlled failover to that node. note For backups and restoration, Aiven uses [MyHoard](https://aiven.io/blog/introducing-myhoard-your-single-solution-to-mysql-backups-and-restoration). ## Single-node Hobbyist and Startup service plans[​](#single-node-hobbyist-and-startup-service-plans "Direct link to Single-node Hobbyist and Startup service plans") Hobbyist and Startup plans provide a single node. When it's lost, Aiven immediately starts the automatic process of creating a new replacement node. The new node starts up, restores its state from the latest available backup, and resumes serving clients. Since there is just a single node providing the service, the service is unavailable for the duration of the restoration. In addition, any write operations made since the backup of the latest binary log file are lost. Typically, this time window is limited to either five minutes or one binary log file. --- # MySQL max\_connections Calculate the total number of simultaneous connections available to all users combined on your Aiven for MySQL® service, and learn why the per-user connection limit is a separate setting. The maximum number of simultaneous connections in Aiven for MySQL® depends on how much RAM your service has. This `max_connections` value applies to all users of the service combined, not to each user individually. note Independent of the size, an `extra_connection` with a value of `1` is added for the system process. ## Under 4 GiB[​](#under-4-gib "Direct link to Under 4 GiB") For services with less than 4 GiB of RAM, the number of allowed connections is per GiB: max\_connections=75×RAM+extra\_connection Example With 2 GiB of RAM, the maximum number of connections is max\_connections=75×2+1 ## 4 GiB or more[​](#4-gib-or-more "Direct link to 4 GiB or more") For services with 4 GiB or more of RAM, the number of allowed connections is per GiB: max\_connections=100×RAM+extra\_connection Example With 7 GiB of RAM, the maximum number of connections is max\_connections=100×7+1 ## Increase the maximum number of connections[​](#increase-the-maximum-number-of-connections "Direct link to Increase the maximum number of connections") `max_connections` isn't a configurable advanced parameter in Aiven for MySQL. Its value is always calculated from your service's plan and RAM using the formulas described earlier. To raise the total number of simultaneous connections available to your service, [upgrade to a plan with more RAM](/docs/products/mysql/howto/change-service-plan.md). ## Per-user connection limits[​](#per-user-connection-limits "Direct link to Per-user connection limits") MySQL supports a separate, per-user limit called `max_user_connections`, which caps how many simultaneous connections a single database user can open. If a client reports an error such as `User 'exampleuser' has exceeded the 'max_user_connections' resource`, that user has hit this per-user limit, not the service-wide `max_connections` limit. Aiven for MySQL doesn't expose `max_user_connections` as a configurable option, and raising your service's `max_connections` value doesn't change any per-user limit. If your application needs more concurrent connections for a specific database user, [distribute connections across multiple database users](/docs/products/mysql/howto/manage-service-users.md) or reduce the number of concurrent connections that user opens. For help with a persistent per-user connection limit, contact [Aiven support](/docs/platform/howto/support.md). Related pages * [Prepare your Aiven for MySQL service for high load](/docs/products/mysql/howto/prepare-for-high-load.md) * [MySQL tuning for concurrency](/docs/products/mysql/concepts/mysql-tuning-and-concurrency.md) * [Change the service plan](/docs/products/mysql/howto/change-service-plan.md) * [Manage service users](/docs/products/mysql/howto/manage-service-users.md) --- # Understand MySQL backups in Aiven Aiven for MySQL databases are automatically backed-up, with full backups daily, and binary logs recorded continuously. The number of stored backups and backup retention time depends on your [Aiven service plan](https://aiven.io/pricing?product=mysql\&tab=plan-comparison). Full backups are version-specific binary backups, which when combined with [binlog](https://dev.mysql.com/doc/internals/en/binary-log-overview.html) allow for consistent recovery to a specific point in time (PITR). important One thing to consider is that you may modify the backup time configuration option in **Advanced configuration** in [Aiven Console](https://console.aiven.io) which will begin shifting the backup schedule to the new time. If there was a recent backup taken, it may take another backup cycle before it starts applying new backup time. note To be able to safely make backups, MySQL INSTANT ALTER TABLE always use the INPLACE or COPY algorithm instead of INSTANT. Specifying ALGORITHM=INSTANT does not fail but automatically falls back to INPLACE or COPY as needed. ## MySQL backups and encryption[​](#mysql-backups-and-encryption "Direct link to MySQL backups and encryption") All Aiven for MySQL backups use the [myhoard software](https://github.com/aiven/myhoard) to perform encryption. Myhoard utilizes [Percona XtraBackup](https://www.percona.com/) internally for taking a full (or [incremental](/docs/products/mysql/howto/use-incremental-backups.md)) snapshot of MySQL. Since [Percona XtraBackup 8.0.23](https://jira.percona.com/browse/PXB-1979) the `--lock-ddl` option is enabled by default. This ensures that DDL changes cannot be performed while a full backup process is ongoing. This is important to guarantee that the backup service is consistent and can be reliably used for restoration. With this feature enabled, if you try to run `CREATE`, `ALTER`, `DROP`, `TRUNCATE` or another command during backup, you may receive the message **Waiting for backup lock**. In this case, wait till the backup is complete for running such operations. ## Configure the backup schedule[​](#configure-the-backup-schedule "Direct link to Configure the backup schedule") Set the time of day when the daily backup is taken. To edit the backup schedule for your service: * Console * Aiven API * Aiven CLI * Terraform 1. In your service, **Backups**. 2. Click **Actions** > **Configure backup settings**. 3. Click **Add configuration options**. 4. Add `backup_hour` and `backup_minute`, and set their values. 5. Click **Save configuration**. Call the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint, and add the following properties to the `user_config` object: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "backup_hour": BACKUP_HOUR, "backup_minute": BACKUP_MINUTE } }' ``` Replace the following: * `SERVICE_NAME`: the name of your service. * `PROJECT_NAME`: the name of your project. * `BACKUP_HOUR`: the hour when the service backup starts. Accepted values are integers between `0` and `23`. * `BACKUP_MINUTE`: the minute when the service backup starts. Accepted values are integers between `0` and `59`. Run the [avn service update](/docs/tools/cli/service-cli.md#avn-cli-service-update) command, and add the following properties to the `user_config` object: ``` avn service update SERVICE_NAME \ --project PROJECT_NAME \ --user-config '{ "backup_hour": BACKUP_HOUR, "backup_minute": BACKUP_MINUTE }' ``` Replace the following: * `SERVICE_NAME`: the name of your service. * `PROJECT_NAME`: the name of your project. * `BACKUP_HOUR`: the hour when the service backup starts. Accepted values are integers between `0` and `23`. * `BACKUP_MINUTE`: the minute when the service backup starts. Accepted values are integers between `0` and `59`. Use the `backup_hour` and `backup_minute` attributes in [your service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs) to set the start time for backups. If a backup was recently made, it can take another backup cycle before the new backup time takes effect. ## More resources[​](#more-resources "Direct link to More resources") * [Use incremental backups](/docs/products/mysql/howto/use-incremental-backups.md) * Our blog post: [MyHoard, your solution to MySQL backups and restoration](https://aiven.io/blog/introducing-myhoard-your-single-solution-to-mysql-backups-and-restoration) * Read about [Aiven cloud security and data encryption](/docs/platform/concepts/cloud-security.md#data-encryption) --- # Aiven for MySQL® free tier Use Aiven for MySQL® for free. You don't need a credit card to sign up and you can use it indefinitely free of charge. ## Features and limitations[​](#features-and-limitations "Direct link to Features and limitations") Free MySQL services include: * A single node * 1 CPU per virtual machine * 1 GB RAM * 1 GB disk storage * Monitoring for metrics and logs * Backups There are some limitations of the free tier: * Cannot create the service in a VPC * No static IPs * No integrations * No forking * `max_connections` limit set to `76` * No support services * Only one service of each service type in your [organization](/docs/platform/concepts/orgs-units-projects.md) * Not covered under Aiven's 99.99% SLA Free services do not have any time limitations. However, Aiven reserves the right to: * Power off free services with no initial usage within the first few hours after the service is running. You can power them back on at any time. * Power off free services with no continuative activity on the service. A notification is sent before the service is powered off. You can power them back on at any time. * Shut down services if Aiven believes they violate the [acceptable use policy](https://aiven.io/terms). * Change the cloud provider, region, or configuration at any time. ## Upgrade or downgrade a free service[​](#upgrade-or-downgrade-a-free-service "Direct link to Upgrade or downgrade a free service") You can upgrade your free service to a paid plan at any time by adding a payment method to the project's billing group. To upgrade a free service: 1. Go to the service **Overview** page. 2. In the **Service plan usage** section, click **Upgrade plan**. The upgrade happens immediately; however, it can take up to 3 hours for Basic tier support to be available. You can also downgrade a paid plan to the free tier as long as: * The data you have in that trial or paid service fits into the smaller instance size. * The free tier is available in the same cloud as the paid plan. --- # Understand MySQL memory usage MySQL memory utilization can appear high, even if the service is relatively idle. All services are subject to operating overhead, but some services, including MySQL, pre-allocate memory. This can lead to a false impression that the service is misbehaving, when it is actually operating under normal conditions. ## The InnoDB buffer pool[​](#the-innodb-buffer-pool "Direct link to The InnoDB buffer pool") Arguably, the most important MySQL component is the InnoDB Buffer Pool. Every time an operation happens to a table (read or write), the page where the records (and indexes) are located is loaded into the Buffer Pool. This means that if the data you read and write the most has its pages in the Buffer Pool, the performance will be better than if you have to read pages from disk. When there are no more free pages in the pool, older pages must be evicted and if they were modified, synchronized back to disk (checkpointing). The [MySQL Reference](https://dev.mysql.com/doc/refman/8.4/en/innodb-buffer-pool.html) says: ``` The buffer pool is an area in main memory where InnoDB caches table and index data as it is accessed. The buffer pool permits frequently used data to be accessed directly from memory, which speeds up processing. ``` And [How MySQL Uses Memory](https://dev.mysql.com/doc/refman/8.4/en/memory-use.html) says: ``` InnoDB allocates memory for the entire buffer pool at server startup, using `malloc()` operations. The `innodb_buffer_pool_size` system variable defines the buffer pool size. Typically, a recommended `innodb_buffer_pool_size` value is 50 to 75 percent of system memory. ``` However, the InnoDB buffer pool isn't the only pre-allocated buffer. ## Global buffers[​](#global-buffers "Direct link to Global buffers") The InnoDB buffer pool is part of the global buffers MySQL allocates to improve performance of database operations. An explanation of these various buffers (or code areas) can be found in the MySQL documentation: [How MySQL Uses Memory](https://dev.mysql.com/doc/refman/8.4/en/memory-use.html). Using a 4 GB service as an example, a view of the global buffers shows what memory has been allocated: ``` SELECT SUBSTRING_INDEX(event_name,'/',2) AS code_area, format_bytes(SUM(current_alloc)) AS current_alloc FROM sys.x$memory_global_by_current_bytes GROUP BY SUBSTRING_INDEX(event_name,'/',2) ORDER BY SUM(current_alloc) DESC; +---------------------------+---------------+ | code_area | current_alloc | +---------------------------+---------------+ | memory/innodb | 1.37 GiB | | memory/performance_schema | 213.22 MiB | | memory/sql | 18.73 MiB | | memory/mysys | 8.82 MiB | | memory/temptable | 1.00 MiB | | memory/mysqld_openssl | 459.71 KiB | | memory/mysqlx | 3.25 KiB | | memory/myisam | 728 bytes | | memory/csv | 120 bytes | | memory/vio | 80 bytes | +---------------------------+---------------+ ``` Although allocated, the buffers may be relatively empty: ``` SELECT CONCAT(FORMAT(A.num * 100.0 / B.num,2),'%') `BufferPool %` FROM (SELECT variable_value num FROM performance_schema.global_status WHERE variable_name = 'Innodb_buffer_pool_pages_data') A, (SELECT variable_value num FROM performance_schema.global_status WHERE variable_name = 'Innodb_buffer_pool_pages_total') B; +--------------+ | BufferPool % | +--------------+ | 0.45% | +--------------+ ``` important **High memory utilization, by itself, does not indicate a performance issue**. In the above example, the global buffers consume \~1.6 GiB of memory; almost half the RAM on a 4 GB service. However, this does not denote any particular issue, but rather, standard operating conditions. When memory issues are suspected or the service encounters out-of-memory conditions, the buffers, queries, and concurrency should be examined to determine if: * The buffer pool is full and checkpointing frequently * The sum of the buffer pools are greater than the available service memory * Queries are generating excessive temporary (spill) files ## Service memory limits[​](#service-memory-limits "Direct link to Service memory limits") The practical memory limit will always be less than the service physical memory limit. **All services are subject to operating overhead:** * A small amount of memory is required by the operating system kernel to manage system resources, including networking functions and disk cache. * Aiven's cloud data platform requires memory to monitor availability, provide metrics, logging and manage backups. A server or node's **usable memory** can be calculated as: `usable memory = RAM - overhead` Where: * `overhead` is 350 MiB (≈ 0.34 GiB). Services may utilize optional components, service integrations, connection pooling, or plug-ins, which are not included in overhead calculations. If a service is overcommitted, the operating system, management layer, backups or availability monitoring, may fail status checks or operations due to resource contention. In severe instances, the node may fail completely with an out-of-memory condition. ## Out of memory conditions[​](#out-of-memory-conditions "Direct link to Out of memory conditions") Many processes request more memory from the kernel than they will ever use or need. In these cases, the kernel overallocates memory. This allows it to satisfy multiple processes requesting more memory than is available, which is not used or is freed by the time any other process actually needs it. However, if enough processes start using all their allocated memory simultaneously there may not be enough physical memory available and an `Out Of Memory` (`OOM`) condition occurs. warning This situation is critical and must be resolved immediately. The solution that the Linux kernel employs is to invoke the `Out of Memory Killer` (or `OOM Killer`). This reviews all running processes and kills one or more of them to free up system memory and keep the system running. The `OOM Killer` selects process to kill based on an `oom_score`; a calculation that balances how much memory the process is using with how long the process has been running. Processes that have been running for a long time are less likely to be killed. Subprocesses are summed with parent processes in terms of memory usage, so a process which forks many subprocesses, but itself does not use a lot of memory, may still be killed. In most instances, the hosted data service, or a child process, will have the highest memory footprint and be a prime candidate for termination when the OOM Killer inspects the running processes. Aiven's cloud data platform leverages kernel namespaces or containers to isolate processes from each other. Isolation has several benefits, including: * A smaller footprint for security‑related concerns * A smaller blast radius for failure * Greater control of system resources Left unchecked, the `OOM Killer` may opt to kill the primary service. This is undesirable as unclean termination of the primary service can lead to data loss, inconsistency, or corrupted backups. Further, if Aiven's management platform detects that the primary service is unavailable for , the service will be marked as down and a failover will occur. To mitigate this scenario, namespaces are used, some with additional memory limits, in combination with an `oom_score_adjust` on the primary process, to coax the `OOM Killer` into selection of less critical processes. This will still result in a service restart, but in a more controlled process, where the database is shut down, rather than killed; exposure to data loss is limited and recovery is faster when the service restarts, often avoiding failover. warning Out of Memory conditions can still lead to unexpected behavior, including data unavailable or data loss conditions. ## Avoid running low on memory[​](#avoid-running-low-on-memory "Direct link to Avoid running low on memory") The OOM killer only runs when the system is critically low on memory. To prevent it from running, either reduce your memory usage or increase the available memory. For most databases, the service memory footprint can often be reduced by: * Reducing concurrency or implementing connection pooling * Tuning queries to limit result sets * Tuning indexes for query load * Dropping unused objects from storage In cases where the working set no longer fits into memory, consider scaling your service. Related pages * [Change the service plan](/docs/products/mysql/howto/change-service-plan.md) * [Scale disk storage](/docs/products/mysql/howto/scale-disk-storage.md) --- # Understand MySQL replication in Aiven Replication in Aiven for MySQL® is always based on replicating logical changes. This means that the replication protocol may contain an actual statement that the target server should apply or it may have an entry saying **update row with these old attributes to have these new attributes**. These are called statement and row formats. The statement format is more compact but can't represent all changes because some statements would yield different results if executed as is on different servers. The row format statement can represent all changes, and it allows using tools like Debezium since the binary log contains full details of all changes in itself. For these reasons, Aiven uses row format by default. Read more about the comparison between MySQL statement based and row based replication on the [Advantages and Disadvantages of Statement-Based and Row-Based Replication](https://dev.mysql.com/doc/refman/8.4/en/replication-sbr-rbr.html) article. The row based replication works very well as long as the tables being replicated have a primary key. MySQL primary key look ups are very fast and the target server can find the rows to update and delete very quickly. However, when the table being replicated is lacking a primary key, the target server needs to make a sequential table scan for each individual update or delete statement and the replication can become extremely slow if the table is large. ``` DELETE FROM nopk WHERE modified_time > '2022-01-13' ``` If a statement like the above matched 500 rows and the table had a million rows altogether, the row based replication format would contain 500 individual delete operations and the target server needed to do sequential scan over the one million rows for each of the individual deletions, which can take tens of minutes. If the table had a primary key, the same statement would likely be replicated in under a second. ## Replication use in Aiven for MySQL[​](#replication-use-in-aiven-for-mysql "Direct link to Replication use in Aiven for MySQL") The considerations presented on the [Replication overview](/docs/products/mysql/concepts/mysql-replication.md) section are not only valid for services that actually have standby nodes or read-only replicas. Whenever the Aiven management platform needs to create a node for a service, the node is first initialized from backup to most recent backed up state. This includes applying the full replication stream that has been created after the most recent full base backup. Once the latest backed up state has been restored the node will connect to the current master, if available, and replicate latest state from that, which is also affected by possible replication slowness. When new nodes are created, it needs to perform replication and having large tables without primary keys may make operations such as replacing failed nodes, upgrading service plan, migrating service to a different cloud provider or region, starting up new read-only replica service, forking a service, and some others to take extremely long time or depending on the situation practically not complete at all without manual operator intervention (for example, new read-only replica might never be able to catch up with existing master because replication is too slow). To work around these issues Aiven operations people may need to resort to operations such as temporarily making `master` read only or promoting a replacement server before it has fully applied the replication stream, resulting in data loss. To make the service operate correctly and avoid such drastic measures you should ensure the primary keys exist for any tables that are not trivially small. You can check the article how to [create missing primary keys](/docs/products/mysql/howto/create-missing-primary-keys.md) to ensure primary keys exists in your Aiven for MySQL service. --- # MySQL tuning for concurrency Determining how much memory is available for queries, and tuning concurrency accordingly, requires calculation of service memory, query analysis, and monitoring. Global buffers, thread buffers, and some uncontrolled memory allocations (`TRIGGERS`, `PROCEDURES` and `FUNCTIONS`), all contribute to the memory MySQL will require for a given workload. There are several key calculations which are fundamental to tuning: * Service memory * Global buffers * Thread buffers * Concurrency important Query output is for reference only. Queries should be run per service for accuracy and re-evaluated periodically for change. ## Service memory[​](#service-memory "Direct link to Service memory") The service memory can be calculated as: where the overhead is currently . ## Global buffers[​](#global-buffers "Direct link to Global buffers") **MySQL pre-allocates global buffers to improve performance of database operations.** An explanation of these various buffers (or code areas) can be found in the MySQL documentation: [How MySQL Uses Memory](https://dev.mysql.com/doc/refman/8.4/en/memory-use.html). ``` SELECT SUBSTRING_INDEX(event_name,'/',2) AS code_area, format_bytes(SUM(current_alloc)) AS current_alloc FROM sys.x$memory_global_by_current_bytes GROUP BY SUBSTRING_INDEX(event_name,'/',2) ORDER BY SUM(current_alloc) DESC; +---------------------------+---------------+ | code_area | current_alloc | +---------------------------+---------------+ | memory/innodb | 1.37 GiB | | memory/performance_schema | 213.22 MiB | | memory/sql | 18.73 MiB | | memory/mysys | 8.82 MiB | | memory/temptable | 1.00 MiB | | memory/mysqld_openssl | 459.71 KiB | | memory/mysqlx | 3.25 KiB | | memory/myisam | 728 bytes | | memory/csv | 120 bytes | | memory/vio | 80 bytes | +---------------------------+---------------+ ``` ## Thread buffers[​](#thread-buffers "Direct link to Thread buffers") **Thread buffers are memory allocated per thread (or connection) to the database.** Queries may use part or all of the allocation. ``` SELECT ( @@read_buffer_size + @@read_rnd_buffer_size + @@sort_buffer_size + @@join_buffer_size + @@binlog_cache_size + @@thread_stack + @@tmp_table_size + 2*@@net_buffer_length ) / (1024 * 1024) AS MEMORY_PER_CON_MB; +-------------------+ | MEMORY_PER_CON_MB | +-------------------+ | 17.9375 | +-------------------+ ``` important The actual amount of memory a query can use is technically unbounded. Uncontrolled memory allocations and temporary table usage can adversely affect memory allocation. The data dictionary size is based on the number of tables, fields and indexes within the database. ## Concurrency[​](#concurrency "Direct link to Concurrency") Aiven configures a default value for the `max_connections` parameter for all MySQL services. The [max\_connections](/docs/products/mysql/concepts/max-number-of-connections.md) parameter is based off the service usable memory. ``` select @@max_connections; +-------------------+ | @@max_connections | +-------------------+ | 226 | +-------------------+ ``` important This parameter should be used as a guideline only. By default, `max_connections` is configured for *optimistic* concurrency using all available memory. In many instances, if the `max connections` are fully utilized, resource overcommitment and Out of memory conditions will occur. At \~18 MB per connection, a 4 GiB service has a potential memory usage of 4068 MB (18 \* 226). This is less than the service RAM, but exceeds the service memory limit. **For performance and stability, the following calculation is recommended:** max\_concurrency= This value may be pessimistic for a workload that does not require the full thread buffer, but is an advisable starting point for concurrency testing and monitoring. Concurrency can be incremented, if service memory permits. --- # Get started with Aiven for MySQL® Start using Aiven for MySQL® by creating a service, connecting to it, and loading sample data. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Console * Terraform - Access to the [Aiven Console](https://console.aiven.io) - [MySQL CLI client](https://dev.mysql.com/doc/refman/8.4/en/mysql.html) installed * [Terraform installed](https://www.terraform.io/downloads) * A [personal token](https://docs.aiven.io/docs/platform/howto/create_authentication_token.html) * [MySQL CLI client](https://dev.mysql.com/doc/refman/8.4/en/mysql.html) installed ## Create a service[​](#create-a-service "Direct link to Create a service") * Console * Terraform 1. In your project, click **Services**. 2. Click **Create service**. 3. Select **MySQL**. 4. Select a **Service tier**. 5. Select a **Cloud**. note You cannot choose a cloud provider or a specific cloud region on the Free tier. 6. Select a **Plan**. note The plans available can vary between cloud providers and regions for the same service. 7. In the **Service details**, enter a name for your service. 8. Optional: Add service tags. 9. In the **Service summary**, click **Create service**. The status of the service is **Rebuilding** during its creation. When the status is **Running**, you can start using the service. This typically takes a couple of minutes and can vary between cloud providers and regions. The following example files are also available in the [Aiven Terraform Provider repository](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/mysql) on GitHub. 1. Create a file named `provider.tf` and add the following: ``` Loading... ``` 2. Create a file named `service.tf` and add the following: ``` Loading... ``` 3. Create a file named `variables.tf` and add the following: ``` Loading... ``` 4. Create a file named `terraform.tfvars` and add values for the variables without defaults: * `aiven_token`: your token * `aiven_project_name`: the name of one of your Aiven projects * `mysql_password`: a password for the service user 5. To output connection details, create a file named `output.tf` and add the following: ``` Loading... ``` To apply your Terraform configuration: 1. Initialize Terraform by running: ``` terraform init ``` The output is similar to the following: ``` Initializing the backend... Initializing provider plugins... - Finding aiven/aiven versions matching ">= 4.0.0, < 5.0.0"... - Installing aiven/aiven v4.9.2... - Installed aiven/aiven v4.9.2 ... Terraform has been successfully initialized! ... ``` 2. To create an execution plan and preview the changes, run: ``` terraform plan ``` 3. To deploy your changes, run: ``` terraform apply --auto-approve ``` ## Configure a service[​](#configure-a-service "Direct link to Configure a service") Edit your service settings if the default service configuration doesn't meet your needs. * Console * Terraform 1. Select the new service from the list of services on the **Services** page. 2. On the **Overview** page, select **Service settings** from the sidebar. 3. In the **Advanced configuration** section, make changes to the service configuration. See the available configuration options in [Advanced parameters for Aiven for MySQL](/docs/products/mysql/reference/advanced-params.md). See [the `aiven_mysql` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/mysql) for the full schema. ## Connect to the service[​](#connect-to-service "Direct link to Connect to the service") * Console * Terraform * mysql 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your organization > project > Aiven for MySQL service. 2. On the **Overview** page of your service, click **Quick connect**. 3. In the **Connect** window, select a tool or language to connect to your service, follow the connection instructions, and click **Done**. ``` mysql --user avnadmin --password=ADMIN_PASSWORD --host mysql-sakila-dev-sandbox.f.aivencloud.com --port 12691 defaultdb ``` Access your new service with [the MySQL client](/docs/products/mysql/howto/connect-from-cli.md) using the outputs. 1. To store the outputs in environment variables, run: ``` MYSQL_HOST="$(terraform output -raw mysql_service_host)" MYSQL_PORT="$(terraform output -raw mysql_service_port)" MYSQL_USER="$(terraform output -raw mysql_service_username)" MYSQL_PASSWORD="$(terraform output -raw mysql_service_password)" ``` 2. To use the environment variables with the MySQL client to connect to the service, run: ``` mysql --host=$MYSQL_HOST --port=$MYSQL_PORT --user=$MYSQL_USER --password=$MYSQL_PASSWORD --database defaultdb ``` [Connect to your new service](/docs/products/mysql/howto/connect-from-cli.md) with [mysql](https://dev.mysql.com/doc/refman/8.4/en/mysql.html). tip Check more tools for connecting to Aiven for MySQL in [Connect to Aiven for MySQL](/docs/products/mysql/howto/list-code-samples.md). ## Load a test dataset[​](#load-a-test-dataset "Direct link to Load a test dataset") `Sakila` is a sample dataset that represents a DVD rental store. It provides a standard schema highlighting MySQL features. 1. Download the `sakila` database archive (`tar` or `zip` format) from the [MySQL example databases](https://dev.mysql.com/doc/sakila/en/sakila-installation.html) page, and extract it to your desired location (for example `/tmp/`). 2. From the folder where you unpacked the archive, [connect to your MySQL service](/docs/products/mysql/howto/connect-from-cli.md), create a `sakila` database, and connect to it: ``` CREATE DATABASE sakila; USE sakila; ``` 3. Populate the database: ``` source sakila-schema.sql; source sakila-data.sql; ``` 4. Verify what objects have been created: ``` SHOW FULL TABLES; ``` Expected output ``` +----------------------------+------------+ | Tables_in_sakila | Table_type | +----------------------------+------------+ | actor | BASE TABLE | | actor_info | VIEW | | address | BASE TABLE | | category | BASE TABLE | | city | BASE TABLE | | country | BASE TABLE | | customer | BASE TABLE | | customer_list | VIEW | | film | BASE TABLE | | film_actor | BASE TABLE | | film_category | BASE TABLE | | film_list | VIEW | | film_text | BASE TABLE | | inventory | BASE TABLE | | language | BASE TABLE | | nicer_but_slower_film_list | VIEW | | payment | BASE TABLE | | rental | BASE TABLE | | sales_by_film_category | VIEW | | sales_by_store | VIEW | | staff | BASE TABLE | | staff_list | VIEW | | store | BASE TABLE | +----------------------------+------------+ 23 rows in set ``` ## Query data[​](#query-data "Direct link to Query data") ### Read data[​](#read-data "Direct link to Read data") Retrieve all the data from a table, for example, from `language`: ``` SELECT * FROM language; ``` Expected output ``` +-------------+----------+---------------------+ | language_id | name | last_update | +-------------+----------+---------------------+ | 1 | English | 2006-02-15 05:02:19 | | 2 | Italian | 2006-02-15 05:02:19 | | 3 | Japanese | 2006-02-15 05:02:19 | | 4 | Mandarin | 2006-02-15 05:02:19 | | 5 | French | 2006-02-15 05:02:19 | | 6 | German | 2006-02-15 05:02:19 | +-------------+----------+---------------------+ 6 rows in set ``` ### Write data[​](#write-data "Direct link to Write data") Add a row to a table, for example, to `category`: ``` INSERT INTO category(category_id,name) VALUES(17,'Thriller'); ``` Expected output ``` Query OK, 1 row affected ``` Check that your new row is there: ``` SELECT * FROM category WHERE name = 'Thriller'; ``` Expected output ``` +-------------+----------+---------------------+ | category_id | name | last_update | +-------------+----------+---------------------+ | 17 | Thriller | 2024-05-22 11:04:03 | +-------------+----------+---------------------+ 1 row in set ``` Related pages * [Connect to Aiven for MySQL with MySQL Workbench](/docs/products/mysql/howto/connect-from-mysql-workbench.md) * [Migrate to Aiven from an external MySQL](/docs/products/mysql/howto/migrate-from-external-mysql.md) * [Create additional databases](/docs/products/mysql/howto/create-database.md) * [Aiven Service Level Agreement](https://aiven.io/sla) --- # AI database optimizer for Aiven for MySQL® Use **Aiven AI Database Optimizer** to receive optimization suggestions to your databases and queries. Aiven's artificial intelligence considers various aspects to suggest optimizations, for example query structure, table size, existing indexes and their cardinality, column types and sizes, and the connections between the tables and columns in the query. To optimize a query automatically: 1. In the [Aiven Console](https://console.aiven.io/login), open your Aiven for MySQL® service. 2. In the **Observe** section, click **AI insights**. 3. For the query of your choice, click **Optimize**. 4. In the **Query optimization report** window, see the optimization suggestion and apply the suggestion by running the provided SQL queries. * To display potential alternative optimization recommendations, click **Advanced options**. * To display the diff view, click **Query diff**. * To display explanations about the optimization, click **Optimization details**. note The quality of the optimization suggestions is proportional to the amount of data collected about the performance of your database. Frequently asked questions **Does Aiven AI Optimizer mask/obfuscate my queries?** Yes, Aiven AI Optimizer provides a non-intrusive solution to optimize your database performance without compromising sensitive data access. It achieves this by gathering information on schema structure, database statistics, and other signals to detect potential performance problems and offer optimization recommendations, without requiring credentials or access to the actual data in the database. To address the possibility of slow query logs containing sensitive data, Aiven offers data masking capabilities that replace sensitive parameters within queries with question marks (`?`). Data masking is enabled by default. note The masking option is not available for the Standalone SQL query optimizer yet. For one-time query optimizations when you do not run an Aiven for PostgreSQL® service, use the [standalone SQL query optimizer](https://aiven.io/tools/sql-query-optimizer). Related pages * [AI DB Optimizer for Aiven for PostgreSQL®](/docs/products/postgresql/howto/ai-insights.md) * [Standalone SQL query optimizer](https://aiven.io/tools/sql-query-optimizer) --- # Back up your Aiven for MySQL® service to another region Copy your Aiven for MySQL® service backups to a secondary region for disaster recovery. In addition to the primary service backup, you can have a secondary backup in an alternative location. important This feature is in [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-). Contact your account team to enable it. Backup to another region (BTAR) is a disaster recovery feature that allows backup files to be copied from the service's primary backup region to an additional (secondary) region. BTAR can bolster data resilience and helps improve data protection against disasters in the primary backup region. When the primary region is down, BTAR allows forking the service from an additional copy of the backup residing in a secondary region. ## Limitations[​](#limitations "Direct link to Limitations") * The cloud provider for your additional backup region must match the cloud provider for your service and the primary backup. * Secondary backup can only be restored in the region where it was stored. For a service that has the backup to another region (BTAR) feature enabled, you can check the service backup status, change the backup region, monitor the replication lag, fork and restore using the cross-region backup, or migrate to another cloud or region. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * At least one Aiven service with BTAR enabled * Access to the [Aiven Console](https://console.aiven.io/) * [Aiven API](/docs/tools/api.md) * [Aiven CLI](/docs/tools/cli.md) ## Change a backup region[​](#change-a-backup-region "Direct link to Change a backup region") 1. Log in to the [Aiven Console](https://console.aiven.io/) and go to your project. 2. On the **Services** page, select an Aiven service on which you'd like to enable BTAR. 3. On your service page, **Backups**. 4. On the **Backups** page, click **Actions** > **Edit secondary backup location**. 5. In the **Edit secondary backup location** window, use the **Backup location** menu to select a region for your additional backup. Confirm your choice by selecting **Save**. important You can change the backup region once in 24 hours. ## Monitor a service with BTAR[​](#monitor-a-service-with-btar "Direct link to Monitor a service with BTAR") There are a few things you may want to check for your Aiven service in the context of BTAR: * What is the status of a secondary backup? * Does your service have a backup in another region? * What is the target region of the secondary backup? * What is the replication lag between data availability in the primary region and the secondary region? ### Check BTAR status[​](#check-btar-status "Direct link to Check BTAR status") To see the availability, the status, and the target region of a secondary (BTAR) backup in the [Aiven Console](https://console.aiven.io/), go to your service page > **Backups** > **Secondary backup location**. ### Determine replication lag[​](#determine-replication-lag "Direct link to Determine replication lag") Determine the target region and the replication lag for a secondary (BTAR) backup of your service, call the [ServiceBackupToAnotherRegionReport](https://api.aiven.io/doc/#tag/Service/operation/ServiceBackupToAnotherRegionReport) endpoint. Configure the call as follows: 1. Enter `YOUR-PROJECT-NAME` and `YOUR-SERVICE-NAME` into the URL. 2. Specify `DESIRED-TIME-PERIOD` depending on the time period you need the metrics for: select one of the following values for the `period` key: `hour`, `day`, `week`, `month`, or `year`. ``` curl --request POST \ --url https://api.aiven.io/v1/project/YOUR-PROJECT-NAME/service/YOUR-SERVICE-NAME/backup_to_another_region/report \ --header 'Authorization: Bearer YOUR-BEARER-TOKEN' \ --header 'content-type: application/json' \ --data '{"period":"DESIRED-TIME-PERIOD"}' ``` As output, you get metrics including replication lags at specific points in time. ## Fork and restore a service with BTAR[​](#fork-and-restore "Direct link to Fork and restore a service with BTAR") You can use the [Aiven Console](/docs/tools/aiven-console.md), [API](/docs/tools/api.md), or [CLI](/docs/tools/cli.md) to recover your service from a backup in another region. To restore your service using BTAR, create a fork of the original service in the region where the secondary backup resides. note When you **fork & restore** from the secondary backup, your new fork service is created in the cloud and region where the secondary backup is located. The fork service gets the same plan that the primary service uses. Backups of the fork service are located in the region where this new service is hosted. * Aiven Console * Aiven CLI * Aiven API 1. Open the [Aiven Console](https://console.aiven.io/) and go to your service homepage. 2. Click **Backups**. 3. On the **Backups** page, select **Fork & restore**. 4. In the **New database fork** window: 1. Set **Backup location** to either **Primary location** or **Secondary location**. 2. Set **Backup version** to one of the following: * **Latest transaction** * **Point in time**: Set it up to no earlier than the time of taking the oldest replicated base backup. 3. Specify a name for the new fork service. 4. Select **Create fork**. Run the [avn service create](/docs/tools/cli/service-cli.md#avn-cli-service-create) command with the `--service-to-fork-from` option and the `--recovery-target-time`option. Set `--recovery-target-time` to no earlier than the time of taking the oldest replicated base backup. ``` avn service create FORK_SERVICE_NAME \ --plan SERVICE_PLAN \ --project PROJECT_NAME \ --service-type SERVICE_TYPE \ --cloud SECONDARY_BACKUP_REGION \ --recovery-target-time "YYYY-MM-DDTHH:MM:SS+00:00" \ --service-to-fork-from PRIMARY_SERVICE_NAME ``` Replace the following with meaningful data: * `FORK_SERVICE_NAME` * `SERVICE_PLAN` * `PROJECT_NAME` * `SERVICE_TYPE` * `SECONDARY_BACKUP_REGION` * `PRIMARY_SERVICE_NAME` * `YYYY-MM-DDTHH:MM:SS+00:00` Use the [ServiceCreate](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) API to create a fork service. When constructing the API request, add the `user_config` object to the request body and nest the `service_to_fork_from` field and the `recovery_target_time` field inside. Set `recovery_target_time` to no earlier than the time of taking the oldest replicated base backup. ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "cloud": "SECONDARY_BACKUP_REGION", "plan": "SERVICE_PLAN", "service_name": "FORK_SERVICE_NAME", "service_type": "SERVICE_TYPE", "user_config": { "service_to_fork_from": "PRIMARY_SERVICE_NAME", "recovery_target_time": "YYYY-MM-DDTHH:MM:SS+00:00" } }' ``` Replace the following with meaningful data: * `FORK_SERVICE_NAME` * `SERVICE_PLAN` * `PROJECT_NAME` * `SERVICE_TYPE` * `SECONDARY_BACKUP_REGION` * `PRIMARY_SERVICE_NAME` * `YYYY-MM-DDTHH:MM:SS+00:00` ## Migrate a service with BTAR[​](#migrate-a-service-with-btar "Direct link to Migrate a service with BTAR") You can migrate a service with BTAR the same way you [migrate a service with a regular backup](/docs/platform/howto/migrate-services-cloud-region.md). note When you migrate your service, locations of service backups, both primary and secondary ones, do not change. ## Delete a cross-region backup[​](#delete-a-cross-region-backup "Direct link to Delete a cross-region backup") Delete an additional service backup created in a region different from your primary backup region. You can delete a cross-region backup using the [Aiven Console](/docs/tools/aiven-console.md), [API](/docs/tools/api.md), or [CLI](/docs/tools/cli.md). When you delete the additional cross-region backup, you still have the default backup located in the primary, service-hosting region. * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. From the **Services** view, select an Aiven service on which you'd like to disable BTAR. 3. On your service's page, click **Backups**. 4. On the **Backups** page, click **Actions** > **Secondary backup location**. 5. In the **Edit secondary backup location** window, select **Disable**. Your additional service backup is no longer visible on your service's **Backups** page in the **Secondary backup location** column. To remove secondary backups for your service, use the [avn service update](/docs/tools/cli/service-cli.md) command to remove all target region names from the `additional_backup_regions` array. ``` avn service update your-sevice-name \ -c additional_backup_regions=\[\] ``` To remove secondary backups for your service, update the service configuration. Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint to remove all target regions names from the `additional_backup_regions` array. ``` curl --request PUT \ --url https://api.aiven.io/v1/project/YOUR_PROJECT_NAME/service/YOUR_SERVICE_NAME \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "additional_backup_regions": [] } }' ``` Related pages * [Backups](/docs/products/mysql/concepts/mysql-backups.md) * [Track restore progress](/docs/products/mysql/howto/track-restore-progress.md) --- # Change the cloud or region for your Aiven for MySQL® service Move your Aiven for MySQL® service to a different cloud provider or region. 1. In your service, click **Service settings** from the sidebar. 2. In the **Cloud and network** section, click **Actions** > **Change cloud**. 3. In the **Cloud** section , select a cloud provider and region, and click **Change**. Your service starts a migration to the new location and remains available during the process. When the migration completes, the service continues running in the new cloud or region. Related pages * [Fork your service](/docs/products/mysql/howto/fork-service.md) * [Migrate to another cloud or region](/docs/platform/howto/migrate-services-cloud-region.md) --- # Change the plan for your Aiven for MySQL® service Change the service plan for your Aiven for MySQL® service to scale resources up or down and optimize costs. Adjust the plan of your services at any time to scale your services as needed and optimize costs. If you can't find a suitable plan, you can [request a custom plan](/docs/platform/concepts/service-pricing.md). tip If you plan to upgrade your service plan, do it immediately after a full backup. This reduces the amount of incremental changes that need to be applied on top of the base backup, which speeds up the upgrade itself. important * When changing a service plan, reserve an additional 25% of disk space. This requirement applies to upgrades and downgrades. * Downgrading to a plan with fewer VMs is supported for most services, including Aiven for Apache Kafka®, Aiven for PostgreSQL®, Aiven for OpenSearch®, Aiven for ClickHouse®, Aiven for MySQL®, Aiven for Metrics, and Aiven for Valkey™. * Changing a service plan triggers a node recycle, service rebuilding, and any pending maintenance updates. - Console - Terraform - CLI 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Change plan**. 3. Select a plan that provides at least 125% of the current disk size and click **Change plan**. Update the `plan` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). To change a service plan in the Aiven CLI, use the [`avn service update --plan `](/docs/tools/cli/service-cli.md#avn-cli-service-update) command. Your service's state becomes **Rebuilding** and remains accessible. When the state switches to **Running**, your new service plan is active. Related pages * [Scale disk storage](/docs/products/mysql/howto/scale-disk-storage.md) * [Prepare for high load](/docs/products/mysql/howto/prepare-for-high-load.md) --- # Connect to Aiven for MySQL® from the command line Connect to your Aiven for MySQL® service via the command line with the following tools: * [mysqlsh shell](/docs/products/mysql/howto/connect-from-cli.md#connect-mysqlsh) * [mysql client](/docs/products/mysql/howto/connect-from-cli.md#connect-mysql) ## Using `mysqlsh`[​](#connect-mysqlsh "Direct link to connect-mysqlsh") ### Variables[​](#variables "Direct link to Variables") These are the placeholders to replace in the code sample: | Variable | Description | | ------------- | --------------------------------------------------------------------------------------------------------------------- | | `SERVICE_URI` | URL for the MySQL connection, from [Aiven Console](https://console.aiven.io/) > the **Overview** page of your service | ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") For this example you need the `mysqlsh` client installed. You can install this by following the [MySQL shell installation documentation](https://dev.mysql.com/doc/mysql-shell/8.4/en/mysql-shell-install.html). ### Code[​](#code "Direct link to Code") From a terminal window to connect to the MySQL database, run: ``` mysqlsh --sql SERVICE_URI ``` Test your setup with a query: ``` MySQL ssl defaultdb SQL> select 1 + 2 as three; +-------+ | three | +-------+ | 3 | +-------+ 1 row in set (0.0539 sec) ``` ## Using `mysql`[​](#using-mysql "Direct link to using-mysql") ### Variables[​](#variables-1 "Direct link to Variables") These are the placeholders to replace in the code sample: | Variable | Description | | --------------- | ------------------------------------------------ | | `USER_HOST` | Hostname for MySQL connection | | `USER_PORT` | Port for MySQL connection | | `USER_PASSWORD` | Password of your Aiven for MySQL connection | | `DB_NAME` | Database Name of your Aiven for MySQL connection | ### Prerequisites[​](#connect-mysql "Direct link to Prerequisites") For this example you need the `mysql` client installed. You can install it by following the [MySQL client installation documentation](https://dev.mysql.com/doc/refman/8.4/en/mysql.html). ### Code[​](#code-1 "Direct link to Code") This step requires to manually specify individual parameters. You can find those parameters in [Aiven Console](https://console.aiven.io) > the **Overview** page of your service. Once you have these parameters, execute the following from a terminal window to connect to the MySQL database: ``` mysql --user avnadmin --password=USER_PASSWORD --host USER_HOST --port USER_PORT DB_NAME ``` warning If you are providing the password via the command line, you must pass it as shown; putting a space between the parameter name and value will cause the password to be parsed incorrectly. --- # Connect to Aiven for MySQL® with MySQL Workbench You can use a graphical client like [MySQL Workbench](https://www.mysql.com/products/workbench/) to connect to Aiven for MySQL® services. ## Connect to Aiven for MySQL®[​](#connect-to-aiven-for-mysql "Direct link to Connect to Aiven for MySQL®") Enter the individual connection parameters as shown in [Aiven Console](https://console.aiven.io/) (the **Overview** page of your service > the **Connection information** section) and also download the SSL CA certificate and specify the file on SSL page. ![Screenshot of the MySQL Workbench settings screen](/docs/assets/images/mysql-workbench-0fd4b95bb86c1fa581aaa2719dc970b6.png) important Using SSL is strongly recommended. To secure your connection, download the CA certificate and configure it in client settings. ## Create an additional database[​](#create-an-additional-database "Direct link to Create an additional database") To create more databases, go to the service's page in [Aiven Console](https://console.aiven.io/). In the **Connect** section, click **Databases**. In the **Databases** view, click **Create database**, enter a name for your database and click **Add database**. ## Add a database user[​](#add-a-database-user "Direct link to Add a database user") To add database users, go to [Aiven Console](https://console.aiven.io/) and select your Aiven for MySQL service from the **Services** page. In your service's page, in the **Connect** section, click **Users**. In the **Users** view, click **Add service user**. In the **Create a service user** window, you can choose the authentication method to use. By default, the web console uses the `caching_sha2_password` authentication mechanism. To successfully connect, your client libraries need to be new enough. If for any reason you are forced to use a client that only supports the older `mysql_native_password` authentication mechanism, select this separately while adding the user. tip You can change this later in the **Users** view for your service ([Aiven Console](https://console.aiven.io/)). note Changing the authentication method for a user resets the password. on top of the authentication method, one more input item required from you in the **Create a service user** window is a name for the user. Enter it into the **Username** field and select **Add service user**. --- # Connect to Aiven for MySQL® using MySQLx with Python Enabling the MySQLx protocol support allows you to use your MySQL instance as a document store. This example shows how to connect to your Aiven for MySQL® instance using MySQLx protocol. note MySQL initially provided support for X-DevAPI (MySQLx) in v5.7.12 as an optional extension that you can install. On the MySQL v8.0+, the X-DevAPI is supported by default. ## Variables[​](#variables "Direct link to Variables") | Variable | Description | | ----------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `SERVICE_URI` | Service URI from [Aiven Console](https://console.aiven.io/) > the **Overview** page of your service > the **Connection information** section > the **MySQLx** tab | | `MYSQLX_USER` | User from [Aiven Console](https://console.aiven.io/) > the **Overview** page of your service > the **Connection information** section > the **MySQLx** tab | | `MYSQLX_PASSWORD` | Password from [Aiven Console](https://console.aiven.io/) > the **Overview** page of your service > the **Connection information** section > the **MySQLx** tab | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Python 3.7 or later * A `mysqlx` python library installed: ``` pip install mysql-connector-python ``` * An Aiven account with an Aiven for MySQL service running * Set environment variable `PROTOCOL_BUFFERS_PYTHON_IMPLEMENTATION=python` to avoid issues described on [Protocol buffers docs](https://developers.google.com/protocol-buffers/docs/news/2022-05-06). If you are running Python from the command line, you can set this in your terminal: ``` export PROTOCOL_BUFFERS_PYTHON_IMPLEMENTATION=python ``` ## Code[​](#code "Direct link to Code") Add the following to `main.py` and replace the placeholders with values for your project: ``` import mysqlx connection_data = f"mysqlx://{MYSQLX_USER}:{MYSQLX_PASSWORD}@{SERVICE_URI}/defaultdb?ssl-mode=REQUIRED" session = mysqlx.get_session(connection_data) # create a test schema schema = session.create_schema("test") # create a new collection in the schema collection = schema.create_collection("food_prices") # add entries to this collection collection.add( {"type": "pizza", "price": "10e"}, {"type": "burger", "price": "5e"}, ).execute() # read it back for doc in collection.find().execute().fetch_all(): print(f"Found document: {doc}") ``` This code creates a MySQL client and connects to the database via the MySQLx protocol. It creates a schema, a collection, inserts some entries, fetches them, and prints the output. If the script runs successfully, the output will be the values that were inserted into the document: ``` Found document: {"_id": "000062c55a6b0000000000000001", "type": "pizza", "price": "10e"} Found document: {"_id": "000062c55a6b0000000000000002", "type": "burger", "price": "5e"} ``` Now that your application is connected, you are all set to use Python with Aiven for MySQL using the MySQLx protocol. --- # Connect to Aiven for MySQL® with DataGrip Use [DataGrip](https://www.jetbrains.com/datagrip/) to connect to your Aiven for MySQL® service. import RelatedPages from "@site/src/components/RelatedPages"; ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to the [Aiven Console](https://console.aiven.io/) * At least one running Aiven for MySQL service * [DataGrip](https://www.jetbrains.com/datagrip/download/) installed on your machine ## Get JDBC URI from Aiven Console[​](#get-jdbc-uri-from-aiven-console "Direct link to Get JDBC URI from Aiven Console") 1. Log in to [Aiven Console](https://console.aiven.io/) and go to your organization > project > Aiven for MySQL service. 2. On the service **Overview** page, click **Quick connect**. 3. In the **Connect** window: 1. Choose to connect with Java using the **Connect with** dropdown menu. 2. Copy the generated JDBC URI. 3. Click **Done**. ## Connect to JDBC URI from DataGrip[​](#connect-to-jdbc-uri-from-datagrip "Direct link to Connect to JDBC URI from DataGrip") 1. Open DataGrip on your machine, and select **File** > **New** > **Data Source** > **MySQL** from the top navigation menu. 2. In the **Data Sources and Drivers** window > **General** tab, paste the URI copied from the [Aiven Console](https://console.aiven.io/). 3. Click **OK** to create and save the connection. The connection to your Aiven for MySQL service has been established and is visible in DataGrip > **Database Explorer**. Related pages * [Connect to Aiven for MySQL](/docs/products/mysql/howto/list-code-samples.md) for more tools you can use for connecting to your service * [DataGrip](https://www.jetbrains.com/datagrip/) * [DataGrip download](https://www.jetbrains.com/datagrip/download/) --- # Connect to Aiven for MySQL® with DBeaver Use [DBeaver](https://dbeaver.com/) to connect to your Aiven for MySQL® service. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to the [Aiven Console](https://console.aiven.io/) * At least one running Aiven for MySQL service * [DBeaver](https://dbeaver.io/download/) installed on your machine ## Get JDBC URI from Aiven Console[​](#get-jdbc-uri-from-aiven-console "Direct link to Get JDBC URI from Aiven Console") 1. Log in to [Aiven Console](https://console.aiven.io/) and go to your organization > project > Aiven for MySQL service. 2. On the service **Overview** page, click **Quick connect**. 3. In the **Connect** window: 1. Choose to connect with Java using the **Connect with** dropdown menu. 2. Copy the generated JDBC URI. 3. Click **Done**. ## Connect to JDBC URI from DBeaver[​](#connect-to-jdbc-uri-from-dbeaver "Direct link to Connect to JDBC URI from DBeaver") 1. Open DBeaver on your machine, and select **Database** > **New Database Connection** from the top navigation menu. 2. In the **Connect to database** window, select MySQL and click **Next**. 3. In the **Connection Settings** window > **Main** tab > **Server** section, choose to connect with URL and paste the URI copied from the [Aiven Console](https://console.aiven.io/). 4. Click **Finish** to create and save the connection. The connection to your Aiven for MySQL service has been established and is visible in DBeaver > **Database Navigator**. Related pages * [Connect to Aiven for MySQL](/docs/products/mysql/howto/list-code-samples.md) for more tools you can use for connecting to your service * [DBeaver](https://dbeaver.com/) * [DBeaver Community](https://dbeaver.io/) --- # Connect to Aiven for MySQL® with Java This example connects your Java application to an Aiven for MySQL® service. ## Variables[​](#variables "Direct link to Variables") Replace the following placeholders in the code sample: | Variable | Description | | ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `MYSQL_HOST` | Host name for the connection, from [Aiven Console](https://console.aiven.io/) > the **Overview** page of your service > the **Connection information** section | | `MYSQL_PORT` | Port number to use, from [Aiven Console](https://console.aiven.io/) > the **Overview** page of your service > the **Connection information** section | | `MYSQL_PASSWORD` | Password for `avnadmin` user | | `MYSQL_DATABASE` | Database to connect | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * JDK 1.8+ * MySQL JDBC Driver, which you can install: * Manually from [MySQL Community Downloads](https://dev.mysql.com/downloads/connector/j/) * Using maven: ``` mvn org.apache.maven.plugins:maven-dependency-plugin:2.8:get -Dartifact=com.mysql:mysql-connector-j:8.4.8:jar -Ddest=mysql-connector-j-8.4.8.jar ``` ## Code[​](#code "Direct link to Code") Add the following to `MySqlExample.java`: ``` import java.sql.Connection; import java.sql.DriverManager; import java.sql.ResultSet; import java.sql.SQLException; import java.sql.Statement; import java.util.Locale; public class MySqlExample { public static void main(String[] args) throws ClassNotFoundException { String host, port, databaseName, userName, password; host = port = databaseName = userName = password = null; for (int i = 0; i < args.length - 1; i++) { switch (args[i].toLowerCase(Locale.ROOT)) { case "-host": host = args[++i]; break; case "-username": userName = args[++i]; break; case "-password": password = args[++i]; break; case "-database": databaseName = args[++i]; break; case "-port": port = args[++i]; break; } } // JDBC allows to have nullable username and password if (host == null || port == null || databaseName == null) { System.out.println("Host, port, database information is required"); return; } Class.forName("com.mysql.cj.jdbc.Driver"); try (final Connection connection = DriverManager.getConnection("jdbc:mysql://" + host + ":" + port + "/" + databaseName + "?sslmode=require", userName, password); final Statement statement = connection.createStatement(); final ResultSet resultSet = statement.executeQuery("SELECT version() AS version")) { while (resultSet.next()) { System.out.println("Version: " + resultSet.getString("version")); } } catch (SQLException e) { System.out.println("Connection failure."); e.printStackTrace(); } } } ``` This code creates a MySQL client and connects to the database. It fetches version of MySQL and prints it the output. Run the code after replacement of the placeholders with values for your project: ``` javac MySqlExample.java && java -cp "mysql-connector-j-8.4.8.jar;." MySqlExample -host MYSQL_HOST -port MYSQL_PORT -database MYSQL_DATABASE -username avnadmin -password MYSQL_PASSWORD ``` If the script runs successfully, the output will be the values that were inserted into the table: ``` Version: 8.4.8 ``` Now that your application is connected, you are all set to use Java with Aiven for MySQL. --- # Connect to Aiven for MySQL® with PHP This example connects to an Aiven for MySQL® service from PHP, making use of the built-in PDO module. ## Variables[​](#variables "Direct link to Variables") These are the placeholders to replace in the code sample: | Variable | Description | | ----------- | ------------------------------------------------------------------------------------------------------------------------- | | `MYSQL_URI` | Service URI for MySQL connection, from [Aiven Console](https://console.aiven.io/) > the **Overview** page of your service | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Download CA certificates](/docs/platform/concepts/tls-ssl-certificates.md#download-ca-certificates) from [Aiven Console](https://console.aiven.io/) > the **Overview** page of your service. This example assumes it is in a local file called `ca.pem`. * Make sure you have read/write permissions to the `ca.pem` file and you add an absolute path to this file into [the code](/docs/products/mysql/howto/connect-with-php.md#connect-mysql-php-code): ``` $conn .= ";sslmode=verify-ca;sslrootcert='D:/absolute/path/to/ssl/certs/ca.pem'" ``` note Your PHP installation needs to include the [MySQL functions](https://www.php.net/manual/en/ref.pdo-pgsql.php) (most installations have this already). ## Code[​](#connect-mysql-php-code "Direct link to Code") Add the following to `index.php` and replace the placeholder with the MySQL URI: ``` query("SELECT VERSION()"); print($stmt->fetch()[0]); } catch (Exception $e) { echo "Error: " . $e->getMessage(); } ``` This code creates a MySQL client and opens a connection to the database. It then runs a query checking the database version and prints the response. note This example replaces the query string parameter to specify `sslmode=verify-ca` to make sure that the SSL certificate is verified, and adds the location of the cert. Run the following code: ``` php index.php ``` If the script runs successfully, the output is the MySQL version running in your service like: ``` 8.4.8 ``` --- # Connect to Aiven for MySQL® with Python This example connects your Python application to an Aiven for MySQL® service, using the [PyMySQL](https://github.com/PyMySQL/PyMySQL) library. ## Variables[​](#variables "Direct link to Variables") These are the placeholders to replace in the code sample: | Variable | Description | | ---------------- | --------------------------------------------------------------------------------------------------------------------- | | `MYSQL_HOST` | Host name for the connection, from [Aiven Console](https://console.aiven.io/) > the **Overview** page of your service | | `MYSQL_PORT` | Port number to use, from [Aiven Console](https://console.aiven.io/) > the **Overview** page of your service | | `MYSQL_USERNAME` | User to connect with | | `MYSQL_PASSWORD` | Password for this user | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") For this example you will need: * Python 3.7 or later * The Python `PyMySQL` library. You can install this with `pip`: ``` pip install pymysql ``` * Install the `cryptography` package: ``` pip install cryptography ``` ## Code[​](#code "Direct link to Code") Add the following to `main.py` and replace the placeholders with values for your project: ``` import pymysql timeout = 10 connection = pymysql.connect( charset="utf8mb4", connect_timeout=timeout, cursorclass=pymysql.cursors.DictCursor, db="defaultdb", host=MYSQL_HOST, password=MYSQL_PASSWORD, read_timeout=timeout, port=MYSQL_PORT, user=MYSQL_USERNAME, write_timeout=timeout, ) try: cursor = connection.cursor() cursor.execute("CREATE TABLE mytest (id INTEGER PRIMARY KEY)") cursor.execute("INSERT INTO mytest (id) VALUES (1), (2)") cursor.execute("SELECT * FROM mytest") print(cursor.fetchall()) finally: connection.close() ``` This code creates a MySQL client and connects to the database. It creates a table, inserts some values, fetches them and prints the output. To run the code: ``` python main.py ``` If the script runs successfully, the output will be the values that were inserted into the table: ``` [{'id': 1}, {'id': 2}] ``` Now that your application is connected, you are all set to use Python with Aiven for MySQL. warning Make sure to create a table with a unique name. If you try to create a table that already exists, an exception will be raised. --- # Create Aiven for MySQL® databases Once you've created your Aiven for MySQL® service, you can add additional databases, for security purposes or to isolate your data per application. To create a MySQL® database: 1. Log in to [Aiven Console](https://console.aiven.io/). 2. In the **Services** page, select an Aiven for MySQL service for where to create a database. 3. In your service's page, in the **Connect** section, click **Databases**. 4. In the **Databases** view, select **Create database**. 5. Enter a name for your database into the **Name** field and select **Add database**. The new database will be visible immediately. tip You can also use the [Aiven client](/docs/tools/cli/service/database.md#avn-service-database-create) or the [MySQL client](/docs/products/mysql/howto/connect-from-cli.md) to create your database from the CLI. --- # Create missing primary keys Learn strategies to create missing primary keys in your Aiven for MySQL® service. They are important [for MySQL replication process](/docs/products/mysql/concepts/mysql-replication.md). ## List tables without primary key[​](#list-tables-without-primary-key "Direct link to List tables without primary key") Once you are connected to the MySQL database, you can determine which tables are missing primary keys by running the following commands: ``` SELECT tab.table_schema AS database_name, tab.table_name AS table_name, tab.table_rows AS table_rows FROM information_schema.tables tab LEFT JOIN information_schema.table_constraints tco ON (tab.table_schema = tco.table_schema AND tab.table_name = tco.table_name AND tco.constraint_type = 'PRIMARY KEY') WHERE tab.table_schema NOT IN ('mysql', 'information_schema', 'performance_schema', 'sys') AND tco.constraint_type IS NULL AND tab.table_type = 'BASE TABLE'; ``` To see the exact table definition for the problematic tables you can run the following command: ``` SHOW CREATE TABLE database_name.table_name; ``` If your table already contains a column or set of columns that can be used as primary key or composite key, then using such columns is recommended. In the next sections, find examples of tables definitions and the guidance on how to create the missing primary keys. ## Example: add primary key[​](#example-add-primary-key "Direct link to Example: add primary key") ``` CREATE TABLE person ( social_security_number VARCHAR(30) NOT NULL, first_name TEXT, last_name TEXT ); ``` You can create the missing primary key by adding the primary key: ``` ALTER TABLE person ADD PRIMARY KEY (social_security_number); ``` You don't have to explicitly define it as UNIQUE, [as the primary key is always unique in MySQL](https://dev.mysql.com/doc/refman/8.4/en/primary-key-optimization.html). ## Example: add a new separate id column[​](#example-add-a-new-separate-id-column "Direct link to Example: add a new separate id column") ``` CREATE TABLE team_membership ( user_id BIGINT NOT NULL, team_id BIGINT NOT NULL ); ``` Add the primary key using the following query: ``` ALTER TABLE team_membership ADD PRIMARY KEY (user_id, team_id); ``` If none of the existing columns or a combination of the existing columns cannot be used as the primary key, add a new separate id column. Check how to deal with it in [Example: alter table error](/docs/products/mysql/howto/create-missing-primary-keys.md#myslq-alter-table-error). ``` ALTER TABLE mytable ADD id BIGINT PRIMARY KEY AUTO_INCREMENT; ``` ## Example: alter table error `mysql.innodb_online_alter_log_max_size`[​](#myslq-alter-table-error "Direct link to myslq-alter-table-error") When executing the `ALTER TABLE` statement for a large table, you may encounter an error similar to the following: ``` Creating index 'PRIMARY' required more than 'mysql.innodb_online_alter_log_max_size' bytes of modification log. Please try again. ``` For the operation to succeed, set a value that is high enough. Depending on the table size, this can be a few gigabytes or even more for very large tables. To edit `mysql.innodb_online_alter_log_max_size` : [Aiven Console](https://console.aiven.io/) > your Aiven for MySQL service's page > the **Service settings** page of the service > the **Advanced configuration** section > **Configure** > **Add configuration options** > `mysql.innodb_online_alter_log_max_size` > set a value > **Save configuration**. Related pages Learn how to [create new tables without primary keys](/docs/products/mysql/howto/create-tables-without-primary-keys.md) in your Aiven for MySQL. --- # Create Aiven for MySQL® read replicas Learn how to create an Aiven for MySQL® read replica to provide a read-only instance of your managed MySQL service in another geographically autonomous region. ## About read replicas[​](#about-read-replicas "Direct link to About read replicas") Aiven for MySQL read-only replicas provide a great way to reduce the load on the primary server by enabling read-only queries to be performed against the replica. It is also a good way to optimise query response times across different geographical locations since, with Aiven, the replica can be placed in different regions or even different cloud providers. Using read-only replicas works as an extra measure to protect your data from the unlikely event that a whole region would go down. It can also improve performance if a read-only replica is placed closer to your end-users that read from the database. ## Create a read replica[​](#create-a-read-replica "Direct link to Create a read replica") 1. On the **Overview** page of your service, go to the **Read replica** section. 2. Click **Create replica**. 3. Enter a name for the replica. 4. Select a **Cloud**. 5. Select a **Plan**. 6. Click **Create**. You can see the read-only replica being created and listed next to other Aiven service in the **Services** page in [Aiven Console](https://console.aiven.io/). --- # Create new tables without primary keys If your Aiven for MySQL® service was created after 2020-06-03, by default it does not allow creating new tables without primary keys. ## Verify if new tables require a primary key[​](#verify-if-new-tables-require-a-primary-key "Direct link to Verify if new tables require a primary key") 1. Log in to [Aiven Console](https://console.aiven.io/). 2. Open your MySQL service and in the sidebar, click **Service settings**. 3. Scroll down to the **Advanced configuration** section and find the `mysql.sql_require_primary_key` parameter and its status. If `mysql.sql_require_primary_key` is enabled, your Aiven for MySQL does not allow you to create new tables without primary keys. If creating tables without primary keys is prevented and the table that you're trying to create is known to be small, you may override this setting and create the table anyway. Read more about the MySQL replication in the [Replication overview](/docs/products/mysql/concepts/mysql-replication.md) article. ## Create a table without primary keys[​](#create-a-table-without-primary-keys "Direct link to Create a table without primary keys") You have two options to create the tables: * Setting `mysql.sql_require_primary_key` to `0` for the current session, programmatically: 1. Run: ``` SET SESSION sql_require_primary_key = 0; ``` 2. Execute the CREATE TABLE or ALTER TABLE statement again in the same session. * Disabling the `mysql.sql_require_primary_key` parameter. To disable the `mysql.sql_require_primary_key` parameter: warning We recommend this approach when the table is created by an external application and using the session variable is not an option. To prevent more problematic tables from being unexpectedly created in the future, you should enable the setting again once you finished creating the tables without primary keys. 1. Log in to [Aiven Console](https://console.aiven.io/). 2. Open your MySQL service and in the sidebar, click **Service settings**. 3. Scroll down to the **Advanced configuration** section and select **Configure**. 4. Find `mysql.sql_require_primary_key` and disable it and click **Save configuration**. Related pages Learn how to [create missing primary keys](/docs/products/mysql/howto/create-missing-primary-keys.md) in your Aiven for MySQL. --- # Disable foreign key checks All Aiven for MySQL® services have foreign key checks enabled by default helping in keeping referential integrity across tables. However, you might want to disable it for a particular session. For example, when migrating to an Aiven for MySQL you may face errors related to foreign key violations similar to: ``` ERROR 3780 (HY000) at line 11596: Referencing column 'g_id' and referenced column 'g_id' in foreign key constraint 'FK_33b11dcfac6148578da087b07c2f388f' are incompatible. ``` The following explains how to temporarily disable Aiven for MySQL foreign key checking for the duration of a session. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * The `mysqlsh` client installed. You can install this by following the MySQL shell installation [documentation](https://dev.mysql.com/doc/mysql-shell/8.4/en/mysql-shell-install.html). * An Aiven account with an Aiven for MySQL service running. ## Variables[​](#variables "Direct link to Variables") There are a few variables to substitute when running the commands. To find the values for the substitution, go to [Aiven Console](https://console.aiven.io/) > your Aiven for MySQL service > **Overview** > the **Connection information** section > the **MySQL** tab. | Variable | Description | | ---------- | ------------------------------------------------ | | `HOST` | Hostname for MySQL connection | | `PORT` | Port for MySQL connection | | `PASSWORD` | Password of your Aiven for MySQL connection | | `DB_NAME` | Database Name of your Aiven for MySQL connection | ## Check the foreign key check flag[​](#check-the-foreign-key-check-flag "Direct link to Check the foreign key check flag") To check the foreign key check flag: * Connect to your Aiven for MySQL service with the following command: ``` mysql --user avnadmin --password=PASSWORD --host HOST --port PORT DB_NAME ``` * Run the following command to check the default configuration for your foreign key checks. ``` SHOW VARIABLES LIKE 'foreign_key_checks'; ``` * Verify that the foreign keys are enabled by default. You can expect to receive the following output: ``` +--------------------+-------+ | Variable_name | Value | +--------------------+-------+ | foreign_key_checks | ON | +--------------------+-------+ 1 row in set (0.05 sec) ``` ## Disable foreign key checks[​](#disable-foreign-key-checks "Direct link to Disable foreign key checks") To disable the foreign key checks for the session, you give an additional parameter when you connect to your Aiven for MySQL using the `mysqlsh`: ``` mysql \ --user avnadmin \ --password=PASSWORD \ --host HOST \ --port PORT DB_NAME \ --init-command="SET @@SESSION.foreign_key_checks = 0;" ``` Once again, we can check the current status of the foreign key checks by running the following: ``` SHOW VARIABLES LIKE 'foreign_key_checks'; ``` As result, we can see that the foreign key checks are disabled for this session: ``` +--------------------+-------+ | Variable_name | Value | +--------------------+-------+ | foreign_key_checks | OFF | +--------------------+-------+ 1 row in set (0.04 sec) ``` The same flag works when running a set of commands saved in a file with extension `.sql`. | Variable | Description | | ---------- | ----------------------------------------------------------------- | | `FILENAME` | File which the extension is `.sql`, for for example, filename.sql | You can paste the following command on your `FILENAME`: ``` SHOW VARIABLES LIKE 'foreign_key_checks'; ``` Now you can set the `init-command` flag to disable the foreign key checks, and run the commands in this file. ``` mysql \ --user avnadmin \ --password=PASSWORD \ --host HOST \ --port PORT DB_NAME \ --init-command="SET @@SESSION.foreign_key_checks = 0;" < FILENAME ``` ## More resources[​](#more-resources "Direct link to More resources") Read the official documentation to understand possible implications that can happen when disabling foreign key checks in your service. * [Foreign Key Checks](https://dev.mysql.com/doc/refman/8.4/en/create-table-foreign-keys.html#foreign-key-checks). * [Server System Variables](https://dev.mysql.com/doc/refman/8.4/en/server-system-variables.html#sysvar_foreign_key_checks). --- # Scale disk storage automatically for your Aiven for MySQL® service Automatically increase the disk storage of your Aiven for MySQL® service when it's running out of space, instead of resizing it manually. Use the Aiven Autoscaler to automatically increase the storage capacity of a service disk when it's running out of space. Disk autoscaler only increases storage, it doesn't scale storage down. ## Why use disk autoscaling[​](#why-use-disk-autoscaling "Direct link to Why use disk autoscaling") * **Cost efficiency**: Start with a regular-sized disk and let Aiven scale it up only when needed, without the risk of running out of disk space. * **Resiliency**: Avoid a service becoming non-functional because it ran out of disk space, including during unexpected spikes in demand. ## How it works[​](#how-it-works "Direct link to How it works") 1. You create an autoscaler integration endpoint in your project, setting the maximum total disk size to allow. 2. You enable an autoscaler integration for your service using that endpoint. 3. Aiven monitors the disk space usage of your service. 4. When disk usage reaches the threshold for your service type, Aiven increases the available storage by at least 10%, using the current used space as a baseline. note The exact increase depends on the service type and cloud provider. Some providers enforce a minimum increase of 10 GB. Autoscale thresholds per service type The threshold that triggers disk autoscaling is a percentage of the available disk storage capacity: * Aiven for OpenSearch®: 75% of the available disk storage capacity * All other supported service types: 85% of the available disk storage capacity 5. The disk increase is recorded in the project event log, and you receive a notification about the added disk space. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * **Maximum storage**: The maximum storage that the autoscaler can allocate for your service is limited by both the maximum disk size set on the autoscaler endpoint and the maximum disk storage supported for your service plan. * **Timing**: Autoscaling takes a moment to complete. In the meantime, the service disk might fill up and the service might enter read-only mode until autoscaling finishes, unless the autoscaler's disk capacity limit is reached. * **Maintenance updates**: Autoscaling works only on fully running services and can't happen during a maintenance update. * **Manual changes**: Changing disk space manually can delay an autoscaling event. * **Terraform**: Don't manage disk space with the Aiven Terraform Provider on a service that uses the autoscaler, to avoid conflicts between the two. * **Performance**: Disk added through autoscaling is slower than the original disk until the next maintenance update applies. This might affect I/O-intensive workloads. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven organization, project, and service that's up and running * The operator role for the organization, project, and service * Dynamic disk sizing support on your service plan and cloud region * One of the following to manage the autoscaler: * [Aiven Console](https://console.aiven.io/) * [Aiven API](https://api.aiven.io/doc/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) ### Enable disk autoscaling[​](#enable-disk-autoscaling "Direct link to Enable disk autoscaling") To enable disk autoscaling, create an autoscaler integration endpoint, then enable an autoscaler integration on your service using that endpoint. * Console * API * CLI * Terraform Create an autoscaler endpoint: 1. Log in to [Aiven Console](https://console.aiven.io/), and go to your organization and project. 2. On the left sidebar, click **Integration endpoints**. 3. Click **Aiven Autoscaler** > **Add new endpoint**. 4. Set the endpoint name and the maximum total disk storage in GB, and click **Add endpoint**. Enable the autoscaler on a service: 1. On the left sidebar, click **Services**, and open your service. 2. On the left sidebar, click **Integrations**. 3. In **Endpoint integrations**, click **Aiven Autoscaler**. 4. Click the endpoint you created, and click **Enable**. 1) Call [ServiceIntegrationEndpointCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointCreate) to create an autoscaler integration endpoint on your project: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "endpoint_name": "ENDPOINT_NAME", "endpoint_type": "autoscaler", "user_config": { "autoscaling": [ { "type": "autoscale_disk", "cap_gb": 300 } ] } }' ``` 2) Call [ServiceIntegrationCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationCreate) to enable the autoscaler integration on your service, using the endpoint ID from the previous response: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "dest_endpoint_id": "ENDPOINT_ID", "integration_type": "autoscaler", "source_project": "PROJECT_NAME", "source_service": "SERVICE_NAME" }' ``` 1. Create an autoscaler integration endpoint using [avn service integration-endpoint-create](/docs/tools/cli.md): ``` avn service integration-endpoint-create \ --project PROJECT_NAME \ --endpoint-name ENDPOINT_NAME \ --endpoint-type autoscaler \ --user-config-json '{"autoscaling": [{"type": "autoscale_disk", "cap_gb": 300}]}' ``` 2. Find the ID of the new endpoint: ``` avn service integration-endpoint-list --project PROJECT_NAME ``` 3. Enable the autoscaler integration on your service, using the endpoint ID from the previous step: ``` avn service integration-create \ --dest-service SERVICE_NAME \ --integration-type autoscaler \ --source-endpoint-id ENDPOINT_ID ``` Use the [`aiven_service_integration_endpoint`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration_endpoint) and [`aiven_service_integration`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration) resources: ``` resource "aiven_service_integration_endpoint" "autoscaler_endpoint" { project = "PROJECT_NAME" endpoint_name = "ENDPOINT_NAME" endpoint_type = "autoscaler" autoscaler_user_config { autoscaling { type = "autoscale_disk" cap_gb = 300 } } } resource "aiven_service_integration" "autoscaler_integration" { project = "PROJECT_NAME" integration_type = "autoscaler" source_service_name = "SERVICE_NAME" destination_endpoint_id = aiven_service_integration_endpoint.autoscaler_endpoint.id } ``` See the [disk autoscaler guide](https://registry.terraform.io/providers/aiven/aiven/latest/docs/guides/disk-autoscaler) for more details. ### Change the maximum disk space for autoscaling[​](#change-the-maximum-disk-space-for-autoscaling "Direct link to Change the maximum disk space for autoscaling") After you enable disk autoscaling, you can update the maximum total disk size at any time. * Console * API * CLI * Terraform 1. On the left sidebar, click **Integration endpoints**. 2. Click **Aiven Autoscaler**. 3. Find your endpoint, click **Actions**, and click the option to edit it. 4. Set a new maximum disk storage value, and save your changes. Call [ServiceIntegrationEndpointUpdate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointUpdate) with the new `cap_gb` value: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint/ENDPOINT_ID \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "user_config": { "autoscaling": [ { "type": "autoscale_disk", "cap_gb": 500 } ] } }' ``` ``` avn service integration-endpoint-update ENDPOINT_ID \ --user-config-json '{"autoscaling": [{"type": "autoscale_disk", "cap_gb": 500}]}' ``` Update the `cap_gb` value in the `autoscaling` block of your [`aiven_service_integration_endpoint`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration_endpoint) resource, then apply the change. ### Turn off disk autoscaling[​](#turn-off-disk-autoscaling "Direct link to Turn off disk autoscaling") To turn off disk autoscaling, remove the autoscaler integration from your service. You can also delete the integration endpoint if you no longer need it. * Console * API * CLI Disconnect the service from the autoscaler: 1. On the left sidebar, click **Services**, and open your service. 2. On the left sidebar, click **Integrations**. 3. In **Endpoint integrations**, find **Aiven Autoscaler**, click **Actions**, and click the option to disconnect it. Delete the autoscaler endpoint, if you no longer need it: 1. On the left sidebar, click **Integration endpoints**. 2. Click **Aiven Autoscaler**. 3. Find your endpoint, and delete it. 1) Call [ServiceIntegrationDelete](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationDelete) to remove the autoscaler integration from your service: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration/INTEGRATION_ID \ --header "Authorization: Bearer TOKEN" ``` 2) Call [ServiceIntegrationEndpointDelete](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointDelete) to delete the autoscaler integration endpoint, if you no longer need it: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint/ENDPOINT_ID \ --header "Authorization: Bearer TOKEN" ``` 1. Find the ID of the integration to remove: ``` avn service integration-list SERVICE_NAME ``` 2. Remove the autoscaler integration from your service: ``` avn service integration-delete INTEGRATION_ID ``` 3. Find the ID of the integration endpoint to delete, if you no longer need it: ``` avn service integration-endpoint-list --project PROJECT_NAME ``` 4. Delete the autoscaler integration endpoint: ``` avn service integration-endpoint-delete ENDPOINT_ID ``` Related pages * [Scale disk storage manually](/docs/products/mysql/howto/scale-disk-storage.md) * [Change the service plan](/docs/products/mysql/howto/change-service-plan.md) * [Memory and out-of-memory conditions](/docs/products/mysql/concepts/mysql-memory-usage.md) --- # Perform pre-migration checks Learn how to find potential errors before starting your database migration process. This can be done by using either the [Aiven CLI](https://github.com/aiven/aiven-client) or the [Aiven REST API](https://api.aiven.io/doc/#section/Introduction). Error example ``` { "migration": { "error": "Migration process failed", "method": "", "seconds_behind_master": null, "source_active": true, "status": "done" }, "migration_detail": [] } -----Response End----- STATUS METHOD ERROR ====== ====== ======================== done Migration process failed ``` ## Aiven CLI[​](#aiven-cli "Direct link to Aiven CLI") **Step 1: Create a task to perform the migration check** You can create the task of migration, for example, from a MySQL DB to an Aiven service (`project`: `MY_PROJECT_NAME`, `service`: `mysql`): ``` avn service task-create --operation migration_check --source-service-uri mysql://user:password@host:port/databasename --project MY_PROJECT_NAME mysql ``` You can see the information about the task including the ID. ``` TASK_TYPE SUCCESS TASK_ID ===================== ======= ==================================== mysql_migration_check null e2df7736-66c5-4696-b6c9-d33a0fc4cbed ``` tip List the options via the -h menu, for example, to ignore certain databases for the check. Note that filter databases are supported by MySQL only at the moment. **Step 2: Retrieve your task's status** You can check the status of your task by running: ``` avn service task-get --task-id e2df7736-66c5-4696-b6c9-d33a0fc4cbed --project MY_PROJECT_NAME mysql ``` It lists whether the operation succeeds and more information about the migration. ``` TASK_TYPE SUCCESS TASK_ID RESULT ===================== ======= ==================================== ==================================================================================== mysql_migration_check true e2df7736-66c5-4696-b6c9-d33a0fc4cbed All pre-checks passed successfully, preferred migration method will be [Replication] ``` ## Aiven REST API[​](#aiven-rest-api "Direct link to Aiven REST API") The same checks can be performed via the REST API. Read more: * [Create a task for service](https://api.aiven.io/doc/#operation/ServiceTaskCreate) * [Get task result](https://api.aiven.io/doc/#operation/ServiceTaskGet) --- # Enable slow query logging You can identify inefficient or time-consuming queries by enabling [slow query log](https://dev.mysql.com/doc/refman/5.7/en/slow-query-log.html) in your Aiven for MySQL® service. warning Since the output of the slow query log is written to the `mysql.slow_log` table on a particular server (of the service or its replica), the slow query logging is not supported on read-only replicas, which don't allow any writes. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") You need an Aiven organization with an Aiven for MySQL service running. ## Configure slow queries in Aiven Console[​](#configure-slow-queries-in-aiven-console "Direct link to Configure slow queries in Aiven Console") Enable your slow queries in your Aiven for MySQL service via [Aiven Console](https://console.aiven.io/): 1. Log in to [Aiven Console](https://console.aiven.io/). 2. In the **Services** page, select your Aiven for MySQL service. 3. In the **Service settings** page of your service, scroll down to the **Advanced configuration** section and select **Configure**. 4. In the **Advanced configuration** window 1. Select **Add configuration options**. From the unfolded list, choose `mysql.slow_query_log`. Enable `mysql.slow_query_log` by toggling it to `On`. By default, `mysql.slow_query_log` is disabled. 2. Select **Add configuration options**. From the unfolded list, choose `mysql.long_query_time`. Set `mysql.long_query_time` according to your specific need. 3. Select **Save configuration**. Your Aiven for MySQL service can now log slow queries. To simulate slow queries to check this feature, check the next section. ## Simulate slow queries[​](#simulate-slow-queries "Direct link to Simulate slow queries") Connect to your Aiven for MySQL using your favorite tool. Make sure you have `mysql.slow_query_log` enabled and set `mysql.long_query_time` to `2` seconds. Now, you can run the following query to simulate a slow query of 3 seconds. ``` select sleep(3); ``` You should see the following output: ``` +----------+ | sleep(3) | +----------+ | 0 | +----------+ 1 row in set (3.03 sec) ``` Now, you can check the logs of your slow query: ``` select convert(sql_text using utf8) as slow_query, query_time from mysql.slow_log; ``` You can expect to receive an output similar to the following: ``` +-----------------+-----------------+ | slow_query | query_time | +-----------------+-----------------+ | select sleep(3) | 00:00:03.000450 | +-----------------+-----------------+ 1 row in set, 1 warning (0.03 sec) ``` warning Disabling the `mysql.slow_query_log` setting truncates the `mysql.slow_query_log` table. Make sure to back up the data from the `mysql.slow_query_log` table in case you need it for further analysis. --- # Fork your Aiven for MySQL® service Fork your Aiven for MySQL® service to create an independent copy for testing, debugging, or development without affecting the original service. Fork an Aiven service to create a complete copy of it from its latest backup. Forked services are independent and don't share resources with or increase the load on the original service. Common use cases for forking include: * Creating a snapshot to analyze an issue. * Creating a development copy of your production environment. * Testing upgrades before applying them to production services. * Creating an instance in a different cloud provider, region, or with a different plan. * Renaming a service. During the forking process, the fork might initially have only one node while backups are being taken. The other nodes appear after the backup process is complete. When you fork a service, its configuration, databases, and service users are copied to the new service. ## Limitations[​](#limitations "Direct link to Limitations") * You can only fork services that have at least one [backup](/docs/products/mysql/concepts/mysql-backups.md). * Service integrations are not copied to the fork. * Cross-project forking is supported only within the same organization. ## Fork a service[​](#fork-a-service "Direct link to Fork a service") * Console * CLI * API * Terraform 1. In your service, in the **Backups** section, click **Backup management**. 2. Click **Fork & restore**. 3. Choose the backup to fork from. 4. Enter a name, and select the cloud and plan. 5. Click **Create fork**. Use the [create service command](/docs/tools/cli/service-cli.md#avn-cli-service-create) with: * `--service-to-fork-from`: the name of the service to use as the source. * `--project-to-fork-from`: to fork a service in a different project, set this to the project name the source service is in. Use the [`ServiceCreate` endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) and in the `user_config` property set: * `service_to_fork_from`: the name of the source service. * `project_to_fork_from`: to fork a service in a different project, set this to the name of the project the source service is in. Use the `service_to_fork_from` attribute in the user config of your service resource. To fork a service in a different project, set the `project_to_fork_from` attribute. More information on the service resources and their configuration options is available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs). ## Fork from a specific point in time[​](#fork-from-a-specific-point-in-time "Direct link to Fork from a specific point in time") * Console * CLI * API * Terraform 1. In your service, **Backups**. 2. Click **Fork & restore**. 3. Choose the point in time to fork from. 4. Enter a name, and choose the cloud and plan. 5. Click **Create fork**. Add the `--recovery-target-time` parameter to the [create service command](/docs/tools/cli/service-cli.md#avn-cli-service-create) and set it to a time between the first and latest available backups. Set the `recovery_target_time` parameter in the `user_config` property of the [`ServiceCreate` endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) to a time between the first and latest available backups. Set the `recovery_target_time` attribute in the user config of your service resource to a time between the first and latest available backups. More information on the service resources and their configuration options is available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Understand MySQL backups](/docs/products/mysql/concepts/mysql-backups.md) * [Rename your Aiven for MySQL® service](/docs/products/mysql/howto/rename-service.md) --- # Identify disk usage issues Aiven for MySQL® is configured to use `innodb_file_per_table=ON`, which means that an `.idb` file is generated per table containing its data and indexes. Over time, when a table receives a lot of inserts and deletions, the amount of space it occupies on disk can become significantly larger than the current data in the table. A classic example of this would be a table containing jobs for a work queue in which rows are repeatedly added to the end of the table and removed from the beginning. This happens because InnoDB does not release the allocated space back to the operating system automatically, in case the table grows larger in the future, but this can cause problems. Since every other table exists in its own `.idb` file, the allocated but unused space is unavailable for the tables to grow. Since every other table exists in its own `.idb` file, the allocated but unused space is unavailable for the tables to grow. ## Find disk usage issues[​](#find-disk-usage-issues "Direct link to Find disk usage issues") To identify tables with significant allocated but unused space, you can run the following query: ``` SELECT TABLES.TABLE_SCHEMA, TABLES.TABLE_NAME, (TABLES.DATA_LENGTH + TABLES.INDEX_LENGTH) / 1024 / 1024 AS "MB used (estimate)", TABLES.DATA_FREE / 1024 / 1024 AS "MB allocated but unused (estimate)", INNODB_TABLESPACES.FILE_SIZE / 1024 / 1024 AS "MB on disk" FROM information_schema.TABLES JOIN information_schema.INNODB_TABLESPACES ON (INNODB_TABLESPACES.NAME = TABLES.TABLE_SCHEMA || '/' || TABLES.TABLE_NAME) WHERE INNODB_TABLESPACES.FILE_SIZE > 10 * 1024 * 1024 ORDER BY TABLES.DATA_FREE DESC; ``` The query results show if a table has significantly more allocated unused space than used space. note By default, statistics in `information_schema.TABLES` are only updated every 24 hours or whenever the `ANALYZE TABLE` command runs. Related pages See [reclaim disk space](/docs/products/mysql/howto/reclaim-disk-space.md) if you are having issues with full disk. --- # Connect to Aiven for MySQL® Connect to the Aiven for MySQL® service using various programming languages or tools. ## [CLI](/docs/products/mysql/howto/connect-from-cli.md) [Connect to your Aiven for MySQL® service via the command line with the following tools:](/docs/products/mysql/howto/connect-from-cli.md) ## [Python](/docs/products/mysql/howto/connect-with-python.md) [This example connects your Python application to an Aiven for MySQL®](/docs/products/mysql/howto/connect-with-python.md) ## [MySQLx with Python](/docs/products/mysql/howto/connect-using-mysqlx-with-python.md) [Enabling the MySQLx protocol support allows you to use your MySQL](/docs/products/mysql/howto/connect-using-mysqlx-with-python.md) ## [PHP](/docs/products/mysql/howto/connect-with-php.md) [This example connects to an Aiven for MySQL® service from PHP, making use of the built-in PDO module.](/docs/products/mysql/howto/connect-with-php.md) ## [Java](/docs/products/mysql/howto/connect-with-java.md) [This example connects your Java application to an Aiven for MySQL® service.](/docs/products/mysql/howto/connect-with-java.md) ## [MySQL Workbench](/docs/products/mysql/howto/connect-from-mysql-workbench.md) [You can use a graphical client like MySQL Workbench to connect to Aiven for MySQL® services.](/docs/products/mysql/howto/connect-from-mysql-workbench.md) ## [DBeaver](/docs/products/mysql/howto/connect-with-dbeaver.md) [Use DBeaver to connect to your Aiven for MySQL® service.](/docs/products/mysql/howto/connect-with-dbeaver.md) ## [DataGrip](/docs/products/mysql/howto/connect-with-datagrip.md) [Use DataGrip to connect to your Aiven for MySQL® service.](/docs/products/mysql/howto/connect-with-datagrip.md) ## [max\_connections](/docs/products/mysql/concepts/max-number-of-connections.md) [Calculate the total number of simultaneous connections available to all users combined on your Aiven for MySQL® service, and learn why the per-user connection limit is a separate setting.](/docs/products/mysql/concepts/max-number-of-connections.md) --- # Maintenance and updates for your Aiven for MySQL® service Manage maintenance updates and set the maintenance window for your Aiven for MySQL® service. ## Maintenance updates[​](#maintenance-updates "Direct link to Maintenance updates") Aiven applies some maintenance updates automatically. The following are the types of updates: * **Mandatory updates:** Security updates, quarterly patch releases, and platform updates that affect reliability or stability of the service nodes. * **Optional updates:** All other updates are initially optional. After six months, they become mandatory and are applied in the next week’s maintenance window. * **Periodic infrastructure updates:** Scheduled automatically for services with nodes active for 180 days and more. These updates are mandatory for all services, except those with maintenance turned off. Critical security updates are applied during the next available maintenance window. For other updates, Aiven gives you at least seven days' notice. Maintenance updates are also automatically applied during service upgrades. To view pending updates: * Console * CLI * API 1. In your service, click **Service settings**. 2. Go to the **Service management** section. Use the [`avn service get`](/docs/tools/cli/service-cli.md#avn_service_get) command. Use the [`service`](https://api.aiven.io/doc/#tag/Service/operation/ServiceGet) endpoint. ## Maintenance window[​](#maintenance-window "Direct link to Maintenance window") The maintenance window is the time period when Aiven can automatically apply maintenance updates to a service. When an update becomes available, Aiven schedules it for the next available maintenance window for each service. The update runs in the first window after it becomes available, and can begin any time after the start time. For example, if a service has a maintenance window of Monday 12:00 UTC, and an update becomes available on Tuesday, the update will be applied on the following Monday. During maintenance, Aiven might restart or replace service nodes. This can cause brief connection interruptions, but services are designed to minimize downtime. Aiven performs maintenance in a rolling-forward style, creating new nodes alongside existing ones and retiring the old nodes after the upgrade completes. Major service upgrades are triggered manually. A manually triggered upgrade starts immediately, regardless of the maintenance window. important You cannot control the order in which services are updated. Each service updates according to its own configured maintenance window, and there is no guaranteed way to control the update sequence. Manual updates and maintenance window adjustments only help for non-critical updates. ## Set the maintenance window[​](#set-the-maintenance-window "Direct link to Set the maintenance window") To set the maintenance window for your service: * Console * Terraform 1. In the Aiven Console, open your service. 2. In the **Maintenance** section, click **Actions** > **Change maintenance window**. 3. Set the day and time. 4. Click **Save changes**. Use the `maintenance_window_dow` and `maintenance_window_time` attributes in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). ## Certificate rotation[​](#certificate-rotation "Direct link to Certificate rotation") Aiven periodically rotates the CA certificate for your project, including for your Aiven for MySQL® service. This rotation uses the same maintenance process described in [Maintenance updates](#maintenance-updates), applied during your service's maintenance window, to update your service to trust and use the new certificate. If you connect using the `VERIFY_CA` or `VERIFY_IDENTITY` SSL mode, update your client to trust the new certificate before the rotation completes. For details on the certificate bundle and rotation process, see [TLS/SSL certificates](/docs/platform/concepts/tls-ssl-certificates.md#certificate-rotation). Related pages * [Version upgrades](/docs/products/mysql/howto/manage-mysql-version.md) * [Change the service plan](/docs/products/mysql/howto/change-service-plan.md) * [TLS/SSL certificates](/docs/platform/concepts/tls-ssl-certificates.md) --- # Manage Aiven for MySQL® versions Aiven for MySQL® supports multiple versions of MySQL running concurrently in the platform. Choose a version that best fits your needs and upgrade your service when ready. ## Supported MySQL versions[​](#supported-mysql-versions "Direct link to Supported MySQL versions") From version 8.4, Aiven for MySQL supports two major upstream MySQL versions at a time. These are the two latest major versions that are stable on the Aiven Platform. You can select either version when you create a service or upgrade an existing service. If you do not select a version, the default is the latest stable version on the Aiven Platform. See the supported versions in the [Aiven for MySQL version reference](/docs/platform/reference/eol-for-major-versions.md#aiven-for-mysql). ## Before you upgrade[​](#before-you-upgrade "Direct link to Before you upgrade") ### Check available versions[​](#check-available-versions "Direct link to Check available versions") Preview versions available for your service in [Aiven Console](https://console.aiven.io): * Major versions: Service **Service settings** page > **Service management** section > **Actions** > **Upgrade version** > Expand the version dropdown list. * Minor versions: Service **Overview** page > **Maintenance** section > See the list of available mandatory and optional updates. For automated upgrades, you get email notifications. ### Check downgrade restrictions[​](#check-downgrade-restrictions "Direct link to Check downgrade restrictions") Downgrading to a previous version is not supported due to data format incompatibilities. Always test upgrades in a non-production environment first. To revert to a previous version: 1. Create a service with the desired version. 2. Restore data from a backup taken before the upgrade. 3. Update your application connection strings. ### Prerequisites for upgrade[​](#prerequisites-for-upgrade "Direct link to Prerequisites for upgrade") Before upgrading your service: * **Test in development**: Test the upgrade in a development environment first, for example, using service forking. * **Backup your data**: Ensure you have recent backups. Backups are automatic, but verify they exist. To upgrade your service version, check that: * Your Aiven for MySQL service is running. * Target version to upgrade to is [available for manual upgrade](/docs/products/mysql/howto/manage-mysql-version.md#check-available-versions). * You can use one of the following tools to upgrade: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) * [Aiven Operator for Kubernetes®](/docs/tools/kubernetes.md) ## Upgrade your service[​](#upgrade-your-service "Direct link to Upgrade your service") * Console * CLI * API * Terraform * Kubernetes 1. In the [Aiven Console](https://console.aiven.io/), go to your Aiven for MySQL service. 2. Open the **Service settings** page from the sidebar, and go to the **Service management** section. 3. Click **Actions** > **Upgrade version**. 4. Expand the version dropdown list, and select a version to upgrade to. warning When you click **Upgrade**: * The system applies the upgrade immediately. * You cannot downgrade the service to a previous version. 5. Click **Upgrade**. Upgrade the service version using the [avn service update](https://aiven.io/docs/tools/cli/service-cli#avn-cli-service-update) command: ``` avn service update SERVICE_NAME -c mysql_version="N.N" ``` Parameters: * `SERVICE_NAME`: Name of your service * `N.N`: Target service version to upgrade to, for example `8.4` Call the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint to set `mysql_version`: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'Content-Type: application/json' \ --data '{ "user_config": { "mysql_version": "N.N" } }' ``` Parameters: * `PROJECT_NAME`: Name of your project * `SERVICE_NAME`: Name of your service * `BEARER_TOKEN`: Your API authentication token * `N.N`: Target service version to upgrade to, for example `8.4` Use the [`aiven_mysql`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/mysql) resource to set [`mysql_version`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/mysql#mysql_version-1): ``` resource "aiven_mysql" "example" { project = var.PROJECT_NAME cloud_name = "CLOUD_NAME" plan = "PLAN_NAME" service_name = "SERVICE_NAME" mysql_user_config { mysql_version = "N.N" } } ``` Parameters: * `PROJECT_NAME`: Name of your project * `CLOUD_NAME`: Cloud region identifier * `PLAN_NAME`: Service plan * `SERVICE_NAME`: Name of your service * `N.N`: Target service version to upgrade to, for example `8.4` Use the [MySQL](https://aiven.github.io/aiven-operator/resources/mysql.html) resource to set [`mysql_version`](https://aiven.github.io/aiven-operator/resources/mysql.html#spec.userConfig.mysql_version-property): ``` apiVersion: aiven.io/v1alpha1 kind: MySQL metadata: name: SERVICE_NAME spec: authSecretRef: name: aiven-token key: token connInfoSecretTarget: name: mysql-connection project: PROJECT_NAME cloudName: CLOUD_NAME plan: PLAN_NAME userConfig: mysql_version: "N.N" ``` Apply the updated configuration: ``` kubectl apply -f mysql-service.yaml ``` Parameters: * `PROJECT_NAME`: Name of your project * `SERVICE_NAME`: Name of your service * `CLOUD_NAME`: Cloud region identifier * `PLAN_NAME`: Service plan * `N.N`: Target service version to upgrade to, for example `8.4` ## Version selection for new services[​](#version-selection-for-new-services "Direct link to Version selection for new services") When creating an Aiven for MySQL service: * **Default version**: The latest stable version on the Aiven Platform is the default version. * **Explicit selection**: You can specify a version using the `mysql_version` parameter. * **Version availability**: Only versions in `available` state can be selected. Example (CLI): ``` avn service create SERVICE_NAME \ --service-type mysql \ --plan PLAN_NAME \ --cloud CLOUD_NAME \ -c mysql_version="N.N" ``` Parameters: * `SERVICE_NAME`: Name of your service * `PLAN_NAME`: Service plan * `CLOUD_NAME`: Cloud region identifier * `N.N`: Service version, for example `8.1` Related pages --- # Manage Aiven for MySQL® service users Create and manage service users in your Aiven for MySQL® service to control access to its databases and tables. Service users only exist in the scope of the Aiven service. They are unique to the service and not shared with any other services. Every service has a default `avnadmin` user with full access to the service. ## Add a service user[​](#add-a-service-user "Direct link to Add a service user") * Aiven Console * Aiven CLI * Aiven API * Terraform 1. In your service, in the **Connect** section, click **Users**. 2. Click **Add service user** or **Create user**. 3. Enter a name for your service user. 4. Set up all the other configuration options. If a password is required, a random password is generated automatically. You can change it later. 5. Click **Add service user**. Run the [avn service user-create](/docs/tools/cli/service/user.md#avn-service-user-create) command: ``` avn service user-create SERVICE_NAME --username USERNAME ``` Replace the following: * `SERVICE_NAME`: the name of your Aiven for MySQL service. * `USERNAME`: the name of the service user to create. Use the [ServiceUserCreate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUserCreate) endpoint: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/user \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{"username": "USERNAME"}' ``` Replace the placeholders with your project name, service name, bearer token, and the username to create. To restrict the privileges granted to the new user, add the optional `mysql_grants` field. See [Restrict privileges for a new user](#restrict-privileges-for-a-new-user). Use the [`aiven_mysql_user` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/mysql_user) to create and manage service users. ## Restrict privileges for a new user[​](#restrict-privileges-for-a-new-user "Direct link to Restrict privileges for a new user") You can restrict the privileges assigned to a new Aiven for MySQL service user when you create it with the Aiven API. By default, a service user gets admin-level privileges, including the ability to create other users. To restrict these privileges, set `mysql_grants` to an array containing only the privileges to assign. Set it to an empty array to create a user with no privileges beyond connecting to the service. Omit the field to keep the default admin-level privileges. ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/user \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{"username": "USERNAME", "mysql_grants": ["SELECT", "INSERT"]}' ``` Replace the placeholders with your project name, service name, bearer token, and the username to create. If `mysql_grants` includes `CREATE USER` or `ROLE_ADMIN`, the created user can also grant every privilege in the list to other users, equivalent to MySQL's `WITH GRANT OPTION`. ### Available privileges[​](#available-privileges "Direct link to Available privileges") `mysql_grants` accepts the following values: | Privilege | Applies to | | ------------------------- | -------------------------------------------------------------- | | `ALTER` | Databases you create | | `ALTER ROUTINE` | Databases you create | | `CREATE` | The service and databases you create | | `CREATE ROUTINE` | Databases you create | | `CREATE TEMPORARY TABLES` | Databases you create | | `CREATE USER` | The service | | `CREATE VIEW` | Databases you create | | `DELETE` | Databases you create | | `DROP` | The service and databases you create | | `EVENT` | Databases you create | | `EXECUTE` | Databases you create | | `INDEX` | Databases you create | | `INSERT` | Databases you create | | `LOCK TABLES` | Databases you create | | `PROCESS` | The service | | `REFERENCES` | Databases you create | | `RELOAD` | The service | | `REPLICATION_APPLIER` | The service | | `REPLICATION CLIENT` | The service | | `REPLICATION SLAVE` | The service | | `ROLE_ADMIN` | The service | | `SELECT` | Databases you create, and read-only access to system databases | | `SHOW DATABASES` | The service | | `SHOW VIEW` | Databases you create | | `TRIGGER` | Databases you create | | `UPDATE` | Databases you create | For privileges that apply to databases, Aiven revokes the privilege from the service's system databases, except `SELECT`. This keeps read access to system information on every user without allowing changes to it. ### Requirements[​](#requirements "Direct link to Requirements") Restricting privileges at user creation requires your Aiven for MySQL service to support granular grants. If your service doesn't support this capability, requests that include `mysql_grants` fail with an HTTP `400 Bad Request` status code. To add support, review and apply pending [maintenance updates](/docs/products/mysql/howto/maintenance-updates.md) on your service. Related pages * [Create a database](/docs/products/mysql/howto/create-database.md) * [Connect to your service](/docs/products/mysql/howto/list-code-samples.md) * [Maintenance updates](/docs/products/mysql/howto/maintenance-updates.md) --- # Backup and restore Aiven for MySQL® with mysqldump or mydumper Copy your Aiven for MySQL® data to a file, back it up to another Aiven for MySQL database, and restore it using [`mysqldump/restore`](https://dev.mysql.com/doc/refman/8.4/en/mysqldump.html) or [`mydumper/myloader`](https://github.com/mydumper/mydumper). mydumper/myloader `mydumper/myloader` is an [early availability feature](/docs/platform/concepts/service-and-feature-releases.md) designed for large database migrations. It offers faster performance and reduced downtime compared to `mysqldump`, which can consume significant resources and time when processing large datasets. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Make sure your service has the required computational power (vCPUs) and memory capacity to process data migration without resource exhaustion or downtime. * mysqldump/restore * mydumper/myloader - [`mysqldump` tool](https://dev.mysql.com/doc/refman/8.4/en/mysqldump.html): [install](https://dev.mysql.com/doc/mysql-shell/8.4/en/mysql-shell-install.html) if missing - Source database to copy your data from: `source-db` - Target database to dump your `source-db` data to: `target-db` [Early availability](/docs/platform/concepts/service-and-feature-releases.md)

* [`mydumper`](https://github.com/mydumper/mydumper) tool ([install](https://mydumper.github.io/mydumper/docs/html/installing.html) if missing) * Source database to copy your data from: `source-db` * Target database to dump your `source-db` data to: `target-db` You can use Aiven for MySQL databases both as `source-db` and as `target-db`. [Create additional Aiven for MySQL® databases](/docs/products/mysql/howto/create-database.md) as needed. warning To avoid conflicts and replication issues, follow these guidelines while your data is being migrated: * Do not write to any tables in the target database that are being processed by the migration tool. * Do not change the replication configuration of the source database manually. Don't modify `binlog_format` or reduce `max_connections`. * Do not make database changes that can disrupt or prevent the connection between the source database and the target database. Do not change the source database's listen address and do not modify or enable firewalls on the databases. ## Back up the data[​](#back-up-the-data "Direct link to Back up the data") * mysqldump * mydumper ### Collect connection details[​](#collect-connection-details "Direct link to Collect connection details") To back up the `source-db` data to the `mydb_backup.sql` file, collect connection details on your Aiven for MySQL `source-db` service: 1. Log in to the [Aiven Console](https://console.aiven.io/) and go to your `source-db` service page. 2. On the **Overview** page, find **Connection information** and note the following: | Variable | Description | | -------------------- | ------------------------------------------- | | `SOURCE_DB_HOST` | **Host** name for the connection | | `SOURCE_DB_USER` | **User** name for the connection | | `SOURCE_DB_PORT` | Connection **Port** number | | `SOURCE_DB_PASSWORD` | Connection **Password** | | `DEFAULTDB` | Database that contains the `source-db` data | ### Back up to a file[​](#back-up-to-a-file "Direct link to Back up to a file") Use the following command to back up your Aiven for MySQL data to the `mydb_backup.sql` file: ``` mysqldump \ -p DEFAULTDB -P SOURCE_DB_PORT \ -h SOURCE_DB_HOST --single-transaction \ -u SOURCE_DB_USER --set-gtid-purged=OFF \ --password > mydb_backup.sql ``` With this command, the password will be requested at the prompt; paste `SOURCE_DB_PASSWORD` to the terminal, then a file named `mydb_backup.sql` will be created with your backup data. Note that having the prompt request for the password is more secure than including the password straight away in the command. The `--single-transaction` [flag](https://dev.mysql.com/doc/refman/8.4/en/mysqldump.html#option_mysqldump_single-transaction) starts a transaction in isolation mode `REPEATABLE READ` before running. This allows `mysqldump` to read the database in its current state at the time of the transaction, ensuring consistency of the data. warning If you are using [Global Transaction Identifiers](https://dev.mysql.com/doc/refman/8.4/en/replication-gtids-concepts.html) (GTIDs) with InnoDB use the `--set-gtid-purged=OFF` [option](https://dev.mysql.com/doc/refman/8.4/en/mysqldump.html#option_mysqldump_set-gtid-purged). The reason is that GTID's are not available with MyISAM. [Early availability](/docs/platform/concepts/service-and-feature-releases.md)

### Collect connection details[​](#collect-connection-details-1 "Direct link to Collect connection details") To backup the `source-db` data to the `mydb_backup_dir` directory, collect connection details on your Aiven for MySQL `source-db` service: 1. Log in to the [Aiven Console](https://console.aiven.io/) and go to your `source-db` service page. 2. On the **Overview** page, find **Connection information** and note the following: | Variable | Description | | -------------------- | ------------------------------------------- | | `SOURCE_DB_HOST` | **Host** name for the connection | | `SOURCE_DB_USER` | **User** name for the connection | | `SOURCE_DB_PORT` | Connection **Port** number | | `SOURCE_DB_PASSWORD` | Connection **Password** | | `DEFAULTDB` | Database that contains the `source-db` data | ### Back up to a directory[​](#back-up-to-a-directory "Direct link to Back up to a directory") To back up your data with `mydumper`, run: ``` mydumper \ --host SOURCE_DB_HOST \ --user SOURCE_DB_USER \ --password SOURCE_DB_PASSWORD \ --port SOURCE_DB_PORT \ --database DEFAULTDB \ --outputdir ./mydb_backup_dir ``` This creates the `mydb_backup_dir` directory containing the backup files. ## Restore the data[​](#restore-the-data "Direct link to Restore the data") * mysqldump/restore * myloader ### Collect connection details[​](#collect-connection-details-2 "Direct link to Collect connection details") To restore the saved data from the file to your `target-db`, collect connection details on your Aiven for MySQL `target-db` service: 1. Log in to the [Aiven Console](https://console.aiven.io/) and go to your `target-db` service page. 2. On the **Overview** page, find **Connection information** and note the following: | Variable | Description | | -------------------- | ------------------------------------------- | | `TARGET_DB_HOST` | **Host** name for the connection | | `TARGET_DB_USER` | **User** name for the connection | | `TARGET_DB_PORT` | Connection **Port** number | | `TARGET_DB_PASSWORD` | Connection **Password** | | `DEFAULTDB` | Database that contains the `target-db` data | ### Restore from the file[​](#restore-from-the-file "Direct link to Restore from the file") Run the following command to load the saved data into your `target-db` service: ``` mysql \ -p DEFAULTDB -P TARGET_DB_PORT \ -h TARGET_DB_HOST \ -u TARGET_DB_USER \ --password < mydb_backup.sql ``` [Early availability](/docs/platform/concepts/service-and-feature-releases.md)

### Collect connection details[​](#collect-connection-details-3 "Direct link to Collect connection details") To restore the saved data from the directory to your `target-db`, collect connection details on your Aiven for MySQL `target-db` service: 1. Log in to the [Aiven Console](https://console.aiven.io/) and go to your `target-db` service page. 2. On the **Overview** page, find **Connection information** and note the following: | Variable | Description | | -------------------- | ------------------------------------------- | | `TARGET_DB_HOST` | **Host** name for the connection | | `TARGET_DB_USER` | **User** name for the connection | | `TARGET_DB_PORT` | Connection **Port** number | | `TARGET_DB_PASSWORD` | Connection **Password** | | `DEFAULTDB` | Database that contains the `target-db` data | ### Restore from the directory[​](#restore-from-the-directory "Direct link to Restore from the directory") Use `myloader` to restore the data from the `mydumper` backup: ``` myloader \ --host TARGET_DB_HOST \ --user TARGET_DB_USER \ --password TARGET_DB_PASSWORD \ --port TARGET_DB_PORT \ --database DEFAULTDB \ --directory ./mydb_backup_dir ``` When the password is requested at the prompt, paste `TARGET_DB_PASSWORD` into the terminal. When the restore or load process is complete and the data is stored in your `target-db`, you can use the [`mysqlcheck` command](https://dev.mysql.com/doc/refman/8.4/en/mysqlcheck.html) to perform data analysis. Related pages * [Migrate to Aiven via CLI](/docs/products/mysql/howto/migrate-from-external-mysql.md) * [Migrate to Aiven via console](/docs/products/mysql/howto/migrate-db-to-aiven-via-console.md) * [Perform pre-migration checks on your Aiven for MySQL® database](/docs/products/mysql/howto/do-check-service-migration.md) --- # Migrate to Aiven for MySQL® via console Use the Aiven Console to migrate MySQL® databases to managed MySQL clusters in your Aiven organization. note To use the [Aiven CLI](/docs/tools/cli.md) to migrate your database, see [Migrate to Aiven via CLI](/docs/products/mysql/howto/migrate-from-external-mysql.md). You can migrate the following: * Existing on-premise MySQL databases * Cloud-hosted MySQL databases * Managed MySQL database clusters on Aiven The console migration tool provides two migration methods: * **(Recommended) Continuous migration:** Used by default in the tool and taken as the reference method. This method uses logical replication so that data transfer is possible not only for existing data in the source database when triggering the migration but also for any data written to the source database during the migration. * **mysqldump/restore:** Exports the current contents of the source database into a text file or backup directory and imports it to the target database. Any changes written to the source database during the migration are **not transferred**. When you trigger the migration setup in the console and initial checks detect that your source database does not support logical replication, you are notified about it via the migration wizard. To continue with the migration, the wizard allows you to select `mysqldump/restore` as an alternative migration method. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * To use the default continuous migration method in the Console tool, you have the logical replication enabled on your source database. * Source database's hostname or IP address are [accessible from the public Internet](/docs/platform/howto/public-access-in-vpc.md). * You have the source database's credentials and reference data: * Public hostname or connection string, or IP address used to connect to the database * Port used to connect to the database * Username (for a user with superuser permissions) * Password * Firewalls protecting the source database and the target databases are open to allow the traffic and connection between the databases (update or disable the firewalls temporarily if needed). ## Pre-configure the source[​](#pre-configure-the-source "Direct link to Pre-configure the source") 1. Allow remote connections on the source database: 1. Log in to the server hosting your database and the MySQL installation. Next, open the network configuration of MySQL with the following command: ``` sudo code /etc/mysql/mysql.conf.d/mysqld.cnf ``` Expected output ``` . . . lc-messages-dir = /usr/share/mysql skip-external-locking # # Instead of skip-networking the default is now to listen only on # localhost which is more compatible and is not less secure. bind-address = 127.0.0.1 . . . ``` 2. Change the value of `bind-address` to a wildcard IP address,`*` or `0.0.0.0`. Expected output ``` . . . lc-messages-dir = /usr/share/mysql skip-external-locking # # Instead of skip-networking the default is now to listen only on # localhost which is more compatible and is not less secure. bind-address = * . . . ``` 3. Save the changes and exit the file. Restart MySQL to apply the changes. ``` sudo systemctl restart mysql ``` note After completing the migration, make sure you revert those changes so that the MySQL database no longer accept remote connections. 2. Enable GTID: 1. Set up GTID on your database so that it can create a unique identifier for each transaction on the source database. See [Enabling GTID Transactions Online](https://dev.mysql.com/doc/refman/5.7/en/replication-mode-change-online-enable-gtids.html). To make sure you have GTID enabled, open your `my.cnf` file in `/etc/my.cnf` or `/etc/mysql/my.cnf`. If this file cannot be found, see [more potential locations in the MySQL documentation](https://dev.mysql.com/doc/refman/8.4/en/option-files.html). 2. Ensure the `my.cnf` file has the `[mysqld]` header. ``` [mysqld] gtid_mode=ON enforce_gtid_consistency=ON ``` 3. Restart MySQL. ``` sudo systemctl restart mysql ``` 3. Enable logical replication. Grant logical replication privileges to the user that you intend to connect to the source database with during the migration: 1. Log in to the database as an administrator and grant the following permission to the user: ``` GRANT ALL ON DATABASE_NAME.* TO USERNAME_CONNECTING_TO_SOURCE_DB; ``` 2. Reload the grant tables to apply the changes to the permissions. ``` FLUSH PRIVILEGES; ``` important After completing the migration, revert those changes so that the user no longer has logical replication privileges. ## Migrate a database[​](#migrate-a-database "Direct link to Migrate a database") 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. On the **Services** page, select the service where your target database is located. 3. From the sidebar on your service's page, select **Service settings**. 4. On the **Service settings** page, go to the **Service management** section, and select **Import database**. 5. Guided by the migration wizard, go through all the migration steps. ### Step 1: Configure[​](#step-1-configure "Direct link to Step 1: Configure") Get familiar **Guidelines for successful database migration** provided in the **MySQL migration configuration guide** window, make sure your configuration is in line with them, and select **Get started**. ### Step 2: Validate[​](#step-2-validate "Direct link to Step 2: Validate") 1. To establish a connection to your source database, enter required source database details into the wizard: * Hostname * Port * Username * Password 2. Select the **SSL encryption recommended** checkbox. 3. In the **Exclude databases** field, enter names of databases that you don't want to migrate (if any). 4. Select **Run checks** to have the connection validated. Unable to use logical replication If your connection check returns the **Unable to use logical replication** warning, either resolve the issues or use the the dump method by selecting **Start the migration using a one-time snapshot (dump method)** > **Run check** > **Start migration**. ### Step 3: Migrate[​](#step-3-migrate "Direct link to Step 3: Migrate") If all the checks pass successfully, trigger the migration by clicking **Start migration**. ### Step 4: Replicating[​](#stop-migration-mysql "Direct link to Step 4: Replicating") While the migration is in progress, you can: * Let it proceed until completed by selecting **Close window**. You can come back to check the status at any time. * Discontinue the migration by selecting **Stop migration**, which retains the data already migrated. To follow up on a stopped migration process, see [Start over](/docs/products/mysql/howto/migrate-db-to-aiven-via-console.md#start-over-mysql). warning To avoid conflicts and replication issues, follow these guidelines while your data is being migrated: * Do not write to any tables in the target database that are being processed by the migration tool. * Do not change the replication configuration of the source database manually. Don't modify `binlog_format` or reduce `max_connections`. * Do not make database changes that can disrupt or prevent the connection between the source database and the target database. Do not change the source database's listen address and do not modify or enable firewalls on the databases. Migration attempt failed? If the migration fails, investigate potential causes of the failure and try to fix the issues. When you're ready, trigger the migration again by selecting **Start over**. When the migration is complete, select one of the following: * **Close connection** to disconnect the databases and stop the replication process if still active. * **Keep replicating** if the replication is still ongoing and you want to keep the connection open for data synchronization. Replication mode active? Your data has been transferred to Aiven but new data is still continuously being synced between the connected databases. ### Step 5: Close the connection[​](#step-5-close-the-connection "Direct link to Step 5: Close the connection") When the migration is completed, and all active replication processes are over, click **Close connection**. All the data in your database has been transferred to Aiven. ## Start over[​](#start-over-mysql "Direct link to Start over") If you [stop a migration process](/docs/products/mysql/howto/migrate-db-to-aiven-via-console.md#stop-migration-mysql), you cannot restart the same process. Still, the data already migrated is retained in the target database. warning If you start a new migration using the same connection details when your *target* database is not empty, the migration tool truncates your *target* database and an existing data set gets overwritten with the new data set. Related pages * [Migrate to Aiven for MySQL from an external MySQL](/docs/products/mysql/howto/migrate-from-external-mysql.md) * [About aiven-db-migrate](/docs/products/postgresql/concepts/aiven-db-migrate.md) --- # Migrate to Aiven for MySQL® via CLI Migrate your external MySQL database to an Aiven-hosted one using either a one-time dump-and-restore or continuous data synchronization through MySQL's built-in replication. note To migrate your database using the Aiven Console, see [Migrate MySQL® databases to Aiven via console](/docs/products/mysql/howto/migrate-db-to-aiven-via-console.md). ## How it works[​](#how-it-works "Direct link to How it works") The Aiven for MySQL migration process begins with an **initial data transfer** and can be followed by **continuous data synchronization** if your setup supports it. ### Initial data transfer[​](#initial-data-transfer "Direct link to Initial data transfer") A bulk copy of your data is first created. This is done using one of the following tools: * `mysqldump` for small and medium-sized databases * `mydumper/myloader` [Early availability](/docs/platform/concepts/service-and-feature-releases.md) for large databases mydumper/myloader `mydumper/myloader` is an [early availability feature](/docs/platform/concepts/service-and-feature-releases.md) designed for large database migrations. It offers faster performance and reduced downtime compared to `mysqldump`, which can consume significant resources and time when processing large datasets. ### Continuous data synchronization[​](#continuous-data-synchronization "Direct link to Continuous data synchronization") After the initial data copy, the Aiven for MySQL service can be configured as a replica of your external database, enabling ongoing data synchronization through MySQL's built-in replication feature. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * The source server is publicly available or accessible via a virtual private cloud (VPC) peering connection between the private networks, and firewalls are open to allow traffic between the source and target servers. * You have a user account on the source server with sufficient privileges to create a user for the replication process. * [GTID](https://dev.mysql.com/doc/refman/8.4/en/replication-gtids.html) is enabled on the source database. Review the current GTID setting by running the following command on the source cluster: ``` show global variables like 'gtid_mode'; ``` * You have a running Aiven for MySQL service with a destination database. If missing, create it in the [Aiven Console](/docs/products/mysql/get-started.md) or the [Aiven CLI](/docs/tools/cli/service-cli.md#avn-cli-service-create). * If you use `mydumper/myloader` for the [initial data transfer](/docs/products/mysql/howto/migrate-from-external-mysql.md#initial-data-transfer), make sure your service has sufficient computational power (multiple vCPUs) and high memory capacity to avoid the resource exhaustion during migration. note If you are migrating from MySQL in Google Cloud, enable backups with [PITR](https://cloud.google.com/sql/docs/mysql/backup-recovery/pitr) for GTID to be set to `on`. ## Collect source and destination details[​](#collect-source-and-destination-details "Direct link to Collect source and destination details") | Variable | Description | | ---------------- | -------------------------------------------------------------------------------------------------- | | `SRC_HOSTNAME` | Hostname for source MySQL connection | | `SRC_PORT` | Port for source MySQL connection | | `SRC_USERNAME` | Username for source MySQL connection | | `SRC_PASSWORD` | Password for source MySQL connection | | `SRC_IGNORE_DBS` | Comma-separated list of databases to ignore in migration | | `SRC_SSL` | SSL setting for source MySQL connection | | `DEST_NAME` | Name of the destination Aiven for MySQL service | | `DEST_PLAN` | Aiven plan for the destination Aiven for MySQL service (for example, `startup-4` or `business-32`) | ## Migrate your database[​](#migrate-your-database "Direct link to Migrate your database") warning To avoid conflicts and replication issues, follow these guidelines while your data is being migrated: * Do not write to any tables in the target database that are being processed by the migration tool. * Do not change the replication configuration of the source database manually. Don't modify `binlog_format` or reduce `max_connections`. * Do not make database changes that can disrupt or prevent the connection between the source database and the target database. Do not change the source database's listen address and do not modify or enable firewalls on the databases. 1. Create a user in the source database with sufficient privileges for the pre-flight checks, the bulk copy (using `mysqldump` or `mydumper` in [Early availability](/docs/platform/concepts/service-and-feature-releases.md)), and the ongoing replication. Replace `%` with the IP address of the Aiven for MySQL database, if already existing. ``` create user 'SRC_USERNAME'@'%' identified by 'SRC_PASSWORD'; grant replication slave on *.* TO 'SRC_USERNAME'@'%'; grant select, process, event on *.* to 'SRC_USERNAME'@'%' ``` 2. Set the migration details using the `avn service update` [Aiven CLI command](/docs/tools/cli/service-cli.md#avn-cli-service-update). * Use your preferred migration tool: * For `mysqldump`, include option `-c migration.dump_tool=mysqldump` in the command. * For `mydumper`, include option `-c migration.dump_tool=mydumper` in the command. * Replace the [placeholders](/docs/products/mysql/howto/migrate-from-external-mysql.md#collect-source-and-destination-details) with meaningful values. ``` avn service update --project PROJECT_NAME \ -c migration.host=SRC_HOSTNAME \ -c migration.port=SRC_PORT \ -c migration.username=SRC_USERNAME \ -c migration.password=SRC_PASSWORD \ -c migration.ignore_dbs=SRC_IGNORE_DBS \ -c migration.ssl=SRC_SSL \ -c migration.dump_tool=mysqldump \ DEST_NAME ``` 3. Check the migration status using the `avn service migration-status` [Aiven CLI command](/docs/tools/cli/service-cli.md#avn-cli-service-migration-status): ``` avn service migration-status --project PROJECT_NAME DEST_SERVICE_NAME ``` When the migration process is ongoing, `migration_detail.status` is `syncing`: ``` { "migration": { "error": null, "method": "replication", "seconds_behind_master": 0, "source_active": true, "status": "done" }, "migration_detail": [ { "dbname": "migration", "error": null, "method": "replication", "status": "syncing" } ] } ``` Ongoing replication starts a few minutes after the initial data copy. tip Monitor the ongoing replication status by running `show replica status` on the destination database. ## Stop the replication[​](#stop-the-replication "Direct link to Stop the replication") After confirming that the migration is complete, stop the ongoing replication by removing the configuration from the destination service via the `avn service update` [Aiven CLI command](/docs/tools/cli/service-cli.md#avn-cli-service-update): ``` avn service update --project PROJECT_NAME --remove-option migration DEST_NAME ``` warning If you don't stop the ongoing replication, you might lose data. For example, if you remove the data on the migration source when the replication is active, the data is also removed on the migration target. --- # Detect and terminate long-running queries in Aiven for MySQL® Aiven does not terminate any customer queries even if they run indefinitely, but long-running queries can cause issues by locking resources and therefore preventing database maintenance tasks such as backups. To identify and terminate such long-running queries, you can do it from either: * [Aiven Console](https://console.aiven.io) * [MySQL shell](/docs/products/mysql/howto/connect-from-cli.md) (`mysql`) ## Terminate long running queries from the Aiven Console[​](#terminate-long-running-queries-from-the-aiven-console "Direct link to Terminate long running queries from the Aiven Console") 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. From the **Services** page, select your Aiven for MySQL® service. 3. In your service's page, in the **Observe** section, click **Current queries**. 4. In the **Current queries** page, you can check the query duration and select **Terminate** to stop any long-running queries. ## Detect and terminate long running queries via CLI[​](#detect-and-terminate-long-running-queries-via-cli "Direct link to Detect and terminate long running queries via CLI") You can [login to your service](/docs/products/mysql/howto/connect-from-cli.md) using `mysqlsh` or `mysql`. Once connected, you can call the following command on the `mysql` shell to view all running queries: ``` SHOW PROCESSLIST WHERE command = 'Query' AND info NOT LIKE '%PROCESSLIST%'; ``` You can learn more about the `SHOW PROCESSLIST` command from the [official documentation](https://dev.mysql.com/doc/refman/8.4/en/show-processlist.html). You can terminate a query manually using: ``` KILL QUERY pid ``` where the `pid` is the process ID output by the `SHOW PROCESSLIST` command above. You can learn more about the `KILL QUERY` command from the [MySQL KILL documentation](https://dev.mysql.com/doc/refman/8.4/en/kill.html). --- # Power on/off and delete your Aiven for MySQL® service Power off your Aiven for MySQL® service to release resources and save credits, power it back on when you need it, or delete it permanently. ## Power off a service[​](#power-off-a-service "Direct link to Power off a service") When you power off a service: * All virtual machines are removed from the public cloud. * The service configuration is stored on the Aiven Platform. * If there are no backups, all service data is lost. * If the service has time-based or point in time recovery backups, the backups remain on the Aiven Platform. Services powered off for more than 180 days are automatically deleted. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power off service**. To power off a service, run: ``` avn service update SERVICE_NAME --power-off ``` ## Power on a service[​](#power-on-a-service "Direct link to Power on a service") When you power on a service: * New virtual machines are created on the service's public cloud. * The service starts with the stored configuration parameters. * The latest time-based backup is restored. * Maintenance updates are automatically applied. * If a point in time recovery backup is available, the database transaction logs are replayed to recover the service data to a specific point in time. The restoration takes from a few minutes to a few hours, depending on the network bandwidth, the disk IOPS allocated to the service, and the size of the backup. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power on service**. To power on a service, run: ``` avn service update SERVICE_NAME --power-on ``` To see when the service is running, run: ``` avn service wait SERVICE_NAME ``` note When you power on an Aiven for MySQL service, the latest backup is restored and the binary log (binlog) is replayed to recover the data to the latest available point in time. note Static IP addresses are not removed when a service is powered off or deleted. They continue to generate the usual costs. To avoid these costs, [remove the static IP addresses](/docs/platform/concepts/static-ips.md). ## Delete a service[​](#delete-a-service "Direct link to Delete a service") * Console * CLI 1. In your project, click **Services**. 2. Open the service to delete, and click **Actions** > **Delete service**. To delete a service, run: ``` avn service terminate SERVICE_NAME ``` Related pages * [Understand MySQL backups](/docs/products/mysql/concepts/mysql-backups.md) * [Fork Aiven for MySQL®](/docs/products/mysql/howto/fork-service.md) --- # Prepare your Aiven for MySQL® service for high load Prepare your Aiven for MySQL® service for higher than usual traffic to avoid outages and keep performance stable. Prepare your services for higher than usual traffic to avoid service outages by doing the following: * **Subscribe to service notifications:** To receive notifications about service health and warnings when resources are low, you can [set service and project contacts](https://aiven.io/docs/platform/howto/technical-emails). You can also view the status of the Aiven Platform and get updates on incidents on the [status page](https://status.aiven.io/). Follow the RSS feed, subscribe to email or SMS notifications, or use the Slack integration to get notifications about incidents. * **Monitor your services:** [Monitor the health of your services](/docs/platform/howto/list-monitoring.md) using metrics, logs, alerts, and dashboards. * **Scale your services:** If you forecast a load that can't be handled by the service, you can scale up your service. * **Set the backup schedule:** To minimize the impact of the higher load during the backup process, schedule backups outside of peak traffic hours. * **Set the maintenance window:** Schedule maintenance updates outside of your peak traffic hours. * **Run load tests on service forks:** To test the impact of high traffic on a production service, fork the service and run your load test on the fork. Additionally, optimizing a service allows it to perform better under stress therefore avoiding the need of an upgrade. The more optimized a service is for your usage, the better you can weather spikes in traffic. Related pages * [Change the service plan](/docs/products/mysql/howto/change-service-plan.md) * [Scale disk storage](/docs/products/mysql/howto/scale-disk-storage.md) --- # Prevent running out of disk space Learn how Aiven prevents running out of disk space from happening and how you can make more space available on your disk when needed. Running out of disk space makes the service start malfunctioning and prevents backups from being properly created. ## Switch to the read-only mode[​](#switch-to-the-read-only-mode "Direct link to Switch to the read-only mode") Aiven automatically detects when your service is running out of free space and prevents further writes to it. This process is done by setting the MySQL `@@GLOBAL.read_only` flag to `1`. The threshold for moving to this state is when your disk usage is at 97% or higher. Once your service is made `read-only`, the service reports errors when you attempt to insert, update, or delete data: ``` ERROR 1290 (HY000): The MySQL server is running with the --read-only option so it cannot execute this statement ``` ## Free up disk space[​](#free-up-disk-space "Direct link to Free up disk space") ### Optimize problem tables[​](#optimize-problem-tables "Direct link to Optimize problem tables") InnoDB does not reclaim unused disk space by default and this can cause a disk to become full. Read the help article [MySQL disk usage](/docs/products/mysql/howto/reclaim-disk-space.md) for more information. ### Upgrade to a larger plan[​](#upgrade-to-a-larger-plan "Direct link to Upgrade to a larger plan") This can be done from within [Aiven Console](https://console.aiven.io/) or with the [Aiven CLI](/docs/tools/cli.md) client. New nodes with more disk capacity are launched, and your existing data is synced to those new nodes. Once the migration is completed, the disk usage drops below the critical level and the read-only state is canceled, allowing writes to be made once more. ### Delete data[​](#delete-data "Direct link to Delete data") As your service is set in `read-only` mode, attempting to free disk space by deleting data can't be done directly. To disable the `read-only` state, use our API to temporarily remove the restriction. You can use our API and send a POST request to: ``` https://api.aiven.io/v1/project//service//enable-writes ``` The output of a successful operation is: ``` { "message": "Writes temporarily enabled", "until": "2022-04-22T13:42:05.385432Z" } ``` This way you can free up space within the next 15 minutes. Related pages See [reclaim disk space](/docs/products/mysql/howto/reclaim-disk-space.md) if you are having issues with a full disk. --- # Reclaim disk space You can configure InnoDB to release disk space back to the operating system by running the `OPTIMIZE TABLE` command. note [Under certain conditions](https://dev.mysql.com/doc/refman/8.4/en/optimize-table.html#optimize-table-innodb-details) (for example, including the presence of a `FULLTEXT` index), command `OPTIMIZE TABLE` [copies](https://dev.mysql.com/doc/refman/8.4/en/alter-table.html) the data to a new table containing just the current data, and drops and renames the new table to match the old one. During this process, data modification is blocked. This requires enough free space to store two copies of the current data at once. To ensure that the space is also reclaimed on standby nodes, run the command as below without any additional modifiers like `NO_WRITE_TO_BINLOG` or `LOCAL`: ``` ``OPTIMIZE TABLE defaultdb.mytable;`` ``` If you do not have enough free space to run the `OPTIMIZE TABLE` command, you can: * Start by **optimizing smaller tables** to free up space. Next, you can proceed with optimizing on larger tables. * **Temporarily upgrade to a larger service plan** to get access to more disk space. You can downgrade your plan again afterward. Read more on how to upgrade your plan for that. note When you perform a temporary upgrade, it may require waiting for a smaller backup to take place before downgrading the plan again. --- # Rename your Aiven for MySQL® service Change the name of your Aiven for MySQL® service by forking it under a new name and deleting the original service. You cannot rename a service after creation. Instead, you can create a fork with the new name and delete the original service. ## Rename a service[​](#rename-a-service "Direct link to Rename a service") 1. Stop writing to the service. 2. Fork the service. 3. Add any integrations or SSO configurations that weren't copied. 4. Connect your clients to the new service. 5. Test the forked service. 6. Delete the original service. Related pages * [Fork Aiven for MySQL®](/docs/products/mysql/howto/fork-service.md) * [Power on/off and delete your Aiven for MySQL® service](/docs/products/mysql/howto/power-cycle-service.md) --- # Scale disk storage for your Aiven for MySQL® service Scale the disk storage of your Aiven for MySQL® service up or down without disrupting the running service. /eol-for-major-versions#aiven-for-flinkAdding or removing disk storage does not disrupt the running service. You pay only for extra storage instead of upgrading compute resources. You can add extra storage when you create a service or after it is running. When you add storage to a running service, the Aiven Platform provisions the extra disk and adds it to the running instances. For a clustered service such as Aiven for Apache Kafka®, Aiven divides extra storage equally between the nodes. For a shared service, each node receives the full extra capacity. ## Limitations[​](#limitations "Direct link to Limitations") * Disk added for extra storage is slower than the original disk until the next maintenance update. The slower disk can reduce performance for I/O-intensive workloads. * Maximum storage depends on the plan, service type, and cloud provider. It can be up to five times the plan's base storage size. * Cloud providers limit how many times you can increase storage between maintenance updates. If you reach the limit, run a maintenance update to optimize performance. * You cannot add storage during a maintenance update. * Dynamic disk sizing (DDS) is not supported on custom service plans. Pricing If you add storage when you create a service, **Additional disk storage** shows an estimated monthly cost. The **Service summary** lists plan storage plus additional storage. The estimated monthly price includes the additional storage cost. If you add storage to a running service, the Aiven Console shows the cost of the additional storage and related backups. The same costs appear on your invoices. ## Add or remove storage[​](#add-or-remove-storage "Direct link to Add or remove storage") ### Add storage when you create a service[​](#add-storage-when-you-create-a-service "Direct link to Add storage when you create a service") To add storage while you create a service: 1. In **Additional disk storage**, set the size with the slider or enter a value in GB. 2. Review the estimated monthly cost. 3. In the **Service summary**, click **Create service**. Change additional storage later on the running service, or enable automatic disk scaling with Aiven Autoscaler. ### Change storage on a running service[​](#change-storage-on-a-running-service "Direct link to Change storage on a running service") You cannot add or remove storage when service nodes are in the rebuilding state, for example during a maintenance update or a service upgrade. If you are removing disk storage: * Make sure the data in your service does not exceed the allocated storage. If it does, you cannot remove the additional storage. * Plan for the time it takes to rebuild the service. The time depends on the service. - Console - CLI - Terraform 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Actions** > **Manage additional storage**. 3. Change the disk storage. note * The price shown for the additional storage includes backup costs. * You can only remove storage that you previously added using this feature. To downgrade further, you can change your service plan. 4. Click **Save Changes**. Use [Aiven CLI](/docs/tools/cli.md) to add or remove additional storage using [`avn service update`](/docs/tools/cli/service-cli.md#avn-cli-service-update) with the `--disk-space-gib` flag to specify the total disk space to provide to your service. For example, if your service has a 80-GiB disk and you would like to add an extra 10-GiB disk, use: ``` avn service update --disk-space-gib 90 --project PROJECT_NAME SERVICE_NAME ``` note * When you perform a horizontal service upgrade or downgrade, remember to include all additional disks the service uses. For example, when switching from `Startup-4` to `Business-4` or from `Business-4` to `Startup-4`, include all the additional disks available for this service. * When you fork an existing service, include all additional disks the service uses. Use the `additional_disk_space` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). If you added storage, the additional storage is available immediately. If you removed additional storage, the service nodes go through a rolling restart. Depending on the service type and configuration, there might be a short downtime for services with no HA capabilities. note Storage optimization is performed at the next maintenance update after a change to the storage size. Due to cloud provider limitations, there is a limit on how many times storage can be increased between two maintenance updates. When this limit is reached, perform a maintenance update for performance optimization. Plan increases to avoid reaching this limit. Related pages * [Disk autoscaler](/docs/products/mysql/howto/disk-autoscaler.md) * [Change the service plan](/docs/products/mysql/howto/change-service-plan.md) * [Memory and out-of-memory conditions](/docs/products/mysql/concepts/mysql-memory-usage.md) --- # Tag your Aiven for MySQL® service Add key-value tags to your Aiven for MySQL® service to organize services and track ownership, cost allocation, and governance. Use tags to add metadata to Aiven services to categorize them or run custom logic on them. Typical uses include: * Tagging for governance to deploy services with specific tags only. * Tagging for internal cost reporting, ownership, allocation, and accountability. A tag is a key/value pair: * **Key**: A case-sensitive string that starts with a letter and consists of letters, numbers, dashes, and underscores. The maximum length for a key is 64 characters. * **Value**: A string value limited to 64 UTF-8 characters. Within a service, the tag keys must be unique. * Console * Terraform 1. In the service, click **Service settings**. 2. In the **Service status** section, click **Actions** > **Add service tags**. 3. Enter a key and value for each tag. 4. Click **Save changes**. Use the `tag` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Power on/off and delete your Aiven for MySQL® service](/docs/products/mysql/howto/power-cycle-service.md) --- # Track restore progress for your Aiven for MySQL® service Track the restore progress of individual nodes in your Aiven for MySQL® service during node replacement, forking, or maintenance, using the Aiven API. You can track restore progress for individual nodes during service node replacement by using the Aiven API. For example, use this endpoint to monitor the restore progress of a forked service or when applying maintenance. The service object exposes restore progress under `node_states[].progress_updates`: * `service.node_states[]` contains per-node state entries. * When a node is restoring or catching up, its `state` is typically `syncing_data`. * When the state is `syncing_data`, the node may include `progress_updates` with one or more phase objects. * Other node states don't include restore progress data. note `progress_updates` may be missing or empty even when a node is in `syncing_data`. This can occur when a restore completes before detailed progress is reported or when the service does not emit detailed progress counters. ## API endpoints[​](#api-endpoints "Direct link to API endpoints") Restore progress fields are part of the standard service response payload. * Get a single service (recommended for polling): `GET /project/{project}/service/{service_name}` * List services in a project: `GET /project/{project}/service` - Request - Response ``` curl -H "Authorization: aivenv1 API_TOKEN" https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME ``` Replace the following placeholders: * `API_TOKEN`: Your Aiven API token. * `PROJECT_NAME`: Your Aiven project name. * `SERVICE_NAME`: The name of your service. ``` { "service": { ... "node_states": [ { "node_name": "...", "state": "syncing_data", "progress_updates": [ { "phase": "basebackup", "completed": false, "current": 3410567, "min": 0, "max": 7569280, "unit": "bytes_uncompressed" } ] } ], ... } } ``` ## Node states[​](#node-states "Direct link to Node states") Common values for `node_states[].state` include: * `setting_up_vm`: The virtual machine is being created or initialized. * `syncing_data`: The node is restoring data or catching up. * `running`: The node is operating normally. * `leaving`: The node is leaving the cluster. * `unknown`: A transient or error state. ## `progress_updates` data model[​](#progress_updates-data-model "Direct link to progress_updates-data-model") `progress_updates` is a list of phase objects. When present, phases appear in the following order: 1. `prepare` 2. `basebackup` 3. `stream` 4. `finalize` Each phase object includes the following fields: ``` { "completed": false, "current": 3410567, "max": 7569280, "min": 0, "phase": "basebackup", "unit": "bytes_uncompressed" } ``` ### Field semantics[​](#field-semantics "Direct link to Field semantics") * `phase`: String, required. The restore phase. Possible values: `prepare`, `basebackup`, `stream`, and `finalize`. * `completed`: Boolean, required. Whether the phase is complete. * `current`: Number or null, optional. The current progress value. This field can be missing or null. * `min`: Number or null, optional. The starting value for the phase. This field can be missing or null. * `max`: Number or null, optional. The expected total value for the phase. This value can be missing, null, or change while the restore is in progress. * `unit`: String or null, optional. The unit for `current`, `min`, and `max`. New unit values can appear over time. Important considerations * Treat `unit` as an opaque identifier. Unknown values can appear. * `max` may change while a restore is in progress. * Not all phases report numeric counters. Some services only indicate phase completion. ## Why `max` values can change[​](#why-max-values-can-change "Direct link to why-max-values-can-change") The `current`, `min`, and `max` values are best-effort progress indicators. They can be based on estimates or on system state that changes over time. Treat `max` as the latest known expected total, not as a fixed guarantee. Common reasons `max` can change include: * The restore process discovers additional work after it starts, such as files, segments, or objects that become visible only after metadata is read. * New data is added on the backend while the node is catching up, which moves the completion point forward. This is common during incremental catch-up phases. * Progress is calculated from system state, such as replication lag, rather than from a fixed work queue. As the system state changes, the value is recalculated. * The service switches restore strategies during the operation, for example from snapshot restore to replication catch-up, which changes what the counters represent. As a result: * Phase percentage can decrease even when the restore operates normally. * Remaining-time estimates based on `max` are unreliable. * Sudden changes in `max` are expected unless the node remains in `syncing_data` longer than expected. ## Restore phase meanings[​](#restore-phase-meanings "Direct link to Restore phase meanings") Phase names are standardized, but the underlying work and the meaning of the counters are service-specific. * `prepare`: Prepares the node for restore. * `basebackup`: Restores the full backup. * `stream`: Applies incremental changes, such as replication or log replay. * `finalize`: Completes final steps before serving traffic. Not all restores include every phase. ## Compute phase progress percentages[​](#compute-phase-progress-percentages "Direct link to Compute phase progress percentages") You cannot reliably compute overall restore progress. You can compute a phase-specific progress percentage when `min`, `max`, and `current` are present and `max != min`. ``` pct = round(((current - min) / (max - min)) * 100, 1) ``` When handling progress values: * If any of `min`, `max`, or `current` is null or missing, display `n/a`. * If `max == min`, treat the percentage as undefined. * Expect the percentage to decrease when `max` changes. * Clamp displayed values to the range `[0, 100]`. ## Polling guidance[​](#polling-guidance "Direct link to Polling guidance") Progress updates are best-effort and refresh every 10 seconds while a node is in `syncing_data`. Poll the service state every 10 to 30 seconds. More frequent polling does not provide additional detail. For each `node_states[]` entry: * If `state` is not `syncing_data`, no restore progress is available. * If `state` is `syncing_data`: * If `progress_updates` is missing or empty, the node is restoring without detailed progress data. * Otherwise, the current phase is the last phase where `completed` is `false`. Stop polling when all nodes reach the `running` state or when a stall is detected. ### Stall detection[​](#stall-detection "Direct link to Stall detection") The API does not provide per-phase timestamps. To detect stalls, use a time-based threshold, such as a node remaining in `syncing_data` longer than expected. Do not rely on counters or `max` values to estimate remaining time. Related pages * [Backups](/docs/products/mysql/concepts/mysql-backups.md) * [Fork your service](/docs/products/mysql/howto/fork-service.md) --- # Use Aiven for MySQL® incremental backups [Early availability](/docs/platform/concepts/service-and-feature-releases.md) Streamline your Aiven for MySQL® backups using incremental backups. The Aiven for MySQL® incremental backups feature is an extension of the [base backups](/docs/products/mysql/concepts/mysql-backups.md) feature. It allows you to back up only the changes (increments) made since your last full backup. While full backups provide a complete snapshot of your database at a specific point in time, incremental backups capture the differences that have occurred between these full snapshots. This feature calculates the difference between the last full backup and the current state of your Aiven for MySQL database and stores it. During a restore, the full backup is applied first, followed by each incremental backup in order, reconstructing your database to the desired point in time. Use Aiven for MySQL® backups for: * **Reduced DDL lock times:** For large databases, full backups can cause prolonged DDL (Data Definition Language) locks. Incremental backups significantly reduce this lock duration, especially when there haven't been many changes. * **Efficient storage:** Databases with fewer changes will benefit from less storage consumption, as incremental backups only store the differences, leading to more efficient use of your storage. note * Restoring from incremental backups might take more time than restoring from a full backup. * Using incremental backups doesn't allow to completely eliminate DDL locks. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Enable Aiven for MySQL® incremental backups as a [early availability feature](/docs/platform/concepts/service-and-feature-releases.md#early-availability-). * Access the [Aiven Console](https://console.aiven.io/) or the [Aiven API](/docs/tools/api.md). * Set up a weekly backup schedule before or when [enabling incremental backups](/docs/products/mysql/howto/use-incremental-backups.md#enable-incremental-backups). ## Enable incremental backups[​](#enable-incremental-backups "Direct link to Enable incremental backups") * Console * API 1. Access your Aiven for MySQL service in the [Aiven Console](https://console.aiven.io/). 2. Click **Service settings** in the sidebar. 3. Scroll to the **Advanced configuration** section and click **Configure**. 4. In the **Advanced configuration** window, click **Add configuration option**. 5. Use the search bar to: 1. Find `mysql_incremental_backup`, and set it to **Enable**. 2. Find `mysql_incremental_backup.full_backup_week_schedule`, and set it to a string of any number of the following comma-separated values: `mon`, `tue`, `wed`, `thu`, `fri`, `sat`, `sun`. 6. Click **Save configuration**. Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) API to add `"mysql_incremental_backup": true` to the `user_config` object: ``` curl -X PUT \ https://api.aiven.io/v1/project/YOUR_PROJECT_NAME/service/YOUR_SERVICE_NAME \ -H "Authorization: Bearer YOUR_BEARER_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "user_config": { "mysql_incremental_backup": { "enabled": true, "full_backup_week_schedule": "mon,wed" } } }' ``` In `"full_backup_week_schedule": "VALID_VALUES"`, replace the `VALID_VALUES` placeholder with a string of any number of the following comma-separated values: `mon`, `tue`, `wed`, `thu`, `fri`, `sat`, `sun`. ## Disable incremental backups[​](#disable-incremental-backups "Direct link to Disable incremental backups") * Console * API 1. Access your Aiven for MySQL service in the [Aiven Console](https://console.aiven.io/). 2. Click **Service settings** in the sidebar. 3. Scroll to the **Advanced configuration** section and click **Configure**. 4. In the **Advanced configuration** window, click **Add configuration option**. 5. Use the search bar to find `mysql_incremental_backup`, and set it to **Disable**. 6. Click **Save configuration**. Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) API to add `"mysql_incremental_backup": false` to the `user_config` object: ``` curl -X PUT \ https://api.aiven.io/v1/project/YOUR_PROJECT_NAME/service/YOUR_SERVICE_NAME \ -H "Authorization: Bearer YOUR_BEARER_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "user_config": { "mysql_incremental_backup": { "enabled": false } } }' ``` Related pages --- # Maintenance and lifecycle in Aiven for MySQL® Keep your Aiven for MySQL® service current by upgrading versions and applying maintenance updates. Related pages * [Manage the MySQL version](/docs/products/mysql/howto/manage-mysql-version.md) --- # Advanced parameters for Aiven for MySQL See the configuration options available for Aiven for MySQL: | Parameter | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | []()[**additional\_backup\_regions**](#additional_backup_regions)`array`Additional Cloud Regions for Backup Replication | | []()[**ip\_filter**](#ip_filter)`array`- default: `0.0.0.0/0,::/0`IP filterAllow incoming connections from CIDR address block, e.g. '10.20.0.0/16' | | []()[**service\_log**](#service_log)`boolean,null`Service loggingStore logs for the service so that they are available in the HTTP API and console. | | []()[**static\_ips**](#static_ips)`boolean`Use static public IP addresses | | []()[**admin\_username**](#admin_username)`string,null`Custom username for admin user. This must be set only when a new service is being created. | | []()[**admin\_password**](#admin_password)`string,null`Custom password for admin user. Defaults to random string. This must be set only when a new service is being created. | | []()[**backup\_hour**](#backup_hour)`integer`- max: `23`The hour of day (in UTC) when backup for the service is started. New backup is only started if previous backup has already completed. | | []()[**backup\_minute**](#backup_minute)`integer`- max: `59`The minute of an hour when backup for the service is started. New backup is only started if previous backup has already completed. | | []()[**migration**](#migration)`object,null`Migrate data from existing servermigration.host string Hostname or IP address of the server where to migrate data from migration.port integer - min: 1 - max: 65535 Port number of the server where to migrate data from migration.password string Password for authentication with the server where to migrate data from migration.ssl boolean - default: true The server where to migrate data from is secured with SSL migration.username string User name for authentication with the server where to migrate data from migration.dbname string Database name for bootstrapping the initial connection migration.ignore\_dbs string Comma-separated list of databases, which should be ignored during migration (supported by MySQL and PostgreSQL only at the moment) migration.ignore\_roles string Comma-separated list of database roles, which should be ignored during migration (supported by PostgreSQL only at the moment) migration.method string The migration method to be used (currently supported only by Redis, Dragonfly, MySQL and PostgreSQL service types) migration.dump\_tool string,null MySQL migration dump tool Experimental! Tool to use for database dump and restore during migration. Default: mysqldump migration.reestablish\_replication boolean Skip dump-restore part and start replication | | []()[**private\_access**](#private_access)`object`Allow access to selected service ports from private networksprivate\_access.mysql boolean Allow clients to connect to mysql with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.mysqlx boolean Allow clients to connect to mysqlx with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.prometheus boolean Allow clients to connect to prometheus with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | []()[**privatelink\_access**](#privatelink_access)`object`Allow access to selected service components through Privatelinkprivatelink\_access.mysql boolean Enable mysql privatelink\_access.mysqlx boolean Enable mysqlx privatelink\_access.prometheus boolean Enable prometheus | | []()[**public\_access**](#public_access)`object`Allow access to selected service ports from the public Internetpublic\_access.mysql boolean Allow clients to connect to mysql from the public internet for service nodes that are in a project VPC or another type of private network public\_access.mysqlx boolean Allow clients to connect to mysqlx from the public internet for service nodes that are in a project VPC or another type of private network public\_access.prometheus boolean Allow clients to connect to prometheus from the public internet for service nodes that are in a project VPC or another type of private network | | []()[**service\_to\_fork\_from**](#service_to_fork_from)`string,null`Name of another service to fork from. This has effect only when a new service is being created. | | []()[**project\_to\_fork\_from**](#project_to_fork_from)`string,null`Name of another project to fork a service from. This has effect only when a new service is being created. | | []()[**mysql\_version**](#mysql_version)`string`MySQL major version | | []()[**recovery\_target\_time**](#recovery_target_time)`string,null`Recovery target time when forking a service. This has effect only when a new service is being created. | | []()[**binlog\_retention\_period**](#binlog_retention_period)`integer`- min: `600`
- max: `9007199254740991`The minimum amount of time in seconds to keep binlog entries before deletion. This may be extended for services that require binlog entries for longer than the default for example if using the MySQL Debezium Kafka connector.Warning: reducing this value can make a large batch of binary logs eligible for purge at once. Depending on the volume, this can sometimes stall the MySQL commit path and block writes until the purge completes. To stay on the safe side, prefer lowering the value gradually in small decrements during a low-traffic window rather than dropping it drastically in one step. | | []()[**mysql\_incremental\_backup**](#mysql_incremental_backup)`object`MySQL incremental backup configurationmysql\_incremental\_backup.enabled boolean Enable periodic incremental backups. When enabled, full\_backup\_week\_schedule must be set. Incremental backups only store changes since the last backup, making them faster and more storage-efficient than full backups. This is particularly useful for large databases where daily full backups would be too time-consuming or expensive. mysql\_incremental\_backup.full\_backup\_week\_schedule string,null Full backup week schedule Comma-separated list of days of the week when full backups should be created. Valid values: mon, tue, wed, thu, fri, sat, sun | | []()[**mysql**](#mysql)`object`mysql.conf configuration valuesmysql.sql\_mode string sql\_mode Global SQL mode. Set to empty to use MySQL server defaults. When creating a new service and not setting this field Aiven default SQL mode (strict, SQL standard compliant) will be assigned. mysql.connect\_timeout integer - min: 2 - max: 3600 connect\_timeout The number of seconds that the mysqld server waits for a connect packet before responding with Bad handshake mysql.default\_time\_zone string default\_time\_zone Default server time zone as an offset from UTC (from -12:00 to +12:00), a time zone name, or 'SYSTEM' to use the MySQL server default. mysql.div\_precision\_increment integer - max: 30 div\_precision\_increment Number of digits by which to increase the scale of the result of division operations performed with the / operator. Default is 4. mysql.eq\_range\_index\_dive\_limit integer - max: 4294967295 eq\_range\_index\_dive\_limit The number of equality ranges in a query at or above which the optimizer switches from index dives to index statistics when estimating the number of qualifying rows. 0 means always use index dives. Default is 200. mysql.group\_concat\_max\_len integer - min: 4 - max: 18446744073709552000 group\_concat\_max\_len The maximum permitted result length in bytes for the GROUP\_CONCAT() function. mysql.information\_schema\_stats\_expiry integer - min: 900 - max: 31536000 information\_schema\_stats\_expiry The time, in seconds, before cached statistics expire mysql.innodb\_adaptive\_hash\_index boolean innodb\_adaptive\_hash\_index Whether InnoDB adaptive hash indexing is enabled. The optimal setting is workload-dependent: it speeds up lookups for some workloads but its internal latch can become a contention point under high concurrency, in which case disabling it can improve throughput. mysql.innodb\_change\_buffer\_max\_size integer - max: 50 innodb\_change\_buffer\_max\_size Maximum size for the InnoDB change buffer, as a percentage of the total size of the buffer pool. Default is 25 mysql.innodb\_flush\_neighbors integer - max: 2 innodb\_flush\_neighbors Specifies whether flushing a page from the InnoDB buffer pool also flushes other dirty pages in the same extent (default is 1): 0 - dirty pages in the same extent are not flushed, 1 - flush contiguous dirty pages in the same extent, 2 - flush dirty pages in the same extent mysql.innodb\_ft\_enable\_stopword boolean innodb\_ft\_enable\_stopword Whether stopword processing is applied when creating or rebuilding an InnoDB FULLTEXT index. Enabled by default. mysql.innodb\_ft\_max\_token\_size integer - min: 10 - max: 84 - Service restart innodb\_ft\_max\_token\_size Maximum length of words that are stored in an InnoDB FULLTEXT index. Changing this parameter will lead to a restart of the MySQL service. mysql.innodb\_ft\_min\_token\_size integer - max: 16 - Service restart innodb\_ft\_min\_token\_size Minimum length of words that are stored in an InnoDB FULLTEXT index. Changing this parameter will lead to a restart of the MySQL service. mysql.innodb\_ft\_num\_word\_optimize integer - min: 1000 - max: 10000 innodb\_ft\_num\_word\_optimize Number of words processed during each OPTIMIZE TABLE operation on an InnoDB FULLTEXT index. Default is 2000. mysql.innodb\_ft\_result\_cache\_limit integer - min: 1000000 - max: 4294967295 innodb\_ft\_result\_cache\_limit Maximum memory in bytes used per query for the InnoDB FULLTEXT search query result cache. Aiven sizes this automatically based on the service plan's memory; setting a value overrides the calculated default. mysql.innodb\_ft\_server\_stopword\_table null,string innodb\_ft\_server\_stopword\_table This option is used to specify your own InnoDB FULLTEXT index stopword list for all InnoDB tables. mysql.innodb\_ft\_user\_stopword\_table null,string innodb\_ft\_user\_stopword\_table This option is used to specify your own InnoDB FULLTEXT index stopword list for specific InnoDB tables. mysql.innodb\_io\_capacity integer - min: 100 - max: 4294967295 The number of I/O operations per second (IOPS) available to InnoDB background tasks, such as flushing pages from the buffer pool and merging data from the change buffer. Set this to a value appropriate for the underlying storage; it must not exceed innodb\_io\_capacity\_max. mysql.innodb\_io\_capacity\_max integer - min: 100 - max: 4294967295 innodb\_io\_capacity\_max The maximum number of I/O operations per second (IOPS) that InnoDB background tasks may perform when flushing falls behind. Defaults to twice innodb\_io\_capacity (minimum 2000). This must be greater than or equal to innodb\_io\_capacity. mysql.innodb\_lock\_wait\_timeout integer - min: 1 - max: 3600 innodb\_lock\_wait\_timeout The length of time in seconds an InnoDB transaction waits for a row lock before giving up. Default is 120. mysql.innodb\_log\_buffer\_size integer - min: 1048576 - max: 4294967296 innodb\_log\_buffer\_size The size in bytes of the buffer that InnoDB uses to write to the log files on disk. Requests above 15% of the RAM provided by your service plan are rejected, because a larger buffer leaves less memory for the buffer pool and client connections. mysql.innodb\_online\_alter\_log\_max\_size integer - min: 65536 - max: 1099511627776 innodb\_online\_alter\_log\_max\_size The upper limit in bytes on the size of the temporary log files used during online DDL operations for InnoDB tables. mysql.innodb\_optimize\_fulltext\_only boolean innodb\_optimize\_fulltext\_only When enabled, OPTIMIZE TABLE on InnoDB tables only updates the FULLTEXT index instead of rebuilding the table. Intended to be enabled temporarily during FULLTEXT index maintenance and disabled afterwards; while enabled, OPTIMIZE TABLE does not reclaim table space. mysql.innodb\_print\_all\_deadlocks boolean innodb\_print\_all\_deadlocks When enabled, information about all deadlocks in InnoDB user transactions is recorded in the error log. Disabled by default. mysql.innodb\_read\_io\_threads integer - min: 1 - max: 64 - Service restart innodb\_read\_io\_threads The number of I/O threads for read operations in InnoDB. Default is 4. Changing this parameter will lead to a restart of the MySQL service. mysql.innodb\_rollback\_on\_timeout boolean - Service restart innodb\_rollback\_on\_timeout When enabled a transaction timeout causes InnoDB to abort and roll back the entire transaction. Changing this parameter will lead to a restart of the MySQL service. mysql.innodb\_thread\_concurrency integer - max: 1000 innodb\_thread\_concurrency Defines the maximum number of threads permitted inside of InnoDB. Default is 0 (infinite concurrency - no limit) mysql.innodb\_write\_io\_threads integer - min: 1 - max: 64 - Service restart innodb\_write\_io\_threads The number of I/O threads for write operations in InnoDB. Default is 4. Changing this parameter will lead to a restart of the MySQL service. mysql.interactive\_timeout integer - min: 30 - max: 604800 interactive\_timeout The number of seconds the server waits for activity on an interactive connection before closing it. mysql.internal\_tmp\_mem\_storage\_engine string internal\_tmp\_mem\_storage\_engine The storage engine for in-memory internal temporary tables. mysql.max\_connections integer - min: 30 - max: 100000 max\_connections The maximum permitted number of simultaneous client connections. Lower this to reserve memory for other work. The value cannot exceed the limit provided by your service plan. Upgrading the plan does not raise a value you have set explicitly, so increase it yourself after an upgrade. mysql.max\_execution\_time integer - max: 4294967295 max\_execution\_time Execution timeout in milliseconds for read-only top-level SELECT statements. 0 (the default) means no timeout. mysql.max\_seeks\_for\_key integer - min: 1 - max: 18446744073709552000 max\_seeks\_for\_key Limit on the assumed maximum number of index seeks when looking up rows based on a key. Lowering this value causes the optimizer to prefer index lookups over table scans. mysql.max\_user\_connections integer - max: 99990 max\_user\_connections The maximum number of simultaneous connections permitted to any single user account. 0, the default, means no per-account limit. Any other value must be at least 10 below max\_connections, so that monitoring and your own admin sessions can still connect when an application saturates its own limit. Aiven's replication and management connections are unaffected however low you set this. mysql.net\_buffer\_length integer - min: 1024 - max: 1048576 - Service restart net\_buffer\_length Start sizes of connection buffer and result buffer. Default is 16384 (16K). Changing this parameter will lead to a restart of the MySQL service. mysql.net\_read\_timeout integer - min: 1 - max: 3600 net\_read\_timeout The number of seconds to wait for more data from a connection before aborting the read. mysql.net\_write\_timeout integer - min: 1 - max: 3600 net\_write\_timeout The number of seconds to wait for a block to be written to a connection before aborting the write. mysql.optimizer\_prune\_level integer - max: 1 optimizer\_prune\_level Controls the heuristics applied during query optimization to prune less-promising partial plans from the optimizer search space. 0 disables heuristics (exhaustive search); 1 prunes plans based on the number of rows retrieved. mysql.optimizer\_search\_depth integer - max: 62 optimizer\_search\_depth Maximum depth of search performed by the query optimizer when choosing a join order. Larger values produce better plans for joins over many tables but take longer to compile; 0 lets the optimizer choose the depth automatically. mysql.optimizer\_switch string optimizer\_switch Comma-separated list of optimizer flag assignments in the form flag=on\|off\|default, or the single value 'default' to reset all flags. Flags not listed keep their current values. Controls query optimizer behaviors such as index merge, hash join and semijoin strategies. mysql.sql\_require\_primary\_key boolean sql\_require\_primary\_key Require primary key to be defined for new tables or old tables modified with ALTER TABLE and fail if missing. It is recommended to always have primary keys because various functionality may break if any large table is missing them. mysql.wait\_timeout integer - min: 1 - max: 2147483 wait\_timeout The number of seconds the server waits for activity on a noninteractive connection before closing it. Requests to set this below 30 are rejected, because a shorter timeout closes your own idle connections between statements. mysql.max\_allowed\_packet integer - min: 102400 - max: 1073741824 max\_allowed\_packet Size of the largest message in bytes that can be received by the server. Default is 67108864 (64M). Statements and rows larger than this are rejected with a packet too large error. mysql.max\_heap\_table\_size integer - min: 1048576 - max: 1073741824 max\_heap\_table\_size Limits the size of internal in-memory tables. Also set tmp\_table\_size. Default is 16777216 (16M) mysql.relay\_log\_space\_limit integer - min: 134217728 - max: 18446744073709552000 - Service restart relay\_log\_space\_limit The maximum amount of space in bytes to use for all relay logs while replicating from an external migration source. When the limit is reached, the replication I/O thread stops fetching relay log events until the SQL thread has caught up. Raise this to give a large migration a bigger relay-log budget; ensure the service disk is sized accordingly. The setting applies only on the node replicating from the external source; standby nodes always use the Aiven-managed default (the smaller of 5 GiB and 30% of the service disk), which is also used when this option is left unset. Changing this parameter will lead to a restart of the MySQL service. mysql.sort\_buffer\_size integer - min: 32768 - max: 1073741824 sort\_buffer\_size Sort buffer size in bytes for ORDER BY optimization. Default is 262144 (256K). Requests above 2% of the RAM provided by your service plan are rejected, because the buffer is allocated per session and its cost multiplies with the connection count. mysql.tmp\_table\_size integer - min: 1048576 - max: 1073741824 tmp\_table\_size Limits the size of internal in-memory tables. Also set max\_heap\_table\_size. Default is 16777216 (16M) mysql.slow\_query\_log boolean Slow query log enables capturing of slow queries. Setting slow\_query\_log to false also truncates the mysql.slow\_log table. mysql.long\_query\_time number - max: 3600 The slow\_query\_logs work as SQL statements that take more than long\_query\_time seconds to execute. mysql.log\_output string log\_output The slow log output destination when slow\_query\_log is ON. To enable MySQL AI Insights, choose INSIGHTS. To use MySQL AI Insights and the mysql.slow\_log table at the same time, choose INSIGHTS,TABLE. To only use the mysql.slow\_log table, choose TABLE. To silence slow logs, choose NONE. mysql.lower\_case\_table\_names integer lower\_case\_table\_names Sets how table and database names are stored and compared. 0 = case-sensitive (default), 1 = names stored lowercase, comparisons are case-insensitive. This option can only be set when creating the service and cannot be changed later. See https\://dev.mysql.com/doc/refman/8.0/en/identifier-case-sensitivity.html for details. mysql.performance\_schema\_events\_statements\_history\_size integer,null - max: 1024 - Service restart performance\_schema\_events\_statements\_history\_size The number of rows per thread in the events\_statements\_history table. Changing this parameter will lead to a restart of the MySQL service. mysql.automatic\_sp\_privileges boolean automatic\_sp\_privileges When enabled, the server automatically grants the EXECUTE and ALTER ROUTINE privileges to the creator of a stored routine and drops them when the routine is dropped. mysql.end\_markers\_in\_json boolean end\_markers\_in\_json Whether optimizer JSON output such as EXPLAIN FORMAT=JSON adds end markers that repeat a structure's key near its closing bracket, making large JSON structures easier to read. mysql.windowing\_use\_high\_precision boolean windowing\_use\_high\_precision Whether window functions are computed to high precision. Disabling this trades exactness for speed in window function evaluation. | --- # Resource capability of Aiven for MySQL® plans When creating or updating an Aiven service, the plan that you choose will drive the specific resources (CPU, memory, disk IOPS, etc.) powering your service. Aiven is a cloud data platform, so the underlying instance types are chosen appropriately for the type of service. Elements like local NVMe SSDs, sufficient memory for the expected workloads, fast access to backup storage, ability to encrypt the disks, contribute to the choice of the instance types to use on the selected cloud platforms. In addition, particular instance types are sometimes not available in a specific cloud region. There is no one-size-fits-all for choosing the optimal instance type, so Aiven takes all of these criteria into account to select the right instance type for a given service. To know how much your service can handle, you can benchmark it with your specific workload, in a representative setup. It will be affected by more than just the instance type: network throughput, latency to your applications, number of connections active at the time, type of TLS encryption in use, and a whole range of things specific to the cloud environment can contribute to a service's expected performance. The best way to know how your application will work is to benchmark it in your setup. You can move between plans, scaling up and down, without any downtime in order to try different sizes. --- # Aiven for MySQL® version lifecycle Learn how Aiven manages Aiven for MySQL® version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. ## Aiven version support and upstream EOL[​](#aiven-version-support-and-upstream-eol "Direct link to Aiven version support and upstream EOL") Aiven aims to follow the EOL schedule set by the original authors and maintainers of the open source software (the upstream projects). Once the upstream project retires a specific version, they do not receive security updates and critical bug fixes anymore by the maintainers. Outdated services don't offer the level of protection you need, so Aiven follows the upstream project's EOL schedule to ensure that Aiven services are always running on supported versions. ## Service version numbering[​](#service-version-numbering "Direct link to Service version numbering") Aiven services inherit the upstream project's software versioning scheme. Depending on the service, a major version can be either a single digit or in the format `major.minor`. The exact version of the service is visible in the [Aiven Console](https://console.aiven.io/) when the service is running. ## Service version EOL policy[​](#service-version-eol-policy "Direct link to Service version EOL policy") Aiven sets an EOL date for each major version of the service. This policy covers both running and powered-off services on affected versions. ## EOL notifications[​](#eol-notifications "Direct link to EOL notifications") When Aiven sets the EOL date for a service major version: * You receive an email notification along with instructions on the next steps. * The [Aiven Console](https://console.aiven.io/) shows an EOL alert for affected services. * You receive email reminders monthly. * In the month of the EOL date, you receive weekly reminders. ## EOL best practices[​](#eol-best-practices "Direct link to EOL best practices") * Use service forking to test the version upgrade before upgrading your production services. * Upgrade to the supported version before the EOL date. This gives you time to test compatibility, resolve any issues, and plan the upgrade on your schedule. After the EOL date: * If the service is powered on, it's automatically upgraded to the latest version when possible, or to another supported version. note If it's not possible to upgrade a powered-on service to a supported version, the service is powered off and ultimately deleted. * If the service is powered off, it's deleted. ## Version EOL dates[​](#version-eol-dates "Direct link to Version EOL dates") | Version | Aiven EOL | Service creation supported until | Service creation supported from | | ------- | ---------- | -------------------------------- | ------------------------------- | | 8.0.x | 2026-10-31 | 2026-04-30 | 2018-05-18 | | 8.4.x | 2032-10-30 | 2032-04-30 | 2026-04-30 | Related pages * [Manage Aiven for MySQL® versions](/docs/products/mysql/howto/manage-mysql-version.md) --- # Scaling and performance in Aiven for MySQL® Tune performance and manage disk usage to keep your Aiven for MySQL® service running efficiently as it grows. Related pages * [Memory usage](/docs/products/mysql/concepts/mysql-memory-usage.md) * [Tuning and concurrency](/docs/products/mysql/concepts/mysql-tuning-and-concurrency.md) * [Identify disk usage issues](/docs/products/mysql/howto/identify-disk-usage-issues.md) --- # Aiven for OpenSearch® Aiven for OpenSearch® is a fully managed distributed search and analytics suite, deployable in the cloud of your choice. Ideal for logs management, application and website search, analytical aggregations and more. OpenSearch is an open source fork derived from Elasticsearch. [OpenSearch®](https://opensearch.org) is an open-source search and analytics suite including a search engine, NoSQL document database, and visualization interface. OpenSearch offers a distributed, full-text search engine based on [Apache Lucene®](https://lucene.apache.org/) with a RESTful API interface and support for JSON documents. Aiven for OpenSearch and Aiven for OpenSearch Dashboards are available on a cloud of your choice. note OpenSearch and OpenSearch Dashboards projects were forked in 2021 from the formerly open source projects Elasticsearch and Kibana. Aiven for OpenSearch includes OpenSearch Dashboards, giving a fully featured user interface and visualization platform for your data. OpenSearch is designed to be robust and scalable, capable of handling various data types and structures. It provides high-performance search functionality for data of any size or type, and with schemaless storage, it can index various sources with different data structures. OpenSearch is widely used for log ingestion and analysis, mainly because it can handle large data volumes, and OpenSearch Dashboards provide a powerful interface to the data, including search, aggregation, and analysis functionality. ## Ways to use OpenSearch[​](#ways-to-use-opensearch "Direct link to Ways to use OpenSearch") OpenSearch is ideal for working with various types of unstructured data. The most common examples include: * **Log ingestion and analysis:** Send your **logs** to OpenSearch so that you can identify and diagnose problems if they arise. tip [Enable the log integration](/docs/products/opensearch/howto/opensearch-log-integration.md) to send logs from a service to your OpenSearch service. * **Document indexing:** Use OpenSearch to index documents to get meaningful **search results** from a large body of knowledge. ## Benefits of using Aiven for OpenSearch®[​](#benefits-of-using-aiven-for-opensearch "Direct link to Benefits of using Aiven for OpenSearch®") Aiven for OpenSearch service has many benefits, such as: * **Easy setup:** With Aiven, you can set up clusters, deploy new nodes, migrate clouds, and fork databases in a single mouse click. * **Open-source alternative:** Aiven for OpenSearch is an open-source alternative to Elasticsearch. * **Search analytics:** OpenSearch has a lot of search analytics capabilities that you can use with Aiven's other services. * **High uptime:** Aiven ensures that you get 99.99% uptime. * **Scalability:** You can scale up or down as needed. Increase your storage, get more nodes, create new clusters, or expand to new regions. * **Rich set of extensions:** Aiven provides a powerful set of default extensions, including SQL support, anomaly detection, and phonetic analysis. ## OpenSearch resources[​](#opensearch-resources "Direct link to OpenSearch resources") Below are a few resources that can help you learn more about OpenSearch and working with your OpenSearch service: * Work with your OpenSearch service [using cURL](/docs/products/opensearch/howto/opensearch-with-curl.md) * Check the [API documentation](https://docs.opensearch.org/latest/api-reference/) for detailed information about the HTTP endpoints. * There's a [list of plugins](/docs/products/opensearch/reference/plugins.md) supported by Aiven for OpenSearch. * Got a question about the OpenSearch project itself? They have an [FAQ](https://opensearch.org/faq/) for that. Related pages * [Aiven.io](https://aiven.io/opensearch) Apache Lucene is a registered trademark or trademark of the Apache Software Foundation in the United States and/or other countries *Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries.* *Kibana is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries.* --- # Access control in Aiven for OpenSearch® Access control is a crucial security measure that allows you to control who can access your data and resources. By setting up access control rules, you can restrict access to sensitive data and prevent unauthorized changes or deletions. important `avnadmin` is the default service user, and its credentials are included in the default connection string for logging into OpenSearch Dashboards. If you require different access permissions for other service users, restrict their permissions as needed with [Access Control Lists (ACLs)](/docs/products/opensearch/get-started.md#secure-access-with-acls). Aiven for OpenSearch® provides the following ways to manage user accounts and access control in OpenSearch®: ## Method 1: Enable access control on the Aiven Console[​](#method-1-enable-access-control-on-the-aiven-console "Direct link to Method 1: Enable access control on the Aiven Console") When you enable [Access control](/docs/products/opensearch/howto/control_access_to_content.md) in the Aiven Console for your Aiven for OpenSearch® service, you can create service users and set their permissions. The Aiven Console has an easy-to-use interface that helps you manage who can access your data. Aiven for OpenSearch supports index-level access control lists (ACLs) to control permissions and API-level rules to restrict access to specific data sets. With access control enabled, you can customize the access control lists for each user by setting up individual "pattern/permission" rules. The "pattern" parameter specifies the indices to which the permission applies and uses glob-style matching, where \* matches any number of characters (including none) and ? matches any single character. note ACLs apply only to indices and do not control access to other OpenSearch APIs, including OpenSearch Dashboards. ## Method 2: Enable OpenSearch® Security management[​](#method-2-enable-opensearch-security-management "Direct link to Method 2: Enable OpenSearch® Security management") Another way to manage user accounts, access control, roles, and permissions for your Aiven for OpenSearch® service is by [enabling OpenSearch® Security management](/docs/products/opensearch/howto/enable-opensearch-security.md). This method lets you use the OpenSearch Dashboard and OpenSearch API to manage all aspects of your service's security. You can use advanced features such as fine-grained access control with OpenSearch Security. This allows you to specify the exact actions that each user can take within your OpenSearch service. OpenSearch Security also supports SAML integrations, which provide single sign-on (SSO) authentication and authorization for your OpenSearch Service. For more information, see [OpenSearch Security for Aiven for OpenSearch®](/docs/products/opensearch/concepts/os-security.md). note ACLs apply only to indices and do not control access to other OpenSearch APIs, including OpenSearch Dashboards. ## Patterns and permissions[​](#patterns-and-permissions "Direct link to Patterns and permissions") Access control in OpenSearch uses patterns and permissions to manage access to indices. Patterns are glob-style strings that specify the indices to which permissions apply, and permissions determine the level of access granted to users for these indices. ### Patterns[​](#patterns "Direct link to Patterns") Patterns use the following syntax: * `*`: Matches any number of characters (including none) * `?`: Matches any single character ### Permissions[​](#permissions "Direct link to Permissions") The available permissions in Aiven for OpenSearch® are: * `deny`: Explicitly denies access * `admin`: Allows unlimited access to the index * `readwrite`: Grants full access to documents * `read`: Allows only searching and retrieving documents * `write`: Allows updating, adding, and deleting documents ### API access[​](#api-access "Direct link to API access") Permissions determine which index APIs users can access, controlling actions like reading, writing, updating, and deleting documents. * `deny`: No access * `admin`: No restrictions * `readwrite`: Allows access to `_search`, `_mget`, `_bulk`, `_mapping`, `_update_by_query`, and `_delete_by_query` APIs * `read`: Allows access to `_search` and `_mget` APIs * `write`: Allows access to `_bulk`, `_mapping`, `_update_by_query`, and `_delete_by_query` APIs note * When no rules match, access is implicitly denied. * The `write` permission allows creating indices that match the rule's index pattern but does not allow deletion. Indices can only be deleted when a matching `admin` permission rule exists. ## Example[​](#example "Direct link to Example") Consider the following set of rules: * `logs_*/read` * `events_*/write` * `logs_2018*/deny` * `logs_201901*/read` * `logs_2019*/admin` This set of rules allows the user to: * Add documents to `events_2018` (second rule) * Retrieve and search documents from `logs_20171230` (first rule) * Gain full access to `logs_20190201` (fifth rule) * Gain full access to `logs_20190115` (fifth rule, as the `admin` permission gets higher priority than the `read` permission in the fourth rule) This same set of rules denies the service user from: * Gain any access to `messages_2019` (no matching rules) * Read or search documents from `events_2018` (the second rule only grants `write` permission) * Write to or use the API `for logs_20171230` (the first rule only grants `read` permission) note These rules apply only to index access and do not affect OpenSearch Dashboards or other OpenSearch APIs. ## Access control for aliases[​](#access-control-for-aliases "Direct link to Access control for aliases") Aliases are virtual indices that reference one or more physical indices, simplifying data management and search. In OpenSearch, you can define access control rules for aliases to ensure proper security and control over data access. When managing aliases in OpenSearch, note that: * Aliases are not automatically expanded in access control, so the ACL must explicitly include a rule that matches the alias pattern. * Only access control rules that match the alias pattern will be applied. Rules matching the physical indices that the alias references will not be used. ## Access to top-level APIs[​](#access-to-top-level-apis "Direct link to Access to top-level APIs") Top-level API access control depends on whether the security plugin is enabled. If the security plugin is [enabled](/docs/products/opensearch/howto/enable-opensearch-security.md), ACLs are not used to control top-level APIs. Instead, the security plugin handles access control. ### Service controlled APIs[​](#service-controlled-apis "Direct link to Service controlled APIs") The following top-level APIs are controlled by the OpenSearch service and not by the ACLs defined by you: * `_cluster` * `_cat` * `_tasks` * `_scripts` * `_snapshot` * `_nodes` [Enabling OpenSearch Security management](/docs/products/opensearch/howto/enable-opensearch-security.md) provides control over the top-level APIs: `_mget`, `_msearch`, and `_bulk`. note **Deprecated \_ \* patterns** When the security plugin is enabled, `_ *` patterns for top-level API access control are ignored. Access is managed by the security plugin settings. You do not need to configure these patterns manually. ## Access control and OpenSearch Dashboards[​](#access-control-and-opensearch-dashboards "Direct link to Access control and OpenSearch Dashboards") Enabling ACLs does not restrict access to OpenSearch Dashboards. However, all requests made by OpenSearch Dashboards are checked against the current user's ACLs. note Service users with read-only access to certain indices might encounter `HTTP 500` internal server errors when viewing dashboards, as these dashboards use the `_msearch` API. To prevent this, add an ACL rule that grants `admin` access to `_msearch` for the affected service user. ## Next steps[​](#next-steps "Direct link to Next steps") Learn how to [enable and manage access control](/docs/products/opensearch/howto/control_access_to_content.md) for your Aiven for OpenSearch® service. --- # Analyze data with Aiven for OpenSearch® aggregations Alongside the search functionality, OpenSearch® offers a powerful analytics engine able to perform summary calculations of your data, and extract statistics and metrics very rapidly. Results of these "aggregations" can be then visualised with OpenSearch Dashboards. Aggregations can be divided into three groups: * **metric aggregation** performs simple calculations on values extracted from the fields of the documents, for example finding minimum or maximum value, calculating average or collecting statistics about field values. * **bucket aggregation** distributes documents over a set of buckets based on provided criteria. For example, based on predefined ranges of values or based on how often a value is encountered in a field. Bucket aggregation is also used to create histograms. * **pipeline aggregations** combine several aggregations in a way that allows using a result of one aggregation as an intermediate step to create a refined output. With pipeline aggregations you can build moving averages, cumulative sums and perform a variety of other mathematical calculations over the data in your documents. --- # OpenSearch® cross-cluster replication [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Cross-cluster replication (CCR) in Aiven for OpenSearch lets you replicate indices, including their data, mappings, and metadata, from one service to another across different regions and cloud providers. This is an `active-passive` model, where the follower service pulls data from the leader service. Setting up [cross-cluster replication](/docs/products/opensearch/howto/setup-cross-cluster-replication-opensearch.md) establishes the remote-cluster connection between the follower and the leader. You then start replication from the follower for selected indices, or use an auto-follow rule to replicate every index that matches a pattern. After replication starts for an index, the follower automatically pulls new data, mappings, and metadata changes for that index from the leader. Follower services can run in different regions and on different cloud providers. Aiven only supports automated backups on the primary leader cluster, which means that follower clusters are not backed up independently. ## Benefits[​](#benefits "Direct link to Benefits") * **Data locality/proximity:** Replicating data to a cluster closer to the user's geographical location helps reduce latency and response time. * **Horizontal scalability:** Splits a query-heavy workload across multiple replica clusters to improve application availability. ## Limitations[​](#ccr-limitatons "Direct link to Limitations") * Cross cluster replication is not available for Free and Startup plans. * During creation, the follower cluster service must have the same service plan as the leader cluster service, or higher. This ensures that the follower cluster service has at least as much memory as the leader cluster. Service plans can be changed later as needed. * To delete the cross-cluster replication integration, **delete** the follower cluster service. * Maintenance upgrade and major version upgrades must be performed manually on leader and follower services. * During a node recycle event, replication will pause until the service is operational again. Related pages [Set up cross-cluster replication for Aiven for OpenSearch®](/docs/products/opensearch/howto/setup-cross-cluster-replication-opensearch.md). --- # Dedicated node roles in Aiven for OpenSearch® Aiven for OpenSearch® supports dedicated node roles, enabling workload isolation across specialized node groups for optimized performance and scaling. The dedicated node roles capability is generally available (GA) for Aiven for OpenSearch® version 2.19 and later. The cluster topology with node roles is available for 9-node and 15-node service plans, for production workloads that require enhanced performance and reliability. ## Benefits and use cases[​](#benefits-and-use-cases "Direct link to Benefits and use cases") The dedicated node roles feature helps achieve the following: * **Improved stability**: Separating cluster management from data operations prevents resource-intensive queries from affecting cluster coordination, reducing the risk of cluster instability. * **Better scalability**: You can scale data nodes independently from cluster manager nodes, adding capacity where needed without over-provisioning management resources. * **Optimized resource allocation**: Each node group can use hardware configurations tailored to its specific workload, improving cost efficiency. * **Enhanced performance**: Dedicated data nodes can focus entirely on query execution and data processing without the overhead of cluster management tasks. The dedicated node roles feature is particularly beneficial for: * **Large-scale deployments**: Clusters with high data volumes or query throughput benefit from isolating coordination overhead from data operations. * **Performance-critical applications**: Preventing resource contention between cluster management and query execution ensures consistent performance. * **Complex cluster topologies**: Larger clusters with many nodes see stability improvements when cluster management runs on dedicated hardware. ## About dedicated node roles[​](#about-dedicated-node-roles "Direct link to About dedicated node roles") By default, OpenSearch nodes perform all roles: cluster management, data storage, and query processing. With dedicated node roles, you can separate these responsibilities across different node groups, each optimized for specific tasks. This architecture separates the cluster control plane from the data plane. It keeps cluster management operations stable during heavy query loads or data ingestion. ### Available node roles[​](#available-node-roles "Direct link to Available node roles") #### Cluster manager nodes[​](#cluster-manager-nodes "Direct link to Cluster manager nodes") Cluster manager nodes handle cluster-wide operations such as: * Managing cluster state and metadata * Coordinating node membership * Creating and deleting indices * Tracking cluster health * Allocating shards to nodes * Orchestrating cluster-wide operations These nodes run on smaller instances optimized for low-latency coordination tasks rather than data storage. Cluster manager nodes do not store data or handle search requests, allowing them to focus on maintaining cluster stability. note Configure cluster manager nodes in odd numbers to ensure quorum for cluster decisions and prevent split-brain scenarios. #### Data nodes[​](#data-nodes "Direct link to Data nodes") Data nodes are responsible for: * Storing and indexing data * Executing search queries * Performing data aggregations * Running ingest pipelines * Handling client requests * Coordinating distributed requests across the cluster * Providing internal DNS routing for cluster traffic In dedicated-role plans, each data node includes the data, ingest, coordinator, and internal DNS roles. Data nodes typically run on larger instances with more storage and compute resources for data-intensive operations. ### Cluster configuration[​](#cluster-configuration "Direct link to Cluster configuration") Dedicated node roles are defined at the service plan level. When you select a plan with dedicated roles: * Cluster manager nodes are configured as a separate node group with their own instance type. * Cluster manager nodes are automatically distributed across different availability zones. * Data nodes form another group optimized for storage and compute. * The configuration is managed automatically by Aiven. * Cluster manager nodes are excluded from DNS routing for client connections. * Node roles are assigned during cluster creation and maintained throughout the cluster lifecycle. * OpenSearch Dashboards is served on every node. * Dedicated dashboard nodes are not included. All standard service operations work with dedicated node roles, including service creation, major version upgrades, plan changes, service forking, and node replacement. The platform handles cluster manager node operations carefully to maintain cluster stability during updates. ### Node replacement and scaling[​](#node-replacement-and-scaling "Direct link to Node replacement and scaling") During maintenance or scaling operations: * To increase data capacity, move to a dedicated-role plan with more or larger data nodes while keeping the cluster manager layout unchanged. * Cluster manager nodes are replaced last during maintenance updates, including version upgrades, to maintain cluster coordination. * Node failures are handled automatically with role-aware replacement. * Disk space validation and additional disk capacity apply only to data nodes, as cluster manager nodes do not store data. This makes adding disk space more cost-efficient compared to scaling disk across all nodes. ## Manage dedicated node roles[​](#manage-dedicated-node-roles "Direct link to Manage dedicated node roles") The dedicated node roles feature is plan-based. ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Upgrade Aiven for OpenSearch®](/docs/products/opensearch/howto/os-version-upgrade.md) to 2.19 or later if your service runs an older version. ### Start using dedicated node roles[​](#start-using-dedicated-node-roles "Direct link to Start using dedicated node roles") Create an Aiven for OpenSearch® service and choose a plan that includes dedicated node roles, available under Cluster plans. ### Scale a cluster plan[​](#scale-a-cluster-plan "Direct link to Scale a cluster plan") To move to another dedicated-role layout, change the service plan to a different eligible plan, available under Cluster plans. ### Disable dedicated node roles[​](#disable-dedicated-node-roles "Direct link to Disable dedicated node roles") Change the service plan to a plan without dedicated node roles. This returns the service to a standard node layout where nodes share roles. Related pages * [High availability in Aiven for OpenSearch®](/docs/products/opensearch/concepts/high-availability-for-opensearch.md) * [Shards and replicas](/docs/products/opensearch/concepts/shards-number.md) * [Service plans](/docs/platform/concepts/service-pricing.md) --- # High availability in Aiven for OpenSearch® Aiven for OpenSearch® is available on a variety of plans, offering different levels of high availability. The selected plan defines the features available, and a summary is provided in the table below: | Plan | High availability features | Backup history | | ------------ | ----------------------------------------------------------- | ---------------------------------------- | | **Free** | Single-node with limited availability | Single backup for disaster recovery | | **Startup** | Single-node with limited availability | 2 days, with hourly backup for 24 hours | | **Business** | Three-node cluster configured for high availability | 14 days, with hourly backup for 24 hours | | **Premium** | Six-node (or more) cluster configured for high availability | 30 days, with hourly backup for 24 hours | ## Failure handling[​](#failure-handling "Direct link to Failure handling") **Minor failures**, such as service process crashes or temporary loss of network access, are handled by Aiven automatically in all plans without any major changes to the service deployment. The service automatically restores normal operation once the crashed process is automatically restarted or when the network access is restored. **Severe failures**, such as loss of a cluster node, require more drastic recovery measures. Aiven platform continuously monitors the health of every node in a cluster. When a node reports failures from its own self-diagnostics or when no response to a health check is returned, Aiven platform starts a replacement node. While the node is being replaced, client requests are rerouted to the other nodes that have replica shards. When the new node joins the cluster, it restores data from the existing nodes or from a backup if no old nodes can be reached anymore, and starts servicing requests when data restoration is completed. ## High Availability in Business and Premium plans[​](#high-availability-in-business-and-premium-plans "Direct link to High Availability in Business and Premium plans") To achieve high availability in business and premium plans, Aiven for OpenSearch configures your indices to have at least one replica. You will want to choose a cloud region with at least three availability zones to ensure that the primary and replica shards will not be down at the same time. The setting that controls the replication factor in index settings is called `number_of_replicas`. Having its value as 1 is enough to have your data resilient to an outage in a single zone, and setting it to higher values will add tolerance for more availability zone failures. Cluster nodes in business and premium plans are distributed evenly across availability zones in a cloud region. Moreover, each node is configured to be aware of its zone. Therefore a primary copy of your data and a replica copy will be allocated on the nodes in different availability zones. Aiven platform constantly monitors all nodes in an OpenSearch cluster. When a node stops responding to a health check for sufficiently long time, Aiven platform starts a replacement node, waits for the node to start, and replaces the failed node with the new node in the cluster's configuration. note The amount of time it takes for a new node to become fully operational depends mainly on the used cloud region and the amount of data that needs to be copied from primary shards. However, in a multi-node cluster on a Business or Premium plan all the nodes with in-sync shards will keep responding to client requests even before the new node is fully operational. Moreover, for every primary shard on the lost node, the cluster promotes one in-sync replica shard. Write requests from the clients will be routed to the promoted shards. All of this is automatic and requires no administrator intervention. ## Single-node Free and Startup service plans[​](#single-node-free-and-startup-service-plans "Direct link to Single-node Free and Startup service plans") Losing the only node from the service starts the automatic process of creating a new replacement node. The new node starts up, restores its state from the latest available backup and resumes serving customers. Since there was just a single node providing the service, the service will be unavailable for the duration of the restore operation. All the write operations made since the last backup are lost. --- # Hot/warm data tiering in Aiven for OpenSearch® [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Hot/warm data tiering lets you store recent, frequently queried data on fast nodes and older data on cheaper nodes—without splitting your cluster or changing how you search. Hot/warm data tiering is in [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-) for Aiven for OpenSearch® version 2.19 and later. It is available on custom plans only. tip [Contact Aiven](https://aiven.io/contact) to request a custom plan with hot/warm data tiering enabled. ## How hot/warm data tiering works[​](#how-hotwarm-data-tiering-works "Direct link to How hot/warm data tiering works") A tiered cluster has two groups of data nodes. Each group carries a `node.attr.temp` attribute: * **Hot nodes** use faster disks and are sized for active writes and frequent queries. * **Warm nodes** use larger, lower-cost disks and are sized for data that is queried less often. OpenSearch shard allocation filtering places index shards on the correct tier using the `index.routing.allocation.require.temp` index setting. Set this to `hot` to pin new indices to hot nodes. When an index ages, change the setting to `warm` to move shards to warm nodes. [Index State Management (ISM)](/docs/products/opensearch/howto/migrate-ism-policies.md) automates these transitions. An ISM policy rolls over an index when it reaches a size or age threshold, migrates it to warm storage after a retention period, and optionally deletes it. ## Dynamic Disk Sizing[​](#dynamic-disk-sizing "Direct link to Dynamic Disk Sizing") Dynamic Disk Sizing adds capacity to both tiers at the same time. You cannot expand a single tier on its own. The added space is distributed proportionally to each tier's base volume size. For example, if hot nodes have 100 GiB of total volume and warm nodes have 200 GiB, adding 90 GiB allocates 30 GiB to the hot tier (100 / 300 × 90) and 60 GiB to the warm tier (200 / 300 × 90). This behavior might change in the future. ## Supported OpenSearch versions[​](#supported-opensearch-versions "Direct link to Supported OpenSearch versions") Hot/warm data tiering requires Aiven for OpenSearch® 2.19 or later. Version 3.3 and later is also supported. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A custom plan with hot and warm node groups. [Contact Aiven](https://aiven.io/contact) to get one configured for your account. * Aiven for OpenSearch® 2.19 or later. Related pages * [Manage hot/warm data tiering](/docs/products/opensearch/howto/hot-warm-tiering.md) * [Dedicated node roles in Aiven for OpenSearch®](/docs/products/opensearch/concepts/dedicated-node-roles.md) * [Index State Management policies](/docs/products/opensearch/howto/migrate-ism-policies.md) --- # Replication factors in Aiven for OpenSearch® The replication factor in Aiven for OpenSearch® determines the number of copies (replicas) of each index shard. Replicas protect against data loss, and OpenSearch can also route search queries to replicas, so more replicas can spread search load across more nodes. The `number_of_replicas` is an index-level setting that defines how many replicas each primary shard has. By default, it is set to `1`, meaning each shard has one replica for redundancy. You can configure this when creating or updating an index in the OpenSearch Dashboard. ## Automatic replication factor adjustment[​](#automatic-replication-factor-adjustment "Direct link to Automatic replication factor adjustment") Aiven for OpenSearch automatically adjusts replication factors in your indexes to maintain data availability. * The maximum `number_of_replicas` is the total number of nodes in your cluster minus one. This ensures that your data is replicated across all available nodes. **Example:** In a 3-node cluster, the maximum `number_of_replicas` is `2`, which replicates all shards across the three nodes. * If you set `number_of_replicas` to `0` in a multi-node cluster, Aiven automatically increases it to `1`. This ensures that your data remains available if one node fails. ## Replication factor 0[​](#replication-factor-0 "Direct link to Replication factor 0") Setting the replication factor (`number_of_replicas`) to `0` means your data has no replicas. This reduces storage usage but significantly increases the risk of data loss if a node in the cluster fails. note Before enabling this configuration, consult with your account manager to discuss your use case and agree on the reduced SLA. ### When to use replication factor 0[​](#when-to-use-replication-factor-0 "Direct link to When to use replication factor 0") Replication factor `0` is generally not recommended for most OpenSearch use cases. In specific cases, you might consider it for: * **Non-critical environments:** QA, testing, or development clusters where data loss is acceptable. * **Temporary data:** Scenarios where data can be recreated and storage costs need to be minimized. ### Risks and considerations[​](#risks-and-considerations "Direct link to Risks and considerations") Setting the replication factor to `0` increases the risk of data loss. If a node failure occurs, Aiven for OpenSearch automatically restores data from the latest snapshot, but any data added after that snapshot will be lost. If no snapshot exists, an empty index is created, resulting in the loss of any unsaved data. note * In a scenario where a node fails before the first snapshot is taken, the system cannot recover the data. Aiven for OpenSearch automatically recreates the missing index, but the recreated index is empty. * In a normal node recovery scenario, Aiven for OpenSearch loads the latest snapshot to restore your data. ### Set replication factor 0[​](#set-replication-factor-0 "Direct link to Set replication factor 0") Before enabling this option, consult with your account manager to discuss your use case and agree on the reduced SLA. Afterward, contact [Aiven support](mailto:support@aiven.io) to enable the option to set the replication factor (`number_of_replicas`) to `0`. --- # Manage indices in Aiven for OpenSearch® Learn how documents, indices, shards, and replicas relate to each other in Aiven for OpenSearch®, and how mapping, aliases, and index lifecycle management keep your indices efficient. ## Documents and indices[​](#documents-and-indices "Direct link to Documents and indices") OpenSearch® stores data as documents, which are JSON records made up of fields and their values. An index is a named collection of documents that share a similar structure and purpose, similar to a table in a relational database. Each index has a single mapping that defines its fields, rather than multiple types within the index. Aiven for OpenSearch doesn't limit the number of indices you can create. warning Avoid storing all your data in a single, continuously growing index with one primary shard. A single-shard index can't spread its data or query load across multiple nodes, and it runs into the memory and recovery-time limits described in [Shards and replicas](#shards-and-replicas) sooner than a properly sized index would. ## Mapping[​](#mapping "Direct link to Mapping") Mapping defines the fields in an index and their data types, similar to a schema in a relational database. When OpenSearch sees a field for the first time in a document, it can create that field automatically. This is known as dynamic mapping. Dynamic mapping causes two problems in production. Every field gets indexed even if you never query most of them, which wastes CPU, increases index size on disk, and uses more memory than a deliberate mapping would. It also locks in a field's type from the first document that introduces it. For example, if the first document sets `user_id` to `12345`, OpenSearch maps `user_id` as a number. A later document that sets `user_id` to `A123` fails to index as a mapping conflict. If you write through the bulk API without checking each item's result, that failure can go unnoticed. Explicit mapping also lets you opt individual fields out of indexing. For example, if a field holds a large value you only need to retrieve, such as a full description, but never search or filter on, set `"index": false` for that field. OpenSearch then stores the value without building search structures for it, which saves CPU, disk space, and memory. For production workloads, define an explicit mapping when you create an index so that field types stay consistent and predictable. For the full mapping syntax, see the [OpenSearch mapping documentation](https://opensearch.org/docs/latest/field-types/). ## Shards and replicas[​](#shards-and-replicas "Direct link to Shards and replicas") Each index is split into primary shards, the units that OpenSearch distributes across the nodes in your cluster. The number of primary shards is set when you create the index and can't be changed later, except by creating a different index with the Split API or by reindexing. A replica is a copy of a primary shard. Replicas protect against data loss, and OpenSearch can also serve search queries from replicas, so more replicas can spread read load across more nodes, not just add redundancy. For details on how Aiven for OpenSearch manages replicas, see [Replication factors in Aiven for OpenSearch®](/docs/products/opensearch/concepts/index-replication.md). Shard count and size directly affect performance: * **Memory**: Aggregations and searches over a large shard build large data structures in memory. Too many large shards can exhaust the memory available to the service. * **Recovery time**: During recovery, such as a version upgrade or node replacement, OpenSearch copies each shard in full. Large shards take longer to recover and generate more disk I/O, which degrades service performance while recovery is in progress. For guidance on choosing a shard count and size, see [Optimal number of shards](/docs/products/opensearch/concepts/shards-number.md). ## When to create an index[​](#when-to-create-an-index "Direct link to When to create an index") The number of indices in your service directly affects performance, so plan your indexing strategy before you create indices at scale. For guidance on when to create a dedicated index instead of reusing an existing one, see [When to create an index](/docs/products/opensearch/concepts/when-create-index.md). ## Indices vs data streams[​](#indices-vs-data-streams "Direct link to Indices vs data streams") For continuously generated, time-series data such as logs and metrics, OpenSearch also offers data streams as an alternative to managing rolling indices manually. A data stream groups a sequence of backing indices under a single name and creates a new backing index through rollover automatically, so you always write to the current one without managing an alias yourself. Use manually managed indices with an alias when you need direct control over rollover timing, mapping changes between generations, or per-index lifecycle policies. Use a data stream when your data is append-only and time-ordered, and you want OpenSearch to manage index creation and rollover for you. For setup and query syntax, see the [OpenSearch data streams documentation](https://opensearch.org/docs/latest/im-plugin/data-streams/). ## Aliases[​](#aliases "Direct link to Aliases") An alias is a name that points to one or more indices. Applications and queries can target an alias instead of a specific index name, so you can change which indices the alias points to without updating client code. Aliases are useful for the following: * **Zero-downtime reindexing**: Point an alias at a new index after reindexing, then switch the alias to it in a single atomic operation. * **Grouping indices**: Query multiple indices, such as `logs-2026.01` and `logs-2026.02`, through a single alias. * **Rollover**: Automatically create an index and switch a write alias to it when the current index reaches an age, size, or document count threshold. Rollover aliases are commonly used together with Index State Management. For a worked example, see [Hot and warm tiering](/docs/products/opensearch/concepts/hot-warm-tiering.md). ## Index lifecycle management[​](#index-lifecycle-management "Direct link to Index lifecycle management") Aiven for OpenSearch includes the Index State Management (ISM) plugin, which automates actions such as rollover, replica changes, and deletion based on policies you define. Historically, some teams managed retention by embedding dates in index names, such as `logs-2018-07-20`, and deleting old indices on a schedule. ISM policies replace this manual approach. Define a policy once, apply it to a rollover alias, and let OpenSearch handle transitions and deletions automatically. Configure ISM policies through the OpenSearch API using the `_plugins/_ism` endpoints. For an example that creates and applies a policy, see [Hot and warm tiering](/docs/products/opensearch/concepts/hot-warm-tiering.md). ### Index retention patterns[​](#index-retention-patterns "Direct link to Index retention patterns") Aiven for OpenSearch also provides an index retention feature you configure directly in the Aiven Console. It caps the number of indices that match a name pattern and deletes the oldest ones once that limit is exceeded. This is a legacy approach that predates ISM being available on Aiven for OpenSearch. It can't combine other lifecycle actions, such as rollover or tiering, into a single policy, and it only deletes indices, so it can't move data to lower-cost storage tiers. Use ISM for any index you're setting up now. Only rely on index retention patterns for existing setups you haven't migrated to ISM yet. For setup steps, see [Index retention patterns](/docs/products/opensearch/howto/set_index_retention_patterns.md). Related pages * [Optimal number of shards](/docs/products/opensearch/concepts/shards-number.md) * [Replication factors in Aiven for OpenSearch®](/docs/products/opensearch/concepts/index-replication.md) * [When to create an index](/docs/products/opensearch/concepts/when-create-index.md) * [Reindex data in Aiven for OpenSearch®](/docs/products/opensearch/howto/reindex-opensearch.md) --- # Aiven for OpenSearch® free tier Get started with Aiven for OpenSearch® at no cost. The Aiven for OpenSearch free tier is a fully managed service for learning, prototyping, and evaluation. No credit card is required. ## When to use the free tier[​](#when-to-use-the-free-tier "Direct link to When to use the free tier") Use the free tier to: * Explore Aiven for OpenSearch concepts with a managed service * Build or test small search applications * Evaluate Aiven for OpenSearch before choosing a paid plan * Test and experiment during early development * Run small proofs of concept or demonstrations The free tier supports limited-scale workloads. For production use, longer retention, or more advanced features, choose a paid plan. ## What the free tier includes[​](#what-the-free-tier-includes "Direct link to What the free tier includes") The free tier features: * Managed Aiven for OpenSearch cluster with a fixed configuration * 4 GB RAM, 20 GB storage * 2 shards * Snapshot retention up to 3 days * Sample dataset tools for testing indexing and search * Basic monitoring for metrics and logs Standard Aiven for OpenSearch clients can connect the same way as they do to paid Aiven for OpenSearch plans. ## Limitations[​](#limitations "Direct link to Limitations") Free tier services have the following restrictions. ### Performance and configuration limits[​](#performance-and-configuration-limits "Direct link to Performance and configuration limits") * Max 20 shards per node * 50 concurrent connections * Fixed storage and performance limits * Fixed cluster settings ### Features not available[​](#features-not-available "Direct link to Features not available") * Service integrations such as logs, metrics, external authentication (SAML, OIDC, JWT) * Custom dictionary files * No security plugin * No import from custom repositories * Custom domains * Dynamic disk sizing ### Service restrictions[​](#service-restrictions "Direct link to Service restrictions") * One free Aiven for OpenSearch service per organization * Paid services cannot be reverted to the free tier * Cloud provider is fixed with a limited set of regions * Cloud region is fixed: You cannot migrate your free tier service to another region. * Free tier services are not covered by an SLA * Maintenance window is fixed. * Additional disk storage cannot be added. * Creation available only through the Aiven Console ## How free tier services operate[​](#how-free-tier-services-operate "Direct link to How free tier services operate") Free tier Aiven for OpenSearch services operate as follows: * **Idle shutdown:** The service powers off automatically if there is no continuative activity on the service. You receive a notification before shutdown and can power on the service from the Aiven Console. * **First-use shutdown:** A new free tier service with no initial usage can power off within the first few hours after the service is running. You can power the service back on from the Aiven Console. * **Alerts:** Platform alerts are not routed to Aiven operators. * **Configuration updates:** Aiven may change the cloud provider, region availability, or configuration of free services. Related pages * [Create a free tier Aiven for OpenSearch service](/docs/products/opensearch/howto/create-free-tier-opensearch.md) * [Aiven free tier services](/docs/platform/concepts/service-pricing.md#free-tier) --- # Key considerations and system adaptation for OpenSearch® Security management Before enabling OpenSearch Security management for Aiven for OpenSearch service, understand its impact on your current system and adapt your infrastructure accordingly. ## Test OpenSearch® Security management[​](#test-opensearch-security-management "Direct link to Test OpenSearch® Security management") If you are an existing Aiven for OpenSearch user, testing OpenSearch Security management before applying it to your running services is advisable. To do this, you have two options: 1. Create a service for testing purposes. 2. Fork one of your existing running services to create a test environment. ## Enabling OpenSearch® Security management[​](#enabling-opensearch-security-management "Direct link to Enabling OpenSearch® Security management") Once OpenSearch Security management is enabled, the following changes will occur: 1. Your current Aiven for OpenSearch users and roles will be transferred to OpenSearch. 2. Aiven for OpenSearch Access Control List (ACL) APIs will be deactivated. The affected API endpoints are: * [ServiceOpenSearchAclGet](https://api.aiven.io/doc/#tag/Service:_OpenSearch/operation/ServiceOpenSearchAclGet) * [ServiceOpenSearchAclSet](https://api.aiven.io/doc/#tag/Service:_OpenSearch/operation/ServiceOpenSearchAclSet) * [ServiceOpenSearchAclUpdate](https://api.aiven.io/doc/#tag/Service:_OpenSearch/operation/ServiceOpenSearchAclUpdate) 3. Aiven for OpenSearch ACL in the Aiven Terraform Provider will be disabled. The affected resources are: * [Opensearch\_acl\_config](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/opensearch_acl_config) * [Opensearch\_acl\_rule](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/opensearch_acl_rule) * [Opensearch\_user](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/opensearch_user) 4. Managing OpenSearch ACL through the [Aiven Console](https://console.aiven.io/) will no longer be available. Instead, use the OpenSearch API and OpenSearch Security dashboard for user and role management. ## Adapting your system, automation, and infrastructure[​](#adapting-your-system-automation-and-infrastructure "Direct link to Adapting your system, automation, and infrastructure") Preparing and adapting your systems, automation processes, and infrastructure is important to accommodate the changes introduced by OpenSearch Security Management for Aiven for OpenSearch. This adaptation ensures that your organization can seamlessly transition to the new security measures and benefit from enhanced protection. 1. Assess the impact of OpenSearch Security management on your existing systems and infrastructure, identifying areas that require modification to support the new security features. 2. Review and update your automation processes to incorporate the security configurations, policy, and control implementation offered by OpenSearch Security, ensuring consistency across your system and infrastructure. 3. Test and validate the updated systems, automation, and infrastructure to ensure compatibility and functionality with the new security features. 4. Provide training and support to your teams to familiarize them with the changes and ensure they can effectively manage the transition. 5. Monitor and adjust your systems, automation, and infrastructure as needed, maintaining optimal performance and security in alignment with your organization's requirements. --- # OpenSearch® vs Elasticsearch OpenSearch® is the open-source version of the Elasticsearch project, which has [a restrictive license](https://www.elastic.co/blog/licensing-change). Third parties cannot offer Elasticsearch as a service. The community (including Aiven) joined forces to create and maintain OpenSearch based on the last open source licensed releases of both Elasticsearch and Kibana (v7.10.2). Version 1.0 release of OpenSearch should be very similar to the Elasticsearch release that it is based on, and Aiven encourages all customers to upgrade at their earliest convenience. This is to ensure that your platforms can continue to receive upgrades in the future. To start exploring Aiven for OpenSearch®, see [Get Started with Aiven for OpenSearch®](/docs/products/opensearch/get-started.md). --- # OpenSearch Security for Aiven for OpenSearch® OpenSearch Security is a powerful feature that enhances the security of your OpenSearch service. By [enabling OpenSearch Security management](/docs/products/opensearch/howto/enable-opensearch-security.md), you can implement fine-grained access controls, SAML authentication, and audit logging to track and analyze activities within your OpenSearch environment. In addition to Role-Based Access Control (**RBAC**), the following external authentication methods are supported for Aiven for OpenSearch Security: * Security Assertion Markup Language (**SAML**) * OpenID Connect (**OIDC**) With OpenSearch Security enabled, you can manage user access and permissions directly from the [OpenSearch Dashboard](/docs/products/opensearch/dashboards.md), giving you full control over your service's security. warning * Once you have enabled OpenSearch Security management, you can no longer use [Aiven Console](https://console.aiven.io/), [Aiven API](https://api.aiven.io/doc/), [Aiven CLI](/docs/tools/cli.md), [Aiven Terraform Provider](/docs/tools/terraform.md) or [Aiven Operator for Kubernetes®](/docs/tools/kubernetes.md) to manage access controls. * You must use the OpenSearch Security Dashboard or OpenSearch Security API for managing user authentication and access control after enabling OpenSearch Security management. * Once enabled, OpenSearch Security management cannot be disabled. If you need assistance disabling OpenSearch Security management, contact [Aiven support](https://aiven.io/support-services). note To implement basic and simplified access control, you can use [Aiven's Access Control Lists (ACL)](/docs/products/opensearch/howto/control_access_to_content.md) to manage user roles and permissions. ## OpenSearch Security use cases[​](#opensearch-security-use-cases "Direct link to OpenSearch Security use cases") OpenSearch Security is a versatile and valuable feature that can meet the needs of a wide range of customers. Some common use cases for this feature include: * **Single Sign-On integration:** If you use an identity management software like Okta or Azure AD, you can use SAML integration to access your OpenSearch Dashboard through your identity provider. With this feature, you won't need to create and manage separate login credentials for OpenSearch, simplifying the authentication process. * **Advanced access control:** If you need different levels of access controls for your employees, OpenSearch Security's Role-Based Access Control (RBAC) can help. With RBAC, you can set up different roles with different access levels and map them to different users. Role mapping is also available with SAML integration, making this feature ideal for enterprises. * **Compliance and audit logging:** If OpenSearch is a critical database in your system, you may need to document a historical record of activity for compliance purposes and other business policy enforcement. OpenSearch Security includes an audit logs feature to help you meet these needs. * **Multi-tenancy:** If you're a reseller or have many smaller departments, you may need different tenants on your OpenSearch Dashboard. OpenSearch Security provides multi-tenancy capabilities, ensuring that each tenant's data is kept separate and secure. ## Key OpenSearch Security features[​](#key-opensearch-security-features "Direct link to Key OpenSearch Security features") OpenSearch Security in Aiven for OpenSearch service offers a range of features to manage the security and access control of your OpenSearch service. These include: * **Role-based access controls:** With OpenSearch Security, you can set up role-based access controls and advanced level security, such as document-level security, field-level security, user and role mapping, and field masking to control access to sensitive data. note User impersonation and cross-cluster search are not supported for the beta release. * **SAML integration:** Aiven for OpenSearch provides basic SAML integration. This allows you to access your OpenSearch Dashboard through the identity provider of your choice. note Aiven for OpenSearch provides basic SAML integration for the beta release, and certain features, such as logout support and request signing, is not be included. * **OpenID Connect integration**: Aiven for OpenSearch enables you to securely authenticate and authorize users through OpenID Connect, enhancing authentication options. * **OpenSearch Dashboard multi-tenancy:** OpenSearch Security provides OpenSearch Dashboard multi-tenancy, which allows you to have different tenants on your OpenSearch Dashboard. * **OpenSearch audit-logs:** Aiven for OpenSearch Security includes OpenSearch Audit-logs, which allow you to document a historical record of activity for compliance purposes and other business policy enforcement. ## OpenSearch Security management changes and impacts[​](#opensearch-security-management-changes-and-impacts "Direct link to OpenSearch Security management changes and impacts") Enabling OpenSearch Security management on your Aiven for OpenSearch service through the Aiven console triggers several changes: * Users and role-based access control will be managed through the OpenSearch Security dashboard or OpenSearch Security API. * The `os-sec-admin` user will initially be mapped to the pre-defined role `service_security_admin_access`, which provides unrestricted access to the service, including the OpenSearch Security API and OpenSearch Security dashboard. * As an `os-sec-admin` user, you can add or remove users from pre-defined roles, and create new roles and assignments, but some pre-defined roles cannot be changed or deleted. * All service users defined before enabling OS Security management are included in OpenSearch's internal users, with the attribute `provider_managed: False`. However, the users `avnadmin` and `os-sec-admin`, are still managed by the service platform and have the attribute `provider_managed:true`. While service platform management of these users is limited to password changes, they can still be assigned to different roles as needed in the OpenSearch Security dashboard. For information on how to enable OpenSearch Security management on Aiven Console, see [Enable OpenSearch® Security management for Aiven for OpenSearch®](/docs/products/opensearch/howto/enable-opensearch-security.md). --- # Memory and out-of-memory conditions in Aiven for OpenSearch® Understand the memory limits and out-of-memory conditions that apply to your Aiven for OpenSearch® service. ## Service memory limits[​](#service-memory-limits "Direct link to Service memory limits") The practical memory limit will always be less than the service physical memory limit. **All services are subject to operating overhead:** * A small amount of memory is required by the operating system kernel to manage system resources, including networking functions and disk cache. * Aiven's cloud data platform requires memory to monitor availability, provide metrics, logging and manage backups. A server or node's **usable memory** can be calculated as: `usable memory = RAM - overhead` Where: * `overhead` is 350 MiB (≈ 0.34 GiB). Services may utilize optional components, service integrations, connection pooling, or plug-ins, which are not included in overhead calculations. If a service is overcommitted, the operating system, management layer, backups or availability monitoring, may fail status checks or operations due to resource contention. In severe instances, the node may fail completely with an out-of-memory condition. ## Out of memory conditions[​](#out-of-memory-conditions "Direct link to Out of memory conditions") Many processes request more memory from the kernel than they will ever use or need. In these cases, the kernel overallocates memory. This allows it to satisfy multiple processes requesting more memory than is available, which is not used or is freed by the time any other process actually needs it. However, if enough processes start using all their allocated memory simultaneously there may not be enough physical memory available and an `Out Of Memory` (`OOM`) condition occurs. warning This situation is critical and must be resolved immediately. The solution that the Linux kernel employs is to invoke the `Out of Memory Killer` (or `OOM Killer`). This reviews all running processes and kills one or more of them to free up system memory and keep the system running. The `OOM Killer` selects process to kill based on an `oom_score`; a calculation that balances how much memory the process is using with how long the process has been running. Processes that have been running for a long time are less likely to be killed. Subprocesses are summed with parent processes in terms of memory usage, so a process which forks many subprocesses, but itself does not use a lot of memory, may still be killed. In most instances, the hosted data service, or a child process, will have the highest memory footprint and be a prime candidate for termination when the OOM Killer inspects the running processes. Aiven's cloud data platform leverages kernel namespaces or containers to isolate processes from each other. Isolation has several benefits, including: * A smaller footprint for security‑related concerns * A smaller blast radius for failure * Greater control of system resources Left unchecked, the `OOM Killer` may opt to kill the primary service. This is undesirable as unclean termination of the primary service can lead to data loss, inconsistency, or corrupted backups. Further, if Aiven's management platform detects that the primary service is unavailable for , the service will be marked as down and a failover will occur. To mitigate this scenario, namespaces are used, some with additional memory limits, in combination with an `oom_score_adjust` on the primary process, to coax the `OOM Killer` into selection of less critical processes. This will still result in a service restart, but in a more controlled process, where the database is shut down, rather than killed; exposure to data loss is limited and recovery is faster when the service restarts, often avoiding failover. warning Out of Memory conditions can still lead to unexpected behavior, including data unavailable or data loss conditions. ## Avoid running low on memory[​](#avoid-running-low-on-memory "Direct link to Avoid running low on memory") The OOM killer only runs when the system is critically low on memory. To prevent it from running, either reduce your memory usage or increase the available memory. For most databases, the service memory footprint can often be reduced by: * Reducing concurrency or implementing connection pooling * Tuning queries to limit result sets * Tuning indexes for query load * Dropping unused objects from storage In cases where the working set no longer fits into memory, consider scaling your service. Related pages * [Change the service plan](/docs/products/opensearch/howto/change-service-plan.md) * [Scale disk storage](/docs/products/opensearch/howto/scale-disk-storage.md) --- # Optimal number of shards A key component of using OpenSearch® is determining the optimal number of shards for your index. Learn how to choose the appropriate number of shards and maximizing performance. For a broader overview of index configuration, including shard sizing guidelines, see [Shards and replicas](/docs/products/opensearch/concepts/indices.md#shards-and-replicas). ## Considerations for optimal shard count[​](#considerations-for-optimal-shard-count "Direct link to Considerations for optimal shard count") The ideal number of shards depends on your data volume, usage patterns, and expected data growth. As a starting point, aim for a shard size of about 30 GB. For example, an index with 100 GB of expected data can start with 3 to 4 shards. A more critical limit than any single index's shard count is the total number of shards your service's memory can support. As a rule of thumb, don't exceed 20 shards per GB of memory available to the service. Going over this significantly increases the risk of the service running out of memory, regardless of how those shards are distributed across indices. For a multi-node OpenSearch® service, Aiven enforces a minimum of one replica per shard to ensure high availability and data redundancy, and replicas also let OpenSearch spread search queries across more nodes. While there is no limit on the number of replicas per shard, adding too many can impact performance and increase disk usage. ## Determining shard count[​](#determining-shard-count "Direct link to Determining shard count") Base your shard count on a target shard size, since the right target depends on how you query the data rather than on a single fixed divisor: * **Search-heavy indices:** Target 10-30 GB per shard. Use the lower end of that range for a smaller total data volume, and the higher end as total data volume grows. * **Write-heavy or seldom-queried indices**, such as logs: Target 30-50 GB per shard. Divide your expected total data volume by your target shard size to get a starting shard count. For example, a 250 GB search-heavy index at 25 GB per shard starts with 10 shards. For a small data volume spread across many indices, start with one shard per index and split the index later if it grows. These are starting points. Monitor disk and CPU usage, and adjust as your usage patterns and data volume evolve. ## Use the plan calculator[​](#use-the-plan-calculator "Direct link to Use the plan calculator") To help configure your shards, the OpenSearch plan calculator is available for online use or download: * [View on Google Docs](https://docs.google.com/spreadsheets/d/1wJwzSdnQiGIADcxb6yx1cFjDR0LEz-pg13U-Mt2PEHc) - Make a copy to your Google drive to use it. * [Download XLSX](https://docs.google.com/spreadsheets/d/1wJwzSdnQiGIADcxb6yx1cFjDR0LEz-pg13U-Mt2PEHc/export) - Download and use it locally. Enter details like the number of nodes, CPUs, RAM, and max shard size to get recommended starting values for your setup. ![Screenshot of the spreadsheet: enter your information and get recommendations.](/docs/assets/images/opensearch-plan-calculator-b21728264e07096fbff1b4ff807ffcae.png) Yellow cells such as `data node count`, `CPUs`, `RAM`, `Max Shard Size` are input fields used to calculate recommended plan sizes. warning Dashboards from Aiven for OpenSearch are not compatible across minor versions of OpenSearch. If your service instance runs an older OpenSearch version, expect downtime during migration or plan changes. ## Adjusting shard count[​](#adjusting-shard-count "Direct link to Adjusting shard count") OpenSearch doesn't let you change the shard count of an existing index directly, so increasing or decreasing it always means creating a different index. **Best practice**: If you're not sure yet what shard count, mappings, or other settings you'll eventually need, create every index behind an alias from the start. Point your application at the alias, not the index name, from day one: ``` PUT /events-v1 { "settings": { "index.number_of_shards": 1, "index.number_of_replicas": 1 }, "aliases": { "events": {} } } ``` Configure your application to use the alias `events`, not the index name `events-v1`, directly. Later, if you need a different shard count, mappings, or other settings, create `events-v2` with the settings you want, copy the data across, and switch the alias to it in a single atomic step, with no application changes. If you didn't set up an alias in advance, use one of the following instead: * **Split the index**: Use the Split API to create an index with a multiple of the current shard count. The new index keeps the source index's mappings and settings automatically, but it gets a new name, so you still need an alias if you want existing clients to keep using the original name. * **Reindex to a new index**: Create an index with the shard count you want and copy the data across with the Reindex API. This works for any index, but unlike the Split API, it doesn't carry over mappings and settings automatically, and it needs more planning around writes that happen during the copy. For step-by-step instructions, see [Manage large shards in Aiven for OpenSearch®](/docs/products/opensearch/howto/resolve-shards-too-large.md). If you're using OpenSearch for daily logs or a similar rolling pattern, you can also change the shard count for new indices going forward, without touching existing ones. OpenSearch automatically rebalances shards across nodes to even out overall disk usage per node. This isn't driven by individual shard size. A single large shard doesn't trigger rebalancing on its own, though its size can limit which nodes have enough free disk space to host it. --- # When to create an index Determine wether to create an index per customer/project/entity or to look for alternatives. Consider creating an index per customer/project/entity when: * You have a very limited number of entities (tens, not hundreds or thousands); * You need to efficiently delete all the data related to a single entity. For example, storing logs or other events on per-date indexes ( `logs_2018-07-20`, `logs_2018-07-21` etc.) adds value assuming old indexes are cleaned up. If you have low-volume logging and want to keep indexes for very long time (years?), consider per-week or per-month indexes instead. It is **not** recommended to create an index per customer/project/entity if: * You have potentially a very large number of entities (thousands), or you have hundreds of entities and need multiple different indexes for each and every one; * You expect a strong growth in number of entities; * You have no other reason than separating different entities from each other; Instead of creating something like `items_project_a`, consider using a single `items` index with a field for project identifier, and query the data with OpenSearch® filtering. This will be far more efficient usage of your OpenSearch service. note Aiven does not place any restrictions on the number of indices in your OpenSearch service. --- # OpenSearch® Dashboards OpenSearch® Dashboards is both a visualisation tool for data in the cluster and a user interface for OpenSearch plugins. It's included with your Aiven for OpenSearch service and accessible through the browser. note * OpenSearch Dashboards was forked in 2021 from the formerly open source project Kibana. * Starting with Aiven for OpenSearch® versions 1.3.13 and 2.10, OpenSearch Dashboards will remain available during a maintenance update that also consists of version updates to your Aiven for OpenSearch service. OpenSearch Dashboards is ideal for data visualisation and monitoring. It allows building dashboards using visual tools and provides an interface to run DSL and SQL queries. ![A screenshot of the OpenSearch Dashboard for flights example dataset.](/docs/assets/images/dashboard-example-9dfa797f561a5617d75bda3335a23d38.png) ## OpenSearch Dashboards resources[​](#opensearch-dashboards-resources "Direct link to OpenSearch Dashboards resources") * Official OpenSearch Dashboards documentation * Dashboards Query Language * How to create reports *Kibana is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries.* --- # Get started with Aiven for OpenSearch® Dashboards To start using **Aiven for OpenSearch® Dashboards**, [create Aiven for OpenSearch® service first](/docs/products/opensearch/get-started.md) and OpenSearch Dashboards service will be added alongside it. Once the Aiven for OpenSearch service is running, the connection information to your OpenSearch Dashboards is displayed in the service overview page. Use your browser to access the OpenSearch Dashboards service. ## Load sample data[​](#load-sample-data "Direct link to Load sample data") OpenSearch Dashboards come with three demonstration datasets included. To add sample data: 1. On the OpenSearch Dashboards landing page click **Add sample data**. 2. Choose one of the available datasets and click **Add data**. 3. Click **View data** to open the dashboard. ## Tools and pages[​](#tools-and-pages "Direct link to Tools and pages") OpenSearch Dashboards have many tools and features for working with data and running queries. ### Discover[​](#discover "Direct link to Discover") **Discover** page provides an interface to work with available data fields and run search queries by using either [the OpenSearch Dashboards Query Language (DQL)](https://opensearch.org/docs/latest/dashboards/dql/) or [Apache Lucene®](https://lucene.apache.org/). Additionally to search queries, you can filter the data by using either a visual interface or [OpenSearch Query DSL](https://opensearch.org/docs/latest/opensearch/query-dsl/index/) tip If the index you're looking at contains a date field, pay attention to the currently selected date range when running a query. ### Visualize[​](#visualize "Direct link to Visualize") **Visualize** page is an interface to create and manage your visualisations. In order to create a visualization: 1. Select visualization type to use. 2. Choose the source of data. 3. Follow the interface to set up metrics and buckets. ### Dashboard[​](#dashboard "Direct link to Dashboard") A set of visualization can be put together on a single dashboard. Search queries and filters applied to the dashboard will refine results for every included visualisation. ### Dev tools console[​](#dev-tools-console "Direct link to Dev tools console") Read how you can use **Dev Tools** to run the queries directly from OpenSearch Dashboards [in a separate article](/docs/products/opensearch/dashboards/howto/dev-tools-usage-example.md) . ### Query Workbench[​](#query-workbench "Direct link to Query Workbench") Query Workbench allows you to use SQL syntax instead of DSL to query the data. For example, you can retrieve the items we just added to the shopping list with: ``` select * from shopping-list ``` Find more on how to work [with SQL Workbench](https://opensearch.org/docs/latest/search-plugins/sql/workbench/) and [how to run SQL queries](https://opensearch.org/docs/latest/search-plugins/sql/index/) in the official documentation. *Apache Lucene is a registered trademark or trademark of the Apache Software Foundation in the United States and/or other countries* --- # Run queries from OpenSearch Dashboards Dev Tools Similarly to how you can work with the OpenSearch® service [using cURL](/docs/products/opensearch/howto/opensearch-with-curl.md) you can run the queries directly from OpenSearch Dashboards **Dev Tools**. The console contains both a request editor and a command output window. To get you started, see the various examples of requests that are included below. You can see that these are same requests as in [on how to use cURL](/docs/products/opensearch/howto/opensearch-with-curl.md), but in a pure [DSL query form](https://opensearch.org/docs/latest/opensearch/query-dsl/index/). This interface is helpful to populate index with sample data, or run a test query. Use **POST** method to add a new item to an index called `shopping-list`. If the index doesn't exist yet, it will be created: ``` POST shopping-list/_doc { "item": "apple", "quantity": 2 } ``` Follow by adding another item with a different set of fields: ``` POST shopping-list/_doc { "item": "bucket", "color": "blue", "quantity": 5, "notes": "the one with the metal handle" } ``` You can see the command output after running the query describing created item: ``` { "_index" : "shopping-list", "_type" : "_doc", "_id" : "jKTeWH0BMffxqtFft8Zv", "_version" : 1, "result" : "created", "_shards" : { "total" : 2, "successful" : 1, "failed" : 0 }, "_seq_no" : 1, "_primary_term" : 1 } ``` Use **GET** method to send a query to find items with an apple: ``` GET _search { "query": { "multi_match" : { "query" : "apple", "fields" : ["item", "notes"] } } } ``` In the output you can see the full response from OpenSearch engine: ``` { "took" : 35, "timed_out" : false, "_shards" : { "total" : 9, "successful" : 9, "skipped" : 0, "failed" : 0 }, "hits" : { "total" : { "value" : 1, "relation" : "eq" }, "max_score" : 0.6931471, "hits" : [ { "_index" : "shopping-list", "_type" : "_doc", "_id" : "i6TdWH0BMffxqtFf7cZv", "_score" : 0.6931471, "_source" : { "item" : "apple", "quantity" : 2 } } ] } } ``` Additionally, you can navigate through the history of queries and run them again. --- # Create alerts with OpenSearch® Dashboards Set up alerts in OpenSearch® Dashboards to send notifications when your data meets specific conditions. The OpenSearch alerting feature monitors data from one or more indexes and sends notifications when conditions are met. You can use alerts to monitor HTTP status codes, CPU load averages, or keyword counts in logs over specific intervals. Configure notifications to be sent through email, Slack, custom webhooks, or other channels. To configure an alert, you need the following: * [Notification channel](/docs/products/opensearch/dashboards/howto/opensearch-alerting-dashboard.md#create-a-notification-channel): a location for notifications to be delivered when an action is triggered * Available channel types are: `Amazon Chime`, `Amazon SNS`, `Slack`, `Custom webhook`, `Email`, or `Microsoft Teams`. * To use `Email`: * Ensure you have an SMTP server configured for a valid domain to deliver email notifications. * [Configure authentication for an email channel](/docs/products/opensearch/dashboards/howto/opensearch-alerting-dashboard.md#configure-authentication-for-an-email-channel) before configuring the email channel itself. * [Monitor](/docs/products/opensearch/dashboards/howto/opensearch-alerting-dashboard.md#create-a-monitor): a job that runs on a defined schedule and queries OpenSearch indexes Available frequency options are: `By interval`, `Daily`, `Weekly`, `Monthly`, or `Custom CRON expression`. * [Data source](/docs/products/opensearch/dashboards/howto/opensearch-alerting-dashboard.md#configure-a-data-source): OpenSearch indexes to query * [Query](/docs/products/opensearch/dashboards/howto/opensearch-alerting-dashboard.md#configure-a-query): the fields to query from indexes and the method for evaluating results * [Trigger](/docs/products/opensearch/dashboards/howto/opensearch-alerting-dashboard.md#create-a-trigger): a defined condition from the query results from the monitor. If a condition is met, the alert is generated. * [Action](/docs/products/opensearch/dashboards/howto/opensearch-alerting-dashboard.md#create-an-action): a notification configured to be sent through a specified channel when trigger conditions are met. You can define multiple actions. Sample requirements for an alert you can create in OpenSearch® Dashboards: * Checks `CPU load` * Uses the `sample-host-health` index as the data source * Uses `Slack` as the notification channel * Triggers when the average `cpu_usage_percentage` over `3 minutes` exceeds `75%` ## Create a notification channel[​](#create-a-notification-channel "Direct link to Create a notification channel") Configure your selected type of the notification channel, for example, [Slack](/docs/products/opensearch/dashboards/howto/opensearch-alerting-dashboard.md#configure-a-slack-channel) or [Email](/docs/products/opensearch/dashboards/howto/opensearch-alerting-dashboard.md#configure-authentication-for-an-email-channel). ### Configure a Slack channel[​](#configure-a-slack-channel "Direct link to Configure a Slack channel") 1. In OpenSearch Dashboards, go to **Notifications** > **Channels**. 2. Click **Create channel**. 3. Enter the following: 1. **Name**: `slack-test` 2. **Channel type**: `Slack` 3. **Slack webhook URL**: Paste your Slack webhook URL. 4. Click **Create**. ### Configure authentication for an email channel[​](#configure-authentication-for-an-email-channel "Direct link to Configure authentication for an email channel") To authenticate the sender account for sending email messages, add their credentials to the OpenSearch keystore: 1. Go to the [Aiven Console](https://console.aiven.io). 1. On the **Service settings** page of your Aiven for OpenSearch® service, go to **Advanced configuration**. 2. Click **Configure** > **Add configuration options**. 3. Add all three of the following configuration options and provide the corresponding details for each field: * `email_sender_name` * `email_sender_username` * `email_sender_password` note Configure all three parameters together. You cannot set them individually or save the configuration with only some of them set. 4. Click **Save configuration**. 2. Go to OpenSearch Dashboards. 1. Go to **Notifications** > **Channels**. 2. Click **Create channel**. 3. Enter the following: 1. **Name**: `email-test` 2. **Channel type**: `Email` 4. Configure a sender: 1. **Sender type**: Select `SMTP sender`. 2. Select an SMTP sender. If no SMTP sender exists, create one: 1. Enter a sender name matching the `email_sender_name` property from the keystore configuration. 2. Click **Create SMTP sender**. 3. Enter the sender details, select **Encryption method** `SSL/TLS`, and click **Create**. 5. Configure default recipients: Select default recipients. If no default recipients exist, create a recipient group: 1. Click **Create recipient group**. 2. Enter the recipient group details, and click **Create**. 6. Click **Create** to save the new channel configuration. ## Access **Alerting** in OpenSearch Dashboards[​](#access-alerting-in-opensearch-dashboards "Direct link to access-alerting-in-opensearch-dashboards") 1. Log in to the [Aiven Console](https://console.aiven.io) and go to your Aiven for OpenSearch service. 2. On the service's **Overview** page, in the **Connection information** section, go to the **OpenSearch Dashboards** tab. 3. Open OpenSearch Dashboards by clicking **Service URI** and logging in. 4. In OpenSearch Dashboards, go to **Alerting**. ## Create a monitor[​](#create-a-monitor "Direct link to Create a monitor") In OpenSearch Dashboards, go to **Alerting** > **Monitors** > **Create monitor**. ### Configure monitor details[​](#configure-monitor-details "Direct link to Configure monitor details") In the **Monitor details** section: 1. **Monitor name**: Enter `High CPU Monitor`. 2. **Monitor type**: Select `Per query monitor` (selected by default). 3. **Monitor defining method**: Select `Visual editor`. 4. **Frequency**: Select `By interval`. 5. **Run every**: Select `1 Minute(s)`. ### Configure a data source[​](#configure-a-data-source "Direct link to Configure a data source") In the **Select data** section, configure a data source: 1. Enter `sample-host-health` as **Indexes**. 2. Enter `timestamp` as **Time field**. ### Configure a query[​](#configure-a-query "Direct link to Configure a query") In the **Query** section, configure a query: 1. Click **Add metric**. 2. **Aggregation**: Select `average()`. 3. **Field**: Select `cpu_usage_percentage`. 4. Click **Save**. 5. **Time range for the last**: Enter `3 minute(s)`. ### Create a trigger[​](#create-a-trigger "Direct link to Create a trigger") In the **Triggers** section, create a trigger: 1. Click **Add trigger**. 2. **Trigger name**: Enter `high_cpu`. 3. **Severity level**: Select `1 (Highest)`. 4. **Trigger condition**: Select `IS ABOVE` and enter `75`. note You can see a visual graph for the trigger with the index data and the defined trigger condition as a red line. ### Create an action[​](#create-an-action "Direct link to Create an action") In the **Triggers** section, configure **Actions** for your trigger. * To use an existing notification channel for your action: 1. **Action name**: Enter `slack`. 2. Select your notification channel. 3. **Message subject**: Enter `High CPU Test Alert`. 4. Enter the message body. * To use a new notification channel for your action: 1. Click either **Manage channels** or **Create channels**, depending on whether you already have notification channels. 2. [Create a channel](/docs/products/opensearch/dashboards/howto/opensearch-alerting-dashboard.md#configure-a-slack-channel). 3. Return to configuring your action: Go to **Alerting** > **Monitors** > **Create monitor** > **Triggers** > **Actions**. 4. **Action name**: Enter `slack`. 5. Select your new notification channel. 6. **Message subject**: Enter `High CPU Test Alert`. 7. Enter the message body. tip Verify your action configuration by using **Preview message** and **Send test message**. Click **Create** to finalize your monitor setup. Related pages * [Alerting monitors configuration](https://opensearch.org/docs/latest/monitoring-plugins/alerting/monitors/) * [Notifications plugin](https://opensearch.org/docs/latest/observing-your-data/notifications/index/) --- # Get started with Aiven for OpenSearch® Learn how to use Aiven for OpenSearch®, create a service, secure access, manage indices, and explore your data. Aiven for OpenSearch® is a fully managed OpenSearch service designed for reliability, scalability, and security. It includes OpenSearch Dashboards for data visualization and supports integrations for logs and monitoring. tip Use an AI assistant connected to [Aiven MCP](/docs/tools/mcp-server.md) to create OpenSearch services, view cluster details, manage indexes, and search data from clients such as Cursor and Claude Code. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Ensure you have the following before getting started: * Console * API * CLI * Terraform - Access to the [Aiven Console](https://console.aiven.io) * [A personal token](https://docs.aiven.io/docs/platform/howto/create_authentication_token.html) - [Aiven CLI](https://github.com/aiven/aiven-client#installation) installed - [A personal token](https://docs.aiven.io/docs/platform/howto/create_authentication_token.html) * [Terraform installed](https://www.terraform.io/downloads) * A [personal token](/docs/platform/howto/create_authentication_token.md) ## Create an Aiven for OpenSearch® service[​](#create-an-aiven-for-opensearch-service "Direct link to Create an Aiven for OpenSearch® service") * Console * API * CLI * Terraform 1. In your project, click **Services**. 2. Click **Create service**. 3. Select **OpenSearch®**. 4. Select a **Cloud**. 5. Select a **Plan**. note The plans available can vary between cloud providers and regions for the same service. 6. In the **Service details**, enter a name for your service. 7. Optional: Add service tags. 8. In the **Service summary**, click **Create service**. The status of the service is **Rebuilding** during its creation. When the status is **Running**, you can start using the service. This typically takes a couple of minutes and can vary between cloud providers and regions. Create the service using the Aiven API, run: ``` curl -X POST https://api.aiven.io/v1/project//service \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "cloud": "google-europe-west1", "plan": "startup-4", "service_name": "example-opensearch", "service_type": "opensearch" }' ``` Parameters: * ``: Your project name. * ``: Your [API token](/docs/platform/howto/create_authentication_token.md). Create the service using the Aiven CLI, run: ``` avn service create \ --service-type opensearch \ --cloud \ --plan ``` Parameters: * ``: Name of your service (for example, `my-opensearch`). * ``: Deployment region (for example, `google-europe-west1`). * ``: Subscription plan (for example, `startup-4`). The following example files are also available in the [Aiven Terraform Provider repository](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/opensearch) on GitHub. 1. Create a file named `provider.tf` and add the following: ``` Loading... ``` 2. Create a file named `service.tf` and add the following: ``` Loading... ``` 3. Create a file named `variables.tf` and add the following: ``` Loading... ``` 4. Create the `terraform.tfvars` file and add the values for your token and project name. To apply your Terraform configuration: 1. Initialize Terraform by running: ``` terraform init ``` The output is similar to the following: ``` Initializing the backend... Initializing provider plugins... - Finding aiven/aiven versions matching ">= 4.0.0, < 5.0.0"... - Installing aiven/aiven v4.9.2... - Installed aiven/aiven v4.9.2 ... Terraform has been successfully initialized! ... ``` 2. To create an execution plan and preview the changes, run: ``` terraform plan ``` 3. To deploy your changes, run: ``` terraform apply --auto-approve ``` ## Secure access with ACLs[​](#secure-access-with-acls "Direct link to Secure access with ACLs") Secure your service by using one of the following options: * **Access control lists (ACLs)**: Manage access to indices by setting patterns (for example, `logs-*`) and permissions (read, write, or all) in the Aiven Console. * **OpenSearch security**: Use OpenSearch Dashboards or APIs for fine-grained access control, including role-based access control (RBAC) and single sign-on (SSO). For more information, see [Access control in Aiven for OpenSearch®](https://aiven.io/docs/products/opensearch/concepts/access_control). ## Manage indices[​](#manage-indices "Direct link to Manage indices") Aiven for OpenSearch® lets you view and manage indices and configure index retention patterns. For detailed steps on creating and managing indices, see the [OpenSearch documentation](https://opensearch.org/docs/latest/opensearch/index-data/). ### View and manage indices[​](#view-and-manage-indices "Direct link to View and manage indices") 1. Open your service in the [Aiven Console](https://console.aiven.io/). 2. In the **Data** section, click **Indexes** to view details such as shards, replicas, size, and health. ### Configure retention patterns[​](#configure-retention-patterns "Direct link to Configure retention patterns") 1. On the **Indexes** page, scroll to **Index retention patterns** 2. Click **Add pattern** and define: * **Pattern**: Specify index patterns (for example, `*_logs_*`). * **Maximum index count**: Set the number of indices to retain. 3. Click **Create** to save. For advanced indexing features, including custom mappings, refer to the [OpenSearch documentation](https://opensearch.org/docs/latest/opensearch/index-data/). ## Access OpenSearch Dashboards[​](#access-opensearch-dashboards "Direct link to Access OpenSearch Dashboards") Use OpenSearch Dashboards to visualize and analyze your data. 1. On the **Overview** page service in the [Aiven Console](https://console.aiven.io/). 2. In the **Connection information** section, click the **OpenSearch Dashboards** tab. 3. Copy or click the **Service URI** to open OpenSearch Dashboards in your browser. 4. Log in with the credentials provided in the **Connection information** section. For more information, see [OpenSearch Dashboards](https://opensearch.org/docs/latest/dashboards/). ## Connect to your service[​](#connect-to-your-service "Direct link to Connect to your service") * Console * cURL * Node.js * Python 1. Go to the **Overview** page of your service in the [Aiven Console](https://console.aiven.io/). 2. Click **Quick connect**. 3. In the **Connect** window, select a **dashboard** or **language** to connect to your service. 4. Complete the actions in the window and click **Done**. See [Use Aiven for OpenSearch® with cURL](https://aiven.io/docs/products/opensearch/howto/opensearch-with-curl) for steps to connect using cURL. See [Connect to OpenSearch® with Node.js](https://aiven.io/docs/products/opensearch/howto/connect-with-nodejs) to connect using Node.js. See [Connect to OpenSearch® with Python](https://aiven.io/docs/products/opensearch/howto/connect-with-python) to connect using Python. ## Manage logs and monitor data[​](#manage-logs-and-monitor-data "Direct link to Manage logs and monitor data") * **Send logs**: To send logs from Aiven services to OpenSearch, see [Enable log integration](https://aiven.io/docs/products/opensearch/howto/opensearch-log-integration). * **Monitor data**: Set up Grafana for monitoring and alerts. See [Integrate with Grafana®](https://aiven.io/docs/products/opensearch/howto/integrate-with-grafana). ## Search and aggregations with Aiven for OpenSearch[​](#search-and-aggregations-with-aiven-for-opensearch "Direct link to Search and aggregations with Aiven for OpenSearch") Aiven for OpenSearch® lets you write and execute search queries, as well as aggregate data using OpenSearch clients like Python and Node.js. ### Write search queries[​](#write-search-queries "Direct link to Write search queries") * **Python**: Learn how to write and run search queries on your Aiven for OpenSearch service using the [Python OpenSearch client](https://github.com/opensearch-project/opensearch-py). For more information, see [Search with OpenSearch® and Python](/docs/products/opensearch/howto/opensearch-search-and-python.md). * **Node.js**: Learn how to write and run search queries on your Aiven for OpenSearch service using the [OpenSearch JavaScript client](https://github.com/opensearch-project/opensearch-js). For more information, see [Search with OpenSearch® and Node.js](/docs/products/opensearch/howto/opensearch-and-nodejs.md). ### Perform aggregations[​](#perform-aggregations "Direct link to Perform aggregations") * **Metric aggregations**: Calculate metrics such as average, minimum, maximum, percentiles, and cardinality on your Aiven for OpenSearch service. or more information, see [Aggregations with OpenSearch® and Node.js](/docs/products/opensearch/howto/opensearch-aggregations-and-nodejs.md#metrics-aggregations). * **Bucket aggregations**: Group data into buckets based on ranges, unique terms, or histograms. or more information, see [Bucket aggregations with OpenSearch®](/docs/products/opensearch/howto/opensearch-aggregations-and-nodejs.md#bucket-aggregations). * **Pipeline aggregations**: Combine results from multiple aggregations, such as calculating moving averages, to analyze trends. Explore examples in [Pipeline aggregations with OpenSearch®](/docs/products/opensearch/howto/opensearch-aggregations-and-nodejs.md#pipeline-aggregations). --- # Enable and manage OpenSearch® Audit logs Aiven for OpenSearch® enables audit logging functionality via the OpenSearch Security dashboard, which allows OpenSearch Security administrators to track system events, security-related events, and user activity. These audit logs contain information about user actions, such as login attempts, API calls, index operations, and other security-related events. Learn how to enable, configure, and visualize OpenSearch audit logs through the OpenSearch Security dashboard. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven for OpenSearch® service * [OpenSearch Security management enabled](/docs/products/opensearch/howto/enable-opensearch-security.md) for the Aiven for OpenSearch service After enabling audit logs in the service's advanced configuration, proceed to enable them in the OpenSearch® Security dashboard. ## Enabling audit logs in Aiven for OpenSearch[​](#enabling-audit-logs-in-aiven-for-opensearch "Direct link to Enabling audit logs in Aiven for OpenSearch") By default, audit logs are disabled in Aiven for OpenSearch. To enable them: 1. Access your Aiven for OpenSearch service in the Aiven Console. 2. From the left sidebar, click **Service settings** 3. Scroll to the **Advanced configuration** and click **Configure**. 4. In the **Advanced configuration** dialog, click **Add configuration to options**. 5. Use the search function to locate the `enable_security_audit` configuration and switch it to the **Enabled** position. 6. Click **Save configuration** to save your changes and enable audit logging. After enabling audit logs in the service's advanced configuration, proceed to enable them in the OpenSearch® Security dashboard. ## Enable audit logs in OpenSearch® Security dashboard[​](#enable-audit-logs-in-opensearch-security-dashboard "Direct link to Enable audit logs in OpenSearch® Security dashboard") To enable audit logs in OpenSearch® Security dashboard, follow these steps: 1. Log in to the OpenSearch® Dashboard using OpenSearch® Security admin credentials. 2. Select **Security** from the left-side menu. 3. Select **Audit logs**. 4. Toggle the switch next to **Enable audit logging** to the on position. note The storage location for audit logs on your Aiven for OpenSearch service is configured to be stored on the cluster's storage and cannot be changed. ## Audit log event types[​](#audit-log-event-types "Direct link to Audit log event types") OpenSearch® enables audit logging for HTTP requests(REST) and the transport layer, capturing various events related to user authentication, privileges, security, and more. The following are the types of audit events recorded by OpenSearch: * `FAILED_LOGIN`: User authentication failure. * `AUTHENTICATED`: User authentication success. * `MISSING_PRIVILEGES`: User does not have request privileges. * `GRANTED_PRIVILEGES`: User's privileges successfully granted. * `SSL_EXCEPTION`: Invalid SSL/TLS certificate in request. * `opensearch_SECURITY_INDEX_ATTEMPT`: Unauthorized security plugin modification attempt. * `BAD_HEADERS`: Attempted request spoofing with internal security headers. ## Configure audit logging[​](#configure-audit-logging "Direct link to Configure audit logging") Customize the audit logging settings in OpenSearch® Security to align with your organization's specific requirements. The configuration process involves two primary sections: General and Compliance settings, each offering distinct options: * **General settings:** Adjust logging for REST and Transport layers, tailor log data for each event, and set preferences to selectively exclude specific users or requests, ensuring log relevance and operational efficiency. * **Compliance settings:** Enable compliance mode to meet regulatory standards and activate tamper-evident logging. You can also enable logging for internal and external configuration changes as well as metadata logging options for a robust security posture. note * You cannot modify the name of the audit log index for your Aiven for OpenSearch service as it is set to the default name `auditlog-YYYY.MM.dd`. * You cannot change the size of the thread pool using `OpenSearch.yml`. ### Optimize audit log configuration[​](#optimize-audit-log-configuration "Direct link to Optimize audit log configuration") Optimize audit log configuration in OpenSearch Security to protect sensitive data from unauthorized access. * Exclude irrelevant event categories. * Consider disabling logging for Rest and Transport layers. * Disable request body logging for sensitive information. * Regularly review and maintain audit log configuration. ## Visualize audit log[​](#visualize-audit-log "Direct link to Visualize audit log") Visualizing audit logs is an effective way to understand the extensive data generated by these logs. Visualization can help identify patterns or anomalies that may indicate security risks or system issues by presenting the information in user-friendly graphical formats. To access and visualize audit logs in OpenSearch: 1. **Create an index pattern**: In the **Stack Management** section, establish an index pattern to organize your log data. 2. **Create a visualization**: Go to the **Visualize** section, where you can design and tailor visualizations to suit your analysis needs. 3. **Save and modify visualization**: Once you've created a visualization, save it for future reference. You can always return to modify and update it as your requirements evolve Related pages * [OpenSearch audit logs documentation](https://opensearch.org/docs/latest/security/audit-logs/index/) --- # Back up your Aiven for OpenSearch® service to another region Copy your Aiven for OpenSearch® service backups to a secondary region for disaster recovery. In addition to the primary service backup, you can have a secondary backup in an alternative location. important This feature is in [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-). Contact your account team to enable it. Backup to another region (BTAR) is a disaster recovery feature that allows backup files to be copied from the service's primary backup region to an additional (secondary) region. BTAR can bolster data resilience and helps improve data protection against disasters in the primary backup region. When the primary region is down, BTAR allows forking the service from an additional copy of the backup residing in a secondary region. ## Limitations[​](#limitations "Direct link to Limitations") * The cloud provider for your additional backup region must match the cloud provider for your service and the primary backup. * Secondary backup can only be restored in the region where it was stored. For a service that has the backup to another region (BTAR) feature enabled, you can check the service backup status, change the backup region, monitor the replication lag, fork and restore using the cross-region backup, or migrate to another cloud or region. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * At least one Aiven service with BTAR enabled * Access to the [Aiven Console](https://console.aiven.io/) * [Aiven API](/docs/tools/api.md) * [Aiven CLI](/docs/tools/cli.md) ## Change a backup region[​](#change-a-backup-region "Direct link to Change a backup region") 1. Log in to the [Aiven Console](https://console.aiven.io/) and go to your project. 2. On the **Services** page, select an Aiven service on which you'd like to enable BTAR. 3. On your service page, in the **Backups** section, click **Backup management**. 4. On the **Backups** page, click **Actions** > **Edit secondary backup location**. 5. In the **Edit secondary backup location** window, use the **Backup location** menu to select a region for your additional backup. Confirm your choice by selecting **Save**. important You can change the backup region once in 24 hours. ## Monitor a service with BTAR[​](#monitor-a-service-with-btar "Direct link to Monitor a service with BTAR") There are a few things you may want to check for your Aiven service in the context of BTAR: * What is the status of a secondary backup? * Does your service have a backup in another region? * What is the target region of the secondary backup? * What is the replication lag between data availability in the primary region and the secondary region? ### Check BTAR status[​](#check-btar-status "Direct link to Check BTAR status") To see the availability, the status, and the target region of a secondary (BTAR) backup in the [Aiven Console](https://console.aiven.io/), go to your service page > **Backups** > **Secondary backup location**. ### Determine replication lag[​](#determine-replication-lag "Direct link to Determine replication lag") Determine the target region and the replication lag for a secondary (BTAR) backup of your service, call the [ServiceBackupToAnotherRegionReport](https://api.aiven.io/doc/#tag/Service/operation/ServiceBackupToAnotherRegionReport) endpoint. Configure the call as follows: 1. Enter `YOUR-PROJECT-NAME` and `YOUR-SERVICE-NAME` into the URL. 2. Specify `DESIRED-TIME-PERIOD` depending on the time period you need the metrics for: select one of the following values for the `period` key: `hour`, `day`, `week`, `month`, or `year`. ``` curl --request POST \ --url https://api.aiven.io/v1/project/YOUR-PROJECT-NAME/service/YOUR-SERVICE-NAME/backup_to_another_region/report \ --header 'Authorization: Bearer YOUR-BEARER-TOKEN' \ --header 'content-type: application/json' \ --data '{"period":"DESIRED-TIME-PERIOD"}' ``` As output, you get metrics including replication lags at specific points in time. ## Fork and restore a service with BTAR[​](#fork-and-restore "Direct link to Fork and restore a service with BTAR") You can use the [Aiven Console](/docs/tools/aiven-console.md), [API](/docs/tools/api.md), or [CLI](/docs/tools/cli.md) to recover your service from a backup in another region. To restore your service using BTAR, create a fork of the original service in the region where the secondary backup resides. note When you **fork & restore** from the secondary backup, your new fork service is created in the cloud and region where the secondary backup is located. The fork service gets the same plan that the primary service uses. Backups of the fork service are located in the region where this new service is hosted. * Aiven Console * Aiven CLI * Aiven API 1. Open the [Aiven Console](https://console.aiven.io/) and go to your service homepage. 2. Click **Backups**. 3. On the **Backups** page, select **Fork & restore**. 4. In the **New database fork** window: 1. Set **Backup location** to either **Primary location** or **Secondary location**. 2. Set **Backup version** to one of the following: * **Latest transaction** * **Point in time**: Set it up to no earlier than the time of taking the oldest replicated base backup. 3. Specify a name for the new fork service. 4. Select **Create fork**. Run the [avn service create](/docs/tools/cli/service-cli.md#avn-cli-service-create) command with the `--service-to-fork-from` option and the `--recovery-target-time`option. Set `--recovery-target-time` to no earlier than the time of taking the oldest replicated base backup. ``` avn service create FORK_SERVICE_NAME \ --plan SERVICE_PLAN \ --project PROJECT_NAME \ --service-type SERVICE_TYPE \ --cloud SECONDARY_BACKUP_REGION \ --recovery-target-time "YYYY-MM-DDTHH:MM:SS+00:00" \ --service-to-fork-from PRIMARY_SERVICE_NAME ``` Replace the following with meaningful data: * `FORK_SERVICE_NAME` * `SERVICE_PLAN` * `PROJECT_NAME` * `SERVICE_TYPE` * `SECONDARY_BACKUP_REGION` * `PRIMARY_SERVICE_NAME` * `YYYY-MM-DDTHH:MM:SS+00:00` Use the [ServiceCreate](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) API to create a fork service. When constructing the API request, add the `user_config` object to the request body and nest the `service_to_fork_from` field and the `recovery_target_time` field inside. Set `recovery_target_time` to no earlier than the time of taking the oldest replicated base backup. ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "cloud": "SECONDARY_BACKUP_REGION", "plan": "SERVICE_PLAN", "service_name": "FORK_SERVICE_NAME", "service_type": "SERVICE_TYPE", "user_config": { "service_to_fork_from": "PRIMARY_SERVICE_NAME", "recovery_target_time": "YYYY-MM-DDTHH:MM:SS+00:00" } }' ``` Replace the following with meaningful data: * `FORK_SERVICE_NAME` * `SERVICE_PLAN` * `PROJECT_NAME` * `SERVICE_TYPE` * `SECONDARY_BACKUP_REGION` * `PRIMARY_SERVICE_NAME` * `YYYY-MM-DDTHH:MM:SS+00:00` ## Migrate a service with BTAR[​](#migrate-a-service-with-btar "Direct link to Migrate a service with BTAR") You can migrate a service with BTAR the same way you [migrate a service with a regular backup](/docs/platform/howto/migrate-services-cloud-region.md). note When you migrate your service, locations of service backups, both primary and secondary ones, do not change. ## Delete a cross-region backup[​](#delete-a-cross-region-backup "Direct link to Delete a cross-region backup") Delete an additional service backup created in a region different from your primary backup region. You can delete a cross-region backup using the [Aiven Console](/docs/tools/aiven-console.md), [API](/docs/tools/api.md), or [CLI](/docs/tools/cli.md). When you delete the additional cross-region backup, you still have the default backup located in the primary, service-hosting region. * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. From the **Services** view, select an Aiven service on which you'd like to disable BTAR. 3. On your service's page, click **Backups**. 4. On the **Backups** page, click **Actions** > **Secondary backup location**. 5. In the **Edit secondary backup location** window, select **Disable**. Your additional service backup is no longer visible on your service's **Backups** page in the **Secondary backup location** column. To remove secondary backups for your service, use the [avn service update](/docs/tools/cli/service-cli.md) command to remove all target region names from the `additional_backup_regions` array. ``` avn service update your-sevice-name \ -c additional_backup_regions=\[\] ``` To remove secondary backups for your service, update the service configuration. Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint to remove all target regions names from the `additional_backup_regions` array. ``` curl --request PUT \ --url https://api.aiven.io/v1/project/YOUR_PROJECT_NAME/service/YOUR_SERVICE_NAME \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "additional_backup_regions": [] } }' ``` Related pages * [Backups](/docs/products/opensearch/howto/restore_opensearch_backup.md) * [Track restore progress](/docs/products/opensearch/howto/track-restore-progress.md) --- # Change the cloud or region for your Aiven for OpenSearch® service Move your Aiven for OpenSearch® service to a different cloud provider or region. 1. In your service, click **Service settings** from the sidebar. 2. In the **Cloud and network** section, click **Actions** > **Change cloud**. 3. In the **Cloud** section , select a cloud provider and region, and click **Change**. Your service starts a migration to the new location and remains available during the process. When the migration completes, the service continues running in the new cloud or region. Related pages * [Fork your service](/docs/products/opensearch/howto/fork-service.md) * [Migrate to another cloud or region](/docs/platform/howto/migrate-services-cloud-region.md) --- # Change the plan for your Aiven for OpenSearch® service Change the service plan for your Aiven for OpenSearch® service to scale resources up or down and optimize costs. Adjust the plan of your services at any time to scale your services as needed and optimize costs. If you can't find a suitable plan, you can [request a custom plan](/docs/platform/concepts/service-pricing.md). tip If you plan to upgrade your service plan, do it immediately after a full backup. This reduces the amount of incremental changes that need to be applied on top of the base backup, which speeds up the upgrade itself. important * When changing a service plan, reserve an additional 25% of disk space. This requirement applies to upgrades and downgrades. * Downgrading to a plan with fewer VMs is supported for most services, including Aiven for Apache Kafka®, Aiven for PostgreSQL®, Aiven for OpenSearch®, Aiven for ClickHouse®, Aiven for MySQL®, Aiven for Metrics, and Aiven for Valkey™. * Changing a service plan triggers a node recycle, service rebuilding, and any pending maintenance updates. - Console - Terraform - CLI 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Change plan**. 3. Select a plan that provides at least 125% of the current disk size and click **Change plan**. Update the `plan` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). To change a service plan in the Aiven CLI, use the [`avn service update --plan `](/docs/tools/cli/service-cli.md#avn-cli-service-update) command. Your service's state becomes **Rebuilding** and remains accessible. When the state switches to **Running**, your new service plan is active. Related pages * [Scale disk storage](/docs/products/opensearch/howto/scale-disk-storage.md) * [Prepare for high load](/docs/products/opensearch/howto/prepare-for-high-load.md) --- # Connect to Aiven for OpenSearch® with NodeJS The most convenient way to work with the cluster when using NodeJS is to rely on [OpenSearch® JavaScript client](https://github.com/opensearch-project/opensearch-js). Follow its `README` file for installation instructions. To connect to the cluster, you'll need `service_uri`, which you can find either in the service overview in the [Aiven console](https://console.aiven.io) or get through the Aiven command line interface [service command](/docs/tools/cli/service-cli.md#avn_service_ca_get). `service_uri` contains credentials, therefore should be treated with care. We strongly recommend using environment variables for credential information. A good way to do this is to use `dotenv`. [See the official docs](https://github.com/motdotla/dotenv), create `.env` file in the project and assign `SERVICE_URI` inside of this file. Add the require line to the top of your file: ``` require("dotenv").config() ``` Now you can refer to the value of `service_uri` as `process.env.SERVICE_URI` in the code. Add the following lines of code to create a client and assign `process.env.SERVICE_URI` to the `node` property. This will be sufficient to connect to the cluster, because `service_uri` already contains credentials. Additionally, when creating a client you can also specify `ssl configuration`, `bearer token`, `CA fingerprint` and other authentication details depending on protocols you use. ``` const { Client } = require('@opensearch-project/opensearch') module.exports.client = new Client({ node: process.env.SERVICE_URI, }); ``` The client will perform request operations on your behalf and return the response in a consistent manner. --- # Connect to Aiven for OpenSearch® with Python You can interact with your cluster with the help of the [Python OpenSearch® client](https://github.com/opensearch-project/opensearch-py). It provides a convenient syntax to send commands to your OpenSearch cluster. Follow its `README` file for installation instructions. To connect with your cluster, you need the **Service URI** of your OpenSearch cluster. Find the connection details in the section **Overview** on [Aiven Console](https://console.aiven.io). Alternatively, you can retrieve it via the `avn service get` command with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). Notice that `service_uri` contains credentials; therefore, should be treated with care. The **Service URI** has information in the following format: ``` [](https://:@::@::@: Manage users and permissions in Aiven for OpenSearch by creating Access Control Lists (ACLs) in the Aiven Console. Service users only exist in the scope of the Aiven service. They are unique to the service and not shared with any other services. Every service has a default `avnadmin` user with full access to the service. Use the **Users** tab in the [Aiven Console](https://console.aiven.io) to manage access control and permissions for your Aiven for OpenSearch service. You can create users, modify their details, and assign index patterns and permissions. Alternatively, enable [OpenSearch Security management](/docs/products/opensearch/howto/enable-opensearch-security.md) to manage users and permissions via the OpenSearch Security dashboard. note ACLs apply only to indices and do not control access to other OpenSearch APIs, including OpenSearch Dashboards. ## Limitations[​](#limitations "Direct link to Limitations") By default, there is a limit of 50 service users per service. To request a higher limit for a service, [create a support ticket](/docs/platform/howto/support.md). ## Create a user without access control[​](#create-a-user-without-access-control "Direct link to Create a user without access control") To create a service user without access control in Aiven for OpenSearch: 1. Open your Aiven for OpenSearch service in the [Aiven Console](https://console.aiven.io). 2. Click **Users** in the sidebar. 3. Click **Create users**. 4. Enter a username, and click **Save**. By default, newly created users are granted **full access rights**. To limit their access, you can enable access control and specify an Access Control List (ACL) that defines the relevant permissions and patterns. ### Create a service user with the CLI, API, or Terraform[​](#create-a-service-user-with-the-cli-api-or-terraform "Direct link to Create a service user with the CLI, API, or Terraform") * Aiven CLI * Aiven API * Terraform Run the [avn service user-create](/docs/tools/cli/service/user.md#avn-service-user-create) command: ``` avn service user-create SERVICE_NAME --username USERNAME ``` Replace the following: * `SERVICE_NAME`: the name of your Aiven for OpenSearch service. * `USERNAME`: the name of the service user to create. Use the [ServiceUserCreate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUserCreate) endpoint: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/user \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{"username": "USERNAME"}' ``` Replace the placeholders with your project name, service name, bearer token, and the username to create. Use the [`aiven_opensearch_user` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/opensearch_user) to create and manage service users. ## Enable access control[​](#enable-access-control "Direct link to Enable access control") To enable access control for the Aiven for OpenSearch service: 1. Open your Aiven for OpenSearch service in the [Aiven Console](https://console.aiven.io). 2. Click **Users** in the sidebar. 3. Toggle the **Access Control** switch to **Enabled** ## Create a user with access control[​](#create-a-user-with-access-control "Direct link to Create a user with access control") To create a service user with access control in Aiven for OpenSearch: 1. Open your Aiven for OpenSearch service in the [Aiven Console](https://console.aiven.io). 2. Click **Users** in the sidebar. 3. Click **Create user**. 4. In the **Create service user** screen, enter a **username**. 5. Specify an **Index pattern** and set the desired **permissions**. 6. To add multiple rules to the user, click **Add another rule**. 7. Click **Save**. note The password for service users is automatically generated and can be reset if necessary. After creating a new service user, log in to the [OpenSearch Dashboard](/docs/products/opensearch/dashboards.md) using the assigned credentials to access and perform actions according to the user’s permissions. ## Manage users[​](#manage-users "Direct link to Manage users") Aiven for OpenSearch provides several additional operations you can perform on service users. To access these operations: 1. Open your Aiven for OpenSearch service in the [Aiven Console](https://console.aiven.io). 2. Click **Users** in the sidebar. 3. Click **Actions** in the respective user row and choose the desired operation: * Click **Show password** to view the password. * Click **Reset password** to reset the password. * Click **Edit ACL rules **to edit ACL rules. * Click **Delete user** to delete the user. warning Deleting a service user immediately terminates all existing database sessions, and the user loses access. ## Disable access control[​](#disable-access-control "Direct link to Disable access control") Disabling access control for your Aiven for OpenSearch service grants admin access to all users and overrides any previously defined user permissions. Carefully consider the outcomes of disabling access control before proceeding. To disable access control: 1. Open your Aiven for OpenSearch service in the [Aiven Console](https://console.aiven.io). 2. Click **Users** in the sidebar. 3. Toggle the **Access Control** switch to **Disable**. 4. Click **Disable** to confirm. --- # Controlled upgrade pipelines for your Aiven for OpenSearch® service [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Link Aiven for OpenSearch® services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. Control when your Aiven managed services receive maintenance updates and test maintenance updates in development or staging environments before they reach production. important Controlled upgrade pipeline is a [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-) feature. [Contact Aiven](https://aiven.io/contact) to request access. Aiven performs automatic service maintenance for security fixes, minor software updates, and other platform changes. The controlled upgrade pipeline feature lets you link services of the same type in an ordered sequence to control when each service receives updates. After a maintenance update upgrades a service at the initial pipeline step, you validate that service version before the update proceeds to the service at the next pipeline step. Validating means approving the new version as safe to roll out to the next service. Validation can be manual or automatic after a configurable delay. ## Why use controlled upgrade pipelines[​](#why-use-controlled-upgrade-pipelines "Direct link to Why use controlled upgrade pipelines") Controlled upgrade pipelines prevent production incidents caused by automatic updates reaching production before teams can test the new version in a lower environment. They give you full oversight of the update process: * **Risk mitigation**: Prevents unexpected maintenance updates from breaking your production environment by ensuring they are tested in a non-production setting first. * **Stability**: Keeps destination services (such as production) on a known-good version until you, or the automatic timer, confirm the new version is safe. * **Process control**: Allows platform teams to standardize their deployment and maintenance lifecycle across environments. ## About controlled upgrade pipelines[​](#about-controlled-upgrade-pipelines "Direct link to About controlled upgrade pipelines") ### Upgrade steps[​](#upgrade-steps "Direct link to Upgrade steps") An upgrade step is a pair of services linked by an upgrade constraint: * **Source service**: The service that receives maintenance updates first * **Destination service**: The service that waits for validation before receiving updates Each destination service can have only one source service. A source service can have multiple destination services. ### Upgrade pipelines[​](#upgrade-pipelines "Direct link to Upgrade pipelines") An upgrade pipeline is a chain of upgrade steps that spans multiple environments. For example: * Single chain: development → staging → production * Multiple destinations: development → production-eu and development → production-na ## How validation works[​](#how-validation-works "Direct link to How validation works") When a maintenance update upgrades your source service: 1. The source service receives the update first. 2. Test the updated source service to verify it works as expected. 3. Validate the update manually using the API or CLI, or wait for automatic validation after the configured delay. The default delay is 7 days. 4. After validation, the destination service becomes eligible for the same maintenance update. 5. The destination service receives the update during its next maintenance window. If one source service has multiple destination services, one validation for the source service applies to all connected destination services. ### Validation and maintenance windows[​](#validation-and-maintenance-windows "Direct link to Validation and maintenance windows") Validation and the maintenance window control different things: * **Validation** controls *what* version the destination service upgrades to. * The maintenance window controls *when* the upgrade happens. After you validate an update, or automatic validation applies, the destination service receives the validated version during its next scheduled maintenance window. Validation does not trigger an immediate upgrade outside the maintenance window. Upgrade pipelines add a constraint on what is installed during a maintenance update; they do not change when maintenance runs. Nodes in the destination service maintain the validated version until a newer version is validated, either when you validate it manually or when automatic validation applies after the configured delay. When a node is recycled, it uses the same validated version, not the latest available version. When you create a step, the destination service keeps the newest version that is already validated at that moment. If the destination service is already applying maintenance during step creation, the in-progress target version becomes the initial validated version. warning A powered-off source service cannot receive maintenance updates, so you cannot validate it. If you power off services earlier in the chain, the destination service upgrades regardless. For example, in a development → staging → production chain, if both development and staging are powered off, production upgrades without testing and validation in the earlier environments. Keep services in the chain powered on to preserve the protection that upgrade pipelines provide. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * **Same service type**: You can only link services of the same type. For example, two Aiven for PostgreSQL services. * **Chain length**: The default maximum chain depth is 3 services, which is 2 steps. If you need a longer chain, [contact Aiven](https://aiven.io/contact). * **No cycles**: You cannot create circular dependencies between services. * **Emergency overrides**: Aiven can apply critical security or stability fixes to a destination service before explicit validation. * **Supported services**: This feature supports all Aiven service types except Aiven for Apache Flink® and Aiven for MySQL. * **Automatic maintenance updates only**: Pipelines apply to automatic maintenance updates, such as minor service version updates and node image updates. Major version upgrades, for example Aiven for PostgreSQL® 15 to 16, require manual action and are not promoted automatically through the pipeline. * **No permanent blocking**: You cannot prevent an update indefinitely. Automatic validation applies after the configured delay, up to the maximum delay. * **No validation rollback**: You cannot undo a validation after it is recorded. ## Use controlled upgrade pipelines[​](#use-controlled-upgrade-pipelines "Direct link to Use controlled upgrade pipelines") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") To use controlled upgrade pipelines, you need the following: * The feature enabled by Aiven ([Limited availability](/docs/platform/concepts/service-and-feature-releases.md)) * Dev tool of your choice: * [Aiven CLI](/docs/tools/cli.md) Install the latest version of the Aiven CLI to access the `upgrade-pipeline` commands. * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) * Set `PROVIDER_AIVEN_ENABLE_BETA=true` before running Terraform. * See the [resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) for full schema, import format, and lifecycle behavior. * [Aiven Operator for Kubernetes](/docs/tools/kubernetes.md) Install the operator and create an Aiven token secret named `aiven-token` that the operator uses to authenticate against the Aiven API. * Write access to the source and destination projects * At least two services of the same type (for example, two Aiven for PostgreSQL® services) * Services can be in different projects in the same organization ### Set up an upgrade pipeline[​](#set-up-an-upgrade-pipeline "Direct link to Set up an upgrade pipeline") Use the Aiven CLI or API to create upgrade steps between your services. note The `upgrade-pipeline` CLI commands require Aiven CLI version 4.x or later. Command names and parameters may change before general availability. #### Create an upgrade step[​](#create-an-upgrade-step "Direct link to Create an upgrade step") Create a step to link a source service and a destination service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ [--source-project SOURCE_PROJECT] SOURCE_SERVICE \ [--destination-project DESTINATION_PROJECT] DESTINATION_SERVICE \ [--auto-validation-delay-days DAYS] ``` **Options** * `--organization-id` is required. * `--source-project` and `--destination-project` are optional. If you omit either project option, Aiven CLI uses the current default project set with `avn project switch`. * `--auto-validation-delay-days` is optional. Defaults to 7 days if not specified. ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "SOURCE_PROJECT_NAME", "source_service_name": "SOURCE_SERVICE_NAME", "destination_project_name": "DESTINATION_PROJECT_NAME", "destination_service_name": "DESTINATION_SERVICE_NAME", "auto_validation_delay_days": 7 }' ``` Use the [`aiven_upgrade_step`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) resource: ``` resource "aiven_upgrade_step" "example" { organization_id = "ORGANIZATION_ID" source_project_name = "SOURCE_PROJECT_NAME" source_service_name = "SOURCE_SERVICE_NAME" destination_project_name = "DESTINATION_PROJECT_NAME" destination_service_name = "DESTINATION_SERVICE_NAME" auto_validation_delay_days = 7 } ``` Apply an `UpgradePipelineStep` manifest with `kubectl`: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: upgrade-step-sample spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: SOURCE_PROJECT_NAME sourceServiceName: SOURCE_SERVICE_NAME destinationProjectName: DESTINATION_PROJECT_NAME destinationServiceName: DESTINATION_SERVICE_NAME autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable after the resource is created. Parameters: * `source_project_name`: Name of the project containing the source service * `source_service_name`: Name of the source service * `destination_project_name`: Name of the project containing the destination service * `destination_service_name`: Name of the destination service * `auto_validation_delay_days`: Optional. Number of days before automatic validation. The value must be at least `1`. The default is 7 days. The maximum delay you can configure is 30 days. #### List upgrade steps[​](#list-upgrade-steps "Direct link to List upgrade steps") View all upgrade steps you have access to: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step list --organization-id ORGANIZATION_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) To list managed upgrade steps, use: ``` terraform state list 'aiven_upgrade_step.*' terraform state show 'aiven_upgrade_step.example' ``` List `UpgradePipelineStep` resources in the current namespace: ``` kubectl get upgradepipelinesteps ``` #### View a specific step[​](#view-a-specific-step "Direct link to View a specific step") Get details about a specific upgrade step, including the last validation: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step get \ --organization-id ORGANIZATION_ID \ STEP_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` terraform state show aiven_upgrade_step.example ``` Show the manifest and full status, including `id`, `conditions`, and `lastValidation`: ``` kubectl describe upgradepipelinestep RESOURCE_NAME kubectl get upgradepipelinestep RESOURCE_NAME -o yaml ``` The step details include `last_validation` values such as `validated_at`, `validated_by_user`, and `comment` when validation exists (available through the API). ### Validate an upgrade[​](#validate-an-upgrade "Direct link to Validate an upgrade") After testing your source service with the new update, validate the version to allow the destination service to receive the same update. #### Manual validation[​](#manual-validation "Direct link to Manual validation") Validate the current version of your source service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step validate-for-service \ --project SOURCE_PROJECT \ SERVICE_NAME \ [--comment "COMMENT"] ``` `--comment` is optional. Use it to record a note about the validation, for example `"Tested and verified in development"`. ``` curl -X POST https://api.aiven.io/v1/project/SOURCE_PROJECT/service/SOURCE_SERVICE/upgrade-validation \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "comment": "Tested and verified in development" }' ``` Terraform manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. The operator manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. #### Automatic validation[​](#automatic-validation "Direct link to Automatic validation") If you do not manually validate an update, the system automatically validates the source service version after the configured delay. Auto-validation starts from when the source service receives the update. ### Manage upgrade steps[​](#manage-upgrade-steps "Direct link to Manage upgrade steps") #### Update a step[​](#update-a-step "Direct link to Update a step") Modify the automatic validation delay for an existing step: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step update \ --organization-id ORGANIZATION_ID \ --auto-validation-delay-days 14 \ STEP_ID ``` ``` curl -X PATCH https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "auto_validation_delay_days": 14 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` resource "aiven_upgrade_step" "example" { # ...required fields... auto_validation_delay_days = 14 # Updated from 7 to 14 } ``` Apply the changes: ``` terraform plan terraform apply ``` Edit `autoValidationDelayDays` in your manifest and re-apply: ``` spec: autoValidationDelayDays: 14 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable. To change them, delete the resource and create a new one. #### Delete a step[​](#delete-a-step "Direct link to Delete a step") Remove an upgrade step to allow the destination service to receive updates independently: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step delete --organization-id ORGANIZATION_ID STEP_ID ``` Find `STEP_ID` from the upgrade step list command. ``` curl -X DELETE https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) Remove the resource from configuration and apply, or destroy it directly: ``` terraform apply terraform destroy -target=aiven_upgrade_step.example ``` Delete the `UpgradePipelineStep` resource: ``` kubectl delete upgradepipelinestep RESOURCE_NAME ``` Deleting a step removes all associated validations. ### Example: Three-environment pipeline[​](#example-three-environment-pipeline "Direct link to Example: Three-environment pipeline") Create a pipeline that promotes updates from development to staging to production: * CLI * API * Terraform * Kubernetes 1. Create a step from development to staging: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project dev-project \ --destination-project staging-project \ --auto-validation-delay-days 3 \ pg-dev pg-staging ``` 2. Create a step from staging to production: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project staging-project \ --destination-project prod-project \ --auto-validation-delay-days 7 \ pg-staging pg-prod ``` 1) Create a step from development to staging: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "dev-project", "source_service_name": "pg-dev", "destination_project_name": "staging-project", "destination_service_name": "pg-staging", "auto_validation_delay_days": 3 }' ``` 2) Create a step from staging to production: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "staging-project", "source_service_name": "pg-staging", "destination_project_name": "prod-project", "destination_service_name": "pg-prod", "auto_validation_delay_days": 7 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` # Step 1: Development → Staging resource "aiven_upgrade_step" "dev_to_staging" { organization_id = "ORGANIZATION_ID" source_project_name = "dev-project" source_service_name = "pg-dev" destination_project_name = "staging-project" destination_service_name = "pg-staging" auto_validation_delay_days = 3 } # Step 2: Staging → Production resource "aiven_upgrade_step" "staging_to_prod" { organization_id = "ORGANIZATION_ID" source_project_name = "staging-project" source_service_name = "pg-staging" destination_project_name = "prod-project" destination_service_name = "pg-prod" auto_validation_delay_days = 7 } ``` Apply the configuration: ``` export PROVIDER_AIVEN_ENABLE_BETA=true terraform init terraform plan terraform apply ``` Define both steps in a single manifest and apply it: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: dev-to-staging spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: dev-project sourceServiceName: pg-dev destinationProjectName: staging-project destinationServiceName: pg-staging autoValidationDelayDays: 3 --- apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: staging-to-prod spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: staging-project sourceServiceName: pg-staging destinationProjectName: prod-project destinationServiceName: pg-prod autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-pipeline.yaml ``` When a maintenance update arrives: 1. The development service receives the update. 2. After testing, validate the development version or wait 3 days for auto-validation. 3. The staging service receives the update during its next maintenance window. 4. After testing, validate the staging version or wait 7 days for auto-validation. 5. The production service receives the update during its next maintenance window. Related pages * [Maintenance and updates for your Aiven for OpenSearch® service](/docs/products/opensearch/howto/maintenance-updates.md) * [Change the service plan](/docs/products/opensearch/howto/change-service-plan.md) * [Service and feature releases](/docs/platform/concepts/service-and-feature-releases.md) * [Aiven CLI](/docs/tools/cli.md) --- # Create a free tier Aiven for OpenSearch® service You can create a free tier Aiven for OpenSearch® service to learn OpenSearch, test indexing and queries, or run small proof-of-concept workloads. For details about Aiven for OpenSearch free tier features, capacity, and restrictions, see [Aiven for OpenSearch free tier](/docs/products/opensearch/concepts/opensearch-free-tier.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Sign-up or sign-in using the [Aiven Console](https://console.aiven.io) * Aiven project important Review [Aiven for OpenSearch free tier limitations](/docs/products/opensearch/concepts/opensearch-free-tier.md#limitations). ## Create a free tier service[​](#create-a-free-tier-service "Direct link to Create a free tier service") 1. In your project, click **Services**. 2. Click **Create service**. 3. Select **OpenSearch®**. 4. In **Service tier**, select **Free**. 5. Review the **Cloud** section. The cloud provider is managed automatically for the free tier. You can select only a region group. 6. In **Service basics**, enter a **Service name**. 7. Review the **Service summary**. 8. Click **Create service**. When the status changes to **Running**, your free tier OpenSearch service is ready to use. To add sample data, use the sample data tools in the Aiven Console. To connect your application, use **Quick connect** to get setup steps and sample code for common programming languages. ## What's next[​](#whats-next "Direct link to What's next") ### Index or query data[​](#index-or-query-data "Direct link to Index or query data") Use your new Aiven for OpenSearch service so that data is indexed or queried within 24 hours from the service creation. warning Your new free tier Aiven for OpenSearch service is powered off automatically after 24 hours if you don't start using it or index or query any data. You receive a notification before shutdown, and you can power on the service from the Aiven Console. ### Upgrade the free tier service[​](#upgrade-the-free-tier-service "Direct link to Upgrade the free tier service") You can upgrade the service at any time to * Enable the following capabilities: * Cloud and region selection * Higher throughput and storage * Advanced OpenSearch features and service integrations * Support for tiered storage * Remove limits and limitations on the following: * Storage * Features To upgrade in the [Aiven Console](https://console.aiven.io) from the service overview: 1. Open your service’s **Overview** page. 2. In **Service plan usage**, click **Upgrade**. To upgrade in the [Aiven Console](https://console.aiven.io) from service settings: 1. In your service, click **Service settings**. 2. In the **Service summary** section, click **Upgrade**. 3. Select a plan and click **Upgrade service**. Upgrades apply immediately. Paid services cannot be downgraded to the free tier. Related pages * [Aiven for OpenSearch free tier overview](/docs/products/opensearch/concepts/opensearch-free-tier.md) * [Aiven free tier services](/docs/platform/concepts/service-pricing.md#free-tier) --- # Custom dictionary files Custom dictionary files are user-defined files that enhance query analysis and improve search relevance in OpenSearch. By adding domain-specific vocabulary and rules, these files refine search results to be more accurate and relevant. Custom dictionary files are categorized into three types: * **Stopwords**: Exclude common words like "the" and "is" to refine search results. * **Synonyms**: Equate similar terms, such as "car" and "automobile," to improve query matching. * **WordNet**: Provide semantic relationships between words, such as synonyms and antonyms. note Ensure your custom dictionary files are in plain text (UTF-8 encoded) format. ## Upload files[​](#upload-files "Direct link to Upload files") Upload new custom dictionary files to your OpenSearch service. * Console * CLI 1. Log in to the [Aiven Console](https://console.aiven.io), select your project, and select your Aiven for OpenSearch service. 2. In the **Data** section, click **Indexes**. 3. Click **Upload file** in the **Custom dictionary files** section. 4. In the **Upload a custom dictionary file** screen: * Select **File type** (Stopwords, Synonyms, WordNet). * Enter a **File name**. * Choose the file from your system and click **Upload**. Run: ``` avn service custom-file upload --project PROJECT_NAME \ --file_type \ --file_path \ --file_name SERVICE_NAME ``` Parameters: * `PROJECT_NAME`: Your Aiven project name. * ``: The type of dictionary file to upload. * ``: Path to the local file on your system. * ``: The name of the file to appear in Aiven for OpenSearch. * `SERVICE_NAME`: Name of your OpenSearch service. ## List files[​](#list-files "Direct link to List files") List all custom dictionary files associated with your OpenSearch service. * Console * CLI In the **Aiven Console**, the **Custom Dictionary Files** section displays all uploaded custom dictionary files, including details such as the file path, type, size, and the most recent upload timestamp. Run: ``` avn service custom-file list --project PROJECT_NAME SERVICE_NAME ``` Parameters: * `PROJECT_NAME`: Your Aiven project name. * `SERVICE_NAME`: Name of your OpenSearch service. ## Replace files[​](#replace-files "Direct link to Replace files") Once you upload a custom dictionary file, you can only replace it, not delete it. To update an existing custom dictionary file, replace it with a new file containing the updated words. * Console * CLI 1. Log in to the [Aiven Console](https://console.aiven.io), select your project, and select your Aiven for OpenSearch service. 2. In the **Data** section, click **Indexes**. 3. In the **Custom dictionary files** section, locate the desired file. 4. Click **Actions** > **Replace file**. 5. Choose the new file from your system and click **Upload**. Run: ``` avn service custom-file update --project PROJECT_NAME \ --file_path \ --file_id SERVICE_NAME ``` Parameters: * `PROJECT_NAME`: Your Aiven project name. * ``: Path to the local file on your system. * ``: ID of the file to replace. Obtain this ID using the [List](#list-files) command. * `SERVICE_NAME`: Name of your OpenSearch service. ## Download files[​](#download-files "Direct link to Download files") Download a custom dictionary file to your local system. * Console * CLI 1. Log in to the [Aiven Console](https://console.aiven.io), select your project, and select your Aiven for OpenSearch service. 2. In the **Data** section, click **Indexes**. 3. In the **Custom dictionary files** section, locate the desired file. 4. Click **Actions** > **Download**. 5. Choose you location and click **Save**. Run: ``` avn service custom-file get --project PROJECT_NAME \ --file_id \ --target_filepath \ --stdout_write SERVICE_NAME ``` Parameters: * `PROJECT_NAME`: Your Aiven project name. * ``: ID of the file to replace to download. Obtain this ID using the [List](#list-files) command. * ``: Path where the file should be saved locally. * `SERVICE_NAME`: Name of your OpenSearch service. ## Limitations[​](#limitations "Direct link to Limitations") * Files cannot be deleted. They can only be replaced. * The file location is fixed and cannot be customized. * If you move to a different cloud or project, files are copied or moved accordingly. * For OpenSearch Cross-Cluster Replication (CCR), files must be uploaded to both services manually. * Use alphanumeric characters and underscores only for file names. ## Example: How to use custom dictionary files with indexes[​](#example-how-to-use-custom-dictionary-files-with-indexes "Direct link to Example: How to use custom dictionary files with indexes") After uploading a custom dictionary file, you can use it in your index settings by specifying custom filters or analyzers. This example demonstrates how to create an index that uses a custom stopwords file. ### Create a stopwords file[​](#create-a-stopwords-file "Direct link to Create a stopwords file") Create a file named `demo_stopwords.txt` with your stopwords. ``` a fox jumps the EOF ``` ### Upload the stopwords file[​](#upload-the-stopwords-file "Direct link to Upload the stopwords file") [Upload](#upload-files) this file using the Aiven Console or CLI. ### Create an index that uses the stopwords file[​](#create-an-index-that-uses-the-stopwords-file "Direct link to Create an index that uses the stopwords file") Create an index using the stopwords file via the OpenSearch Dashboards or the API. * OpenSearch Dashboards * API 1. Log in to the [Aiven Console](https://console.aiven.io), select your project, and select your Aiven for OpenSearch service. 2. Access the **OpenSearch Dashboards** tab in the **Connection information** section. 3. Use the **Service URI** to access OpenSearch Dashboards in a browser. 4. Log in with the provided **User** and **Password**. 5. Click **Index Management** > **Indices** > **Create Index**. 6. Enter the details for the index. 7. Expand the **Advanced settings** section and insert the following JSON configuration to use the stopwords file: ``` { "index.analysis.analyzer.default.filter": [ "custom_stop_words_filter" ], "index.analysis.analyzer.default.tokenizer": "whitespace", "index.analysis.filter.custom_stop_words_filter.ignore_case": "true", "index.analysis.filter.custom_stop_words_filter.stopwords_path": "custom/stopwords/nofox", "index.analysis.filter.custom_stop_words_filter.type": "stop", "index.number_of_replicas": "1", "index.number_of_shards": "1" } ``` 8. Click **Create**. Alternatively, you can use the API by replacing `${SERVICE_URL}` with your service URL and running the following command: ``` curl -X PUT -H "Content-Type: application/json" -d'{ "settings": { "analysis": { "analyzer": { "default": { "tokenizer": "whitespace", "filter": ["custom_stop_words_filter"] } }, "filter": { "custom_stop_words_filter": { "type": "stop", "ignore_case": true, "stopwords_path": "custom/stopwords/nofox" } } } } }' ${SERVICE_URL}/demo-index?pretty ``` ### Verify the stopwords filter[​](#verify-the-stopwords-filter "Direct link to Verify the stopwords filter") Verify the stopwords filter by using the `_analyze` API. * OpenSearch Dashboards * API 1. Go to **Dev Tools** in OpenSearch Dashboards. 2. Use the `_analyze` API to verify that the stopwords filter is working. ``` POST customdictionarytest/_analyze { "text": "a quick brown fox jumps over the lazy dog" } ``` Alternatively, use the API by replacing `${SERVICE_URL}` with your service URL and running the following command: ``` curl -H 'Content-Type: application/json' -d'{ "text": "a quick brown fox jumps over the lazy dog" } ' ${SERVICE_URL}/demo-index/_analyze?pretty ``` Related pages * [Indices](/docs/products/opensearch/concepts/indices.md) * [OpenSearch text analysis](https://opensearch.org/docs/2.13/analyzers/) * [Analyze API](https://opensearch.org/docs/latest/api-reference/analyze-apis/) --- # Manage Aiven for OpenSearch® custom repositories in the Aiven Console or API Use the Aiven Console or API for configuring custom repositories in Aiven for OpenSearch to store [snapshots](/docs/products/opensearch/howto/manage-snapshots.md) in your cloud storage. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven Console * Aiven API - Running Aiven for OpenSearch service - Access to the [Aiven Console](https://console.aiven.io/) - Access to a supported object storage service (AWS S3, GCS, or Azure) - Credentials for the selected storage provider * Running Aiven for OpenSearch service * [Aiven API](/docs/tools/api.md) and authentication [token](/docs/platform/howto/create_authentication_token.md) * Access to a supported object storage service (AWS S3, GCS, or Azure) * Credentials for the selected storage provider note * Repository credentials are omitted from API responses for security reasons. * Aiven API provides a direct interface to the OpenSearch snapshot API. ## Limitations[​](#limitations "Direct link to Limitations") * Aiven Console * Aiven API You can configure custom repositories for the following object storage services: * Amazon S3 or any S3-compatible object storage service * Google Cloud Storage (GCS) * Microsoft Azure Blob Storage - Supported storage services * Amazon S3 or any S3-compatible object storage service * Google Cloud Storage (GCS) * Microsoft Azure Blob Storage - To [edit repository details](/docs/products/opensearch/howto/custom-repositories.md#view-or-edit-repository-details), you can only use the Aiven Console. - The following operations are not supported via Aiven API: * [Remove a repository](/docs/products/opensearch/howto/custom-repositories.md#remove-a-repository) * [View or edit repository details](/docs/products/opensearch/howto/custom-repositories.md#view-or-edit-repository-details) ## Create custom repositories[​](#create-custom-repositories "Direct link to Create custom repositories") Each repository requires a unique name, a storage type (such as S3, Azure, or GCS), and the appropriate settings for the selected storage provider. * Aiven Console * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/), go to your project, and open your service's page. 2. In the **Backups** section, click **Snapshots**. 3. On the **Snapshots** page, click **Add repository**. 4. In the **Add custom repository** window: 1. Enter a repository name. 2. Select a storage provider. 3. Give provider-specific details required for accessing the storage. 4. Click **Add**. Custom repositories are configured in the `user_config` of your Aiven for OpenSearch service. Use the following API request to configure custom repositories: ``` curl -s --url "https://api.aiven.io/v1/project/{project_name}/service/{service_name}" \ --header "Authorization: Bearer $TOKEN" \ --header "Content-Type: application/json" \ -X PUT -d '{ "user_config": { "custom_repos": [ { "name": "azure-repo", "type": "azure", "settings": { "account": "AZURE_ACCOUNT", "base_path": "your/path", "container": "AZURE_CONTAINER", "sas_token": "AZURE_SAS_TOKEN", "readonly": false } }, { "name": "aws-repo", "type": "s3", "settings": { "access_key": "AWS_ACCESS_KEY", "secret_key": "AWS_SECRET_KEY", "base_path": "your/path", "bucket": "AWS_BUCKET", "region": "AWS_REGION", "endpoint": "S3_ENDPOINT", "server_side_encryption": true, "readonly": false } } ] } }' ``` note Aiven for OpenSearch supports Amazon S3 and any S3-compatible object storage service. To connect to a service other than Amazon S3, set `endpoint` to that service's S3 API endpoint URL. ## List custom repositories[​](#list-custom-repositories "Direct link to List custom repositories") * Aiven Console * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/), go to your project, and open your service's page. 2. In the **Backups** section, click **Snapshots**. Find your custom repositories listed on the **Snapshots** page. ``` curl -s --url "https://api.aiven.io/v1/project/{project_name}/service/{service_name}/opensearch/_snapshot"\ --header "Authorization: Bearer $TOKEN" \ --header "Content-Type: application/json" ``` Example response: ``` { "repositories": [ { "name": "aws-repo", "settings": { "base_path": "test/path", "bucket": "testbucket", "endpoint": "http://s3.eu-north-1.amazonaws.com", "region": "eu-north-1", "server_side_encryption": true, "readonly": false }, "type": "s3" }, { "name": "azure-repo", "settings": { "base_path": "test/path", "container": "testcontainer", "readonly": false }, "type": "azure" } ] } ``` ## View or edit repository details[​](#view-or-edit-repository-details "Direct link to View or edit repository details") 1. Log in to the [Aiven Console](https://console.aiven.io/), go to your project, and open your service's page. 2. In the **Backups** section, click **Snapshots**. 3. On the **Snapshots** page, find your custom repository and click **Actions** > **Edit repository**. 4. Edit repository details and save your changes by clicking **Update**. ## Remove a repository[​](#remove-a-repository "Direct link to Remove a repository") 1. Log in to the [Aiven Console](https://console.aiven.io/), go to your project, and open your service's page. 2. In the **Backups** section, click **Snapshots**. 3. On the **Snapshots** page, find your custom repository and click **Actions** > **Remove repository** > **Remove**. ## Error handling[​](#error-handling "Direct link to Error handling") The Aiven API returns OpenSearch errors as they are. **Exceptions:** * 502: OpenSearch did not respond. * 409: The service is not powered on or does not support this feature. Related pages [Manage Aiven for OpenSearch® custom repositories in OpenSearch® API](/docs/products/opensearch/howto/manage-custom-repo/custom-repositories-os-api.md) --- # Aiven for OpenSearch® metrics sent to Datadog Send Aiven for OpenSearch® metrics to Datadog, and choose which metric categories the integration collects. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running Aiven for OpenSearch® service on a paid plan. Service integrations, including Datadog, aren't available on [free tier](/docs/products/opensearch/concepts/opensearch-free-tier.md) services. * A Datadog account * A Datadog [API key](https://docs.datadoghq.com/account_management/api-app-keys/) * A [Datadog Metrics integration](/docs/integrations/datadog/datadog-metrics.md) enabled for your service ## Metrics sent to Datadog[​](#metrics-sent-to-datadog "Direct link to Metrics sent to Datadog") When you enable the Datadog Metrics integration for your Aiven for OpenSearch service, Aiven runs Datadog's [Elasticsearch integration](https://docs.datadoghq.com/integrations/elasticsearch/?tab=host) check against your service. The check always collects node-level and cluster health metrics. Four additional metric categories are off by default. You enable each one independently through the `opensearch` configuration on the Datadog integration. | Category | Configuration parameter | What it collects | | ------------------------ | ---------------------------- | --------------------------------------------------------- | | Cluster monitoring | `cluster_stats_enabled` | Cluster-wide disk usage metrics | | Index monitoring | `index_stats_enabled` | Per-index metrics, such as document counts and store size | | Pending task monitoring | `pending_task_stats_enabled` | Metrics for tasks waiting in the cluster task queue | | Primary shard monitoring | `pshard_stats_enabled` | Primary shard and index count metrics | For the full list of metrics in each category, see [Metrics](https://docs.datadoghq.com/integrations/elasticsearch/?tab=host#metrics) in the Datadog Elasticsearch integration documentation. ## Enable a metric category[​](#enable-a-metric-category "Direct link to Enable a metric category") Enable metric categories through the Datadog integration for your Aiven for OpenSearch service. 1. Find the ID of the Datadog Metrics integration for your service by running the [avn service integration-list](/docs/tools/cli/service/integration.md#avn_service_integration_list) command: ``` avn service integration-list --project PROJECT_NAME SERVICE_NAME ``` Use the `service_integration_id` value from the output as `INTEGRATION_ID` in the following commands. 2. Set the categories to collect to `true`. This example enables all four categories: ``` avn service integration-update --project PROJECT_NAME \ --user-config-json '{ "opensearch": { "cluster_stats_enabled": true, "index_stats_enabled": true, "pending_task_stats_enabled": true, "pshard_stats_enabled": true } }' \ INTEGRATION_ID ``` Include only the categories to change. Setting one category doesn't affect the others. 3. Check that the configuration is set correctly: ``` avn service integration-list SERVICE_NAME \ --project PROJECT_NAME \ --json | jq '.[] | select(.integration_type=="datadog").user_config' ``` Expect output similar to the following: ``` { "opensearch": { "cluster_stats_enabled": true, "index_stats_enabled": true, "pending_task_stats_enabled": true, "pshard_stats_enabled": true } } ``` 4. Find the collected metrics in the Datadog Metrics Explorer under the `elasticsearch.` prefix. Related pages * [Datadog and Aiven](/docs/integrations/datadog.md) * [Send metrics to Datadog](/docs/integrations/datadog/datadog-metrics.md) * [Aiven for OpenSearch® metrics available via Prometheus](/docs/products/opensearch/howto/os-metrics.md) * [Aiven for OpenSearch® free tier](/docs/products/opensearch/concepts/opensearch-free-tier.md) --- # Scale disk storage automatically for your Aiven for OpenSearch® service Automatically increase the disk storage of your Aiven for OpenSearch® service when it's running out of space, instead of resizing it manually. Use the Aiven Autoscaler to automatically increase the storage capacity of a service disk when it's running out of space. Disk autoscaler only increases storage, it doesn't scale storage down. ## Why use disk autoscaling[​](#why-use-disk-autoscaling "Direct link to Why use disk autoscaling") * **Cost efficiency**: Start with a regular-sized disk and let Aiven scale it up only when needed, without the risk of running out of disk space. * **Resiliency**: Avoid a service becoming non-functional because it ran out of disk space, including during unexpected spikes in demand. ## How it works[​](#how-it-works "Direct link to How it works") 1. You create an autoscaler integration endpoint in your project, setting the maximum total disk size to allow. 2. You enable an autoscaler integration for your service using that endpoint. 3. Aiven monitors the disk space usage of your service. 4. When disk usage reaches the threshold for your service type, Aiven increases the available storage by at least 10%, using the current used space as a baseline. note The exact increase depends on the service type and cloud provider. Some providers enforce a minimum increase of 10 GB. Autoscale thresholds per service type The threshold that triggers disk autoscaling is a percentage of the available disk storage capacity: * Aiven for OpenSearch®: 75% of the available disk storage capacity * All other supported service types: 85% of the available disk storage capacity 5. The disk increase is recorded in the project event log, and you receive a notification about the added disk space. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * **Maximum storage**: The maximum storage that the autoscaler can allocate for your service is limited by both the maximum disk size set on the autoscaler endpoint and the maximum disk storage supported for your service plan. * **Timing**: Autoscaling takes a moment to complete. In the meantime, the service disk might fill up and the service might enter read-only mode until autoscaling finishes, unless the autoscaler's disk capacity limit is reached. * **Maintenance updates**: Autoscaling works only on fully running services and can't happen during a maintenance update. * **Manual changes**: Changing disk space manually can delay an autoscaling event. * **Terraform**: Don't manage disk space with the Aiven Terraform Provider on a service that uses the autoscaler, to avoid conflicts between the two. * **Performance**: Disk added through autoscaling is slower than the original disk until the next maintenance update applies. This might affect I/O-intensive workloads. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven organization, project, and service that's up and running * The operator role for the organization, project, and service * Dynamic disk sizing support on your service plan and cloud region * One of the following to manage the autoscaler: * [Aiven Console](https://console.aiven.io/) * [Aiven API](https://api.aiven.io/doc/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) ### Enable disk autoscaling[​](#enable-disk-autoscaling "Direct link to Enable disk autoscaling") To enable disk autoscaling, create an autoscaler integration endpoint, then enable an autoscaler integration on your service using that endpoint. * Console * API * CLI * Terraform Create an autoscaler endpoint: 1. Log in to [Aiven Console](https://console.aiven.io/), and go to your organization and project. 2. On the left sidebar, click **Integration endpoints**. 3. Click **Aiven Autoscaler** > **Add new endpoint**. 4. Set the endpoint name and the maximum total disk storage in GB, and click **Add endpoint**. Enable the autoscaler on a service: 1. On the left sidebar, click **Services**, and open your service. 2. On the left sidebar, click **Integrations**. 3. In **Endpoint integrations**, click **Aiven Autoscaler**. 4. Click the endpoint you created, and click **Enable**. 1) Call [ServiceIntegrationEndpointCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointCreate) to create an autoscaler integration endpoint on your project: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "endpoint_name": "ENDPOINT_NAME", "endpoint_type": "autoscaler", "user_config": { "autoscaling": [ { "type": "autoscale_disk", "cap_gb": 300 } ] } }' ``` 2) Call [ServiceIntegrationCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationCreate) to enable the autoscaler integration on your service, using the endpoint ID from the previous response: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "dest_endpoint_id": "ENDPOINT_ID", "integration_type": "autoscaler", "source_project": "PROJECT_NAME", "source_service": "SERVICE_NAME" }' ``` 1. Create an autoscaler integration endpoint using [avn service integration-endpoint-create](/docs/tools/cli.md): ``` avn service integration-endpoint-create \ --project PROJECT_NAME \ --endpoint-name ENDPOINT_NAME \ --endpoint-type autoscaler \ --user-config-json '{"autoscaling": [{"type": "autoscale_disk", "cap_gb": 300}]}' ``` 2. Find the ID of the new endpoint: ``` avn service integration-endpoint-list --project PROJECT_NAME ``` 3. Enable the autoscaler integration on your service, using the endpoint ID from the previous step: ``` avn service integration-create \ --dest-service SERVICE_NAME \ --integration-type autoscaler \ --source-endpoint-id ENDPOINT_ID ``` Use the [`aiven_service_integration_endpoint`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration_endpoint) and [`aiven_service_integration`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration) resources: ``` resource "aiven_service_integration_endpoint" "autoscaler_endpoint" { project = "PROJECT_NAME" endpoint_name = "ENDPOINT_NAME" endpoint_type = "autoscaler" autoscaler_user_config { autoscaling { type = "autoscale_disk" cap_gb = 300 } } } resource "aiven_service_integration" "autoscaler_integration" { project = "PROJECT_NAME" integration_type = "autoscaler" source_service_name = "SERVICE_NAME" destination_endpoint_id = aiven_service_integration_endpoint.autoscaler_endpoint.id } ``` See the [disk autoscaler guide](https://registry.terraform.io/providers/aiven/aiven/latest/docs/guides/disk-autoscaler) for more details. ### Change the maximum disk space for autoscaling[​](#change-the-maximum-disk-space-for-autoscaling "Direct link to Change the maximum disk space for autoscaling") After you enable disk autoscaling, you can update the maximum total disk size at any time. * Console * API * CLI * Terraform 1. On the left sidebar, click **Integration endpoints**. 2. Click **Aiven Autoscaler**. 3. Find your endpoint, click **Actions**, and click the option to edit it. 4. Set a new maximum disk storage value, and save your changes. Call [ServiceIntegrationEndpointUpdate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointUpdate) with the new `cap_gb` value: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint/ENDPOINT_ID \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "user_config": { "autoscaling": [ { "type": "autoscale_disk", "cap_gb": 500 } ] } }' ``` ``` avn service integration-endpoint-update ENDPOINT_ID \ --user-config-json '{"autoscaling": [{"type": "autoscale_disk", "cap_gb": 500}]}' ``` Update the `cap_gb` value in the `autoscaling` block of your [`aiven_service_integration_endpoint`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration_endpoint) resource, then apply the change. ### Turn off disk autoscaling[​](#turn-off-disk-autoscaling "Direct link to Turn off disk autoscaling") To turn off disk autoscaling, remove the autoscaler integration from your service. You can also delete the integration endpoint if you no longer need it. * Console * API * CLI Disconnect the service from the autoscaler: 1. On the left sidebar, click **Services**, and open your service. 2. On the left sidebar, click **Integrations**. 3. In **Endpoint integrations**, find **Aiven Autoscaler**, click **Actions**, and click the option to disconnect it. Delete the autoscaler endpoint, if you no longer need it: 1. On the left sidebar, click **Integration endpoints**. 2. Click **Aiven Autoscaler**. 3. Find your endpoint, and delete it. 1) Call [ServiceIntegrationDelete](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationDelete) to remove the autoscaler integration from your service: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration/INTEGRATION_ID \ --header "Authorization: Bearer TOKEN" ``` 2) Call [ServiceIntegrationEndpointDelete](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointDelete) to delete the autoscaler integration endpoint, if you no longer need it: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint/ENDPOINT_ID \ --header "Authorization: Bearer TOKEN" ``` 1. Find the ID of the integration to remove: ``` avn service integration-list SERVICE_NAME ``` 2. Remove the autoscaler integration from your service: ``` avn service integration-delete INTEGRATION_ID ``` 3. Find the ID of the integration endpoint to delete, if you no longer need it: ``` avn service integration-endpoint-list --project PROJECT_NAME ``` 4. Delete the autoscaler integration endpoint: ``` avn service integration-endpoint-delete ENDPOINT_ID ``` Related pages * [Scale disk storage manually](/docs/products/opensearch/howto/scale-disk-storage.md) * [Change the service plan](/docs/products/opensearch/howto/change-service-plan.md) * [Memory and out-of-memory conditions](/docs/products/opensearch/concepts/service-memory.md) --- # Enable OpenSearch Security management for Aiven for OpenSearch® [OpenSearch Security](/docs/products/opensearch/concepts/os-security.md) provides a range of security features, including fine-grained access controls, SAML authentication, and audit logging to monitor activity within your Aiven for OpenSearch® service. By enabling this, you can manage user permissions, roles, and other security aspects through the OpenSearch Dashboard. info In addition to Role-Based Access Control (**RBAC**), the following external authentication methods are supported for Aiven for OpenSearch Security: * Security Assertion Markup Language (**SAML**) * OpenID Connect (**OIDC**) ## Considerations before enabling OpenSearch Security management[​](#considerations-before-enabling-opensearch-security-management "Direct link to Considerations before enabling OpenSearch Security management") Before enabling OpenSearch Security management on your Aiven for OpenSearch service, note the following: * OpenSearch Security management cannot be disabled once enabled. Therefore, ensure that you thoroughly understand the security features and implications before proceeding. If you need assistance disabling OpenSearch Security management, contact [Aiven support](https://aiven.io/support-services). * Fine-grained user access control can be managed through the OpenSearch Dashboard after enabling OpenSearch Security management for the service. * Any existing user roles and permissions will be automatically transferred to the OpenSearch Dashboard. * To ensure the security of your OpenSearch service, managing the security features of OpenSearch is limited only to a dedicated administrator role. * Once you have enabled OpenSearch Security management, you can no longer use [Aiven Console](https://console.aiven.io/), [Aiven API](https://api.aiven.io/doc/), [Aiven CLI](/docs/tools/cli.md), [Aiven Terraform Provider](/docs/tools/terraform.md) or [Aiven Operator for Kubernetes®](/docs/tools/kubernetes.md) to manage access controls. ## Enable OpenSearch Security[​](#enable-opensearch-security "Direct link to Enable OpenSearch Security") To activate OpenSearch Security management for your Aiven for OpenSearch service: 1. Log in to the [Aiven Console](https://console.aiven.io/) and access the Aiven for OpenSearch service for which to enable security. 2. On the service page, click **Users** in the sidebar. 3. On the **Users** page, click **Enable OpenSearch Security**. 4. Review the information in the **OpenSearch Security management** window, confirm you understand and want to proceed by selecting the checkbox, and click **Continue**. 5. Create your administrator user by entering a password for this user. note * OpenSearch Security administrator username set by default cannot be changed. * To reset the password later, contact [Aiven Support](mailto:support@aiven.io). 6. Click **Enable OpenSearch Security** to create the administrator user and activate OpenSearch Security management. After activating OpenSearch Security management, you are redirected to the **Users** page, where you can verify that the security feature is enabled. To manage user permissions and other security settings, access OpenSearch Security management by logging in to the [OpenSearch Dashboard](/docs/products/opensearch/dashboards.md) using your security admin credentials. --- # Enable slow query logging Identify inefficient or time-consuming queries by enabling [slow query logging](https://docs.opensearch.org/latest/install-and-configure/configuring-opensearch/logs/#search-request-slow-logs) in your Aiven for OpenSearch® service. Slow query logging records queries that exceed a specified time threshold, helping you diagnose performance issues and optimize query patterns. This is an advanced feature for power users who need to deep dive into query performance analysis. You configure slow query logging using [advanced parameters](/docs/products/opensearch/reference/advanced-params.md) that control the logging behavior at the cluster level. important Both the log `level` and its corresponding `threshold` must be configured for slow query logging to work. `level` controls which `threshold` is applied. Setting only the log `level` without a `threshold` will not generate any logs. ``` "slowlog": { "level": "info", "threshold": { "trace": "1s", "debug": "10s", "info": "30s", "warn": "60s" } } ``` In this example, `info` and `warn` level messages can appear in logs. note Slow query logging can impact CPU performance. Start with a higher threshold value and monitor your service performance after enabling logging. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven for OpenSearch service * Access to manage the service configuration: * [Aiven Console](https://console.aiven.io/) or * [Aiven CLI](/docs/tools/cli.md) with a [personal token](/docs/platform/howto/create_authentication_token.md) or * [Aiven API](/docs/tools/api.md) with a [personal token](/docs/platform/howto/create_authentication_token.md) or * [Aiven Provider for Terraform](/docs/tools/terraform.md) or * [Aiven Operator for Kubernetes®](/docs/tools/kubernetes.md) ## About the configuration parameters[​](#about-the-configuration-parameters "Direct link to About the configuration parameters") Slow query logging uses two advanced configuration parameters: * **Log level** (`opensearch.cluster.search.request.slowlog.level`): Determines the severity level for logged queries. Choose from `debug`, `info`, `trace`, or `warn`. * **Threshold** (`opensearch.cluster.search.request.slowlog.threshold.`): Sets the time limit for queries. Queries exceeding this time are logged. The threshold parameters must match the log level you choose. ## Enable slow query logging[​](#enable-slow-query-logging "Direct link to Enable slow query logging") * Console * CLI * API * Terraform * Kubernetes 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. On the **Services** page, select your Aiven for OpenSearch service. 3. On the **Service settings** page, scroll to the **Advanced configuration** section and click **Configure**. 4. In the **Advanced configuration** window: 1. Click **Add configuration options**. From the list, select `opensearch.cluster.search.request.slowlog.level`. 2. Set the value to one of the following: `debug`, `info`, `trace`, or `warn`. 3. Click **Add configuration options**. From the list, select threshold configuration options: * `opensearch.cluster.search.request.slowlog.threshold.debug` * `opensearch.cluster.search.request.slowlog.threshold.info` * `opensearch.cluster.search.request.slowlog.threshold.trace` * `opensearch.cluster.search.request.slowlog.threshold.warn` 4. Set each threshold value as a number followed by a time unit with no space. Queries exceeding this time will be logged. Start with a higher value like `10s` or `20s` and adjust based on your needs. note * Default value: `-1` (disabled) * Allowed units: `s` (seconds), `m` (minutes), `h` (hours), `d` (days), `nanos` (nanoseconds), `ms` (milliseconds), `micros` (microseconds) * Example values: `1s`, `500ms`, `2m` 5. Click **Save configuration**. Use the [`avn service update`](/docs/tools/cli/service-cli.md#avn-cli-service-update) command to configure slow query logging: ``` avn service update SERVICE_NAME \ -c opensearch.cluster.search.request.slowlog.level=LEVEL_A \ -c opensearch.cluster.search.request.slowlog.threshold.LEVEL_A=THRESHOLD_A \ -c opensearch.cluster.search.request.slowlog.threshold.LEVEL_B=THRESHOLD_B \ ``` Parameters: * `SERVICE_NAME`: Your Aiven for OpenSearch service name * `LEVEL`: Log level (`debug`, `info`, `trace`, or `warn`) * `THRESHOLD`: Time threshold (for example, `1s`, `500ms`, `2m`) Example: ``` avn service update my-opensearch \ -c opensearch.cluster.search.request.slowlog.level=info \ -c opensearch.cluster.search.request.slowlog.threshold.info=10s \ -c opensearch.cluster.search.request.slowlog.threshold.warn=30s ``` Call the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint to configure slow query logging: ``` curl --request PUT \ --url "https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME" \ --header "Authorization: Bearer API_TOKEN" \ --header "Content-Type: application/json" \ --data '{ "user_config": { "opensearch": { "cluster_search_request_slowlog": { "level": "LEVEL", "threshold": { "warn": "TIME_LIMIT", "info": "TIME_LIMIT", "trace": "TIME_LIMIT", "debug": "TIME_LIMIT", } } } } } }' ``` Parameters: * `PROJECT_NAME`: Your Aiven project name * `SERVICE_NAME`: Your Aiven for OpenSearch service name * `API_TOKEN`: Your [personal token](/docs/platform/howto/create_authentication_token.md) * `LEVEL`: Log level (`debug`, `info`, `trace`, or `warn`) * `THRESHOLD`: Time threshold (for example, `1s`, `500ms`, `2m`) Example: ``` curl --request PUT \ --url "https://api.aiven.io/v1/project/my-project/service/my-opensearch" \ --header "Authorization: Bearer your-api-token" \ --header "Content-Type: application/json" \ --data '{ "user_config": { "opensearch": { "cluster_search_request_slowlog": { "level": "warn", "threshold": { "warn": "10s" } } } } }' ``` Add the slow query logging configuration to your `aiven_opensearch` resource: ``` resource "aiven_opensearch" "example_opensearch" { project = var.aiven_project_name cloud_name = "google-europe-west1" plan = "startup-4" service_name = "my-opensearch" opensearch_user_config { opensearch { cluster_search_request_slowlog { level = "warn" threshold { warn = "10s" } } } } } ``` For more configuration options, see the [`aiven_opensearch` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/opensearch#nested-schema-for-opensearch_user_configopensearchcluster_search_request_slowlog). Add the slow query logging configuration to your `OpenSearch` resource: ``` apiVersion: aiven.io/v1alpha1 kind: OpenSearch metadata: name: my-opensearch spec: project: PROJECT_NAME cloudName: google-europe-west1 plan: startup-4 userConfig: opensearch: cluster_search_request_slowlog: level: warn threshold: warn: 10s ``` For more configuration options, see the [OpenSearch resource documentation](https://aiven.github.io/aiven-operator/resources/opensearch.html#spec.userConfig.opensearch.cluster.search.request.slowlog). ## View slow query logs[​](#view-slow-query-logs "Direct link to View slow query logs") After configuring slow query logging, view the logs in the [Aiven Console](https://console.aiven.io/): 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. On the **Services** page, select your Aiven for OpenSearch service. 3. In the **Observe** section, click **Logs**. 4. Search for slow query entries using the search field or filter by log level. Slow query log entries include: * Query execution time * Query details * Index name * Number of shards queried To send logs to another Aiven for OpenSearch service, see [Enable logs integration](/docs/products/opensearch/howto/opensearch-log-integration.md#enable-log-integration). ## Adjust the threshold[​](#adjust-the-threshold "Direct link to Adjust the threshold") To capture more or fewer slow queries, adjust the threshold value: * **Lower the threshold** to capture more queries (for example, change from `10s` to `5s`) * **Raise the threshold** to capture only slower queries (for example, change from `10s` to `30s`) tip Start with a higher threshold (such as `20s` to `30s`) and gradually lower it to avoid generating excessive logs. ## Disable slow query logging[​](#disable-slow-query-logging "Direct link to Disable slow query logging") To disable slow query logging, set the threshold to `-1`: * Console * CLI * API * Terraform * Kubernetes 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. On the **Services** page, select your Aiven for OpenSearch service. 3. Go to **Service settings** > **Advanced configuration**. 4. Locate the threshold parameters you configured. 5. Change its value to `-1`. 6. Click **Save configuration**. ``` avn service update SERVICE_NAME \ -c opensearch.cluster.search.request.slowlog.threshold.LEVEL=-1 ``` Replace `SERVICE_NAME` with your service name and `LEVEL` with the log level you configured (for example, `warn`, `info`). Call the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint: ``` curl --request PUT \ --url "https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME" \ --header "Authorization: Bearer API_TOKEN" \ --header "Content-Type: application/json" \ --data '{ "user_config": { "opensearch": { "cluster_search_request_slowlog": { "threshold": { "LEVEL": "-1" } } } } }' ``` Replace `PROJECT_NAME`, `SERVICE_NAME`, `API_TOKEN`, and `LEVEL` with your values. Update the threshold in your `aiven_opensearch` resource: ``` resource "aiven_opensearch" "example_opensearch" { project = var.aiven_project_name cloud_name = "google-europe-west1" plan = "startup-4" service_name = "my-opensearch" opensearch_user_config { opensearch { cluster_search_request_slowlog { level = "warn" threshold { warn = "-1" } } } } } ``` Update the threshold in your `OpenSearch` resource: ``` apiVersion: aiven.io/v1alpha1 kind: OpenSearch metadata: name: my-opensearch spec: project: PROJECT_NAME cloudName: google-europe-west1 plan: startup-4 userConfig: opensearch: cluster_search_request_slowlog: level: warn threshold: warn: "-1" ``` ## Alternative: Query Insights plugin[​](#alternative-query-insights-plugin "Direct link to Alternative: Query Insights plugin") For continuous query monitoring with less manual configuration, consider using the [Query Insights plugin](https://docs.opensearch.org/latest/observing-your-data/query-insights/index/). This plugin provides: * Automatic tracking of top N queries by CPU usage, latency, or memory * Query grouping and aggregation * Integration with OpenSearch Dashboards for visualization * Lower performance overhead compared to detailed slow query logging Configure Query Insights using the `search.insights.top_queries` advanced configuration parameters. Related pages * [OpenSearch slow query logging documentation](https://docs.opensearch.org/latest/install-and-configure/configuring-opensearch/logs/#search-request-slow-logs) * [Advanced parameters for Aiven for OpenSearch](/docs/products/opensearch/reference/advanced-params.md) * [OpenSearch Query Insights plugin](https://docs.opensearch.org/latest/observing-your-data/query-insights/index/) * [OpenSearch top N queries](https://docs.opensearch.org/latest/observing-your-data/query-insights/top-n-queries/) --- # Fork your Aiven for OpenSearch® service Fork your Aiven for OpenSearch® service to create an independent copy for testing, debugging, or development without affecting the original service. Fork an Aiven service to create a complete copy of it from its latest backup. Forked services are independent and don't share resources with or increase the load on the original service. Common use cases for forking include: * Creating a snapshot to analyze an issue. * Creating a development copy of your production environment. * Testing upgrades before applying them to production services. * Creating an instance in a different cloud provider, region, or with a different plan. * Renaming a service. During the forking process, the fork might initially have only one node while backups are being taken. The other nodes appear after the backup process is complete. When you fork a service, its configuration, indices, and service users are copied to the new service. ## Limitations[​](#limitations "Direct link to Limitations") * You can only fork services that have at least one [backup](/docs/products/opensearch/howto/restore_opensearch_backup.md). * Service integrations are not copied to the fork. * Cross-project forking is supported only within the same organization. * Single sign-on (SSO) methods are not copied to forked services because they are linked to specific URLs and endpoints that change during forking. If you don't reconfigure the SSO methods for the forked service, user access can be disrupted. ## Fork a service[​](#fork-a-service "Direct link to Fork a service") * Console * CLI * API * Terraform 1. In your service, in the **Backups** section, click **Backup management**. 2. Click **Fork & restore**. 3. Choose the backup to fork from. 4. Enter a name, and select the cloud and plan. 5. Click **Create fork**. Use the [create service command](/docs/tools/cli/service-cli.md#avn-cli-service-create) with: * `--service-to-fork-from`: the name of the service to use as the source. * `--project-to-fork-from`: to fork a service in a different project, set this to the project name the source service is in. Use the [`ServiceCreate` endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) and in the `user_config` property set: * `service_to_fork_from`: the name of the source service. * `project_to_fork_from`: to fork a service in a different project, set this to the name of the project the source service is in. Use the `service_to_fork_from` attribute in the user config of your service resource. To fork a service in a different project, set the `project_to_fork_from` attribute. More information on the service resources and their configuration options is available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Restore an Aiven for OpenSearch® backup](/docs/products/opensearch/howto/restore_opensearch_backup.md) * [Rename your Aiven for OpenSearch® service](/docs/products/opensearch/howto/rename-service.md) --- # Handle low disk space in Aiven for OpenSearch® Free up disk space in Aiven for OpenSearch® and recover from the flood stage watermark when a node runs low on disk. OpenSearch relies on [watermarks](/docs/products/opensearch/reference/low-space-watermarks.md) to respond to low disk space. When you're running low on disk space, use one of the following options: * Upgrade to a larger plan using the [Aiven Console](https://console.aiven.io/) or the [Aiven CLI](https://github.com/aiven/aiven-client). * Clean up unnecessary indices. For logs, create separate daily indices so you can clean up the oldest data efficiently. ## Recover after cleaning up space[​](#recover-after-cleaning-up-space "Direct link to Recover after cleaning up space") If OpenSearch exceeded only the low or high watermark, no further action is needed once you free up space: OpenSearch continues to allow writes. If OpenSearch exceeded the flood stage watermark, it also sets `index.blocks.read_only_allow_delete` on every index with a shard on the affected node, not just the one you're cleaning up. Freeing up space doesn't clear this setting. Unset it on all affected indices in one request using `_all` as the target: ``` curl -X PUT "https://USER:PASSWORD@HOST:PORT/_all/_settings" \ -H 'Content-Type: application/json' \ -d '{ "index.blocks.read_only_allow_delete": null }' ``` To unset it on a single index instead, replace `_all` with that index's name. If the affected node is still over the flood stage watermark after you clear this setting, OpenSearch sets it again automatically. Confirm disk usage has actually dropped below the watermark before you try to clear the setting. note `index.blocks.read_only_allow_delete` applies at the index level. While it's set, you can't delete individual documents from an index to shrink it, only delete the whole index. Free up space by deleting entire indices or removing data elsewhere, then clear the setting. Replace the following: * `USER`: the username for the OpenSearch cluster. * `PASSWORD`: the password for the OpenSearch cluster. * `HOST`: the hostname for the connection. * `PORT`: the port number for the connection. note Aiven for OpenSearch doesn't unset `index.blocks.read_only_allow_delete` automatically, to avoid the index flipping between read-only and read-write as disk usage fluctuates near the threshold. Related pages * [Disk watermarks in Aiven for OpenSearch®](/docs/products/opensearch/reference/low-space-watermarks.md) * [Manage indices in Aiven for OpenSearch®](/docs/products/opensearch/concepts/indices.md) * [Scale disk storage](/docs/products/opensearch/howto/scale-disk-storage.md) --- # Manage hot/warm data tiering in Aiven for OpenSearch® [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Set up Index State Management (ISM) policies to automate index lifecycle management across hot and warm data nodes in Aiven for OpenSearch®. This is a [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) feature. See [hot/warm data tiering](/docs/products/opensearch/concepts/hot-warm-tiering.md) for a feature overview. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A custom plan with hot and warm data nodes. [Contact Aiven](https://aiven.io/contact) to request one. * Aiven for OpenSearch® 2.19 or later. ## Verify node tier attributes[​](#verify-node-tier-attributes "Direct link to Verify node tier attributes") Before configuring ISM policies, confirm that all data nodes have their tier attribute set: ``` GET _cat/nodeattrs?v&h=node,attr,value ``` The response must show `temp=hot` on all hot data nodes and `temp=warm` on all warm data nodes. Cluster manager, coordinator, and Dashboards nodes have no `temp` attribute—that is expected. ## Bootstrap the first index[​](#bootstrap-the-first-index "Direct link to Bootstrap the first index") Create your first index with the shard allocation requirement pinned to the hot tier and attach a rollover alias. Applications write to the alias. ISM rolls over and creates new indices that inherit the same settings from your index template. ``` PUT logs-000001 { "aliases": { "logs": { "is_write_index": true } }, "settings": { "index.routing.allocation.require.temp": "hot", "index.number_of_shards": 3, "index.number_of_replicas": 1 } } ``` Create an index template for `logs-*` so that every index created by rollover also receives `index.routing.allocation.require.temp: hot`. ## Create an ISM policy[​](#create-an-ism-policy "Direct link to Create an ISM policy") The following sections show two common ISM policy patterns. Apply a policy to the rollover alias so every new index enters the lifecycle automatically. ### Time-series logs: hot for 7 days, warm for 23 days, then delete[​](#time-series-logs-hot-for-7-days-warm-for-23-days-then-delete "Direct link to Time-series logs: hot for 7 days, warm for 23 days, then delete") Use this pattern for application logs, metrics, or audit trails with a fixed 30-day retention window. ``` PUT _plugins/_ism/policies/logs-hot-warm-policy { "policy": { "description": "Logs stay hot for 7 days, warm for 23 days, then deleted", "default_state": "hot", "states": [ { "name": "hot", "actions": [ { "rollover": { "min_index_age": "1d", "min_primary_shard_size": "40gb" } } ], "transitions": [ { "state_name": "warm", "conditions": { "min_index_age": "7d" } } ] }, { "name": "warm", "actions": [ { "index_priority": { "priority": 50 } }, { "allocation": { "require": { "temp": "warm" }, "include": {}, "exclude": {}, "wait_for": true } }, { "force_merge": { "max_num_segments": 1 } } ], "transitions": [ { "state_name": "delete", "conditions": { "min_index_age": "30d" } } ] }, { "name": "delete", "actions": [{ "delete": {} }] } ], "ism_template": [ { "index_patterns": ["logs-*"], "priority": 100 } ] } } ``` What this policy does: 1. Creates an index on hot nodes (`temp=hot` required). 2. Rolls over when the index is 1 day old or a primary shard reaches 40 GB. The previous index stops receiving writes. 3. Transitions to warm 7 days after creation. ISM updates `require.temp` to `warm`. OpenSearch relocates shards to warm nodes. `wait_for: true` holds the transition until relocation completes. 4. Force-merges to 1 segment per shard on warm. Warm nodes do less write work, so this is low-cost. 5. Deletes the index 30 days after creation. ### Search-focused data: hot for 30 days, then warm indefinitely[​](#search-focused-data-hot-for-30-days-then-warm-indefinitely "Direct link to Search-focused data: hot for 30 days, then warm indefinitely") Use this pattern for product catalogs, document search, or any workload where all historical data must be queryable but older data can tolerate slower response times. This pattern does not delete data. ``` PUT _plugins/_ism/policies/catalog-hot-warm-policy { "policy": { "description": "Catalog indices live hot for 30 days, then move to warm permanently", "default_state": "hot", "states": [ { "name": "hot", "actions": [], "transitions": [ { "state_name": "warm", "conditions": { "min_index_age": "30d" } } ] }, { "name": "warm", "actions": [ { "allocation": { "require": { "temp": "warm" }, "wait_for": true } }, { "force_merge": { "max_num_segments": 1 } }, { "index_priority": { "priority": 20 } } ], "transitions": [] } ], "ism_template": [ { "index_patterns": ["catalog-*"], "priority": 100 } ] } } ``` Search requests that use aliases or `catalog-*` wildcards hit both hot and warm shards. Warm shards respond more slowly. For most workloads, letting OpenSearch aggregate across all tiers is the correct behavior. ## Move an index to warm manually[​](#move-an-index-to-warm-manually "Direct link to Move an index to warm manually") Use manual transitions for ops work such as responding to an incident or decommissioning a dataset without waiting for the policy schedule. 1. Attach a policy to an index: ``` POST _plugins/_ism/add/logs-000042 { "policy_id": "logs-hot-warm-policy" } ``` 2. Force-transition the index to warm: ``` POST _plugins/_ism/change_policy/logs-000042 { "policy_id": "logs-hot-warm-policy", "state": "warm" } ``` 3. Check the ISM state: ``` GET _plugins/_ism/explain/logs-000042 ``` The response shows which state the index is in, when it last transitioned, and what is blocking any pending action—for example, shards still relocating. ## Size warm nodes[​](#size-warm-nodes "Direct link to Size warm nodes") Warm nodes must hold all data within the warm retention window. Calculate warm storage as follows: ``` daily_ingest_rate × warm_retention_days × (1 + number_of_replicas) ``` Add at least 20% headroom above OpenSearch's high-watermark threshold. By default, OpenSearch applies the same watermarks (85% / 90% / 95%) to both tiers. Monitor warm disk usage separately from hot. Dynamic Disk Sizing adds capacity to both tiers at the same time, distributed proportionally to each tier's base volume size, so account for that when you size the warm tier. ## Troubleshoot[​](#troubleshoot "Direct link to Troubleshoot") ### Shards not moving to warm[​](#shards-not-moving-to-warm "Direct link to Shards not moving to warm") The most common cause is a missing `node.attr.temp` attribute on a node. Verify with: ``` GET _cat/nodeattrs?v&h=node,attr,value ``` All hot data nodes must show `temp=hot` and all warm data nodes must show `temp=warm`. If an attribute is missing, [contact Aiven support](https://aiven.io/support). ### ISM transition completed but shards remain on hot nodes[​](#ism-transition-completed-but-shards-remain-on-hot-nodes "Direct link to ISM transition completed but shards remain on hot nodes") Without `"wait_for": true` in the allocation action, ISM marks the transition complete when it updates the setting—before shards relocate. Always pair an `allocation` action with `"wait_for": true` when followed by `force_merge`, which assumes shards are on their final tier. ### Allocation filter reference[​](#allocation-filter-reference "Direct link to Allocation filter reference") The following table describes the allocation filter settings for tiering: | Setting | Behavior | | --------- | -------------------------------------------------------------------------------------- | | `require` | Shard must be on a node matching all listed attributes. Use this for tiering. | | `include` | Shard may be on any node matching at least one attribute. | | `exclude` | Shard must not be on matching nodes. Use this for "anywhere but hot" during backfills. | Most tiering configurations need only `require`. ## API reference[​](#api-reference "Direct link to API reference") The following table lists useful API endpoints for managing tiered clusters: | Purpose | Endpoint | | -------------------------- | -------------------------------------------- | | List node tier attributes | `GET _cat/nodeattrs?v&h=node,attr,value` | | See where each shard lives | `GET _cat/shards?v&h=index,shard,node,state` | | Get per-index ISM state | `GET _plugins/_ism/explain/` | | Debug unallocated shards | `GET _cluster/allocation/explain` | | List all ISM policies | `GET _plugins/_ism/policies` | | Retry a failed ISM action | `POST _plugins/_ism/retry/` | Related pages * [Hot/warm data tiering in Aiven for OpenSearch®](/docs/products/opensearch/concepts/hot-warm-tiering.md) * [Index State Management policies](/docs/products/opensearch/howto/migrate-ism-policies.md) * [Resolve low disk space issues](/docs/products/opensearch/howto/handle-low-disk-space.md) --- # Copy data from OpenSearch to Aiven for OpenSearch® using elasticsearch-dump Backup your OpenSearch® data into Aiven for Opensearch. To copy the index data, we will be using `elasticsearch-dump` [tool](https://github.com/elasticsearch-dump/elasticsearch-dump). You can read the [instructions on GitHub](https://github.com/elasticsearch-dump/elasticsearch-dump/blob/master/README.md) on how to install it. From this library, we will use `elasticdump` command to copy the input index data to an specific output. ## Prerequisites[​](#copy-data-from-os-to-os "Direct link to Prerequisites") * `elasticsearch-dump` [tool](https://github.com/elasticsearch-dump/elasticsearch-dump) installed * OpenSearch cluster as the `input` (can be in Aiven or elsewhere) * Aiven for OpenSearch cluster as the `output` note The `input` and `ouput` can be either an OpenSearch URI or a file path (local or remote file storage). In this particular case, we are using both URLs, one from an **OpenSearch cluster** and the other one from **Aiven for OpenSearch cluster**. To copy your data, collect this information: OpenSearch cluster: * `INPUT_SERVICE_URI`: OpenSearch cluster URI, in the format `https://user:password@host:port` * `INPUT_INDEX_NAME`: the index that you aim to copy from your input source. Aiven for OpenSearch: * `OUTPUT_SERVICE_URI`: your output OpenSearch service URI. You can find it in Aiven's dashboard. * `OUTPUT_INDEX_NAME`: the index to have in your output with the copied data. note Use `export` command to assign your variables in the command line before running `elasticdump` command. For example, suppose your `INPUT_SERVICE_URI` is `myexample`: ``` export INPUT_INDEX_NAME=myexample ``` ## Import mapping[​](#import-mapping "Direct link to Import mapping") The process of defining how a document and the fields are stored and indexed is called mapping. When no data structure is specified, we rely on OpenSearch to automatically detect the fields using dynamic mapping. However, we can set our data mapping before the data is sent: ``` elasticdump \ --input=$INPUT_SERVICE_URI/$INPUT_INDEX_NAME \ --output=$OUTPUT_SERVICE_URI/$OUTPUT_INDEX_NAME \ --type=mapping ``` ## Import data[​](#import-data "Direct link to Import data") This is how you can copy your index data from an OpenSearch cluster (can be in Aiven or elsewhere) to an Aiven for OpenSearch one. ``` elasticdump \ --input=$INPUT_SERVICE_URI/$INPUT_INDEX_NAME \ --output=$OUTPUT_SERVICE_URI/$OUTPUT_INDEX_NAME \ --type=data ``` When the dump is completed, you can check that the index is available in the OpenSearch service you send it to. You will be able to find it under **Indexes** in the **Data** section in your Aiven Console. --- # Copy data from Aiven for OpenSearch® to AWS S3 using elasticsearch-dump Backup your OpenSearch® data into an AWS S3 bucket. To copy the index data, we will be using `elasticsearch-dump` [tool](https://github.com/elasticsearch-dump/elasticsearch-dump). You can read the [instructions on GitHub](https://github.com/elasticsearch-dump/elasticsearch-dump/blob/master/README.md) on how to install it. From this library, we will use `elasticdump` command to copy the input index data to an specific output. ## Prerequisites[​](#copy-data-from-os-to-s3 "Direct link to Prerequisites") * `elasticsearch-dump` [tool](https://github.com/elasticsearch-dump/elasticsearch-dump) installed * Aiven for OpenSearch cluster as the `input` * AWS S3 bucket as the `output` Collect the following information about your Aiven for OpenSearch cluster and your AWS service: Aiven for OpenSearch: * `SERVICE_URI`: OpenSearch service URI, see it in the Aiven dashboard. * `INPUT_INDEX_NAME`: the index that you aim to copy from your input source. AWS S3: * AWS credentials (`ACCESS_KEY_ID` and `SECRET_ACCESS_KEY`). * S3 file path. This includes bucket name and file name, for for example, `s3://${BUCKET_NAME}/${FILE_NAME}.json` For more information about AWS credentials, see the [AWS documentation](https://docs.aws.amazon.com/general/latest/gr/aws-sec-cred-types). ## Export OpenSearch index data to S3[​](#export-opensearch-index-data-to-s3 "Direct link to Export OpenSearch index data to S3") Use `elasticsearch-dump` command to copy the data from your **Aiven for OpenSearch cluster** to your **AWS S3 bucket**. Use your Aiven for OpenSearch `SERVICE_URI` for the `input`. For the `output`, choose an AWS S3 file path including the file name that you want for your document. ``` elasticdump \ --s3AccessKeyId "${ACCESS_KEY_ID}" \ --s3SecretAccessKey "${SECRET_ACCESS_KEY}" \ --input=SERVICE_URI/INPUT_INDEX_NAME --output "s3://${BUCKET_NAME}/${FILE_NAME}.json" ``` --- # Integrate with Grafana® You can monitor and set up alerts for the data in your Aiven for OpenSearch® service with Grafana®. This feature is especially powerful if you're sending your Aiven service logs to an OpenSearch instance using [log integration](/docs/products/opensearch/howto/opensearch-log-integration.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") 1. Aiven for OpenSearch service. 2. Aiven for Grafana service, see how to [get started with Aiven for Grafana](/docs/products/grafana/get-started.md). ## Variables[​](#variables "Direct link to Variables") We'll use these values later in the set up. They can be found in your Aiven for OpenSearch service page, in the connection information. | Variable | Description | | --------------------- | --------------------------------------- | | `OPENSEARCH_URI` | Service URI of your OpenSearch service. | | `OPENSEARCH_USER` | Username to access OpenSearch service. | | `OPENSEARCH_PASSWORD` | Password to access OpenSearch service. | ## Integration steps[​](#integration-steps "Direct link to Integration steps") 1. [Log in to Aiven for Grafana](/docs/products/grafana/get-started.md#log-in-to-grafana). 2. In **Configuration menu**, select **Data sources**. 3. Click to **Add data source**. 4. Find **OpenSearch** in the list and select it. You'll see a panel with list of settings to fill in. 5. Use your preferred in the *Name* field. You'll use it later for creating dashboards and alerts. 6. Set *URL* to `OPENSEARCH_URI`. 7. In *Auth* section enable **Basic auth** and **With Credentials**. 8. In *Basic Auth Details* set your `OPENSEARCH_USER` and `OPENSEARCH_PASSWORD`. 9. Scroll down to *OpenSearch details* and set the index name or an index pattern (for example, `logs-*`). 10. Set the time field name (in case you use [the log integration](/docs/products/opensearch/howto/opensearch-log-integration.md) it will be `timestamp`). 11. Press on **Save & test**. In case of errors, verify that the data source information is set correctly. ## Create dashboards and alerts[​](#create-dashboards-and-alerts "Direct link to Create dashboards and alerts") Using the interface of Grafana, you can now create dashboards and alerts. Select **Create** from the menu on the left and select to create either a dashboard or an alert rule. --- # Enable JSON Web Token authentication on Aiven for OpenSearch® Configure JSON Web Token (JWT) authentication to enable secure, stateless authentication for Aiven for OpenSearch®. ## How it works[​](#how-it-works "Direct link to How it works") JWT authentication allows you to access Aiven for OpenSearch using tokens issued by your existing identity provider, eliminating the need to manage separate OpenSearch credentials or store session state on the server. When you make a request to your Aiven for OpenSearch service: 1. **Token is validated**: Aiven for OpenSearch validates the JWT token using the signing key you configure. 2. **User is identified**: Aiven for OpenSearch extracts the username from the token's subject claim. 3. **Role is assigned**: If configured, user roles are extracted from the token or managed separately in Aiven for OpenSearch. 4. **Access is granted**: Valid tokens provide seamless access to your Aiven for OpenSearch service. JWT authentication in Aiven for OpenSearch uses the known-signing-keys validation method, where Aiven for OpenSearch uses a trusted public key to verify a digital signature from your identity provider, ensuring that data is coming from a known and authentic source. JWT authentication in Aiven for OpenSearch is configured through Aiven's user configuration API rather than directly through the OpenSearch Security API. Supported `user_config` options are as follows: | Option | Data type | More information | | -------------------------------------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------- | | `jwt.enabled` | Boolean | Enables or disables JWT authentication. Set to `true` to activate JWT authentication. | | `jwt.signing_key` | String | Base64-encoded signing key used to verify JWT tokens. Can be PEM-formatted RSA/ECDSA public key or HMAC secret key. | | `jwt.jwt_header` | String | HTTP header name containing the JWT token. Default is `Authorization`. | | `jwt.jwt_url_parameter` | String | URL parameter name for passing JWT token as a query parameter. Optional alternative to header-based authentication. | | `jwt.subject_key` | String | JWT claim key that contains the username/subject. Default is `sub`. | | `jwt.roles_key` | String | JWT claim key that contains user roles. If not specified, roles must be managed separately in OpenSearch. | | `jwt.required_audience` | String | Required audience (`aud`) claim value that must be present in JWT tokens. Optional but recommended for security. | | `jwt.required_issuer` | String | Required issuer (`iss`) claim value that must be present in JWT tokens. Optional but recommended for security. | | `jwt.jwt_clock_skew_tolerance_seconds` | Integer | Clock skew tolerance in seconds for JWT token validation. Accounts for time differences between systems. Default is typically 30 seconds. | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven for OpenSearch® version 2.4 or later * OpenSearch Dashboards version 2.19 or later * OpenSearch Security management [enabled](/docs/products/opensearch/howto/enable-opensearch-security.md) on your service * Base64-encoded signing key (PEM-formatted RSA/ECDSA public key or HMAC secret key) * Tool for enabling and configuring the JWT authentication: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) ## Enable JWT authentication[​](#enable-jwt-authentication "Direct link to Enable JWT authentication") * Console * CLI * API 1. In the [Aiven Console](https://console.aiven.io/), access your Aiven for OpenSearch service where to enable the JWT authentication. 2. Click **Users** in the sidebar. 3. In the **SSO authentication** section, click **Add method** > **JWT**. 4. In the **Configure JWT authentication** window, set up the following: * **Signing algorithm**: Choose `RSA/ECDSA` or `HMAC`. * **Signing key**: Enter your public key to verify your JWT signature when using RSA/ECDSA. * **HTTP header name**: Provide it if your JWT is transmitted as an HTTP header. * **URL parameter name**: Provide it if your JWT is transmitted as a URL parameter. * **JWT claim key for subject**: Enter the JWT payload key that contains the user's subject identifier to override the `sub` default. * **JWT claim key for roles**: Enter the JWT payload key that contains the user's roles to have them extracted from the JWT for authorization. * **Required JWT audience**: Provide a value for the `aud` claim in the JWT to restrict its audience. * **Required JWT issuer**: Provide a value for the `iss` claim in the JWT to restrict its issuer. * **JWT Clock Skew Tolerance (seconds)**: Specify the maximum time difference between the JWT's issuer's clock and the OpenSearch server's clock. 5. Click **Enable** to complete the setup and activate the configuration. Run [avn service update](/docs/tools/cli/service-cli.md#avn-cli-service-update): ``` avn service update SERVICE_NAME \ -c jwt.enabled=true \ -c jwt.signing_key='SIGNING_KEY' \ -c jwt.jwt_url_parameter=JWT_URL_PARAMETER ``` Replace the following placeholders with your data: * `SERVICE_NAME` with the name of your Aiven for OpenSearch service, for example, `os2-jwt` * `SIGNING_KEY` with your base64-encoded signing key (PEM-formatted RSA/ECDSA public key or HMAC secret key) * `JWT_URL_PARAMETER` with the URL parameter name for JWT token, for example, `token` Call the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint: ``` curl 'https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME' \ -X 'PUT' \ -H 'authorization: aivenv1 BEARER_TOKEN' \ -H 'content-type: application/json' \ --data-raw '{ "user_config": { "jwt": { "enabled": true, "signing_key": "SIGNING_KEY", "jwt_clock_skew_tolerance_seconds": 45, "jwt_url_parameter": "JWT_URL_PARAMETER" } } }' ``` Replace the following placeholders with meaningful data: * `PROJECT_NAME` with your project name, for example, `my-project` * `SERVICE_NAME` with your Aiven for OpenSearch service name, for example, `os2-jwt` * `BEARER_TOKEN` with your Aiven API authentication token * `SIGNING_KE`Y with your base64-encoded signing key (PEM-formatted RSA/ECDSA public key or HMAC secret key) * `JWT_URL_PARAMETER` with the URL parameter name for JWT token, for example, `token` Related pages [Upstream OpenSearch JWT authentication documentation](https://docs.opensearch.org/latest/security/authentication-backends/jwt/) --- # Connect to Aiven for OpenSearch® Connect to the Aiven for OpenSearch® service using various programming languages or tools. ## [cURL](/docs/products/opensearch/howto/opensearch-with-curl.md) [Connect to your Aiven for OpenSearch® service with cURL.](/docs/products/opensearch/howto/opensearch-with-curl.md) ## [NodeJS](/docs/products/opensearch/howto/connect-with-nodejs.md) [The most convenient way to work with the cluster when using NodeJS is to rely on OpenSearch® JavaScript client.](/docs/products/opensearch/howto/connect-with-nodejs.md) ## [Python](/docs/products/opensearch/howto/connect-with-python.md) [You can interact with your cluster with the help of the Python OpenSearch® client.](/docs/products/opensearch/howto/connect-with-python.md) ## [Upgrade ES clients to OS](/docs/products/opensearch/howto/upgrade-clients-to-opensearch.md) [Elasticsearch has introduced breaking changes into their client libraries as early as 7.13.\\\*, meaning newer Elasticsearch clients won't work with OpenSearch®.](/docs/products/opensearch/howto/upgrade-clients-to-opensearch.md) --- # OpenSearch® Security management in Aiven for OpenSearch® Using OpenSearch Security can significantly strengthen the security of your service. By enabling and leveraging this feature for your Aiven for OpenSearch service, you gain access to a wide range of advanced functionalities that will allow you to manage security and access control effectively. ## [OS Security overview](/docs/products/opensearch/concepts/os-security.md) [OpenSearch Security is a powerful feature that enhances the security of](/docs/products/opensearch/concepts/os-security.md) ## [Prepare for OS Security](/docs/products/opensearch/concepts/opensearch-security-considerations.md) [Before enabling OpenSearch Security management for Aiven for OpenSearch service, understand its impact on your current system and adapt your infrastructure accordingly.](/docs/products/opensearch/concepts/opensearch-security-considerations.md) ## [Enable OS Security](/docs/products/opensearch/howto/enable-opensearch-security.md) [OpenSearch Security provides a range of security features, including fine-grained access controls, SAML authentication, and audit logging to monitor activity within your Aiven for OpenSearch® service.](/docs/products/opensearch/howto/enable-opensearch-security.md) ## [SAML authentication](/docs/products/opensearch/howto/saml-sso-authentication.md) [SAML (Security Assertion Markup Language) is a standard protocol for exchanging authentication and authorization data between an identity provider (IdP) and a Service Provider (SP).](/docs/products/opensearch/howto/saml-sso-authentication.md) ## [OIDC authentication](/docs/products/opensearch/howto/oidc-authentication.md) [OpenID Connect (OIDC) is an authentication protocol that builds on top of the OAuth 2.0 protocol.](/docs/products/opensearch/howto/oidc-authentication.md) ## [JWT authentication](/docs/products/opensearch/howto/jwt-authentication.md) [Configure JSON Web Token (JWT) authentication to enable secure, stateless authentication for Aiven for OpenSearch®.](/docs/products/opensearch/howto/jwt-authentication.md) ## [Manage OS audit logs](/docs/products/opensearch/howto/audit-logs.md) [Aiven for OpenSearch® enables audit logging functionality via the OpenSearch Security dashboard, which allows OpenSearch Security administrators to track system events, security-related events, and user activity.](/docs/products/opensearch/howto/audit-logs.md) --- # Search and aggregations with Aiven for OpenSearch® Learn how to write and execute search queries and aggregate data using OpenSearch clients in two widely used programming languages: Python and NodeJS. ## [Aggregations overview](/docs/products/opensearch/concepts/aggregations.md) [Alongside the search functionality, OpenSearch® offers a powerful](/docs/products/opensearch/concepts/aggregations.md) ## [Search queries with Python](/docs/products/opensearch/howto/opensearch-search-and-python.md) [Learn how to write and run search queries on your OpenSearch cluster using a Python OpenSearch client.](/docs/products/opensearch/howto/opensearch-search-and-python.md) ## [Search queries with NodeJS](/docs/products/opensearch/howto/opensearch-and-nodejs.md) [Learn how the OpenSearch® JavaScript client gives a clear and useful interface to communicate with an OpenSearch cluster and run search queries.](/docs/products/opensearch/howto/opensearch-and-nodejs.md) ## [Aggregations with NodeJS](/docs/products/opensearch/howto/opensearch-aggregations-and-nodejs.md) [Learn how to aggregate data using OpenSearch and its NodeJS client.](/docs/products/opensearch/howto/opensearch-aggregations-and-nodejs.md) ## [Custom dictionary files](/docs/products/opensearch/howto/custom-dictionary-files.md) [Custom dictionary files are user-defined files that enhance query analysis and improve search relevance in OpenSearch. By adding domain-specific vocabulary and rules, these files refine search results to be more accurate and relevant.](/docs/products/opensearch/howto/custom-dictionary-files.md) ## [Enable slow query logs](/docs/products/opensearch/howto/enable-slow-query-log.md) [Identify inefficient or time-consuming queries by enabling slow query logging in your Aiven for OpenSearch® service.](/docs/products/opensearch/howto/enable-slow-query-log.md) --- # Maintenance and updates for your Aiven for OpenSearch® service Manage maintenance updates and set the maintenance window for your Aiven for OpenSearch® service. ## Maintenance updates[​](#maintenance-updates "Direct link to Maintenance updates") Aiven applies some maintenance updates automatically. The following are the types of updates: * **Mandatory updates:** Security updates, quarterly patch releases, and platform updates that affect reliability or stability of the service nodes. * **Optional updates:** All other updates are initially optional. After six months, they become mandatory and are applied in the next week’s maintenance window. * **Periodic infrastructure updates:** Scheduled automatically for services with nodes active for 180 days and more. These updates are mandatory for all services, except those with maintenance turned off. Critical security updates are applied during the next available maintenance window. For other updates, Aiven gives you at least seven days' notice. Maintenance updates are also automatically applied during service upgrades. To view pending updates: * Console * CLI * API 1. In your service, click **Service settings**. 2. Go to the **Service management** section. Use the [`avn service get`](/docs/tools/cli/service-cli.md#avn_service_get) command. Use the [`service`](https://api.aiven.io/doc/#tag/Service/operation/ServiceGet) endpoint. ## Maintenance window[​](#maintenance-window "Direct link to Maintenance window") The maintenance window is the time period when Aiven can automatically apply maintenance updates to a service. When an update becomes available, Aiven schedules it for the next available maintenance window for each service. The update runs in the first window after it becomes available, and can begin any time after the start time. For example, if a service has a maintenance window of Monday 12:00 UTC, and an update becomes available on Tuesday, the update will be applied on the following Monday. During maintenance, Aiven might restart or replace service nodes. This can cause brief connection interruptions, but services are designed to minimize downtime. Aiven performs maintenance in a rolling-forward style, creating new nodes alongside existing ones and retiring the old nodes after the upgrade completes. Major service upgrades are triggered manually. A manually triggered upgrade starts immediately, regardless of the maintenance window. important You cannot control the order in which services are updated. Each service updates according to its own configured maintenance window, and there is no guaranteed way to control the update sequence. Manual updates and maintenance window adjustments only help for non-critical updates. ## Set the maintenance window[​](#set-the-maintenance-window "Direct link to Set the maintenance window") To set the maintenance window for your service: * Console * Terraform 1. In the Aiven Console, open your service. 2. In the **Maintenance** section, click **Actions** > **Change maintenance window**. 3. Set the day and time. 4. Click **Save changes**. Use the `maintenance_window_dow` and `maintenance_window_time` attributes in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Version upgrades](/docs/products/opensearch/howto/os-version-upgrade.md) * [Change the service plan](/docs/products/opensearch/howto/change-service-plan.md) --- # Manage Aiven for OpenSearch® custom repositories in OpenSearch® API Use the OpenSearch® API for configuring custom repositories in Aiven for OpenSearch to store [snapshots](/docs/products/opensearch/howto/manage-snapshots.md) in your cloud storage. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Maintenance updates applied for your service * [Security management enabled](/docs/products/opensearch/howto/enable-opensearch-security.md) for your service * [Snapshot permissions](https://docs.opensearch.org/docs/latest/security/access-control/permissions/#snapshot-permissions) and [snapshot repository permissions](https://docs.opensearch.org/docs/latest/security/access-control/permissions/#snapshot-repository-permissions) configured * [Storage credentials](/docs/products/opensearch/howto/snapshot-credentials.md) ## Limitations[​](#limitations "Direct link to Limitations") * Supported storage services * Amazon S3 or any S3-compatible object storage service * Google Cloud Storage (GCS) * Microsoft Azure Blob Storage * The following operations are not supported via native OpenSearch API: * [Remove a repository](/docs/products/opensearch/howto/custom-repositories.md#remove-a-repository) * [Edit repository details](/docs/products/opensearch/howto/custom-repositories.md#view-or-edit-repository-details) * [List custom repositories](/docs/products/opensearch/howto/custom-repositories.md#list-custom-repositories) * Using the native OpenSearch API requires providing [storage credentials](/docs/products/opensearch/howto/snapshot-credentials.md). ## Register custom repositories[​](#register-custom-repositories "Direct link to Register custom repositories") Each repository requires a unique name, a storage type (such as S3, Azure, or GCS), and the appropriate settings for the selected storage provider. Use the [Register Snapshot Repository](https://docs.opensearch.org/docs/latest/api-reference/snapshots/create-repository/) native OpenSearch API endpoint. ## View repository details[​](#view-repository-details "Direct link to View repository details") To view details on a repository, use the [Get Snapshot Repository](https://docs.opensearch.org/docs/latest/api-reference/snapshots/get-snapshot-repository/) native OpenSearch API endpoint. ## Error handling[​](#error-handling "Direct link to Error handling") The Aiven API returns OpenSearch errors as they are. **Exceptions:** * 502: OpenSearch did not respond. * 409: The service is not powered on or does not support this feature. Related pages [OpenSearch snapshot API reference](https://opensearch.org/docs/latest/api-reference/snapshots/index/) --- # Manage Aiven for OpenSearch® custom repositories Set up custom repositories in Aiven for OpenSearch® in the Aiven Console, the Aiven API, or the native OpenSearch API. tip Unless you have a good reason to use the native OpenSearch API, choose either the Aiven Console or the Aiven API for speed and simplicity. ## [In Aiven Console or API](/docs/products/opensearch/howto/custom-repositories.md) [Use the Aiven Console or API for configuring custom repositories in Aiven for OpenSearch to store snapshots in your cloud storage.](/docs/products/opensearch/howto/custom-repositories.md) ## [In OpenSearch® API](/docs/products/opensearch/howto/manage-custom-repo/custom-repositories-os-api.md) [Use the OpenSearch® API for configuring custom repositories in Aiven for OpenSearch to store snapshots in your cloud storage.](/docs/products/opensearch/howto/manage-custom-repo/custom-repositories-os-api.md) ## [Manage credentials](/docs/products/opensearch/howto/snapshot-credentials.md) [Use custom\_keystores in Aiven for OpenSearch® to store object storage credentials in Amazon S3, Google Cloud Storage, or Azure.](/docs/products/opensearch/howto/snapshot-credentials.md) --- # Create and manage snapshots in Aiven for OpenSearch® Create, list, retrieve, or delete snapshots in your Aiven for OpenSearch [custom repositories](/docs/products/opensearch/howto/custom-repositories.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven Console * Aiven API * OpenSearch API - Running Aiven for OpenSearch service - Access to the [Aiven Console](https://console.aiven.io/) - Access to a supported object storage service (AWS S3, GCS, or Azure) - Credentials for the selected storage provider - Configured [custom repository](/docs/products/opensearch/howto/custom-repositories.md) * Running Aiven for OpenSearch service * [Aiven API](/docs/tools/api.md) and authentication [token](/docs/platform/howto/create_authentication_token.md) * Access to a supported object storage service (AWS S3, GCS, or Azure) * Credentials for the selected storage provider * Configured [custom repository](/docs/products/opensearch/howto/custom-repositories.md) - Configured [custom repository](/docs/products/opensearch/howto/manage-custom-repo/list-manage-custom-repo.md) - Maintenance updates applied for your service - [Security management enabled](/docs/products/opensearch/howto/enable-opensearch-security.md) for your service - [Snapshot permissions](https://docs.opensearch.org/docs/latest/security/access-control/permissions/#snapshot-permissions) and [snapshot repository permissions](https://docs.opensearch.org/docs/latest/security/access-control/permissions/#snapshot-repository-permissions) configured important Automatic snapshot scheduling is not supported. Create, list, and delete snapshots manually. ## Limitations[​](#limitations "Direct link to Limitations") See [Aiven for OpenSearch limits and limitations](/docs/products/opensearch/reference/opensearch-limitations.md#snapshot-management) for the tool-agnostic list of snapshot management restrictions. * Aiven Console * Aiven API * OpenSearch API Only the following storage services are supported: * Amazon S3 * Google Cloud Storage (GCS) * Microsoft Azure Blob Storage - You cannot access the `/_snapshot` API directly without configuring [custom repositories](/docs/products/opensearch/howto/custom-repositories.md). - Only the following storage services are supported: * Amazon S3 * Google Cloud Storage (GCS) * Microsoft Azure Blob Storage See [Aiven for OpenSearch limits and limitations](/docs/products/opensearch/reference/opensearch-limitations.md#api-restrictions) for more Aiven for OpenSearch API restrictions. * You cannot access the `/_snapshot` API directly without configuring [custom repositories](/docs/products/opensearch/howto/custom-repositories.md). * Only the following storage services are supported: * Amazon S3 * Google Cloud Storage (GCS) * Microsoft Azure Blob Storage * [Restore from snapshot](/docs/products/opensearch/howto/manage-snapshots.md#restore-from-snapshots) has a couple of [security-related restrictions](https://docs.opensearch.org/latest/tuning-your-cluster/availability-and-recovery/snapshots/snapshot-restore/#security-considerations). * [Create a snapshot](/docs/products/opensearch/howto/manage-snapshots.md#create-a-snapshot) and [delete a snapshot](/docs/products/opensearch/howto/manage-snapshots.md#delete-a-snapshot) are not supported for snapshots in Aiven-managed repositories (prefixed with `aiven_repo`). See [Aiven for OpenSearch limits and limitations](/docs/products/opensearch/reference/opensearch-limitations.md#api-restrictions) for more Aiven for OpenSearch API restrictions. ## Create a snapshot[​](#create-a-snapshot "Direct link to Create a snapshot") Create a snapshot in a custom repository. * Aiven Console * Aiven API * OpenSearch API 1. Log in to the [Aiven Console](https://console.aiven.io/), go to your project, and open your service's page. 2. In the **Backups** section, click **Snapshots**. 3. On the **Snapshots** page, do one of the following: * Click **Create snapshot**. * Find your custom repository and click **Actions** > **Create snapshot**. 4. In the **Create snapshot** window: 1. Select a destination repository. 2. Enter a snapshot name. 3. Specify which indices to include. 4. Optionally, enable the following: * **Ignore unavailable indices** * **Include global state** * **Partial snapshot** 5. Click **Create**. Your snapshot is being created. Monitor its status until it shows **Success**. ``` curl -s -X POST \ --url "https://api.aiven.io/v1/project/{project_name}/service/{service_name}/opensearch/_snapshot/{repository_name}/{snapshot_name}/_restore" \ --header "Authorization: Bearer $TOKEN" \ --header "Content-Type: application/json" \ -d '{"indices": "test*"}' ``` Example response: ``` { "accepted": true } ``` Use the [Create Snapshot](https://docs.opensearch.org/docs/latest/api-reference/snapshots/create-snapshot/) native OpenSearch API endpoint. ## Restore from snapshots[​](#restore-from-snapshots "Direct link to Restore from snapshots") important Refrain from actions such as updating firewalls, changing index settings, or modifying security configurations during the restore process as it can cause restore failures. * Aiven Console * Aiven API * OpenSearch API 1. Log in to the [Aiven Console](https://console.aiven.io/), go to your project, and open your service's page. 2. In the **Backups** section, click **Snapshots**. 3. On the **Snapshots** page, find your custom repository and click to expand the list of snapshots inside. 4. Find the snapshot to restore and click **Actions** > **Restore to this service**. 5. In the **Restore snapshot** window: 1. In the **Indices** field, enter the indices to include in the snapshot, separated by commas. These indices will be closed automatically before the restore begins. Existing data may be overwritten as a result of the restoration process. To prevent this, use **advanced configuration** to add suffixes to the included indices. 2. Toggle **Enable advanced configuration** to set the following: * Rename pattern: a regular expression that matches the original index name from the snapshot * Rename replacement: a string that replaces the matched part of the index name * Ignore unavailable indices * Include aliases 3. Click **Continue** > **Close indices**. This triggers the closing of the selected indices. Wait for **Indices closed** to be displayed, and click **Continue**. 4. Check the box labeled **I understand the effects of this action**. warning The restoration process you're about to start cannot be interrupted. Refrain from changes during this process: Updating firewalls, index settings, or security configuration during restore may cause failures. 5. Click **Start restore**. This triggers the restoration process. It may take time, and it's length depends on the snapshot size. 6. Shut down the **Restore snapshot** window by clicking **Close** either during the restore process or when it completes. tip If you close the **Restore snapshot** window before the restore process is complete, you can preview its status on the **Snapshots** page. ``` curl -s -X POST \ "https://api.aiven.io/v1/project/{project_name}/service/{service_name}/opensearch/_snapshot/{repository_name}/{snapshot_name}/_restore" \ -H "Authorization: Bearer $TOKEN" \ -H "Content-Type: application/json" \ -d '{ "indices": "test-index-*", "include_aliases": true, "ignore_unavailable": false, "rename_pattern": "index_(.+)", "rename_replacement": "restored_index_$1" }' ``` Example response: ``` { "accepted": true } ``` To restore data from a snapshot, use the [Restore Snapshot](https://docs.opensearch.org/docs/latest/api-reference/snapshots/restore-snapshot/) native OpenSearch API endpoint. ## List snapshots in progress[​](#list-snapshots-in-progress "Direct link to List snapshots in progress") Preview snapshots that are still being created in a repository. * Aiven Console * Aiven API * OpenSearch API 1. Log in to the [Aiven Console](https://console.aiven.io/), go to your project, and open your service's page. 2. In the **Backups** section, click **Snapshots**. 3. On the **Snapshots** page, find your custom repository and click to expand the list of snapshots inside. 4. In the column with snapshot status information, find **in progress** values. ``` curl -s --url "https://api.aiven.io/v1/project/{project_name}/service/{service_name}/opensearch/_snapshot/aws-repo/_status" \ --header "Authorization: Bearer $TOKEN" --header "Content-Type: application/json" ``` Example response: ``` { "snapshots": [ { "repository": "aws-repo", "snapshot": "second-snapshot", "state": "SUCCESS", "shards_stats": { "done": 1, "failed": 0, "total": 1 }, "uuid": "osmCbdF-RMyyUKpWD-4bJA" } ] } ``` Use one of the following variants of the native OpenSearch [Get Snapshot Status](https://docs.opensearch.org/latest/api-reference/snapshots/get-snapshot-status/) API: * For snapshot statuses in a specified repository: `GET /_snapshot/REPOSITORY_NAME/_status` * For snapshot statuses in all repositories: `GET /_snapshot/_status` ## List snapshots in a repository[​](#list-snapshots-in-a-repository "Direct link to List snapshots in a repository") Preview all snapshots, including completed and failed ones. * Aiven Console * Aiven API * OpenSearch API 1. Log in to the [Aiven Console](https://console.aiven.io/), go to your project, and open your service's page. 2. In the **Backups** section, click **Snapshots**. 3. On the **Snapshots** page, find your custom repository and click to expand the list of snapshots inside. ``` curl -s --url "https://api.aiven.io/v1/project/{project_name}/service/{service_name}/opensearch/_snapshot/aws-repo/_all" \ --header "Authorization: Bearer $TOKEN" --header "Content-Type: application/json" ``` Example response: ``` { "snapshots": [ { "snapshot": "first-snapshot", "state": "SUCCESS", "indices": ["test"], "uuid": "7cdWedW7RC6FMSktlZTCDw" }, { "snapshot": "second-snapshot", "state": "SUCCESS", "indices": ["test"], "uuid": "osmCbdF-RMyyUKpWD-4bJA" } ] } ``` Use native OpenSearch API endpoint `GET /_snapshot/REPOSITORY_NAME/_all`, replacing the `REPOSITORY_NAME` with the actual name of your repository. Example response: ``` { "snapshots" : [ { "snapshot" : "opensearch-123qaz456wsx789edc123q-infrequent", "uuid" : "-abCabCabC-abCabCabCab", "version_id" : 123456789, "version" : "N.NN.N", "remote_store_index_shallow_copy" : false, "indices" : [ ".plugins-ml-config", ".opensearch-sap-log-types-config", ".kibana_N", ".opensearch-observability" ], "data_streams" : [ ], "include_global_state" : true, "state" : "SUCCESS", "start_time" : "YYY-MM-DDTHH:MM:SS.619Z", "start_time_in_millis" : 1234567891234, "end_time" : "YYY-MM-DDTHH:MM:SS.624Z", "end_time_in_millis" : 1234567891234, "duration_in_millis" : 1234, "failures" : [ ], "shards" : { "total" : 4, "failed" : 0, "successful" : 4 } }, { "snapshot" : "opensearch-123qaz456wsx789edc123q-frequent", "uuid" : "-abCabCabC-abCabCabCab", "version_id" : 123456789, "version" : "N.NN.N", "remote_store_index_shallow_copy" : false, "indices" : [ ".plugins-ml-config", ".opensearch-sap-log-types-config", ".kibana_N", ".opensearch-observability" ], "data_streams" : [ ], "include_global_state" : true, "state" : "SUCCESS", "start_time" : "YYY-MM-DDTHH:MM:SS.219Z", "start_time_in_millis" : 12345678912345, "end_time" : "YYY-MM-DDTHH:MM:SS.220Z", "end_time_in_millis" : 1234567891234, "duration_in_millis" : 1234, "failures" : [ ], "shards" : { "total" : 4, "failed" : 0, "successful" : 4 } }, { "snapshot" : "opensearch-123qaz456wsx789edc123q-frequent", "uuid" : "-abCabCabC-abCabCabCabQ", "version_id" : 123456789, "version" : "N.NN.N", "remote_store_index_shallow_copy" : false, "indices" : [ ".plugins-ml-config", ".opensearch-sap-log-types-config", ".kibana_N", ".opensearch-observability" ], "data_streams" : [ ], "include_global_state" : true, "state" : "SUCCESS", "start_time" : "YYY-MM-DDTHH:MM:SS.088Z", "start_time_in_millis" : 12345678912345, "end_time" : "YYY-MM-DDTHH:MM:SS.890Z", "end_time_in_millis" : 1234567891234, "duration_in_millis" : 1234, "failures" : [ ], "shards" : { "total" : 4, "failed" : 0, "successful" : 4 } } ] } ``` ## View snapshot details[​](#view-snapshot-details "Direct link to View snapshot details") Get details of a specific snapshot. * Aiven Console * Aiven API * OpenSearch API 1. Log in to the [Aiven Console](https://console.aiven.io/), go to your project, and open your service's page. 2. In the **Backups** section, click **Snapshots**. 3. On the **Snapshots** page, find your custom repository, click to expand the list of snapshots inside, find a snapshot to be previewed, and click **Actions** > **View snapshot details**. ``` curl -s --url "https://api.aiven.io/v1/project/{project_name}/service/{service_name}/opensearch/_snapshot/aws-repo/first-snapshot" \ --header "Authorization: Bearer $TOKEN" --header "Content-Type: application/json" ``` Example response: ``` { "snapshots": [ { "snapshot": "first-snapshot", "state": "SUCCESS", "indices": ["test"], "uuid": "7cdWedW7RC6FMSktlZTCDw" } ] } ``` Use the native OpenSearch API endpoints: [Get Snapshot](https://docs.opensearch.org/docs/latest/api-reference/snapshots/get-snapshot/) or [Get Snapshot Status](https://docs.opensearch.org/docs/latest/api-reference/snapshots/get-snapshot-status/). ## Delete a snapshot[​](#delete-a-snapshot "Direct link to Delete a snapshot") Delete a snapshot from a repository. * Aiven Console * Aiven API * OpenSearch API 1. Log in to the [Aiven Console](https://console.aiven.io/), go to your project, and open your service's page. 2. In the **Backups** section, click **Snapshots**. 3. On the **Snapshots** page, find your custom repository, click to expand the list of snapshots inside, find a snapshot to be deleted, click **Actions** > **Delete snapshot** > **Delete**. ``` curl -s -X DELETE \ --url "https://api.aiven.io/v1/project/{project_name}/service/{service_name}/opensearch/_snapshot/aws-repo/first-snapshot" \ --header "Authorization: Bearer $TOKEN" --header "Content-Type: application/json" ``` Example response: ``` { "acknowledged": true } ``` Use the [Delete Snapshot](https://docs.opensearch.org/docs/latest/api-reference/snapshots/delete-snapshot/) native OpenSearch API endpoint. ## Error handling[​](#error-handling "Direct link to Error handling") The Aiven API returns OpenSearch errors as they are. **Exceptions:** * 502: OpenSearch did not respond. * 409: The service is not powered on or does not support this feature. Related pages [OpenSearch snapshot API reference](https://opensearch.org/docs/latest/api-reference/snapshots/index/) --- # Migrate external OpenSearch or Elasticsearch snapshots to Aiven Migrate an existing OpenSearch or Elasticsearch® snapshot to Aiven for OpenSearch® with minimal downtime and data integrity. The migration process uses [custom repositories](/docs/products/opensearch/howto/manage-custom-repo/list-manage-custom-repo.md) and consists of the following phases: 1. [Configure a custom repository](/docs/products/opensearch/howto/manage-custom-repo/list-manage-custom-repo.md) where your migrated data will reside. 2. [Migrate the data](/docs/products/opensearch/howto/manage-snapshots.md#restore-from-snapshots) by restoring it from external snapshots stored on supported platforms like Amazon S3, Google Cloud Storage (GCS), Microsoft Azure, or other S3-compatible services. 3. Optional: [Reapply Index State Management (ISM) policies](/docs/products/opensearch/howto/migrate-ism-policies.md): Use a [script](https://github.com/aiven/aiven-examples/blob/main/solutions/reapply-ism-policies/avn-re-apply-ism-policies.py) to migrate ISM policies to maintain consistent index lifecycle management, including tasks like index rollover, retention, and deletion. 4. Optional: [Recreate security configuration](/docs/products/opensearch/howto/migrate-opendistro-security-config-aiven.md): Use a [script](https://github.com/aiven/aiven-examples/blob/main/solutions/migrate-opendistro-security-to-aiven-for-opensearch/avn-migrate-os-security-config.py) to migrate user roles, permissions, and access controls to preserve security settings and ensure a smooth user experience after migration. --- # Reapply ISM policies after snapshot restore Reapply Index State Management (ISM) policies to Aiven for OpenSearch® using a script. After restoring your snapshot, ISM policies that manage index rollover, retention, and deletion must be reapplied to your indices. These policies are stored in the `.opendistro-ism-config` index, but the assignments between indices and policies must be reapplied using a script. ## What is restored[​](#what-is-restored "Direct link to What is restored") The `.opendistro-ism-config` index stores ISM policy configurations and is restored with the snapshot, but policy assignments to specific indices are stored in the cluster metadata and must be reapplied. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A machine with network access to Aiven for OpenSearch services * Python 3.11 or higher installed * Ensure all data indices and the `.opendistro-ism-config` index are restored from the snapshot * Take the snapshot with `restore_global_state: true`. By default, this setting is `false`, and you must enable it to restore ISM policy assignments. warning * **Snapshot must include global state**: Ensure the snapshot is **created** with `restore_global_state: true`. If the snapshot was created without the global state, the ISM policy assignments will not be available in the cluster metadata, and the script will fail to reapply the policies. * **Script can only be run once**: After ISM policies are applied, OpenSearch clears the cluster metadata that holds ISM policy assignments. This means you can only run the script once. ## Validate index sync before reapplying ISM policies[​](#validate-index-sync-before-reapplying-ism-policies "Direct link to Validate index sync before reapplying ISM policies") Before reapplying ISM policies, ensure the indices are synchronized between the source and target services. Check document counts to confirm they match. For more details, see the [verify the migration](/docs/products/opensearch/howto/manage-snapshots.md#restore-from-snapshots) section in [Migrate data to Aiven for OpenSearch® using snapshots](/docs/products/opensearch/howto/manage-snapshots.md#restore-from-snapshots). ## Reapply ISM policies[​](#reapply-ism-policies "Direct link to Reapply ISM policies") The script retrieves the ISM policy assignments stored in the cluster state and reapplies them to the corresponding indices. warning **Potential ISM policy restore error** Restoring certain ISM policies may cause an error, such as when a rollover index (for example, `x-001` to `x-002`) already exists from a previous snapshot reload. This error can prevent secondary ISM policies, like deleting index `x-001`, from being applied. Ensure indices from previous rollovers or other similar situations are handled correctly before reapplying ISM policies. To reapply ISM policies to indices in Aiven for OpenSearch: 1. Download the script from the [Aiven examples GitHub repository](https://github.com/aiven/aiven-examples/blob/main/solutions/reapply-ism-policies/avn-re-apply-ism-policies.py). 2. Create a JSON configuration file with the connection details for your Aiven for OpenSearch service. Use `avnadmin` as the `user` and replace `host`, `port`, and `password` with your service information: ``` { "host": "target-ip-or-fqdn", "port": target-port-number, "user": "avnadmin", "password": "the password" } ``` 3. Once your configuration file is ready, run the script. ``` python avn-re-apply-ism-policies.py --config path-to-config-file ``` ## Re-running the ISM script[​](#re-running-the-ism-script "Direct link to Re-running the ISM script") You can rerun the ISM script if needed. Use the --force option to bypass the check that prevents it from running more than once. note Run the ISM script only after completing all data migration. ## Monitor ISM task progress[​](#monitor-ism-task-progress "Direct link to Monitor ISM task progress") Once ISM policies are reapplied, index lifecycle management tasks like rollovers, retention, and deletion resume automatically. To monitor ISM task progress and ensure policies are enforced correctly, run the following command and replace `SERVICE_URL` with your Aiven for OpenSearch service's URL: ``` curl -X GET --insecure "$SERVICE_URL/_plugins/_ism/explain?pretty&size=100" ``` Alternatively, you can verify the status of individual indices: ``` curl -X GET --insecure "$SERVICE_URL/_plugins/_ism/explain/?pretty" ``` Related pages * [Migrate data to Aiven for OpenSearch® using snapshots](/docs/products/opensearch/howto/manage-snapshots.md#restore-from-snapshots) * [Migrate Opendistro security configuration to Aiven for OpenSearch](/docs/products/opensearch/howto/migrate-opendistro-security-config-aiven.md) --- # Migrate Aiven for OpenSearch® k-NN indices off the nmslib engine Identify indices that use the deprecated `nmslib` k-NN engine, and reindex them to `faiss` or `lucene` before upgrading Aiven for OpenSearch® from version 2.19 to 3.x. ## Why migrate off nmslib[​](#why-migrate-off-nmslib "Direct link to Why migrate off nmslib") Upstream, the k-NN plugin deprecated the `nmslib` engine as of version 2.19.0 and scheduled it for removal. `faiss` is the current default k-NN engine. To give you time to migrate before `nmslib` is removed, Aiven for OpenSearch blocks upgrades from version 2.19 to 3.x for any service with an index that has a `knn_vector` field using `method.engine: "nmslib"`. The upgrade request fails with `403 Forbidden`, and the response lists the affected index names. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * You have an Aiven for OpenSearch service running version 2.19. * You have the service connection credentials. * You have required permissions to create, reindex, and delete indices. note In the examples, * `$OS_URI` is used for the service connection URL (for example, `https://USER:PASSWORD@HOST:PORT`). * `$OLD_INDEX_NAME` is used for the index using the `nmslib` engine. ## Identify indices using nmslib[​](#identify-indices-using-nmslib "Direct link to Identify indices using nmslib") List the indices with a `knn_vector` field whose `method.engine` is `nmslib`: ``` curl -s "$OS_URI/_all/_mapping?filter_path=*.mappings.properties.*.method.engine" | \ jq -r 'to_entries[] as $idx | $idx.value.mappings.properties // {} | to_entries[] | select(.value.method.engine? == "nmslib") | $idx.key' | sort -u ``` ## Choose a destination engine[​](#choose-a-destination-engine "Direct link to Choose a destination engine") | Engine | When to choose | Notes | | -------- | ------------------------------------------------------------ | ----------------------------------------------------------------------------- | | `faiss` | Most workloads; matches the current default engine | Supports the same `hnsw` method as `nmslib`, plus filtering and radial search | | `lucene` | You prefer a pure Java implementation with no native library | Uses Lucene's native filtering | ## Create the destination mapping[​](#create-the-destination-mapping "Direct link to Create the destination mapping") Create an index with the same `knn_vector` field definitions, but with `method.engine` set to `faiss` or `lucene`. Keep the same `method.name` and `space_type` as the source index so search behavior stays consistent. * nmslib to faiss * nmslib to lucene PUT /my-index-v2 ``` { "settings": { "index.knn": true }, "mappings": { "properties": { "embedding": { "type": "knn_vector", "dimension": 768, "method": { "name": "hnsw", "engine": "faiss", "space_type": "l2", "parameters": { "m": 16, "ef_construction": 100 } } } } } } ``` PUT /my-index-v2 ``` { "settings": { "index.knn": true }, "mappings": { "properties": { "embedding": { "type": "knn_vector", "dimension": 768, "method": { "name": "hnsw", "engine": "lucene", "space_type": "cosinesimil", "parameters": { "m": 16, "ef_construction": 100 } } } } } } ``` Adjust `dimension`, `space_type`, and the other field definitions to match your source index mapping. Get the full source mapping with: ``` curl -s "$OS_URI/$OLD_INDEX_NAME/_mapping" ``` ## Reindex and switch over[​](#reindex-and-switch-over "Direct link to Reindex and switch over") Reindex data into the new index, verify the results, and switch your application over to it by following [Reindex Aiven for OpenSearch data on a newer version](/docs/products/opensearch/howto/reindex-opensearch.md#reindex-earlier-version-indices). Exporting settings, running the reindex, verifying document counts, and swapping aliases work the same way for an engine migration. ## Complete the upgrade[​](#complete-the-upgrade "Direct link to Complete the upgrade") After reindexing all indices that use the `nmslib` engine: 1. Confirm no indices remain with `method.engine: "nmslib"` using the query in [Identify indices using nmslib](#identify-indices-using-nmslib). 2. [Upgrade your service](/docs/products/opensearch/howto/os-version-upgrade.md) to OpenSearch 3.x. Related pages * [Upgrade Aiven for OpenSearch](/docs/products/opensearch/howto/os-version-upgrade.md) * [Reindex Aiven for OpenSearch data on a newer version](/docs/products/opensearch/howto/reindex-opensearch.md) * [Available plugins for Aiven for OpenSearch](/docs/products/opensearch/reference/plugins.md) * [OpenSearch® k-NN plugin documentation](https://docs.opensearch.org/latest/search-plugins/knn/knn-index/) --- # Migrate OpenDistro security configuration to Aiven for OpenSearch Migrate your security configuration from an OpenDistro service to Aiven for OpenSearch® using a migration script. The `.opendistro_security` index, which stores security settings, cannot be restored directly from an external snapshot. Instead, you use a migration script that interacts with the security REST API of both services. ## What is migrated[​](#what-is-migrated "Direct link to What is migrated") The migration script transfers the security configuration from the source OpenDistro service to Aiven for OpenSearch. The following configurations are migrated: * Internal users (including passwords) * Action groups * Roles * Backend roles * Tenants The following configurations are **not** migrated: * Reserved, static, or hidden entries * Authentication methods and backend configurations, as these are configured differently in Aiven for OpenSearch ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting the migration, ensure the following: * A machine with network access to OpenDistro and Aiven for OpenSearch services * Python 3.11 or higher installed * [Security management](/docs/products/opensearch/howto/enable-opensearch-security.md) enabled on Aiven for OpenSearch * Admin user certificate and key in PEM format from the source OpenDistro service * OpenDistro security configuration version 2 * HTTPS enabled on both services (the script does not verify server certificates) note * You need the admin user key and certificate to migrate internal users and their passwords. The password hash is fetched directly from the `.opendistro_security` index because the standard API endpoint doesn’t return it. * Ensure HTTPS is enabled on the source and target services, as the script assumes this setup. ## Migration steps[​](#migration-steps "Direct link to Migration steps") Migrate your security configuration from OpenDistro to Aiven for OpenSearch. ### Access the migration script[​](#access-the-migration-script "Direct link to Access the migration script") Find the migration script in the [Aiven examples GitHub repository](https://github.com/aiven/aiven-examples/blob/main/solutions/migrate-opendistro-security-to-aiven-for-opensearch/avn-migrate-os-security-config.py). ### Create the migration configuration file[​](#create-the-migration-configuration-file "Direct link to Create the migration configuration file") Create a JSON configuration file with the connection details for the source (OpenDistro) and target (Aiven for OpenSearch) services. Following is an example: ``` { "source": { "host": "source-ip-or-fqdn", "port": source-port-number, "key": "path-to-admin-key.pem", "certificate": "path-to-admin-cert.pem" }, "target": { "host": "target-ip-or-fqdn", "port": target-port-number, "password": "os-sec-admin-user-password" } } ``` ### Run the migration script[​](#run-the-migration-script "Direct link to Run the migration script") Once your configuration file is ready, execute the migration: ``` python avn-migrate-os-security-config.py --config path-to-config-file ``` The script connects to both services and starts migrating security configurations. It logs which entries (roles, role mappings, action groups, tenants, and internal users) are added or updated. If an entry already exists on the target service, the script skips it and marks it as unchanged. ### Re-running the script[​](#re-running-the-script "Direct link to Re-running the script") You can run the migration script multiple times, but keep the following in mind: * Previously migrated entries are not deleted from the target service. * Internal user data is updated each time the script is run because the OpenDistro API does not return password hashes. As a result, the script treats the data as if it has changed, even when it has not. Related pages * [Migrate data to Aiven for OpenSearch® using snapshots](/docs/products/opensearch/howto/manage-snapshots.md#restore-from-snapshots) * [Reapply ISM policies after snapshot restore](/docs/products/opensearch/howto/migrate-ism-policies.md) --- # Migrate Elasticsearch data to Aiven for OpenSearch® To migrate Elasticsearch data to Aiven for OpenSearch®, reindex from a remote Elasticsearch cluster. This method can also be used to migrate data from Aiven for OpenSearch to a self-hosted Elasticsearch service. tip To migrate a large number of indexes, consider automating the process with a script. As Aiven for OpenSearch does not support joining external Elasticsearch servers to the same cluster, online migration is not currently possible. important Migrating from Elasticsearch to OpenSearch can impact connectivity between client applications and services. Some clients or tools may check the service version, which can lead to compatibility issues with OpenSearch. For more details, refer to the following OpenSearch resources: * [OpenSearch release notes](https://github.com/opensearch-project/OpenSearch/blob/main/release-notes/opensearch.release-notes-1.0.0.md) * [OpenSearch Dashboards release notes](https://github.com/opensearch-project/OpenSearch-Dashboards/blob/main/release-notes/opensearch-dashboards.release-notes-1.0.0.md) * [Frequently asked questions about OpenSearch](https://opensearch.org/faq/) ## Migrate data[​](#migrate-data "Direct link to Migrate data") 1. [Create an Aiven for OpenSearch service](/docs/products/opensearch/get-started.md#create-an-aiven-for-opensearch-service). 2. Set the `reindex.remote.whitelist` parameter to point to your source Elasticsearch service using the following [Aiven CLI](https://github.com/aiven/aiven-client) command: ``` avn service update your-aiven-service \ -c 'opensearch.reindex_remote_whitelist=["your-non-aiven-service:port"]' ``` Replace `port` with the port number your source Elasticsearch service is using. 3. Wait for the cluster to restart. This process might take a few minutes as the service attempts a rolling restart to minimize downtime. 4. Start migrating the indexes. For each index: 1. Stop writes to the index. This step is optional if testing the process. 2. Export the index mapping from the source Elasticsearch instance. For example, using `curl`: ``` curl https://avnadmin:yourpassword@os-123-demoprj.aivencloud.com:23125/logs-2024-09-21/_mapping > mapping.json ``` 3. Edit `mapping.json`: * With jq * Manual update If you have `jq`, run: ``` jq .[].mappings mapping.json > src_mapping.json ``` To edit `mapping.json` manually: * Remove the wrapping `{"logs-2024-09-21":{"mappings": ... }}`. * Keep `{"properties":...}}`. 4. Create the empty index on your destination Aiven for OpenSearch service. ``` curl -XPUT https://avnadmin:yourpassword@os-123-demoprj.aivencloud.com:23125/logs-2024-09-21 ``` 5. Import the mapping to the destination Aiven for OpenSearch index. ``` curl -XPUT https://avnadmin:yourpassword@os-123-demoprj.aivencloud.com:23125/logs-2024-09-21/_mapping \ -H 'Content-type: application/json' -T src_mapping.json ``` 6. Submit the reindexing request. ``` curl -XPOST https://avnadmin:yourpassword@os-123-demoprj.aivencloud.com:23125/_reindex \ -H 'Content-type: application/json' \ -d '{"source": {"index": "logs-2024-09-21", "remote": {"username": "your-remote-username", "password": "your-remote-password", "host": "https://your.non-aiven-service.example.com:9200" } }, "dest": {"index": "logs-2024-09-21"} }' ``` 7. Wait for the reindexing process to complete. If you receive a response message such as: ``` [your.non-aiven-service.example.com:9200] not whitelisted in reindex.remote.whitelist ``` Verify the hostname and port match those set earlier. The time required for reindexing can vary depending on the amount of data. 8. Update clients to use the new index on Aiven for OpenSearch for both read and write operations, then resume any paused write activity. 9. Delete the source index if necessary. --- # Enable OpenID Connect authentication on Aiven for OpenSearch® OpenID Connect (OIDC) is an authentication protocol that builds on top of the OAuth 2.0 protocol. It provides a simple and secure way to verify the identity of a user and obtain basic profile information about them. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven for OpenSearch® version 2 or later is required. If you are using an earlier version, upgrade to the latest version. * OpenSearch Security management must be [enabled](/docs/products/opensearch/howto/enable-opensearch-security.md) on the Aiven for OpenSearch® service. * An OpenID provider (IdP) that supports the OpenID Connect protocol. ## Requirements to enable OpenID Connect[​](#requirements-to-enable-openid-connect "Direct link to Requirements to enable OpenID Connect") To enable OpenID Connect authentication for Aiven for OpenSearch, you must configure OpenID Connect with an Identity Provider (IdP). Aiven for OpenSearch integrates with various OpenID Connect IdPs, and the exact steps to achieve this differ depending on your chosen IdP. Refer to your Identity Provider's official documentation for specific configuration steps. To successfully set up OpenID Connect authentication, the following parameters from your IdP: * **IdP URL**: The URL of your Identity Provider (IdP), which will be used to authenticate users. * **Client ID**: Credentials your IdP provides when registering Aiven for OpenSearch as a client application. This credential is used to authenticate your Aiven for OpenSearch client application against the IdP and facilitate secure communication. * **Client Secret**: Credentials your IdP provides when you register Aiven for OpenSearch as a client application. * **Scope**: The scope of the authentication request specifies the permissions to request from the Identity Provider (IdP). The available and required scopes may vary depending on the IdP you are using. Some common scopes include `openid`, `profile`, `email`. * **Roles key and subject key**: Keys that help Aiven for OpenSearch Dashboards understand which part of the returned token contains role information and which part contains the user's identity or name. note The **Redirect URL** is automatically generated and available in the Aiven Console. This is the URL to which the Identity Provider (IdP) will redirect users after successful authentication. For more information on how to obtain this URL, see the next section. ## Enable OpenID Connect authentication via Aiven Console[​](#enable-openid-connect-authentication-via-aiven-console "Direct link to Enable OpenID Connect authentication via Aiven Console") 1. In the [Aiven Console](https://console.aiven.io/), access your Aiven for OpenSearch service where to enable OpenID Connect. 2. Select **Users** from the left sidebar. 3. In the **SSO authentication** section, use the **Add method** drop-down and select **OpenID**. 4. On **Configure OpenID Connect authentication** screen, * **Redirect URL**: This URL is auto-populated. It is the URL that users will be redirected to after they have successfully authenticated through the IdP. * **IdP URL**: Enter the URL of your OpenID Connect Identity Provider. This is the URL that Aiven will use to redirect users to your IdP for authentication. * **Client ID**: Enter the ID you obtained during your IdP registration. This ID is used to authenticate your Aiven application with your IdP. * **Client Secret**: Enter the secret associated with the Client ID. This secret is used to encrypt communications between your Aiven application and your IdP. * **Scope**: The scope of the claims. This is the set of permissions that you are requesting from your IdP. For example, you can request the `openid`, `profile`, and `email` scopes to get the user's identity, profile information, and email address. * **Roles key**: The key in the returned JSON that stores the user's roles. This key is used by Aiven to determine the user's permissions. * **Subject key**: This refers to the specific key within the returned JSON that holds the user's name or identifying subject. Aiven uses this key to recognize and authenticate the user. By default, this key is labeled as `Subject`. 5. Optional: **Enable advanced configuration** to fine-tune the authentication process. Aiven for OpenSearch provides the following advanced configuration options: * **Token handling**: Choose your preferred method for handling the authentication token. * **HTTP Header**: The authentication token is passed in an HTTP header. * **Header name**: Enter the specific name of the HTTP header that will contain the authentication token. The default header name is Authorization. * **URL parameter**: The authentication token is passed as a URL parameter. You can specify the parameter name. * **Parameter name**: Enter the specific URL parameter that will carry the authentication token. * **Refresh limit count**: Maximum number of unrecognized JWT key IDs allowed within 10 seconds. Enter the value for the Refresh Limit Count parameter. The default value is 10. * **Refresh limit window (ms)**: This is the interval, measured in milliseconds, during which the system will verify unrecognized JWT key IDs. Enter the value for the Refresh Limit Window parameter. The default value is 10,000 (10 seconds). 6. Select **Enable** to complete the setup and activate the configuration. ## Additional resources[​](#additional-resources "Direct link to Additional resources") * [OpenSearch OpenID Connect documentation](https://opensearch.org/docs/latest/security/authentication-backends/openid-connect/) --- # Use Aggregations with OpenSearch® and NodeJS Learn how to aggregate data using OpenSearch and its NodeJS client. In this tutorial we'll look at different types of aggregations, write and execute requests to learn more about the data in our dataset. note If you're new to OpenSearch® and its JavaScript client, see [how to write search queries with OpenSearch with NodeJS](/docs/products/opensearch/howto/opensearch-and-nodejs.md). ## Prepare the playground[​](#prepare-the-playground "Direct link to Prepare the playground") You can create an OpenSearch cluster either with the visual interface or with the command line. Depending on your preference follow the instructions for [getting started with the console for Aiven for Opensearch](/docs/products/opensearch/get-started.md) or see [how to create a service with the help of Aiven command line interface](/docs/tools/cli/service-cli.md). note You can also clone the final demo project from [GitHub repository](https://github.com/aiven/demo-open-search-node-js). ### File structure and GitHub repository[​](#file-structure-and-github-repository "Direct link to File structure and GitHub repository") To organise our development space we'll use these files: * `config.js` to keep necessary basis to connect to the cluster, * `index.js` to hold methods which manipulate the index, * `helpers.js` to contain utilities for logging responses, * `search.js` and `aggregation.js` for methods specific to search and aggregation requests. we'll be adding code into these files and running the methods from the command line. ### Connect to the cluster and load data[​](#connect-to-the-cluster-and-load-data "Direct link to Connect to the cluster and load data") Follow instructions on how to [connect to the cluster with a NodeJS client](/docs/products/opensearch/howto/connect-with-nodejs.md) and add the necessary code to `config.js`. Once you're connected [load a sample data set](/docs/products/opensearch/howto/sample-dataset.md#load-data-with-nodejs) and [retrieve the data mapping](/docs/products/opensearch/howto/sample-dataset.md#get-mapping-with-nodejs) to understand the structure of the created index. note In the code snippets we'll keep error handling somewhat simple and use `console.log` to print information into the terminal. Now you're ready to start aggregating the data. ## Aggregations[​](#aggregations "Direct link to Aggregations") We'll write and run examples for three different types of aggregations: metric, bucket and pipeline. You can read more about aggregations in [a concept article](/docs/products/opensearch/concepts/aggregations.md). ### Structure and syntax[​](#structure-and-syntax "Direct link to Structure and syntax") To calculate an aggregation, create a request and send it to the `search()` endpoint (the same we used for search queries). The request body should include an object with the key `aggs` (or you can use a longer word `aggregations`). Specify in this object **a name** of your aggregation (we might want to reference each aggregation individually), **an aggregation type** and **a set of properties** specific for the given aggregation type. note In most cases when we deal with aggregations we are not interested in individual hits, that's why it is common to set `size` property to zero. The structure of a simple request looks like this: ``` client.search( { index, body: { aggs: { 'GIVE-IT-A-NAME': { // aggregation name 'SPECIFY-TYPE': { // one of the supported aggregation types ... // list of properties, such as a field on which to perform the aggregation }, }, }, }, size: 0, // we're not interested in `hits` } ); ``` The best way to learn more about each type of aggregations is to try them out. Therefore, it's time to make our hands dirty and do some coding. Create `aggregate.js` file, this is where we'll be adding our code. At the top of the file import client and index name, we'll need them to send requests to the cluster. ``` const { client, indexName: index } = require("./config"); ``` ## Metrics aggregations[​](#metrics-aggregations "Direct link to Metrics aggregations") ### Average value[​](#average-value "Direct link to Average value") The simplest form of an aggregation is perhaps a calculation of a single-value metric, such as finding an average across values in a field. Using the draft structure of an aggregation we can create a method to calculate the average of the recipe ratings: ``` /** * Calculate average rating of all documents * run-func aggregate averageRating */ module.exports.averageRating = () => { client.search( { index, body: { aggs: { "average-rating": { // aggregation name avg: { // one of the supported aggregation types field: "rating", // list of properties for the aggregation }, }, }, }, size: 0, // ignore `hits` }, (error, result) => { // callback to log the output if (error) { console.error(error); } else { console.log(result.body.aggregations["average-rating"]); } } ); }; ``` Run the method from the command line: ``` run-func aggregate averageRating ``` You'll see a calculated numeric value, the average of all values from the rating field across the documents. ``` { value: 3.7130597014925373 } ``` `avg` is one of many metric aggregation functions offered by OpenSearch. We can also use `max`, `min`, `sum` and others. To have a possibility to easily change aggregation function and aggregation field we can do couple of simplifications in the method we created: * move the aggregation type and aggregation field to the method parameters, so that different values can be passed as arguments * generate name dynamically based on field name * separate the callback function and use the dynamically generated name to print out the result With these changes our method looks like this: ``` const logAggs = (field, error, result) => { if (error) { console.error(error); } else { console.log(result.body.aggregations[field]); } }; /** * Get metric aggregations for the field * Examples: avg, min, max, stats, extended_stats, percentiles, terms * run-func aggregate metric avg rating */ module.exports.metric = (metric, field) => { const body = { aggs: { [`aggs-for-${field}`]: { // aggregation name, which you choose [metric]: { // one of the supported aggregation types field, }, }, }, }; client.search( { index, body, size: 0, // ignore `hits` }, logAggs.bind(this, `aggs-for-${field}`) // callback to log the aggregation output ); }; ``` Run the method to make sure that we still can calculate the average rating: ``` run-func aggregate metric avg rating ``` And because we like clean code, move and export the `logAggs` function from `helpers.js` and reference it in `aggregate.js`. ``` const { logAggs } = require("./helpers"); ``` ### Other simple metrics[​](#other-simple-metrics "Direct link to Other simple metrics") We can use the method we created to run other types of metric aggregations, for example, to find what the minimum sodium value is, in any of the recipes: ``` run-func aggregate metric min sodium ``` Try out other fields and simple functions such as `min`, `max`, `avg`, `sum`, `count`, `value_count` and see what results you will get. ### Cardinality[​](#cardinality "Direct link to Cardinality") Another interesting single-value metric is `cardinality`. Cardinality is an estimated number of distinct values found in a field of a document. For example, by calculating the cardinality of the rating field, you will learn that there are only eight distinct rating values over all 20k recipes. Which makes me suspect that the rating data was added artificially later into the data set. The cardinality of `calories`, `sodium` and `fat` field contain more realistic diversity: ``` run-func aggregate metric cardinality rating ``` ``` { value: 8 } ``` Calculating cardinality for sodium and other fields and see what conclusions you can make! ### Field statistics[​](#field-statistics "Direct link to Field statistics") A multi-value aggregation returns an object rather than a single value. An example of such aggregation are statistics and we can continue using the method we created to explore different types of computed statistics. Get a set of metrics (`avg`, `count`, `max`, `min` and `sum`) by using `stats` aggregation type: ``` run-func aggregate metric stats rating ``` ``` { count: 20100, min: 0, max: 5, avg: 3.7130597014925373, sum: 74632.5 } ``` To get additional information, such as standard deviation, variance and bounds, use `extended_stats`: ``` run-func aggregate metric extended_stats rating ``` ``` { count: 20100, min: 0, max: 5, avg: 3.7130597014925373, sum: 74632.5, sum_of_squares: 313374.21875, variance: 1.803944804893444, variance_population: 1.803944804893444, variance_sampling: 1.8040345578565216, std_deviation: 1.3431101238891188, std_deviation_population: 1.3431101238891188, std_deviation_sampling: 1.3431435358354376, std_deviation_bounds: { upper: 6.399279949270775, lower: 1.0268394537142997, upper_population: 6.399279949270775, lower_population: 1.0268394537142997, upper_sampling: 6.399346773163412, lower_sampling: 1.0267726298216622 } } ``` ### Percentiles[​](#percentiles "Direct link to Percentiles") Another example of a multi-value aggregation are `percentiles`. Percentiles are used to interpret and understand data indicating how a given data point compares to other values in a data set. For example, if you take a test and score on the 80th percentile, it means that you did better than 80% of participants. Similarly, when a provider measures internet usage and peaks, the 90th percentile indicates that 90% of time the usage falls below that amount. Calculate percentiles for `calories`: ``` run-func aggregate metric percentiles calories ``` ``` { values: { '1.0': 17.503999999999998, '5.0': 62, '25.0': 197.65254901960782, '50.0': 331.2031703590527, '75.0': 585.5843561472852, '95.0': 1317.4926233766223, '99.0': 3256.4999999999945 } } ``` From the returned result you can see that 50% of recipes have less than 331 calories. Interestingly, only one percent of the meals is more than 3256 calories. You must be curious what falls within that last percentile ;) Now that we know the value to look for, we can use [a range query](/docs/products/opensearch/howto/opensearch-and-nodejs.md#find-fields-with-a-value-within-a-range) to find the recipes. Set the minimum value, but keep the maximum empty to allow no bounds: ``` run-func search range calories 3256 ``` ``` [ 'Ginger Crunch Cake with Strawberry Sauce ', 'Apple, Pear, and Cranberry Coffee Cake ', 'Roast Lobster with Pink Butter Sauce ', 'Birthday Party Paella ', 'Clementine-Salted Turkey with Redeye Gravy ', 'Roast Goose with Garlic, Onion and Sage Stuffing ', 'Chocolate Plum Cake ', 'Carrot Cake with Cream Cheese-Lemon Zest Frosting ', 'Lemon Cream Pie ', 'Rice Pilaf with Lamb, Carrots, and Raisins ' ] ``` ## Bucket aggregations[​](#bucket-aggregations "Direct link to Bucket aggregations") ### Buckets based on ranges[​](#buckets-based-on-ranges "Direct link to Buckets based on ranges") You can aggregate data by dividing it into a set of buckets. We can either predefine these buckets, or create them dynamically to fit the data. To understand how this works, we'll create a method to aggregate recipes into buckets based on sodium ranges. We use `range` aggregation and add a property `ranges` to describe how we want to split the data across buckets: ``` /** * Group recipes into bucket based on sodium levels * run-func aggregate sodiumRange */ module.exports.sodiumRange = () => { client.search( { index, body: { aggs: { "sodium-ranges": { // aggregation name range: { // range aggregation field: "sodium", // field to use for the aggregation ranges: [ // the buckets we want { to: 500.0 }, { from: 500.0, to: 1000.0 }, { from: 1000.0 }, ], }, }, }, }, size: 0, }, (error, result) => { // callback to output the result if (error) { console.error(error); } else { console.log(result.body.aggregations["sodium-ranges"]); } } ); }; ``` Run it with: ``` run-func aggregate sodiumRange ``` and check the results: ``` { buckets: [ { key: '*-500.0', to: 500, doc_count: 10411 }, { key: '500.0-1000.0', from: 500, to: 1000, doc_count: 2938 }, { key: '1000.0-*', from: 1000, doc_count: 2625 } ] } ``` By looking at `doc_count` we can say how many recipes fall into each of the buckets. However, our method is narrowed down to a specific scenario. We want to refactor it a bit to use for other fields and different sets of ranges. To achieve this we'll: * move aggregation field and bucket ranges to the list of method parameters * use the rest parameter syntax to collect range values * transform the list of range values into `ranges` object in a format `from X]{.title-ref} / [to Y` expected by OpenSearch API * use the `logAggs` function, which we already created, to log the results * separate `body` into a variable for better readability ``` /** * Group recipes into bucket based on the provided field and set of ranges * run-func aggregate range sodium 500 1000 */ module.exports.range = (field, ...values) => { // map values to list of ranges // in format 'from X'/'to Y' const ranges = values.map((value, index) => ({ from: values[index - 1], to: value, })); // account for the last item 'from X to infinity' ranges.push({ from: values[values.length - 1], }); const body = { aggs: { [`range-aggs-for-${field}`]: { range: { field, ranges, }, }, }, }; client.search( { index, body, size: 0, }, logAggs.bind(this, `range-aggs-for-${field}`) ); }; ``` To make sure that the upgraded function works just like the one one, run: ``` run-func aggregate range sodium 500 1000 ``` Now you can run the method with other fields and custom ranges, for example, split recipes into buckets based on values in the field `fat`: ``` run-func aggregate range fat 1 5 10 30 50 100 ``` The returned buckets are: ``` { buckets: [ { key: '*-1.0', to: 1, doc_count: 1230 }, { key: '1.0-5.0', from: 1, to: 5, doc_count: 1609 }, { key: '5.0-10.0', from: 5, to: 10, doc_count: 1916 }, { key: '10.0-30.0', from: 10, to: 30, doc_count: 6526 }, { key: '30.0-50.0', from: 30, to: 50, doc_count: 2404 }, { key: '50.0-100.0', from: 50, to: 100, doc_count: 1648 }, { key: '100.0-*', from: 100, doc_count: 575 } ] } ``` Why not experiment more with the range aggregation? We still have `protein` values, and can also play with the values for the ranges to learn more about recipes from our dataset. ### Buckets for every unique value[​](#buckets-for-every-unique-value "Direct link to Buckets for every unique value") Sometimes we want to divide the data into buckets, where each bucket corresponds to a unique value present in a field. This type of aggregations is called `terms` aggregation and is helpful when we need to have more granular understanding of a dataset. For example, we can learn how many recipes belong to each category. The structure of the method for `terms aggregation` will be similar to what we wrote for the ranges, with a couple of differences: * use aggregation type `terms` * use an optional property `size`, which specifies the upper limit of the buckets we want to create. ``` /** * Group recipes into buckets for every unique value * `run-func aggregate terms categories.keyword 20` */ module.exports.terms = (field, size) => { const body = { aggs: { [`terms-aggs-for-${field}`]: { terms: { // aggregate data by unique terms field, size, // max number of buckets generated, default value is 10 }, }, }, }; client.search( { index, body, size: 0, }, logAggs.bind(this, `terms-aggs-for-${field}`) ); }; ``` To get the buckets created for different categories run: ``` run-func aggregate terms categories.keyword ``` Here are the resulting delicious categories: ``` { doc_count_error_upper_bound: 0, sum_other_doc_count: 175719, buckets: [ { key: 'Bon Appétit', doc_count: 9355 }, { key: 'Peanut Free', doc_count: 8390 }, { key: 'Soy Free', doc_count: 8088 }, { key: 'Tree Nut Free', doc_count: 7044 }, { key: 'Vegetarian', doc_count: 6846 }, { key: 'Gourmet', doc_count: 6648 }, { key: 'Kosher', doc_count: 6175 }, { key: 'Pescatarian', doc_count: 6042 }, { key: 'Quick & Easy', doc_count: 5372 }, { key: 'Wheat/Gluten-Free', doc_count: 4906 } ] } ``` We can see a couple of interesting things in the response. First, there were just 10 buckets created, each of which contains `doc_count` indicating number of recipes within particular category. Second, `sum_other_doc_count` is the sum of documents which are left out of response, this number is high because almost every recipe is assigned to more than one category. We can increase the number of created buckets by using the `size` property: ``` run-func aggregate terms categories.keyword 30 ``` Now the list of buckets contains 30 items. ### Find least frequent items[​](#find-least-frequent-items "Direct link to Find least frequent items") Did you notice that the buckets created with the help of `terms` aggregation are sorted by their size in descending order? You might wonder how you can find the least frequent items? You can use the `rare_terms` aggregation! This creates a set of buckets sorted by number of documents in ascending order. As a result, the most rarely used items will be at the top of the response. `rare_terms` request is very similar to `terms`, however, instead of `size` property which defines total number of created buckets, `rare_terms` relies on `max_doc_count`, which sets upper limit for number of documents per bucket. ``` /** * Group recipes into buckets to find the most rare items * `run-func aggregate rareTerms categories.keyword 3` */ module.exports.rareTerms = (field, max) => { const body = { aggs: { [`rare-terms-aggs-for-${field}`]: { rare_terms: { field, max_doc_count: max, // get buckets that contain no more than max items }, }, }, }; client.search( { index, body, size: 0, }, logAggs.bind(this, `rare-terms-aggs-for-${field}`) ); }; ``` ``` run-func aggregate rareTerms categories.keyword 3 ``` The result will return us all the categories with at most three documents each. ### Histograms[​](#histograms "Direct link to Histograms") The story of bucket aggregations won't be complete without speaking about histograms. Histograms aggregate date based on provided interval. And since we have a `date` property, we'll build a date histogram. The format of the histogram aggregation is similar to what we saw so far, so we can create a new method almost identical to previous ones: ``` /** * Date histogram with a time interval * `run-func aggregate dateHistogram date year` */ module.exports.dateHistogram = (field, interval) => { const body = { aggs: { [`histogram-for-${field}`]: { date_histogram: { // aggregation type field, interval, // such as minute, hour, day, month or year }, }, }, }; client.search( { index, body, size: 0, }, logAggs.bind(this, `histogram-for-${field}`) ); }; ``` Values for the interval field can be from `minute` up to a `year`. ``` run-func aggregate dateHistogram date year ``` The results when we use a year: ``` { buckets: [ { key_as_string: '1996-01-01T00:00:00.000Z', key: 820454400000, doc_count: 1 }, ... { key_as_string: '2004-01-01T00:00:00.000Z', key: 1072915200000, doc_count: 11576 }, ... ] } ``` You should see a list of buckets, one per each year starting at 1996 and up to 2016, with `doc_count` indicating how many recipes belong to each year. Most of the data items are marked by year 2004. Now that we have seen examples of metric and bucket aggregations, it is time to learn some more advanced concepts of pipeline aggregations. ## Pipeline aggregations[​](#pipeline-aggregations "Direct link to Pipeline aggregations") ### Calculate moving average[​](#calculate-moving-average "Direct link to Calculate moving average") When working with continuously incoming data we might want to understand the trends and changes in the figures. This is convenient in many situations, such as helping to see the changes in sales over a given time, noticing the divergence in the activity of users or learn about other trends. OpenSearch allows "piping" the results of one aggregation into the different one to achieve more granular analysis through an intermediate step. To demonstrate an example of pipeline aggregations, we'll look at the moving average of number of recipes added throughout the years. With the help of what we learned so far and a couple of new tools we can do the following: 1. Create a date histogram to divide documents across years (we name it `date_histogram`) 2. Create a metric aggregation to count documents added per year (we name it `new_recipes`) 3. Use a moving function, a pipeline feature, to glue theses aggregations together 4. Use a built-in function `unweightedAvg` to calculate average value within a window 5. Use `shift` property to move window one step forward and include the current year (by default the current data position is excluded from the calculated year) 6. Set `window` property to define the size of moving window When put these pieces together we can write this method: ``` /** * Calculating the moving average of number of added recipes across years * `run-func aggregate movingAverage` */ module.exports.movingAverage = () => { const body = { aggs: { recipes_per_year: { // 1. date histogram date_histogram: { field: "date", interval: "year", }, aggs: { recipes_count: { // 2. metric aggregation to count new recipes value_count: { // aggregate by number of documents with field 'date' field: "date" }, }, moving_average: { moving_fn: { // 3. glue the aggregations script: "MovingFunctions.unweightedAvg(values)", // 4. a built-in function shift: 1, // 5. take into account the existing year as part of the window window: 3, // 6. set size of the moving window buckets_path: "recipes_count", gap_policy: "insert_zeros", // account for years where no recipes were // added and replace null value with zeros }, }, }, }, }, }; client.search( { index, body, size: 0, }, (error, result) => { if (error) { console.error(error); } else { console.log(result.body.aggregations["recipes_per_year"].buckets); } } ); }; ``` Run it on the command line: ``` run-func aggregate movingAverage ``` The returned data for every year including a value `moving_average`: ``` [ { key_as_string: '1996-01-01T00:00:00.000Z', key: 820454400000, doc_count: 1, count: { value: 1 }, moving_average: { value: 1 } }, { key_as_string: '1997-01-01T00:00:00.000Z', key: 852076800000, doc_count: 0, count: { value: 0 }, moving_average: { value: 0.5 } }, { key_as_string: '1998-01-01T00:00:00.000Z', key: 883612800000, doc_count: 3, count: { value: 3 }, moving_average: { value: 1.3333333333333333 } }, { key_as_string: '1999-01-01T00:00:00.000Z', key: 915148800000, doc_count: 4, count: { value: 4 }, moving_average: { value: 2.3333333333333335 } }, ... ] ``` Pay attention to the values of `count` and `moving_average`. To understand better how those numbers were calculated, we can compute first several values on our own: | Year | Added documents | Moving average | | ---- | --------------- | ------------------------------------------ | | 1996 | 1 | 1 (no previous years to make a comparison) | | 1997 | 0 | (1 + 0) / 2 = 0.5 (we had only two years) | | 1998 | 3 | (1 + 0 + 3) / 3 = 1.3(3) | | 1999 | 4 | (0 + 3 + 4) / 3 = 2.3(3) | | 2000 | 0 | (3 + 4 + 0) / 3 = 2.3(3) | | ... | ... | and so on | Making sense of the `moving_average` result For every data point (a year in our case) we take the count of added recipes, add number of recipes added over last two years and divide the result by three (according to the size of our window). For the first and second year we divide by the number of available years (1 and 2 respectively). And this is how moving average is calculated. If you compare numbers from the table with the numbers returned in the `moving_average` field of the response body, you can see they are same. ### Other moving functions[​](#other-moving-functions "Direct link to Other moving functions") We used one of existing built-in functions `MovingFunctions.unweightedAvg(values)`, which as its name says calculates unweighted average. Unweighted in this context means that the function does not perform any time-dependent weighting. You can also use other functions such as max(), min(), stdDev() and sum(). Additionally, you can write your own functions, such as ``` moving_fn: { script: "return values.length === 1 ? 1 : 0" } ``` Try replacing the script with `MovingFunctions.min(values)`, `MovingFunctions.max(values)` or custom scripts, changing the window size and shift, and see how thus affects the outcome! ## What's next[​](#whats-next "Direct link to What's next") This was a long ride, hopefully you have a better understanding now how to use aggregations with OpenSearch and its NodeJS client. The best way to deepen the knowledge on these concepts is to play and experiment with different types of aggregations. We covered some examples but [OpenSearch documentation](https://opensearch.org/docs/latest/opensearch/aggregations/) contains many more. Check OpenSearch docs, as well as other resources listed below to learn more. ## Resources[​](#resources "Direct link to Resources") * [Demo GitHub repository](https://github.com/aiven/demo-open-search-node-js) - where all the examples we run in this tutorial can be found * [Previous chapter of the tutorial](/docs/products/opensearch/howto/opensearch-and-nodejs.md) - learn how to use OpenSearch with NodeJS to make search queries * [How to use OpenSearch with curl](/docs/products/opensearch/howto/opensearch-with-curl.md) * [GitHub repository for OpenSearch JavaScript client](https://github.com/opensearch-project/opensearch-js) * [Official OpenSearch documentation](https://opensearch.org) * [Metric aggregations](https://opensearch.org/docs/latest/opensearch/metric-agg/) * [Bucket aggregations](https://opensearch.org/docs/latest/opensearch/bucket-agg/) * [Pipeline aggregations](https://opensearch.org/docs/latest/opensearch/pipeline-agg/) --- # Create alerts with OpenSearch® API OpenSearch® alerting feature sends notifications when data from one or more indices meets certain conditions that can be customized. Use case examples are such as monitoring for HTTP status code 503, CPU load average above certain percentage or watch for counts of a specific keyword in logs for a specific amount of interval, notification to be configured to be sent via email, slack or custom webhooks and other destination, in this example we are using slack as the destination. In the following example, we are creating an alert programmatically by using OpenSearch Alerting API. We are using a `sample-host-health` index as datasource to create a simple alert to check cpu load, action will be triggered when average of `cpu_usage_percentage` over `3` minutes is above `75%` OpenSearch API Alerting API URL can be copied from Aiven console: Click the **Overview** tab > **OpenSearch** under `Connection Information` > **Service URI** append `_plugins/_alerting/monitors` to the **Service URI**. Example: `https://username:password@os-name-myproject.aivencloud.com:24947/_plugins/_alerting/monitors` Save the JSON below into `cpu_alert.json` ``` { "name": "High CPU Monitor", "type": "monitor", "monitor_type": "query_level_monitor", "enabled": true, "schedule": { "period": { "unit": "MINUTES", "interval": 1 } }, "inputs": [ { "search": { "indices": ["sample-host-health"], "query": { "size": 0, "aggregations": { "metric": { "avg": { "field": "cpu_usage_percentage" } } }, "query": { "bool": { "filter": [ { "range": { "timestamp": { "gte": "{{period_end}}||-3m", "lte": "{{period_end}}", "format": "epoch_millis" } } } ] } } } } } ], "triggers": [ { "query_level_trigger": { "id": "lNbSt30BZGFGbIUYx2bb", "name": "high_cpu", "severity": "1", "condition": { "script": { "source": "return ctx.results[0].aggregations.metric.value == null ? false : ctx.results[0].aggregations.metric.value > 75", "lang": "painless" } }, "actions": [ { "id": "ldbSt30BZGFGbIUYx2bb", "name": "slack", "destination_id": "gkQgp30BRvA_n4QUwZDL", "message_template": { "source": "Monitor {{ctx.monitor.name}} just entered alert status. Please investigate the issue.\n - Trigger: {{ctx.trigger.name}}\n - Severity: {{ctx.trigger.severity}}\n - Period start: {{ctx.periodStart}}\n - Period end: {{ctx.periodEnd}}", "lang": "mustache" }, "throttle_enabled": false, "subject_template": { "source": "High CPU Test Alert", "lang": "mustache" } } ] } } ], "ui_metadata": { "schedule": { "timezone": null, "frequency": "interval", "period": { "unit": "MINUTES", "interval": 1 }, "daily": 0, "weekly": { "tue": false, "wed": false, "thur": false, "sat": false, "fri": false, "mon": false, "sun": false }, "monthly": { "type": "day", "day": 1 }, "cronExpression": "0 */1 * * *" }, "search": { "searchType": "graph", "timeField": "timestamp", "aggregations": [ { "aggregationType": "avg", "fieldName": "cpu_usage_percentage" } ], "groupBy": [], "bucketValue": 3, "bucketUnitOfTime": "m", "where": { "fieldName": [], "fieldRangeEnd": 0, "fieldRangeStart": 0, "fieldValue": "", "operator": "is" } }, "monitor_type": "query_level_monitor" } } ``` Use `curl` to create the alert ``` curl -XPOST \ https://username:password@os-name-myproject.aivencloud.com:24947/_plugins/_alerting/monitors \ -H 'Content-type: application/json' -T cpu_alert.json ``` * The required JSON request format can be found in [OpenSearch Alerting API documentation](https://opensearch.org/docs/latest/observing-your-data/alerting/api/) --- # Write search queries with OpenSearch® and NodeJS Learn how the OpenSearch® JavaScript client gives a clear and useful interface to communicate with an OpenSearch cluster and run search queries. To make it more delicious we'll be using a recipe dataset from Kaggle. ## Prepare the playground[​](#prepare-the-playground "Direct link to Prepare the playground") You can create an OpenSearch cluster either with the visual interface or with the command line. Depending on your preference follow the instructions for [getting started with the console for Aiven for Opensearch](/docs/products/opensearch/get-started.md) or see [how to create a service with the help of Aiven command line interface](/docs/tools/cli/service-cli.md). note You can also clone the final demo project from [GitHub repository](https://github.com/aiven/demo-open-search-node-js). ### File structure and GitHub repository[​](#file-structure-and-github-repository "Direct link to File structure and GitHub repository") To organise our development space we'll use these files: * `config.js` to keep necessary basis to connect to the cluster, * `index.js` to hold methods which manipulate the index, * `helpers.js` to contain utilities for logging responses, * `search.js` for methods specific to search requests. we'll be adding code into these files and running the methods from the command line. ### Connect to the cluster and load data[​](#connect-to-the-cluster-and-load-data "Direct link to Connect to the cluster and load data") Follow instructions on how to [connect to the cluster with a NodeJS client](/docs/products/opensearch/howto/connect-with-nodejs.md) and add the necessary code to `config.js`. Once you're connected [load a sample data set](/docs/products/opensearch/howto/sample-dataset.md#load-data-with-nodejs) and [retrieve the data mapping](/docs/products/opensearch/howto/sample-dataset.md#get-mapping-with-nodejs) to understand the structure of the created index. ### Extra helpers[​](#extra-helpers "Direct link to Extra helpers") To render the response, add the following helper method to your `helpers.js` file. ``` /** * Parsing and logging list of titles from the result, used in callbacks. */ const logTitles = (error, result) => { if (error) { console.error(error); } else { const hits = result.body.hits.hits; console.log(`Number of returned results is ${hits.length}`); console.log(hits.map(hit => hit._source.title)); } }; ``` note In the code snippets we'll keep error handling somewhat simple and use `console.log` to print information into the terminal. Now you're ready to start querying the data. ## Query the data[​](#query-the-data "Direct link to Query the data") Now that we have data in the OpenSearch cluster, we're ready to construct and run search queries. We will use `search` method which is provided by the OpenSearch JavaScript client. The following code goes into `search.js`, you'll need connection configuration and helpers methods. Therefore, include them at the top of your `search.js` file with ``` const { client, indexName } = require("./config"); const { logTitles } = require("./helpers"); ``` The `search` method expects three optional parameters: `params`, `options` and `callback`. The query details are into the `params` object, which contains the name of the index (`index`), the maximum number of results to be returned (`size`), if the response is paginated (`size` and `from`), by which fields to sort the data (`sort`) and others. we'll pay a closer attention to two of these parameters - `q` - a query defined in the Lucene query string syntax and `body` - a query based on Query DSL (Domain Specific Language). These are two main methods to construct a query. The query string syntax is a powerful tool which can be used for a variety of requests. It is especially convenient for cURL requests, since it is a very compact string. However, as the complexity of a request grows, it becomes more difficult to read and maintain these types of queries. ``` //example of using a query syntax client.search({ index: 'recipes', q: 'ingredients:broccoli AND calories:(>=100 AND <200)' }) ``` A query with a request `body` might look bulky at first glance, but its structure makes it easier to read, understand and modify the content. Unlike `q`, which expects a string, `body` is an object allowing a variety of granular parameters. ``` //example of using a request body client.search({ index: indexName, body: { query: { match: { property: 'value' } } } }) ``` Let's focus on Query DSL and its three main groups of requests: term-level, full-text and boolean. You will also see how to use the Lucene query string syntax inside Query DSL. * Term-level queries are handy when we need to find **exact matches** for numbers, dates or tags and don't need to sort the results by relevance. Term-level queries use search terms as they are without additional analysis. * Full-text queries allow a smarter search for matches in analysed text fields and return results sorted by relevance. * Boolean queries are useful to combine multiple queries together. It supports boolean clauses such as `must`, `filter`, `should` and `must_not`. ### Find matching field values[​](#find-matching-field-values "Direct link to Find matching field values") One of the examples of a term-level query is searching for all entries containing a particular value in a field. To construct a body request we use `term` property which defines an object, where the name is a field and the value is a term we're searching in this field. ``` /** * Searching for exact matches of a value in a field. * run-func search term sodium 0 */ module.exports.term = (field, value) => { console.log(`Searching for values in the field ${field} equal to ${value}`); const body = { query: { term: { [field]: value, }, }, }; client.search( { index: indexName, body, }, logTitles ); }; ``` ``` run-func search term sodium 0 ``` Try to replace "sodium" with other fields we have, such as "calories" or "fat". ### Find fields with a value within a range[​](#find-fields-with-a-value-within-a-range "Direct link to Find fields with a value within a range") When dealing with numeric values, naturally we want to be able to search for certain ranges of values. To find all documents that contain terms in a specific field within a given range, use `range` property. It expects an object, where the name is set to the field name and the body defines the upper and lower bounds: `gt` (greater than), `gte` (greater than or equal to), `lt` (less than) and `lte` (less than or equal to). ``` /** * Searching for a range of values in a field. * run-func search range sodium 0 10 */ module.exports.range = (field, gte, lte) => { console.log( `Searching for values in the ${field} ranging from ${gte} to ${lte}` ); const body = { query: { range: { [field]: { gte, lte, }, }, }, }; client.search( { index: indexName, body, }, logTitles ); }; ``` ``` run-func search range sodium 0 10 ``` Try your own term query. How about a search for food with a particular rating value, or finding all meals with zero calories? ### Find fields with fuzzy text matching[​](#find-fields-with-fuzzy-text-matching "Direct link to Find fields with fuzzy text matching") When searching for terms inside text fields, we can take into account typos and misspellings. We measure such "deviations" by a minimum number of single-character edits necessary to convert one word into another. Such types of queries are called `fuzzy` and the property `fuzziness` specifies the maximum edit distance. ``` /** * Specifying fuzziness to account for typos and misspelling. * run-func search fuzzy title pinapple 2 */ module.exports.fuzzy = (field, value, fuzziness) => { console.log( `Search for ${value} in the ${field} with fuzziness set to ${fuzziness}` ); const query = { query: { fuzzy: { [field]: { value, fuzziness, }, }, }, }; client.search( { index: indexName, body: query, }, logTitles ); }; ``` See if you can find recipes with misspelled pineapple 🍍 ``` run-func search fuzzy title pinapple 2 ``` Even though there is a typo in the word "pineapple", you still got relevant results. Try other search terms and different values for `fuzziness` to understand better how fuzzy queries work. What is your favourite food ingredient typo? ### Find best match with multiple search words[​](#find-best-match-with-multiple-search-words "Direct link to Find best match with multiple search words") A standard way to perform a full-text query is to use `match` property inside a request. `match` expects an object, the name of which is set to a specific field, and its body contains a search query in a form of a string. To see `match` in action use the method below to search for "Tomato garlic soup with dill". ``` /** * Finding matches sorted by relevance. * run-func search match title 'Tomato-garlic soup with dill' */ module.exports.match = (field, query) => { console.log(`Searching for ${query} in the field ${field}`); const body = { query: { match: { [field]: { query, }, }, }, }; client.search( { index: indexName, body, }, logTitles ); }; ``` ``` run-func search match title 'Tomato-garlic soup with dill' ``` In the response you should see different recipes of soups sorted by how close they are to "Tomato-garlic soup with dill" according to OpenSearch engine. What are your favourite recipes? Try searching for them and see if you find some new and unusual recipe combinations. ### Find matching phrases[​](#find-matching-phrases "Direct link to Find matching phrases") When the order of the words is important, use `match_phrase` instead of `match`. An additional power of `match_phrase` is that it allows to define how far search words can be from each other to still be considered a match. This parameter is called `slop` and its default value is `0`. The format of `match_phrase` is almost identical to `match`: ``` /** * Specifying a slop - a distance between search words. * run-func search slop directions "pizza pineapple" 10 */ module.exports.slop = (field, query, slop) => { console.log( `Searching for ${query} with slop value ${slop} in the field ${field}` ); const body = { query: { match_phrase: { [field]: { query, slop, }, }, }, }; client.search( { index: indexName, body, }, logTitles ); }; ``` We can use this method to find some recipes for pizza with pineapple. I learned from my Italian colleague that this considered a combination only for tourists, not a true pizza recipe. we'll do it by searching the `directions` field for words "pizza" and "pineapple" with top-most distance of 10 words in between. ``` run-func search slop directions "pizza pineapple" 10 ``` Oh look: "Pan-Fried Hawaiian Pizza" (don't tell my colleague). So far all the requests we tried returned us at most 10 results. Why 10? Because it is a default `size` value. It can be increased by setting `size` property to a higher number when making the request. we'll include this in the next example. ### Search with query string syntax[​](#search-with-query-string-syntax "Direct link to Search with query string syntax") Remember the Lucene query string syntax we talked about earlier, in relation to `q` parameter? We can also use it inside of Query DSL by defining `query_string` object. It requires its own `query` parameter and, optionally, we can specify `default_field` or `fields` properties to indicate the search fields. This example also sets `size` to demonstrate how we can get more than 10 results. ``` /** * Using special operators within a query string and a size parameter. * run-func search query ingredients "(salmon|tuna) +tomato -onion" 100 */ module.exports.query = (field, query, size) => { console.log( `Searching for ${query} in the field ${field} and returning maximum ${size} results` ); const body = { query: { query_string: { default_field: field, query, }, }, }; client.search( { index: indexName, body, size, }, logTitles ); }; ``` To find recipes with tomato, salmon or tuna and no onion, run: ``` run-func search query ingredients "(salmon|tuna) +tomato -onion" 100 ``` Now, experiment with your recipe search by including and excluding different ingredients. ### Combine queries to improve results[​](#combine-queries-to-improve-results "Direct link to Combine queries to improve results") The boolean clause types each affect the document relevance score differently. Both `must` and `should` positively contribute to the score, affecting the relevance of matches; `must_not` sets the score to 0, ensuring that the document won't appear in the results. `filter` clause is similar to `must`, however it has no effect on the relevance score. In the next method we combine what we learned so far, using both term-level and full-search queries to find recipes to make a quick and easy dish, with no garlic, low sodium and high protein. ``` /** * Combining several queries together * run-func search boolean */ module.exports.boolean = () => { console.log( `Searching for quick and easy recipes without garlic with low sodium and high protein` ); const body = { query: { bool: { must: { match: { categories: "Quick & Easy" } }, must_not: { match: { ingredients: "garlic" } }, filter: [ { range: { sodium: { lte: 50 } } }, { range: { protein: { gte: 5 } } }, ], }, }, }; client.search( { index: indexName, body, }, logTitles ); }; ``` ``` run-func search boolean ``` Now create your own boolean query, using what we learned to find recipes with particular nutritional values and ingredients. Experiment using different clauses to see how they affects the results. Related pages * [Aggregation tutorial](/docs/products/opensearch/howto/opensearch-aggregations-and-nodejs.md). * [Demo repository](https://github.com/aiven/demo-open-search-node-js). All the examples we run in this tutorial can be found in: * [OpenSearch JavaScript client](https://github.com/opensearch-project/opensearch-js) * [How to use OpenSearch with curl](/docs/products/opensearch/howto/opensearch-with-curl.md) * [Official OpenSearch documentation](https://opensearch.org) * [Term-level queries](https://opensearch.org/docs/latest/opensearch/query-dsl/term/) * [Full-text queries](https://opensearch.org/docs/latest/opensearch/query-dsl/full-text/) * [Boolean queries](https://opensearch.org/docs/latest/opensearch/query-dsl/bool/) --- # Set up OpenSearch® Dashboard multi-tenancy Aiven for OpenSearch® provides support for multi-tenancy through OpenSearch Security Dashboard. Multi-tenancy in OpenSearch Security enables multiple users or groups to securely access the same OpenSearch cluster while maintaining their distinct permissions and data access levels. With multi-tenancy, each tenant has its own isolated space for working with indexes, visualizations, dashboards, and other OpenSearch objects, ensuring tenant-specific data and resources are protected from unauthorized access. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven for OpenSearch * Administrative access to both the Aiven for OpenSearch service and OpenSearch Dashboard ## Optional: Enabling security management[​](#optional-enabling-security-management "Direct link to Optional: Enabling security management") Enabling OpenSearch Security management is optional if you are using the default tenants (Private and Global) in OpenSearch Dashboard without the need to restrict access to the Global tenant. However, if you intend to create custom tenants or require advanced authentication and authorization features, you must [enable OpenSearch Security management](/docs/products/opensearch/howto/enable-opensearch-security.md). ## Configure multi-tenancy in OpenSearch® Dashboard[​](#configure-multi-tenancy-in-opensearch-dashboard "Direct link to Configure multi-tenancy in OpenSearch® Dashboard") This section provides information on configuring multi-tenancy in OpenSearch Dashboard, which involves enabling OpenSearch Security management for custom tenant, creating custom tenants, assigning roles, and mapping roles to users. ### Step 1: Create a tenant[​](#step-1-create-a-tenant "Direct link to Step 1: Create a tenant") A tenant is a logical grouping of users and data, each with its own set of users, roles, and permissions. OpenSearch users can access two default tenants: *Global* and *Private*, which are available even without enabling OpenSearch Security management. All users share the Global Tenant, and the Private Tenant is exclusively available to a single user and cannot be shared. If you have enabled OpenSearch Security management and wish to create a custom tenant: 1. Log in to OpenSearch Dashboard with administrative access. 2. From the left navigation menu, select **Security** and select **Tenants**. 3. Select **Create tenant** to create a tenant. 4. In the **Create Tenant** screen, enter a name and description for your new tenant. 5. Select **Create** to save your new custom tenant. ### Step 2: Assign tenant to roles[​](#step-2-assign-tenant-to-roles "Direct link to Step 2: Assign tenant to roles") After creating a tenant, assign it to a role. A role is a collection of permissions for a specific tenant that can be granted to users. To assign a tenant to a role: 1. In the OpenSearch dashboard, go to the **Security** section in the left-hand navigation menu, then select **Roles**. 2. Choose whether to create a role or modify an existing one to include the tenant. 3. To create a role: * Select **Create role** and enter a name for your new role. * Select the permissions to grant to this role. * In the **Tenant permissions** section, choose the tenant you want to assign to the role from the dropdown menu. Then, select the tenant permissions for the role, such as read and/or write permissions. * Select **Create** to save your new role with the assigned tenant. 4. To modify an existing role: * Search for the role to edit and select it to view its permissions screen. * Select **Edit role** and add the required tenant in the **Tenant permissions** section. Additionally, select the tenant permissions for the role, such as read and/or write permissions. * Select **Update** to save your changes. ### Step 3: Map roles to users[​](#step-3-map-roles-to-users "Direct link to Step 3: Map roles to users") After assigning tenants to roles and setting the required permissions, the next step is associating each user with a specific role, granting them access to the tenant and its resources. The level of access and control a user has over the tenant's data and resources will be determined by their assigned role. To map roles to internal users: 1. In the OpenSearch dashboard, go to the **Security** section in the left-hand navigation menu, then select **Roles**. 2. Search for the role to assign a user and select it to view its details. 3. Select the **Mapped Users** tab and select **Map users** (or **Manage mapping** if users are already mapped). 4. In the **Users** section, choose the internal user you wish to assign to the role from the dropdown list. 5. Select **Map** to add the selected user to the mapped user list. note If you have enabled SAML SSO authentication in your Aiven for OpenSearch service, you can use SAML integration to map users roles. ### Step 4: Manage tenants[​](#step-4-manage-tenants "Direct link to Step 4: Manage tenants") To manage tenants in the OpenSearch dashboard, you can follow these steps: * **Access the tenant list**: Go to the Security section in the OpenSearch dashboard and select the Tenant option to view the available tenants and create new ones. * **Switch between tenants**: Besides creating new tenants, you can switch between them by selecting the checkbox next to the desired tenant and using the **Actions** dropdown to select **Switch the selected tenant**. * **View and create Index Patterns**: To view and create *Index Patterns*, *Saved Objects*, and manage *Advanced settings*, select **View Dashboards** or **View Visualisations** for a specific tenant. * **Edit, delete, or duplicate tenants**: To manage existing tenants, select them from the list and use the **Actions** dropdown to edit, delete, or duplicate them according to your needs. Related pages * [OpenSearch Dashboards multi-tenancy](https://opensearch.org/docs/2.6/security/multi-tenancy/tenant-index/) --- # Manage OpenSearch® log integration Aiven provides a log integration that sends logs from any Aiven service, such as Aiven for Apache Kafka®, Aiven for PostgreSQL®, and Aiven for Grafana®, to Aiven for OpenSearch®, so you can search, analyze, and monitor all your logs in one place. tip See this [video tutorial](https://www.youtube.com/watch?v=f4y9nPadO-M) for an end-to-end example of how to enable your Aiven for OpenSearch® log integration. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Console * API * CLI * Terraform * Kubernetes - User profile in the [Aiven Console](https://console.aiven.io/) - Source service: Aiven-managed service producing logs to be sent to Aiven for OpenSearch - Destination service: Aiven for OpenSearch service that receives the logs * [Personal token](/docs/platform/howto/create_authentication_token.md) for use with the Aiven API, CLI, Terraform, or other applications * [Aiven API](/docs/tools/api.md) * Source service: Aiven-managed service producing logs to be sent to Aiven for OpenSearch * Destination service: Aiven for OpenSearch service that receives the logs - [Personal token](/docs/platform/howto/create_authentication_token.md) for use with the Aiven API, CLI, Terraform, or other applications - [Aiven CLI](/docs/tools/cli.md) - Source service: Aiven-managed service producing logs to be sent to Aiven for OpenSearch - Destination service: Aiven for OpenSearch service that receives the logs * [Personal token](/docs/platform/howto/create_authentication_token.md) for use with the Aiven API, CLI, Terraform, or other applications * [Aiven Provider for Terraform](/docs/tools/terraform.md) * Source service: Aiven-managed service producing logs to be sent to Aiven for OpenSearch * Destination service: Aiven for OpenSearch service that receives the logs - [Personal token](/docs/platform/howto/create_authentication_token.md) for use with the Aiven API, CLI, Terraform, or other applications - [Aiven Operator for Kubernetes®](/docs/tools/kubernetes.md) - Source service: Aiven-managed service producing logs to be sent to Aiven for OpenSearch - Destination service: Aiven for OpenSearch service that receives the logs ## Enable log integration[​](#enable-log-integration "Direct link to Enable log integration") Enable logs integration to send logs from your service to Aiven for OpenSearch®: * Console * API * CLI * Terraform * Kubernetes 1. Log in to the [Aiven Console](https://console.aiven.io/) and go to the service that produces the logs to be sent to Aiven for OpenSearch. 2. In the **Observe** section, click **Logs**. 3. On the **Logs** page, click **Enable logs integration**. 4. In the **Logs integration** window, select an existing Aiven for OpenSearch service or create one, and click **Continue**. note If you choose to select an existing service and you are a member of more than one Aiven project with *operator* or *admin* access rights, select a project before selecting an Aiven for OpenSearch service. 5. In the **Configure logs integration** window, set up the `index prefix` and `index retention limit` parameters, and click **Enable**. note To effectively disable the `index retention limit`, set it to its maximum value of `10000` days. Call the [ServiceIntegrationCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationCreate) endpoint to enable log integration: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration \ --header "Authorization: Bearer API_TOKEN" \ --header "Content-Type: application/json" \ --data-raw '{ "integration_type": "logs", "source_service": "SOURCE_SERVICE_NAME", "dest_service": "OPENSEARCH_SERVICE_NAME", "user_config": { "elasticsearch_index_prefix": "INDEX_PREFIX", "elasticsearch_index_days_max": INDEX_RETENTION_DAYS } }' ``` Parameters: * `PROJECT_NAME`: Your Aiven project name * `API_TOKEN`: Your [personal token](/docs/platform/howto/create_authentication_token.md) * `SOURCE_SERVICE_NAME`: The service producing logs * `OPENSEARCH_SERVICE_NAME`: Your Aiven for OpenSearch service name * `INDEX_PREFIX`: Prefix for the index name * `INDEX_RETENTION_DAYS`: Number of days to keep logs Example: ``` curl --request POST \ --url https://api.aiven.io/v1/project/dev-sandbox/integration \ --header "Authorization: Bearer 123abc456def789ghi" \ --header "Content-Type: application/json" \ --data-raw '{ "integration_type": "logs", "source_service": "my-postgresql", "dest_service": "my-opensearch", "user_config": { "elasticsearch_index_prefix": "logs", "elasticsearch_index_days_max": 3 } }' ``` Use the [avn service integration-create](/docs/tools/cli/service/integration.md#avn_service_integration_create) command to enable log integration: ``` avn service integration-create \ --project PROJECT_NAME \ --source-service SOURCE_SERVICE_NAME \ --dest-service OPENSEARCH_SERVICE_NAME \ --integration-type logs \ -c elasticsearch_index_prefix=INDEX_PREFIX \ -c elasticsearch_index_days_max=INDEX_RETENTION_DAYS ``` Parameters: * `PROJECT_NAME`: Your Aiven project name * `SOURCE_SERVICE_NAME`: The service producing logs * `OPENSEARCH_SERVICE_NAME`: Your Aiven for OpenSearch service name * `INDEX_PREFIX`: Prefix for the index name * `INDEX_RETENTION_DAYS`: Number of days to keep logs Example: ``` avn service integration-create \ --project my-project \ --source-service my-kafka-service \ --dest-service my-opensearch \ --integration-type logs \ -c elasticsearch_index_prefix=logs \ -c elasticsearch_index_days_max=7 ``` Add a `aiven_service_integration` resource to enable log integration: ``` resource "aiven_service_integration" "logs_integration" { project = "PROJECT_NAME" integration_type = "logs" source_service_name = "SOURCE_SERVICE_NAME" destination_service_name = "OPENSEARCH_SERVICE_NAME" logs_user_config { elasticsearch_index_prefix = "INDEX_PREFIX" elasticsearch_index_days_max = INDEX_RETENTION_DAYS } } ``` Parameters: * `PROJECT_NAME`: Your Aiven project name * `SOURCE_SERVICE_NAME`: The service producing logs * `OPENSEARCH_SERVICE_NAME`: Your Aiven for OpenSearch service name * `INDEX_PREFIX`: Prefix for the index name, for example, `logs` * `INDEX_RETENTION_DAYS`: Number of days to keep logs, for example, `3` Run `terraform init`, `terraform plan`, and `terraform apply`. For more information, see the [aiven-service-integration resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration). Add a `ServiceIntegration` resource to enable log integration: ``` apiVersion: aiven.io/v1alpha1 kind: ServiceIntegration metadata: name: INTEGRATION_RESOURCE_NAME spec: authSecretRef: name: aiven-token key: TOKEN_NAME project: PROJECT_NAME integrationType: logs sourceServiceName: SOURCE_SERVICE_NAME destinationServiceName: OPENSEARCH_SERVICE_NAME logs: elasticsearch_index_prefix: INDEX_PREFIX elasticsearch_index_days_max: INDEX_RETENTION_DAYS ``` Apply the resource with `kubectl apply -f FILE_NAME.yaml`. For more information, see the [ServiceIntegration resource documentation](https://aiven.github.io/aiven-operator/resources/serviceintegration.html#spec.logs). ## Configure log integration[​](#configure-log-integration "Direct link to Configure log integration") There are two parameters that you can adjust when integrating logs to your OpenSearch service: * `index prefix`, specifies the prefix part of the index name * `index retention limit`, number of days to preserve the daily indexes warning The service's logs are sent from the selected service to your OpenSearch cluster. When the `index retention limit` is reached, those indexes are deleted from the OpenSearch cluster. You can change the configuration of the `index prefix` and `index retention limit` after the integration is enabled. * Console * API * CLI * Terraform * Kubernetes 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your Aiven for OpenSearch service. 2. In the **Data** section, click **Integrations**. 3. On the **Integrations** page, find the integrated service to configure. 4. Click **Actions** > **Edit**. 5. After updating `index prefix` or `index retention limit`, click **Edit**. 1) Get `INTEGRATION_ID` by calling the [ServiceGet](https://api.aiven.io/doc/#tag/Service/operation/ServiceGet) endpoint: ``` curl --request GET \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header "Authorization: Bearer API_TOKEN" ``` In the output, under `service_integrations`, find your integration and its `service_integration_id`. 2) Call the [ServiceIntegrationUpdate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationUpdate) endpoint to update the log integration configuration: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration/INTEGRATION_ID \ --header "Authorization: Bearer API_TOKEN" \ --header "Content-Type: application/json" \ --data-raw '{ "user_config": { "elasticsearch_index_prefix": "UPDATED_INDEX_PREFIX", "elasticsearch_index_days_max": UPDATED_INDEX_RETENTION_DAYS } }' ``` Parameters: * `PROJECT_NAME`: Your Aiven project name * `INTEGRATION_ID`: The integration ID * `API_TOKEN`: Your [personal token](/docs/platform/howto/create_authentication_token.md) * `UPDATED_INDEX_PREFIX`: Updated prefix for the index name * `UPDATED_INDEX_RETENTION_DAYS`: Updated number of days to keep logs Example: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/my-project/integration/123abc456def789ghi-123abc456def789ghi \ --header "Authorization: Bearer 123abc456def789ghi" \ --header "Content-Type: application/json" \ --data-raw '{ "user_config": { "elasticsearch_index_prefix": "new", "elasticsearch_index_days_max": 100 } }' ``` 1. Get `INTEGRATION_ID` by running the [`avn service integration-list`](/docs/tools/cli/service/integration.md#avn_service_integration_list) command: ``` avn service integration-list SOURCE_SERVICE_NAME --project PROJECT_NAME ``` Parameters: * `PROJECT_NAME`: Your Aiven project name * `SOURCE_SERVICE_NAME`: The service producing logs In the output table, find `SERVICE_INTEGRATION_ID` for the integrated services listed in columns `SOURCE` and `DEST`. 2. Use the [avn service integration-update](/docs/tools/cli/service/integration.md#avn_service_integration_update) command to update the log integration configuration: ``` avn service integration-update INTEGRATION_ID \ --project PROJECT_NAME \ -c elasticsearch_index_prefix=UPDATED_INDEX_PREFIX \ -c elasticsearch_index_days_max=UPDATED_INDEX_RETENTION_DAYS ``` Parameters: * `INTEGRATION_ID`: The integration ID * `PROJECT_NAME`: Your Aiven project name * `UPDATED_INDEX_PREFIX`: Updated prefix for the index name * `UPDATED_INDEX_RETENTION_DAYS`: Updated number of days to keep logs Example: ``` avn service integration-update 123abc456def789ghi-123abc456def789ghi \ --project my-project \ -c elasticsearch_index_prefix=new \ -c elasticsearch_index_days_max=100 ``` Update the `logs_user_config` block in your `aiven_service_integration` resource: ``` resource "aiven_service_integration" "logs_integration" { project = "PROJECT_NAME" integration_type = "logs" source_service_name = "SOURCE_SERVICE_NAME" destination_service_name = "OPENSEARCH_SERVICE_NAME" logs_user_config { elasticsearch_index_prefix = "UPDATED_INDEX_PREFIX" elasticsearch_index_days_max = UPDATED_INDEX_RETENTION_DAYS } } ``` Parameters: * `PROJECT_NAME`: Your Aiven project name * `SOURCE_SERVICE_NAME`: The service producing logs * `OPENSEARCH_SERVICE_NAME`: Your Aiven for OpenSearch service name * `UPDATED_INDEX_PREFIX`: Prefix for the index name, for example, `logs` * `UPDATED_INDEX_RETENTION_DAYS`: Number of days to keep logs, for example, `3` Run `terraform apply` to apply the changes. Update the `logs` section in your [`ServiceIntegration` resource](https://aiven.github.io/aiven-operator/resources/serviceintegration.html#spec.logs): ``` apiVersion: aiven.io/v1alpha1 kind: ServiceIntegration metadata: name: INTEGRATION_RESOURCE_NAME spec: authSecretRef: name: aiven-token key: TOKEN_NAME project: PROJECT_NAME integrationType: logs sourceServiceName: SOURCE_SERVICE_NAME destinationServiceName: OPENSEARCH_SERVICE_NAME logs: elasticsearch_index_prefix: UPDATED_INDEX_PREFIX elasticsearch_index_days_max: UPDATED_INDEX_RETENTION_DAYS ``` Apply the resource with `kubectl apply -f FILE_NAME.yaml`. ## Disable logs integration[​](#disable-logs-integration "Direct link to Disable logs integration") To stop sending logs from your service to Aiven for OpenSearch, disable the integration: * Console * API * CLI * Terraform * Kubernetes 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your integrated Aiven for OpenSearch service. 2. In the **Data** section, click **Integrations**. 3. On the **Integrations** page, find the service sending its logs to your Aiven for OpenSearch service. 4. Click **Actions** > **Disconnect**. 1) Get `INTEGRATION_ID` by calling the [ServiceGet](https://api.aiven.io/doc/#tag/Service/operation/ServiceGet) endpoint: ``` curl --request GET \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header "Authorization: Bearer API_TOKEN" ``` In the output, under `service_integrations`, find your integration and its `service_integration_id`. 2) Call the [ServiceIntegrationDelete](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationDelete) endpoint to delete the log integration: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration/INTEGRATION_ID \ --header "Authorization: Bearer API_TOKEN" ``` Parameters: * `PROJECT_NAME`: Your Aiven project name * `INTEGRATION_ID`: The integration ID * `API_TOKEN`: Your [personal token](/docs/platform/howto/create_authentication_token.md) Example: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/dev-sandbox/integration/123abc456def789ghi-123abc456def789ghi \ --header "Authorization: Bearer 123abc456def789ghi" ``` 1. Get `INTEGRATION_ID` by running the [`avn service integration-list`](/docs/tools/cli/service/integration.md#avn_service_integration_list) command: ``` avn service integration-list SOURCE_SERVICE_NAME --project PROJECT_NAME ``` Parameters: * `PROJECT_NAME`: Your Aiven project name * `SOURCE_SERVICE_NAME`: The service producing logs In the output table, find `SERVICE_INTEGRATION_ID` for the integrated services listed in columns `SOURCE` and `DEST`. 2. Use the [avn service integration-delete](/docs/tools/cli/service/integration.md#avn-service-integration-delete) command to delete the log integration: ``` avn service integration-delete INTEGRATION_ID \ --project PROJECT_NAME ``` Parameters: * `INTEGRATION_ID`: The integration ID * `PROJECT_NAME`: Your Aiven project name Example: ``` avn service integration-delete 123abc456def789ghi-123abc456def789ghi \ --project my-project ``` Remove the `aiven_service_integration` resource from your Terraform configuration: ``` # Remove or comment out this resource # resource "aiven_service_integration" "logs_integration" { # project = "PROJECT_NAME" # integration_type = "logs" # source_service_name = "SOURCE_SERVICE_NAME" # destination_service_name = "OPENSEARCH_SERVICE_NAME" # # logs_user_config { # elasticsearch_index_prefix = "INDEX_PREFIX" # elasticsearch_index_days_max = INDEX_RETENTION_DAYS # } # } ``` Run `terraform apply` to delete the integration. Delete the [`ServiceIntegration` resource](https://aiven.github.io/aiven-operator/resources/serviceintegration.html#spec.logs): ``` kubectl delete serviceintegration INTEGRATION_RESOURCE_NAME ``` or ``` kubectl delete -f FILE_NAME.yaml ``` --- # Write search queries with OpenSearch® and Python Learn how to write and run search queries on your OpenSearch cluster using a [Python OpenSearch client](https://github.com/opensearch-project/opensearch-py). For our data, we use a food recipe dataset [from Kaggle](https://www.kaggle.com/hugodarwood/epirecipes?select=full_format_recipes.json). After injecting this data into our cluster, we will write search queries to find different food recipes. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") ### GitHub repository[​](#github-repository "Direct link to GitHub repository") The code be found [in a GitHub repository](https://github.com/aiven/demo-opensearch-python). The files are organized according to their functions: * [config.py](https://github.com/aiven/demo-opensearch-python/blob/main/config.py), information to connect to the cluster * [index.py](https://github.com/aiven/demo-opensearch-python/blob/main/index.py), methods that manipulate the index * [search.py](https://github.com/aiven/demo-opensearch-python/blob/main/search.py), customized search query methods * [helpers.py](https://github.com/aiven/demo-opensearch-python/blob/main/helpers.py), response handler of search requests We use `Typer` Python [library](https://typer.tiangolo.com/) to create CLI commands to run from the terminal. To get the code on your machine and try the commands: 1. Clone the repository, go to the project directory, and install the dependencies ``` git clone https://github.com/aiven/demo-opensearch-python cd demo-opensearch-python pip install -r requirements.txt ``` note The repository includes a recipe dataset (`full_format_recipes.json`) from [Kaggle](https://www.kaggle.com/hugodarwood/epirecipes?select=full_format_recipes.json). If you encounter an `extra data` error when you load the data, download a fresh copy from Kaggle. Replace the file in the repository. ### Connect to the OpenSearch cluster with Python[​](#connect-to-the-opensearch-cluster-with-python "Direct link to Connect to the OpenSearch cluster with Python") Make sure to update the `SERVICE_URI` to your cluster `SERVICE_URI` in the `.env` [file](https://github.com/aiven/demo-opensearch-python/blob/main/.env) as explained in the [README](https://github.com/aiven/demo-opensearch-python). Once the environment variables are set, create an OpenSearch Python client to connect to your OpenSearch cluster using the [connection instructions](/docs/products/opensearch/howto/connect-with-python.md). You can see find the whole code sample in the [config.py](https://github.com/aiven/demo-opensearch-python/blob/main/config.py): ``` import os from dotenv import load_dotenv from opensearchpy import OpenSearch load_dotenv() INDEX_NAME = "epicurious-recipes" SERVICE_URI = os.getenv("SERVICE_URI") client = OpenSearch(SERVICE_URI, use_ssl=True) ``` tip The `SERVICE_URI` value can be found in the Aiven Console dashboard. After creating a client with a valid `SERVICE_URI`, you're set to interact with your cluster. ### Upload data to OpenSearch using Python[​](#upload-data-to-opensearch-using-python "Direct link to Upload data to OpenSearch using Python") Once you're connected, the next step should be to [inject data into our cluster](/docs/products/opensearch/howto/sample-dataset.md#load-data-with-python). This is done in our demo with the [`load_data` function](https://github.com/aiven/demo-opensearch-python/blob/main/index.py). You can inject the data to your cluster by running: ``` python index.py load-data ``` Once the data is loaded, we can [retrieve the data mapping](/docs/products/opensearch/howto/sample-dataset.md#get-mapping-with-python) to explore the structure of the data, with their respective fields and types. Find the code implementation in the [`get_mapping` function](https://github.com/aiven/demo-opensearch-python/blob/main/index.py). Check the structure of your data by running: ``` python index.py get-mapping ``` You should be able to see the fields' output: ``` [ 'calories', 'categories', 'date', 'desc', 'directions', 'fat', 'ingredients', 'protein', 'rating', 'sodium', 'title' ] ``` And the mapping with the fields and their respective types. ``` {'calories': {'type': 'float'}, 'categories': {'fields': {'keyword': {'ignore_above': 256, 'type': 'keyword'}}, 'type': 'text'}, 'date': {'type': 'date'}, 'desc': {'fields': {'keyword': {'ignore_above': 256, 'type': 'keyword'}}, 'type': 'text'}, 'directions': {'fields': {'keyword': {'ignore_above': 256, 'type': 'keyword'}}, 'type': 'text'}, 'fat': {'type': 'float'}, 'ingredients': {'fields': {'keyword': {'ignore_above': 256, 'type': 'keyword'}}, 'type': 'text'}, 'protein': {'type': 'float'}, 'rating': {'type': 'float'}, 'sodium': {'type': 'float'}, 'title': {'fields': {'keyword': {'ignore_above': 256, 'type': 'keyword'}}, 'type': 'text'}} ``` All set to start writing your search queries. ## Query the data[​](#query-the-data "Direct link to Query the data") ### Use the `search()` method[​](#use-the-search-method "Direct link to use-the-search-method") You have an OpenSearch client and data injected in your cluster, so you can start writing search queries. Python OpenSearch client has a handy method called `search()`, which we'll use to run our queries. We can check the method signature to understand the function and which parameters we'll use. As you can see, all the parameters are optional in the `search()` method. Find below the method signature: ``` client.search: (body=None, index=None, doc_type=None, params=None, headers=None) ``` To run the search queries, we'll use two of these parameters - `index` and `body`: * `index`, parameter refers to the name of the index we used to load the data. Therefore, it does not change. * `body`, parameter refers to the search query specifications. we'll modify it according to our query purpose. ### Lucene query and query DSL[​](#lucene-query-and-query-dsl "Direct link to Lucene query and query DSL") OpenSearch supports the **Lucene query syntax** to perform searches by using the `q` parameter. The `q` parameter expects a string with your query specifications, for example: ``` client.search({ index: 'recipes', q: 'ingredients:broccoli AND calories:(>=100 AND <200)' }) ``` For users, who prefer to work with nested objects and familiar structures like JSON (equivalent to Python dictionaries), OpenSearch supports the [query domain-specific language (DSL)](https://opensearch.org/docs/latest/opensearch/query-dsl/index/). For the **Query DSL**, the field `body` expects a dictionary object which can facilitate the construction of more complex queries depending on your use case, for example: ``` query_body = { "query": { "multi_match": { "query": "Garlic-Lemon", "fields": [ "title", "ingredients" ] } } } ``` In this example, we are searching for "Garlic-Lemon" across `title` and `ingredients` fields. Try out yourself using our demo: ``` python search.py multi-match title ingredients Garlic-Lemon ``` Check what comes out from this interesting combination 🧄 🍋 : ``` [ 'Garlic-Lemon Potatoes ', 'Lemon Garlic Mayonnaise ', 'Lemon Garlic Mayonnaise ', 'Garlic-Lemon Croutons ', 'Lemon-Garlic Vinaigrette ', 'Lemon-Garlic Lamb Chops ', 'Lemon Pepper Garlic Vinaigrette ', 'Lemon-Garlic Baked Shrimp ', 'Lemon-Herb Turkey with Lemon-Garlic Gravy ', 'Garlic, Oregano, and Lemon Vinaigrette ' ] ``` For this tutorial, we focus on the query DSL syntax to construct queries modifying the `body` parameter. In the method `search()`, one of the optional fields is the `size` field, which is defined as the number of results returned in the search. note The default value of the `size` field is 10, and we're using the default value in this tutorial. ## Write common queries[​](#write-common-queries "Direct link to Write common queries") In the next section, we cover some of the more common queries. Time to start querying 🔎 ### Create `match` query[​](#match-query "Direct link to match-query") The `match` query helps you to find the best matches with multiple search words. It is the default option for a [full-text search](https://opensearch.org/docs/latest/opensearch/query-dsl/full-text/). You can build your match query based on a `field` and the `query` that you are searching for. The DSL defaults to the "or" `operator`. ``` query_body = { "query": { "match": { field: { "query": query, "operator": operator } } } } ``` Thinking about how the match query works, if we run this query, it will return matches. This could be confusing because in our cluster the field `fat` corresponds to a value `float`, not a `string`. ``` query_body = { "query": { "match": { "fat": { "query": "0" } } } } ``` This is possible because [full-text queries](https://opensearch.org/docs/latest/opensearch/query-dsl/full-text/), such as the match query, use an analyzer to make the data optimized for search. As we have not specified an analyzer when we searched, the default standard analyzer is used: ``` query_body = { "query": { "match": { "fat": { "query": "0", "analyzer": "standard", } } } } ``` The default standard analyzer drops most punctuation, breaks up text into individual words, and lower cases them to optimize the search. If you want to choose a different analyzer, check out the available ones in the [OpenSearch documentation](https://opensearch.org/docs/latest/query-dsl/full-text/match/). You can find out how a customized match query can be written with your Python OpenSearch client in the [search\_match()](https://github.com/aiven/demo-opensearch-python/blob/main/search.py) function. You can run yourself the code to explore the `match` function. For example, if you want to find out recipes with the name "Spring" on them: ``` python search.py match title Spring ``` As a result of the "Spring" search recipes, you'll find: ``` [ 'Spring Fever ', 'Spring Rolls ', 'Spring Feeling ', 'Spring Fever ', 'Spring Rolls ', 'Spring Feeling ', 'Spring Vegetable Sauté ', 'Spring-Onion Cocktail ', 'Braised Spring Legumes ', 'Asian Spring Rolls ' ] ``` Find out more about [match queries](https://opensearch.org/docs/latest/query-dsl/full-text/match/). ### Use a `multi_match` query[​](#use-a-multi_match-query "Direct link to use-a-multi_match-query") One useful query when you want to align the `match` query properties but expand it to search in more fields is the `multi_match` query. You can add several fields in the `fields` property, to search for the `query` string across all those fields included in the list. ``` query_body = { "query": { "multi_match": { "query": query, "fields": [field1, field2 ...] } } } ``` In our demo, we have a function called [search\_multi\_match()](https://github.com/aiven/demo-opensearch-python/blob/main/search.py) that build customized multi match queries in Python. You can use our demo with `multi-match` keyword followed by the `fields` and the `query` to explore this type of query. Suppose you are looking for citrus recipes 🍋. For example, recipes with ingredients and lemon in the title, you can run your query from our [demo](https://github.com/aiven/demo-opensearch-python/) as: ``` python search.py multi-match title ingredients lemon ``` ### Match with phrases[​](#match-phrase-query "Direct link to Match with phrases") This query can be used to match **exact phrases** in a field. Where the `query` is the phrase that is being searched in a certain `field`: ``` query_body = { "query": { "match_phrase": { field: { "query": query } } } } ``` If you know exactly which phrases you're looking for, you can try out our `match-phrase` [search\_match\_phrase()](https://github.com/aiven/demo-opensearch-python/blob/main/search.py). note If you misspell the searched word, the query will not return any results as the purpose is to look for **exact phrases**. The lowercase and uppercase can bring your results according to the relevance For example, try searching for `pannacotta with lemon marmalade` in the title: ``` python search.py match-phrase title "Pannacotta with lemon marmalade" ``` If you just have a rough idea of the phrase you're looking for, you can make your match phrase query more flexible with the `slop` parameter as explained in the section [match phrase with slop query](/docs/products/opensearch/howto/opensearch-search-and-python.md#match-phrase-slop) section. ### Match phrases and add some `slop`[​](#match-phrase-slop "Direct link to match-phrase-slop") You can use the `slop` parameter to create more flexible searches. Suppose you're searching for `pannacotta marmalade` with the `match_phrase` query, and no results are found. This happens because you are looking for exact phrases, as discussed in [match phrase query](/docs/products/opensearch/howto/opensearch-search-and-python.md#match-phrase-query) section. You can expand your searches by configuring the `slop` parameter. The default value for the `slop` parameter is 0. The `slop` parameter allows to control the degree of disorder in your search as explained in the [OpenSearch documentation for the slop feature](https://opensearch.org/docs/latest/query-dsl/full-text/match/): > `slop` is the number of other words allowed between words in the query phrase. For example, to switch the order of two words requires two moves (the first move places the words atop one another), so to permit re-orderings of phrases, the slop must be at least two. A value of zero requires an exact match. You can construct a query and add some `slop` like this: ``` query_body = { "query": { "match_phrase": { field: { "query": query "slop": slop # integer or float } } } } ``` In the demo, you can find the [search\_slop()](https://github.com/aiven/demo-opensearch-python/blob/main/search.py) function where this query is used. Suppose you're looking for `pannacotta marmalade` phrase. To find more results rather than exact phrases, you should allow a certain degree. You can configure the `slop` to 2 , so it can find matches skipping **two words** between the searched ones. This is how you can run this query yourself: ``` python search.py slop "title" "pannacotta marmalade" 2 ``` Your result should look like this: ``` ['Lemon Pannacotta with Lemon Marmalade '] ``` So with `slop` parameter adjusted, you're may be able to find results even with other words in between the ones you searched. Read more about `slop` parameter on the [OpenSearch project specifications](https://opensearch.org/docs/latest/query-dsl/full-text/index/). ### Use a `term` query[​](#use-a-term-query "Direct link to use-a-term-query") If you want results with a precise value in a `field`, the [term query](https://opensearch.org/docs/latest/query-dsl/term/term/) is the right choice. The term query can be used to find documents according to a precise value such as a price or product ID, for example. This query can be constructed as: ``` query_body = { "query": { "term": { field: value } } } ``` In this query, the term is matched as it is, which means that no analyzer is applied to the search term. If you are searching for text field values, it is recommended to use [match query](/docs/products/opensearch/howto/opensearch-search-and-python.md#match-query) instead. You can look the [search\_term()](https://github.com/aiven/demo-opensearch-python/blob/main/search.py) function, which uses this query to build customized term queries. Run the search query yourself to find recipes with zero sodium on it, for example: ``` python search.py term sodium 0 ``` ### Search with a `range` query[​](#search-with-a-range-query "Direct link to search-with-a-range-query") This query helps to find documents that the field is within a provided range. This can be handy if you're dealing with **numerical values** and are interested **in ranges** instead of specific values. The queries can be constructed as: ``` query_body = { "query": { "range": { field: { "gte": gte, "lte": lte } } } } ``` You can construct range queries with combinations of inclusive and exclusive parameters as can be seen in the table: | Parameter | Behavior | | --------- | ------------------------ | | `gte` | Greater than or equal to | | `gt` | Greater than | | `lt` | Less than | | `lte` | Less than or equal to | Try to find recipes in a certain range of sodium, for example: ``` python search.py range sodium 0 10 ``` See more about the range query in the [OpenSearch documentation](https://opensearch.org/docs/latest/query-dsl/term/range/). ### Write fuzzy queries[​](#fuzzy-query "Direct link to Write fuzzy queries") This query looks for documents that have **similar term** to the searched term. This similarity is calculated by the `Levenshtein` [edit distance](https://en.wikipedia.org/wiki/Levenshtein_distance). This distance refers to the minimum number of single-character edits between two words. Some of those changes: * Change of a character: `post` → `lost` * Removal of a character: `eggs` → `ggs` * Insertion of a character: `edi` → `edit` * Transposition of two adjacent characters: `act` → `cat` The queries can be constructed as: ``` query_body = { "query": { "fuzzy": { field: { "value": value "fuzziness": fuzziness, } } } } ``` We can try out looking for a misspelled word and allowing some `fuzziness`. Writing a fuzzy query with a **misspelled word**, such as `pinapple` and setting `fuzziness` to zero. Running it, will bring no results: ``` python search.py fuzzy "title" "pinapple" 0 ``` To correct `pinapple` → `Pineapple` word, we only need to change one letter. So we can try again to search this word setting the `fuzziness` to one and run the search again. ``` python search.py fuzzy "title" "pinapple" 1 ``` As you can see, this search returns results 🍍: ``` [ 'Pineapple "Lasagna" ', 'Pineapple Bowl ', 'Pineapple Paletas ', 'Pineapple "Salsa" ', 'Pineapple Sangria ', 'Pineapple Tart ', 'Pineapple Split ', 'Roasted Pineapple with Star Anise Pineapple Sorbet ', 'Pineapple-Apricot Salsa ', 'Pineapple Papaya Relish ' ] ``` It is your turn, try out more combinations to better understand the fuzzy query. Related pages Want to try out OpenSearch with other clients? You can learn how to write search queries with NodeJS client, see [our tutorial](/docs/products/opensearch/howto/opensearch-and-nodejs.md). We created an OpenSearch cluster, connected to it, and tried out different types of search queries. Now, you can explore more resources to help you to learn other features of OpenSearch and its Python client. * [Demo repository](https://github.com/aiven/demo-opensearch-python), contains all code from this tutorial * [OpenSearch Python client](https://opensearch.org/docs/latest/clients/python/) * [How to use OpenSearch with curl](/docs/products/opensearch/howto/opensearch-with-curl.md) * [Official OpenSearch documentation](https://opensearch.org) --- # Connect to Aiven for OpenSearch® with cURL Connect to your Aiven for OpenSearch® service with [cURL](https://curl.se/). ## Variables[​](#variables "Direct link to Variables") These are the placeholders you will need to replace in the code samples: | Variable | Description | | ---------------- | -------------------------------------------------------------------------------------- | | `OPENSEARCH_URI` | Service URI, including username and password, available from the service overview page | ## Connect to OpenSearch[​](#connect-to-opensearch "Direct link to Connect to OpenSearch") Connect to your service with: ``` curl OPENSEARCH_URI ``` If the connection is successful, one of the nodes in your cluster will respond with some information including: * `name` The node name you are connected to (this will be the name of your service, with the node identifier suffix) * `version` The version information includes both the `distribution` and the `version` of the distribution that is running ## Manage indices[​](#manage-indices "Direct link to Manage indices") OpenSearch groups data into an index rather than a table. ### Create an index[​](#create-an-index "Direct link to Create an index") Create an index by making a `PUT` call to it: ``` curl -X PUT OPENSEARCH_URI/shopping-list ``` The response should have status 200 and the body data will have `acknowledged` set to true. If you already know something about the fields that will be in the documents you'll store, you can create an index with mappings to describe those known fields: ``` curl -X PUT -H "Content-Type: application/json" \ OPENSEARCH_URI/shopping-list \ -d '{ "mappings": { "properties": { "item": { "type": "keyword" }, "quantity": { "type": "integer" } } } }' ``` This example creates the shopping list example but adds information to help the indexer know how to handle the expected fields. ### List of indices[​](#list-of-indices "Direct link to List of indices") To list the indices do: ``` curl OPENSEARCH_URI/_cat/indices ``` ### Add an item to the index[​](#add-an-item-to-the-index "Direct link to Add an item to the index") OpenSearch is a document database so there is no enforced schema structure for the data you store. To add an item, `POST` the JSON data that should be stored: ``` curl -H "Content-Type: application/json" \ OPENSEARCH_URI/shopping-list/_doc \ -d '{ "item": "apple", "quantity": 2 }' ``` Other data fields don't need to match in format: ``` curl -H "Content-Type: application/json" \ OPENSEARCH_URI/shopping-list/_doc \ -d '{ "item": "bucket", "color": "blue", "quantity": 5, "notes": "the one with the metal handle" }' ``` These documents are stored in the `shopping-list` index. ## Search or retrieve data[​](#search-or-retrieve-data "Direct link to Search or retrieve data") OpenSearch is designed to make stored information easy to search and access (the clue is in the name). See the following examples to get you started, and refer to the official [OpenSearch Documentation](https://opensearch.org/docs/opensearch/index/) for more information. ### Search all items[​](#search-all-items "Direct link to Search all items") To list everything, run: ``` curl OPENSEARCH_URI/_search ``` Search results include some key fields to look at when you try this example: * `hits` holds the main payload of the response * `hits.hits` is a collection of results that matched the query. This includes: * `_index` the index that the document was stored in * `_score` how good a match the document is (on a scale of 0 to 1) * `_source` the document that matched ### Simple search[​](#simple-search "Direct link to Simple search") For the most simple search to match a string, you can use: ``` curl OPENSEARCH_URI/_search?q=apple ``` ### Advanced search options[​](#advanced-search-options "Direct link to Advanced search options") For more advanced searches, you can send a more detailed payload to specify which fields to search among other options: ``` curl -H "Content-Type: application/json" \ OPENSEARCH_URI/_search \ -d '{ "query": { "multi_match" : { "query" : "apple", "fields" : ["item", "notes"] } } }' ``` --- # Aiven for OpenSearch® metrics available via Prometheus Monitor and optimize your Aiven for OpenSearch service with metrics available via Prometheus. These metrics help track cluster health, replication status, and overall performance. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Enable Prometheus integration](/docs/platform/howto/integrations/prometheus-metrics.md). * Note the Prometheus **username** and **password** in the **Integration endpoints** section of the [Aiven Console](https://console.aiven.io/). ## Access Prometheus metrics[​](#access-prometheus-metrics "Direct link to Access Prometheus metrics") * View in browser * Retrieve via cURL 1. Open your service's **Overview** page in the [Aiven Console](https://console.aiven.io/). 2. In the **Connection information** section, click the **Prometheus** tab. 3. Copy the **Service URI**. 4. Paste the Service URI into your browser's address bar. 5. When prompted, enter your Prometheus credentials. 6. Click **Login**. To retrieve metrics, run the following `curl` command: ``` curl --user 'USERNAME:PASSWORD' PROMETHEUS_URL/metrics ``` Replace `USERNAME:PASSWORD` with your Prometheus credentials and `PROMETHEUS_URL` with the Service URI from the **Connection information** section. ## Host metrics[​](#host-metrics "Direct link to Host metrics") Host metrics provide insights into system-level performance, including CPU, memory, disk, and network usage. ### CPU utilization[​](#cpu-utilization "Direct link to CPU utilization") CPU utilization metrics offer insights into CPU usage. These metrics include time spent on different processes, system load, and overall uptime. | Metric | Description | | ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | | `cpu_usage_guest` | CPU time spent running a virtual CPU for guest operating systems | | `cpu_usage_guest_nice` | CPU time running low-priority virtual CPUs for guest operating systems; interrupted by higher-priority tasks and measured in hundredths of a second | | `cpu_usage_idle` | Time the CPU spends doing nothing | | `cpu_usage_iowait` | Time waiting for I/O to complete | | `cpu_usage_irq` | Time servicing interrupts | | `cpu_usage_nice` | Time running user-niced processes | | `cpu_usage_softirq` | Time servicing softirqs | | `cpu_usage_steal` | Time spent in other operating systems when running in a virtualized environment | | `cpu_usage_system` | Time spent running system processes | | `cpu_usage_user` | Time spent running user processes | | `system_load1` | System load average for the last minute | | `system_load15` | System load average for the last 15 minutes | | `system_load5` | System load average for the last 5 minutes | | `system_n_cpus` | Number of CPU cores available | | `system_n_users` | Number of users logged in | | `system_uptime` | Time for which the system has been up and running | ### Disk space utilization[​](#disk-space-utilization "Direct link to Disk space utilization") Disk space utilization metrics provide a snapshot of disk usage. These metrics include information about free and used disk space, as well as `inode` usage and total disk capacity. | Metric | Description | | ------------------- | ----------------------------- | | `disk_free` | Amount of free disk space | | `disk_inodes_free` | Number of free inodes | | `disk_inodes_total` | Total number of inodes | | `disk_inodes_used` | Number of used inodes | | `disk_total` | Total disk space | | `disk_used` | Amount of used disk space | | `disk_used_percent` | Percentage of disk space used | ### Disk input and output[​](#disk-input-and-output "Direct link to Disk input and output") Metrics such as `diskio_io_time` and `diskio_iops_in_progress` provide insights into disk I/O operations. These metrics cover read/write operations, the duration of these operations, and the number of bytes read/written. | Metric | Description | | ------------------------- | ------------------------------------------------------------------------------- | | `diskio_io_time` | Total time spent on I/O operations | | `diskio_iops_in_progress` | Number of I/O operations currently in progress | | `diskio_merged_reads` | Number of read operations that were merged | | `diskio_merged_writes` | Number of write operations that were merged | | `diskio_read_bytes` | Total bytes read from disk | | `diskio_read_time` | Total time spent on read operations | | `diskio_reads` | Total number of read operations | | `diskio_weighted_io_time` | Weighted time spent on I/O operations, considering their duration and intensity | | `diskio_write_bytes` | Total bytes written to disk | | `diskio_write_time` | Total time spent on write operations | | `diskio_writes` | Total number of write operations | ### Generic memory[​](#generic-memory "Direct link to Generic memory") The following metrics, including `mem_active` and `mem_available`, provide insights into your system's memory usage. | Metric | Description | | ----------------------- | ------------------------------------------------------------- | | `mem_active` | Amount of actively used memory | | `mem_available` | Amount of available memory | | `mem_available_percent` | Percentage of available memory | | `mem_buffered` | Amount of memory used for buffering I/O | | `mem_cached` | Amount of memory used for caching | | `mem_commit_limit` | Maximum amount of memory that can be committed | | `mem_committed_as` | Total amount of committed memory | | `mem_dirty` | Amount of memory waiting to be written to disk | | `mem_free` | Amount of free memory | | `mem_high_free` | Amount of free memory in the high memory zone | | `mem_high_total` | Total amount of memory in the high memory zone | | `mem_huge_pages_free` | Number of free huge pages | | `mem_huge_page_size` | Size of huge pages | | `mem_huge_pages_total` | Total number of huge pages | | `mem_inactive` | Amount of inactive memory | | `mem_low_free` | Amount of free memory in the low memory zone | | `mem_low_total` | Total amount of memory in the low memory zone | | `mem_mapped` | Amount of memory mapped into the process's address space | | `mem_page_tables` | Amount of memory used by page tables | | `mem_shared` | Amount of memory shared between processes | | `mem_slab` | Amount of memory used by the kernel for data structure caches | | `mem_swap_cached` | Amount of swap memory cached | | `mem_swap_free` | Amount of free swap memory | | `mem_swap_total` | Total amount of swap memory | | `mem_total` | Total amount of memory | | `mem_used` | Amount of used memory | | `mem_used_percent` | Percentage of used memory | | `mem_vmalloc_chunk` | Largest contiguous block of vmalloc memory available | | `mem_vmalloc_total` | Total amount of vmalloc memory | | `mem_vmalloc_used` | Amount of used vmalloc memory | | `mem_wired` | Amount of wired memory | | `mem_write_back` | Amount of memory being written back to disk | | `mem_write_back_tmp` | Amount of temporary memory being written back to disk | ### Network[​](#network "Direct link to Network") The following metrics, including `net_bytes_recv` and `net_packets_sent`, provide insights into your system's network operations. | Metric | Description | | ----------------------------- | ----------------------------------------------------------------- | | `net_bytes_recv` | Total bytes received on the network interfaces | | `net_bytes_sent` | Total bytes sent on the network interfaces | | `net_drop_in` | Incoming packets dropped | | `net_drop_out` | Outgoing packets dropped | | `net_err_in` | Incoming packets with errors | | `net_err_out` | Outgoing packets with errors | | `net_icmp_inaddrmaskreps` | Number of ICMP address mask replies received | | `net_icmp_inaddrmasks` | Number of ICMP address mask requests received | | `net_icmp_incsumerrors` | Number of ICMP checksum errors | | `net_icmp_indestunreachs` | Number of ICMP destination unreachable messages received | | `net_icmp_inechoreps` | Number of ICMP echo replies received | | `net_icmp_inechos` | Number of ICMP echo requests received | | `net_icmp_inerrors` | Number of ICMP messages received with errors | | `net_icmp_inmsgs` | Total number of ICMP messages received | | `net_icmp_inparmprobs` | Number of ICMP parameter problem messages received | | `net_icmp_inredirects` | Number of ICMP redirect messages received | | `net_icmp_insrcquenchs` | Number of ICMP source quench messages received | | `net_icmp_intimeexcds` | Number of ICMP time exceeded messages received | | `net_icmp_intimestampreps` | Number of ICMP timestamp reply messages received | | `net_icmp_intimestamps` | Number of ICMP timestamp request messages received | | `net_icmpmsg_intype3` | Number of ICMP type 3 (destination unreachable) messages received | | `net_icmpmsg_intype8` | Number of ICMP type 8 (echo request) messages received | | `net_icmpmsg_outtype0` | Number of ICMP type 0 (echo reply) messages sent | | `net_icmpmsg_outtype3` | Number of ICMP type 3 (destination unreachable) messages sent | | `net_icmp_outaddrmaskreps` | Number of ICMP address mask reply messages sent | | `net_icmp_outaddrmasks` | Number of ICMP address mask request messages sent | | `net_icmp_outdestunreachs` | Number of ICMP destination unreachable messages sent | | `net_icmp_outechoreps` | Number of ICMP echo reply messages sent | | `net_icmp_outechos` | Number of ICMP echo request messages sent | | `net_icmp_outerrors` | Number of ICMP messages sent with errors | | `net_icmp_outmsgs` | Total number of ICMP messages sent | | `net_icmp_outparmprobs` | Number of ICMP parameter problem messages sent | | `net_icmp_outredirects` | Number of ICMP redirect messages sent | | `net_icmp_outsrcquenchs` | Number of ICMP source quench messages sent | | `net_icmp_outtimeexcds` | Number of ICMP time exceeded messages sent | | `net_icmp_outtimestampreps` | Number of ICMP timestamp reply messages sent | | `net_icmp_outtimestamps` | Number of ICMP timestamp request messages sent | | `net_icmp_outratelimitglobal` | Number of globally rate-limited ICMP messages sent | | `net_icmp_outratelimithost` | Number of ICMP messages rate-limited per host | | `net_ip_defaultttl` | Default time-to-live for IP packets | | `net_ip_forwarding` | Indicates if IP forwarding is enabled | | `net_ip_forwdatagrams` | Number of forwarded IP datagrams | | `net_ip_fragcreates` | Number of IP fragments created | | `net_ip_fragfails` | Number of failed IP fragmentations | | `net_ip_fragoks` | Number of successful IP fragmentations | | `net_ip_inaddrerrors` | Number of incoming IP packets with address errors | | `net_ip_indelivers` | Number of incoming IP packets delivered to higher layers | | `net_ip_indiscards` | Number of incoming IP packets discarded | | `net_ip_inhdrerrors` | Number of incoming IP packets with header errors | | `net_ip_inreceives` | Total number of incoming IP packets received | | `net_ip_inunknownprotos` | Number of incoming IP packets with unknown protocols | | `net_ip_outdiscards` | Number of outgoing IP packets discarded | | `net_ip_outnoroutes` | Number of outgoing IP packets with no route available | | `net_ip_outrequests` | Total number of outgoing IP packets requested to be sent | | `net_ip_outtransmits` | Number of IP packets transmitted successfully | | `net_ip_reasmfails` | Number of failed IP reassembly attempts | | `net_ip_reasmoks` | Number of successful IP reassembly attempts | | `net_ip_reasmreqds` | Number of IP fragments received needing reassembly | | `net_ip_reasmtimeout` | Number of IP reassembly timeouts | | `net_packets_recv` | Total number of packets received on the network interfaces | | `net_packets_sent` | Total number of packets sent on the network interfaces | | `netstat_tcp_close` | Number of TCP connections in the CLOSE state | | `netstat_tcp_close_wait` | Number of TCP connections in the CLOSE\_WAIT state | | `netstat_tcp_closing` | Number of TCP connections in the CLOSING state | | `netstat_tcp_established` | Number of TCP connections in the ESTABLISHED state | | `netstat_tcp_fin_wait1` | Number of TCP connections in the FIN\_WAIT\_1 state | | `netstat_tcp_fin_wait2` | Number of TCP connections in the FIN\_WAIT\_2 state | | `netstat_tcp_last_ack` | Number of TCP connections in the LAST\_ACK state | | `netstat_tcp_listen` | Number of TCP connections in the LISTEN state | | `netstat_tcp_none` | Number of TCP connections in the NONE state | | `netstat_tcp_syn_recv` | Number of TCP connections in the SYN\_RECV state | | `netstat_tcp_syn_sent` | Number of TCP connections in the SYN\_SENT state | | `netstat_tcp_time_wait` | Number of TCP connections in the TIME\_WAIT state | | `netstat_udp_socket` | Number of UDP sockets | | `net_tcp_activeopens` | Number of active TCP open connections | | `net_tcp_attemptfails` | Number of failed TCP connection attempts | | `net_tcp_currestab` | Number of currently established TCP connections | | `net_tcp_estabresets` | Number of established TCP connections reset | | `net_tcp_incsumerrors` | Number of TCP checksum errors in incoming packets | | `net_tcp_inerrs` | Number of incoming TCP packets with errors | | `net_tcp_insegs` | Number of TCP segments received | | `net_tcp_maxconn` | Maximum number of TCP connections supported | | `net_tcp_outrsts` | Number of TCP reset packets sent | | `net_tcp_outsegs` | Number of TCP segments sent | | `net_tcp_passiveopens` | Number of passive TCP open connections | | `net_tcp_retranssegs` | Number of TCP segments retransmitted | | `net_tcp_rtoalgorithm` | TCP retransmission timeout algorithm | | `net_tcp_rtomax` | Maximum TCP retransmission timeout | | `net_tcp_rtomin` | Minimum TCP retransmission timeout | | `net_udp_ignoredmulti` | Number of UDP multicast packets ignored | | `net_udp_incsumerrors` | Number of UDP checksum errors in incoming packets | | `net_udp_indatagrams` | Number of UDP datagrams received | | `net_udp_inerrors` | Number of incoming UDP packets with errors | | `net_udp_memerrors` | Number of UDP packets dropped due to memory errors | | `net_udplite_ignoredmulti` | Number of UDP-Lite multicast packets ignored | | `net_udplite_incsumerrors` | Number of UDP-Lite checksum errors in incoming packets | | `net_udplite_indatagrams` | Number of UDP-Lite datagrams received | | `net_udplite_inerrors` | Number of incoming UDP-Lite packets with errors | | `net_udplite_memerrors` | Number of UDP-L | ### Kernel[​](#kernel "Direct link to Kernel") The metrics listed below, such as `kernel_boot_time` and `kernel_context_switches`, provide insights into the operations of your system's kernel. | Metric | Description | | ------------------------- | ----------------------------------------------------------- | | `kernel_boot_time` | Time at which the system was last booted | | `kernel_context_switches` | Number of context switches that have occurred in the kernel | | `kernel_entropy_avail` | Amount of available entropy in the kernel's entropy pool | | `kernel_interrupts` | Number of interrupts that have occurred | | `kernel_processes_forked` | Number of processes that have been forked | ### Process[​](#process "Direct link to Process") Metrics such as `processes_running` and `processes_zombies` provide insights into the management of the system's processes. | Metric | Description | | ------------------------- | ------------------------------------------------------------------------ | | `processes_blocked` | Number of processes that are blocked | | `processes_dead` | Number of processes that have terminated | | `processes_idle` | Number of processes that are idle | | `processes_paging` | Number of processes that are paging | | `processes_running` | Number of processes currently running | | `processes_sleeping` | Number of processes that are sleeping | | `processes_stopped` | Number of processes that are stopped | | `processes_total` | Total number of processes | | `processes_total_threads` | Total number of threads across all processes | | `processes_unknown` | Number of processes in an unknown state | | `processes_zombies` | Number of zombie processes (terminated but not reaped by parent process) | ### Swap usage[​](#swap-usage "Direct link to Swap usage") Metrics such as `swap_free` and `swap_used` provide insights into the usage of the system's swap memory. | Metric | Description | | ------------------- | ----------------------------------- | | `swap_free` | Amount of free swap memory | | `swap_in` | Amount of data swapped in from disk | | `swap_out` | Amount of data swapped out to disk | | `swap_total` | Total amount of swap memory | | `swap_used` | Amount of used swap memory | | `swap_used_percent` | Percentage of swap memory used | ## OpenSearch-specific metrics[​](#opensearch-specific-metrics "Direct link to OpenSearch-specific metrics") These metrics provide insights into the performance and health of your Aiven for OpenSearch service. ### Node statistics[​](#node-statistics "Direct link to Node statistics") Track node metrics such as CPU and memory usage, disk I/O, and JVM statistics. For more information, see the [node stats API](https://opensearch.org/docs/latest/api-reference/nodes-apis/nodes-stats/). ### Cluster statistics[​](#cluster-statistics "Direct link to Cluster statistics") Track cluster-level metrics, including the number of indices, shard distribution, and memory usage. For more information, see the [cluster stats API](https://opensearch.org/docs/latest/api-reference/cluster-api/cluster-stats/). ### Cluster health[​](#cluster-health "Direct link to Cluster health") Monitor cluster health with granular metrics available at the index level. Use `local` with `level=index` to view index-specific health status. For more information, see the [cluster health API](https://opensearch.org/docs/latest/api-reference/cluster-api/cluster-health/). ### Cross-cluster replication (CCR) stats [Limited availability](/docs/platform/concepts/service-and-feature-releases.md)[​](#cross-cluster-replication-ccr-stats- "Direct link to cross-cluster-replication-ccr-stats-") * **Leader stats:** Monitor replication metrics from the leader cluster, including replication lag and synchronization status. For more information, see the [leader cluster stats API](https://opensearch.org/docs/latest/tuning-your-cluster/replication-plugin/api/#get-leader-cluster-stats). * **Follower stats:** Monitor follower cluster metrics, including replication delays and error rates, to maintain data consistency. For more information, see the [follower cluster stats API](https://opensearch.org/docs/latest/tuning-your-cluster/replication-plugin/api/#get-follower-cluster-stats). --- # Upgrade Aiven for OpenSearch® Aiven for OpenSearch® allows you to choose the version that best fits your needs and upgrade when ready. ## Multi-version: how it works[​](#multi-version-how-it-works "Direct link to Multi-version: how it works") * Aiven for OpenSearch supports the [two latest upstream OpenSearch major versions](/docs/platform/reference/eol-for-major-versions.md#aiven-for-opensearch). * When Aiven releases a new minor version within a major version, it decommissions the previous minor version after six months, giving you time to test and migrate. * Patch version upgrades are automatic. For example, Aiven upgrades 3.3.1 to 3.3.2. ### Default version[​](#default-version "Direct link to Default version") When [creating an Aiven for OpenSearch service](/docs/products/opensearch/get-started.md#create-an-aiven-for-opensearch-service), you can select the starting version. See all available versions in [Versions of Aiven-managed services and tools](/docs/platform/reference/eol-for-major-versions.md#aiven-for-opensearch). ### Voluntary manual upgrades[​](#voluntary-manual-upgrades "Direct link to Voluntary manual upgrades") Except for a few [specific cases](/docs/products/opensearch/howto/os-version-upgrade.md#mandatory-automatic-upgrades), upgrades are voluntary and require [your action](/docs/products/opensearch/howto/os-version-upgrade.md#upgrade-your-service). * Upgrading between minor versions of the **same major version**: [upgrade](/docs/products/opensearch/howto/os-version-upgrade.md#upgrade-your-service) directly to the target version. * Upgrading between minor versions of two **different major versions**: 1. [Upgrade](/docs/products/opensearch/howto/os-version-upgrade.md#upgrade-your-service) to the latest major version. 2. [Upgrade](/docs/products/opensearch/howto/os-version-upgrade.md#upgrade-your-service) to a minor version of the latest major version. ### Mandatory automatic upgrades[​](#mandatory-automatic-upgrades "Direct link to Mandatory automatic upgrades") Although mandatory upgrades don't require your action, you can still [run](/docs/products/opensearch/howto/os-version-upgrade.md#upgrade-your-service) them yourself as soon as they become [available](/docs/products/opensearch/howto/os-version-upgrade.md#available-or-upcoming-upgrades). Otherwise, they are applied automatically during the maintenance window. * **Version patches**: `major.minor.patch1` > `major.minor.patch2`, for example, `2.19.1` > `2.19.2` * Scheduled and applied during your maintenance window * Possible after any voluntary manual minor version upgrade to ensure the patch version consistency between the cluster nodes * Cluster node version upgrades after **node replacements** in disaster recovery scenarios to ensure the version consistency between the cluster nodes * Upgrades due to **[versions reaching end-of-life](/docs/platform/reference/eol-for-major-versions.md#aiven-for-opensearch)** on the Aiven Platform ## Before you start[​](#before-you-start "Direct link to Before you start") ### Available or upcoming upgrades[​](#available-or-upcoming-upgrades "Direct link to Available or upcoming upgrades") Track upgrades for your service via: * [Aiven Console](https://console.aiven.io): Service **Overview** page > the **Maintenance** section > List of available mandatory and optional upgrades * Email: notifications for automated upgrades ### Downgrade restriction[​](#downgrade-restriction "Direct link to Downgrade restriction") Downgrades are not supported: You cannot revert to a previous version or change to a lower ### k-NN field requirements[​](#k-nn-field-requirements "Direct link to k-NN field requirements") If your indices use [k-NN fields](/docs/products/opensearch/reference/plugins.md), Aiven for OpenSearch checks them during specific upgrades and blocks the upgrade (`403 Forbidden`) if a check fails. The error response lists the affected index names. * **OpenSearch 1.x to 2.x**: Fails if any index has the legacy `index.knn: true` setting. Reindex the affected indices with an explicit `knn_vector` field mapping before upgrading. * **OpenSearch 1.x to 2.x**: Fails if any `knn_vector` field name contains a space or one of `" * \ < | , > / ?`. Rename the field or reindex the affected indices with valid field names before upgrading. * **OpenSearch 2.19 to 3.x**: Fails if any `knn_vector` field uses the deprecated `nmslib` engine. To resolve this, see [Migrate off the nmslib engine](/docs/products/opensearch/howto/migrate-knn-nmslib-engine.md). ### Prerequisites for upgrade[​](#prerequisites-for-upgrade "Direct link to Prerequisites for upgrade") To upgrade your service version, check that: * Your Aiven for OpenSearch service is running. * Target version to upgrade to is available for [manual upgrade](/docs/products/opensearch/howto/os-version-upgrade.md#available-or-upcoming-upgrades). * You can use one of the following tools to upgrade: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) * [Aiven Operator for Kubernetes®](/docs/tools/kubernetes.md) ## Upgrade your service[​](#upgrade-your-service "Direct link to Upgrade your service") * Console * CLI * API * Terraform * Kubernetes 1. In the [Aiven Console](https://console.aiven.io), go to your service. 2. On the **Overview** page, go to the **Maintenance** section. 3. Click **Actions** > **Upgrade version**. 4. Select a version to upgrade to, and click **Upgrade**. 1) Optional: Check available service versions using the [avn service versions](https://aiven.io/docs/tools/cli/service-cli#avn-service-versions) command. 2) Upgrade the service version using the [avn service update](https://aiven.io/docs/tools/cli/service-cli#avn-cli-service-update) command. Replace placeholders `PROJECT_NAME`, `SERVICE_NAME`, and `NEW_VERSION` as needed. ``` avn service update --project PROJECT_NAME SERVICE_NAME -c service_version=NEW_VERSION ``` Call the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint to set `service_version`. Replace placeholders `PROJECT_NAME`, `SERVICE_NAME`, `API_TOKEN`, and `NEW_VERSION` as needed. ``` curl --request PUT \ --url "https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME" \ --header "Authorization: Bearer API_TOKEN" \ --header "Content-Type: application/json" \ --data '{ "user_config": { "service_version": "NEW_VERSION" } }' ``` Use the [`aiven_opensearch`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/opensearch) resource to set [`opensearch_version`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/opensearch#opensearch_version-1). Use the [OpenSearch](https://aiven.github.io/aiven-operator/resources/opensearch.html) resource to set [`opensearch_version`](https://aiven.github.io/aiven-operator/resources/opensearch.html#spec.userConfig.opensearch_version-property). Related pages * [Reindex Aiven for OpenSearch data on a newer version](/docs/products/opensearch/howto/reindex-opensearch.md) * [Migrate off the nmslib k-NN engine](/docs/products/opensearch/howto/migrate-knn-nmslib-engine.md) * [Available plugins for Aiven for OpenSearch](/docs/products/opensearch/reference/plugins.md) * [Versions of Aiven-managed services and tools](/docs/platform/reference/eol-for-major-versions.md) --- # Power on/off and delete your Aiven for OpenSearch® service Power off your Aiven for OpenSearch® service to release resources and save credits, power it back on when you need it, or delete it permanently. ## Power off a service[​](#power-off-a-service "Direct link to Power off a service") When you power off a service: * All virtual machines are removed from the public cloud. * The service configuration is stored on the Aiven Platform. * If there are no backups, all service data is lost. * If the service has time-based or point in time recovery backups, the backups remain on the Aiven Platform. Services powered off for more than 180 days are automatically deleted. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power off service**. To power off a service, run: ``` avn service update SERVICE_NAME --power-off ``` ## Power on a service[​](#power-on-a-service "Direct link to Power on a service") When you power on a service: * New virtual machines are created on the service's public cloud. * The service starts with the stored configuration parameters. * The latest time-based backup is restored. * Maintenance updates are automatically applied. * If a point in time recovery backup is available, the database transaction logs are replayed to recover the service data to a specific point in time. The restoration takes from a few minutes to a few hours, depending on the network bandwidth, the disk IOPS allocated to the service, and the size of the backup. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power on service**. To power on a service, run: ``` avn service update SERVICE_NAME --power-on ``` To see when the service is running, run: ``` avn service wait SERVICE_NAME ``` note When you power on an Aiven for OpenSearch service, your indices are restored from the latest available backup. note Static IP addresses are not removed when a service is powered off or deleted. They continue to generate the usual costs. To avoid these costs, [remove the static IP addresses](/docs/platform/concepts/static-ips.md). ## Delete a service[​](#delete-a-service "Direct link to Delete a service") * Console * CLI 1. In your project, click **Services**. 2. Open the service to delete, and click **Actions** > **Delete service**. To delete a service, run: ``` avn service terminate SERVICE_NAME ``` Related pages * [Restore an Aiven for OpenSearch® backup](/docs/products/opensearch/howto/restore_opensearch_backup.md) * [Fork Aiven for OpenSearch®](/docs/products/opensearch/howto/fork-service.md) --- # Prepare your Aiven for OpenSearch® service for high load Prepare your Aiven for OpenSearch® service for higher than usual traffic to avoid outages and keep performance stable. Prepare your services for higher than usual traffic to avoid service outages by doing the following: * **Subscribe to service notifications:** To receive notifications about service health and warnings when resources are low, you can [set service and project contacts](https://aiven.io/docs/platform/howto/technical-emails). You can also view the status of the Aiven Platform and get updates on incidents on the [status page](https://status.aiven.io/). Follow the RSS feed, subscribe to email or SMS notifications, or use the Slack integration to get notifications about incidents. * **Monitor your services:** [Monitor the health of your services](/docs/platform/howto/list-monitoring.md) using metrics, logs, alerts, and dashboards. * **Scale your services:** If you forecast a load that can't be handled by the service, you can scale up your service. * **Set the backup schedule:** To minimize the impact of the higher load during the backup process, schedule backups outside of peak traffic hours. * **Set the maintenance window:** Schedule maintenance updates outside of your peak traffic hours. * **Run load tests on service forks:** To test the impact of high traffic on a production service, fork the service and run your load test on the fork. Additionally, optimizing a service allows it to perform better under stress therefore avoiding the need of an upgrade. The more optimized a service is for your usage, the better you can weather spikes in traffic. Related pages * [Change the service plan](/docs/products/opensearch/howto/change-service-plan.md) * [Scale disk storage](/docs/products/opensearch/howto/scale-disk-storage.md) --- # Reindex Aiven for OpenSearch® data on a newer version When upgrading Aiven for OpenSearch® to a newer version, reindex indices created with an earlier version to ensure compatibility with the target version. ## Why reindexing is required[​](#why-reindexing-is-required "Direct link to Why reindexing is required") In a production environment, reindexing is a fundamental administrative task for changing the underlying structure of your data, not only for version upgrades. Because Aiven for OpenSearch uses immutable Lucene segments, certain changes require creating a new index and moving data into it. ### Version upgrades[​](#version-upgrades "Direct link to Version upgrades") Newer Aiven for OpenSearch versions can introduce compatibility requirements where indices require a minimum version. Upgrading to a newer version with indices created in an incompatible earlier version can cause the upgrade to fail. Version compatibility rules An index version cannot be more than one major version lower than the target service version. For example: * An index with version `1.3.2` is compatible with service version `2.19.4` but not with `3.3.0` because the difference in the major version is greater than one. * Upgrading a service to version 3.x requires all indices to have version `>=2.0.0`. * Services with old Elasticsearch® indices at versions `6.x` or `7.x` require first upgrading those indices to OpenSearch version `2.19` before upgrading the service to version 3.x. This process is the same as for `1.x` indices. To upgrade when reindexing is required: 1. Upgrade your service to an intermediate compatible version if needed. 2. Reindex all indices created with incompatible earlier versions. 3. Upgrade to the target version. ### Mapping transformations[​](#mapping-transformations "Direct link to Mapping transformations") You might need reindexing when changing a field type, for example, when converting a text field to a keyword for exact matching. ### Static setting updates[​](#static-setting-updates "Direct link to Static setting updates") Common reasons for reindexing are changing the number of primary shards and updating custom analyzers and tokenizers. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * You have an Aiven for OpenSearch service at an intermediate version. * You have identified indices created with earlier versions that need reindexing. * You have the service connection credentials. * Your cluster health is green, and all applications writing to the index are stopped if possible. * You have enough storage space for both the source index and the destination index, plus transient merge space required by the segments. * You have required permissions to create and delete indices. note In the examples, * `$OS_URI` is used for the service connection URL (for example, `https://USER:PASSWORD@HOST:PORT`). * `$OLD_INDEX_NAME` is used for the index to be reindexed. * `$NEW_INDEX_NAME` is used for the target index. ## Identify indices requiring reindexing[​](#identify-indices-requiring-reindexing "Direct link to Identify indices requiring reindexing") Check which indices were created with an earlier version of Aiven for OpenSearch: ``` curl "https://USER:PASSWORD@HOST:PORT/_all/_settings?filter_path=*.settings.index.version.created_string,*.settings.index.creation_date_string&human=true" ``` Replace `USER`, `PASSWORD`, `HOST`, and `PORT` with your service connection details. This lists the indices with their creation version and date. Identify indices with versions that are incompatible with your target upgrade version. ## Choose your reindexing strategy[​](#choose-your-reindexing-strategy "Direct link to Choose your reindexing strategy") There are two primary patterns for reindexing. Your choice depends on whether you can use aliases. ### Strategy A: blue-green swap[​](#strategy-a-blue-green-swap "Direct link to Strategy A: blue-green swap") Create an index alongside your existing index and move a pointer (alias) once the data is synchronized. * **How it works**: Data is copied to a new version, for example, v2. Once verified, the application alias is updated to point to v2 in a single, atomic operation. * **Pros**: Instant rollback, low downtime, and source data remains untouched until the end. * **Use when**: Your application uses aliases to reference indices. ### Strategy B: double reindex[​](#strategy-b-double-reindex "Direct link to Strategy B: double reindex") Use when you cannot use aliases. * **How it works**: Data is moved to a temporary buffer index, the original index is deleted and recreated with new settings, and data is reindexed back to the original name. * **Pros**: No application code changes required. * **Cons**: Involves a window of downtime. * **Use when**: Your application references indices by name and cannot use aliases. ## Check storage availability[​](#check-storage-availability "Direct link to Check storage availability") Before starting a reindex, verify you have enough storage space by checking how Aiven for OpenSearch views its disk watermarks. ### Check disk usage per node[​](#check-disk-usage-per-node "Direct link to Check disk usage per node") The `_cat/allocation` API provides a view of available space across your cluster: ``` curl -s "$OS_URI/_cat/allocation?v&s=disk.avail:desc" ``` Check these values: * `disk.indices`: the amount of space taken by your data * `disk.avail`: the remaining space on the node * `disk.percent`: your current usage percentage ### Check flood stage watermarks[​](#check-flood-stage-watermarks "Direct link to Check flood stage watermarks") Aiven for OpenSearch blocks all writes (including reindexing) if a node hits the flood stage watermark. By default, this is 95%. Check if your cluster has custom settings: ``` curl -s "$OS_URI/_cluster/settings?include_defaults=true" | jq ' .defaults.cluster.routing.allocation.disk.watermark, .persistent.cluster.routing.allocation.disk.watermark, .transient.cluster.routing.allocation.disk.watermark ' ``` ### Check index-specific storage[​](#check-index-specific-storage "Direct link to Check index-specific storage") See the primary storage and total storage (primaries plus replicas): ``` curl -s "$OS_URI/_cat/indices/$OLD_INDEX_NAME?v&h=index,docs.count,pri.store.size,store.size" ``` Check these values: * `pri.store.size`: the size of your unique data (the primary shards) * `store.size`: the total space on disk, including replicas ### Verify sufficient storage[​](#verify-sufficient-storage "Direct link to Verify sufficient storage") If current disk usage plus 1.5 times the index size (with replicas) does not push the disk usage over the watermark levels (specifically the flood stage), you can proceed. warning As a rule of thumb, if your `disk.percent` is already above 70%, do not start a reindex of a large index without first increasing your plan's storage. The reindex process creates new segments before deleting the old ones, causing a temporary storage spike. ## Reindex earlier-version indices[​](#reindex-earlier-version-indices "Direct link to Reindex earlier-version indices") For each index created with an earlier version of Aiven for OpenSearch, follow these steps: ### 1. Export source index configuration[​](#1-export-source-index-configuration "Direct link to 1. Export source index configuration") Reindexing does not automatically preserve the settings and mappings of the source index. Capture the current settings as the source of truth. Export the complete definition of your existing index: ``` curl -s "$OS_URI/$OLD_INDEX_NAME" > original_state.json ``` ### 2. Create configuration for the new index[​](#2-create-configuration-for-the-new-index "Direct link to 2. Create configuration for the new index") Keep your custom mappings and analyzers while stripping out system-generated metadata that belongs only to the old index instance: ``` jq '.[0][value] | { settings: { index: ( .settings.index | del( .uuid, .version, .creation_date, .provided_name, .creation_date_string, .store ) ) }, mappings: .mappings }' original_state.json > new_index_request.json ``` Alternatively, manually create an index with updated settings: ``` curl -X PUT "$OS_URI/$NEW_INDEX_NAME" \ -H 'Content-Type: application/json' \ -d '{ "settings": { "number_of_shards": 1, "number_of_replicas": 1 } }' ``` Adjust `number_of_shards` and `number_of_replicas` based on your requirements. ### 3. Initialize the destination index[​](#3-initialize-the-destination-index "Direct link to 3. Initialize the destination index") Use the sanitized settings to create the destination index: ``` curl -s -X PUT "$OS_URI/$NEW_INDEX_NAME" \ -H 'Content-Type: application/json' \ -d @new_index_request.json ``` For large indices, consider changing the index refresh interval to `-1` to speed up reindexing. Change it back once reindexing completes. If you created the index manually, apply the mapping. Extract only the `properties` object from the mapping response: ``` curl -X PUT "$OS_URI/$NEW_INDEX_NAME/_mapping" \ -H 'Content-Type: application/json' \ -d '{ "properties": { "example_field": { "type": "text" } } }' ``` ### 4. Make the source index read-only[​](#4-make-the-source-index-read-only "Direct link to 4. Make the source index read-only") Optionally, prevent the moving target problem by making the source index read-only. This prevents writes during reindexing. ``` curl -s -X PUT "$OS_URI/$OLD_INDEX_NAME/_settings" \ -H 'Content-Type: application/json' \ -d '{ "index.blocks.write": true }' ``` warning Applications trying to write to the source index will receive a `403 Forbidden` error. ### 5. Reindex the data[​](#5-reindex-the-data "Direct link to 5. Reindex the data") Use the Reindex API to copy data from the old index to the new index. For large indices that might take a long time, use asynchronous reindexing to prevent request timeouts: ``` TASK_ID=$(curl -s -X POST "$OS_URI/_reindex?wait_for_completion=false&slices=auto" \ -H 'Content-Type: application/json' \ -d "{ \"source\": {\"index\": \"$OLD_INDEX_NAME\"}, \"dest\": {\"index\": \"$NEW_INDEX_NAME\"} }" | jq -r '.task') echo "Reindex Task started: $TASK_ID" ``` The `wait_for_completion=false` parameter allows the reindex to continue in the background. The reindex returns a task ID for monitoring progress. For synchronous reindexing of smaller indices: ``` curl -s -X POST "$OS_URI/_reindex" \ -H 'Content-Type: application/json' \ -d '{ "source": { "index": "$OLD_INDEX_NAME" }, "dest": { "index": "$NEW_INDEX_NAME" } }' ``` For large indices, consider using these additional parameters: * Slicing for performance * Monitor async reindexing * Batch size control Use slicing to parallelize the reindexing process: ``` curl -s -X POST "$OS_URI/_reindex?slices=5&refresh" \ -H 'Content-Type: application/json' \ -d '{ "source": { "index": "$OLD_INDEX_NAME" }, "dest": { "index": "$NEW_INDEX_NAME" } }' ``` The `slices` parameter splits the reindexing into multiple subtasks. Use a value equal to the number of shards for optimal performance. Monitor the reindex task by periodically running: ``` curl -s "$OS_URI/_tasks/$TASK_ID" | jq '.task.status' ``` When the task completes, it disappears from the `_tasks` endpoint (the request returns `404`). If the task is successful, its result is stored in the `.tasks` index for a short period. You can also use: ``` GET /_tasks/TASK_ID ``` Control the batch size to manage memory usage: ``` curl -s -X POST "$OS_URI/_reindex" \ -H 'Content-Type: application/json' \ -d '{ "source": { "index": "$OLD_INDEX_NAME", "size": 1000 }, "dest": { "index": "$NEW_INDEX_NAME" } }' ``` The `size` parameter specifies how many documents to process in each batch. ### 6. Verify the reindexing[​](#6-verify-the-reindexing "Direct link to 6. Verify the reindexing") Check that all documents are copied successfully: ``` curl -s "$OS_URI/$OLD_INDEX_NAME/_count" curl -s "$OS_URI/$NEW_INDEX_NAME/_count" ``` The document counts should match. ### 7. Finalize the reindex[​](#7-finalize-the-reindex "Direct link to 7. Finalize the reindex") Complete the reindexing process based on your chosen strategy. #### Blue-green approach[​](#blue-green-approach "Direct link to Blue-green approach") Update aliases to point to the newly created index: ``` curl -s -X POST "$OS_URI/_aliases" \ -H 'Content-Type: application/json' \ -d '{ "actions": [ { "remove": { "index": "$OLD_INDEX_NAME", "alias": "my_alias" } }, { "add": { "index": "$NEW_INDEX_NAME", "alias": "my_alias" } } ] }' ``` If you modified `refresh_interval`, set it back to the original value on the target index. After verifying that your application works correctly with the new index, delete the old index: ``` curl -s -X DELETE "$OS_URI/$OLD_INDEX_NAME" ``` #### Double reindex approach[​](#double-reindex-approach "Direct link to Double reindex approach") Before repeating the reindexing, ensure no applications are doing write or delete operations targeting the index. 1. Optionally, clone the original source before deleting it (the index must be read-only): ``` curl -s -X POST "$OS_URI/$OLD_INDEX_NAME/_clone/${OLD_INDEX_NAME}_backup" ``` 2. Delete the original source index and redo the reindex steps using the freshly created new index as the source and the original source index as the target. 3. After verifying that the second reindexing succeeded, remove the temporary index and the backup. ## Complete the upgrade[​](#complete-the-upgrade "Direct link to Complete the upgrade") After reindexing all indices created with earlier versions: 1. Verify all indices have a compatible version. 2. [Upgrade your service](/docs/products/opensearch/howto/os-version-upgrade.md) to the target version. ## ISM plugin caveats[​](#ism-plugin-caveats "Direct link to ISM plugin caveats") Reindexing does not consider lifecycle management provided by the ISM plugin. Issues that can arise after reindexing: ### Orphaned index[​](#orphaned-index "Direct link to Orphaned index") When you create an index and move data into it, the ISM plugin sees it as a new entity. Unless you explicitly attach a policy during creation (using a template or a manual API call), the new index has no lifecycle management. With the double reindex approach, the deletion of the original source index purges all ISM metadata associated with that name. Even if you recreate the index with the same name, the ISM plugin doesn't recognize it. Re-run the `_plugins/_ism/add` command to ensure the index is managed. ### Clock reset[​](#clock-reset "Direct link to Clock reset") Most ISM policies calculate the age of data based on `index.creation_date`. Reindexing creates an index today, resetting this date. ### Policy state reset[​](#policy-state-reset "Direct link to Policy state reset") ISM policies are stateful. A policy might be in a warm state waiting to move to cold. You cannot migrate the state of a policy from one index to another. The new index starts at the initial state (usually hot). Related pages * [Upgrade Aiven for OpenSearch](/docs/products/opensearch/howto/os-version-upgrade.md) * [Manage large shards in Aiven for OpenSearch](/docs/products/opensearch/howto/resolve-shards-too-large.md) * [OpenSearch® reindex API documentation](https://opensearch.org/docs/latest/api-reference/document-apis/reindex/) --- # Rename your Aiven for OpenSearch® service Change the name of your Aiven for OpenSearch® service by forking it under a new name and deleting the original service. You cannot rename a service after creation. Instead, you can create a fork with the new name and delete the original service. ## Rename a service[​](#rename-a-service "Direct link to Rename a service") 1. Stop writing to the service. 2. Fork the service. 3. Add any integrations or SSO configurations that weren't copied. 4. Connect your clients to the new service. 5. Test the forked service. 6. Delete the original service. note Single sign-on (SSO) methods are not copied to the forked service. Reconfigure them on the new service to avoid disrupting user access. Related pages * [Fork Aiven for OpenSearch®](/docs/products/opensearch/howto/fork-service.md) * [Power on/off and delete your Aiven for OpenSearch® service](/docs/products/opensearch/howto/power-cycle-service.md) --- # Manage large shards in Aiven for OpenSearch® Resolve the large shard size alert in Aiven for OpenSearch® by deleting old data, splitting an index, or reindexing it with more shards. OpenSearch doesn't enforce a shard size limit, but for guidance on the recommended range, see [Optimal number of shards](/docs/products/opensearch/concepts/shards-number.md). Shards that grow too large can fail to relocate or recover, which risks data loss. Aiven for OpenSearch monitors shard sizes for all services. If a shard exceeds the recommended size, you get a notification through the `user_alert_resource_usage_es_shard_too_large` alert. Use one of the following options to resolve it. ## Delete old records and force merge[​](#delete-old-records-and-force-merge "Direct link to Delete old records and force merge") If your application allows it, permanently delete old or unnecessary records from the index. For example, delete records older than five days: ``` curl -X POST "https://USER:PASSWORD@HOST:PORT/INDEX_NAME/_delete_by_query" \ -H 'Content-Type: application/json' \ -d '{ "query": { "range": { "@timestamp": { "lte": "now-5d" } } } }' ``` Deleting documents doesn't reduce disk usage on its own. OpenSearch only marks matching documents as deleted; it removes them from disk the next time it merges the underlying segments. To reclaim the space immediately, force a merge after the deletion completes and write traffic to the index has stopped: ``` curl -X POST "https://USER:PASSWORD@HOST:PORT/INDEX_NAME/_forcemerge?max_num_segments=1" ``` warning Both operations temporarily increase disk usage before they reduce it. Deleting the records creates new segments to record the deletions, and the force merge rewrites the remaining segments into new ones before removing the old ones. Setting `max_num_segments` to `1` can temporarily double the shard's disk usage. Confirm you have enough free disk space for this spike before you start. ## Split the index[​](#split-the-index "Direct link to Split the index") Use the Split API to create an index with a multiple of the current shard count. The new index keeps the source index's mappings and settings automatically, so you don't need to recreate them. 1. Make the source index read-only: ``` curl -X PUT "https://USER:PASSWORD@HOST:PORT/INDEX_NAME/_settings" \ -H 'Content-Type: application/json' \ -d '{ "index.blocks.write": true }' ``` 2. Split it into a new index. The target shard count must be a multiple of the source shard count. For example, 2 shards can split into 4, 6, or 8, and 3 shards can split into 6, 9, or 12: ``` curl -X POST "https://USER:PASSWORD@HOST:PORT/INDEX_NAME/_split/NEW_INDEX_NAME" \ -H 'Content-Type: application/json' \ -d '{ "settings": { "index.number_of_shards": 4, "index.blocks.write": null } }' ``` 3. The split creates an index under a new name. If your application references the index by name rather than through an alias, add an alias so existing clients keep working without changes. Verify that INDEX\_NAME and NEW\_INDEX\_NAME have the same document count first, because the `remove_index` action below deletes INDEX\_NAME outright rather than only detaching the alias: ``` curl -X POST "https://USER:PASSWORD@HOST:PORT/_aliases" \ -H 'Content-Type: application/json' \ -d '{ "actions": [ { "remove_index": { "index": "INDEX_NAME" } }, { "add": { "index": "NEW_INDEX_NAME", "alias": "INDEX_NAME" } } ] }' ``` Both actions apply in the same request, so there's no window where INDEX\_NAME resolves to nothing, and existing clients keep reading and writing through INDEX\_NAME unchanged. For more information, see [Aliases](/docs/products/opensearch/concepts/indices.md#aliases). If the security plugin is enabled, ACL rules matching INDEX\_NAME don't automatically extend to its alias. For more information, see [Access control for aliases](/docs/products/opensearch/concepts/access_control.md#access-control-for-aliases). ## Reindex with more shards[​](#reindex-with-more-shards "Direct link to Reindex with more shards") You can also create an index with the shard count you want and copy the data across with the Reindex API. Unlike the Split API, reindexing doesn't carry over the source index's mappings and settings automatically. Capture and reapply them yourself, or follow the full procedure in [Reindex Aiven for OpenSearch® data on a newer version](/docs/products/opensearch/howto/reindex-opensearch.md), which covers this end to end, including exporting and reapplying the original settings. Related pages * [Optimal number of shards](/docs/products/opensearch/concepts/shards-number.md) * [Manage indices in Aiven for OpenSearch®](/docs/products/opensearch/concepts/indices.md) * [Reindex Aiven for OpenSearch® data on a newer version](/docs/products/opensearch/howto/reindex-opensearch.md) --- # Restore an OpenSearch® backup Depending on your service plan, you can restore OpenSearch® backups from a specific day or hour. To restore a backup: 1. In your service, in the **Backups** section, click **Backup management**. 2. Click **Fork & restore**. 3. Choose the backup to fork from. 4. Enter a name, and select the cloud and plan. 5. Click **Create fork**. Once the new service is running, change your application's connection settings to point to it and power off the original service. --- # Enable SAML authentication on Aiven for OpenSearch® SAML (Security Assertion Markup Language) is a standard protocol for exchanging authentication and authorization data between an identity provider (IdP) and a Service Provider (SP). SAML enables users to authenticate themselves to a service provider with credentials from a trusted third-party identity provider without the need to create and manage separate user accounts for each service provider. SAML authentication on Aiven for OpenSearch® can enhance the authentication process for users, providing increased security and a more streamlined experience. OpenSearch can delegate authentication and authorization to a trusted external identity provider, reducing security risks and simplifying user management. Additionally, this allows for Single Sign-On (SSO) functionality, enabling users to access several OpenSearch instances without the need to log in multiple times. important When you fork an Aiven for OpenSearch® service, any Single Sign-On (SSO) methods configured at the service level, such as SAML, must be explicitly reconfigured for the forked service. SSO configurations are linked to specific URLs and endpoints, which change during forking. Failing to reconfigure SSO methods for the forked service can lead to authentication problems and potentially disrupt user access. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven for OpenSearch® version 2.4 or later is required. If you are using an earlier version, upgrade to the latest version. * OpenSearch Security management must be [enabled](/docs/products/opensearch/howto/enable-opensearch-security.md) on the Aiven for OpenSearch® service. * You will need a SAML identity provider (IdP), the Metadata URL, and IdP entity ID. ## Configure SAML on IdP[​](#configure-saml-on-idp "Direct link to Configure SAML on IdP") To enable SAML SSO Authentication for Aiven for OpenSearch, configure SAML with an Identity Provider (IdP). As Aiven for OpenSearch is designed to work with various IdPs, the configuration steps may differ depending on your IdP. Refer to your Identity Provider's documentation for detailed instructions on configuring SAML applications. To enable SAML SSO authentication, make sure you have correctly configured your IdP and have the following two critical parameters: * **IdP Metadata URL**: The IdP metadata URL provides essential metadata about your IdP, including the certificate used for signing the SAML response. * **IdP Entity ID** : The IdP Entity ID is the identifier that the IdP uses to recognize itself. To establish trust between Aiven and OpenSearch, Aiven uses the Entity ID value for OpenSearch. You can find the *IdP Entity ID* in your Identity Provider's metadata or configuration settings. ## Enable SAML SSO authentication via Aiven Console[​](#enable-saml-sso-authentication-via-aiven-console "Direct link to Enable SAML SSO authentication via Aiven Console") To enable SAML authentication for your Aiven for OpenSearch service: 1. On your Aiven for OpenSearch service, select **Users** from the left sidebar. 2. In the **SAML SSO Authentication** section, select **Enable SAML**. 3. On the **Configure SAML Authentication** screen, enter the following details: * **SSO URL**: This is a distinct URL assigned to each Aiven for OpenSearch service and serves as the destination where the Identity Provider (IdP) sends SAML responses after a user has been authenticated successfully. It is a fixed URL that cannot be modified. * **IdP Metadata URL**: Enter the URL of your SAML Identity Provider's (IdP) metadata that Aiven will use to authenticate users. * **IdP Entity ID**: Enter the unique identifier assigned to your Identity Provider (IdP). This identifier assists Aiven in distinguishing between various IdPs. * **SP Entity ID**: Enter a unique identifier the IdP uses to recognize and authenticate Aiven for OpenSearch as the service provider. note The SP Entity ID can be any arbitrary value defined by the user. Additionally, OpenSearch suggests creating a new application for OpenSearch Dashboards and using the URL of your OpenSearch Dashboards as the SP entity ID. * **SAML roles key**: This is an optional field that allows you to map SAML roles to Aiven for OpenSearch roles. * **SAML subject key**: This is also an optional field that allows you to map SAML subject to Aiven for OpenSearch users. 4. Select **Enable**. 5. In the **SAML SSO Authentication** section, you can see the SAML method configured with the status set to **Enabled**. --- # Sample dataset Databases are more fun with data, so to get you started on your OpenSearch® journey we picked this open data set of recipes as a great example you can try out yourself. ## Epicurious recipes[​](#epicurious-recipes "Direct link to Epicurious recipes") A dataset from [Kaggle](https://www.kaggle.com/hugodarwood/epirecipes) with recipes, rating and nutrition information from [Epicurious](https://www.epicurious.com). Let's take a look at a sample recipe document: ``` { "title": "A very nice Vegan dish", "desc": "A beautiful description of the recipe", "date": "2015-05-01T04:00:00.000Z", "categories": [ "Vegan", "Tree Nut Free", "Soy Free", "No Sugar Added" ], "ingredients": [ "list", "of", "ingredients" ], "directions": [ "list", "of", "steps", "to prepare the dish" ], "calories": 32.0, "fat": 1.0, "protein": 1.0, "rating": 5.0, "sodium": 959.0, } ``` ## Load the data with Python[​](#load-data-with-python "Direct link to Load the data with Python") 1. Download and unzip the [full\_format\_recipes.json](https://www.kaggle.com/hugodarwood/epirecipes?select=full_format_recipes.json) file from the dataset in your current directory. 2. Install the Python dependencies: ``` pip install opensearch-py==1.0.0 ``` 3. In this step you will create the script that reads the data file you downloaded and puts the records into the OpenSearch service. Create a file named `epicurious_recipes_import.py`, and add the following code; you will need to edit it to add the connection details for your OpenSearch service. Find the `SERVICE_URI` on Aiven's dashboard. ``` import json from opensearchpy import helpers, OpenSearch SERVICE_URI = 'YOUR_SERVICE_URI_HERE' INDEX_NAME = 'epicurious-recipes' os_client = OpenSearch(hosts=SERVICE_URI, ssl_enable=True) def load_data(): with open('full_format_recipes.json', 'r') as f: data = json.load(f) for recipe in data: yield {'_index': INDEX_NAME, '_source': recipe} ``` OpenSearch Python client offers a helper called bulk() which allows us to send multiple documents in one API call. ``` helpers.bulk(os_client, load_data()) ``` 4. Run the script with the following command, and wait for it to complete: ``` python epicurious_recipes_import.py ``` ## Get data mapping with Python[​](#get-mapping-with-python "Direct link to Get data mapping with Python") When no data structure is specified, which is our case as shown on [load the data with Python](/docs/products/opensearch/howto/sample-dataset.md#load-data-with-python), OpenSearch uses dynamic mapping to automatically detect the fields. To check the mapping definition of your data, OpenSearch client provides a function called `get_mapping` as shown: ``` import pprint INDEX_NAME = 'epicurious-recipes' mapping_data = os_client.indices.get_mapping(INDEX_NAME) # Find index doc_type doc_type = list(mapping_data[INDEX_NAME]["mappings"].keys())[0] schema = mapping_data[INDEX_NAME]["mappings"][doc_type] fields = list(schema.keys()) pprint(fields) pprint(schema) ``` You should be able to see the fields' output: ``` ['calories', 'categories', 'date', 'desc', 'directions', 'fat', 'ingredients', 'protein', 'rating', 'sodium', 'title'] ``` And the mapping with the fields and their respective types. ``` {'calories': {'type': 'float'}, 'categories': {'fields': {'keyword': {'ignore_above': 256, 'type': 'keyword'}}, 'type': 'text'}, 'date': {'type': 'date'}, 'desc': {'fields': {'keyword': {'ignore_above': 256, 'type': 'keyword'}}, 'type': 'text'}, 'directions': {'fields': {'keyword': {'ignore_above': 256, 'type': 'keyword'}}, 'type': 'text'}, 'fat': {'type': 'float'}, 'ingredients': {'fields': {'keyword': {'ignore_above': 256, 'type': 'keyword'}}, 'type': 'text'}, 'protein': {'type': 'float'}, 'rating': {'type': 'float'}, 'sodium': {'type': 'float'}, 'title': {'fields': {'keyword': {'ignore_above': 256, 'type': 'keyword'}}, 'type': 'text'}} ``` Read more about OpenSearch mapping in the [official OpenSearch documentation](https://opensearch.org/docs/latest/opensearch/rest-api/index-apis/put-mapping/). ## Load the data with NodeJS[​](#load-data-with-nodejs "Direct link to Load the data with NodeJS") To load data with NodeJS we'll use [OpenSearch JavaScript client](https://github.com/opensearch-project/opensearch-js) Download [full\_format\_recipes.json](https://www.kaggle.com/hugodarwood/epirecipes?select=full_format_recipes.json), unzip and put it into the project folder. It is possible to index values either one by one or by using a bulk operation. Because we have a file containing a long list of recipes we'll use a bulk operation. A bulk endpoint expects a request in a format of a list where an action and an optional document are followed one after another: * Action and metadata * Optional document * Action and metadata * Optional document * and so on To achieve this expected format, use a flat map to create a flat list of such pairs instructing OpenSearch to index the documents. ``` module.exports.recipes = require("./full_format_recipes.json"); /** * Indexing data from json file with recipes. */ module.exports.indexData = () => { console.log(`Ingesting data: ${recipes.length} recipes`); const body = recipes.flatMap((doc) => [ { index: { _index: indexName } }, doc, ]); client.bulk({ refresh: true, body }, console.log(result.body)); }; ``` Run this method to load the data and wait till it's done. We're injecting over 20k recipes, so it can take 10-15 seconds. ## Get data mapping with NodeJS[​](#get-mapping-with-nodejs "Direct link to Get data mapping with NodeJS") We didn't specify any particular structure for the recipes data when we uploaded it. Even though we could have set explicit mapping beforehand, we opted to rely on OpenSearch to derive the structure from the data and use dynamic mapping. To see the mapping definitions use the `getMapping` method and provide the index name as a parameter. ``` /** * Retrieving mapping for the index. */ module.exports.getMapping = () => { console.log(`Retrieving mapping for the index with name ${indexName}`); client.indices.getMapping({ index: indexName }, (error, result) => { if (error) { console.error(error); } else { console.log(result.body.recipes.mappings.properties); } }); }; ``` You should be able to see the following structure: ``` { calories: { type: 'long' }, categories: { type: 'text', fields: { keyword: [Object] } }, date: { type: 'date' }, desc: { type: 'text', fields: { keyword: [Object] } }, directions: { type: 'text', fields: { keyword: [Object] } }, fat: { type: 'long' }, ingredients: { type: 'text', fields: { keyword: [Object] } }, protein: { type: 'long' }, rating: { type: 'float' }, sodium: { type: 'long' }, title: { type: 'text', fields: { keyword: [Object] } } } ``` These are the fields you can play with. You can find information on dynamic mapping types [in the documentation](https://opensearch.org/docs/latest/field-types/#dynamic-mapping). ## Sample queries with HTTP client[​](#sample-queries-with-http-client "Direct link to Sample queries with HTTP client") With the data in place, we can start trying some queries against your OpenSearch service. Since it has a simple HTTP interface, you can use your favorite HTTP client. In these examples, we will use [httpie](https://github.com/httpie/httpie) because it's one of our favorites. First, export the `SERVICE_URI` variable with your OpenSearch service URI address and index name from the previous script: ``` export SERVICE_URI="YOUR_SERVICE_URI_HERE/epicurious-recipes" ``` 1. Execute a basic search for the word `vegan` across all documents and fields: ``` http "$SERVICE_URI/_search?q=vegan" ``` 2. Search for `vegan` in the `desc` or `title` fields only: ``` http POST "$SERVICE_URI/_search" <<< ' { "query": { "multi_match": { "query": "vegan", "fields": ["desc", "title"] } } } ' ``` 3. Search for recipes published only in 2013: ``` http POST "$SERVICE_URI/_search" <<< ' { "query": { "range" : { "date": { "gte": "2013-01-01", "lte": "2013-12-31" } } } } ' ``` --- # Scale disk storage for your Aiven for OpenSearch® service Scale the disk storage of your Aiven for OpenSearch® service up or down without disrupting the running service. /eol-for-major-versions#aiven-for-flinkAdding or removing disk storage does not disrupt the running service. You pay only for extra storage instead of upgrading compute resources. You can add extra storage when you create a service or after it is running. When you add storage to a running service, the Aiven Platform provisions the extra disk and adds it to the running instances. For a clustered service such as Aiven for Apache Kafka®, Aiven divides extra storage equally between the nodes. For a shared service, each node receives the full extra capacity. ## Limitations[​](#limitations "Direct link to Limitations") * Disk added for extra storage is slower than the original disk until the next maintenance update. The slower disk can reduce performance for I/O-intensive workloads. * Maximum storage depends on the plan, service type, and cloud provider. It can be up to five times the plan's base storage size. * Cloud providers limit how many times you can increase storage between maintenance updates. If you reach the limit, run a maintenance update to optimize performance. * You cannot add storage during a maintenance update. * Dynamic disk sizing (DDS) is not supported on custom service plans. Pricing If you add storage when you create a service, **Additional disk storage** shows an estimated monthly cost. The **Service summary** lists plan storage plus additional storage. The estimated monthly price includes the additional storage cost. If you add storage to a running service, the Aiven Console shows the cost of the additional storage and related backups. The same costs appear on your invoices. ## Add or remove storage[​](#add-or-remove-storage "Direct link to Add or remove storage") ### Add storage when you create a service[​](#add-storage-when-you-create-a-service "Direct link to Add storage when you create a service") To add storage while you create a service: 1. In **Additional disk storage**, set the size with the slider or enter a value in GB. 2. Review the estimated monthly cost. 3. In the **Service summary**, click **Create service**. Change additional storage later on the running service, or enable automatic disk scaling with Aiven Autoscaler. ### Change storage on a running service[​](#change-storage-on-a-running-service "Direct link to Change storage on a running service") You cannot add or remove storage when service nodes are in the rebuilding state, for example during a maintenance update or a service upgrade. If you are removing disk storage: * Make sure the data in your service does not exceed the allocated storage. If it does, you cannot remove the additional storage. * Plan for the time it takes to rebuild the service. The time depends on the service. - Console - CLI - Terraform 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Actions** > **Manage additional storage**. 3. Change the disk storage. note * The price shown for the additional storage includes backup costs. * You can only remove storage that you previously added using this feature. To downgrade further, you can change your service plan. 4. Click **Save Changes**. Use [Aiven CLI](/docs/tools/cli.md) to add or remove additional storage using [`avn service update`](/docs/tools/cli/service-cli.md#avn-cli-service-update) with the `--disk-space-gib` flag to specify the total disk space to provide to your service. For example, if your service has a 80-GiB disk and you would like to add an extra 10-GiB disk, use: ``` avn service update --disk-space-gib 90 --project PROJECT_NAME SERVICE_NAME ``` note * When you perform a horizontal service upgrade or downgrade, remember to include all additional disks the service uses. For example, when switching from `Startup-4` to `Business-4` or from `Business-4` to `Startup-4`, include all the additional disks available for this service. * When you fork an existing service, include all additional disks the service uses. Use the `additional_disk_space` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). If you added storage, the additional storage is available immediately. If you removed additional storage, the service nodes go through a rolling restart. Depending on the service type and configuration, there might be a short downtime for services with no HA capabilities. note Storage optimization is performed at the next maintenance update after a change to the storage size. Due to cloud provider limitations, there is a limit on how many times storage can be increased between two maintenance updates. When this limit is reached, perform a maintenance update for performance optimization. Plan increases to avoid reaching this limit. Related pages * [Disk autoscaler](/docs/products/opensearch/howto/disk-autoscaler.md) * [Change the service plan](/docs/products/opensearch/howto/change-service-plan.md) * [Memory and out-of-memory conditions](/docs/products/opensearch/concepts/service-memory.md) --- # Index retention patterns Learn to set index retention patterns and manage maximum indices in your Aiven for OpenSearch® instance. Index retention patterns are a legacy approach for capping the number of indices that match a name pattern, from before Index State Management (ISM) was available on Aiven for OpenSearch. Use ISM instead for any index you're setting up now. For more information, see [Index lifecycle management](/docs/products/opensearch/concepts/indices.md#index-lifecycle-management). ## Set index retention patterns[​](#set-index-retention-patterns "Direct link to Set index retention patterns") To define index retention policies for your OpenSearch indices: * Console * API 1. Log in to the [Aiven Console](https://console.aiven.io), select your project, and select your Aiven for OpenSearch service. 2. In the **Data** section, click **Indexes**. The **Indexes** section lists the patterns that are currently in use. 3. Click **Add pattern**. 4. Enter the pattern to use and the maximum index count for the pattern. 5. Click **Create**. Alternatively, you can use the [API](https://api.aiven.io/doc/) with a request similar to the following: ``` curl -X PUT --data '{ "user_config": { "index_patterns": [ {"pattern": "logs*", "max_index_count": 2}, {"pattern": "test.?", "max_index_count": 3} ] } }' \ --header "content-type: application/json" \ --header "authorization: aivenv1 " \ https://api.aiven.io/v1beta/project//service/ ``` Parameters: * ``: Your Aiven project name. * ``: Name of your Aiven for OpenSearch service. Related pages * [Indices in Aiven for OpenSearch](/docs/products/opensearch/concepts/indices.md) --- # Set up cross-cluster replication for Aiven for OpenSearch® [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Set up cross-cluster replication (CCR) for your Aiven for OpenSearch service to synchronize data across regions and cloud providers efficiently. note Cross cluster replication is not available for the Free and Startup plans. ## Steps to set up CCR[​](#steps-to-set-up-ccr "Direct link to Steps to set up CCR") * Aiven Console * Aiven API * Terraform 1. Log in to the [Aiven Console](https://console.aiven.io/), and select the Aiven for OpenSearch service. 2. On the service's **Overview** page, scroll to the **Cross cluster replication** section. 3. Click **+ Create follower**. 4. In the **Create follower** dialog, select how to set up the follower: **To use an existing service:** 1. Select **Existing service**. 2. Select the **Project name** and the **Service name**. note Only services running the same OpenSearch version as the leader can be selected as a follower. For example, a 2.19 leader requires a 2.19 follower. 3. Click **Create follower**. **To create a new service:** 1. Select **New service**. 2. Select the **Project name**. 3. Click **Create service**. The service creation page opens. 4. On the service creation page: * Enter a name for the follower service. * Select the cloud provider, region, and service plan. * Add additional disk storage if required. note The follower service must use the same service plan as the leader service during creation to ensure sufficient memory. You can change the service plan later. 5. Click **Create**. To set up cross-cluster replication using the Aiven API, create the follower service and include the `opensearch_cross_cluster_replication` integration in the service creation request. For more information, see [Create service](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate). ``` curl -X POST https://api.aiven.io/v1/project/PROJECT_NAME/service \ -H "Authorization: Bearer API_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "service_type": "opensearch", "plan": "business-4", "cloud": "gcp-us-east1", "service_name": "follower-os-cluster", "service_integrations": [ { "integration_type": "opensearch_cross_cluster_replication", "source_service": "LEADER_SERVICE_NAME", "user_config": {} } ] }' ``` Parameters: * `PROJECT_NAME`: Aiven project name * `API_TOKEN`: API authentication token * `LEADER_SERVICE_NAME`: Leader service name Use the [`aiven_opensearch`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/opensearch) and [`aiven_service_integration`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration) resources to create the follower service and configure replication. 1. Add the following to your Terraform configuration: ``` # Leader service resource "aiven_opensearch" "leader" { project = var.project_name cloud_name = "google-us-east1" plan = "business-4" service_name = "leader-os-cluster" } # Follower service resource "aiven_opensearch" "follower" { project = var.project_name cloud_name = "aws-us-east-1" plan = "business-4" service_name = "follower-os-cluster" } # Cross-cluster replication integration resource "aiven_service_integration" "ccr" { project = var.project_name integration_type = "opensearch_cross_cluster_replication" source_service_name = aiven_opensearch.leader.service_name destination_service_name = aiven_opensearch.follower.service_name depends_on = [ aiven_opensearch.leader, aiven_opensearch.follower, ] } ``` note The follower service must use the same service plan as the leader during creation to ensure sufficient memory. You can change the plan later. 2. Apply the configuration: ``` terraform apply ``` note To learn about the limitations with cross-cluster replications for Aiven for OpenSearch, see the [Limitations](/docs/products/opensearch/concepts/cross-cluster-replication-opensearch.md#ccr-limitatons) section. ## Start replication[​](#start-replication "Direct link to Start replication") The `opensearch_cross_cluster_replication` integration sets up the remote-cluster connection between the follower and the leader and maps the replication security roles. It does not start replicating any data. After the integration is created, start replication from the follower using the [OpenSearch replication API](https://docs.opensearch.org/latest/tuning-your-cluster/replication-plugin/api/), either for a single index or for all indices that match a pattern. The integration exposes the leader as a remote-cluster alias named `LEADER_PROJECT_LEADER_SERVICE`, formed from the leader's project name and service name joined by an underscore. Use this value as `leader_alias` in the requests that follow. note Aiven maps the replication roles as `ccr_leader_full_access` and `ccr_follower_full_access`. The OpenSearch documentation uses the built-in role names `cross_cluster_replication_leader_full_access` and `cross_cluster_replication_follower_full_access`. Use the Aiven role names on Aiven services. note The replication API requests that send a JSON body require the `Content-Type: application/json` header. For example, with `curl`, add `-H 'Content-Type: application/json'` to the request. ### Replicate a single index[​](#replicate-a-single-index "Direct link to Replicate a single index") Run the following request against the follower to start replicating one index: ``` PUT https://FOLLOWER_HOST/_plugins/_replication/FOLLOWER_INDEX/_start { "leader_alias": "LEADER_PROJECT_LEADER_SERVICE", "leader_index": "LEADER_INDEX", "use_roles": { "leader_cluster_role": "ccr_leader_full_access", "follower_cluster_role": "ccr_follower_full_access" } } ``` Replace the following: * `FOLLOWER_HOST`: connection URI of the follower service. * `FOLLOWER_INDEX`: name of the index to create on the follower. * `LEADER_PROJECT_LEADER_SERVICE`: remote-cluster alias of the leader service. * `LEADER_INDEX`: name of the index to replicate from the leader. For more information, see [Start replication](https://docs.opensearch.org/latest/tuning-your-cluster/replication-plugin/getting-started/) in the OpenSearch documentation. ### Replicate indices by pattern[​](#replicate-indices-by-pattern "Direct link to Replicate indices by pattern") An auto-follow rule replicates every existing and future index on the leader that matches a wildcard pattern, so you do not start replication for each index separately. Run the following request against the follower: ``` POST https://FOLLOWER_HOST/_plugins/_replication/_autofollow { "leader_alias": "LEADER_PROJECT_LEADER_SERVICE", "name": "REPLICATION_RULE_NAME", "pattern": "movies*", "use_roles": { "leader_cluster_role": "ccr_leader_full_access", "follower_cluster_role": "ccr_follower_full_access" } } ``` The `pattern` field is the index mask and supports wildcards. The following table shows example masks and the indices they match: | Mask | Matches | | ----------- | ---------------------------------------- | | `movies*` | `movies`, `movies-0001`, `movies-2024` | | `logs-*` | `logs-2024.06.18`, `logs-app` | | `index-01*` | `index-01`, `index-012`, `index-01-prod` | | `*` | every user index on the leader | note A `*` pattern matches only user indices. Auto-follow skips system and hidden dot-prefixed indices such as `.kibana_1`, `.opendistro_security`, and `.tasks`. Dot-prefixed indices that match the pattern are listed under `failed_indices` in `autofollow_stats` instead of being replicated. To verify that replication is running, query the auto-follow statistics and the status of a replicated index: ``` GET https://FOLLOWER_HOST/_plugins/_replication/autofollow_stats GET https://FOLLOWER_HOST/_plugins/_replication/INDEX_NAME/_status ``` The `_status` request returns `SYNCING` when replication is active. To stop auto-following new indices, delete the rule: ``` DELETE https://FOLLOWER_HOST/_plugins/_replication/_autofollow { "leader_alias": "LEADER_PROJECT_LEADER_SERVICE", "name": "REPLICATION_RULE_NAME" } ``` Deleting the rule stops auto-following new indices. Indices that are already replicating continue until you stop each one. For more information, see [Auto-follow](https://docs.opensearch.org/latest/tuning-your-cluster/replication-plugin/auto-follow/) and [Replication permissions](https://docs.opensearch.org/latest/tuning-your-cluster/replication-plugin/permissions/) in the OpenSearch documentation. ## View follower services[​](#view-follower-services "Direct link to View follower services") To view the follower services configured for your Aiven for OpenSearch service: 1. Go to the **Overview** page of your service. 2. Scroll to the **Cross cluster replication** section. The section lists all follower services with their version, region, project, and replication status. The **Follower** badge identifies each follower service. Click a follower service name to open it. ## Promote a follower service to a standalone service[​](#promote-a-follower-service-to-a-standalone-service "Direct link to Promote a follower service to a standalone service") You can promote a follower service to standalone status to make it work independently, without replicating data from a leader service. This is helpful in disaster recovery situations where replication needs to stop, and the service must function on its own. note Promoting a follower service to standalone stops replication and deletes the replication integration. * Aiven Console * Aiven API * Terraform 1. Log in to the [Aiven Console](https://console.aiven.io/), and select the Aiven for OpenSearch service. 2. On the service's **Overview** page, scroll to the **Cross-cluster replica status** section. 3. Click the follower Aiven for OpenSearch service to promote. 4. On the follower service's **Overview** page, click **Promote to standalone** in the **Cross-cluster replica status**. 5. Click **Confirm** to complete the promotion. The follower service is now a standalone service and can accept writes. You can set up replication again if needed. To promote a follower service to standalone using the Aiven API, delete the `opensearch_cross_cluster_replication` integration from the service. ``` curl -X DELETE https://api.aiven.io/v1/project//integration/ \ -H "Authorization: Bearer " ``` Parameters: * ``: Aiven project name. * ``: ID of the `opensearch_cross_cluster_replication` integration. * ``: API authentication token. Removing the integration transitions the follower service to a standalone service. To promote a follower service to standalone, remove the `aiven_service_integration` resource for the `opensearch_cross_cluster_replication` integration from your Terraform configuration and apply the change. 1. Delete the `aiven_service_integration` block for `opensearch_cross_cluster_replication` from your configuration. 2. Apply the change: ``` terraform apply ``` Removing the integration transitions the follower service to a standalone service. Related pages * [Cross-cluster replication for Aiven for OpenSearch®](/docs/products/opensearch/concepts/cross-cluster-replication-opensearch.md) * [OpenSearch® cross-cluster replication via the OpenSearch API](https://opensearch.org/docs/latest/replication-plugin/get-started/) --- # Store and manage snapshot repository credentials in Aiven for OpenSearch® Use `custom_keystores` in Aiven for OpenSearch® to store object storage credentials in Amazon S3, Google Cloud Storage, or Azure. You can then use these credentials when registering snapshot repositories through the native OpenSearch API. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A running Aiven for OpenSearch service * [Aiven API](/docs/tools/api.md) access with a valid [authentication token](/docs/platform/howto/create_authentication_token.md) * Credentials for one of the following supported storage services: * Amazon S3 * Google Cloud Storage (GCS) * Microsoft Azure Blob Storage * [Security management enabled](/docs/products/opensearch/howto/enable-opensearch-security.md) for your OpenSearch service, with the `os-sec-admin` password set * OpenSearch snapshot API [enabled](#enable-the-snapshot-api) * OpenSearch user with [permissions](https://docs.opensearch.org/docs/2.19/security/access-control/permissions/#snapshot-repository-permissions) to modify snapshot repositories note To create, restore, or delete snapshots using the registered repositories, assign additional [snapshot permissions](https://docs.opensearch.org/docs/2.19/security/access-control/permissions/#snapshot-permissions) to the OpenSearch user. ## Limitations[​](#limitations "Direct link to Limitations") * Up to 10 custom keystores are allowed per service. * Each keystore name must be unique and must not conflict with names reserved by Aiven. * Credentials are not validated when added. If the credentials are invalid, OpenSearch returns an error when you use them, such as when registering a snapshot repository. * Only supported via the native OpenSearch API. * Not available in the Aiven Console. ## How it works[​](#how-it-works "Direct link to How it works") When you add a `custom_keystores` entry to your service's `user_config`, Aiven stores the credentials on each Aiven for OpenSearch node using the [OpenSearch keystore mechanism](https://docs.opensearch.org/docs/latest/security/configuration/opensearch-keystore/). Each credential is stored using the following format: ``` PROVIDER.client.KEYSTORE_NAME.KEY = VALUE ``` For example: ``` s3.client.MY_S3_KEYS.access_key = AKIA... s3.client.MY_S3_KEYS.secret_key = d6pD... ``` To use these credentials, specify the keystore name in the `client` field when registering a snapshot repository using the OpenSearch API. caution Sensitive values (such as `secret_key`, `sas_token`, and `credentials`) are excluded from API responses for security reasons. ## Enable the snapshot API[​](#enable-the-snapshot-api "Direct link to Enable the snapshot API") By default, the snapshot API is disabled. You can enable it using the Aiven Console or Aiven CLI. * Console * CLI 1. Access your Aiven for OpenSearch service in the [Aiven Console](https://console.aiven.io/). 2. Click **Service settings** in the sidebar. 3. Scroll to the **Advanced configuration** section and click **Configure**. 4. In the **Advanced configuration** window, click **Add configuration option**. 5. Use the search bar to find `opensearch.enable_snapshot_api`, and set it to **Enable**. 6. Click **Save configuration**. ``` avn service update SERVICE_NAME --project PROJECT_NAME -c opensearch.enable_snapshot_api=true ``` Replace `SERVICE_NAME` and `PROJECT_NAME` with your own values. ## Configure keystores[​](#configure-keystores "Direct link to Configure keystores") Custom keystores are configured in the `user_config` of your Aiven for OpenSearch service. Use the following API request to store credentials: ``` curl -s --url "https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME" \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ -X PUT -d '{ "user_config": { "custom_keystores": [ { "name": "MY_S3_KEYS", "type": "s3", "settings": { "access_key": "AWS_ACCESS_KEY", "secret_key": "AWS_SECRET_KEY" } } ] } }' ``` Replace each placeholder with the appropriate value for your environment. ## Register a repository using the OpenSearch API[​](#register-a-repository-using-the-opensearch-api "Direct link to Register a repository using the OpenSearch API") After storing your credentials, use the [OpenSearch snapshot API](https://opensearch.org/docs/latest/api-reference/snapshots/create-repository/) to register a repository. Set the `client` field to the name of the corresponding keystore. ``` curl -X PUT "https://SERVICE_URI/_snapshot/MY_S3_REPO" \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "type": "s3", "settings": { "bucket": "SNAPSHOT_BUCKET", "region": "AWS_REGION", "base_path": "backups/opensearch", "client": "MY_S3_KEYS" } }' ``` ## Update or delete credentials[​](#update-or-delete-credentials "Direct link to Update or delete credentials") To update or remove stored credentials, modify the `custom_keystores` field using the same API endpoint: ``` PUT https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/update ``` Add the updated `custom_keystores` list to the request body and send the API request. The changes are applied automatically to all Aiven for OpenSearch nodes. ## Example keystore configurations[​](#example-keystore-configurations "Direct link to Example keystore configurations") Use the following examples to configure credentials for each supported storage provider. ### Azure with SAS token[​](#azure-with-sas-token "Direct link to Azure with SAS token") ``` { "name": "MY_AZURE_KEYS", "type": "azure", "settings": { "account": "AZURE_ACCOUNT", "sas_token": "AZURE_SAS_TOKEN" } } } ``` ### Google Cloud Storage with service account credentials[​](#google-cloud-storage-with-service-account-credentials "Direct link to Google Cloud Storage with service account credentials") ``` { "name": "MY_GCS_KEYS", "type": "gcs", "settings": { "credentials": { "type": "service_account", "project_id": "PROJECT_ID", "private_key_id": "KEY_ID", "private_key": "PRIVATE_KEY", "client_email": "SERVICE_ACCOUNT_EMAIL", "client_id": "CLIENT_ID", "auth_uri": "https://accounts.google.com/o/oauth2/auth", "token_uri": "https://oauth2.googleapis.com/token", "auth_provider_x509_cert_url": "https://www.googleapis.com/oauth2/v1/certs", "client_x509_cert_url": "CERT_URL" } } } ``` ### AWS S3[​](#aws-s3 "Direct link to AWS S3") ``` { "name": "MY_S3_KEYS", "type": "s3", "settings": { "access_key": "AWS_ACCESS_KEY", "secret_key": "AWS_SECRET_KEY" } } ``` Related pages * [Create and manage custom repositories](/docs/products/opensearch/howto/custom-repositories.md) * [OpenSearch snapshot API](https://opensearch.org/docs/latest/api-reference/snapshots/index/) * [OpenSearch keystore mechanism](https://docs.opensearch.org/docs/latest/security/configuration/opensearch-keystore/) --- # Tag your Aiven for OpenSearch® service Add key-value tags to your Aiven for OpenSearch® service to organize services and track ownership, cost allocation, and governance. Use tags to add metadata to Aiven services to categorize them or run custom logic on them. Typical uses include: * Tagging for governance to deploy services with specific tags only. * Tagging for internal cost reporting, ownership, allocation, and accountability. A tag is a key/value pair: * **Key**: A case-sensitive string that starts with a letter and consists of letters, numbers, dashes, and underscores. The maximum length for a key is 64 characters. * **Value**: A string value limited to 64 UTF-8 characters. Within a service, the tag keys must be unique. * Console * Terraform 1. In the service, click **Service settings**. 2. In the **Service status** section, click **Actions** > **Add service tags**. 3. Enter a key and value for each tag. 4. Click **Save changes**. Use the `tag` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Fork Aiven for OpenSearch®](/docs/products/opensearch/howto/fork-service.md) --- # Track restore progress for your Aiven for OpenSearch® service Track the restore progress of individual nodes in your Aiven for OpenSearch® service during node replacement, forking, or maintenance, using the Aiven API. You can track restore progress for individual nodes during service node replacement by using the Aiven API. For example, use this endpoint to monitor the restore progress of a forked service or when applying maintenance. The service object exposes restore progress under `node_states[].progress_updates`: * `service.node_states[]` contains per-node state entries. * When a node is restoring or catching up, its `state` is typically `syncing_data`. * When the state is `syncing_data`, the node may include `progress_updates` with one or more phase objects. * Other node states don't include restore progress data. note `progress_updates` may be missing or empty even when a node is in `syncing_data`. This can occur when a restore completes before detailed progress is reported or when the service does not emit detailed progress counters. ## API endpoints[​](#api-endpoints "Direct link to API endpoints") Restore progress fields are part of the standard service response payload. * Get a single service (recommended for polling): `GET /project/{project}/service/{service_name}` * List services in a project: `GET /project/{project}/service` - Request - Response ``` curl -H "Authorization: aivenv1 API_TOKEN" https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME ``` Replace the following placeholders: * `API_TOKEN`: Your Aiven API token. * `PROJECT_NAME`: Your Aiven project name. * `SERVICE_NAME`: The name of your service. ``` { "service": { ... "node_states": [ { "node_name": "...", "state": "syncing_data", "progress_updates": [ { "phase": "basebackup", "completed": false, "current": 3410567, "min": 0, "max": 7569280, "unit": "bytes_uncompressed" } ] } ], ... } } ``` ## Node states[​](#node-states "Direct link to Node states") Common values for `node_states[].state` include: * `setting_up_vm`: The virtual machine is being created or initialized. * `syncing_data`: The node is restoring data or catching up. * `running`: The node is operating normally. * `leaving`: The node is leaving the cluster. * `unknown`: A transient or error state. ## `progress_updates` data model[​](#progress_updates-data-model "Direct link to progress_updates-data-model") `progress_updates` is a list of phase objects. When present, phases appear in the following order: 1. `prepare` 2. `basebackup` 3. `stream` 4. `finalize` Each phase object includes the following fields: ``` { "completed": false, "current": 3410567, "max": 7569280, "min": 0, "phase": "basebackup", "unit": "bytes_uncompressed" } ``` ### Field semantics[​](#field-semantics "Direct link to Field semantics") * `phase`: String, required. The restore phase. Possible values: `prepare`, `basebackup`, `stream`, and `finalize`. * `completed`: Boolean, required. Whether the phase is complete. * `current`: Number or null, optional. The current progress value. This field can be missing or null. * `min`: Number or null, optional. The starting value for the phase. This field can be missing or null. * `max`: Number or null, optional. The expected total value for the phase. This value can be missing, null, or change while the restore is in progress. * `unit`: String or null, optional. The unit for `current`, `min`, and `max`. New unit values can appear over time. Important considerations * Treat `unit` as an opaque identifier. Unknown values can appear. * `max` may change while a restore is in progress. * Not all phases report numeric counters. Some services only indicate phase completion. ## Why `max` values can change[​](#why-max-values-can-change "Direct link to why-max-values-can-change") The `current`, `min`, and `max` values are best-effort progress indicators. They can be based on estimates or on system state that changes over time. Treat `max` as the latest known expected total, not as a fixed guarantee. Common reasons `max` can change include: * The restore process discovers additional work after it starts, such as files, segments, or objects that become visible only after metadata is read. * New data is added on the backend while the node is catching up, which moves the completion point forward. This is common during incremental catch-up phases. * Progress is calculated from system state, such as replication lag, rather than from a fixed work queue. As the system state changes, the value is recalculated. * The service switches restore strategies during the operation, for example from snapshot restore to replication catch-up, which changes what the counters represent. As a result: * Phase percentage can decrease even when the restore operates normally. * Remaining-time estimates based on `max` are unreliable. * Sudden changes in `max` are expected unless the node remains in `syncing_data` longer than expected. ## Restore phase meanings[​](#restore-phase-meanings "Direct link to Restore phase meanings") Phase names are standardized, but the underlying work and the meaning of the counters are service-specific. * `prepare`: Prepares the node for restore. * `basebackup`: Restores the full backup. * `stream`: Applies incremental changes, such as replication or log replay. * `finalize`: Completes final steps before serving traffic. Not all restores include every phase. ## Compute phase progress percentages[​](#compute-phase-progress-percentages "Direct link to Compute phase progress percentages") You cannot reliably compute overall restore progress. You can compute a phase-specific progress percentage when `min`, `max`, and `current` are present and `max != min`. ``` pct = round(((current - min) / (max - min)) * 100, 1) ``` When handling progress values: * If any of `min`, `max`, or `current` is null or missing, display `n/a`. * If `max == min`, treat the percentage as undefined. * Expect the percentage to decrease when `max` changes. * Clamp displayed values to the range `[0, 100]`. ## Polling guidance[​](#polling-guidance "Direct link to Polling guidance") Progress updates are best-effort and refresh every 10 seconds while a node is in `syncing_data`. Poll the service state every 10 to 30 seconds. More frequent polling does not provide additional detail. For each `node_states[]` entry: * If `state` is not `syncing_data`, no restore progress is available. * If `state` is `syncing_data`: * If `progress_updates` is missing or empty, the node is restoring without detailed progress data. * Otherwise, the current phase is the last phase where `completed` is `false`. Stop polling when all nodes reach the `running` state or when a stall is detected. ### Stall detection[​](#stall-detection "Direct link to Stall detection") The API does not provide per-phase timestamps. To detect stalls, use a time-based threshold, such as a node remaining in `syncing_data` longer than expected. Do not rely on counters or `max` values to estimate remaining time. Related pages * [Backups](/docs/products/opensearch/howto/restore_opensearch_backup.md) * [Fork your service](/docs/products/opensearch/howto/fork-service.md) --- # Upgrade Elasticsearch clients to OpenSearch® Elasticsearch has introduced breaking changes into their client libraries as early as **7.13.\***, meaning newer Elasticsearch clients won't work with OpenSearch®. ## Migration steps[​](#migration-steps "Direct link to Migration steps") To upgrade the Elasticsearch clients to OpenSearch: 1. Pin your Elasticsearch libraries to version **7.10.2** (latest version under the open-source license). 2. Switch from **Elasticsearch 7.10.2** to the fully compatible **OpenSearch 1.0.0**. 3. Update OpenSearch libraries till their latest version. note You can migrate your cluster from Elasticsearch to OpenSearch either before or after switching the clients. ## Postpone the upgrade[​](#postpone-the-upgrade "Direct link to Postpone the upgrade") If you cannot upgrade immediately, we recommend locking the client version of Elasticsearch to **7.10.2** and use this version till you can proceed with the migration steps outlined above. This is true for application libraries as well as for the supporting ecosystem of tooling like File Beats and Logstash. note See [OpenSearch compatibility documentation](https://opensearch.org/docs/latest/clients/index/) as the source of truth. ## Client migration examples[​](#client-migration-examples "Direct link to Client migration examples") To help you with the migration, see some [code migration examples](https://github.com/aiven/opensearch-migration-examples). ### Java and Spring Boot[​](#java-and-spring-boot "Direct link to Java and Spring Boot") [Java client](https://opensearch.org/docs/latest/clients/java-rest-high-level/) is a fork of the Elasticsearch library. To migrate change the Maven or Gradle dependencies, and update the import statements for the DOA layer. See an [example code change](https://github.com/aiven/opensearch-migration-examples/commit/7453d659c06b234ae7f28f801a074e459c2f31c8) in our repository. ### NodeJS[​](#nodejs "Direct link to NodeJS") [Node library](https://opensearch.org/docs/latest/clients/javascript/) is a fork of the Elasticsearch library. The only required change should be the dependency in `package.json` and the `require` or `import` statements. See an [example migration code](https://github.com/aiven/opensearch-migration-examples/tree/main/node-client-migration) in our repository. ### Python[​](#python "Direct link to Python") [Python client library](https://opensearch.org/docs/latest/clients/python) is a fork of the Elasticsearch libraries. The only required change should be the dependencies and the `require` or `import` statements. See an [example migration code](https://github.com/aiven/opensearch-migration-examples/tree/main/python-client-migration) in our repository. *Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries.* --- # Maintenance and lifecycle in Aiven for OpenSearch® Keep your Aiven for OpenSearch® service current by upgrading versions and applying maintenance updates. Related pages * [Upgrade the OpenSearch version](/docs/products/opensearch/howto/os-version-upgrade.md) --- # Advanced parameters for Aiven for OpenSearch® See the configuration options available for Aiven for OpenSearch®: | Parameter | | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | []()[**additional\_backup\_regions**](#additional_backup_regions)`array`Additional Cloud Regions for Backup Replication | | []()[**opensearch\_version**](#opensearch_version)`string,null`OpenSearch version | | []()[**elasticsearch\_version**](#elasticsearch_version)`string,null`OpenSearch version | | []()[**disable\_replication\_factor\_adjustment**](#disable_replication_factor_adjustment)`boolean,null`Disable automatic replication factor adjustment for multi-node services. By default, Aiven ensures all indexes are replicated at least to two nodes. Note: Due to potential data loss in case of losing a service node, this setting can not be activated unless specifically allowed for the project. | | []()[**custom\_domain**](#custom_domain)`string,null`Serve the web frontend using a custom CNAME pointing to the Aiven DNS name. When you set a custom domain for a service deployed in a VPC, the service certificate is only created for the public-\* hostname and the custom domain. | | []()[**custom\_repos**](#custom_repos)`array`Allow to register object storage repositories in OpenSearch | | []()[**custom\_keystores**](#custom_keystores)`array`Allow to register custom keystores in OpenSearch | | []()[**ip\_filter**](#ip_filter)`array`- default: `0.0.0.0/0,::/0`IP filterAllow incoming connections from CIDR address block, e.g. '10.20.0.0/16' | | []()[**service\_log**](#service_log)`boolean,null`Service loggingStore logs for the service so that they are available in the HTTP API and console. | | []()[**static\_ips**](#static_ips)`boolean`Use static public IP addresses | | []()[**saml**](#saml)`object`OpenSearch SAML configurationsaml.enabled boolean - default: true Enable or disable OpenSearch SAML authentication Enables or disables SAML-based authentication for OpenSearch. When enabled, users can authenticate using SAML with an Identity Provider. saml.idp\_metadata\_url string The URL of the SAML metadata for the Identity Provider (IdP). This is used to configure SAML-based authentication with the IdP. saml.idp\_entity\_id string The unique identifier for the Identity Provider (IdP) entity that is used for SAML authentication. This value is typically provided by the IdP. saml.sp\_entity\_id string The unique identifier for the Service Provider (SP) entity that is used for SAML authentication. This value is typically provided by the SP. saml.subject\_key string,null Optional. Specifies the attribute in the SAML response where the subject identifier is stored. If not configured, the NameID attribute is used by default. saml.roles\_key string,null Optional. Specifies the attribute in the SAML response where role information is stored, if available. Role attributes are not required for SAML authentication, but can be included in SAML assertions by most Identity Providers (IdPs) to determine user access levels or permissions. saml.idp\_pemtrustedcas\_content string,null This parameter specifies the PEM-encoded root certificate authority (CA) content for the SAML identity provider (IdP) server verification. The root CA content is used to verify the SSL/TLS certificate presented by the server. | | []()[**openid**](#openid)`object`OpenSearch OpenID Connect Configurationopenid.enabled boolean - default: true Enable or disable OpenSearch OpenID Connect authentication Enables or disables OpenID Connect authentication for OpenSearch. When enabled, users can authenticate using OpenID Connect with an Identity Provider. openid.connect\_url string The URL of your IdP where the Security plugin can find the OpenID Connect metadata/configuration settings. openid.roles\_key string,null The key in the JSON payload that stores the user’s roles. The value of this key must be a comma-separated list of roles. Required only if you want to use roles in the JWT openid.subject\_key string,null The key in the JSON payload that stores the user’s name. If not defined, the subject registered claim is used. Most IdP providers use the preferred\_username claim. Optional. openid.jwt\_header string,null The HTTP header that stores the token. Typically the Authorization header with the Bearer schema: Authorization: Bearer \. Optional. Default is Authorization. openid.jwt\_url\_parameter string,null URL JWT token. If the token is not transmitted in the HTTP header, but as an URL parameter, define the name of the parameter here. Optional. openid.refresh\_rate\_limit\_count integer,null - min: 10 - max: 9223372036854776000 - default: 10 The maximum number of unknown key IDs in the time frame. Default is 10. Optional. openid.refresh\_rate\_limit\_time\_window\_ms integer,null - min: 10000 - max: 9223372036854776000 - default: 10000 The time frame to use when checking the maximum number of unknown key IDs, in milliseconds. Optional.Default is 10000 (10 seconds). openid.client\_id string The ID of the OpenID Connect client configured in your IdP. Required. openid.client\_secret string The client secret of the OpenID Connect client configured in your IdP. Required. openid.scope string The scope of the identity token issued by the IdP. Optional. Default is openid profile email address phone. openid.header string - default: Authorization HTTP header name of the JWT token. Optional. Default is Authorization. | | []()[**jwt**](#jwt)`object`OpenSearch JWT Configurationjwt.enabled boolean Enable or disable OpenSearch JWT authentication Enables or disables JWT-based authentication for OpenSearch. When enabled, users can authenticate using JWT tokens. jwt.signing\_key string JWT signing key The secret key used to sign and verify JWT tokens. This should be a secure, randomly generated key HMAC key or public RSA/ECDSA key. jwt.jwt\_header string,null - default: Authorization The HTTP header name where the JWT token is transmitted. Typically 'Authorization' for Bearer tokens. jwt.jwt\_url\_parameter string,null If the JWT token is transmitted as a URL parameter instead of an HTTP header, specify the parameter name here. jwt.subject\_key string,null The key in the JWT payload that contains the user's subject identifier. If not specified, the 'sub' claim is used by default. jwt.roles\_key string,null JWT claim key for roles The key in the JWT payload that contains the user's roles. If specified, roles will be extracted from the JWT for authorization. jwt.required\_audience string,null Required JWT audience If specified, the JWT must contain an 'aud' claim that matches this value. This provides additional security by ensuring the JWT was issued for the expected audience. jwt.required\_issuer string,null Required JWT issuer If specified, the JWT must contain an 'iss' claim that matches this value. This provides additional security by ensuring the JWT was issued by the expected issuer. jwt.jwt\_clock\_skew\_tolerance\_seconds integer,null - max: 300 - default: 20 JWT clock skew tolerance in seconds The maximum allowed time difference in seconds between the JWT issuer's clock and the OpenSearch server's clock. This helps prevent token validation failures due to minor time synchronization issues. | | []()[**azure\_migration**](#azure_migration)`object`Azure migration settingsazure\_migration.snapshot\_name string The snapshot name to restore from azure\_migration.restore\_global\_state boolean If true, restore the cluster state. Defaults to false azure\_migration.include\_aliases boolean Include aliases Whether to restore aliases alongside their associated indexes. Default is true. azure\_migration.indices string A comma-delimited list of indices to restore from the snapshot. Multi-index syntax is supported. azure\_migration.base\_path string The path to the repository data within its container. The value of this setting should not start or end with a / azure\_migration.compress boolean when set to true metadata files are stored in compressed format azure\_migration.chunk\_size string Chunk size Big files can be broken down into chunks during snapshotting if needed. Should be the same as for the 3rd party repository azure\_migration.max\_snapshot\_bytes\_per\_sec string Snapshot bytes per second Throttles the snapshot rate per node. Defaults to 40mb. Note that if the recovery settings for managed services are set, this value is overridden by the recovery settings. Value should be a byte size with unit, e.g. 40mb, 100kb, 1gb azure\_migration.max\_restore\_bytes\_per\_sec string Restore bytes per second Throttles the restore rate per node. Defaults to unlimited. Note that if the recovery settings for managed services are set, this value is overridden by the recovery settings. Value should be a byte size with unit, e.g. 40mb, 100kb, 1gb azure\_migration.account string Account name azure\_migration.key string Azure account secret key. One of key or sas\_token should be specified azure\_migration.sas\_token string SAS token A shared access signatures (SAS) token. One of key or sas\_token should be specified azure\_migration.container string Azure container name azure\_migration.endpoint\_suffix string Endpoint suffix Defines the DNS suffix for Azure Storage endpoints. azure\_migration.readonly boolean - default: true Whether the repository is read-only. | | []()[**gcs\_migration**](#gcs_migration)`object`Google Cloud Storage migration settingsgcs\_migration.snapshot\_name string The snapshot name to restore from gcs\_migration.restore\_global\_state boolean If true, restore the cluster state. Defaults to false gcs\_migration.include\_aliases boolean Include aliases Whether to restore aliases alongside their associated indexes. Default is true. gcs\_migration.indices string A comma-delimited list of indices to restore from the snapshot. Multi-index syntax is supported. gcs\_migration.base\_path string The path to the repository data within its container. The value of this setting should not start or end with a / gcs\_migration.compress boolean when set to true metadata files are stored in compressed format gcs\_migration.chunk\_size string Chunk size Big files can be broken down into chunks during snapshotting if needed. Should be the same as for the 3rd party repository gcs\_migration.max\_snapshot\_bytes\_per\_sec string Snapshot bytes per second Throttles the snapshot rate per node. Defaults to 40mb. Note that if the recovery settings for managed services are set, this value is overridden by the recovery settings. Value should be a byte size with unit, e.g. 40mb, 100kb, 1gb gcs\_migration.max\_restore\_bytes\_per\_sec string Restore bytes per second Throttles the restore rate per node. Defaults to unlimited. Note that if the recovery settings for managed services are set, this value is overridden by the recovery settings. Value should be a byte size with unit, e.g. 40mb, 100kb, 1gb gcs\_migration.bucket string Google Cloud Storage bucket name The path to the repository data within its container gcs\_migration.credentials string Google Cloud Storage credentials file content gcs\_migration.readonly boolean - default: true Whether the repository is read-only. | | []()[**s3\_migration**](#s3_migration)`object`AWS S3 / AWS S3 compatible migration settingss3\_migration.snapshot\_name string The snapshot name to restore from s3\_migration.restore\_global\_state boolean If true, restore the cluster state. Defaults to false s3\_migration.include\_aliases boolean Include aliases Whether to restore aliases alongside their associated indexes. Default is true. s3\_migration.indices string A comma-delimited list of indices to restore from the snapshot. Multi-index syntax is supported. s3\_migration.base\_path string The path to the repository data within its container. The value of this setting should not start or end with a / s3\_migration.compress boolean when set to true metadata files are stored in compressed format s3\_migration.chunk\_size string Chunk size Big files can be broken down into chunks during snapshotting if needed. Should be the same as for the 3rd party repository s3\_migration.max\_snapshot\_bytes\_per\_sec string Snapshot bytes per second Throttles the snapshot rate per node. Defaults to 40mb. Note that if the recovery settings for managed services are set, this value is overridden by the recovery settings. Value should be a byte size with unit, e.g. 40mb, 100kb, 1gb s3\_migration.max\_restore\_bytes\_per\_sec string Restore bytes per second Throttles the restore rate per node. Defaults to unlimited. Note that if the recovery settings for managed services are set, this value is overridden by the recovery settings. Value should be a byte size with unit, e.g. 40mb, 100kb, 1gb s3\_migration.access\_key string AWS Access key s3\_migration.secret\_key string AWS secret key s3\_migration.bucket string S3 bucket name s3\_migration.region string S3 region s3\_migration.endpoint string The S3 service endpoint to connect to. If you are using an S3-compatible service then you should set this to the service’s endpoint s3\_migration.server\_side\_encryption boolean When set to true files are encrypted on server side s3\_migration.readonly boolean - default: true Whether the repository is read-only. | | []()[**index\_patterns**](#index_patterns)`array`Index patterns | | []()[**max\_index\_count**](#max_index_count)`integer`- max: `9223372036854776000`Maximum index countDEPRECATED: use index\_patterns instead | | []()[**keep\_index\_refresh\_interval**](#keep_index_refresh_interval)`boolean`Don't reset index.refresh\_interval to the default valueAiven automation resets index.refresh\_interval to default value for every index to be sure that indices are always visible to search. If it doesn't fit your case, you can disable this by setting up this flag to true. | | []()[**opensearch\_dashboards**](#opensearch_dashboards)`object`OpenSearch Dashboards settingsopensearch\_dashboards.enabled boolean - default: true Enable or disable OpenSearch Dashboards opensearch\_dashboards.max\_old\_space\_size integer - min: 64 - max: 4096 - default: 128 Limits the maximum amount of memory (in MiB) the OpenSearch Dashboards process can use. This sets the max\_old\_space\_size option of the nodejs running the OpenSearch Dashboards. Note: the memory reserved by OpenSearch Dashboards is not available for OpenSearch. opensearch\_dashboards.opensearch\_request\_timeout integer - min: 5000 - max: 120000 - default: 30000 Timeout in milliseconds for requests made by OpenSearch Dashboards towards OpenSearch opensearch\_dashboards.multiple\_data\_source\_enabled boolean - default: true Enable or disable multiple data sources in OpenSearch Dashboards opensearch\_dashboards.session\_keepalive boolean - default: true Determines whether the session TTL resets (is “kept alive”) on each user activity. Optional. Default is true. opensearch\_dashboards.session\_ttl string - default: 1h session\_ttl Defines the time-to-live (TTL) for user sessions. The value should be a time value with unit, e.g. 1m, 5s, 1h, 3d, 100ms. Default is 1 hour. | | []()[**index\_rollup**](#index_rollup)`object`Index rollup settingsindex\_rollup.rollup\_search\_backoff\_millis integer - min: 1 plugins.rollup.search.backoff\_millis The backoff time between retries for failed rollup jobs. Defaults to 1000ms. index\_rollup.rollup\_search\_backoff\_count integer - min: 1 plugins.rollup.search.backoff\_count How many retries the plugin should attempt for failed rollup jobs. Defaults to 5. index\_rollup.rollup\_search\_search\_all\_jobs boolean plugins.rollup.search.all\_jobs Whether OpenSearch should return all jobs that match all specified search terms. If disabled, OpenSearch returns just one, as opposed to all, of the jobs that matches the search terms. Defaults to false. index\_rollup.rollup\_dashboards\_enabled boolean plugins.rollup.dashboards.enabled Whether rollups are enabled in OpenSearch Dashboards. Defaults to true. index\_rollup.rollup\_enabled boolean plugins.rollup.enabled Whether the rollup plugin is enabled. Defaults to true. | | []()[**opensearch**](#opensearch)`object`OpenSearch settingsopensearch.reindex\_remote\_whitelist array,null reindex\_remote\_allowlist Whitelisted addresses for reindexing. Changing this value will cause all OpenSearch instances to restart. opensearch.http\_max\_content\_length integer - min: 1048576 - max: 2147483647 http.max\_content\_length Maximum content length for HTTP requests to the OpenSearch HTTP API, in bytes. opensearch.http\_max\_header\_size integer - min: 1024 - max: 262144 http.max\_header\_size The max size of allowed headers, in bytes opensearch.http\_max\_initial\_line\_length integer - min: 1024 - max: 65536 http.max\_initial\_line\_length The max length of an HTTP URL, in bytes opensearch.indices\_query\_bool\_max\_clause\_count integer - min: 64 - max: 4096 indices.query.bool.max\_clause\_count Maximum number of clauses Lucene BooleanQuery can have. The default value (1024) is relatively high, and increasing it may cause performance issues. Investigate other approaches first before increasing this value. opensearch.search\_max\_buckets integer,null - min: 1 - max: 1000000 search.max\_buckets Maximum number of aggregation buckets allowed in a single response. OpenSearch default value is used when this is not defined. opensearch.indices\_fielddata\_cache\_size integer,null - min: 3 - max: 100 indices.fielddata.cache.size Relative amount. Maximum amount of heap memory used for field data cache. This is an expert setting; decreasing the value too much will increase overhead of loading field data; too much memory used for field data cache will decrease amount of heap available for other operations. opensearch.indices\_memory\_index\_buffer\_size integer - min: 3 - max: 40 indices.memory.index\_buffer\_size Percentage value. Default is 10%. Total amount of heap used for indexing buffer, before writing segments to disk. This is an expert setting. Too low value will slow down indexing; too high value will increase indexing performance but causes performance issues for query performance. opensearch.indices\_memory\_min\_index\_buffer\_size integer - min: 3 - max: 2048 indices.memory.min\_index\_buffer\_size Absolute value. Default is 48mb. Doesn't work without indices.memory.index\_buffer\_size. Minimum amount of heap used for query cache, an absolute indices.memory.index\_buffer\_size minimal hard limit. opensearch.indices\_memory\_max\_index\_buffer\_size integer - min: 3 - max: 2048 indices.memory.max\_index\_buffer\_size Absolute value. Default is unbound. Doesn't work without indices.memory.index\_buffer\_size. Maximum amount of heap used for query cache, an absolute indices.memory.index\_buffer\_size maximum hard limit. opensearch.indices\_queries\_cache\_size integer - min: 3 - max: 40 indices.queries.cache.size Percentage value. Default is 10%. Maximum amount of heap used for query cache. This is an expert setting. Too low value will decrease query performance and increase performance for other operations; too high value will cause issues with other OpenSearch functionality. opensearch.indices\_recovery\_max\_bytes\_per\_sec integer - min: 40 - max: 400 indices.recovery.max\_bytes\_per\_sec Limits total inbound and outbound recovery traffic for each node. Applies to both peer recoveries as well as snapshot recoveries (i.e., restores from a snapshot). Defaults to 40mb opensearch.indices\_recovery\_max\_concurrent\_file\_chunks integer - min: 2 - max: 5 indices.recovery.max\_concurrent\_file\_chunks Number of file chunks sent in parallel for each recovery. Defaults to 2. opensearch.action\_auto\_create\_index\_enabled boolean action.auto\_create\_index Explicitly allow or block automatic creation of indices. Defaults to true opensearch.plugins\_alerting\_filter\_by\_backend\_roles boolean plugins.alerting.filter\_by\_backend\_roles Enable or disable filtering of alerting by backend roles. Requires Security plugin. Defaults to false opensearch.knn\_memory\_circuit\_breaker\_limit integer - max: 100 knn.memory.circuit\_breaker.limit Maximum amount of memory in percentage that can be used for the KNN index. Defaults to 50% of the JVM heap size. 0 is used to set it to null which can be used to invalidate caches. opensearch.knn\_memory\_circuit\_breaker\_enabled boolean - Service restart knn.memory.circuit\_breaker.enabled Enable or disable KNN memory circuit breaker. Defaults to true. opensearch.ml\_commons\_only\_run\_on\_ml\_node boolean plugins.ml\_commons.only\_run\_on\_ml\_node Enable or disable running ML Commons tasks only on ML nodes. When enabled, ML tasks will only execute on nodes designated as ML nodes. Defaults to true. opensearch.ml\_commons\_model\_access\_control\_enabled boolean plugins.ml\_commons.model\_access\_control.enabled Enable or disable model access control for ML Commons. When enabled, access to ML models is controlled by security permissions. Defaults to false. opensearch.ml\_commons\_native\_memory\_threshold integer - min: 1 - max: 100 plugins.ml\_commons.native\_memory\_threshold Native memory threshold percentage for ML Commons. Controls the maximum percentage of native memory that can be used by ML Commons operations. Defaults to 90%. opensearch.ml\_commons\_connector\_access\_control\_enabled boolean plugins.ml\_commons.connector\_access\_control\_enabled When set to true, the setting allows admins to control access and permissions to the connector API using backend\_roles. Defaults to false. opensearch.ml\_commons\_trusted\_connector\_endpoints\_regex array plugins.ml\_commons.trusted\_connector\_endpoints\_regex Adds the trusted endpoints to the cluster settings. Supports Java regex expressions. opensearch.auth\_failure\_listeners object Opensearch Security Plugin Settings opensearch.auth\_failure\_listeners.internal\_authentication\_backend\_limiting object opensearch.auth\_failure\_listeners.internal\_authentication\_backend\_limiting.type string internal\_authentication\_backend\_limiting.type The type of rate limiting opensearch.auth\_failure\_listeners.internal\_authentication\_backend\_limiting.authentication\_backend string internal\_authentication\_backend\_limiting.authentication\_backend The internal backend. Enter internal opensearch.auth\_failure\_listeners.internal\_authentication\_backend\_limiting.allowed\_tries integer - min: 1 - max: 32767 internal\_authentication\_backend\_limiting.allowed\_tries The number of login attempts allowed before login is blocked opensearch.auth\_failure\_listeners.internal\_authentication\_backend\_limiting.time\_window\_seconds integer - max: 2147483647 internal\_authentication\_backend\_limiting.time\_window\_seconds The window of time in which the value for allowed\_tries is enforced opensearch.auth\_failure\_listeners.internal\_authentication\_backend\_limiting.block\_expiry\_seconds integer - max: 2147483647 internal\_authentication\_backend\_limiting.block\_expiry\_seconds The duration of time that login remains blocked after a failed login opensearch.auth\_failure\_listeners.internal\_authentication\_backend\_limiting.max\_blocked\_clients integer - max: 2147483647 internal\_authentication\_backend\_limiting.max\_blocked\_clients The maximum number of blocked IP addresses opensearch.auth\_failure\_listeners.internal\_authentication\_backend\_limiting.max\_tracked\_clients integer - max: 2147483647 internal\_authentication\_backend\_limiting.max\_tracked\_clients The maximum number of tracked IP addresses that have failed login opensearch.enable\_security\_audit boolean Enable/Disable security audit opensearch.enable\_snapshot\_api boolean Enable/Disable snapshot API for custom repositories, this requires security management to be enabled opensearch.thread\_pool\_search\_size integer - min: 1 - max: 128 search thread pool size Size for the thread pool. See documentation for exact details. Do note this may have maximum value depending on CPU count - value is automatically lowered if set to higher than maximum value. opensearch.thread\_pool\_search\_throttled\_size integer - min: 1 - max: 128 search\_throttled thread pool size Size for the thread pool. See documentation for exact details. Do note this may have maximum value depending on CPU count - value is automatically lowered if set to higher than maximum value. opensearch.thread\_pool\_get\_size integer - min: 1 - max: 128 Size for the thread pool. See documentation for exact details. Do note this may have maximum value depending on CPU count - value is automatically lowered if set to higher than maximum value. opensearch.thread\_pool\_analyze\_size integer - min: 1 - max: 128 analyze thread pool size Size for the thread pool. See documentation for exact details. Do note this may have maximum value depending on CPU count - value is automatically lowered if set to higher than maximum value. opensearch.thread\_pool\_write\_size integer - min: 1 - max: 128 write thread pool size Size for the thread pool. See documentation for exact details. Do note this may have maximum value depending on CPU count - value is automatically lowered if set to higher than maximum value. opensearch.thread\_pool\_force\_merge\_size integer - min: 1 - max: 128 force\_merge thread pool size Size for the thread pool. See documentation for exact details. Do note this may have maximum value depending on CPU count - value is automatically lowered if set to higher than maximum value. opensearch.thread\_pool\_search\_queue\_size integer - min: 10 - max: 2000 search thread pool queue size Size for the thread pool queue. See documentation for exact details. opensearch.thread\_pool\_search\_throttled\_queue\_size integer - min: 10 - max: 2000 search\_throttled thread pool queue size Size for the thread pool queue. See documentation for exact details. opensearch.thread\_pool\_get\_queue\_size integer - min: 10 - max: 2000 Size for the thread pool queue. See documentation for exact details. opensearch.thread\_pool\_analyze\_queue\_size integer - min: 10 - max: 2000 analyze thread pool queue size Size for the thread pool queue. See documentation for exact details. opensearch.thread\_pool\_write\_queue\_size integer - min: 10 - max: 2000 write thread pool queue size Size for the thread pool queue. See documentation for exact details. opensearch.action\_destructive\_requires\_name boolean,null Require explicit index names when deleting opensearch.cluster\_max\_shards\_per\_node integer - min: 10 - max: 10000 cluster.max\_shards\_per\_node Controls the number of shards allowed in the cluster per data node opensearch.override\_main\_response\_version boolean (DEPRECATED) compatibility.override\_main\_response\_version Compatibility mode sets OpenSearch to report its version as 7.10 so clients continue to work. Default is false. Deprecated and ignored for service version 3.3 and higher. opensearch.script\_max\_compilations\_rate string Script max compilation rate - circuit breaker to prevent/minimize OOMs Script compilation circuit breaker limits the number of inline script compilations within a period of time. Default is use-context opensearch.cluster\_routing\_allocation\_node\_concurrent\_recoveries integer - min: 2 - max: 16 How many concurrent incoming/outgoing shard recoveries (normally replicas) are allowed to happen on a node. Defaults to node cpu count \* 2. opensearch.email\_sender\_name string This should be identical to the Sender name defined in Opensearch dashboards opensearch.email\_sender\_username string Sender username for Opensearch alerts opensearch.email\_sender\_password string Sender password for Opensearch alerts to authenticate with SMTP server opensearch.ism\_enabled boolean Specifies whether ISM is enabled or not opensearch.ism\_history\_enabled boolean Specifies whether audit history is enabled or not. The logs from ISM are automatically indexed to a logs document. opensearch.ism\_history\_max\_age integer - min: 1 - max: 2147483647 The maximum age before rolling over the audit history index in hours opensearch.ism\_history\_max\_docs integer - min: 1 - max: 9223372036854776000 The maximum number of documents before rolling over the audit history index. opensearch.ism\_history\_rollover\_check\_period integer - min: 1 - max: 2147483647 The time between rollover checks for the audit history index in hours. opensearch.ism\_history\_rollover\_retention\_period integer - min: 1 - max: 2147483647 How long audit history indices are kept in days. opensearch.search\_backpressure object Search Backpressure Settings opensearch.search\_backpressure.mode string The search backpressure mode. Valid values are monitor\_only, enforced, or disabled. Default is monitor\_only opensearch.search\_backpressure.node\_duress object Node duress settings opensearch.search\_backpressure.node\_duress.cpu\_threshold number - max: 1 The CPU usage threshold (as a percentage) required for a node to be considered to be under duress. Default is 0.9 opensearch.search\_backpressure.node\_duress.heap\_threshold number - max: 1 The heap usage threshold (as a percentage) required for a node to be considered to be under duress. Default is 0.7 opensearch.search\_backpressure.node\_duress.num\_successive\_breaches integer - min: 1 The number of successive limit breaches after which the node is considered to be under duress. Default is 3 opensearch.search\_backpressure.search\_task object Search task settings opensearch.search\_backpressure.search\_task.cancellation\_burst number - min: 1 The maximum number of search tasks to cancel in a single iteration of the observer thread. Default is 5.0 opensearch.search\_backpressure.search\_task.cancellation\_rate number The maximum number of search tasks to cancel per millisecond of elapsed time. Default is 0.003 opensearch.search\_backpressure.search\_task.cancellation\_ratio number - max: 1 The maximum number of search tasks to cancel, as a percentage of successful search task completions. Default is 0.1 opensearch.search\_backpressure.search\_task.cpu\_time\_millis\_threshold integer The CPU usage threshold (in milliseconds) required for an individual parent task before it is considered for cancellation. Default is 30000 opensearch.search\_backpressure.search\_task.elapsed\_time\_millis\_threshold integer The elapsed time threshold (in milliseconds) required for an individual parent task before it is considered for cancellation. Default is 45000 opensearch.search\_backpressure.search\_task.heap\_moving\_average\_window\_size integer The window size used to calculate the rolling average of the heap usage for the completed parent tasks. Default is 10 opensearch.search\_backpressure.search\_task.heap\_percent\_threshold number - max: 1 The heap usage threshold (as a percentage) required for an individual parent task before it is considered for cancellation. Default is 0.2 opensearch.search\_backpressure.search\_task.heap\_variance number The heap usage variance required for an individual parent task before it is considered for cancellation. A task is considered for cancellation when taskHeapUsage is greater than or equal to heapUsageMovingAverage \* variance. Default is 2.0 opensearch.search\_backpressure.search\_task.total\_heap\_percent\_threshold number - max: 1 The heap usage threshold (as a percentage) required for the sum of heap usages of all search tasks before cancellation is applied. Default is 0.5 opensearch.search\_backpressure.search\_shard\_task object Search shard settings opensearch.search\_backpressure.search\_shard\_task.cancellation\_burst number - min: 1 The maximum number of search tasks to cancel in a single iteration of the observer thread. Default is 10.0 opensearch.search\_backpressure.search\_shard\_task.cancellation\_rate number The maximum number of tasks to cancel per millisecond of elapsed time. Default is 0.003 opensearch.search\_backpressure.search\_shard\_task.cancellation\_ratio number - max: 1 The maximum number of tasks to cancel, as a percentage of successful task completions. Default is 0.1 opensearch.search\_backpressure.search\_shard\_task.cpu\_time\_millis\_threshold integer The CPU usage threshold (in milliseconds) required for a single search shard task before it is considered for cancellation. Default is 15000 opensearch.search\_backpressure.search\_shard\_task.elapsed\_time\_millis\_threshold integer The elapsed time threshold (in milliseconds) required for a single search shard task before it is considered for cancellation. Default is 30000 opensearch.search\_backpressure.search\_shard\_task.heap\_moving\_average\_window\_size integer The number of previously completed search shard tasks to consider when calculating the rolling average of heap usage. Default is 100 opensearch.search\_backpressure.search\_shard\_task.heap\_percent\_threshold number - max: 1 The heap usage threshold (as a percentage) required for a single search shard task before it is considered for cancellation. Default is 0.5 opensearch.search\_backpressure.search\_shard\_task.heap\_variance number The minimum variance required for a single search shard task’s heap usage compared to the rolling average of previously completed tasks before it is considered for cancellation. Default is 2.0 opensearch.search\_backpressure.search\_shard\_task.total\_heap\_percent\_threshold number - max: 1 The heap usage threshold (as a percentage) required for the sum of heap usages of all search shard tasks before cancellation is applied. Default is 0.5 opensearch.shard\_indexing\_pressure object Shard indexing back pressure settings opensearch.shard\_indexing\_pressure.enabled boolean Enable or disable shard indexing backpressure. Default is false opensearch.shard\_indexing\_pressure.enforced boolean Run shard indexing backpressure in shadow mode or enforced mode. In shadow mode (value set as false), shard indexing backpressure tracks all granular-level metrics, but it doesn’t actually reject any indexing requests. In enforced mode (value set as true), shard indexing backpressure rejects any requests to the cluster that might cause a dip in its performance. Default is false opensearch.shard\_indexing\_pressure.primary\_parameter object Primary parameter opensearch.shard\_indexing\_pressure.primary\_parameter.node object opensearch.shard\_indexing\_pressure.primary\_parameter.node.soft\_limit number Node soft limit Define the percentage of the node-level memory threshold that acts as a soft indicator for strain on a node. Default is 0.7 opensearch.shard\_indexing\_pressure.primary\_parameter.shard object opensearch.shard\_indexing\_pressure.primary\_parameter.shard.min\_limit number Shard min limit Specify the minimum assigned quota for a new shard in any role (coordinator, primary, or replica). Shard indexing backpressure increases or decreases this allocated quota based on the inflow of traffic for the shard. Default is 0.001 opensearch.shard\_indexing\_pressure.operating\_factor object Operating factor opensearch.shard\_indexing\_pressure.operating\_factor.lower number Specify the lower occupancy limit of the allocated quota of memory for the shard. If the total memory usage of a shard is below this limit, shard indexing backpressure decreases the current allocated memory for that shard. Default is 0.75 opensearch.shard\_indexing\_pressure.operating\_factor.optimal number Specify the optimal occupancy of the allocated quota of memory for the shard. If the total memory usage of a shard is at this level, shard indexing backpressure doesn’t change the current allocated memory for that shard. Default is 0.85 opensearch.shard\_indexing\_pressure.operating\_factor.upper number Specify the upper occupancy limit of the allocated quota of memory for the shard. If the total memory usage of a shard is above this limit, shard indexing backpressure increases the current allocated memory for that shard. Default is 0.95 opensearch.search.insights.top\_queries object opensearch.search.insights.top\_queries.cpu object Top N queries monitoring by CPU opensearch.search.insights.top\_queries.cpu.enabled boolean Enable or disable top N query monitoring by the metric opensearch.search.insights.top\_queries.cpu.top\_n\_size integer - min: 1 Specify the value of N for the top N queries by the metric opensearch.search.insights.top\_queries.cpu.window\_size string The window size of the top N queries by the metric Configure the window size of the top N queries. The value should be a time value with unit, e.g. 1m, 5s, 1h. opensearch.search.insights.top\_queries.latency object Top N queries monitoring by latency opensearch.search.insights.top\_queries.latency.enabled boolean Enable or disable top N query monitoring by the metric opensearch.search.insights.top\_queries.latency.top\_n\_size integer - min: 1 Specify the value of N for the top N queries by the metric opensearch.search.insights.top\_queries.latency.window\_size string The window size of the top N queries by the metric Configure the window size of the top N queries. The value should be a time value with unit, e.g. 1m, 5s, 1h. opensearch.search.insights.top\_queries.memory object Top N queries monitoring by memory opensearch.search.insights.top\_queries.memory.enabled boolean Enable or disable top N query monitoring by the metric opensearch.search.insights.top\_queries.memory.top\_n\_size integer - min: 1 Specify the value of N for the top N queries by the metric opensearch.search.insights.top\_queries.memory.window\_size string The window size of the top N queries by the metric Configure the window size of the top N queries. The value should be a time value with unit, e.g. 1m, 5s, 1h. opensearch.cluster.routing.allocation.balance.prefer\_primary boolean cluster.routing.allocation.balance.prefer\_primary When set to true, OpenSearch attempts to evenly distribute the primary shards between the cluster nodes. Enabling this setting does not always guarantee an equal number of primary shards on each node, especially in the event of a failover. Changing this setting to false after it was set to true does not invoke redistribution of primary shards. Default is false. opensearch.disk\_watermarks object Watermark settings opensearch.disk\_watermarks.low integer Low watermark (percentage) The low watermark for disk usage. opensearch.disk\_watermarks.high integer The high watermark for disk usage. opensearch.disk\_watermarks.flood\_stage integer The flood stage watermark for disk usage. opensearch.segrep object Segment Replication Backpressure Settings opensearch.segrep.pressure.enabled boolean segrep.pressure.enabled Enables the segment replication backpressure mechanism. Default is false. opensearch.segrep.pressure.time.limit string - default: 5m The maximum amount of time that a replica shard can take to copy from the primary shard. Once segrep.pressure.time.limit is breached along with segrep.pressure.checkpoint.limit, the segment replication backpressure mechanism is initiated. Default is 5 minutes. opensearch.segrep.pressure.checkpoint.limit integer - default: 4 The maximum number of indexing checkpoints that a replica shard can fall behind when copying from primary. Once segrep.pressure.checkpoint.limit is breached along with segrep.pressure.time.limit, the segment replication backpressure mechanism is initiated. Default is 4 checkpoints. opensearch.segrep.pressure.replica.stale.limit number - max: 1 - default: 0.5 The maximum number of stale replica shards that can exist in a replication group. Once segrep.pressure.replica.stale.limit is breached, the segment replication backpressure mechanism is initiated. Default is .5, which is 50% of a replication group. opensearch.cluster.remote\_store object opensearch.cluster.remote\_store.translog.buffer\_interval string The default value of the translog buffer interval used when performing periodic translog updates. This setting is only effective when the index setting index.remote\_store.translog.buffer\_interval is not present. Defaults to 650ms. opensearch.cluster.remote\_store.translog.max\_readers integer - min: 100 - max: 2147483647 Sets the maximum number of open translog files for remote-backed indexes. This limits the total number of translog files per shard. After reaching this limit, the remote store flushes the translog files. Default is 1000. The minimum required is 100. opensearch.cluster.remote\_store.state.global\_metadata.upload\_timeout string The amount of time to wait for the cluster state upload to complete. Defaults to 20s. opensearch.cluster.remote\_store.state.metadata\_manifest.upload\_timeout string The amount of time to wait for the manifest file upload to complete. The manifest file contains the details of each of the files uploaded for a single cluster state, both index metadata files and global metadata files. Defaults to 20s. opensearch.remote\_store object opensearch.remote\_store.segment.pressure.enabled boolean Enables remote segment backpressure. Default is true opensearch.remote\_store.segment.pressure.consecutive\_failures.limit integer - min: 1 - max: 2147483647 The minimum consecutive failure count for activating remote segment backpressure. Defaults to 5. opensearch.remote\_store.segment.pressure.bytes\_lag.variance\_factor number - min: 1 The variance factor that is used together with the moving average to calculate the dynamic bytes lag threshold for activating remote segment backpressure. Defaults to 10. opensearch.remote\_store.segment.pressure.time\_lag.variance\_factor number - min: 1 The variance factor that is used together with the moving average to calculate the dynamic time lag threshold for activating remote segment backpressure. Defaults to 10. opensearch.cluster.filecache.remote\_data\_ratio integer,number,null - max: 100 Defines a limit of how much total remote data can be referenced as a ratio of the size of the disk reserved for the file cache. This is designed to be a safeguard to prevent oversubscribing a cluster. Defaults to 0. opensearch.node.search.cache.size string,null Defines a limit of how much total remote data can be referenced as a ratio of the size of the disk reserved for the file cache. This is designed to be a safeguard to prevent oversubscribing a cluster. Defaults to 5gb. Requires restarting all OpenSearch nodes. opensearch.cluster.search.request.slowlog object opensearch.cluster.search.request.slowlog.level string - default: trace Log level opensearch.cluster.search.request.slowlog.threshold object opensearch.cluster.search.request.slowlog.threshold.debug string Debug threshold for total request took time. The value should be in the form count and unit, where unit one of (s,m,h,d,nanos,ms,micros) or -1. Default is -1 opensearch.cluster.search.request.slowlog.threshold.info string Info threshold for total request took time. The value should be in the form count and unit, where unit one of (s,m,h,d,nanos,ms,micros) or -1. Default is -1 opensearch.cluster.search.request.slowlog.threshold.trace string Trace threshold for total request took time. The value should be in the form count and unit, where unit one of (s,m,h,d,nanos,ms,micros) or -1. Default is -1 opensearch.cluster.search.request.slowlog.threshold.warn string Warning threshold for total request took time. The value should be in the form count and unit, where unit one of (s,m,h,d,nanos,ms,micros) or -1. Default is -1 opensearch.enable\_remote\_backed\_storage boolean Enable remote-backed storage opensearch.enable\_searchable\_snapshots boolean Enable searchable snapshots | | []()[**index\_template**](#index_template)`object`Template settings for all new indexesindex\_template.mapping\_nested\_objects\_limit integer,null - max: 100000 (DEPRECATED) index.mapping.nested\_objects.limit The maximum number of nested JSON objects that a single document can contain across all nested types. This limit helps to prevent out of memory errors when a document contains too many nested objects. Default is 10000. Deprecated, use an index template instead. index\_template.number\_of\_shards integer,null - min: 1 - max: 1024 (DEPRECATED) index.number\_of\_shards The number of primary shards that an index should have. Deprecated, use an index template instead. index\_template.number\_of\_replicas integer,null - max: 29 (DEPRECATED) index.number\_of\_replicas The number of replicas each primary shard has. Deprecated, use an index template instead. | | []()[**private\_access**](#private_access)`object`Allow access to selected service ports from private networksprivate\_access.opensearch boolean Allow clients to connect to opensearch with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.opensearch\_dashboards boolean Allow clients to connect to opensearch\_dashboards with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.prometheus boolean Allow clients to connect to prometheus with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | []()[**privatelink\_access**](#privatelink_access)`object`Allow access to selected service components through Privatelinkprivatelink\_access.opensearch boolean Enable opensearch privatelink\_access.opensearch\_dashboards boolean Enable opensearch\_dashboards privatelink\_access.prometheus boolean Enable prometheus | | []()[**public\_access**](#public_access)`object`Allow access to selected service ports from the public Internetpublic\_access.opensearch boolean Allow clients to connect to opensearch from the public internet for service nodes that are in a project VPC or another type of private network public\_access.opensearch\_dashboards boolean Allow clients to connect to opensearch\_dashboards from the public internet for service nodes that are in a project VPC or another type of private network public\_access.prometheus boolean Allow clients to connect to prometheus from the public internet for service nodes that are in a project VPC or another type of private network | | []()[**recovery\_basebackup\_name**](#recovery_basebackup_name)`string`Name of the basebackup to restore in forked service | | []()[**service\_to\_fork\_from**](#service_to_fork_from)`string,null`Name of another service to fork from. This has effect only when a new service is being created. | | []()[**project\_to\_fork\_from**](#project_to_fork_from)`string,null`Name of another project to fork a service from. This has effect only when a new service is being created. | --- # Plugin versions per OpenSearch release Plugin availability and versions in Aiven for OpenSearch® vary by OpenSearch major version. Each plugin version corresponds to the OpenSearch core version. OpenSearch\_core\_version=OpenSearch\_plugin\_version Example OpenSearch 2.19 uses plugins version 2.19. ## OpenSearch 3 plugins[​](#opensearch-3-plugins "Direct link to OpenSearch 3 plugins") | Plugin name | Supported version | | ------------------------------------ | ----------------- | | analysis-icu | 3.3, 3.6 | | analysis-kuromoji | 3.3, 3.6 | | analysis-phonetic | 3.3, 3.6 | | ingest-attachment | 3.3, 3.6 | | mapper-size | 3.3, 3.6 | | opensearch-alerting | 3.3, 3.6 | | opensearch-anomaly-detection | 3.3, 3.6 | | opensearch-asynchronous-search | 3.3, 3.6 | | opensearch-cross-cluster-replication | 3.3, 3.6 | | opensearch-custom-codecs | 3.3, 3.6 | | opensearch-flow-framework | 3.3, 3.6 | | opensearch-geospatial | 3.3, 3.6 | | opensearch-index-management | 3.3, 3.6 | | opensearch-job-scheduler | 3.3, 3.6 | | opensearch-knn | 3.3, 3.6 | | opensearch-ml | 3.3, 3.6 | | opensearch-ltr | 3.3, 3.6 | | opensearch-neural-search | 3.3, 3.6 | | opensearch-notifications | 3.3, 3.6 | | opensearch-notifications-core | 3.3, 3.6 | | opensearch-observability | 3.3, 3.6 | | opensearch-reports-scheduler | 3.3, 3.6 | | opensearch-security | 3.3, 3.6 | | opensearch-security-analytics | 3.3, 3.6 | | opensearch-skills | 3.3, 3.6 | | opensearch-sql | 3.3, 3.6 | | opensearch-system-templates | 3.3, 3.6 | | query-insights | 3.3, 3.6 | | opensearch-ubi | 3.6 | | repository-azure | 3.3, 3.6 | | repository-gcs | 3.3, 3.6 | | repository-s3 | 3.3, 3.6 | ## OpenSearch 2 plugins[​](#opensearch-2-plugins "Direct link to OpenSearch 2 plugins") | Plugin name | Supported version | | ------------------------------------ | ----------------- | | analysis-icu | 2.19 | | analysis-kuromoji | 2.19 | | analysis-phonetic | 2.19 | | ingest-attachment | 2.19 | | mapper-size | 2.19 | | opensearch-alerting | 2.19 | | opensearch-anomaly-detection | 2.19 | | opensearch-asynchronous-search | 2.19 | | opensearch-cross-cluster-replication | 2.19 | | opensearch-custom-codecs | 2.19 | | opensearch-flow-framework | 2.19 | | opensearch-geospatial | 2.19 | | opensearch-index-management | 2.19 | | opensearch-job-scheduler | 2.19 | | opensearch-knn | 2.19 | | opensearch-ltr | 2.19 | | opensearch-ml | 2.19 | | opensearch-neural-search | 2.19 | | opensearch-notifications | 2.19 | | opensearch-notifications-core | 2.19 | | opensearch-observability | 2.19 | | opensearch-reports-scheduler | 2.19 | | opensearch-security | 2.19 | | opensearch-security-analytics | 2.19 | | opensearch-skills | 2.19 | | opensearch-sql | 2.19 | | opensearch-system-templates | 2.19 | | query-insights | 2.19 | | repository-azure | 2.19 | | repository-gcs | 2.19 | | repository-s3 | 2.19 | --- # Low disk space watermarks OpenSearch® relies on three indicators to identify and respond to low disk space. ## Disk allocation low watermark[​](#disk-allocation-low-watermark "Direct link to Disk allocation low watermark") Defined by parameter `cluster.routing.allocation.disk.watermark.low` and the default value is set to 85% of the disk space. When this limit is exceeded, OpenSearch starts avoiding allocating new shards to the server. On a single-server OpenSearch, this has no effect. On a multi-server cluster, amount of data is not always equally distributed, in which case this will help balancing disk usage between servers. ## Disk allocation high watermark[​](#disk-allocation-high-watermark "Direct link to Disk allocation high watermark") Defined by parameter `cluster.routing.allocation.disk.watermark.high` and the default value is set to 90% of the disk space. When this limit is exceeded, OpenSearch will actively try to move shards to other servers with more available disk space. On a single-server OpenSearch, this has no effect. ## Disk allocation flood stage watermark[​](#disk-allocation-flood-stage-watermark "Direct link to Disk allocation flood stage watermark") Defined by parameter `cluster.routing.allocation.disk.watermark.flood_stage` and the default value is set to 95% of the disk space. When this limit is exceeded, OpenSearch will mark all indices hosted on the server exceeding this limit as read-only, allowing deletes (`index.blocks.read_only_allow_delete`). --- # Aiven for OpenSearch® limits and limitations Aiven for OpenSearch® has configuration, API, and feature restrictions that differ from upstream OpenSearch to maintain service stability and security. ## Configuration restrictions[​](#configuration-restrictions "Direct link to Configuration restrictions") You cannot directly modify OpenSearch configuration files or settings in Aiven for OpenSearch. These restrictions apply: | Restriction | Description | | -------------------------- | -------------------------------------------------------------------------------------------- | | **No shell access** | You cannot access or modify YAML configuration files | | **JVM tuning** | You cannot modify JVM options directly | | **Advanced configuration** | Only supported options are available through **Advanced configuration** in the Aiven Console | | **Configuration files** | You cannot access or modify static configuration files | To request support for additional configuration options, [contact Aiven support](http://support.aiven.io/). ## Connection requirements[​](#connection-requirements "Direct link to Connection requirements") All connections to Aiven for OpenSearch must meet these requirements: | Requirement | Details | | ------------------ | ------------------------------------------------------------------------------------- | | **Protocol** | HTTPS only | | **Authentication** | User authentication always required | | **Authorization** | Managed using Aiven ACLs or OpenSearch Security (when security management is enabled) | ## API restrictions[​](#api-restrictions "Direct link to API restrictions") Aiven restricts access to certain OpenSearch APIs to maintain service stability and security. Attempting to access blocked endpoints returns a `403 Forbidden - Request forbidden by administrative rules` error. | API endpoint | Allowed methods | Restrictions | | -------------------- | --------------- | ------------------------------------------------------------------------------------- | | `/_cluster/*` | `GET` only | Limited to specific read-only endpoints; all other `/_cluster/` endpoints are blocked | | `/_tasks` | `GET` only | View tasks only; you cannot cancel tasks using `/_tasks/_cancel` | | `/_nodes` | `GET` only | Read-only access to node information | | `/_snapshot` | None | Automated by Aiven; no direct access | | `/_cat/repositories` | None | No access allowed | ### Allowed cluster endpoints[​](#allowed-cluster-endpoints "Direct link to Allowed cluster endpoints") You can access these read-only cluster endpoints: * `/_cluster/allocation/explain/` * `/_cluster/health/` * `/_cluster/pending_tasks/` * `/_cluster/stats/` * `/_cluster/state/` * `/_cluster/settings/` ## Snapshot management[​](#snapshot-management "Direct link to Snapshot management") | Feature | Behavior | | ------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Automated snapshots** | Daily or hourly snapshots managed automatically by Aiven | | **API access** | You cannot access the snapshot API directly without configuring custom repositories | | **OpenSearch API** | - [Restore from snapshot](/docs/products/opensearch/howto/manage-snapshots.md#restore-from-snapshots) has [security-related restrictions](https://docs.opensearch.org/latest/tuning-your-cluster/availability-and-recovery/snapshots/snapshot-restore/#security-considerations).
- [Create a snapshot](/docs/products/opensearch/howto/manage-snapshots.md#create-a-snapshot) and [delete a snapshot](/docs/products/opensearch/howto/manage-snapshots.md#delete-a-snapshot) are not supported for snapshots in Aiven-managed repositories (prefixed with `aiven_repo`). | | **Dashboard limitations** | Dashboard suggestions for snapshot management that require configuration file changes cannot be completed | See [snapshot management limitations](/docs/products/opensearch/howto/manage-snapshots.md#limitations) for details. ## Plugin restrictions[​](#plugin-restrictions "Direct link to Plugin restrictions") You can only use pre-approved plugins with Aiven for OpenSearch. | Aspect | Details | | --------------------- | ----------------------------------------------------------------------- | | **Supported plugins** | Only a defined set of plugins is available | | **Custom plugins** | You cannot install custom plugins | | **Plugin list** | See [available plugins](/docs/products/opensearch/reference/plugins.md) | To request support for additional plugins, [contact Aiven support](http://support.aiven.io/). ## Access control models[​](#access-control-models "Direct link to Access control models") Aiven for OpenSearch supports two access control models with different limitations: * [Security management disabled (default)](/docs/products/opensearch/reference/opensearch-limitations.md#security-management-disabled) * [Security management enabled](/docs/products/opensearch/reference/opensearch-limitations.md#security-management-enabled) note To turn on security management, see [Enable security management](/docs/products/opensearch/howto/enable-opensearch-security.md). ### Security management disabled[​](#security-management-disabled "Direct link to Security management disabled") | Feature | Behavior | | --------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **User management** | You manage users through Aiven API, CLI, Console, or Terraform | | **Access control** | You configure access using Aiven ACLs | | **Permission scope** | Index-level access only | | **User equality** | All service users have equal privileges within their ACL permissions | | **Dashboard tenancy** | Private dashboards per user plus global dashboards | | **Password changes** | You change passwords using the Aiven Console. Password changes you make in the OpenSearch dashboard are overwritten during service configuration updates, which occur daily | ### Security management enabled[​](#security-management-enabled "Direct link to Security management enabled") | Feature | Behavior | | --------------------------- | ------------------------------------------------------------------------------------------------- | | **User management** | You manage users directly in OpenSearch using OpenSearch Security API or dashboard | | **Access control** | You configure access using OpenSearch Security roles and permissions | | **Permission scope** | Document-level access control available | | **Dashboard tenancy** | Full multi-tenancy support | | **External authentication** | SAML and OpenID Connect supported | | **Password changes** | You manage passwords directly in OpenSearch. Password changes in the Aiven Console have no effect | | **Aiven API support** | Limited; displays state at enablement time only | warning You cannot reverse security management after you enable it. Once enabled, you manage all users and permissions directly in OpenSearch. note The security plugin is always present in Aiven for OpenSearch. Security management is an additional feature you can enable to gain full control over security configurations. ## ACL limitations[​](#acl-limitations "Direct link to ACL limitations") When [security management](/docs/products/opensearch/concepts/os-security.md) is disabled, Aiven ACLs control access to your service. | Limitation | Description | | ---------------------------- | ----------------------------------------------------------------------------------------------------------------------- | | **Index patterns only** | You can only define ACL rules using index patterns; rules for top-level APIs like `_bulk` or `_search` are not enforced | | **Index-level access** | You can control access to indices but not to OpenSearch Dashboards | | **Predefined action groups** | ACL access levels are fixed; you cannot create custom permission sets | note When you [enable security management](/docs/products/opensearch/howto/enable-opensearch-security.md), Aiven ACLs no longer apply. You manage all permissions using OpenSearch Security roles. ### ACL access levels[​](#acl-access-levels "Direct link to ACL access levels") Use these access levels when configuring ACLs: | ACL level | Permissions granted | | ----------- | ----------------------------------------- | | `admin` | Full access to matching indices | | `read` | Read-only access to matching indices | | `write` | Write access to matching indices | | `readwrite` | Read and write access to matching indices | ## Reserved users[​](#reserved-users "Direct link to Reserved users") Aiven creates and manages these special users. You cannot delete or modify their permissions. | Username | Purpose | | ---------------------- | ------------------------------------------------------------------------ | | `avnadmin` | Default administrator user for your service | | `metrics_user_datadog` | Metrics collection by Datadog integration | | `osd_internal_user` | Internal OpenSearch Dashboards operations | | `replication_user` | Cross-cluster replication | | `os-sec-admin` | Security management access (created when you enable security management) | ## Reserved roles[​](#reserved-roles "Direct link to Reserved roles") When [security management](/docs/products/opensearch/concepts/os-security.md) is disabled, you cannot modify the reserved roles. When you [enable security management](/docs/products/opensearch/howto/enable-opensearch-security.md), you can modify the `provider_*` roles but not the `service_security_admin_access` role. | Role name | Purpose | | --------------------------------------- | ------------------------------------------------------ | | `service_security_admin_access` | Grants access to security management API and dashboard | | `provider_service_user` | Base permissions for all service users | | `provider_index_all_access` | Full index access (when ACLs are disabled) | | `provider_managed_user_role_` | Individual user permissions (when ACLs are enabled) | ## Known issues and limitations[​](#known-issues-and-limitations "Direct link to Known issues and limitations") ### Security dashboard[​](#security-dashboard "Direct link to Security dashboard") | Issue | Description | | ----------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Get started** page | Most content is not applicable to Aiven for OpenSearch; only the **multi-tenancy** section applies | | **Configuration file instructions** | Dashboard help text references configuration file modifications that you cannot perform in managed services | | **Password changes** | When security management is disabled, you change passwords using the Aiven Console. Password changes you make in the OpenSearch dashboard are overwritten during service configuration updates, which occur daily. When security management is enabled, you change passwords directly in OpenSearch and Aiven Console password changes have no effect | ### Security management[​](#security-management "Direct link to Security management") | Issue | Description | Solution | | ------------------------- | ------------------------------------------------- | ---------------------------------------------------------------------------------- | | **REST API permissions** | You cannot create roles with REST API permissions | Map your users to the `service_security_admin_access` role | | **Self-lockout** | You can unmap yourself from security admin role | [Contact Aiven support](http://support.aiven.io/) to remap the `os-sec-admin` user | | **os-sec-admin deletion** | You cannot delete the `os-sec-admin` user | User remains but you can unmap it from admin role | ### Permissions model[​](#permissions-model "Direct link to Permissions model") | Behavior | Description | | ------------------ | ------------------------------------------------------------------------ | | **Index creation** | Writing to non-existent index requires both write and create permissions | | **Error messages** | Permission errors specify the missing permission in `error.root_cause` | ## Differences from upstream OpenSearch[​](#differences-from-upstream-opensearch "Direct link to Differences from upstream OpenSearch") | Feature | Upstream OpenSearch | Aiven for OpenSearch | | ----------------------- | ---------------------------- | ------------------------------------------------------------------------------------------------------------------- | | **Configuration files** | Direct file access | You manage configuration using Advanced configuration options | | **Snapshot management** | Full API access | Automated; you cannot access the API directly | | **Security plugin** | Optional | Always enabled | | **User management** | Direct configuration | You manage users using Aiven tools or Security API (when security management is enabled) | | **Cluster settings** | Full API access | Limited to approved settings using [advanced configuration](/docs/products/opensearch/reference/advanced-params.md) | | **Plugin installation** | Install any plugin | Only selected [plugins available](/docs/products/opensearch/reference/plugins.md) | | **API access** | Full access to all APIs | Restricted access to certain management APIs | | **JVM tuning** | Direct access to JVM options | Not available | ## Elasticsearch compatibility[​](#elasticsearch-compatibility "Direct link to Elasticsearch compatibility") Aiven for OpenSearch diverged from Elasticsearch 7 and is not compatible with Elasticsearch-specific features. | Aspect | Details | | -------------------- | ------------------------------------------------------ | | **Client libraries** | You must use OpenSearch-compatible client libraries | | **APIs** | Elasticsearch-specific APIs are not supported | | **Migration** | Verify compatibility when migrating from Elasticsearch | ## Service tiers and quotas[​](#service-tiers-and-quotas "Direct link to Service tiers and quotas") For information about service-specific limits based on your plan, see: * [Quotas for Business and Premium plans](https://aiven.io/pricing?tab=plan-pricing\&product=opensearch) * [Plans comparison](https://aiven.io/pricing?tab=plan-comparison\&product=opensearch) Related pages * [OpenSearch security](/docs/products/opensearch/concepts/os-security.md) * [Enable security management](/docs/products/opensearch/howto/enable-opensearch-security.md) * [Access control](/docs/products/opensearch/concepts/access_control.md) * [Manage access and access control lists](/docs/products/opensearch/howto/control_access_to_content.md) * [Available plugins](/docs/products/opensearch/reference/plugins.md) --- # Plugins available with Aiven for OpenSearch® Aiven for OpenSearch® includes a standard set of plugins. In addition to the plugins that were previously available in Aiven for Elasticsearch, Aiven for OpenSearch also includes plugins that are designed and developed specifically for OpenSearch. ## List of plugins[​](#list-of-plugins "Direct link to List of plugins") note Each version of Aiven for OpenSearch comes with its own set of supported extension versions. You can only use the plugin versions supported for the Aiven for OpenSearch version that your service uses. For details on particular plugins, see the [OpenSearch documentation](https://docs.opensearch.org/docs/latest/about/) or [Elasticsearch documentation](https://www.elastic.co/docs/reference/elasticsearch/plugins/), and make sure you preview the documentation version corresponding to the Aiven for OpenSearch version that your service uses. Depending on the Aiven for OpenSearch version your service runs on, the following plugins are available: * [Anomaly Detection](https://github.com/opensearch-project/anomaly-detection) * [Asynchronous search](https://github.com/opensearch-project/asynchronous-search) * [ICU Analysis](https://github.com/opensearch-project/OpenSearch/tree/main/plugins/analysis-icu) * [Index Management](https://github.com/opensearch-project/index-management) * [Ingest Attachment](https://github.com/opensearch-project/OpenSearch/tree/2.15/plugins/ingest-attachment) * [Job Scheduler](https://github.com/opensearch-project/job-scheduler) * [k-NN](https://github.com/opensearch-project/k-NN) * [Kuromoji (Japanese Analysis)](https://github.com/opensearch-project/OpenSearch/tree/main/plugins/analysis-kuromoji) * [Learning to Rank](https://github.com/opensearch-project/opensearch-learning-to-rank-base) (OpenSearch 2.19.5 and later) * [Mapper Size](https://github.com/opensearch-project/OpenSearch/tree/main/plugins/mapper-size) * [Neural Search](https://github.com/opensearch-project/neural-search) * [Notebooks](https://github.com/opensearch-project/dashboards-notebooks) * [OpenSearch Dashboards Alerting](https://github.com/opensearch-project/alerting-dashboards-plugin) * [OpenSearch Dashboards Gantt Charts](https://github.com/opensearch-project/dashboards-visualizations) * [OpenSearch Dashboards Reports](https://github.com/opensearch-project/dashboards-reporting) * [OpenSearch Dashboards Trace Analytics](https://github.com/opensearch-project/trace-analytics) * [OpenSearch Notifications](https://github.com/opensearch-project/notifications) * [OpenSearch Observability](https://github.com/opensearch-project/dashboards-observability) * [OpenSearch security](/docs/products/opensearch/concepts/os-security.md) for RBAC, SAML, and OIDC * [OpenSearch Security Analytics](https://github.com/opensearch-project/security-analytics) * [OpenSearch SQL](https://github.com/opensearch-project/sql) * [Phonetic analysis](https://github.com/opensearch-project/OpenSearch/tree/main/plugins/analysis-phonetic) * [Query Insights](https://github.com/opensearch-project/query-insights) * [Scheduler for Dashboards Reports](https://github.com/opensearch-project/dashboards-reporting) * [User Behavior Insights](https://github.com/opensearch-project/user-behavior-insights) note The **Notebooks** and **OpenSearch Notifications** plugins are part of **OpenSearch Observability**. ## Plugins per release[​](#plugins-per-release "Direct link to Plugins per release") Plugin availability and versions in Aiven for OpenSearch® vary by OpenSearch major version. Each plugin version corresponds to the OpenSearch core version. OpenSearch\_core\_version=OpenSearch\_plugin\_version Example OpenSearch 2.19 uses plugins version 2.19. See [plugins and plugin versions](/docs/products/opensearch/reference/list-of-plugins-for-each-version.md) available in each supported Aiven for OpenSearch release. ## Request a new plugin[​](#request-a-new-plugin "Direct link to Request a new plugin") To request a new plugin, submit an new idea through our [Aiven Ideas](https://ideas.aiven.io/) portal. We're always looking to expand our offerings, and your feedback and ideas are important in shaping the future updates of our product. --- # Aiven for OpenSearch® version lifecycle Learn how Aiven manages Aiven for OpenSearch® version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. ## Aiven version support and upstream EOL[​](#aiven-version-support-and-upstream-eol "Direct link to Aiven version support and upstream EOL") Aiven aims to follow the EOL schedule set by the original authors and maintainers of the open source software (the upstream projects). Once the upstream project retires a specific version, they do not receive security updates and critical bug fixes anymore by the maintainers. Outdated services don't offer the level of protection you need, so Aiven follows the upstream project's EOL schedule to ensure that Aiven services are always running on supported versions. ## Service version numbering[​](#service-version-numbering "Direct link to Service version numbering") Aiven services inherit the upstream project's software versioning scheme. Depending on the service, a major version can be either a single digit or in the format `major.minor`. The exact version of the service is visible in the [Aiven Console](https://console.aiven.io/) when the service is running. ## Service version EOL policy[​](#service-version-eol-policy "Direct link to Service version EOL policy") Aiven sets an EOL date for each major version of the service. This policy covers both running and powered-off services on affected versions. ## EOL notifications[​](#eol-notifications "Direct link to EOL notifications") When Aiven sets the EOL date for a service major version: * You receive an email notification along with instructions on the next steps. * The [Aiven Console](https://console.aiven.io/) shows an EOL alert for affected services. * You receive email reminders monthly. * In the month of the EOL date, you receive weekly reminders. ## EOL best practices[​](#eol-best-practices "Direct link to EOL best practices") * Use service forking to test the version upgrade before upgrading your production services. * Upgrade to the supported version before the EOL date. This gives you time to test compatibility, resolve any issues, and plan the upgrade on your schedule. After the EOL date: * If the service is powered on, it's automatically upgraded to the latest version when possible, or to another supported version. note If it's not possible to upgrade a powered-on service to a supported version, the service is powered off and ultimately deleted. * If the service is powered off, it's upgraded and starts running the new version when powered on. ## Version EOL dates[​](#version-eol-dates "Direct link to Version EOL dates") Aiven for OpenSearch® is the open source continuation of the original Elasticsearch service. The EOL for Aiven for OpenSearch® is generally dependent on the upstream project. Some major versions are designated long-term support (LTS) releases, marked as such in the following table. | Version | Aiven EOL | After EOL | Service creation supported until | Service creation supported from | | ---------- | ------------ | ---------------------------------------- | -------------------------------- | ------------------------------- | | 1.3.x | 2026-07-26 | Automatic upgrade to 2.19 | 2026-07-26 | 2022-05-19 | | 2.17.x | 2026-07-26 | Automatic upgrade to 2.19 | 2026-07-26 | 2024-10-15 | | 2.19.x LTS | Date not set | Automatic upgrade to a supported version | Date not set | 2025-09-15 | | 3.3.x | 2027-02-01 | Automatic upgrade to a supported version | 2027-02-01 | 2026-01-20 | | 3.6.x LTS | Date not set | Automatic upgrade to a supported version | Date not set | 2026-06-23 | Related pages * [Upgrade Aiven for OpenSearch®](/docs/products/opensearch/howto/os-version-upgrade.md) --- # Scaling and performance in Aiven for OpenSearch® Scale resources and tune performance to keep your Aiven for OpenSearch® service running efficiently as it grows. Related pages * [Change the service plan](/docs/products/opensearch/howto/change-service-plan.md) * [Scale disk storage](/docs/products/opensearch/howto/scale-disk-storage.md) * [Prepare for high load](/docs/products/opensearch/howto/prepare-for-high-load.md) --- # OpenSearch® Dashboards incompatible version issues OpenSearch® Dashboards version must match your OpenSearch cluster version. If OpenSearch Dashboards is unavailable, you see the following error message in the OpenSearch Dashboards logs: ``` This version of OpenSearch Dashboards (v1.3.2) is incompatible with the following OpenSearch nodes in your cluster: v1.2.4 @ opensearch-searchdex-1.aiven.local/[${ipaddress}]:9200 (${ip_address}) ``` This situation typically occurs when a new version of OpenSearch is released. When a new version of OpenSearch is available, any newly created nodes for the service will automatically use the latest version. The Aiven platform will identify any variance in versions among the service nodes and initiate the process of upgrading outdated ones to the new version. Following the successful completion of this upgrade, OpenSearch Dashboards will be accessible. --- # Aiven for PostgreSQL® Aiven for PostgreSQL® is a fully managed and hosted relational database service. It's a high-performance data warehouse that offers maximum flexibility and functionality with a variety of advanced extensions out of the box. PostgreSQL® is an open source database ideal for organisations that need a well-organised tabular datastore. On top of the strict table and columns formats, PostgreSQL also offers solutions for nested datasets with the native `jsonb` format and advanced set of extensions including [PostGIS](https://postgis.net/), a spatial database extender for location queries. Aiven for PostgreSQL is the perfect fit for your relational data. A scalable SQL database solution that can be up and running within a few minutes. ## [Get started](/docs/products/postgresql/get-started.md) [4 items](/docs/products/postgresql/get-started.md) ## [Connect to service](/docs/products/postgresql/howto/list-code-samples.md) [3 items](/docs/products/postgresql/howto/list-code-samples.md) ## [Query and analyze data](/docs/products/postgresql/howto/pg-studio.md) [6 items](/docs/products/postgresql/howto/pg-studio.md) ## [Service management](/docs/products/postgresql/howto/power-cycle-service.md) [7 items](/docs/products/postgresql/howto/power-cycle-service.md) ## [Database management](/docs/products/postgresql/database-management.md) [5 items](/docs/products/postgresql/database-management.md) ## [Scaling and performance](/docs/products/postgresql/scaling-performance.md) [8 items](/docs/products/postgresql/scaling-performance.md) ## [Maintenance and lifecycle](/docs/products/postgresql/maintenance-lifecycle.md) [4 items](/docs/products/postgresql/maintenance-lifecycle.md) ## [High availability and disaster recovery](/docs/products/postgresql/concepts/high-availability.md) [6 items](/docs/products/postgresql/concepts/high-availability.md) ## [Backups and migration](/docs/products/postgresql/backups-migration.md) [6 items](/docs/products/postgresql/backups-migration.md) ## [Observability and monitoring](/docs/products/postgresql/reference/pg-metrics.md) [9 items](/docs/products/postgresql/reference/pg-metrics.md) ## [Integrations and extensions](/docs/products/postgresql/reference/list-of-extensions.md) [5 items](/docs/products/postgresql/reference/list-of-extensions.md) ## [User and schema](/docs/products/postgresql/concepts/dba-tasks-pg.md) [3 items](/docs/products/postgresql/concepts/dba-tasks-pg.md) Related pages * Learn more about Aiven for PostgreSQL on [Aiven.io](https://aiven.io/postgresql). * Find a comprehensive per-release documentation on the [PostgreSQL project website](https://www.postgresql.org/). * Read about location queries and the spatial extension on the [PostGIS website](https://postgis.net/). --- # Backups and migration in Aiven for PostgreSQL® Back up, restore, and migrate your Aiven for PostgreSQL® service data. Related pages * [Aiven for PostgreSQL® backups](/docs/products/postgresql/concepts/pg-backups.md) * [Create manual backups](/docs/products/postgresql/howto/create-manual-backups.md) * [Restore from a backup](/docs/products/postgresql/howto/restore-backup.md) --- # Prepare for migrating PostgreSQL® to Aiven using aiven-db-migrate The `aiven-db-migrate` tool is the recommended approach for migrating your PostgreSQL® to Aiven. It supports both logical replication and also using a dump and restore process. Find out more about the tool [on GitHub](https://github.com/aiven/aiven-db-migrate). Logical replication is the default method and once successfully set up, this keeps the two databases synchronized until the replication is interrupted. If the preconditions for logical replication are not met for a database, the migration falls back to using `pg_dump`. Regardless of the migration method used, the migration tool first performs a schema dump and migration to ensure schema compatibility. note Logical replication also works when migrating from AWS RDS PostgreSQL® 10+ and [Google CloudSQL PostgreSQL](https://cloud.google.com/sql/docs/release-notes#August_30_2021). ## Migration requirements[​](#aiven-db-migrate-migration-requirements "Direct link to Migration requirements") The following are the two basic requirements for a migration: 1. The source server is publicly available or there is a virtual private cloud (VPC) peering connection between the private networks 2. A user account with access to the destination cluster from an external IP, as configured in `pg_hba.conf` on the source cluster is present. Additionally to perform a **logical replication**, the following need to be valid: 1. PostgreSQL® version 10 or newer 2. Credentials with superuser access to the source cluster or the `aiven-extras` extension installed (see also: [Aiven Extras on GitHub](https://github.com/aiven/aiven-extras)) note The `aiven_extras` extension allows you to perform publish/subscribe-style logical replication without a superuser account, and it is preinstalled on Aiven for PostgreSQL servers. * An available replication slot on the destination cluster for each database migrated from the source cluster. * `wal_level` setting on the source cluster to `logical`. ## Migration pre-checks[​](#migration-pre-checks "Direct link to Migration pre-checks") The `aiven-db-migrate` migration tool checks the following requirements before it starts the actual migration: 1. A connection can be established to both source and target servers 2. The source and target server are not the same 3. A connection can be established to all source databases, ignoring `template0`, `template1`, and admin databases 4. There is enough disk space on the target for 130% of the total size of the source databases (ignoring `template0`, `template1`, and admin databases) 5. There are at least as many free logical replication slots as there are databases to migrate 6. A connection can be established to any target database that already exists on both source and target servers 7. The target version is the same as or newer than the source version 8. All languages installed in the source are available in the target 9. All extensions in the source are installed or available in the target 10. Source extensions not installed in the target can be installed: * by connecting as a superuser * without requiring superuser access * by inclusion in the allowed extensions list * by being a trusted extension (for PostgreSQL® 13 and newer) 11. In addition, for using the logical replication method: * The source version is PostgreSQL 10 or newer * The source `wal_level` is set to `logical` * The user connecting to the source has superuser access or the `aiven_extra` extension is or can be installed on each database ## Next steps[​](#next-steps "Direct link to Next steps") There's a detailed guide for performing the migration: [Migrate to Aiven for PostgreSQL® with aiven-db-migrate](/docs/products/postgresql/howto/migrate-aiven-db-migrate.md) --- # Perform DBA-type tasks in Aiven for PostgreSQL® Aiven doesn't allow superuser access to Aiven for PostgreSQL® services. However, most DBA-type actions are still available through other methods. ## `avnadmin` user privileges[​](#avnadmin-user-privileges "Direct link to avnadmin-user-privileges") By default, in every PostgreSQL instance, an `avnadmin` database user is created, with permissions to perform most of the usual DB management operations. It can manage: * Databases (`CREATE DATABASE`, `DROP DATABASE`) * Database users (`CREATE USER/ROLE`, `DROP USER/ROLE`) * Extensions (`CREATE EXTENSION`), you can also view the [list of available extensions](/docs/products/postgresql/reference/list-of-extensions.md) * Access permissions (`GRANT`, `REVOKE`) * Logical replication with the `REPLICATION` privilege tip You can also manage databases and users in the Aiven Console or though our [REST API](/docs/tools/api.md). ## `aiven_extras` extension[​](#aiven_extras_extension "Direct link to aiven_extras_extension") The `aiven_extras` extension, developed and maintained by Aiven, enables the `avnadmin` to perform superuser-like functionalities like: * Manage [subscriptions](https://www.postgresql.org/docs/current/catalog-pg-subscription.html) * Manage `auto_explain` [functionality](https://www.postgresql.org/docs/current/auto-explain.html) * Manage [publications](https://www.postgresql.org/docs/current/sql-createpublication.html) * [Claim public schema ownership](/docs/products/postgresql/howto/claim-public-schema-ownership.md) You can install the `aiven_extras` extension executing the following command with the `avnadmin` user: ``` CREATE EXTENSION aiven_extras CASCADE; ``` For more information about `aiven_extras` check the [GitHub repository](https://github.com/aiven/aiven-extras) for the project. --- # High availability of Aiven for PostgreSQL® Aiven for PostgreSQL® is available on a variety of plans, offering different levels of service availability. The selected plan defines the features available. | Plan | Service availability features | Backup history | | ------------ | -------------------------------------------------------------------------- | -------------- | | **Hobbyist** | Single node (limited availability) | 2 days | | **Startup** | Single node (limited availability) | 2 days | | **Business** | 1 primary node and 1 standby node (higher availability) | 14 days | | **Premium** | 1 primary node and 2 standby nodes (top high availability characteristics) | 30 days | ## Primary and standby nodes[​](#primary-and-standby-nodes "Direct link to Primary and standby nodes") Aiven's Business and Premium plans offer a [primary node](/docs/products/postgresql/reference/terminology.md) and [standby nodes](/docs/products/postgresql/reference/terminology.md). A standby node is useful for multiple reasons: * Provides another physical copy of the data in case of hardware, software, or network failures * Typically reduces the data loss window in disaster scenarios * Provides a quicker database time to restore with a controlled failover in case of failures as the standby is already installed, running, and synchronised with the data * Can be used for read-only queries to reduce the load on the primary server ## Failure handling[​](#failure-handling "Direct link to Failure handling") ### Minor failures[​](#minor-failures "Direct link to Minor failures") Minor failures, such as service process crashes or temporary loss of network access, are handled automatically by Aiven in all plans without any major changes to the service deployment. The service automatically restores normal operation once the crashed process is automatically restarted or when network access is restored. ### Severe failures[​](#severe-failures "Direct link to Severe failures") Severe failures, such as losing a node entirely in case of hardware or severe software problems, require radical recovery measures. The Aiven monitoring infrastructure automatically detects a failing node when the node starts reporting issues in the self-diagnostics or when it stops communicating. In such cases, the monitoring infrastructure automatically schedules a new replacement node to be created. note In the event of database failover, the **Service URI** of your service remains the same; only the IP address changes to point to the new primary node. ## Highly available Business and Premium service plans[​](#highly-available-business-and-premium-service-plans "Direct link to Highly available Business and Premium service plans") When a standby node fails, the primary node keeps running normally and provides a normal service level to the client applications. Once the new replacement standby node is ready and synchronised with the primary node, it starts replicating the primary node in real time as the situation gets back to normal. When a primary node fails, the combined information from the Aiven monitoring infrastructure and the standby node is used to make a failover decision. The standby node is promoted as the new primary and immediately starts serving clients. A new replacement node is automatically scheduled and becomes the new standby node. If the primary node and all the standby nodes fail at the same time, new nodes are automatically scheduled for creation to become the new primary and standby. The primary node is restored from the latest available backup, which can involve some degree of data loss. Any write operations made since the backup of the latest [WAL](/docs/products/postgresql/reference/terminology.md) file are lost. Typically, this time window is limited to either five minutes or one [WAL](/docs/products/postgresql/reference/terminology.md) file. note The amount of time required to replace a failed node depends mainly on the selected cloud region and the amount of data to be restored. In the case of partial loss of the cluster, the surviving node keeps on serving clients even during the recreation of the other node. This happens automatically and requires no administrator intervention. **Premium** plans operate in a similar way as **Business** plans. The main difference comes when one of the standby nodes or the primary node fails. Premium plans have an additional, redundant standby node available, providing platform availability even in the event of losing two nodes. If the primary node fails, the Aiven monitoring infrastructure, using [PGLookout](/docs/products/postgresql/reference/terminology.md), determines which of the standby nodes is the furthest along in replication (has the least potential for data loss) and does a controlled failover to that node. note For backups and restoration, Aiven utilises the popular Open Source backup daemon [PGHoard](/docs/products/postgresql/reference/terminology.md), which Aiven maintains. It makes real-time copies of [WAL](/docs/products/postgresql/reference/terminology.md) files to an object store in compressed and encrypted format. ## Single-node Hobbyist and Startup service plans[​](#single-node-hobbyist-and-startup-service-plans "Direct link to Single-node Hobbyist and Startup service plans") Hobbyist and Startup plans provide a single node. When it's lost, Aiven immediately starts the automatic process of creating a new replacement node. The new node starts up, restores its state from the latest available backup, and resumes serving clients. Since there is just a single node providing the service, the service is unavailable for the duration of the restoration. In addition, any write operations made since the backup of the latest [WAL](/docs/products/postgresql/reference/terminology.md) file are lost. Typically, this time window is limited to either five minutes or one [WAL](/docs/products/postgresql/reference/terminology.md) file. For more information, go to [Upgrade and failover procedures](/docs/products/postgresql/concepts/upgrade-failover.md). --- # Aiven for PostgreSQL® audit logging The path to optimal data security, compliance, incident management, and system performance starts with [collecting robust audit logs](/docs/products/postgresql/howto/use-pg-audit-logging.md). ## About audit logging[​](#about-audit-logging "Direct link to About audit logging") The audit logging feature allows you to monitor and track activities within relational database systems, such as Aiven for PostgreSQL®. Learn about multiple applications of this feature in [Why use audit logging](#why-use-pgaudit). ## Why use audit logging[​](#why-use-pgaudit "Direct link to Why use audit logging") Data Security * Monitor user activities to identify unusual or suspicious behavior * Detect unauthorized access attempts to critical data or systems * Identify intrusion attempts or unauthorized activities within the organization's IT environment Compliance * Use audit logs as regulatory compliance evidence to demonstrate that the organization meets industry or state regulations during audits * Track access to sensitive data to comply with data privacy regulations Accountability * Have specific actions attributed to individual users to hold them accountable for their activities within the system * Track changes to databases and systems to hold users accountable for alterations or configurations Operational security * Proactively identify and resolve security incidents * Detect and respond to potential security threats Incident management and root cause analysis * Investigate an incident with a detailed trail of events leading up to it * Analyze the root cause of an incident with audit logs providing data on actions and events that may have led to the incident System performance optimization * Monitor and analyze system performance to identify bottlenecks * Analyzing audit logs to assess resource utilization patterns and optimize the system configuration Data recovery and disaster planning * Use audit logs for data restoration in case of data loss or system failure * Analyze audit logs to improve system resilience and disaster planning strategies by identifying potential points of failure Change management and version control * Use audit logs to keep a record of changes made to databases, software, and configurations, ensuring proper version control ## Use cases[​](#use-cases "Direct link to Use cases") The audit logging feature has application in the following industries: * Finance and banking Ensuring compliance with regulatory requirements, tracking financial transactions, and detecting fraudulent activities * Healthcare Maintaining the confidentiality and integrity of patient records as well as complying with privacy regulations * Government and public sector Tracking changes in critical systems, securing sensitive data, and meeting legal and regulatory requirements * Information technology (IT) and software companies Monitoring access to the systems, tracking software changes, and identifying potential security breaches * Retail and e-commerce Tracking customer data, transactions, and inventory management to ensure data integrity and prevent unauthorized access * Manufacturing Tracking changes to production processes, monitoring equipment performance, and maintaining data integrity for quality control * Education Protecting sensitive student data, tracking changes to academic records, and monitoring system access for security purposes ## Limitations[​](#limitations "Direct link to Limitations") Aiven for PostgreSQL® audit logging requires the following: * Aiven for PostgreSQL version 11 or later * `avnadmin` superuser role * [psql](https://www.postgresql.org/docs/current/app-psql.html) for advanced configuration ## How it works[​](#how-it-works "Direct link to How it works") ### Activation with predefined settings[​](#activation-with-predefined-settings "Direct link to Activation with predefined settings") To use the audit logging on your service (database) for collecting logs in Aiven for PostgreSQL, [enable and configure this feature](/docs/products/postgresql/howto/use-pg-audit-logging.md) using the [Aiven Console](https://console.aiven.io), the [Aiven CLI](/docs/tools/cli.md), or [psql](https://www.postgresql.org/docs/current/app-psql.html). ### Configuration options[​](#configuration-options "Direct link to Configuration options") When enabled on your service, the audit logging can be [configured](/docs/products/postgresql/howto/use-pg-audit-logging.md) to match your use case. Audit logging parameters for fine-tuning the feature are the following: * `pgaudit.log` (default: none) Classes of statements to be logged by the session audit logging * `pgaudit.log_catalog` (default: on) Whether the session audit logging should be enabled for a statement with all relations in `pg_catalog` * `pgaudit.log_client` Whether log messages should be visible to a client process, such as `psql` * `pgaudit.log_level` Log level that should be used for log entries * `pgaudit.log_parameter` (default: off) Whether audit logs should include the parameters passed with the statement * `pgaudit.log_parameter_max_size` Maximum size (in bytes) of a parameter's value that can be logged * `pgaudit.log_relation` (default: off) Whether a separate log entry for each relation (for example, TABLE or VIEW) referenced in a SELECT or DML statement should be created * `pgaudit.log_rows` Whether the audit logging should include the rows retrieved or affected by a statement with the rows field located after the parameter field * `pgaudit.log_statement` (default: on) Whether the audit logging should include the statement text and parameters * `pgaudit.log_statement_once` (default: off) Whether the audit logging should include the statement text and parameters in the first log entry for a statement/sub-statement combination as opposed to including them in all the entries * `pgaudit.role` Master role to use for an object audit logging Full list of audit logging parameters For information on all the configuration parameters, preview [Settings](https://github.com/pgaudit/pgaudit/tree/6afeae52d8e4569235bf6088e983d95ec26f13b7#readme). ### Collecting and visualizing logs[​](#collecting-and-visualizing-logs "Direct link to Collecting and visualizing logs") [Access and visualize collected audit logs](/docs/products/postgresql/howto/use-pg-audit-logging.md) using one of the supported methods: | Integration method | Accessing `pgaudit` logs | Visualizing `pgaudit` logs | | ---------------------------------- | ----------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | | Aiven for OpenSearch integration | [Aiven for OpenSearch®](/docs/products/opensearch.md) | [OpenSearch Dashboards](/docs/products/opensearch/dashboards.md) | | Aiven for Apache Kafka integration | Kafka topic in Aiven for Apache Kafka | Requires a separate downstream tool to consume the log data from Aiven for Kafka and provide visualization | | External syslog integration | `rsyslog` | Third-party platforms: Datadog, Google Cloud Logging, Amazon CloudWatch Logs, and other syslog-compatible tools | ### Disabling audit logging[​](#disabling-audit-logging "Direct link to Disabling audit logging") To [disable the audit logging on your service (database)](/docs/products/postgresql/howto/use-pg-audit-logging.md), modify your service's advanced configuration with the [Aiven Console](https://console.aiven.io), the [Aiven CLI](/docs/tools/cli.md), or [psql](https://www.postgresql.org/docs/current/app-psql.html). ## What's next[​](#whats-next "Direct link to What's next") [Set up the audit logging on your Aiven for PostgreSQL service and start collecting audit logs](/docs/products/postgresql/howto/use-pg-audit-logging.md). --- # Aiven for PostgreSQL® backups Aiven for PostgreSQL® databases are automatically backed up, with **full backups** made daily, and **write-ahead logs (WAL)** copied at 5 minute intervals, or for every new file generated. All backups are encrypted using [`pghoard`](https://github.com/aiven/pghoard), an open source tool developed and maintained by Aiven. The time of day when the daily backups are made is initially randomly selected, but can be customised by setting the `backup_hour` and `backup_minute` advanced parameters, see [Advanced parameters for Aiven for PostgreSQL®](/docs/products/postgresql/reference/advanced-params.md). note The size of backups and the Aiven backup size shown on the Aiven web console differ, in some cases significantly. The backup sizes shown in the Aiven Console are for daily backups, before encryption and compression. ## Backup retention time by plan[​](#backup-retention-time-by-plan "Direct link to Backup retention time by plan") The number of stored backups and the backup retention time depends on the service plan that you have selected. | Plan Type | Backup Retention Time | | --------- | --------------------- | | Hobbyist | None | | Startup | 2 days | | Business | 14 days | | Premium | 30 days | ## Differences between logical and full backups[​](#differences-between-logical-and-full-backups "Direct link to Differences between logical and full backups") **Full backups** are version-specific binary backups that, when combined with WAL, allow consistent recovery to a point in time (PITR). **Logical backups** contain the SQL statements used to create the database schema and fill it with data, and are not tied to a specific version. | Feature | Full backup | Logical backup | | ------------------ | ------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- | | Version-specific | Full backups can be restored with the same PostgreSQL version as the backup was created from | Cross-version compatibility | | Database selection | The entire PostgreSQL instance is backed up. No option to backup (or restore) as single database | Single database objects can be backed up. Easy to backup and restore any item down to a single table | | Data | Contains uncommitted transactions and deleted and updated rows that have not been cleaned up by the PostgreSQL `VACUUM` process | Contains only the current committed content of the tables | | Indexes | Contains all data from indexes | Contains the queries needed to recreate indexes | | PITR capabilities | Previous database status can be restored using the WAL | Only "as-of-backup" status can be restored | | Restore time | Almost instantaneous, restore backup and replay delta WAL files | Long restoration process, replay of all SQL statements is needed to generate schema object and insert data | ## Delta base backups[​](#delta-base-backups "Direct link to Delta base backups") Aiven for PostgreSQL uses delta base backups, which allows to store data files that have been changed since the last backup and leave out the unchanged files. It's particularly beneficial for databases that include considerable portions of static data. Compared to regular base backups, delta base backups are more efficient, bringing improved performance and speeding up backup operations (unless all the data is updated constantly). Since delta base backups don't take all the data files, they are faster and easier to perform on large databases with huge volumes of data. Because performing a delta base backup doesn't last long, Aiven can back up data more frequently if required and applicable to specific datasets. With the increased backup frequency, service restoration and node replacement potentially can be faster for highly updated services because fewer WAL files need to be restored since the last backup (WAL restoration in PostgreSQL is single-threaded and, therefore, slow). ## Configure the backup schedule[​](#configure-the-backup-schedule "Direct link to Configure the backup schedule") Set the time of day when the daily backup is taken. To edit the backup schedule for your service: * Console * Aiven API * Aiven CLI * Terraform 1. In your service, in the **Backups** section, click **Backup management**. 2. Click **Actions** > **Configure backup settings**. 3. Click **Add configuration options**. 4. Add `backup_hour` and `backup_minute`, and set their values. 5. Click **Save configuration**. Call the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint, and add the following properties to the `user_config` object: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "backup_hour": BACKUP_HOUR, "backup_minute": BACKUP_MINUTE } }' ``` Replace the following: * `SERVICE_NAME`: the name of your service. * `PROJECT_NAME`: the name of your project. * `BACKUP_HOUR`: the hour when the service backup starts. Accepted values are integers between `0` and `23`. * `BACKUP_MINUTE`: the minute when the service backup starts. Accepted values are integers between `0` and `59`. Run the [avn service update](/docs/tools/cli/service-cli.md#avn-cli-service-update) command, and add the following properties to the `user_config` object: ``` avn service update SERVICE_NAME \ --project PROJECT_NAME \ --user-config '{ "backup_hour": BACKUP_HOUR, "backup_minute": BACKUP_MINUTE }' ``` Replace the following: * `SERVICE_NAME`: the name of your service. * `PROJECT_NAME`: the name of your project. * `BACKUP_HOUR`: the hour when the service backup starts. Accepted values are integers between `0` and `23`. * `BACKUP_MINUTE`: the minute when the service backup starts. Accepted values are integers between `0` and `59`. Use the `backup_hour` and `backup_minute` attributes in [your service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs) to set the start time for backups. If a backup was recently made, it can take another backup cycle before the new backup time takes effect. Related pages * [Create manual backups](/docs/products/postgresql/howto/create-manual-backups.md) * [Restore from a backup](/docs/products/postgresql/howto/restore-backup.md) * [Advanced parameters for Aiven for PostgreSQL®](/docs/products/postgresql/reference/advanced-params.md) --- # Aiven for PostgreSQL® connection pooling with PgBouncer Connection pooling in Aiven for PostgreSQL® services allows you to maintain very large numbers of connections to a database while minimizing the consumption of server resources. note Connection pooling requires a startup plan or higher. Verify your password encryption method If you use PGBouncer connection pooling, [verify your password encryption method compatibility](/docs/products/postgresql/troubleshooting/pg-password-encryption-upgrade.md) to ensure successful connections. You may need to migrate to `SCRAM-SHA-256` to maintain compatibility as the MD5 password encryption will be deprecated in PostgreSQL 19. ## About connection pooling[​](#about-connection-pooling "Direct link to About connection pooling") Aiven for PostgreSQL connection pooling uses [PgBouncer](https://www.pgbouncer.org/) to manage the database connection. Unlike when you connect directly to the PostgreSQL® server, each client connection does not require a separate backend process on the server. PgBouncer automatically inserts the client queries and only uses a limited number of actual backend connections, leading to lower resource usage on the server and better total performance. ## Maximum number of client connections[​](#maximum-number-of-client-connections "Direct link to Maximum number of client connections") How many client connections your service can handle depends on the RAM size that your service plan supports. * Each gigabyte of RAM allows 500 connections. * Minimum number of client connections per service is 5000. * Maximum number of client connections per service is 50000. ### Calculate `max_client_connections`[​](#calculate-max_client_connections "Direct link to calculate-max_client_connections") Use the following formula to calculate how many client connections your service can handle: min(max(n∗500,5000),50000) Where: * `n` is the number of RAM GB that a service plan supports. * n∗500 is `intermedia_max_connections`. * 5000≤intermedia\_max\_connections≤50000 * If `intermedia_max_connections` is less than 5000, lower bound 5000 applies. * If `intermedia_max_connections` is greater than 50000, upper bound 50000 applies. ### Examples[​](#examples "Direct link to Examples") * Startup-4 service plan (4 GB RAM) n=4 n∗500=2000 min(max(2000,5000),50000)=5000 For a startup-4 machine, `pgbouncer_max_client_connections` is 5000. * Business-16 service plan (16 GB RAM) n=16 n∗500=8000 min(max(8000,5000),50000)=8000 For a business-16 machine, `pgbouncer_max_client_connections` is 8000. * Business-120 service plan (120 GB RAM) n=120 n∗500=60000 min(max(60000,5000),50000)=50000 For a business-120 machine, `pgbouncer_max_client_connections` is 50000. ## Connection pooling benefits[​](#connection-pooling-benefits "Direct link to Connection pooling benefits") A high number of backend connections can become a problem with PostgreSQL, as the resource cost per connection is quite high due to how PostgreSQL manages client connections. PostgreSQL creates a separate backend process for each connection, and the unnecessary memory usage caused by the processes will start affecting the total throughput of the system at some point. Moreover, if each connection is very active, the performance can be affected by the high number of parallel executing tasks. It makes sense to have enough connections so that each CPU core on the server has something to do (each connection can only utilise a single CPU core), but a hundred connections per CPU core may be too much. All this is workload-specific, but often a good number of connections to have is roughly 3-5 times the CPU core count. Aiven enforces [connection limits](/docs/products/postgresql/reference/pg-connection-limits.md) to avoid overloading the PostgreSQL database. note Since 9.6, PostgreSQL offers parallelization support enabling to [run queries in parallel](https://www.postgresql.org/docs/current/parallel-query.html) on multiple CPU cores. Without a connection pooler, the database connections are handled directly by PostgreSQL backend processes, with one process per connection: Adding a PgBouncer pooler that utilizes fewer backend connections frees up server resources for more important uses, such as disk caching: Instead of having dedicated connections per client, now PgBouncer manages the connections assignment optimising them based on client request and settings like the pooling modes. tip Many frameworks and libraries (ORMs, Django, Rails, etc.) support client-side pooling, which solves much the same problem. However, when there are many distributed applications or devices accessing the same database, a server-side solution is a better approach. ## Connection pooling modes[​](#pooling-modes "Direct link to Connection pooling modes") Aiven for PostgreSQL supports three different operational pool modes: `transaction`, `session` and `statement`. * The default and recommended setting option is `transaction` pooling mode allows each client connection to take their turn in using a backend connection for the duration of a single transaction. After the transaction is committed, the backend connection is returned back into the pool and the next waiting client connection gets to reuse the same connection immediately. In practice, this provides quick response times for queries as long as the typical execution times for transactions are not excessively long. This is the most commonly used PgBouncer mode and also the default pooling mode in Aiven for PostgreSQL. warning Several PostgreSQL features, described in the [official PgBouncer features page](https://www.pgbouncer.org/features), are known to be **broken** by the default transaction-based pooling and **must not be used by the application when in this mode**. You must carefully consider the design of the client applications connecting to PgBouncer, otherwise the application may not work as expected. note Named prepared statements are one of the features affected by transaction and statement pooling. To allow PgBouncer to track them in these modes, set the [`pgbouncer.max_prepared_statements`](/docs/products/postgresql/reference/advanced-params.md#pgbouncer_max_prepared_statements) advanced parameter to a non-zero value. It defaults to 100 and can be set up to 3000. * The `session` pooling mode means that once a client connection is granted access to a PostgreSQL server-side connection, it can hold it until the client disconnects from the pooler. After this, the server connection is added back onto the connection pooler's free connection list to wait for its next client connection. Client connections are accepted (at TCP level), but their queries only proceed once another client disconnects and frees up its backend connection back into the pool. This mode can be helpful in some cases for providing a wait queue for incoming connections while keeping the server memory usage low, but is of limited use under most common scenarios due to the slow recycling of the backend connections. * The `statement` operational pooling mode, similar to the `transaction` pool mode, except that instead of allowing a full transaction to run, it cycles the server-side connections after each and every database statement (`SELECT`, `INSERT`, `UPDATE`, `DELETE` statements, etc.). Transactions containing multiple SQL statements are not allowed in this mode. This mode is sometimes used, for example when running specialised sharding frontend proxies. --- # About PostgreSQL® disk usage When you create your first Aiven for PostgreSQL® service, you may see the disk usage gradually increasing even though you are not yet inserting or updating much data. ![Disk space usage graph showing % growing over an hour](/docs/assets/images/initial-disk-usage-3e1bffbb25f66966b396237283cb6de4.png) This is completely normal within the first 24 hours of operation for Aiven for PostgreSQL services because of our WAL (Write-Ahead Logging) archiving settings. To prevent loss of data due to node failure on services with only one node, we set the `archive_timeout` configuration to write WAL segments to disk at regular intervals. The WAL segments are then backed up to cloud storage. Even if your service is idle, each WAL segment occupies the same amount of space on disk, which is why the disk usage grows steadily over time. After the first 24 hours of service, the system begins archiving old WAL segments and deleting them from disk to free up space. From this point onward, new WAL segments no longer have such a high impact on disk usage as the service reaches a steady state for low-traffic services. You can read more about WAL archiving [in the PostgreSQL manual](https://www.postgresql.org/docs/current/runtime-config-wal.html#RUNTIME-CONFIG-WAL-ARCHIVING). ## High disk usage discrepancy[​](#high-disk-usage-discrepancy "Direct link to High disk usage discrepancy") There can be instances when you notice your disk space usage increasing contrary to the amount of data is being written to your database. This can be caused due to an inactive replication slot present in your database. You can list the replication slot subscriptions created on your service with: ``` SELECT * FROM aiven_extras.pg_list_all_subscriptions(); ``` Inactive replication slots can cause indefinite increase in WAL log size. To resolve the issue, you can perform one of the following options: * Have a client connect to the replication slot so that the WAL files would be rotated and purged * Manually remove the unused replication slots (see our "Remove unused replication setup" under [Setup logical replication slots](/docs/products/postgresql/howto/setup-logical-replication.md) documentation) --- # Aiven for PostgreSQL® free tier Use Aiven for PostgreSQL® for free. You don't need a credit card to sign up and you can use it indefinitely free of charge. ## Features and limitations[​](#features-and-limitations "Direct link to Features and limitations") Free PostgreSQL services include: * A single node * 1 CPU per virtual machine * 1 GB RAM * 1 GB disk storage * Monitoring for metrics and logs * Backups * Support for PostgreSQL extensions There are some limitations of the free tier: * Cannot create a service in a VPC * Cannot fork a service to a free plan * No static IPs * No integrations * No connection pooling * `max_connections` limit set to `20` * No support services * Only one service of each service type in your [organization](/docs/platform/concepts/orgs-units-projects.md) * Not covered under Aiven's 99.99% SLA Free services do not have any time limitations. However, Aiven reserves the right to: * Power off free services with no initial usage within the first few hours after the service is running. You can power them back on at any time. * Power off free services with no continuative activity on the service. A notification is sent before the service is powered off. You can power them back on at any time. * Shut down services if Aiven believes they violate the [acceptable use policy](https://aiven.io/terms). * Change the cloud provider, region, or configuration at any time. ## Upgrade or downgrade a free service[​](#upgrade-or-downgrade-a-free-service "Direct link to Upgrade or downgrade a free service") You can upgrade your free service to a paid plan at any time by adding a payment method to the project's billing group. To upgrade a free service: 1. Go to the service **Overview** page. 2. In the **Service plan usage** section, click **Upgrade plan**. The upgrade happens immediately; however, it can take up to 3 hours for Basic tier support to be available. You can also downgrade a paid plan to the free tier as long as: * The data you have in that trial or paid service fits into the smaller instance size. * The free tier is available in the same cloud as the paid plan. Related pages * [Get started with Aiven for PostgreSQL®](/docs/products/postgresql/get-started.md) * [Supported extensions](/docs/products/postgresql/reference/list-of-extensions.md) * [Service pricing](/docs/platform/concepts/service-pricing.md) --- # Aiven for PostgreSQL® shared buffers Use shared buffers to share memory over multiple sessions. Discover how to inspect the database cache performance and the query cache performance and learn how to put data into cache manually. There are two primary memory allocations in Aiven for PostgreSQL that drastically impact the performance of queries: `shared_buffers` (the amount of RAM used for shared memory buffers) and `work_mem` (the maximum amount of memory to be used by a query operation before writing to temporary disk files). ### Purpose of shared buffers[​](#purpose-of-shared-buffers "Direct link to Purpose of shared buffers") The `shared_buffers` parameter controls the amount of memory allocated to the database server for disk page caching. The primary purpose of `shared_buffers` is to share memory over multiple sessions that may want to access the same blocks concurrently. Managing the access using the memory helps avoid unnecessary locking. note Even with `shared_buffers`, Aiven for PostgreSQL relies on the filesystem cache for optimization so reading from the disk is still needed. ### Allocation and setup[​](#allocation-and-setup "Direct link to Allocation and setup") `shared_buffers` is allocated only once at startup. It's not allocated per-session or per-user. It is shared among all sessions and useable for each worker and, therefore, each query. To obtain a good performance of a database server with 1 GB or more of RAM, it is usually necessary to set this value to \~25% of the system memory. For systems with less than 1 GB of RAM, a smaller portion of RAM is preferable to allow for the operating system. Allocating a lot of memory to `shared_buffers` is not always optimal for your configuration since the remaining free memory is allocated to workers (queries) and the filesystem cache. There are workloads where large `shared_buffers` are effective, but the allocation of more than 40% is unlikely to work better than a smaller-amount allocation. The optimal setting for any service depends on the available RAM, the working data set, and the workload applied. To examine the current `shared_buffers` value, run the following query: ``` SHOW shared_buffers; ``` ## Tuning guidelines[​](#tuning-guidelines "Direct link to Tuning guidelines") Aiven for PostgreSQL actively tracks data access patterns, updates the `shared_buffers` with frequently accessed data, and ejects Least Recently Used (LRU) when necessary. With many applications, only a fraction of the entire data set is accessed regularly. This data set fraction can be referred to as a *working set* or a *frequent read set* and often follows the 80/20 rule: 80% of the reads is in 20% of the data. tip For optimal performance, a `shared_buffers` cache hit rate of 97-99% is ideal. See an overview of `shared_buffers` cache hit rates using `pg_statio_user_tables`: ``` SELECT * FROM pg_statio_user_tables; relid | schemaname | relname | heap_blks_read | heap_blks_hit | idx_blks_read | idx_blks_hit | toast_blks_read | toast_blks_hit | tidx_blks_read | tidx_blks_hit -------+------------+---------+----------------+---------------+---------------+--------------+-----------------+----------------+----------------+--------------- 16415 | public | records | 1042770 | 88157826 | 184280 | 40282404 | 0 | 0 | 0 | 0 (1 row) ``` Calculate the database cache hit rate with: ``` SELECT sum(heap_blks_read) as heap_read, sum(heap_blks_hit) as heap_hit, sum(heap_blks_hit) / (sum(heap_blks_hit) + sum(heap_blks_read)) as hit_ratio FROM pg_statio_user_tables; heap_read | heap_hit | ratio -----------+----------+------------------------ 6942770 | 88157826 | 0.9883098315 (1 row) ``` If the cache hit rate is significantly lower than 95%, this may be an indicator of several issues: * Insufficient data activity to generate accurate statistics (new database) * Current `shared_buffers_percentage` value too low * Size of the working set larger than the maximum available `shared_buffers_percentage` (60%) To achieve an optimal performance, the working set needs to fit in `shared buffers`. While the `shared_buffers_percentage` has a maximum value of `60%]{.title-ref}, exceeding a value of [40%` suggests more RAM is required. tip In many cases, the Aiven default value of 20% requires no further modification. ## Inspecting the database cache performance[​](#inspecting-the-database-cache-performance "Direct link to Inspecting the database cache performance") For a deeper examination into the contents of the `shared_buffers`, enable the `pg_buffercache` extension: ``` CREATE EXTENSION pg_buffercache; ``` Calculate how many blocks from tables (r), indexes (i), sequences (S), and other objects are currently cached using the following query: ``` SELECT c.relname, c.relkind , pg_size_pretty(count(*) * 8193) as buffered , round(100.0 * count(*) / ( SELECT setting FROM pg_settings WHERE name='shared_buffers')::integer,1) AS buffers_percent , round(100.0 * count(*) * 8192 / pg_relation_size(c.oid),1) AS percent_of_relation FROM pg_class c INNER JOIN pg_buffercache b ON b.relfilenode = c.relfilenode INNER JOIN pg_database d ON (b.reldatabase = d.oid AND d.datname = current_database()) WHERE c.oid >= 16384 AND pg_relation_size(c.oid) > 0 GROUP BY c.oid, c.relname ORDER BY 3 DESC LIMIT 10; relname | relkind | buffered | buffers_percent | percent_of_relation ---------+---------+----------+-----------------+--------------------- records | r | 781 MB | 99.7 | 27.2 ``` Relations with object IDs (`oid`) below `16384` are reserved system objects. ## Inspecting the query cache performance[​](#inspecting-the-query-cache-performance "Direct link to Inspecting the query cache performance") Queries can also be inspected for the cache hit performance using `EXPLAIN`: ``` EXPLAIN (ANALYZE, BUFFERS, VERBOSE) SELECT * from records; QUERY PLAN -------------------------------------------------------------------------------------------------------------------------------- Seq Scan on public.records (cost=0.00..480095.20 rows=11207220 width=77) (actual time=0.158..16863.051 rows=11600000 loops=1) Output: id, "timestamp", data Buffers: shared hit=92345 read=275678 dirtied=10938 Query Identifier: 2582883386000135492 Planning: Buffers: shared hit=30 dirtied=2 Planning Time: 1.081 ms Execution Time: 17467.342 ms (8 rows) ``` Using `hit / (hit + read)]{.title-ref} shows [\~25%` of this full table scan was in the `shared_buffers` ## Putting data into the cache manually[​](#putting-data-into-the-cache-manually "Direct link to Putting data into the cache manually") You may want to prewarm the `shared_buffers` in anticipation of a specific workload, such as a large analytical query set used for reporting. This can be accomplished using the `pg_prewarm` extension. ``` CREATE EXTENSION pg_prewarm; ``` Example Call the `pg_prewarm` function and pass the name of a desired table. ``` SELECT * FROM pg_prewarm('public.records'); pg_prewarm ------------ 368023 SELECT pg_size_pretty(pg_relation_size('public.records')); pg_size_pretty ---------------- 2875 MB ``` 368023 pages have been read into the cache (or \~2875 MB). If the `shared buffers` size is less than pre-loaded data, only the tailing end of the data is cached as the earlier data encounters a forced ejection. ## Service memory limits[​](#service-memory-limits "Direct link to Service memory limits") The practical memory limit will always be less than the service physical memory limit. **All services are subject to operating overhead:** * A small amount of memory is required by the operating system kernel to manage system resources, including networking functions and disk cache. * Aiven's cloud data platform requires memory to monitor availability, provide metrics, logging and manage backups. A server or node's **usable memory** can be calculated as: `usable memory = RAM - overhead` Where: * `overhead` is 350 MiB (≈ 0.34 GiB). Services may utilize optional components, service integrations, connection pooling, or plug-ins, which are not included in overhead calculations. If a service is overcommitted, the operating system, management layer, backups or availability monitoring, may fail status checks or operations due to resource contention. In severe instances, the node may fail completely with an out-of-memory condition. ## Out of memory conditions[​](#out-of-memory-conditions "Direct link to Out of memory conditions") Many processes request more memory from the kernel than they will ever use or need. In these cases, the kernel overallocates memory. This allows it to satisfy multiple processes requesting more memory than is available, which is not used or is freed by the time any other process actually needs it. However, if enough processes start using all their allocated memory simultaneously there may not be enough physical memory available and an `Out Of Memory` (`OOM`) condition occurs. warning This situation is critical and must be resolved immediately. The solution that the Linux kernel employs is to invoke the `Out of Memory Killer` (or `OOM Killer`). This reviews all running processes and kills one or more of them to free up system memory and keep the system running. The `OOM Killer` selects process to kill based on an `oom_score`; a calculation that balances how much memory the process is using with how long the process has been running. Processes that have been running for a long time are less likely to be killed. Subprocesses are summed with parent processes in terms of memory usage, so a process which forks many subprocesses, but itself does not use a lot of memory, may still be killed. In most instances, the hosted data service, or a child process, will have the highest memory footprint and be a prime candidate for termination when the OOM Killer inspects the running processes. Aiven's cloud data platform leverages kernel namespaces or containers to isolate processes from each other. Isolation has several benefits, including: * A smaller footprint for security‑related concerns * A smaller blast radius for failure * Greater control of system resources Left unchecked, the `OOM Killer` may opt to kill the primary service. This is undesirable as unclean termination of the primary service can lead to data loss, inconsistency, or corrupted backups. Further, if Aiven's management platform detects that the primary service is unavailable for , the service will be marked as down and a failover will occur. To mitigate this scenario, namespaces are used, some with additional memory limits, in combination with an `oom_score_adjust` on the primary process, to coax the `OOM Killer` into selection of less critical processes. This will still result in a service restart, but in a more controlled process, where the database is shut down, rather than killed; exposure to data loss is limited and recovery is faster when the service restarts, often avoiding failover. warning Out of Memory conditions can still lead to unexpected behavior, including data unavailable or data loss conditions. ## Avoid running low on memory[​](#avoid-running-low-on-memory "Direct link to Avoid running low on memory") The OOM killer only runs when the system is critically low on memory. To prevent it from running, either reduce your memory usage or increase the available memory. For most databases, the service memory footprint can often be reduced by: * Reducing concurrency or implementing connection pooling * Tuning queries to limit result sets * Tuning indexes for query load * Dropping unused objects from storage In cases where the working set no longer fits into memory, consider scaling your service. Related pages For more information on shared buffers, see [Resource Consumption](https://www.postgresql.org/docs/current/runtime-config-resource.html) in the PostgreSQL documentation. --- # pgvector for AI-powered search in Aiven for PostgreSQL® In machine learning (ML) models, all data items in a particular data set are mapped into one unified n-dimensional vector space, no matter how big the input data set is. This optimized way of data representation allows for high performance of AI algorithms. Mapping regular data into a vector space requires so called data vectorizing, which is transforming data items into vectors (data structures with at least two components: magnitude and direction). On the vectorized data, you can perform AI-powered operations using different instruments, one of them being pgvector. Discover the pgvector extension to Aiven for PostgreSQL® and learn how it works. Check why you might need it and what benefits you get using it. ## About pgvector[​](#about-pgvector "Direct link to About pgvector") pgvector is an open-source vector extension for similarity search. It's available as an extension to your Aiven for PostgreSQL® services. pgvector introduces capabilities to store and search over data of the vector type (ML-generated embeddings). Applying a specific index type for querying a table, the extension enables you to search for vector's exact nearest or approximate nearest neighbors (data items). ### Vector embeddings[​](#vector-embeddings "Direct link to Vector embeddings") In machine learning, real-world objects and concepts (text, images, video, or audio) are represented as a set of continuous numbers residing in a high-dimensional vector space. These numerical representations are called vector embeddings, and the process of transformation into numerical representations is called vector embedding. Vector embedding allows ML algorithms to identify semantic and syntactic relationships between data, find patterns, and make predictions. Vector representations have different applications, for example, information retrieval, image classification, sentiment analysis, natural language processing, or similarity search. ### Vector similarity[​](#vector-similarity "Direct link to Vector similarity") Since on vector embeddings you can use AI tools for capturing relationships between objects (vector representations), you are also able to identify similarities between them in a computable and scalable manner. A vector usually represents a data point, and components of the vector correspond to attributes of the data point. In most cases, vector similarity calculations use distance metrics, for example, by measuring the straight-line distance between two vectors or the cosine of the angle between two vectors. The greater the resulting value of the similarity calculation is, the more similar the vectors are, with 0 as the minimum value and 1 as the maximum value. ## How pgvector works[​](#how-pgvector-works "Direct link to How pgvector works") * Enabling pgvector: You enable the extension on your database. * Vectorizing data: You generate embeddings for your data, for example, for a products catalog using tools such as the [OpenAI API](https://platform.openai.com/docs/api-reference/embeddings/create) client. * Storing embeddings: You store the embeddings in Aiven for PostgreSQL using the pgvector extension. * Querying embeddings: You use the embeddings for the vector similarity search on the products catalog. * Adding indices: By default, pgvector executes the *exact* nearest neighbor search, which gives the perfect recall. If you add an index to use the *approximate* nearest neighbor search, you can speed up your search, trading off some recall for performance. ## Why use pgvector[​](#why-use-pgvector "Direct link to Why use pgvector") With the pgvector extension, you can perform the vector similarity search and use embedding techniques directly in Aiven for PostgreSQL. pgvector allows for efficient handling of high-dimensional vector data within the Aiven for PostgreSQL database for tasks such as similarity search, model training, data augmentation, or machine learning. pgvector helps you optimize and personalize the similarity search experience by improving searching speed and accuracy (also by adding indices). ## Typical use cases[​](#typical-use-cases "Direct link to Typical use cases") There are multiple industry applications for similarity searches over vector embeddings: * e-commerce * recommendation systems * fraud detection Examples * AI-powered tools can find similarities between products or transactions, which can be used to produce product recommendations or detect potential scams or frauds. * Sentiment analysis: words represented with similar vector embeddings have similar sentiment scores. Related pages * [Enable and use pgvector on Aiven for PostgreSQL®](/docs/products/postgresql/howto/use-pgvector.md) * [pgvector README on GitHub](https://github.com/pgvector/pgvector/blob/master/README.md) --- # Use TimescaleDB with Aiven for PostgreSQL® [TimescaleDB](https://github.com/timescale/timescaledb) is an open-source database designed to make your existing relational database scalable for time series data. TimescaleDB is available as a PostgreSQL® extension on Aiven. A time series indexes a series of data points in chronological order, usually as a sequence over regular intervals. Examples of a time series include: * the temperature of a home during a day * the position of a satellite during a day The data in these examples consists of a measured value (temperature or position) corresponding to the time at which the reading of the value took place. ## Enable TimescaleDB on Aiven for PostgreSQL[​](#enable-timescaledb-on-aiven-for-postgresql "Direct link to Enable TimescaleDB on Aiven for PostgreSQL") TimescaleDB is available as an extension; you can enable it by running: ``` CREATE EXTENSION timescaledb CASCADE; ``` After enabling the extension, you can create TimescaleDB hypertables and make use of its features for working with time-series data. For further information, have a look at the [Hypertables and chunks](https://www.tigerdata.com/docs/reference/timescaledb/hypertables) guide from Timescale. More information about [how to install and manage extensions](/docs/products/postgresql/howto/manage-extensions.md) is also available. ## TSL-licensed features[​](#tsl-licensed-features "Direct link to TSL-licensed features") The majority of TimescaleDB functionality is open source, however since TimescaleDB 1.2.0 some features have been restricted to the Timescale License, which explicitly prohibits them from being made available in any database-as-a-service offering. Therefore some features, such as `time_bucket_gapfill`, are not available on the Aiven hosted TimescaleDB service. When you try using these features, you will see an error similar to the following: `ERROR 0A000 (feature_not_supported) function "" is not supported under the current license "ApacheOnly"` Since Aiven only offers open source licensed platforms, these features cannot be made available. --- # Aiven for PostgreSQL® upgrade and failover procedures Aiven for PostgreSQL® Business and Premium plans include **standby read-replica** servers. If the primary server fails, a standby replica server is automatically promoted as new primary server. warning Standby read-replica servers available on PostgreSQL Business and Premium plans are substantially different from manually created read-replica services since the latter are not promoted if the primary server fails. There are two cases when failover or switchover occurs: 1. Uncontrolled primary/replica disconnection 2. Controlled switchover during rolling-forward upgrades warning For Hobbyist and Startup plans, due to missing standby read-replica servers, uncontrolled disconnections can only be mitigated by restoring data from a backup, and can result in data loss of the database changes since the latest backup data that was uploaded to object storage. ## Uncontrolled primary/replica disconnection[​](<#Failover PGUncontrolled> "Direct link to Uncontrolled primary/replica disconnection") When a server unexpectedly disconnects, there is no certain way to know whether it really disappeared or whether there is a temporary glitch in the cloud provider's network. Aiven's management platform has different procedures in case of primary or replica nodes disconnections. ### Primary server disconnection[​](#primary-server-disconnection "Direct link to Primary server disconnection") If the primary server disappears, Aiven's management platform uses an initial 60-second timeout before marking the server as down and promoting a replica server as new primary. During this 60-second timeout, the master is unavailable (`servicename-projectname.aivencloud.com` does not respond), and `replica-servicename-projectname.aivencloud.com` works fine (in read-only mode). After the replica promotion, `servicename-projectname.aivencloud.com` would point to the new primary server, while `replica-servicename-projectname.aivencloud.com` becomes unreachable. Finally, a new replica server is created, and after the synchronisation with the primary, the `replica-servicename-projectname.aivencloud.com` DNS is switched to point to the new replica server. Recovery time can take up to 180 seconds depending on client side settings and when in the DNS TTL cycle the disconnection occurs. ### Replica server disconnection[​](#replica-server-disconnection "Direct link to Replica server disconnection") If the replica server disappears, Aiven's management platform uses an initial 60-second timeout before marking the server as down and creating a new replica server. note Each Aiven for PostgreSQL® Business plan supports one replica server only, which is why the service's read replica endpoint `replica-SERVICE_NAME-PROJECT_NAME.aivencloud.com` remains unavailable and queries to this endpoint time-out until a new replica is available. tip For higher availability on a service's read replica endpoint, you can upgrade to a Premium plan with two standby servers used as read replicas. The DNS record pointing to primary server `SERVICE_NAME-PROJECT_NAME.aivencloud.com` remains unchanged during the recovery of the replica server. Recovery time can take up to 180 seconds depending on client side settings and when in the DNS TTL cycle the disconnection occurs. ## Controlled switchover during upgrades or migrations[​](#controlled-switchover-during-upgrades-or-migrations "Direct link to Controlled switchover during upgrades or migrations") note This section doesn't apply to major version upgrades with `pg_upgrade`. For more information, see [Perform a PostgreSQL® major version upgrade](/docs/products/postgresql/howto/upgrade.md). During maintenance updates, cloud migrations, or plan changes, the below procedure is followed: 1. For each of the **replica** nodes (available only on Business and Premium plans), a new server is created, and data restored from a backup. Then the new server starts following the existing primary server. After the new server is up and running and data up-to-date, `replica-servicename-projectname.aivencloud.com` DNS entry is changed to point to it, and the old replica server is deleted. 2. An additional server is created, and data restored from a backup. Then the new server is synced up to the old primary server. 3. Cluster replication is changed to **quorum commit synchronous** to avoid data loss when changing primary server. note At this stage, one extra server is running: the old primary server, and N+1 replica servers (2 for Business and 3 for Premium plans). 4. The old primary server is scheduled for termination, and one of the new replica servers is immediately promoted as a primary server. `servicename-projectname.aivencloud.com` DNS is updated to point to the new primary server. The new primary server is removed from the `replica-servicename-projectname.aivencloud.com` DNS record. note * The old primary server is kept alive for a short period of time (minimum 60 seconds) with a TCP forwarding setup pointing to the new primary server allowing clients to connect before learning the new IP address. * If the service plan is changed from a business plan that has two nodes to a startup plan which only has one node of the same tier (for example, business-8 to startup-8), the standby node is removed while the primary node is retained, and connections to the primary are not affected by the downgrade. Similarly, upgrading the service plan from a startup one to a business one adds a standby node to the service cluster, and connections to the primary node are unaffected. ## Recreation of replication slots[​](#recreation-of-replication-slots "Direct link to Recreation of replication slots") In case of failover or controlled switchover of an Aiven for PostgreSQL service, the replication slots from the old primary server are automatically recreated in the new primary server. note The recreation of replication slots feature is enabled automatically and doesn't require restarting the nodes for services that have been created or updated as of January 2023. Additional details are outlined in [our blog post](https://aiven.io/blog/aiven-for-pg-recreates-logical-replication-slots). important * Replication slots are not recreated during power-on or power-off events. * Replication slots are not recovered after major version upgrades of Aiven for PostgreSQL. * To prevent the loss of a replication slot when promoting a new primary server, all changes from the old primary must be fully consumed to ensure that its `catalog_xmin` doesn't predate the standby server. If the slot is not consumed within 30 minutes, the syncing process will still be marked as complete, and there is a risk that the logical slot will be lost when the new node is promoted. ### One-node cluster[​](#one-node-cluster "Direct link to One-node cluster") Before replacing a node in the one-node cluster, the new node acquires information on replication slots on the original service, re-creates them, and only then the failover is performed. ### Multi-node cluster[​](#multi-node-cluster "Direct link to Multi-node cluster") For multi-node setups, replication slots from the primary are synchronized to the standbys periodically. At regular time intervals * Dependencies for newly created slots are installed in the corresponding databases (currently, every 30 seconds). When the new slot is created on a database and we want to re-create this slot on a standby, we use a functionality from the `aiven_extras` extension, which needs to be installed in the database. Therefore, every 30 seconds there is a job checking that this extension is installed on the databases with logical replication slots. * Positions (`confirmed_flush_lsn`) of the slots are synchronized between the primary and the standbys. When a failover to a standby occurs, the standby node already has replication slots with an up-to-date (maximum 5-second delay) positions from the primary. warning Uncontrolled failover ramifications * Slots created up to 30 seconds before the failover might be lost. * If due to a cloud provider failure, a node from the one-node cluster disappears, replication slots on a new replacement node cannot be restored since the replication slots information is lost. note * Position of recovered replication slots might be up to several seconds older than on the original primary. Therefore, when re-connecting to PostgreSQL and reading from replication slots, it's recommended to use start positions known to the client until which the data was already received. Otherwise, the client might receive duplicate entries. * In case of failover with a huge lag between the primary node and the standby node (for example, when a master disappears), the position of the replication slot restored on a new master is not newer than the position on the standby node, even though the position of that slot on the old master was newer. --- # Cross-region disaster recovery in Aiven for PostgreSQL® ## [CRDR overview](/docs/products/postgresql/crdr/crdr-overview.md) [The cross-region disaster recovery (CRDR) feature ensures your business continuity by recovering your workloads to a remote region in the event of a region-wide](/docs/products/postgresql/crdr/crdr-overview.md) ## [CRDR setup](/docs/products/postgresql/crdr/enable-crdr.md) [Enable cross-region disaster recovery (CRDR) in Aiven for PostgreSQL® by creating a recovery service, which takes over from a primary service in case of a region outage.](/docs/products/postgresql/crdr/enable-crdr.md) ## [Failover and failback](/docs/products/postgresql/crdr/failover/list-failover.md) [2 items](/docs/products/postgresql/crdr/failover/list-failover.md) ## [Switchover and switchback](/docs/products/postgresql/crdr/switchover/list-switchover.md) [2 items](/docs/products/postgresql/crdr/switchover/list-switchover.md) ## Related pages[​](#related-pages "Direct link to Related pages") * [Backups in Aiven for PostgreSQL®](/docs/products/postgresql/concepts/pg-backups.md) * [Read-only replicas in Aiven for PostgreSQL®](/docs/products/postgresql/howto/create-read-replica.md) * [High availability in Aiven for PostgreSQL®](/docs/products/postgresql/concepts/high-availability.md) * [Upgrade and failover procedures in in Aiven for PostgreSQL®](/docs/products/postgresql/concepts/upgrade-failover.md) --- # Cross-region disaster recovery in Aiven for PostgreSQL® [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) The cross-region disaster recovery (CRDR) feature ensures your business continuity by recovering your workloads to a remote region in the event of a region-wide failure. ## Region-wide outage[​](#region-wide-outage "Direct link to Region-wide outage") CRDR allows you to cope with the primary region failure by initiating a recovery transition to another region. To identify a region outage, look into the region status: * Check your monitoring and alerts, and watch the following metrics: * Instances, nodes, services failures * Connectivity loss, latency spikes, packet drops * High error rates, timeouts, 5xx server errors * Check your cloud provider's status page: * [AWS](https://health.aws.amazon.com) * [Google Cloud](https://status.cloud.google.com) * [Azure](https://status.azure.com) * Test connectivity and DNS resolution for your instances or services. ## CRDR overview[​](#crdr-overview "Direct link to CRDR overview") The CRDR setup is a pair of integrated multi-node services, sharing credentials and a DNS name but located in different regions. CRDR peer services can be hosted on 1-3 nodes. * **Primary service** hosted in the primary region is your original service you use on regular basis. It hands over to the recovery service when you initiate [a failover or a switchover](/docs/products/postgresql/crdr/crdr-overview.md#recovery-transition). When you initiate [a failback or a switchback](/docs/products/postgresql/crdr/crdr-overview.md#recovery-reversion), the primary service takes back control from the recovery service as soon as the infrastructure is up and running again. * **Recovery service** hosted in the recovery region is the service you create for disaster recovery purposes. It takes over from the primary service when you initiate [a failover or a switchover](/docs/products/postgresql/crdr/crdr-overview.md#recovery-transition). When you initiate [a failback or a switchback](/docs/products/postgresql/crdr/crdr-overview.md#recovery-reversion), the recovery service hands over to the primary service as soon as the infrastructure is up and running again. The CRDR cycle is a sequence of actions involving CRDR peer services aimed at enabling and executing CRDR as well as resuming the original service operation. Throughout the CRDR cycle, CRDR peer services or service nodes go into the following states: * **Active**: A CRDR peer service is *active* when it runs on a node that is replicating data to CRDR standby nodes. * Primary service is active during normal operations, when a region is up and running. * Recovery service is active after taking over from primary service in the event of a region outage. * **Passive**: A CRDR peer service is *passive* when it runs on CRDR standby nodes only. Either CRDR peer service can be passive depending on a phase of the CRDR cycle. * **Failed**: A CRDR peer service is *failed* when it's defunct or unreachable after failing over in the event of a region outage. Only a primary service can be failed. ## Limitations[​](#limitations "Direct link to Limitations") * **Service plan requirements**: To set up CRDR, your primary service must use at least a Startup plan. Hobbyist, Free, and Developer plans are not supported. Upgrading your plan If your Aiven for PostgreSQL service uses a Hobbyist, Free, or Developer plan, upgrade to at least a Startup plan. * **Console restrictions**: When creating a recovery service through the [Aiven Console](https://console.aiven.io/), you must use the same service plan and cloud provider as your primary service. Alternative setup methods For different service plans or cloud providers, create your recovery service using the [Aiven CLI](/docs/tools/cli.md), the [Aiven API](/docs/tools/api.md), or the [Aiven Provider for Terraform](https://registry.terraform.io/providers/aiven/aiven/latest/docs). ## How it works[​](#how-it-works "Direct link to How it works") The CRDR feature is eligible for all startup, business, and premium service plans. ![Ready for CRDR](/docs/assets/images/ready-for-crdr-83268e8a684162680a7cc7d92719184c.png) ### CRDR setup[​](#crdr-setup "Direct link to CRDR setup") You [enable CRDR by creating a recovery service](/docs/products/postgresql/crdr/enable-crdr.md). The CRDR setup completes as soon as the recovery service is created and in sync with the primary service. At that point, the primary service is the **Active** service receiving incoming traffic and replicating to the recovery service, and the recovery service is the **Passive** service replicating from the primary service. ![CRDR setup](/docs/assets/images/crdr-setup-90715f42cce11c02eeac39095dc0e580.png) ### Recovery transition[​](#recovery-transition "Direct link to Recovery transition") CRDR supports two types of the recovery transition: * [Failover](/docs/products/postgresql/crdr/crdr-overview.md#failover-to-the-recovery-region) * **Triggered by you** typically in the event of a region-wide outage * **Destroys the primary service** and requires the primary service recreation to fail back. * [Switchover](/docs/products/postgresql/crdr/crdr-overview.md#switchover-to-the-recovery-region) * **Triggered by you** for any purposes other than a region-wide outage * Leaves the **primary service intact** with no need for recreating it to switch back. #### Failover to the recovery region[​](#failover-to-the-recovery-region "Direct link to Failover to the recovery region") You typically trigger a [failover to the recovery region](/docs/products/postgresql/crdr/failover/crdr-failover-to-recovery.md) in the event of a region-wide outage. This destroys the primary service, which becomes **Failed**, and promotes the recovery service to **Active**. To fail back to the primary service, it needs to be recreated first. ![CRDR failover](/docs/assets/images/crdr-failover-d3176dc257518b1181be9fd73858a7cc.png) #### Switchover to the recovery region[​](#switchover-to-the-recovery-region "Direct link to Switchover to the recovery region") You trigger a [switchover to the recovery service](/docs/products/postgresql/crdr/switchover/crdr-switchover.md) for testing, simulating a disaster scenario, or verifying the disaster resilience of your infrastructure. This demotes the primary service to **Passive** and promotes the recovery service to **Active**. To switch back to the primary service, no service recreation is needed. ![CRDR switchover](/docs/assets/images/crdr-switchover-342202c26afbcd5575923df705a08afc.png) ### Recovery reversion[​](#recovery-reversion "Direct link to Recovery reversion") You trigger a recovery reversion to shift your workload back to the primary region and restore the CRDR setup to its original configuration. There are two types of the recovery reversion: * [Failback](/docs/products/postgresql/crdr/crdr-overview.md#failback-to-the-primary-region) * Reverts a [failover](/docs/products/postgresql/crdr/crdr-overview.md#failover-to-the-recovery-region). * Recreates the primary service. * [Switchback](/docs/products/postgresql/crdr/crdr-overview.md#switchback-to-the-primary-region) * Reverts a [switchover](/docs/products/postgresql/crdr/crdr-overview.md#switchover-to-the-recovery-region). * No need to recreate the primary service. #### Failback to the primary region[​](#failback-to-the-primary-region "Direct link to Failback to the primary region") The failback process consists of two steps you initiate at your convenience: 1. [Primary service recreation](/docs/products/postgresql/crdr/failover/crdr-revert-to-primary.md) You initiate this step to restore primary service nodes from the local backups and to synchronize (replicate) the most recent data from the active service (recovery service). When completed, the primary service is restored and in near real-time sync with the recovery service. 2. [Primary service takeover](/docs/products/postgresql/crdr/failover/crdr-revert-to-primary.md) You initiate a takeover as soon as the primary service is recreated. This switches the direction of the replication to effectively route the traffic back to the primary region. When completed, both the primary service and the recovery service are up and running again: the primary service as an active service, and the recovery service as a passive service. ![CRDR revert](/docs/assets/images/crdr-revert-675a2f1d8d8f28dd8dfb14674f84e287.png) #### Switchback to the primary region[​](#switchback-to-the-primary-region "Direct link to Switchback to the primary region") You initiate a switchback at your convenience to switch the direction of the replication and route the traffic back to the primary region. When completed, both the primary service and the recovery service are up and running again: the primary service as an active service, and the recovery service as a passive service. ![CRDR switchback](/docs/assets/images/crdr-switchback-80f0e027c6d711a4e34db5139a796561.png) ## DNS name and service URI[​](#dns-name-and-service-uri "Direct link to DNS name and service URI") ### Active service DNS name[​](#active-service-dns-name "Direct link to Active service DNS name") CRDR allows you to access your active service always using the same **Service URI**, which doesn't change in the event of a failover to the recovery region. note **Service URI** is a locator that is shared between the primary service and the recovery service. It always points to the replicating node of the active service. This node is the only read-write node in both CRDR regions. The **Service URI** of an active service can remain unchanged in the event of a region outage because the DNS record of this **Service URI** is updated to point to the active service. This allows your applications to work uninterrupted and adapt to the change automatically without updating its code or data. ### Standby nodes DNS names[​](#standby-nodes-dns-names "Direct link to Standby nodes DNS names") Regardless of the CRDR cycle phase, you can always connect and access separately each standby node in the CRDR peer services. This can help you compensate for potential network delays by using the service geographically closer to your applications. Standby nodes in the CRDR service pair can have two different URIs, depending on the CRDR service (region) they belong to: * For the **primary service standby URI**, the DNS record always points to the standby nodes in the primary region. * For the **recovery service standby URI**, the DNS record always points to the standby nodes in the recovery region. Both the primary service standby URI and the recovery service standby URI are dedicated, not shared, and read-only. ## Backups in the recovery region[​](#backups-in-the-recovery-region "Direct link to Backups in the recovery region") After a failover to the recovery region in the event of a primary region outage, service backups start to be taken in the recovery region. You can use this backup history for operations and data resiliency purposes. Related pages * [Aiven for PostgreSQL high availability](/docs/products/postgresql/concepts/high-availability.md) * [Aiven for PostgreSQL backups](/docs/products/postgresql/concepts/pg-backups.md) * [Aiven for PostgreSQL read-only replica](/docs/products/postgresql/howto/create-read-replica.md) --- # Set up cross-region disaster recovery in Aiven for PostgreSQL® [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Enable [cross-region disaster recovery (CRDR)](/docs/products/postgresql/crdr/crdr-overview.md) in Aiven for PostgreSQL® by creating a recovery service, which takes over from a primary service in case of a region outage. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven for PostgreSQL service on at least a Startup plan - [see limitations](/docs/products/postgresql/crdr/crdr-overview.md#limitations) * One of the following tools for operating CRDR: * [Aiven Console](https://console.aiven.io/) - [see limitations](/docs/products/postgresql/crdr/crdr-overview.md#limitations) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](https://registry.terraform.io/providers/aiven/aiven/latest/docs) ## Set up a recovery service[​](#set-up-a-recovery-service "Direct link to Set up a recovery service") Create a [CRDR setup](/docs/products/postgresql/crdr/crdr-overview.md#crdr-setup) using a tool of your choice: * Console * CLI * API * Terraform 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your primary Aiven for PostgreSQL service. 2. In the **Backups** section, click **Disaster recovery**. 3. On the **Cross region disaster recovery** page, click **Create recovery service**. 4. In **Create recovery service** wizard: 1. Provide a service name. 2. Select a cloud region. 3. Click **Create recovery service**. Throughout the process of creating the recovery service, the recovery service is in the **Rebuilding** state. As soon as the recovery service is ready, its status changes to **Passive**, which means your CRDR setup is up and running. Run [avn service create](/docs/tools/cli/service-cli.md#avn-cli-service-create): ``` avn service create RECOVERY_SERVICE_NAME \ --service-type pg \ --plan SERVICE_PLAN \ --cloud CLOUD_PROVIDER_REGION \ --disaster-recovery-copy-for PRIMARY_SERVICE_NAME ``` Replace the following: * `RECOVERY_SERVICE_NAME` with the name of the recovery service, for example, `pg-demo-recovery` * `SERVICE_PLAN` with the plan to use for the recovery service, for example, `startup-4` * `CLOUD_PROVIDER_REGION` with the cloud region where to host the recovery service, for example, `google-europe-west-4` * `PRIMARY_SERVICE_NAME` with the name of the primary service, for example, `pg-demo` Call the [ServiceCreate endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) to create a recovery service and enable the `disaster_recovery` service integration between the recovery service and the primary service, for example: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service \ -H 'accept: application/json, text/plain, */*' \ -H 'Authorization: Bearer BEARER_TOKEN' \ -H 'content-type: application/json' \ --data-raw '{ "service_name": "RECOVERY_SERVICE_NAME", "cloud": "CLOUD_PROVIDER_REGION", "plan": "SERVICE_PLAN", "service_type": "SERVICE_TYPE", "disk_space_mb": DISK_SIZE, "service_integrations": [ { "integration_type": "disaster_recovery", "source_service": "PRIMARY_SERVICE_NAME", "user_config": {} } ] }' ``` Replace the following placeholders with your values: * `PROJECT_NAME`, for example `crdr-test` * `BEARER_TOKEN` * `RECOVERY_SERVICE_NAME`, for example `pg-dr-test` * `CLOUD_PROVIDER_REGION`, for example `google-europe-west10` * `SERVICE_PLAN`, for example `startup-4` * `SERVICE_TYPE`, for example `pg` * `DISK_SIZE` in MiB, for example `81920` * `PRIMARY_SERVICE_NAME`, for example `pg-primary-test` After sending the request, you can check the CRDR status on each of the CRDR peer services: * Primary service status ``` avn service get PRIMARY_SERVICE_NAME --project PROJECT_NAME --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "RUNNING", "disaster_recovery_role": "active" } ``` * Recovery service status ``` avn service get RECOVERY_SERVICE_NAME --project PROJECT_NAME --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "REBUILDING", "disaster_recovery_role": "passive" } ``` 1. Use the [`aiven_service_integration`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration) resource to create the disaster recovery integration between your primary service and the recovery service. Set `integration_type` to `disaster_recovery`. ``` # Primary PostgreSQL service resource "aiven_postgresql" "primary" { project = var.project_name service_name = var.primary_service_name plan = var.service_plan cloud_name = var.primary_cloud_region } # Recovery PostgreSQL service resource "aiven_postgresql" "recovery" { project = var.project_name service_name = var.recovery_service_name plan = var.service_plan cloud_name = var.recovery_cloud_region } # Disaster recovery integration resource "aiven_service_integration" "disaster_recovery" { project = var.project_name integration_type = "disaster_recovery" source_service_name = aiven_postgresql.primary.service_name destination_service_name = aiven_postgresql.recovery.service_name depends_on = [ aiven_postgresql.primary, aiven_postgresql.recovery ] } ``` 2. Apply the configuration: ``` terraform apply ``` 3. Monitor the setup status: ``` terraform show aiven_service_integration.disaster_recovery ``` Related pages * [Aiven for PostgreSQL® CRDR failover to the recovery region](/docs/products/postgresql/crdr/failover/crdr-failover-to-recovery.md) * [Aiven for PostgreSQL® CRDR revert to the primary region](/docs/products/postgresql/crdr/failover/crdr-revert-to-primary.md) --- # Perform Aiven for PostgreSQL® failover to the recovery region [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Move your workload to another region for disaster recovery or testing purposes. A [failover](/docs/products/postgresql/crdr/crdr-overview.md#failover-to-the-recovery-region) allows you to respond to a region outage or simulate a disaster and test the resilience of your infrastructure. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [CRDR setup](/docs/products/postgresql/crdr/enable-crdr.md) up and running * One of the following tools for operating CRDR: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](https://registry.terraform.io/providers/aiven/aiven/latest/docs) ## Perform a failover[​](#perform-a-failover "Direct link to Perform a failover") Initiate a [failover](/docs/products/postgresql/crdr/crdr-overview.md#failover-to-the-recovery-region) using a tool of your choice: * Console * CLI * API * Terraform 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your primary Aiven for PostgreSQL service. 2. In the **Backups** section, click **Disaster recovery**. 3. On the **Cross-region disaster recovery** page, click **Manage**. 4. In the **Service recovery cycle** wizard, click **Initiate failover** > **Close**. When the failover process is completed, your primary service is **Failed**, and the recovery service is **Active**, which means the recovery service is in control over your workloads. Run [avn service update](/docs/tools/cli/service-cli.md#avn-cli-service-update): ``` avn service update PRIMARY_SERVICE_NAME \ --disaster-recovery-role failed ``` Replace `PRIMARY_SERVICE_NAME` with the name of the primary service, for example, `pg-demo`. Call the [ServiceUpdate endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) to change the `disaster_recovery_role` of the primary service to `failed`: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/PRIMARY_SERVICE_NAME \ -H 'Authorization: Bearer BEARER_TOKEN' \ -H 'content-type: application/json' \ --data '{"disaster_recovery_role": "failed"}' ``` Replace the following placeholders with meaningful data: * `PROJECT_NAME`, for example `crdr-test` * `PRIMARY_SERVICE_NAME`, for example `pg-primary-test` * `BEARER_TOKEN` After sending the request, you can check the CRDR status on each of the CRDR peer services: * Primary service status ``` avn service get pg-primary --project $PROJECT_NAME --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "POWEROFF", "disaster_recovery_role": "failed" } ``` * Recovery service status ``` avn service get pg-primary-dr --project $PROJECT_NAME --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "RUNNING", "disaster_recovery_role": "active" } ``` The [`aiven_service_integration`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration) resource with the `disaster_recovery` type manages the active-passive relationship between services. CRDR operations are performed by manipulating this integration. To trigger failover and promote the recovery service to active: 1. Remove the existing disaster recovery integration from the Terraform state. ``` terraform state rm aiven_service_integration.disaster_recovery ``` 2. If primary service is completely unreachable, remove it from the Terraform state. ``` terraform state rm aiven_postgresql.primary ``` The recovery service is now promoted to active and can handle traffic. Related pages [Aiven for PostgreSQL® CRDR revert to the primary region](/docs/products/postgresql/crdr/failover/crdr-revert-to-primary.md) --- # Perform Aiven for PostgreSQL® failback to the primary region [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Shift your workloads back to the primary region, where your service was hosted originally before failing over to the recovery region. Restore your CRDR setup. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [CRDR failover](/docs/products/postgresql/crdr/failover/crdr-failover-to-recovery.md) completed * One of the following tools for operating CRDR: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](https://registry.terraform.io/providers/aiven/aiven/latest/docs) ## Revert to the primary region[​](#revert-to-the-primary-region "Direct link to Revert to the primary region") Initiate a [revert process](/docs/products/postgresql/crdr/crdr-overview.md#failback-to-the-primary-region) using a tool of your choice: * Console * CLI * API * Terraform 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your primary Aiven for PostgreSQL service. 2. In the **Backups** section, click **Disaster recovery**. 3. On the **Cross-region disaster recovery** page, click **Manage**. 4. In **Service recovery cycle** wizard: 1. Click **Restore primary service**. Wait until the status of your primary service changes from **Rebuilding** to **Passive**. This recreates your primary service in the primary region. 2. Click **Promote primary service**. Wait until the status of the primary service changes from **Rebuilding** to **Active** and the status of the recovery service changes from **Rebalancing** to **Passive**. This switches traffic and replication direction back to the recreated primary service, restoring your CRDR setup to its original configuration. 3. Click **Close**. 1) Restore the primary service by running [avn service update](/docs/tools/cli/service-cli.md#avn-cli-service-update): ``` avn service update PRIMARY_SERVICE_NAME \ --disaster-recovery-role passive ``` Replace `PRIMARY_SERVICE_NAME` with the name of the primary service, for example, `pg-demo`. 2) Promote the primary service to active by running [avn service update](/docs/tools/cli/service-cli.md#avn-cli-service-update): ``` avn service update PRIMARY_SERVICE_NAME \ --disaster-recovery-role active ``` Replace `PRIMARY_SERVICE_NAME` with the name of the primary service, for example, `pg-demo`. 1. Trigger the recreation of the primary service by calling the [ServiceUpdate endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) to change the `disaster_recovery_role` of the primary service to `passive`: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/PRIMARY_SERVICE_NAME \ -H 'Authorization: Bearer BEARER_TOKEN' \ -H 'content-type: application/json' \ --data '{"disaster_recovery_role": "passive"}' ``` Replace the following placeholders with meaningful data: * `PROJECT_NAME`, for example `crdr-test` * `PRIMARY_SERVICE_NAME`, for example `pg-primary-test` * `BEARER_TOKEN` After sending the request, you can check the CRDR status on each of the CRDR peer services: * Primary service status ``` avn service get pg-primary --project $PROJECT_NAME --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "REBUILDING", "disaster_recovery_role": "passive" } ``` * Recovery service status ``` avn service get pg-primary-dr --project $PROJECT_NAME --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "RUNNING", "disaster_recovery_role": "active" } ``` 2. Promote the primary service as active by calling the [ServiceUpdate endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) to change `disaster_recovery_role` of the primary service to `active`: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/PRIMARY_SERVICE_NAME \ -H 'Authorization: Bearer BEARER_TOKEN' \ -H 'content-type: application/json' \ --data '{"disaster_recovery_role": "active"}' ``` Replace the following placeholders with meaningful data: * `PROJECT_NAME`, for example `crdr-test` * `PRIMARY_SERVICE_NAME`, for example `pg-primary-test` * `BEARER_TOKEN` After sending the request, you can check the CRDR status on each of the CRDR peer services: * Primary service status ``` avn service get pg-primary --project $PROJECT_NAME --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "RUNNING", "disaster_recovery_role": "active" } ``` * Recovery service status ``` avn service get pg-primary-dr --project $PROJECT_NAME --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "RUNNING", "disaster_recovery_role": "passive" } ``` The [`aiven_service_integration`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration) resource with the `disaster_recovery` type manages the active-passive relationship between services. CRDR operations are performed by manipulating this integration. To get back to the original primary-recovery setup: 1. Ensure both services exist and are healthy. ``` resource "aiven_postgresql" "primary" { project = var.project_name service_name = var.primary_service_name plan = var.service_plan cloud_name = var.primary_cloud_region } resource "aiven_postgresql" "recovery" { project = var.project_name service_name = var.recovery_service_name plan = var.service_plan cloud_name = var.recovery_cloud_region } ``` If the services were removed from the Terraform state during the disaster, re-import them: ``` terraform import aiven_postgresql.primary PROJECT_NAME/PRIMARY_SERVICE_NAME terraform import aiven_postgresql.recovery PROJECT_NAME/RECOVERY_SERVICE_NAME ``` 2. Re-establish CRDR with the original primary as active: ``` resource "aiven_service_integration" "disaster_recovery_restored" { project = var.project_name integration_type = "disaster_recovery" source_service_name = aiven_postgresql.primary.service_name # Original primary back to active destination_service_name = aiven_postgresql.recovery.service_name # Back to passive role depends_on = [ aiven_postgresql.primary, aiven_postgresql.recovery ] } ``` 3. Apply to restore the original CRDR setup: ``` terraform apply ``` Related pages [Aiven for PostgreSQL® CRDR failover to the recovery region](/docs/products/postgresql/crdr/failover/crdr-failover-to-recovery.md) --- # Failover & failback Perform a [failover](/docs/products/postgresql/crdr/crdr-overview.md#failover-to-the-recovery-region) to the recovery region, and later [revert](/docs/products/postgresql/crdr/crdr-overview.md#failback-to-the-primary-region) to the primary region when ready. ## [Failover](/docs/products/postgresql/crdr/failover/crdr-failover-to-recovery.md) [Move your workload to another region for disaster recovery or testing purposes.](/docs/products/postgresql/crdr/failover/crdr-failover-to-recovery.md) ## [Failback](/docs/products/postgresql/crdr/failover/crdr-revert-to-primary.md) [Shift your workloads back to the primary region, where your service was hosted originally before failing over to the recovery region. Restore your CRDR setup.](/docs/products/postgresql/crdr/failover/crdr-revert-to-primary.md) --- # Perform Aiven for PostgreSQL® switchback to the primary region [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Shift your workloads back to the primary region, where your service was hosted originally before [switching over to the recovery region](/docs/products/postgresql/crdr/crdr-overview.md#switchover-to-the-recovery-region). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [CRDR switchover](/docs/products/postgresql/crdr/switchover/crdr-switchover.md) completed * One of the following tools for operating CRDR: * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](https://registry.terraform.io/providers/aiven/aiven/latest/docs) ## Switch back[​](#switch-back "Direct link to Switch back") * CLI * API * Terraform Run [avn service update](/docs/tools/cli/service-cli.md#avn-cli-service-update) to promote the primary service back to active: ``` avn service update PRIMARY_SERVICE_NAME \ --disaster-recovery-role active ``` Replace `PRIMARY_SERVICE_NAME` with the name of the primary service, for example, `pg-demo`. Verify the switchback by checking both services: * Primary service status: ``` avn service get PRIMARY_SERVICE_NAME \ --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "RUNNING", "disaster_recovery_role": "active" } ``` * Recovery service status: ``` avn service get RECOVERY_SERVICE_NAME \ --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "RUNNING", "disaster_recovery_role": "passive" } ``` Call the [ServiceUpdate endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) to change the `disaster_recovery_role` of the primary service to `active`: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/PRIMARY_SERVICE_NAME \ -H 'Authorization: Bearer BEARER_TOKEN' \ -H 'content-type: application/json' \ --data '{"disaster_recovery_role": "active"}' ``` Replace the following: * `PROJECT_NAME`, for example `crdr-test` * `PRIMARY_SERVICE_NAME`, for example `pg-demo` * `BEARER_TOKEN` After sending the request, verify the status of each service: * Primary service status: ``` avn service get PRIMARY_SERVICE_NAME \ --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "RUNNING", "disaster_recovery_role": "active" } ``` * Recovery service status: ``` avn service get RECOVERY_SERVICE_NAME \ --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "RUNNING", "disaster_recovery_role": "passive" } ``` The Terraform provider manages the `disaster_recovery` integration as a static topology declaration. All integration fields are immutable, so a switchback requires destroying the switched integration and restoring the original one. This approach has a window during which the DR integration is offline. For a lower-risk switchback, use the Aiven CLI or API instead. To use Terraform: 1. Remove the switched disaster recovery integration from Terraform state. ``` terraform destroy -target=aiven_service_integration.disaster_recovery_switched ``` 2. Restore the original integration with the primary service as the source. ``` resource "aiven_service_integration" "disaster_recovery" { project = var.project_name integration_type = "disaster_recovery" source_service_name = aiven_postgresql.primary.service_name destination_service_name = aiven_postgresql.recovery.service_name } ``` 3. Apply the restored configuration: ``` terraform apply ``` After the switchback completes, your primary service is **Active**, and the recovery service is **Passive**, which means the primary service is in control over your workloads. Related pages * [Perform Aiven for PostgreSQL® failover to the recovery region](/docs/products/postgresql/crdr/failover/crdr-failover-to-recovery.md) * [Perform Aiven for PostgreSQL® switchover to the recovery region](/docs/products/postgresql/crdr/switchover/crdr-switchover.md) --- # Perform Aiven for PostgreSQL® switchover to the recovery region [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Perform a planned promotion of your recovery service while the primary service is healthy. [Switch over](/docs/products/postgresql/crdr/crdr-overview.md#switchover-to-the-recovery-region) to your Aiven for PostgreSQL® recovery service for planned maintenance, simulating a disaster, or testing the resilience of your infrastructure. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [CRDR setup](/docs/products/postgresql/crdr/enable-crdr.md) up and running * One of the following tools for operating CRDR: * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](https://registry.terraform.io/providers/aiven/aiven/latest/docs) ## Switch over[​](#switch-over "Direct link to Switch over") * CLI * API * Terraform Run [avn service update](/docs/tools/cli/service-cli.md#avn-cli-service-update) to promote the recovery service to active: ``` avn service update RECOVERY_SERVICE_NAME \ --disaster-recovery-role active ``` Replace `RECOVERY_SERVICE_NAME` with the name of the recovery service, for example, `pg-demo-recovery`. Verify the switchover by checking both services: * Recovery service status: ``` avn service get RECOVERY_SERVICE_NAME \ --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "RUNNING", "disaster_recovery_role": "active" } ``` * Primary service status: ``` avn service get PRIMARY_SERVICE_NAME \ --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "RUNNING", "disaster_recovery_role": "passive" } ``` Call the [ServiceUpdate endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) to change the `disaster_recovery_role` of the recovery service to `active`: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/RECOVERY_SERVICE_NAME \ -H 'Authorization: Bearer BEARER_TOKEN' \ -H 'content-type: application/json' \ --data '{"disaster_recovery_role": "active"}' ``` Replace the following: * `PROJECT_NAME`, for example `crdr-test` * `RECOVERY_SERVICE_NAME`, for example `pg-demo-recovery` * `BEARER_TOKEN` After sending the request, verify the status of each service: * Recovery service status: ``` avn service get RECOVERY_SERVICE_NAME \ --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "RUNNING", "disaster_recovery_role": "active" } ``` * Primary service status: ``` avn service get PRIMARY_SERVICE_NAME \ --json | jq '{state: .state, disaster_recovery_role: .disaster_recovery_role}' ``` Expect the following output: ``` { "state": "RUNNING", "disaster_recovery_role": "passive" } ``` The Terraform provider manages the `disaster_recovery` integration as a static topology declaration. All integration fields are immutable, so a switchover requires destroying the existing integration and creating a new one with the source and destination reversed. This approach has a window during which the DR integration is offline. For a lower-risk switchover, use the Aiven CLI or API instead. To use Terraform: 1. Remove the existing disaster recovery integration from Terraform state. ``` terraform destroy -target=aiven_service_integration.disaster_recovery ``` 2. Create an integration with the recovery service as the source. ``` resource "aiven_service_integration" "disaster_recovery_switched" { project = var.project_name integration_type = "disaster_recovery" source_service_name = aiven_postgresql.recovery.service_name destination_service_name = aiven_postgresql.primary.service_name } ``` 3. Apply the new configuration: ``` terraform apply ``` After the switchover completes, your primary service is **Passive**, and the recovery service is **Active**, which means the recovery service is in control over your workloads. Related pages * [Perform Aiven for PostgreSQL® switchback to the primary region](/docs/products/postgresql/crdr/switchover/crdr-switchback.md) * [Perform Aiven for PostgreSQL® failover to the recovery region](/docs/products/postgresql/crdr/failover/crdr-failover-to-recovery.md) --- # Switchover & switchback Perform a [switchover](/docs/products/postgresql/crdr/crdr-overview.md#switchover-to-the-recovery-region) to the recovery region, and later [revert](/docs/products/postgresql/crdr/crdr-overview.md#switchback-to-the-primary-region) to the primary region. ## [Switchover](/docs/products/postgresql/crdr/switchover/crdr-switchover.md) [Perform a planned promotion of your recovery service while the primary service is healthy.](/docs/products/postgresql/crdr/switchover/crdr-switchover.md) ## [Switchback](/docs/products/postgresql/crdr/switchover/crdr-switchback.md) [Shift your workloads back to the primary region, where your service was hosted originally before switching over to the recovery region.](/docs/products/postgresql/crdr/switchover/crdr-switchback.md) --- # Database management in Aiven for PostgreSQL® Create and manage databases, extensions, and routine maintenance tasks in your Aiven for PostgreSQL® service. Related pages * [Create a database](/docs/products/postgresql/howto/create-database.md) * [Enable JIT compilation](/docs/products/postgresql/howto/enable-jit.md) * [Use the pg\_repack extension](/docs/products/postgresql/howto/use-pg-repack-extension.md) * [Avoid transaction ID wraparound](/docs/products/postgresql/howto/check-avoid-transaction-id-wraparound.md) --- # Get started with Aiven for PostgreSQL® Start using Aiven for PostgreSQL® by creating a service, connecting to it, and loading sample data. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Console * Terraform * MCP - Access to the [Aiven Console](https://console.aiven.io) - [psql](https://www.postgresql.org/download/) command line tool installed * [Terraform installed](https://www.terraform.io/downloads) * A [personal token](https://docs.aiven.io/docs/platform/howto/create_authentication_token.html) * [psql](https://www.postgresql.org/download/) command line tool installed - An MCP-compatible client such as Cursor, Claude Code, Claude Desktop, VS Code, or Gemini CLI - A configured Aiven MCP server in your client ## Create a service[​](#create-a-service "Direct link to Create a service") * Console * Terraform * MCP 1. In your project, click **Services**. 2. Click **Create service**. 3. Select **PostgreSQL®**. 4. Select a **Service tier**. 5. Select a **Cloud**. note You cannot choose a cloud provider or a specific cloud region on the Free tier. 6. Select a **Plan**. note The plans available can vary between cloud providers and regions for the same service. 7. In the **Service details**, enter a name for your service. 8. Optional: Add service tags. 9. In the **Service summary**, click **Create service**. The status of the service is **Rebuilding** during its creation. When the status is **Running**, you can start using the service. This typically takes a couple of minutes and can vary between cloud providers and regions. The following example files are also available in the [Aiven Terraform Provider repository](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/postgres) on GitHub. 1. Create a file named `provider.tf` and add the following: ``` Loading... ``` 2. Create a file named `service.tf` and add the following: ``` Loading... ``` 3. Create a file named `variables.tf` and add the following: ``` Loading... ``` 4. Create the `terraform.tfvars` file and add the values for your token and project name. To apply your Terraform configuration: 1. Initialize Terraform by running: ``` terraform init ``` The output is similar to the following: ``` Initializing the backend... Initializing provider plugins... - Finding aiven/aiven versions matching ">= 4.0.0, < 5.0.0"... - Installing aiven/aiven v4.9.2... - Installed aiven/aiven v4.9.2 ... Terraform has been successfully initialized! ... ``` 2. To create an execution plan and preview the changes, run: ``` terraform plan ``` 3. To deploy your changes, run: ``` terraform apply --auto-approve ``` Open your AI assistant and describe the service you want to create. For example: > Create an Aiven for PostgreSQL service named `my-pg` in the `google-europe-west1` region using the `hobbyist` plan. The assistant uses the Aiven MCP server to create the service. For setup instructions, see [Aiven MCP server](/docs/tools/mcp-server.md). ## Configure a service[​](#configure-a-service "Direct link to Configure a service") Edit your service settings if the default service configuration doesn't meet your needs. * Console * Terraform 1. Select the new service from the list of services on the **Services** page. 2. On the **Overview** page, select **Service settings** from the sidebar. 3. In the **Advanced configuration** section, make changes to the service configuration. See the available configuration options in [Advanced parameters for Aiven for PostgreSQL](/docs/products/postgresql/reference/advanced-params.md). See [the `aiven_pg` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/pg) for the full schema. ## Connect to the service[​](#connect-to-the-service "Direct link to Connect to the service") * Console * Terraform * psql 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to your organization > project > Aiven for PostgreSQL service. 2. On the **Overview** page of your service, click **Quick connect**. 3. In the **Connect** window, select a tool or language to connect to your service, follow the connection instructions, and click **Done**. ``` psql 'postgres://ADMIN_PASSWORD@vine-pg-test.a.aivencloud.com:12691/defaultdb?sslmode=require' ``` Access your service with [the psql client](/docs/products/postgresql/howto/connect-psql.md) using the `postgresql_service_uri` Terraform output. ``` psql "$(terraform output -raw postgresql_service_uri)" ``` The output of the command is similar to the following: ``` psql (13.2) SSL connection (protocol: TLSv1.3, cipher: TLS_AES_256_GCM_SHA384, bits: 256, compression: off) Type "help" for help. ``` [Connect to your new service](/docs/products/postgresql/howto/connect-psql.md) with [psql](https://www.postgresql.org/download/) CLI tool. tip Check more tools for connecting to Aiven for PostgreSQL in [Connect to Aiven for PostgreSQL](/docs/products/postgresql/howto/list-code-samples.md). ## Load a test dataset[​](#load-a-test-dataset "Direct link to Load a test dataset") `dellstore2` is a standard store dataset with products, orders, inventory, and customer information. 1. Download the `dellstore2-normal-1.0.tar.gz` file from the [PostgreSQL website](https://www.postgresql.org/ftp/projects/pgFoundry/dbsamples/dellstore2/dellstore2-normal-1.0/) and unzip it. 2. From the folder where you unzipped the file, [connect to your PostgreSQL instance](/docs/products/postgresql/howto/connect-psql.md), create a `dellstore` database, and connect to it: ``` CREATE DATABASE dellstore; \c dellstore ``` 3. Populate the database: ``` \i dellstore2-normal-1.0.sql ``` 4. Verify what objects have been created: ``` \d ``` Expected output ``` List of relations Schema | Name | Type | Owner --------+--------------------------+----------+---------- public | categories | table | avnadmin public | categories_category_seq | sequence | avnadmin public | cust_hist | table | avnadmin public | customers | table | avnadmin public | customers_customerid_seq | sequence | avnadmin public | inventory | table | avnadmin public | orderlines | table | avnadmin public | orders | table | avnadmin public | orders_orderid_seq | sequence | avnadmin public | products | table | avnadmin public | products_prod_id_seq | sequence | avnadmin public | reorder | table | avnadmin (12 rows) ``` ## Query data[​](#query-data "Direct link to Query data") ### Read data[​](#read-data "Direct link to Read data") Retrieve all the data from a table, for example, from `orders`: ``` SELECT * FROM orders; ``` Expected output ``` orderid | orderdate | customerid | netamount | tax | totalamount ---------+------------+------------+-----------+-------+------------- 1 | 2004-01-27 | 7888 | 313.24 | 25.84 | 339.08 2 | 2004-01-01 | 4858 | 54.90 | 4.53 | 59.43 3 | 2004-01-17 | 15399 | 160.10 | 13.21 | 173.31 4 | 2004-01-28 | 17019 | 106.67 | 8.80 | 115.47 5 | 2004-01-09 | 14771 | 256.00 | 21.12 | 277.12 6 | 2004-01-11 | 13734 | 382.59 | 31.56 | 414.15 7 | 2004-01-05 | 17622 | 256.44 | 21.16 | 277.60 8 | 2004-01-18 | 8331 | 67.85 | 5.60 | 73.45 9 | 2004-01-06 | 14902 | 29.82 | 2.46 | 32.28 10 | 2004-01-18 | 15112 | 20.78 | 1.71 | 22.49 ... (20000 rows) ``` ### Write data[​](#write-data "Direct link to Write data") Add a row to a table, for example, to `customers`: ``` INSERT INTO customers(customerid,firstname,lastname,address1,city,country,region,email,creditcardtype,creditcard,creditcardexpiration,username,password,age,gender) VALUES(20001,'John','Doe','WEDEBTRTBD','NY','US',11,'john.doe@mailbox.com',3,1879279217775922,2025/11,'user20001','password',44,'M'); ``` Expected output ``` INSERT 0 1 ``` Check that your new row is there: ``` SELECT * FROM customers WHERE firstname = 'John'; ``` Expected output ``` customerid | firstname | lastname | address1 | address2 | city | state | zip | country | region | email | phone | creditcardtype | creditcard | creditcardexpiration | username | password | age | income | gender ------------+-----------+----------+------------+----------+------+-------+-----+---------+--------+----------------------+-------+----------------+------------------+----------------------+-----------+------------+-----+--------+-------- 20001 | John | Doe | WEDEBTRTBD | | NY | | | US | 11 | john.doe@mailbox.com | | 3 | 1879279217775922 | 184 | user20001 | password | 44 | | M (1 row) ``` Related pages * [Connection pooling](/docs/products/postgresql/concepts/pg-connection-pooling.md) and [Pgbouncer](/docs/products/postgresql/howto/pgbouncer-stats.md) * [High availability](/docs/products/postgresql/concepts/high-availability.md) * [Restrict access](/docs/products/postgresql/howto/readonly-user.md) * [Migrate your PostgreSQL to Aiven](/docs/products/postgresql/concepts/aiven-db-migrate.md) * [Aiven Service Level Agreement](https://aiven.io/sla) --- # AI database optimizer for Aiven for PostgreSQL® Use **Aiven AI Database Optimizer** to receive optimization suggestions to your databases and queries. Aiven's artificial intelligence considers various aspects to suggest optimizations, for example query structure, table size, existing indexes and their cardinality, column types and sizes, and the connections between the tables and columns in the query. To optimize a query automatically: 1. In the [Aiven Console](https://console.aiven.io/login), open your Aiven for PostgreSQL® service. 2. In the **Observe** section, click **AI insights**. 3. For the query of your choice, click **Optimize**. 4. In the **Query optimization report** window, see the optimization suggestion and apply the suggestion by running the provided SQL queries. * To display potential alternative optimization recommendations, click **Advanced options**. * To display the diff view, click **Query diff**. * To display explanations about the optimization, click **Optimization details**. note The quality of the optimization suggestions is proportional to the amount of data collected about the performance of your database. Frequently asked questions **Does Aiven AI Optimizer mask/obfuscate my queries?** Yes, Aiven AI Optimizer provides a non-intrusive solution to optimize your database performance without compromising sensitive data access. It achieves this by gathering information on schema structure, database statistics, and other signals to detect potential performance problems and offer optimization recommendations, without requiring credentials or access to the actual data in the database. To address the possibility of slow query logs containing sensitive data, Aiven offers data masking capabilities that replace sensitive parameters within queries with question marks (`?`). Data masking is enabled by default. note The masking option is not available for the Standalone SQL query optimizer yet. For one-time query optimizations when you do not run an Aiven for PostgreSQL® service, use the [standalone SQL query optimizer](https://aiven.io/tools/sql-query-optimizer). Related pages * [AI DB Optimizer for Aiven for MySQL®](/docs/products/mysql/howto/ai-insights.md) * [Standalone SQL query optimizer](https://aiven.io/tools/sql-query-optimizer) --- # Report and analyze with Google Looker Studio Google Looker Studio (previously Data Studio) allows you to create reports and visualisations of the data in your Aiven for PostgreSQL® database, and combine these with data from many other data sources. ## Variables[​](#variables "Direct link to Variables") These are the values you will need to connect to Google Looker Studio: | Variable | Description | | ---------- | --------------------------------------------------------------------------- | | `HOSTNAME` | Hostname for the PostgreSQL connection, from the service overview page | | `PORT` | Port for the PostgreSQL connection, from the service overview page | | `DATABASE` | Database Name for the PostgreSQL connection, from the service overview page | | `PASSWORD` | `avnadmin` password, from the service overview page | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") 1. You will need a Google account, to access Google Looker Studio. 2. On the Aiven Console service page for your PostgreSQL database, download the CA certificate. The default filename is `ca.pem`. ## Connect your Aiven for PostgreSQL data source to Google Looker Studio[​](#connect-your-aiven-for-postgresql-data-source-to-google-looker-studio "Direct link to Connect your Aiven for PostgreSQL data source to Google Looker Studio") 1. Login to Google and open [Google Looker Studio](https://lookerstudio.google.com/overview) . 2. Select **Create** and choose **Data source**. 3. Fill in the requested information and agree to the Google terms and conditions. 4. Select the **PostgreSQL** Google Connector. 5. On the **Basic** tab, set * **Host name** to the `HOSTNAME` * **Port**: to the `PORT` * **Database** to the `DATABASE` * **Username** to `avnadmin` * **Password** to the `PASSWORD` 6. Select **Enable SSL** and upload your server certificate file, `ca.pem`. 7. Click **AUTHENTICATE**. 8. Choose the table to be queried, or select **CUSTOM QUERY** to create an SQL query. 9. Click **CONNECT** You can then proceed to create reports and create visualisations. --- # Back up your Aiven for PostgreSQL® service to another region Copy your Aiven for PostgreSQL® service backups to a secondary region for disaster recovery. In addition to the primary service backup, you can have a secondary backup in an alternative location. important This feature is in [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-). Contact your account team to enable it. Backup to another region (BTAR) is a disaster recovery feature that allows backup files to be copied from the service's primary backup region to an additional (secondary) region. BTAR can bolster data resilience and helps improve data protection against disasters in the primary backup region. When the primary region is down, BTAR allows forking the service from an additional copy of the backup residing in a secondary region. ## Limitations[​](#limitations "Direct link to Limitations") * The cloud provider for your additional backup region must match the cloud provider for your service and the primary backup. * Secondary backup can only be restored in the region where it was stored. For a service that has the backup to another region (BTAR) feature enabled, you can check the service backup status, change the backup region, monitor the replication lag, fork and restore using the cross-region backup, or migrate to another cloud or region. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * At least one Aiven service with BTAR enabled * Access to the [Aiven Console](https://console.aiven.io/) * [Aiven API](/docs/tools/api.md) * [Aiven CLI](/docs/tools/cli.md) ## Change a backup region[​](#change-a-backup-region "Direct link to Change a backup region") 1. Log in to the [Aiven Console](https://console.aiven.io/) and go to your project. 2. On the **Services** page, select an Aiven service on which you'd like to enable BTAR. 3. On your service page, in the **Backups** section, click **Backup management**. 4. On the **Backups** page, click **Actions** > **Edit secondary backup location**. 5. In the **Edit secondary backup location** window, use the **Backup location** menu to select a region for your additional backup. Confirm your choice by selecting **Save**. important You can change the backup region once in 24 hours. ## Monitor a service with BTAR[​](#monitor-a-service-with-btar "Direct link to Monitor a service with BTAR") There are a few things you may want to check for your Aiven service in the context of BTAR: * What is the status of a secondary backup? * Does your service have a backup in another region? * What is the target region of the secondary backup? * What is the replication lag between data availability in the primary region and the secondary region? ### Check BTAR status[​](#check-btar-status "Direct link to Check BTAR status") To see the availability, the status, and the target region of a secondary (BTAR) backup in the [Aiven Console](https://console.aiven.io/), go to your service page > **Backups** > **Secondary backup location**. ### Determine replication lag[​](#determine-replication-lag "Direct link to Determine replication lag") Determine the target region and the replication lag for a secondary (BTAR) backup of your service, call the [ServiceBackupToAnotherRegionReport](https://api.aiven.io/doc/#tag/Service/operation/ServiceBackupToAnotherRegionReport) endpoint. Configure the call as follows: 1. Enter `YOUR-PROJECT-NAME` and `YOUR-SERVICE-NAME` into the URL. 2. Specify `DESIRED-TIME-PERIOD` depending on the time period you need the metrics for: select one of the following values for the `period` key: `hour`, `day`, `week`, `month`, or `year`. ``` curl --request POST \ --url https://api.aiven.io/v1/project/YOUR-PROJECT-NAME/service/YOUR-SERVICE-NAME/backup_to_another_region/report \ --header 'Authorization: Bearer YOUR-BEARER-TOKEN' \ --header 'content-type: application/json' \ --data '{"period":"DESIRED-TIME-PERIOD"}' ``` As output, you get metrics including replication lags at specific points in time. ## Fork and restore a service with BTAR[​](#fork-and-restore "Direct link to Fork and restore a service with BTAR") You can use the [Aiven Console](/docs/tools/aiven-console.md), [API](/docs/tools/api.md), or [CLI](/docs/tools/cli.md) to recover your service from a backup in another region. To restore your service using BTAR, create a fork of the original service in the region where the secondary backup resides. note When you **fork & restore** from the secondary backup, your new fork service is created in the cloud and region where the secondary backup is located. The fork service gets the same plan that the primary service uses. Backups of the fork service are located in the region where this new service is hosted. * Aiven Console * Aiven CLI * Aiven API 1. Open the [Aiven Console](https://console.aiven.io/) and go to your service homepage. 2. Click **Backups**. 3. On the **Backups** page, select **Fork & restore**. 4. In the **New database fork** window: 1. Set **Backup location** to either **Primary location** or **Secondary location**. 2. Set **Backup version** to one of the following: * **Latest transaction** * **Point in time**: Set it up to no earlier than the time of taking the oldest replicated base backup. 3. Specify a name for the new fork service. 4. Select **Create fork**. Run the [avn service create](/docs/tools/cli/service-cli.md#avn-cli-service-create) command with the `--service-to-fork-from` option and the `--recovery-target-time`option. Set `--recovery-target-time` to no earlier than the time of taking the oldest replicated base backup. ``` avn service create FORK_SERVICE_NAME \ --plan SERVICE_PLAN \ --project PROJECT_NAME \ --service-type SERVICE_TYPE \ --cloud SECONDARY_BACKUP_REGION \ --recovery-target-time "YYYY-MM-DDTHH:MM:SS+00:00" \ --service-to-fork-from PRIMARY_SERVICE_NAME ``` Replace the following with meaningful data: * `FORK_SERVICE_NAME` * `SERVICE_PLAN` * `PROJECT_NAME` * `SERVICE_TYPE` * `SECONDARY_BACKUP_REGION` * `PRIMARY_SERVICE_NAME` * `YYYY-MM-DDTHH:MM:SS+00:00` Use the [ServiceCreate](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) API to create a fork service. When constructing the API request, add the `user_config` object to the request body and nest the `service_to_fork_from` field and the `recovery_target_time` field inside. Set `recovery_target_time` to no earlier than the time of taking the oldest replicated base backup. ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "cloud": "SECONDARY_BACKUP_REGION", "plan": "SERVICE_PLAN", "service_name": "FORK_SERVICE_NAME", "service_type": "SERVICE_TYPE", "user_config": { "service_to_fork_from": "PRIMARY_SERVICE_NAME", "recovery_target_time": "YYYY-MM-DDTHH:MM:SS+00:00" } }' ``` Replace the following with meaningful data: * `FORK_SERVICE_NAME` * `SERVICE_PLAN` * `PROJECT_NAME` * `SERVICE_TYPE` * `SECONDARY_BACKUP_REGION` * `PRIMARY_SERVICE_NAME` * `YYYY-MM-DDTHH:MM:SS+00:00` ## Migrate a service with BTAR[​](#migrate-a-service-with-btar "Direct link to Migrate a service with BTAR") You can migrate a service with BTAR the same way you [migrate a service with a regular backup](/docs/platform/howto/migrate-services-cloud-region.md). note When you migrate your service, locations of service backups, both primary and secondary ones, do not change. ## Delete a cross-region backup[​](#delete-a-cross-region-backup "Direct link to Delete a cross-region backup") Delete an additional service backup created in a region different from your primary backup region. You can delete a cross-region backup using the [Aiven Console](/docs/tools/aiven-console.md), [API](/docs/tools/api.md), or [CLI](/docs/tools/cli.md). When you delete the additional cross-region backup, you still have the default backup located in the primary, service-hosting region. * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. From the **Services** view, select an Aiven service on which you'd like to disable BTAR. 3. On your service's page, click **Backups**. 4. On the **Backups** page, click **Actions** > **Secondary backup location**. 5. In the **Edit secondary backup location** window, select **Disable**. Your additional service backup is no longer visible on your service's **Backups** page in the **Secondary backup location** column. To remove secondary backups for your service, use the [avn service update](/docs/tools/cli/service-cli.md) command to remove all target region names from the `additional_backup_regions` array. ``` avn service update your-sevice-name \ -c additional_backup_regions=\[\] ``` To remove secondary backups for your service, update the service configuration. Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint to remove all target regions names from the `additional_backup_regions` array. ``` curl --request PUT \ --url https://api.aiven.io/v1/project/YOUR_PROJECT_NAME/service/YOUR_SERVICE_NAME \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "additional_backup_regions": [] } }' ``` Related pages * [Backups](/docs/products/postgresql/concepts/pg-backups.md) * [Track restore progress](/docs/products/postgresql/howto/track-restore-progress.md) --- # Change the plan for your Aiven for PostgreSQL® service Change the service plan for your Aiven for PostgreSQL® service to scale resources up or down and optimize costs. Adjust the plan of your services at any time to scale your services as needed and optimize costs. If you can't find a suitable plan, you can [request a custom plan](/docs/platform/concepts/service-pricing.md). tip If you plan to upgrade your service plan, do it immediately after a full backup. This reduces the amount of incremental changes that need to be applied on top of the base backup, which speeds up the upgrade itself. important * When changing a service plan, reserve an additional 25% of disk space. This requirement applies to upgrades and downgrades. * Downgrading to a plan with fewer VMs is supported for most services, including Aiven for Apache Kafka®, Aiven for PostgreSQL®, Aiven for OpenSearch®, Aiven for ClickHouse®, Aiven for MySQL®, Aiven for Metrics, and Aiven for Valkey™. * Changing a service plan triggers a node recycle, service rebuilding, and any pending maintenance updates. - Console - Terraform - CLI 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Change plan**. 3. Select a plan that provides at least 125% of the current disk size and click **Change plan**. Update the `plan` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). To change a service plan in the Aiven CLI, use the [`avn service update --plan `](/docs/tools/cli/service-cli.md#avn-cli-service-update) command. Your service's state becomes **Rebuilding** and remains accessible. When the state switches to **Running**, your new service plan is active. Related pages * [Scale disk storage](/docs/products/postgresql/howto/scale-disk-storage.md) * [Prepare for high load](/docs/products/postgresql/howto/prepare-for-high-load.md) --- # Check and avoid transaction ID wraparound The PostgreSQL® transaction control mechanism assigns a transaction ID to every row that is modified in the database; these IDs control the visibility of that row to other concurrent transactions. The transaction ID is a 32-bit number, where 2 billion (2 thousand million) IDs are always in the "visible past" and the remainder (about 2.2 billion) are reserved for future transactions and not visible to the running transaction. To avoid a transaction wraparound and having old, existing rows invisible when more transactions are created, PostgreSQL® requires an occasional cleanup and "freezing" of old rows. ## Automatic or manual vacuuming[​](#automatic-or-manual-vacuuming "Direct link to Automatic or manual vacuuming") You can manually trigger a cleanup by executing `VACUUM FREEZE`, but the autovacuum also does this automatically once a configured number of transactions have been created since the last freeze. ## Check the `autovacuum` frequency[​](#check-the-autovacuum-frequency "Direct link to check-the-autovacuum-frequency") Aiven for PostgreSQL® sets that number to scale according to the database size, up to 1.5 billion transactions (which leaves 500 million transaction IDs available before a forced freeze), to avoid unnecessary churn for stable data in existing tables. To check your transaction freeze limits, run the following command in your PostgreSQL® instance: ``` show autovacuum_freeze_max_age ``` This shows you the number of transactions that trigger autovacuum to start freezing old rows. ## Configuring client applications[​](#configuring-client-applications "Direct link to Configuring client applications") Some applications may not automatically adjust their configuration based on the actual PostgreSQL® configuration and may show unnecessary warnings. For example, [PgHero](https://github.com/ankane/pghero)'s default settings trigger an alert once 500 million transactions have been created, while the correct behavior might be to trigger an alert after 1.5 billion transactions. The `transaction_id_danger` setting controls this behavior, and changing the value from 1500000000 (1,500,000,000 or 1.5 billion) to 500000000 (500,000,000 or 500 million) would make it warn you when appropriate. Related pages * [25.1.5. Preventing Transaction ID Wraparound Failures](https://www.postgresql.org/docs/current/routine-vacuuming.html#VACUUM-FOR-WRAPAROUND) * [Table 9.76. Transaction ID and Snapshot Information Functions](https://www.postgresql.org/docs/14/functions-info.html#FUNCTIONS-PG-SNAPSHOT) --- # Claim public schema ownership When an Aiven for PostgreSQL® instance is created, the `public` schema is owned by the `postgres` user that is available only to Aiven for management purposes. If changes to the `public` schema are required, you can claim the ownership using the `aiven_extras` extension as the `avnadmin` database user. 1. Enable the `aiven_extras` extension: ``` CREATE EXTENSION aiven_extras CASCADE; ``` 2. Claim the public schema ownership with the dedicated `claim_public_schema_ownership` function: ``` SELECT * FROM aiven_extras.claim_public_schema_ownership(); ``` Now the `avnadmin` user owns the public schema and can modify it. --- # Connect to Aiven for PostgreSQL® with DataGrip Use [DataGrip](https://www.jetbrains.com/datagrip/) to connect to your Aiven for PostgreSQL® service. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to the [Aiven Console](https://console.aiven.io/) * At least one running Aiven for PostgreSQL service * [DataGrip](https://www.jetbrains.com/datagrip/download/) installed on your machine ## Get JDBC URI from Aiven Console[​](#get-jdbc-uri-from-aiven-console "Direct link to Get JDBC URI from Aiven Console") 1. Log in to [Aiven Console](https://console.aiven.io/) and go to your organization > project > Aiven for PostgreSQL service. 2. On the service **Overview** page, select **Quick connect**. 3. In the **Connect** window 1. Choose to connect with Java using the **Connect with** dropdown menu. 2. Copy the generated JDBC URI. 3. Select **Done**. ## Connect to JDBC URI from DataGrip[​](#connect-to-jdbc-uri-from-datagrip "Direct link to Connect to JDBC URI from DataGrip") 1. Open DataGrip on your machine, and select **File** > **New** > **Data Source** > **PostgreSQL** from the top navigation menu. 2. In the **Data Sources and Drivers** window > **General** tab, paste the URI copied from the [Aiven Console](https://console.aiven.io/). 3. Select **OK** to create and save the connection. ![Connect to Aiven for PostgreSQL with DataGrip](/docs/assets/images/datagrip-create-connection-4a7d16dce5b297b748616cc4e3e545b0.png) The connection to your Aiven for PostgreSQL service has been established and is visible in DataGrip > **Database Explorer**. Related pages * [Connect to Aiven for PostgreSQL](/docs/products/postgresql/howto/list-code-samples.md) for more tools you can use for connecting to your service * [DataGrip](https://www.jetbrains.com/datagrip/) * [DataGrip download](https://www.jetbrains.com/datagrip/download/) --- # Connect to Aiven for PostgreSQL® with DBeaver Use [DBeaver](https://dbeaver.com/) to connect to your Aiven for PostgreSQL® service. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to the [Aiven Console](https://console.aiven.io/) * At least one running Aiven for PostgreSQL service * [DBeaver](https://dbeaver.io/download/) installed on your machine ## Get JDBC URI from Aiven Console[​](#get-jdbc-uri-from-aiven-console "Direct link to Get JDBC URI from Aiven Console") 1. Log in to [Aiven Console](https://console.aiven.io/) and navigate to your organization > project > Aiven for PostgreSQL service. 2. On the service **Overview** page, select **Quick connect**. 3. In the **Connect** window 1. Choose to connect with Java using the **Connect with** dropdown menu. 2. Copy the generated JDBC URI. 3. Select **Done**. ## Connect to JDBC URI from DBeaver[​](#connect-to-jdbc-uri-from-dbeaver "Direct link to Connect to JDBC URI from DBeaver") 1. Open DBeaver on your machine, and select **Database** > **New Database Connection** from the top navigation menu. 2. In the **Connect to database** window, select PostgreSQL and click **Next**. 3. In the **Connection Settings** window > **Main** tab > **Server** section, choose to connect with URL and paste the URI copied from the [Aiven Console](https://console.aiven.io/). 4. Select **Finish** to create and save the connection. ![Connect to Aiven for PostgreSQL with DBeaver](/docs/assets/images/dbeaver-create-connection-6623e3045c2086c87af31983dbf36a33.png) The connection to your Aiven for PostgreSQL service has been established and is visible in DBeaver > **Database Navigator**. Related pages * [Connect to Aiven for PostgreSQL](/docs/products/postgresql/howto/list-code-samples.md) for more tools you can use for connecting to your service * [DBeaver](https://dbeaver.com/) * [DBeaver Community](https://dbeaver.io/) --- # Connect to Aiven for PostgreSQL® with Go This example connects to PostgreSQL® service from Go, making use of the `pg` library. ## Variables[​](#variables "Direct link to Variables") These are the placeholders you will need to replace in the code sample: | Variable | Description | | ---------------- | ------------------------------------------------------------- | | `POSTGRESQL_URI` | URL for PostgreSQL connection, from the service overview page | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") For this example you will need: * The Go `pq` library: ``` go get github.com/lib/pq ``` * [Download CA certificates](/docs/platform/concepts/tls-ssl-certificates.md#download-ca-certificates) from the service overview page, this example assumes it is in a local file called `ca.pem`. ## Code[​](#code "Direct link to Code") Add the following to `main.go` and replace the placeholder with the PostgreSQL URI: ``` package main import ( "database/sql" "fmt" "log" "net/url" _ "github.com/lib/pq" ) func main() { serviceURI := "POSTGRESQL_URI" conn, _ := url.Parse(serviceURI) conn.RawQuery = "sslmode=verify-ca;sslrootcert=ca.pem" db, err := sql.Open("postgres", conn.String()) if err != nil { log.Fatal(err) } defer db.Close() rows, err := db.Query("SELECT version()") if err != nil { panic(err) } for rows.Next() { var result string err = rows.Scan(&result) if err != nil { panic(err) } fmt.Printf("Version: %s\n", result) } } ``` This code creates a PostgreSQL client and opens a connection to the database. Then runs a query checking the database version and prints the response note This example replaces the query string parameter to specify `sslmode=verify-ca` to make sure that the SSL certificate is verified, and adds the location of the cert. To run the code: ``` go run main.go ``` If the script runs successfully, the outputs should be the PostgreSQL version running in your service like: ``` Version: PostgreSQL PG_VERSION_NUMBER on x86_64-pc-linux-gnu, compiled by gcc, a 68c5366192 p 6520304dc1, 64-bit ``` --- # Connect to Aiven for PostgreSQL® with Java This example connects to PostgreSQL® service from Java, making use of JDBC Driver. ## Variables[​](#variables "Direct link to Variables") These are the placeholders you will need to replace in the code sample: | Variable | Description | | ---------- | ----------------------------------------------------------------------- | | `HOSTNAME` | Hostname for PostgreSQL connection, from the service overview page | | `PORT` | Port for PostgreSQL connection, from the service overview page | | `DATABASE` | Database Name for PostgreSQL connection, from the service overview page | | `PASSWORD` | `avnadmin` password, from the service overview page | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") For this example you will need the PostgreSQL Driver: 1. If you have maven version >= 2+, run: ``` mvn org.apache.maven.plugins:maven-dependency-plugin:2.8:get -Dartifact=org.postgresql:postgresql:42.3.2:jar -Ddest=postgresql-42.3.2.jar ``` 2. Download the jar at ## Code[​](#code "Direct link to Code") Add the following to `PostgresqlExample.java` and replace the placeholder with the PostgreSQL URI: ``` import java.sql.Connection; import java.sql.DriverManager; import java.sql.ResultSet; import java.sql.SQLException; import java.sql.Statement; import java.util.Locale; public class PostgresqlExample { public static void main(String[] args) throws ClassNotFoundException { String host, port, databaseName, userName, password; host = port = databaseName = userName = password = null; for (int i = 0; i < args.length - 1; i++) { switch (args[i].toLowerCase(Locale.ROOT)) { case "-host": host = args[++i]; break; case "-username": userName = args[++i]; break; case "-password": password = args[++i]; break; case "-database": databaseName = args[++i]; break; case "-port": port = args[++i]; break; } } // JDBC allows to have nullable username and password if (host == null || port == null || databaseName == null) { System.out.println("Host, port, database information is required"); return; } Class.forName("org.postgresql.Driver"); try (final Connection connection = DriverManager.getConnection("jdbc:postgresql://" + host + ":" + port + "/" + databaseName + "?sslmode=require", userName, password); final Statement statement = connection.createStatement(); final ResultSet resultSet = statement.executeQuery("SELECT version()")) { while (resultSet.next()) { System.out.println("Version: " + resultSet.getString("version")); } } catch (SQLException e) { System.out.println("Connection failure."); e.printStackTrace(); } } } ``` This code creates a PostgreSQL client and opens a connection to the database. Then runs a query checking the database version and prints the response Before running the code, change: * **HOST** to `HOSTNAME` * **PORT**: to `PORT` * **DATABASE** to `DATABASE` * **PASSWORD** to `PASSWORD` To run the code: ``` javac PostgresqlExample.java && java -cp postgresql-42.2.24.jar:. PostgresqlExample -host HOST -port PORT -database DATABASE -username avnadmin -password PASSWORD ``` If the script runs successfully, the outputs should be the PostgreSQL version running in your service like: ``` Version: PostgreSQL PG_VERSION_NUMBER on x86_64-pc-linux-gnu, compiled by gcc, a cdda7373b4 p 9751fce1e6, 64-bit ``` --- # Connect to Aiven for PostgreSQL® with LibreDB Studio Use [LibreDB Studio](https://libredb.org/) to connect to your Aiven for PostgreSQL® service from a browser. LibreDB Studio is an open source SQL client that you host yourself, so a team connects through one URL instead of installing a client on every machine. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Access to the [Aiven Console](https://console.aiven.io/) * At least one running Aiven for PostgreSQL service * LibreDB Studio running on your machine or your network, for example with Docker: ``` docker run -p 3000:3000 -v libredb:/app/data \ -e STORAGE_PROVIDER=sqlite ghcr.io/libredb/libredb-studio:0.16.1 ``` `STORAGE_PROVIDER=sqlite` keeps saved connections on the server instead of in the browser, and the volume keeps them when you replace the container. On the first run, LibreDB Studio prints the admin password to the container log, so read it with `docker logs` before you sign in. ## Get the service URI from the Aiven Console[​](#get-the-service-uri-from-the-aiven-console "Direct link to Get the service URI from the Aiven Console") 1. Log in to the Aiven Console and go to your organization > project > Aiven for PostgreSQL service. 2. On the **Overview** page, go to the **Connection information** section. 3. Copy the **Service URI**. It carries the host, port, user, password, database name, and `sslmode=require`. ## Connect to the service URI from LibreDB Studio[​](#connect-to-the-service-uri-from-libredb-studio "Direct link to Connect to the service URI from LibreDB Studio") 1. Open LibreDB Studio, sign in, and create a connection. 2. Click **Paste URL**, paste the service URI, and click **Parse**. LibreDB Studio fills in the host, port, user, password, and database name, and sets **SSL** to **`REQUIRE`** because the URI carries `sslmode=require`. 3. Click **Test Connection** to verify the settings, and click **Establish Connection** to save the connection. The object browser lists your tables, and the query editor and `EXPLAIN` plans run against the service. ## Connection limits on smaller plans[​](#connection-limits-on-smaller-plans "Direct link to Connection limits on smaller plans") LibreDB Studio keeps a connection pool open while a connection is active, so it uses several of the connections your plan allows. Plans with a low connection limit, such as the free plan with a limit of 20, leave less room for other clients. Check **Connections** on the **Overview** page if several tools connect at the same time. Related pages * [Connect to Aiven for PostgreSQL](/docs/products/postgresql/howto/list-code-samples.md) for more tools you can use for connecting to your service * [LibreDB Studio](https://libredb.org/) * [LibreDB Studio on GitHub](https://github.com/libredb/libredb-studio) --- # Connect to Aiven for PostgreSQL® with NodeJS This example connects to PostgreSQL® service from NodeJS, making use of the `pg` package. ## Variables[​](#variables "Direct link to Variables") These are the placeholders you will need to replace in the code sample: | Variable | Description | | ---------- | ----------------------------------------------------------------------- | | `USER` | PostgreSQL username, from the service overview page | | `PASSWORD` | PostgreSQL password, from the service overview page | | `HOST` | Hostname for PostgreSQL connection, from the service overview page | | `PORT` | Port for PostgreSQL connection, from the service overview page | | `DATABASE` | Database Name for PostgreSQL connection, from the service overview page | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") For this example you will need: * The npm `pg` package: ``` npm install pg --save ``` * [Download CA certificates](/docs/platform/concepts/tls-ssl-certificates.md#download-ca-certificates) from the service overview page, this example assumes it is in a local file called `ca.pem`. ## Code[​](#code "Direct link to Code") Add the following to `index.js` and replace the connection parameters with the ones from the service overview page: ``` const fs = require("fs"); const pg = require("pg"); const config = { user: "USER", password: "PASSWORD", host: "HOST", port: "PORT", database: "DATABASE", ssl: { rejectUnauthorized: true, ca: fs.readFileSync("./ca.pem").toString(), }, }; const client = new pg.Client(config); client.connect(function (err) { if (err) throw err; client.query("SELECT VERSION()", [], function (err, result) { if (err) throw err; console.log(result.rows[0]); client.end(function (err) { if (err) throw err; }); }); }); ``` This code creates a PostgreSQL client and opens a connection to the database. Then runs a query checking the database version and prints the response. To run the code: ``` node index.js ``` If the script runs successfully, the outputs should be the PostgreSQL version running in your service like: ``` PostgreSQL PG_VERSION_NUMBER on x86_64-pc-linux-gnu, compiled by gcc, a 68c5366192 p 6520304dc1, 64-bit ``` --- # Connect to Aiven for PostgreSQL® with pgAdmin [pgAdmin](https://www.pgadmin.org/) is one of the most popular PostgreSQL® clients. Use it to manage and query your database. ## Variables[​](#variables "Direct link to Variables") These are the placeholders you will need to replace in the code sample: | Variable | Description | | ---------- | ----------------------------------------------------------------------- | | `HOSTNAME` | Hostname for PostgreSQL connection, from the service overview page | | `PORT` | Port for PostgreSQL connection, from the service overview page | | `DATABASE` | Database Name for PostgreSQL connection, from the service overview page | | `PASSWORD` | `avnadmin` password, from the service overview page | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") For this example you'll need pgAdmin already installed on your computer, for installation instructions follow the [pgAdmin website](https://www.pgadmin.org/download/) ## Connect to PostgreSQL[​](#connect-to-postgresql "Direct link to Connect to PostgreSQL") 1. Open pgAdmin and click **Create New Server**. 2. In the **General** Tab give the connection a name, for example `MyDatabase`. 3. In the **Connection** tab, set: * **Host name/address** to `HOSTNAME` * **Port**: to `PORT` * **Maintenance database** to `DATABASE` * **Username** to `avnadmin` * **Password** to `PASSWORD` 4. In the **SSL** tab, set **SSL mode** to `Require` 5. Click **Save** tip If you experience a SSL error while connecting, add the service CA certificate as the **Root certificate**. 1. Download the CA Certificate file to your computer. 2. In the pgAdmin connection settings, click the SSL tab and select the CA certificate file you downloaded. Save the settings. Your connection to PostgreSQL should now be opened, with a **Dashboard** page showing activity metrics on your PostgreSQL database. ![Screenshot of a pgAdmin Dashboard window](/docs/assets/images/pg-pgadmin-activity-b8786900052dc31ec92fed5df4802907.png) --- # Connect to Aiven for PostgreSQL® with PHP This example connects to PostgreSQL® service from PHP, making use of the built-in PDO module. ## Variables[​](#variables "Direct link to Variables") These are the placeholders you will need to replace in the code sample: | Variable | Description | | ---------------- | ------------------------------------------------------------- | | `POSTGRESQL_URI` | URL for PostgreSQL connection, from the service overview page | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") For this example you will need: * [Download CA certificates](/docs/platform/concepts/tls-ssl-certificates.md#download-ca-certificates) from the service overview page, this example assumes it is in a local file called `ca.pem`. note Your PHP installation will need to include the [PostgreSQL functions](https://www.php.net/manual/en/ref.pdo-pgsql.php) (most installations will have this already). ## Code[​](#code "Direct link to Code") Add the following to `index.php` and replace the placeholder with the PostgreSQL URI: ``` query("SELECT VERSION()") as $row) { print($row[0]); } ``` This code creates a PostgreSQL client and opens a connection to the database. Then runs a query checking the database version and prints the response note This example replaces the query string parameter to specify `sslmode=verify-ca` to make sure that the SSL certificate is verified, and adds the location of the cert. To run the code: ``` php index.php ``` If the script runs successfully, the outputs should be the PostgreSQL version running in your service like: ``` PostgreSQL PG_VERSION_NUMBER on x86_64-pc-linux-gnu, compiled by gcc, a 68c5366192 p 6520304dc1, 64-bit ``` --- # Connect to Aiven for PostgreSQL® with psql `psql` is a command line tool for PostgreSQL®, useful to manage and query your database. ## Variables[​](#variables "Direct link to Variables") These are the placeholders you will need to replace in the code sample: | Variable | Description | | ---------------- | ------------------------------------------------------------- | | `POSTGRESQL_URI` | URL for PostgreSQL connection, from the service overview page | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") For this example you'll need `psql` already installed on your computer ## Connect to PostgreSQL[​](#connect-to-postgresql "Direct link to Connect to PostgreSQL") From your terminal, execute the following code: ``` psql POSTGRESQL_URI ``` The output should look like the following if the connection is successful: ``` psql (PG_VERSION_NUMBER, server VERSION_NUMBER) SSL connection (protocol: TLSv1.3, cipher: TLS_AES_256_GCM_SHA384, bits: 256, compression: off) Type "help" for help. defaultdb=> ``` To confirm that the connection is working, issue the following code checking the PostgreSQL version: ``` select version(); ``` The result will be similar to the following: ``` version -------------------------------------------------------------------------------------------- PostgreSQL PG_VERSION_NUMBER on x86_64-pc-linux-gnu, compiled by gcc, a 68c5366192 p 6520304dc1, 64-bit (1 row) ``` --- # Connect to Aiven for PostgreSQL® with Python This example connects to a PostgreSQL® service from Python, making use of the `psycopg2` library. ## Variables[​](#variables "Direct link to Variables") These are the placeholders you will need to replace in the code sample: | Variable | Description | | ---------------- | ------------------------------------------------------------- | | `POSTGRESQL_URI` | URL for PostgreSQL connection, from the service overview page | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") For this example you will need: * Python 3.6 or later * The Python `psycopg2` library. You can install this with `pip`: ``` pip install psycopg2 ``` ## Code[​](#code "Direct link to Code") Add the following to `main.py` and replace the placeholders with values for your project: ``` import psycopg2 def main(): conn = psycopg2.connect('POSTGRESQL_URI') query_sql = 'SELECT VERSION()' cur = conn.cursor() cur.execute(query_sql) version = cur.fetchone()[0] print(version) if __name__ == "__main__": main() ``` This code creates a PostgreSQL client and connects to the database. Then runs a query checking the database version and prints the response note By default, the connection string specifies `sslmode=require` which does not verify the CA certificate. A better approach for production would be to change it to `sslmode=verify-ca` and include the certificate. To run the code: ``` python main.py ``` If the script runs successfully, the outputs should be the PostgreSQL version running in your service like: ``` PostgreSQL PG_VERSION_NUMBER on x86_64-pc-linux-gnu, compiled by gcc, a 68c5366192 p 6520304dc1, 64-bit ``` --- # Connect to Aiven for PostgreSQL® with Rivery Configure PostgreSQL® as a connection in [Rivery](https://rivery.io/), a fully managed solution for data ingestion, transformation, orchestration, reverse ETL and more. ## Variables[​](#variables "Direct link to Variables") These are the placeholders you will need to replace from your PostgreSQL® connection information: | Variable | Description | | ---------- | --------------------------------------------------------------------------- | | `HOSTNAME` | Hostname for the PostgreSQL connection, from the service overview page | | `PORT` | Port for the PostgreSQL connection, from the service overview page | | `DATABASE` | Database Name for the PostgreSQL connection, from the service overview page | | `USERNAME` | Username for the PostgreSQL connection | | `PASSWORD` | Password for the above specified username | ## Connect to PostgreSQL®[​](#connect-to-postgresql "Direct link to Connect to PostgreSQL®") 1. In skyvia workspace > **Connections** > **Connector** > **PostgreSQL**. 2. In the **General** Tab give the connection a name, for example `MyDatabase`. 3. In the **Connection** tab, set: * **Host** to `HOSTNAME` * **Port**: to `PORT` * **Database** to `DATABASE` * **User name** to `USERNAME` * **Password** to `PASSWORD` 4. Click **SSL Options** to expand the settings and set: * **SSL** Mode set to `Require` 5. Click **Save**. --- # Connect to Aiven for PostgreSQL® with Skyvia [Skyvia](https://skyvia.com/) is an universal cloud data platform. This example shows how to configure PostgreSQL® as a connection in skyvia. ## Variables[​](#variables "Direct link to Variables") These are the placeholders you will need to replace your PostgreSQL® connection information: | Variable | Description | | ---------- | --------------------------------------------------------------------------- | | `HOSTNAME` | Hostname for the PostgreSQL connection, from the service overview page | | `PORT` | Port for the PostgreSQL connection, from the service overview page | | `DATABASE` | Database Name for the PostgreSQL connection, from the service overview page | | `USERNAME` | Username for the PostgreSQL connection | | `PASSWORD` | Password for the above specified username | ## Connect to PostgreSQL®[​](#connect-to-postgresql "Direct link to Connect to PostgreSQL®") 1. In skyvia workspace > **Connections** > **Connector** > **PostgreSQL**. 2. In the **General** Tab give the connection a name, for example `MyDatabase`. 3. In the **Connection** tab, set: * **Server** to `HOSTNAME` * **Port**: to `PORT` * **User ID** to `USERNAME` * **Password** to `PASSWORD` * **Database** to `DATABASE` 4. Click **Advanced Settings** to expand the settings and set: * **SSL** Mode set to `Require` * In **SSL CA Cert** copy and paste `CA Certificate` from the [Aiven Console](https://console.aiven.io/) * **SSL Cert** and **SSL Key** empty. * **SSL TLS Protocol** to `1.2`. 5. Click **Save Connection**. --- # Connect to Aiven for PostgreSQL® with Zapier [Zapier](https://zapier.com/) is an automation platform that connects your work apps and does repetitive tasks for you. This example shows how to configure PostgreSQL® as a connection in zapier. ## Variables[​](#variables "Direct link to Variables") These are the placeholders you will need to replace in the code sample: | Variable | Description | | ---------- | --------------------------------------------------------------------------- | | `HOSTNAME` | Hostname for the PostgreSQL connection, from the service overview page | | `PORT` | Port for the PostgreSQL connection, from the service overview page | | `DATABASE` | Database Name for the PostgreSQL connection, from the service overview page | | `SCHEMA` | Default `public` schema or a specific schema | | `USERNAME` | Username for the PostgreSQL connection | | `PASSWORD` | Password for the above specified username | ## Connect to PostgreSQL[​](#connect-to-postgresql "Direct link to Connect to PostgreSQL") 1. In skyvia workspace > **Connections** > **Connector** > **PostgreSQL**. 2. In the **General** Tab give the connection a name, for example `MyDatabase`. 3. In the **Connection** tab, set: * **Host** to `HOSTNAME` * **Port**: to `PORT` * **Database** to `DATABASE` * **Schema** to `SCHEMA` * **Username** to `USERNAME` * **Password** to `PASSWORD` 4. Click **Yes, Continue**. --- # Controlled upgrade pipelines for your Aiven for PostgreSQL® service [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Link Aiven for PostgreSQL® services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. Control when your Aiven managed services receive maintenance updates and test maintenance updates in development or staging environments before they reach production. important Controlled upgrade pipeline is a [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-) feature. [Contact Aiven](https://aiven.io/contact) to request access. Aiven performs automatic service maintenance for security fixes, minor software updates, and other platform changes. The controlled upgrade pipeline feature lets you link services of the same type in an ordered sequence to control when each service receives updates. After a maintenance update upgrades a service at the initial pipeline step, you validate that service version before the update proceeds to the service at the next pipeline step. Validating means approving the new version as safe to roll out to the next service. Validation can be manual or automatic after a configurable delay. ## Why use controlled upgrade pipelines[​](#why-use-controlled-upgrade-pipelines "Direct link to Why use controlled upgrade pipelines") Controlled upgrade pipelines prevent production incidents caused by automatic updates reaching production before teams can test the new version in a lower environment. They give you full oversight of the update process: * **Risk mitigation**: Prevents unexpected maintenance updates from breaking your production environment by ensuring they are tested in a non-production setting first. * **Stability**: Keeps destination services (such as production) on a known-good version until you, or the automatic timer, confirm the new version is safe. * **Process control**: Allows platform teams to standardize their deployment and maintenance lifecycle across environments. ## About controlled upgrade pipelines[​](#about-controlled-upgrade-pipelines "Direct link to About controlled upgrade pipelines") ### Upgrade steps[​](#upgrade-steps "Direct link to Upgrade steps") An upgrade step is a pair of services linked by an upgrade constraint: * **Source service**: The service that receives maintenance updates first * **Destination service**: The service that waits for validation before receiving updates Each destination service can have only one source service. A source service can have multiple destination services. ### Upgrade pipelines[​](#upgrade-pipelines "Direct link to Upgrade pipelines") An upgrade pipeline is a chain of upgrade steps that spans multiple environments. For example: * Single chain: development → staging → production * Multiple destinations: development → production-eu and development → production-na ## How validation works[​](#how-validation-works "Direct link to How validation works") When a maintenance update upgrades your source service: 1. The source service receives the update first. 2. Test the updated source service to verify it works as expected. 3. Validate the update manually using the API or CLI, or wait for automatic validation after the configured delay. The default delay is 7 days. 4. After validation, the destination service becomes eligible for the same maintenance update. 5. The destination service receives the update during its next maintenance window. If one source service has multiple destination services, one validation for the source service applies to all connected destination services. ### Validation and maintenance windows[​](#validation-and-maintenance-windows "Direct link to Validation and maintenance windows") Validation and the maintenance window control different things: * **Validation** controls *what* version the destination service upgrades to. * The maintenance window controls *when* the upgrade happens. After you validate an update, or automatic validation applies, the destination service receives the validated version during its next scheduled maintenance window. Validation does not trigger an immediate upgrade outside the maintenance window. Upgrade pipelines add a constraint on what is installed during a maintenance update; they do not change when maintenance runs. Nodes in the destination service maintain the validated version until a newer version is validated, either when you validate it manually or when automatic validation applies after the configured delay. When a node is recycled, it uses the same validated version, not the latest available version. When you create a step, the destination service keeps the newest version that is already validated at that moment. If the destination service is already applying maintenance during step creation, the in-progress target version becomes the initial validated version. warning A powered-off source service cannot receive maintenance updates, so you cannot validate it. If you power off services earlier in the chain, the destination service upgrades regardless. For example, in a development → staging → production chain, if both development and staging are powered off, production upgrades without testing and validation in the earlier environments. Keep services in the chain powered on to preserve the protection that upgrade pipelines provide. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * **Same service type**: You can only link services of the same type. For example, two Aiven for PostgreSQL services. * **Chain length**: The default maximum chain depth is 3 services, which is 2 steps. If you need a longer chain, [contact Aiven](https://aiven.io/contact). * **No cycles**: You cannot create circular dependencies between services. * **Emergency overrides**: Aiven can apply critical security or stability fixes to a destination service before explicit validation. * **Supported services**: This feature supports all Aiven service types except Aiven for Apache Flink® and Aiven for MySQL. * **Automatic maintenance updates only**: Pipelines apply to automatic maintenance updates, such as minor service version updates and node image updates. Major version upgrades, for example Aiven for PostgreSQL® 15 to 16, require manual action and are not promoted automatically through the pipeline. * **No permanent blocking**: You cannot prevent an update indefinitely. Automatic validation applies after the configured delay, up to the maximum delay. * **No validation rollback**: You cannot undo a validation after it is recorded. ## Use controlled upgrade pipelines[​](#use-controlled-upgrade-pipelines "Direct link to Use controlled upgrade pipelines") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") To use controlled upgrade pipelines, you need the following: * The feature enabled by Aiven ([Limited availability](/docs/platform/concepts/service-and-feature-releases.md)) * Dev tool of your choice: * [Aiven CLI](/docs/tools/cli.md) Install the latest version of the Aiven CLI to access the `upgrade-pipeline` commands. * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) * Set `PROVIDER_AIVEN_ENABLE_BETA=true` before running Terraform. * See the [resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) for full schema, import format, and lifecycle behavior. * [Aiven Operator for Kubernetes](/docs/tools/kubernetes.md) Install the operator and create an Aiven token secret named `aiven-token` that the operator uses to authenticate against the Aiven API. * Write access to the source and destination projects * At least two services of the same type (for example, two Aiven for PostgreSQL® services) * Services can be in different projects in the same organization ### Set up an upgrade pipeline[​](#set-up-an-upgrade-pipeline "Direct link to Set up an upgrade pipeline") Use the Aiven CLI or API to create upgrade steps between your services. note The `upgrade-pipeline` CLI commands require Aiven CLI version 4.x or later. Command names and parameters may change before general availability. #### Create an upgrade step[​](#create-an-upgrade-step "Direct link to Create an upgrade step") Create a step to link a source service and a destination service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ [--source-project SOURCE_PROJECT] SOURCE_SERVICE \ [--destination-project DESTINATION_PROJECT] DESTINATION_SERVICE \ [--auto-validation-delay-days DAYS] ``` **Options** * `--organization-id` is required. * `--source-project` and `--destination-project` are optional. If you omit either project option, Aiven CLI uses the current default project set with `avn project switch`. * `--auto-validation-delay-days` is optional. Defaults to 7 days if not specified. ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "SOURCE_PROJECT_NAME", "source_service_name": "SOURCE_SERVICE_NAME", "destination_project_name": "DESTINATION_PROJECT_NAME", "destination_service_name": "DESTINATION_SERVICE_NAME", "auto_validation_delay_days": 7 }' ``` Use the [`aiven_upgrade_step`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) resource: ``` resource "aiven_upgrade_step" "example" { organization_id = "ORGANIZATION_ID" source_project_name = "SOURCE_PROJECT_NAME" source_service_name = "SOURCE_SERVICE_NAME" destination_project_name = "DESTINATION_PROJECT_NAME" destination_service_name = "DESTINATION_SERVICE_NAME" auto_validation_delay_days = 7 } ``` Apply an `UpgradePipelineStep` manifest with `kubectl`: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: upgrade-step-sample spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: SOURCE_PROJECT_NAME sourceServiceName: SOURCE_SERVICE_NAME destinationProjectName: DESTINATION_PROJECT_NAME destinationServiceName: DESTINATION_SERVICE_NAME autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable after the resource is created. Parameters: * `source_project_name`: Name of the project containing the source service * `source_service_name`: Name of the source service * `destination_project_name`: Name of the project containing the destination service * `destination_service_name`: Name of the destination service * `auto_validation_delay_days`: Optional. Number of days before automatic validation. The value must be at least `1`. The default is 7 days. The maximum delay you can configure is 30 days. #### List upgrade steps[​](#list-upgrade-steps "Direct link to List upgrade steps") View all upgrade steps you have access to: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step list --organization-id ORGANIZATION_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) To list managed upgrade steps, use: ``` terraform state list 'aiven_upgrade_step.*' terraform state show 'aiven_upgrade_step.example' ``` List `UpgradePipelineStep` resources in the current namespace: ``` kubectl get upgradepipelinesteps ``` #### View a specific step[​](#view-a-specific-step "Direct link to View a specific step") Get details about a specific upgrade step, including the last validation: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step get \ --organization-id ORGANIZATION_ID \ STEP_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` terraform state show aiven_upgrade_step.example ``` Show the manifest and full status, including `id`, `conditions`, and `lastValidation`: ``` kubectl describe upgradepipelinestep RESOURCE_NAME kubectl get upgradepipelinestep RESOURCE_NAME -o yaml ``` The step details include `last_validation` values such as `validated_at`, `validated_by_user`, and `comment` when validation exists (available through the API). ### Validate an upgrade[​](#validate-an-upgrade "Direct link to Validate an upgrade") After testing your source service with the new update, validate the version to allow the destination service to receive the same update. #### Manual validation[​](#manual-validation "Direct link to Manual validation") Validate the current version of your source service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step validate-for-service \ --project SOURCE_PROJECT \ SERVICE_NAME \ [--comment "COMMENT"] ``` `--comment` is optional. Use it to record a note about the validation, for example `"Tested and verified in development"`. ``` curl -X POST https://api.aiven.io/v1/project/SOURCE_PROJECT/service/SOURCE_SERVICE/upgrade-validation \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "comment": "Tested and verified in development" }' ``` Terraform manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. The operator manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. #### Automatic validation[​](#automatic-validation "Direct link to Automatic validation") If you do not manually validate an update, the system automatically validates the source service version after the configured delay. Auto-validation starts from when the source service receives the update. ### Manage upgrade steps[​](#manage-upgrade-steps "Direct link to Manage upgrade steps") #### Update a step[​](#update-a-step "Direct link to Update a step") Modify the automatic validation delay for an existing step: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step update \ --organization-id ORGANIZATION_ID \ --auto-validation-delay-days 14 \ STEP_ID ``` ``` curl -X PATCH https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "auto_validation_delay_days": 14 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` resource "aiven_upgrade_step" "example" { # ...required fields... auto_validation_delay_days = 14 # Updated from 7 to 14 } ``` Apply the changes: ``` terraform plan terraform apply ``` Edit `autoValidationDelayDays` in your manifest and re-apply: ``` spec: autoValidationDelayDays: 14 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable. To change them, delete the resource and create a new one. #### Delete a step[​](#delete-a-step "Direct link to Delete a step") Remove an upgrade step to allow the destination service to receive updates independently: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step delete --organization-id ORGANIZATION_ID STEP_ID ``` Find `STEP_ID` from the upgrade step list command. ``` curl -X DELETE https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) Remove the resource from configuration and apply, or destroy it directly: ``` terraform apply terraform destroy -target=aiven_upgrade_step.example ``` Delete the `UpgradePipelineStep` resource: ``` kubectl delete upgradepipelinestep RESOURCE_NAME ``` Deleting a step removes all associated validations. ### Example: Three-environment pipeline[​](#example-three-environment-pipeline "Direct link to Example: Three-environment pipeline") Create a pipeline that promotes updates from development to staging to production: * CLI * API * Terraform * Kubernetes 1. Create a step from development to staging: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project dev-project \ --destination-project staging-project \ --auto-validation-delay-days 3 \ pg-dev pg-staging ``` 2. Create a step from staging to production: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project staging-project \ --destination-project prod-project \ --auto-validation-delay-days 7 \ pg-staging pg-prod ``` 1) Create a step from development to staging: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "dev-project", "source_service_name": "pg-dev", "destination_project_name": "staging-project", "destination_service_name": "pg-staging", "auto_validation_delay_days": 3 }' ``` 2) Create a step from staging to production: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "staging-project", "source_service_name": "pg-staging", "destination_project_name": "prod-project", "destination_service_name": "pg-prod", "auto_validation_delay_days": 7 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` # Step 1: Development → Staging resource "aiven_upgrade_step" "dev_to_staging" { organization_id = "ORGANIZATION_ID" source_project_name = "dev-project" source_service_name = "pg-dev" destination_project_name = "staging-project" destination_service_name = "pg-staging" auto_validation_delay_days = 3 } # Step 2: Staging → Production resource "aiven_upgrade_step" "staging_to_prod" { organization_id = "ORGANIZATION_ID" source_project_name = "staging-project" source_service_name = "pg-staging" destination_project_name = "prod-project" destination_service_name = "pg-prod" auto_validation_delay_days = 7 } ``` Apply the configuration: ``` export PROVIDER_AIVEN_ENABLE_BETA=true terraform init terraform plan terraform apply ``` Define both steps in a single manifest and apply it: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: dev-to-staging spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: dev-project sourceServiceName: pg-dev destinationProjectName: staging-project destinationServiceName: pg-staging autoValidationDelayDays: 3 --- apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: staging-to-prod spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: staging-project sourceServiceName: pg-staging destinationProjectName: prod-project destinationServiceName: pg-prod autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-pipeline.yaml ``` When a maintenance update arrives: 1. The development service receives the update. 2. After testing, validate the development version or wait 3 days for auto-validation. 3. The staging service receives the update during its next maintenance window. 4. After testing, validate the staging version or wait 7 days for auto-validation. 5. The production service receives the update during its next maintenance window. Related pages * [Maintenance and updates for your Aiven for PostgreSQL® service](/docs/products/postgresql/howto/maintenance-updates.md) * [Change the service plan](/docs/products/postgresql/howto/change-service-plan.md) * [Service and feature releases](/docs/platform/concepts/service-and-feature-releases.md) * [Aiven CLI](/docs/tools/cli.md) --- # Create PostgreSQL® databases Once you've created your Aiven for PostgreSQL® service, you can add additional databases, whether for security purposes or to isolate your data per application. To create a PostgreSQL® database: 1. In the [Aiven Console](https://console.aiven.io/), on the **Services** page, click the Aiven for PostgreSQL service name for which to create a database. 2. In your service's page, in the **Connect** section, click **Databases**. 3. In the **Databases** page, select **Create database**. 4. In the **Create a database** window, enter a name for your database into the **Name** field, and select **Add database**. The new database is visible immediately. tip You can also use the [Aiven client](/docs/tools/cli/service/database.md#avn-service-database-create) or the [PostgreSQL client](/docs/products/postgresql/howto/connect-psql.md) to create your database from the CLI. --- # Create manual PostgreSQL® backups with pg\_dump Aiven provides [fully automated backup management for PostgreSQL](/docs/products/postgresql/concepts/pg-backups.md). All backups are encrypted with service-specific keys, and point-in-time recovery is supported to allow recovering the system to any point within the backup window. Aiven stores the backups to the closest available cloud storage to enhance restore speed. Perform a backup of your database using the standard PostgreSQL `pg_dump` command. See the [pg\_dump docs](https://www.postgresql.org/docs/current/app-pgdump.html), but a typical command looks like: ``` pg_dump 'POSTGRESQL_URI' \ -f backup_folder \ -j 2 \ -F directory ``` Where: * `POSTGRESQL_URI` is the URL for PostgreSQL connection, from the service overview page. This command creates a backup in `directory` format (ready for use with `pg_restore`) using 2 concurrent jobs and storing the output to a folder called `backup_folder`. tip `pg_dump` can be run against any **standby** node, using the *Replica URI* from the Aiven Console. Creating more jobs via the `-j` option can be useful, since it may not be a problem to add extra load to the standby node. --- # Create and use Aiven for PostgreSQL® read-only replicas Use Aiven for PostgreSQL® read-only replicas to reduce the load on the primary server and optimize query response times across different geographical locations. You can run read-only queries against Aiven for PostgreSQL read-only replicas. These replicas can be hosted in different regions or on different cloud providers. note Services that have standby nodes available in a high availability setup support read-only queries to reduce the effect of slow queries on the primary node. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Running Aiven for PostgreSQL service * Access the [Aiven Console](https://console.aiven.io/) * Optionally, for creating a read-only replica programmatically: * [Aiven Operator for Kubernetes®](https://aiven.io/docs/tools/kubernetes) * [Aiven API](https://api.aiven.io/doc/) * [Aiven Provider for Terraform](https://registry.terraform.io/providers/aiven/aiven/latest/docs) important Creating a read-only replica programmatically, you can only use the Startup plan for the replica. ## Create a replica[​](#create-a-replica "Direct link to Create a replica") * Console * API * Terraform * Kubernetes 1. On the **Overview** page of your service, go to the **Read replica** section. 2. Click **Create replica**. 3. Enter a name for the replica. 4. Select a **Cloud**. 5. Select a **Plan**. 6. Click **Create**. The read-only replica is created and added to the list of services in your project. The **Overview** page of the replica indicates the name of the primary service for the replica. Use the [ServiceCreate](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) endpoint and configure the `service_integrations` object so that: * `integration_type` is set to `read_replica`. * `source_service` is set as needed. Use [the `aiven_service_integration` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration) resource. Set `integration_type` to `read_replica` and `source_service` as needed. [Create an Aiven for PostgreSQL read-only replica with K8s](https://aiven.github.io/aiven-operator/examples/postgresql.html#create-a-postgresql-read-only-replica). Read-only replicas can be manually promoted to become the primary database. For more high availability and failover scenarios, check the [related documentation](/docs/products/postgresql/concepts/high-availability.md). tip You can promote a read-replica to master using the API endpoint to [delete the service integration](https://api.aiven.io/doc/#operation/ServiceIntegrationDelete) and passing the `integration_id` of the replica service. After deleting the integration that comes with `integration_type` of value `read_replica`, the service is no longer a read-replica and, hence, becomes the master. ## Use a replica[​](#use-a-replica "Direct link to Use a replica") To use a read only replica: 1. Log in to the Aiven Console and select your Aiven for PostgreSQL service. 2. In the **Overview** page, copy the **Replica URI** an use it to connect via `psql`: ``` psql POSTGRESQL_REPLICA_URI ``` ## Identify replica status[​](#identify-replica-status "Direct link to Identify replica status") To check whether you are connected to a primary or replica node, run the following command within a `psql` terminal already connected to a database: ``` SELECT * FROM pg_is_in_recovery(); ``` If the above command returns `TRUE` if you are connected to the replica, and `FALSE` if you are connected to the primary server. warning Aiven for PostgreSQL uses asynchronous replication and so a small lag is expected. When running an `INSERT` operation on the primary node, a minimal delay (usually less than a second) can be expected for the change to be propagated to the replica and to be visible there. ## Read-replica for disaster recovery[​](#read-replica-for-disaster-recovery "Direct link to Read-replica for disaster recovery") High availability enables data distribution across availability zones within a single region. To do this without a default multi-region service with node allocation spanning multiple regions: 1. Establish a high-availability Aiven for PostgreSQL service within a single region. 2. Configure a remote read-only replica in a different region or even on an alternate cloud platform. As a result, you introduce an additional node in the distinct region/cloud. Since this node does not work as a hot standby node, you might want to promote it manually to the primary role, which makes it operate as an independent standalone service. --- # Data API for Aiven for PostgreSQL® [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Data API turns your Aiven for PostgreSQL® database into a backend by exposing its tables as secure REST endpoints, without backend code. note Data API is a [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) feature. To request access, [contact Aiven](https://aiven.io/contact). To access Data API, open your Aiven for PostgreSQL® service in the [Aiven Console](https://console.aiven.io) and click **Data** > **Data API**. It auto-generates an API directly from your database schema, powered by [PostgREST](https://postgrest.org). ## What Data API offers[​](#what-data-api-offers "Direct link to What Data API offers") Data API provides the following: * **Instant REST endpoints**: Expose a database as REST endpoints from the Aiven Console, without writing or hosting an API server. * **Schema-driven endpoints**: Each table gets endpoints for the `GET`, `POST`, `PATCH`, and `DELETE` methods, based on your database schema. * **Authentication with your identity provider**: Authenticate requests with the JSON Web Tokens (JWTs) issued by your own identity provider (IdP) and verified against your JWKS URL. * **Authorization with PostgreSQL roles**: Control access with standard PostgreSQL roles and table privileges. ## How it works[​](#how-it-works "Direct link to How it works") When you enable Data API for a database, Aiven deploys a dedicated [Aiven Runtime application](/docs/products/runtime.md) that runs PostgREST and connects it to the selected database. By default, the application runs in the same cloud and region as your PostgreSQL service, but you can choose a different one when you set up Data API. Aiven Runtime isn't available in all clouds and regions yet, so the cloud and region you can choose from may be more limited than for your PostgreSQL service. PostgREST reads the database schema and publishes a REST endpoint for each table. By default, endpoints are published for the `public` schema. To access tables in other schemas, include the `Accept-Profile` header with the schema name in your requests. Clients call these endpoints over HTTPS and authenticate with a bearer token. Each database that you expose runs as an independent Aiven Runtime application with its own status and base URL. You can enable Data API for more than one database in the same service. ## Limitations[​](#limitations "Direct link to Limitations") * Each Data API serves one database. To expose more databases, enable Data API for each one separately. * Each Data API uses a single identity provider, set by one JWKS URL. Multiple identity providers per service aren't supported. * Endpoints reflect the database schema captured when you enable Data API. They don't refresh automatically when the schema changes, but you can refresh the schema cache from the Aiven Console. * Each Data API runs as a dedicated Aiven Runtime application that is billed separately from your PostgreSQL service. ## Related pages[​](#related-pages "Direct link to Related pages") ## [Enable Data API](/docs/products/postgresql/howto/data-api/get-started.md) [Expose an Aiven for PostgreSQL database as REST endpoints.](/docs/products/postgresql/howto/data-api/get-started.md) ## [Authentication](/docs/products/postgresql/howto/data-api/authentication.md) [Authenticate Data API requests with JWTs from your own identity provider, and authorize with PostgreSQL roles.](/docs/products/postgresql/howto/data-api/authentication.md) ## [Use endpoints](/docs/products/postgresql/howto/data-api/use-endpoints.md) [Find your API URL and call your database over HTTPS with bearer token authentication.](/docs/products/postgresql/howto/data-api/use-endpoints.md) ## [Manage Data API](/docs/products/postgresql/howto/data-api/manage.md) [Check status, expose more databases, and remove the Data API.](/docs/products/postgresql/howto/data-api/manage.md) --- # Configure authentication for Aiven for PostgreSQL® Data API [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Data API authenticates every request with a bearer token in the `Authorization` header. The token is a JWT issued by your own identity provider (IdP) and verified against your JWKS URL. The token carries a role, and your Aiven for PostgreSQL® database enforces that role's privileges. note Data API is a [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) feature. ## Authenticate with your identity provider[​](#authenticate-with-your-identity-provider "Direct link to Authenticate with your identity provider") End users authenticate with the JWTs issued by your IdP. You provide your JWKS URL when you [enable Data API](/docs/products/postgresql/howto/data-api/get-started.md), and Data API verifies each token against the public keys at that URL. note Data API uses one identity provider per database, set by a single JWKS URL. You can update the JWKS URL or audience later from the **Data API** page in the Aiven Console. ### Configure the JWKS URL[​](#configure-the-jwks-url "Direct link to Configure the JWKS URL") The *JSON Web Key Set (JWKS) URL* is the endpoint where your IdP publishes the public keys used to verify token signatures. The URL is required and must use HTTPS. The URL format depends on your IdP. The following are common patterns: * **Auth0**: `https://TENANT_NAME.us.auth0.com/.well-known/jwks.json` * **Okta**: `https://OKTA_DOMAIN/oauth2/default/v1/keys` * **Microsoft Entra ID**: `https://login.microsoftonline.com/TENANT_ID/discovery/v2.0/keys` Replace `TENANT_NAME`, `OKTA_DOMAIN`, and `TENANT_ID` with the values from your IdP. For the exact URL, see your IdP's documentation. When you enable Data API or update the JWKS URL, the Aiven Console fetches the URL and checks that it returns a valid JWKS document: an HTTP 200 response with a JSON body that contains a non-empty `keys` array, where each key has a `kty` field. If the check fails, the console rejects the value and shows the reason, for example that the endpoint is unreachable or returned an error. note This check confirms that the URL is reachable and returns a well-formed JWKS document when you save it. It doesn't detect a JWKS URL that becomes unreachable later, or one that's reachable but serves the wrong IdP's keys; either causes requests to fail at runtime. Because Data API reads the keys from the JWKS URL, key rotation is automatic. When your IdP rotates its signing keys, Data API picks up the new keys from the same URL. You don't need to update any configuration in the Aiven Console. Data API refreshes the keys periodically, about every 12 hours, so allow time for new keys to take effect and keep the previous keys valid during the overlap. ### Configure the audience[​](#configure-the-audience "Direct link to Configure the audience") The *audience* identifies the intended recipient of a token, such as a specific API or tenant. The audience is optional. If you set it, use the same value in your IdP and in the **Audience** field when you enable Data API; Data API then rejects any token whose `aud` claim doesn't match. If you leave it blank, Data API doesn't check the `aud` claim. ## Authorize requests with PostgreSQL roles[​](#authorize-requests-with-postgresql-roles "Direct link to Authorize requests with PostgreSQL roles") Data API uses standard PostgreSQL roles and table privileges for authorization. The token carries a `role` claim that names the role to use, and PostgreSQL enforces that role's privileges. ### Roles Data API creates automatically[​](#roles-data-api-creates-automatically "Direct link to Roles Data API creates automatically") When you enable Data API for a database, Aiven creates two PostgreSQL roles for it. You don't need to create either role yourself: * **`postgrest_authenticator`**: The role Data API uses to connect to your database, instead of your service's admin user. It can only assume roles that you explicitly grant to it, so a request can never access more than what you've granted. * **`web_anon`**: The default role for requests whose token doesn't include a `role` claim. It has no privileges. note If a token doesn't include a `role` claim, the request runs as `web_anon`. Include a `role` claim in every token that needs to access data. ### Create a role and grant privileges[​](#create-a-role-and-grant-privileges "Direct link to Create a role and grant privileges") Connect to your database and create a role with the privileges to expose, then grant it to `postgrest_authenticator` so Data API can assume it: ``` -- Create the role CREATE ROLE api_worker NOLOGIN; -- Grant schema access GRANT USAGE ON SCHEMA public TO api_worker; -- Grant table and sequence privileges GRANT SELECT, INSERT, UPDATE, DELETE ON ALL TABLES IN SCHEMA public TO api_worker; GRANT USAGE, SELECT ON ALL SEQUENCES IN SCHEMA public TO api_worker; -- Link the role to the PostgREST authenticator GRANT api_worker TO postgrest_authenticator; ``` Adjust the granted privileges to match what each role should be able to do. A role with `SELECT` only can read data, while a leaked token for that role can't modify it. ### Add the role to your IdP tokens[​](#add-the-role-to-your-idp-tokens "Direct link to Add the role to your IdP tokens") Configure your IdP to include a `role` claim in its tokens, set to the PostgreSQL role name, such as `api_worker`. The following example adds the claim in Auth0: 1. In Auth0, go to **Actions** > **Library**, then click **Create Action** > **Build from Scratch**. 2. Name the action, set the trigger to **Machine to Machine**, and add the following code: ``` exports.onExecuteCredentialsExchange = async (event, api) => { // Replace 'api_worker' with the name of your PostgreSQL role api.accessToken.setCustomClaim('role', 'api_worker'); }; ``` 3. Click **Save Draft**, then **Deploy**. 4. Go to **Actions** > **Triggers**, click **Credentials Exchange**, and add the action to the flow between **Start** and **Complete**. 5. Click **Apply**. When requesting a token, include the audience parameter so the IdP issues a token valid for Data API. tip Grant each role only the privileges it needs. The token controls which role runs the query, and PostgreSQL enforces the privileges of that role. --- # Enable Aiven for PostgreSQL® Data API [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Enable Data API to expose a database in your Aiven for PostgreSQL® service as REST endpoints. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") To enable Data API, you need the following: * [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) access to Data API. To request access, [contact Aiven](https://aiven.io/contact). * [Aiven Runtime](/docs/products/runtime.md) enabled for your project, since Data API deploys as a Runtime application. If it isn't, the Aiven Console shows **Data API requires Aiven Runtime**. * Data API available for your service's plan and cloud. If it isn't, the Aiven Console shows **The data API is not available for your service**. * The `project:services:write` permission. * An identity provider (IdP) that issues JWTs and publishes a [JWKS URL](/docs/products/postgresql/howto/data-api/authentication.md) over HTTPS. Auth0, Okta, and Microsoft Entra ID are common options. ## Enable Data API for a database[​](#enable-data-api-for-a-database "Direct link to Enable Data API for a database") 1. In the [Aiven Console](https://console.aiven.io/login), open your Aiven for PostgreSQL service. 2. Click **Data** > **Data API**. 3. In the **Database** list, select the database to expose. 4. Click **Set up API**. 5. In the **Data API for \[database]** dialog, configure the following: * Under **Identity provider**: * **JWKS URL**: Enter the HTTPS URL where your IdP publishes its public keys. * **Audience** (optional): Enter the API identifier configured in your IdP. * If Aiven has a recommendation for this service, choose a deployment mode under **Settings**: * **Recommended** (default): Deploys the underlying [Aiven Runtime application](/docs/products/runtime.md) on the cheapest available paid plan, using the same cloud and region as your PostgreSQL service when possible, or the nearest available region otherwise. * **Custom**: Configure your own cloud, region, and plan under **Cloud and plan**. * Under **Cloud and plan**, shown when there's no recommendation or you choose **Custom**: * **Cloud**: Defaults to the same cloud and region as your PostgreSQL service. You can select a different cloud and region that supports the Aiven Runtime application. * **Plan**: Select a plan for the Aiven Runtime application. Free-tier plans aren't available for Data API, so choose a paid plan. 6. Review the **Summary** panel on the right, which shows the cloud, plan, and estimated monthly price for the Aiven Runtime application. 7. Click **Confirm and deploy**. If the cloud, region, and plan you select don't support the Aiven Runtime application, Data API shows an error message so that you can pick a different combination. Data API starts deploying and the **Status** shows **Deploying**. When the app is healthy, the status changes to **Running** and the endpoints become available. While the service is still being provisioned, setup is unavailable and the Aiven Console shows **Set up your data API** with a note that the service is still being provisioned. For details on the JWKS URL and audience fields, see [Configure authentication](/docs/products/postgresql/howto/data-api/authentication.md). ## Next steps[​](#next-steps "Direct link to Next steps") * [Configure authentication](/docs/products/postgresql/howto/data-api/authentication.md) and authorization for the Data API. * [Call the endpoints](/docs/products/postgresql/howto/data-api/use-endpoints.md) with code snippets. * [Manage your Data API](/docs/products/postgresql/howto/data-api/manage.md), including exposing more databases and removing the API. --- # Manage Aiven for PostgreSQL® Data API [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) After you enable Data API, you can monitor it, expose more databases, and remove it for a database you no longer need. note Data API is a [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) feature. To manage Data API, open your Aiven for PostgreSQL® service in the [Aiven Console](https://console.aiven.io/login) and click **Data** > **Data API**. ## Check the status[​](#check-the-status "Direct link to Check the status") Each database that you expose runs as an independent [Aiven Runtime application](/docs/products/runtime.md). The **Data API** page shows a **Status** chip for each database: * **Deploying**: The application is deploying or applying a change. **API URL** shows **API building...**, and **Refresh cache** and the edit icons next to **JWKS URL** and **Audience** aren't available yet. * **Running**: The application is healthy and serving requests. **API URL** shows the real endpoint, and **Refresh cache** and the edit icons become available. * **Error**: The deployment failed. An error message appears at the top of the page. For next steps, see [Troubleshooting](#troubleshooting). * **Disabled**: The underlying application is powered off. A warning message appears at the top of the page. For next steps, see [Troubleshooting](#troubleshooting). ## View the underlying Aiven Runtime app[​](#view-the-underlying-aiven-runtime-app "Direct link to View the underlying Aiven Runtime app") After enabling Data API, the **Data API** page shows a **Runtime application** row with a link to the dedicated app running PostgREST. Aiven deploys the app in the cloud and region you chose when you set up Data API, and bills it separately. Aiven Runtime isn't available in all clouds and regions yet. You can also find the app in your project's **Runtime** list, tagged **Data API** for easy identification. ## Expose more databases[​](#expose-more-databases "Direct link to Expose more databases") A Data API serves one database. To expose another database in the same service, select it in the **Database** list and set up Data API for it. Each database keeps its own status, API URL, and authentication settings. ## Rotate identity provider keys[​](#rotate-identity-provider-keys "Direct link to Rotate identity provider keys") Key rotation is automatic. Data API reads your IdP's public keys from the JWKS URL and picks up rotated keys from the same URL, refreshing them about every 12 hours by default. For more information, see [Configure the JWKS URL](/docs/products/postgresql/howto/data-api/authentication.md#configure-the-jwks-url). ## Change authentication settings[​](#change-authentication-settings "Direct link to Change authentication settings") To change the JWKS URL or audience, open the **Data API** page and select the database. Next to **JWKS URL** or **Audience**, click the edit icon, enter the new value, and save. If you change the JWKS URL, the console [validates it](/docs/products/postgresql/howto/data-api/authentication.md#configure-the-jwks-url) before saving. A confirmation message confirms the update. You don't need to remove Data API to update these settings, but the edit icons are available only while the application is running. ## Refresh the schema cache[​](#refresh-the-schema-cache "Direct link to Refresh the schema cache") Endpoints reflect the database schema captured when you enable Data API. They don't refresh automatically when the schema changes. To pick up new or changed tables, click **Refresh cache** on the **Data API** page. **Refresh cache** is available only while the application is running. A confirmation message confirms the refresh. Refreshing updates the PostgREST schema cache without restarting the service. ## Remove Data API[​](#remove-data-api "Direct link to Remove Data API") Remove Data API to stop serving endpoints for a database. On the **Data API** page, click the database, then click **Remove Data API**. In the **Delete Data API** confirmation dialog, click **Delete**. The endpoints stop responding after the underlying Aiven Runtime application is deleted. Removing Data API doesn't change the data in your database, but it permanently deletes the Data API configuration for that database. To turn Data API back on for the same database, set up the JWKS URL and audience again. note If you delete the PostgreSQL service, Aiven also deletes all associated Data API apps. The apps are no longer accessible from the **Runtime** list or anywhere else. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") ### Data API is not available for the service[​](#data-api-is-not-available-for-the-service "Direct link to Data API is not available for the service") Data API must be available for your service's plan and cloud. If it isn't, the Aiven Console shows **The data API is not available for your service**. ### The service is still being provisioned[​](#the-service-is-still-being-provisioned "Direct link to The service is still being provisioned") Setup is unavailable while the service is still being provisioned. Wait until the service is **Running**, then enable Data API. ### The deployment fails[​](#the-deployment-fails "Direct link to The deployment fails") If the deployment fails, the setup dialog shows an error message and you can try again without losing your entered values. If the issue persists, confirm that the service meets the [prerequisites](/docs/products/postgresql/howto/data-api/get-started.md#prerequisites), then remove and re-enable Data API. ### The selected cloud, region, or plan isn't available[​](#the-selected-cloud-region-or-plan-isnt-available "Direct link to The selected cloud, region, or plan isn't available") If you select a cloud, region, or plan that doesn't support the Aiven Runtime application, the setup dialog shows an error message instead of failing partway through. Select a different cloud, region, or plan and try again. ### The underlying application is powered off[​](#the-underlying-application-is-powered-off "Direct link to The underlying application is powered off") If someone powers off the [Aiven Runtime application](/docs/products/runtime.md) that runs your Data API from the **Runtime** list, the **Data API** page shows a warning message. Click **Go to app to power it on** in the warning, then power on the application. Data API resumes once the application is running again. ### Endpoints don't reflect schema changes[​](#endpoints-dont-reflect-schema-changes "Direct link to Endpoints don't reflect schema changes") Endpoints reflect the database schema captured when you enabled Data API, and don't refresh automatically when the schema changes. To pick up new or changed tables, click **Refresh cache** on the **Data API** page. For more information, see [Refresh the schema cache](#refresh-the-schema-cache). --- # Call the Aiven for PostgreSQL® Data API endpoints [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) After you enable Data API for a database, you can find your API URL and call your endpoints over HTTPS. note Data API is a [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) feature. ## Find the base URL[​](#find-the-base-url "Direct link to Find the base URL") 1. In the [Aiven Console](https://console.aiven.io/login), open your Aiven for PostgreSQL® service. 2. Click **Data** > **Data API**. 3. Select the database with Data API enabled. The **Data API** page shows the **API URL** for the database. All endpoints are relative to the API URL. For example, a `products` table is available at the `/products` path under the API URL. ## Call an endpoint[​](#call-an-endpoint "Direct link to Call an endpoint") Send the bearer token in the `Authorization` header. The following examples use `REST_API_BASE_URL` for the API URL and `TOKEN` for the JWT issued by your IdP. Read rows from the `products` table and select specific columns: ``` curl "https://REST_API_BASE_URL/products?select=id,name,price" \ -H "Authorization: Bearer TOKEN" ``` Insert a row: ``` curl -X POST "https://REST_API_BASE_URL/products" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer TOKEN" \ -d '{"name": "Notebook", "price": 9.99}' ``` Update a row that matches a filter: ``` curl -X PATCH "https://REST_API_BASE_URL/products?id=eq.1" \ -H "Content-Type: application/json" \ -H "Authorization: Bearer TOKEN" \ -d '{"price": 12.99}' ``` For the full query syntax, including filtering, ordering, and pagination, see the [PostgREST documentation](https://postgrest.org/en/stable/references/api.html). --- # Connect two PostgreSQL® services via datasource integration There are two types of datasource integrations you can use with Aiven for PostgreSQL®: * [Aiven for Grafana®](/docs/products/postgresql/howto/visualize-grafana.md) * Another Aiven for PostgreSQL® service If you are connecting two PostgreSQL® services together, perhaps to [query across them](/docs/products/postgresql/howto/use-dblink-extension.md), but still want to have a restricted IP allow-list, use the `IP Allow-List` service integration. Whenever a service node needs to be recycled, for example, for maintenance, a new node is created with a new IP address. As the new IP address cannot be predicted, to maintain a connection between two PostgreSQL services your choices are either to have a very broad IP allow-list (which might be acceptable in the private IP-range of a project VPC) or to use the `IP Allow-List` service integration to dynamically create an IP allow-list entry for the other PostgreSQL service. ## Integrate two PostgreSQL services[​](#integrate-two-postgresql-services "Direct link to Integrate two PostgreSQL services") 1. On the **Overview** page for your Aiven for PostgreSQL service, go to **Manage integrations** and choose the **IP Allow-List** option. 2. Choose either a new or existing Aiven for PostgreSQL service. Now your Aiven for PostgreSQL allows IP traffic from the chosen PostgreSQL service, regardless of what its IP address is. --- # Scale disk storage automatically for your Aiven for PostgreSQL® service Automatically increase the disk storage of your Aiven for PostgreSQL® service when it's running out of space, instead of resizing it manually. Use the Aiven Autoscaler to automatically increase the storage capacity of a service disk when it's running out of space. Disk autoscaler only increases storage, it doesn't scale storage down. ## Why use disk autoscaling[​](#why-use-disk-autoscaling "Direct link to Why use disk autoscaling") * **Cost efficiency**: Start with a regular-sized disk and let Aiven scale it up only when needed, without the risk of running out of disk space. * **Resiliency**: Avoid a service becoming non-functional because it ran out of disk space, including during unexpected spikes in demand. ## How it works[​](#how-it-works "Direct link to How it works") 1. You create an autoscaler integration endpoint in your project, setting the maximum total disk size to allow. 2. You enable an autoscaler integration for your service using that endpoint. 3. Aiven monitors the disk space usage of your service. 4. When disk usage reaches the threshold for your service type, Aiven increases the available storage by at least 10%, using the current used space as a baseline. note The exact increase depends on the service type and cloud provider. Some providers enforce a minimum increase of 10 GB. Autoscale thresholds per service type The threshold that triggers disk autoscaling is a percentage of the available disk storage capacity: * Aiven for OpenSearch®: 75% of the available disk storage capacity * All other supported service types: 85% of the available disk storage capacity 5. The disk increase is recorded in the project event log, and you receive a notification about the added disk space. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * **Maximum storage**: The maximum storage that the autoscaler can allocate for your service is limited by both the maximum disk size set on the autoscaler endpoint and the maximum disk storage supported for your service plan. * **Timing**: Autoscaling takes a moment to complete. In the meantime, the service disk might fill up and the service might enter read-only mode until autoscaling finishes, unless the autoscaler's disk capacity limit is reached. * **Maintenance updates**: Autoscaling works only on fully running services and can't happen during a maintenance update. * **Manual changes**: Changing disk space manually can delay an autoscaling event. * **Terraform**: Don't manage disk space with the Aiven Terraform Provider on a service that uses the autoscaler, to avoid conflicts between the two. * **Performance**: Disk added through autoscaling is slower than the original disk until the next maintenance update applies. This might affect I/O-intensive workloads. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven organization, project, and service that's up and running * The operator role for the organization, project, and service * Dynamic disk sizing support on your service plan and cloud region * One of the following to manage the autoscaler: * [Aiven Console](https://console.aiven.io/) * [Aiven API](https://api.aiven.io/doc/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) ### Enable disk autoscaling[​](#enable-disk-autoscaling "Direct link to Enable disk autoscaling") To enable disk autoscaling, create an autoscaler integration endpoint, then enable an autoscaler integration on your service using that endpoint. * Console * API * CLI * Terraform Create an autoscaler endpoint: 1. Log in to [Aiven Console](https://console.aiven.io/), and go to your organization and project. 2. On the left sidebar, click **Integration endpoints**. 3. Click **Aiven Autoscaler** > **Add new endpoint**. 4. Set the endpoint name and the maximum total disk storage in GB, and click **Add endpoint**. Enable the autoscaler on a service: 1. On the left sidebar, click **Services**, and open your service. 2. On the left sidebar, click **Integrations**. 3. In **Endpoint integrations**, click **Aiven Autoscaler**. 4. Click the endpoint you created, and click **Enable**. 1) Call [ServiceIntegrationEndpointCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointCreate) to create an autoscaler integration endpoint on your project: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "endpoint_name": "ENDPOINT_NAME", "endpoint_type": "autoscaler", "user_config": { "autoscaling": [ { "type": "autoscale_disk", "cap_gb": 300 } ] } }' ``` 2) Call [ServiceIntegrationCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationCreate) to enable the autoscaler integration on your service, using the endpoint ID from the previous response: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "dest_endpoint_id": "ENDPOINT_ID", "integration_type": "autoscaler", "source_project": "PROJECT_NAME", "source_service": "SERVICE_NAME" }' ``` 1. Create an autoscaler integration endpoint using [avn service integration-endpoint-create](/docs/tools/cli.md): ``` avn service integration-endpoint-create \ --project PROJECT_NAME \ --endpoint-name ENDPOINT_NAME \ --endpoint-type autoscaler \ --user-config-json '{"autoscaling": [{"type": "autoscale_disk", "cap_gb": 300}]}' ``` 2. Find the ID of the new endpoint: ``` avn service integration-endpoint-list --project PROJECT_NAME ``` 3. Enable the autoscaler integration on your service, using the endpoint ID from the previous step: ``` avn service integration-create \ --dest-service SERVICE_NAME \ --integration-type autoscaler \ --source-endpoint-id ENDPOINT_ID ``` Use the [`aiven_service_integration_endpoint`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration_endpoint) and [`aiven_service_integration`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration) resources: ``` resource "aiven_service_integration_endpoint" "autoscaler_endpoint" { project = "PROJECT_NAME" endpoint_name = "ENDPOINT_NAME" endpoint_type = "autoscaler" autoscaler_user_config { autoscaling { type = "autoscale_disk" cap_gb = 300 } } } resource "aiven_service_integration" "autoscaler_integration" { project = "PROJECT_NAME" integration_type = "autoscaler" source_service_name = "SERVICE_NAME" destination_endpoint_id = aiven_service_integration_endpoint.autoscaler_endpoint.id } ``` See the [disk autoscaler guide](https://registry.terraform.io/providers/aiven/aiven/latest/docs/guides/disk-autoscaler) for more details. ### Change the maximum disk space for autoscaling[​](#change-the-maximum-disk-space-for-autoscaling "Direct link to Change the maximum disk space for autoscaling") After you enable disk autoscaling, you can update the maximum total disk size at any time. * Console * API * CLI * Terraform 1. On the left sidebar, click **Integration endpoints**. 2. Click **Aiven Autoscaler**. 3. Find your endpoint, click **Actions**, and click the option to edit it. 4. Set a new maximum disk storage value, and save your changes. Call [ServiceIntegrationEndpointUpdate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointUpdate) with the new `cap_gb` value: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint/ENDPOINT_ID \ --header "Authorization: Bearer TOKEN" \ --header "Content-Type: application/json" \ --data '{ "user_config": { "autoscaling": [ { "type": "autoscale_disk", "cap_gb": 500 } ] } }' ``` ``` avn service integration-endpoint-update ENDPOINT_ID \ --user-config-json '{"autoscaling": [{"type": "autoscale_disk", "cap_gb": 500}]}' ``` Update the `cap_gb` value in the `autoscaling` block of your [`aiven_service_integration_endpoint`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/service_integration_endpoint) resource, then apply the change. ### Turn off disk autoscaling[​](#turn-off-disk-autoscaling "Direct link to Turn off disk autoscaling") To turn off disk autoscaling, remove the autoscaler integration from your service. You can also delete the integration endpoint if you no longer need it. * Console * API * CLI Disconnect the service from the autoscaler: 1. On the left sidebar, click **Services**, and open your service. 2. On the left sidebar, click **Integrations**. 3. In **Endpoint integrations**, find **Aiven Autoscaler**, click **Actions**, and click the option to disconnect it. Delete the autoscaler endpoint, if you no longer need it: 1. On the left sidebar, click **Integration endpoints**. 2. Click **Aiven Autoscaler**. 3. Find your endpoint, and delete it. 1) Call [ServiceIntegrationDelete](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationDelete) to remove the autoscaler integration from your service: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration/INTEGRATION_ID \ --header "Authorization: Bearer TOKEN" ``` 2) Call [ServiceIntegrationEndpointDelete](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationEndpointDelete) to delete the autoscaler integration endpoint, if you no longer need it: ``` curl --request DELETE \ --url https://api.aiven.io/v1/project/PROJECT_NAME/integration_endpoint/ENDPOINT_ID \ --header "Authorization: Bearer TOKEN" ``` 1. Find the ID of the integration to remove: ``` avn service integration-list SERVICE_NAME ``` 2. Remove the autoscaler integration from your service: ``` avn service integration-delete INTEGRATION_ID ``` 3. Find the ID of the integration endpoint to delete, if you no longer need it: ``` avn service integration-endpoint-list --project PROJECT_NAME ``` 4. Delete the autoscaler integration endpoint: ``` avn service integration-endpoint-delete ENDPOINT_ID ``` Related pages * [Scale disk storage manually](/docs/products/postgresql/howto/scale-disk-storage.md) * [Change the service plan](/docs/products/postgresql/howto/change-service-plan.md) * [Memory and out-of-memory conditions](/docs/products/postgresql/concepts/pg-shared-buffers.md) --- # Enable JIT in PostgreSQL® PostgreSQL® 11 introduces a new component in the execution engine, a [Just-in-Time (JIT) expression compiler](https://www.postgresql.org/docs/current/jit-reason.html). By default, the JIT feature is disabled for PostgreSQL 11 and enabled for all the subsequent PostgreSQL versions that Aiven supports. You can change JIT settings in Aiven for PostgreSQL on the global level or just for a specific database. ## Enable JIT on the global level[​](#enable-jit-on-the-global-level "Direct link to Enable JIT on the global level") You can enable JIT for the complete Aiven for PostgreSQL service both via [Aiven Console](https://console.aiven.io/) and [Aiven CLI](/docs/tools/cli.md). To enable JIT in the [Aiven Console](https://console.aiven.io/): 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. From the **Services** page, select the Aiven for PostgreSQL service where to enable JIT. 3. From the sidebar on your service's page, select **Service settings**. 4. On the **Service settings** page, go to the **Advanced configuration** section, and select **Configure**. 5. In the **Advanced configuration** window, select **Add configuration options**. 6. Select parameter `pg.jit`, and switch the toggle to `Enabled`. 7. Select **Save configuration**. To enable JIT via [Aiven CLI](/docs/tools/cli.md), you can use the [service update command](/docs/tools/cli/service-cli.md#avn-cli-service-update): ``` avn service update --project PROJECT_NAME -c pg.jit=true PG_SERVICE_NAME ``` ## Enable JIT for a specific database[​](#enable-jit-for-a-specific-database "Direct link to Enable JIT for a specific database") You might not want to use JIT for most simple queries since it would increase the cost. JIT can also be enabled for a single database: 1. Connect to the database where to enable JIT, for example, with `psql` and the service URI available in [Aiven Console](https://console.aiven.io/) > your Aiven for PostgreSQL service > the **Overview** page. ``` psql PG_CONNECTION_URI ``` 2. Alter the database (in the example `mytestdb`) and enable JIT ``` alter database mytestdb set jit=on; ``` note The above setting enables JIT by default for a logical database. The default is only applied to new client sessions. ## Enable JIT for a specific user[​](#enable-jit-for-a-specific-user "Direct link to Enable JIT for a specific user") JIT can be enabled also for a specific user: 1. Connect to the database where to enable JIT using, for example, `psql` and the service URI available in [Aiven Console](https://console.aiven.io/) > the **Overview** page of your Aiven for PostgreSQL service. ``` psql PG_CONNECTION_URI ``` 2. Alter the role (in the example: `mytestrole`), and enable JIT. ``` alter role mytestrole set jit=on; ``` note The above setting enables JIT by default for a logical database. The default is only applied to new client sessions. 3. Start a new session with the role, and check that JIT is running. ``` show jit; ``` The result should be: ``` jit ----- on (1 row) ``` 4. Run a simple query to test JIT is applied properly. ``` defaultdb=> explain analyze select sum(row) from table; QUERY PLAN ------------------------------------------------------------------------------------------------------------------------------ ----------- Finalize Aggregate (cost=10633.55..10633.56 rows=1 width=8) (actual time=299.417..299.418 rows=1 loops=1) -> Gather (cost=10633.33..10633.54 rows=2 width=8) (actual time=299.111..307.748 rows=3 loops=1) Workers Planned: 2 Workers Launched: 2 -> Partial Aggregate (cost=9633.33..9633.34 rows=1 width=8) (actual time=178.676..178.676 rows=1 loops=3) -> Parallel Seq Scan on bigone (cost=0.00..8591.67 rows=416667 width=4) (actual time=0.022..89.465 rows=33333 3 loops=3) Planning Time: 0.087 ms JIT: Functions: 12 Options: Inlining false, Optimization false, Expressions true, Deforming true Timing: Generation 1.878 ms, Inlining 0.000 ms, Optimization 4.438 ms, Emission 44.926 ms, Total 51.243 ms Execution Time: 308.777 ms (12 rows) ``` In the above example, a separate JIT section is shown after the planning time. tip The last row of the `explain analyze` command output above shows the execution time, which can be useful for a benchmark comparison. --- # Fork your Aiven for PostgreSQL® service Fork your Aiven for PostgreSQL® service to create an independent copy for testing, debugging, or development without affecting the original service. Fork an Aiven service to create a complete copy of it from its latest backup. Forked services are independent and don't share resources with or increase the load on the original service. Common use cases for forking include: * Creating a snapshot to analyze an issue. * Creating a development copy of your production environment. * Testing upgrades before applying them to production services. * Creating an instance in a different cloud provider, region, or with a different plan. * Renaming a service. During the forking process, the fork might initially have only one node while backups are being taken. The other nodes appear after the backup process is complete. When you fork a service, its configuration, databases, service users, and connection pools are copied to the new service. ## Limitations[​](#limitations "Direct link to Limitations") * You can only fork services that have at least one [backup](/docs/products/postgresql/concepts/pg-backups.md). * Service integrations are not copied to the fork. * Cross-project forking is supported only within the same organization. ## Fork a service[​](#fork-a-service "Direct link to Fork a service") * Console * CLI * API * Terraform 1. In your service, in the **Backups** section, click **Backup management**. 2. Click **Fork & restore**. 3. Choose the backup to fork from. 4. Enter a name, and select the cloud and plan. 5. Click **Create fork**. Use the [create service command](/docs/tools/cli/service-cli.md#avn-cli-service-create) with: * `--service-to-fork-from`: the name of the service to use as the source. * `--project-to-fork-from`: to fork a service in a different project, set this to the project name the source service is in. Use the [`ServiceCreate` endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) and in the `user_config` property set: * `service_to_fork_from`: the name of the source service. * `project_to_fork_from`: to fork a service in a different project, set this to the name of the project the source service is in. Use the `service_to_fork_from` attribute in the user config of your service resource. To fork a service in a different project, set the `project_to_fork_from` attribute. More information on the service resources and their configuration options is available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs). ## Fork from a specific point in time[​](#fork-from-a-specific-point-in-time "Direct link to Fork from a specific point in time") * Console * CLI * API * Terraform 1. In your service, in the **Backups** section, click **Backup management**. 2. Click **Fork & restore**. 3. Choose the point in time to fork from. 4. Enter a name, and choose the cloud and plan. 5. Click **Create fork**. Add the `--recovery-target-time` parameter to the [create service command](/docs/tools/cli/service-cli.md#avn-cli-service-create) and set it to a time between the first and latest available backups. Set the `recovery_target_time` parameter in the `user_config` property of the [`ServiceCreate` endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) to a time between the first and latest available backups. Set the `recovery_target_time` attribute in the user config of your service resource to a time between the first and latest available backups. More information on the service resources and their configuration options is available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Aiven for PostgreSQL® backups](/docs/products/postgresql/concepts/pg-backups.md) * [Rename your Aiven for PostgreSQL® service](/docs/products/postgresql/howto/rename-service.md) --- # Identify PostgreSQL® slow queries with pg\_stat\_statements Use the PostgreSQL® `pg_stat_statements` [extension](https://www.postgresql.org/docs/current/pgstatstatements.html) to find slow queries. ## Identify slow queries in the Console[​](#identify-slow-queries-in-the-console "Direct link to Identify slow queries in the Console") Use **Aiven AI Database Optimizer** to list and optimize slow queries. [Learn more](/docs/products/postgresql/howto/optimize-pg-slow-queries.md). ## Use `pg_stat_statements`[​](#use-pg_stat_statements "Direct link to use-pg_stat_statements") Query statistics deduced via the `pg_stat_statements` are the following: | Column Type | Description | | ------------- | ------------------------------------------------------------------- | | `Query` | Text of a representative statement | | `Rows` | Total number of rows retrieved or affected by the statement | | `Calls` | Number of times the statement was executed | | `Min (ms)` | Minimum time spent executing the statement | | `Max (ms)` | Maximum time spent executing the statement | | `Mean (ms)` | Mean time spent executing the statement | | `Stddev (ms)` | Population standard deviation of time spent executing the statement | | `Total (ms)` | Total time spent executing the statement | You can also create custom queries using the `pg_stat_statements` view and use all the [available columns](https://www.postgresql.org/docs/current/pgstatstatements.html) to investigate your use case. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") To query the `pg_stat_statements` view, create the `pg_stat_statements` extension (included in the [list of available extensions](/docs/products/postgresql/reference/list-of-extensions.md)): ``` CREATE EXTENSION pg_stat_statements; ``` tip Use an AI assistant connected to [Aiven MCP](/docs/tools/mcp-server.md) to inspect slow queries and suggest possible optimizations. For example: > Show the queries with the highest execution time on `my-pg-service` from `pg_stat_statements`, and suggest indexes that might improve them. ## Discover slow queries[​](#discover-slow-queries "Direct link to Discover slow queries") Display the `pg_stat_statements` view and all the columns contained: ``` \d pg_stat_statements; ``` Expect to receive the following output: ``` View "public.pg_stat_statements" Column | Type | Collation | Nullable | Default --------------------+------------------+-----------+----------+--------- userid | oid | | | dbid | oid | | | toplevel | boolean | | | queryid | bigint | | | query | text | | | plans | bigint | | | total_plan_time | double precision | | | min_plan_time | double precision | | | max_plan_time | double precision | | | mean_plan_time | double precision | | | stddev_plan_time | double precision | | | calls | bigint | | | total_exec_time | double precision | | | min_exec_time | double precision | | | max_exec_time | double precision | | | mean_exec_time | double precision | | | stddev_exec_time | double precision | | | rows | bigint | | | shared_blks_hit | bigint | | | shared_blks_read | bigint | | | shared_blks_dirtied | bigint | | | shared_blks_written | bigint | | | local_blks_hit | bigint | | | local_blks_read | bigint | | | local_blks_dirtied | bigint | | | local_blks_written | bigint | | | temp_blks_read | bigint | | | temp_blks_written | bigint | | | blk_read_time | double precision | | | blk_write_time | double precision | | | wal_records | bigint | | | wal_fpi | bigint | | | wal_bytes | numeric | | | ``` On older PostgreSQL versions, you might find different column names (for example, the column previously named `max_time` is now `max_exec_time`). Always refer to the [PostgreSQL® official documentation](https://www.postgresql.org/docs/current/pgstatstatements.html) with the version you are using for accurate column matching. tip You can write custom queries to `pg_stat_statements` to help you analyze recently run queries in your database. ## Sort database queries based on `total_exec_time`[​](#sort-database-queries-based-on-total_exec_time "Direct link to sort-database-queries-based-on-total_exec_time") The following query, inspired by a [GitHub repository](https://github.com/heroku/heroku-pg-extras/blob/ece431777dd34ff6c2a8dfb790b24db99f114165/commands/outliers.js), uses the `pg_stat_statements` view, shows the running queries sorted descending by `total_exec_time`, re-formats the `calls` column, and deduces the `prop_exec_time` and `sync_io_time`: ``` SELECT interval '1 millisecond' * total_exec_time AS total_exec_time, to_char((total_exec_time/sum(total_exec_time) OVER()) * 100, 'FM90D0') || '%' AS prop_exec_time, to_char(calls, 'FM999G999G999G990') AS calls, interval '1 millisecond' * (blk_read_time + blk_write_time) AS sync_io_time, query AS query FROM pg_stat_statements WHERE userid = ( SELECT usesysid FROM pg_user WHERE usename = current_user LIMIT 1 ) ORDER BY total_exec_time DESC LIMIT 10; ``` Run the above commands on your own PostgreSQL® to gather more information about how the recent queries are performing. tip To discard the `pg_stat_statements` previously gathered statistics, run ``` SELECT pg_stat_statements_reset() ``` ## Find top queries with high I/O activity[​](#find-top-queries-with-high-io-activity "Direct link to Find top queries with high I/O activity") The following SQL shows queries with their `id` and mean time in seconds. The result set is ordered based on the sum of `blk_read_time` and `blk_write_time` so that queries with the highest read/write are shown at the top. ``` SELECT userid::regrole, dbid, query, queryid, mean_time/1000 as mean_time_seconds FROM pg_stat_statements ORDER by (blk_read_time+blk_write_time) DESC LIMIT 10; ``` ## See top time-consuming queries[​](#see-top-time-consuming-queries "Direct link to See top time-consuming queries") Aside from relevant information about the database, the following SQL retrieves * Number of calls * Consumption time as `total_time_seconds` (in milliseconds) * Minimum time (in milliseconds) * Maximum time (in milliseconds) * Mean times (in milliseconds) The result set is ordered in descending order by `mean_time`, showing the queries with the longest consumption time first. ``` SELECT userid::regrole, dbid, query, calls, total_time/1000 as total_time_seconds, min_time/1000 as min_time_seconds, max_time/1000 as max_time_seconds, mean_time/1000 as mean_time_seconds FROM pg_stat_statements ORDER by mean_time desc LIMIT 10; ``` ## Check queries with high memory usage[​](#check-queries-with-high-memory-usage "Direct link to Check queries with high memory usage") The following SQL retrieves the query, its `id`, and relevant information about the database. The result set is ordered by showing the queries with the highest memory usage at the top, summing the number of shared memory blocks returned from the cache (`shared_blks_hit`) and the number of shared memory blocks marked as "dirty" during a request needed to be written to the disk (`shared_blks_dirtied`). ``` SELECT userid::regrole, dbid, queryid, query FROM pg_stat_statements ORDER by (shared_blks_hit+shared_blks_dirtied) DESC limit 10; ``` ## Next steps[​](#next-steps "Direct link to Next steps") Once you have identified slow queries: * Inspect the query plan and execution using [EXPLAIN ANALYZE](https://www.postgresql.org/docs/current/using-explain.html) to understand how to optimise your design to improve the performance. * [Optimize slow PostgreSQL® queries](/docs/products/postgresql/howto/optimize-pg-slow-queries.md). --- # Connect to Aiven for PostgreSQL® services Connect to the Aiven for PostgreSQL® service using various programming languages or tools. All connections to PostgreSQL are encrypted and protected with TLS. For a connection to be established, `sslmode` can be set as follows: * **By default**, `sslmode` needs to be set to `require`. This ensures that TLS is used and data is encrypted while in-transit. This doesn't require or verify a certificate. * **For more security**, `sslmode` can be set either to `verify-ca` or to `verify-full`. Each of these modes requires supplying a certificate (`ca.pem`) and verifies it. Read more about [TLS/SSL certificates](/docs/platform/concepts/tls-ssl-certificates.md). ## [Go](/docs/products/postgresql/howto/connect-go.md) [This example connects to PostgreSQL® service from Go, making use of the](/docs/products/postgresql/howto/connect-go.md) ## [Java](/docs/products/postgresql/howto/connect-java.md) [This example connects to PostgreSQL® service from Java, making use of JDBC Driver.](/docs/products/postgresql/howto/connect-java.md) ## [NodeJS](/docs/products/postgresql/howto/connect-node.md) [This example connects to PostgreSQL® service from NodeJS, making use of the pg package.](/docs/products/postgresql/howto/connect-node.md) ## [PHP](/docs/products/postgresql/howto/connect-php.md) [This example connects to PostgreSQL® service from PHP, making use of the](/docs/products/postgresql/howto/connect-php.md) ## [Python](/docs/products/postgresql/howto/connect-python.md) [This example connects to a PostgreSQL® service from Python, making use](/docs/products/postgresql/howto/connect-python.md) ## [psql](/docs/products/postgresql/howto/connect-psql.md) [psql is a command line tool for PostgreSQL®, useful to manage and](/docs/products/postgresql/howto/connect-psql.md) ## [pgAdmin](/docs/products/postgresql/howto/connect-pgadmin.md) [pgAdmin is one of the most popular PostgreSQL® clients. Use it to manage and query your database.](/docs/products/postgresql/howto/connect-pgadmin.md) ## [Rivery](/docs/products/postgresql/howto/connect-rivery.md) [Configure PostgreSQL® as a connection in Rivery, a fully managed solution for data ingestion, transformation, orchestration, reverse ETL and more.](/docs/products/postgresql/howto/connect-rivery.md) ## [Skyvia](/docs/products/postgresql/howto/connect-skyvia.md) [Skyvia is an universal cloud data platform. This](/docs/products/postgresql/howto/connect-skyvia.md) ## [Zapier](/docs/products/postgresql/howto/connect-zapier.md) [Zapier is an automation platform that connects](/docs/products/postgresql/howto/connect-zapier.md) ## [DataGrip](/docs/products/postgresql/howto/connect-datagrip.md) [Use DataGrip to connect to your Aiven for](/docs/products/postgresql/howto/connect-datagrip.md) ## [DBeaver](/docs/products/postgresql/howto/connect-dbeaver.md) [Use DBeaver to connect to your Aiven for PostgreSQL® service.](/docs/products/postgresql/howto/connect-dbeaver.md) ## [LibreDB Studio](/docs/products/postgresql/howto/connect-libredb-studio.md) [Use LibreDB Studio to connect to your Aiven for PostgreSQL®](/docs/products/postgresql/howto/connect-libredb-studio.md) --- # pgaudit logging ## [pgaudit overview](/docs/products/postgresql/concepts/pg-audit-logging.md) [The path to optimal data security, compliance, incident management, and system performance starts with collecting robust audit logs.](/docs/products/postgresql/concepts/pg-audit-logging.md) ## [Collect pgaudit logs](/docs/products/postgresql/howto/use-pg-audit-logging.md) [Enable and configure the Aiven for PostgreSQL® audit logging feature on your service. Access and visualize your logs to monitor activities on your databases.](/docs/products/postgresql/howto/use-pg-audit-logging.md) --- # Enable logical replication on Amazon Aurora PostgreSQL® If you have not enabled logical replication on Aurora already, the following instructions shows how to set the `rds.logical_replication` parameter to `1` (true) in the parameter group. 1. Create a DB Cluster parameter group for your Aurora database. ![Aurora PostgreSQL cluster parameter group](/docs/assets/images/migrate-aurora-pg-parameter-group-c9e09afb01690842194f6f3709bda7ec.png) 2. Set the `rds.logical_replication` parameter to `1` (true) in the parameter group. ![Aurora PostgreSQL cluster parameter value](/docs/assets/images/migrate-aurora-pg-parameter-value-00ea545714cbbb60d15207a3a1694405.png) 3. Modify Database options to use the new DB Cluster parameter group - `RDS` > `Databases` > `Modify`. ![Aurora PostgreSQL cluster parameter modify](/docs/assets/images/migrate-aurora-pg-parameter-modify-31e2e336b1b146b5db2a060bdc40893d.png) warning Apply immediately or reboot is required to see configuration change reflected to `wal_level`. --- # Enable logical replication on Amazon RDS PostgreSQL® If you have not enabled logical replication on RDS already, the following instructions shows how to set the `rds.logical_replication` parameter to `1` (true) in the parameter group. 1. Create a parameter group for your RDS database. ![RDS PostgreSQL parameter group](/docs/assets/images/migrate-rds-pg-parameter-group-aaa4a18a1a1640f80e903a1ce9697c39.png) 2. Set the `rds.logical_replication` parameter to `1` (true) in the parameter group ![RDS PostgreSQL parameter value](/docs/assets/images/migrate-rds-pg-parameter-value-bf6cbf64047500a410876cf18536ca9b.png) 3. Modify Database options to use the new DB parameter group: `RDS` > `Databases` > `Modify` ![RDS PostgreSQL parameter modify](/docs/assets/images/migrate-rds-pg-parameter-modify-7e9de45fdf2c0a0d17d0bdb7a8a38316.png) Apply immediately or a reboot is required to reflect the configuration change into `wal_level`. --- # Enable logical replication on Google Cloud SQL If you have not enabled logical replication on Google Cloud SQL PostgreSQL® already, set the `cloudsql.logical_decoding` parameter to `On`: 1. Set the logical replication parameter for your Cloud SQL PostgreSQL® database. ![Cloud SQL PostgreSQL flags](/docs/assets/images/migrate-cloudsql-flags-a9c2f610b9e36679540f9bffa894d36b.png) 2. Authorize the Aiven for PostgresSQL® IP to connect to Cloud SQL, using the network CIDR. ![Cloud SQL PostgreSQL network](/docs/assets/images/migrate-cloudsql-network-672d8092c549d66453dd3f3659f12acb.png) 3. Set replication role to PostgreSQL user (or the user will be used for migration) in Cloud SQL PostgreSQL: ``` ALTER ROLE postgres REPLICATION; ``` --- # Maintenance and updates for your Aiven for PostgreSQL® service Manage maintenance updates and set the maintenance window for your Aiven for PostgreSQL® service. ## Maintenance updates[​](#maintenance-updates "Direct link to Maintenance updates") Aiven applies some maintenance updates automatically. The following are the types of updates: * **Mandatory updates:** Security updates, quarterly patch releases, and platform updates that affect reliability or stability of the service nodes. * **Optional updates:** All other updates are initially optional. After six months, they become mandatory and are applied in the next week’s maintenance window. * **Periodic infrastructure updates:** Scheduled automatically for services with nodes active for 180 days and more. These updates are mandatory for all services, except those with maintenance turned off. Critical security updates are applied during the next available maintenance window. For other updates, Aiven gives you at least seven days' notice. Maintenance updates are also automatically applied during service upgrades. To view pending updates: * Console * CLI * API 1. In your service, click **Service settings**. 2. Go to the **Service management** section. Use the [`avn service get`](/docs/tools/cli/service-cli.md#avn_service_get) command. Use the [`service`](https://api.aiven.io/doc/#tag/Service/operation/ServiceGet) endpoint. ## Maintenance window[​](#maintenance-window "Direct link to Maintenance window") The maintenance window is the time period when Aiven can automatically apply maintenance updates to a service. When an update becomes available, Aiven schedules it for the next available maintenance window for each service. The update runs in the first window after it becomes available, and can begin any time after the start time. For example, if a service has a maintenance window of Monday 12:00 UTC, and an update becomes available on Tuesday, the update will be applied on the following Monday. During maintenance, Aiven might restart or replace service nodes. This can cause brief connection interruptions, but services are designed to minimize downtime. Aiven performs maintenance in a rolling-forward style, creating new nodes alongside existing ones and retiring the old nodes after the upgrade completes. Major service upgrades are triggered manually. A manually triggered upgrade starts immediately, regardless of the maintenance window. important You cannot control the order in which services are updated. Each service updates according to its own configured maintenance window, and there is no guaranteed way to control the update sequence. Manual updates and maintenance window adjustments only help for non-critical updates. ## Set the maintenance window[​](#set-the-maintenance-window "Direct link to Set the maintenance window") To set the maintenance window for your service: * Console * Terraform 1. In the Aiven Console, open your service. 2. In the **Maintenance** section, click **Actions** > **Change maintenance window**. 3. Set the day and time. 4. Click **Save changes**. Use the `maintenance_window_dow` and `maintenance_window_time` attributes in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). ## Certificate rotation[​](#certificate-rotation "Direct link to Certificate rotation") Aiven periodically rotates the CA certificate for your project, including for your Aiven for PostgreSQL® service. This rotation uses the same maintenance process described in [Maintenance updates](#maintenance-updates), applied during your service's maintenance window, to update your service to trust and use the new certificate. If you connect using `sslmode=verify-ca` or `verify-full`, update your client to trust the new certificate before the rotation completes. For details on the certificate bundle and rotation process, see [TLS/SSL certificates](/docs/platform/concepts/tls-ssl-certificates.md#certificate-rotation). Related pages * [Version upgrades](/docs/products/postgresql/howto/upgrade.md) * [Change the service plan](/docs/products/postgresql/howto/change-service-plan.md) * [TLS/SSL certificates](/docs/platform/concepts/tls-ssl-certificates.md) --- # Manage Aiven for PostgreSQL® extensions Install, update, and remove PostgreSQL® extensions on Aiven for PostgreSQL using SQL commands. Aiven for PostgreSQL supports a curated set of extensions that you install, update, and remove using SQL commands. All database users can manage extensions, including the default `avnadmin` user and any other database user created through the Aiven Console, Aiven CLI, Aiven API, or Aiven Provider for Terraform. Aiven for PostgreSQL applies an extension allowlist at the service level, so managing extensions doesn't require elevated database privileges. tip Instead of running SQL commands, you can also manage extensions using an AI assistant connected to [Aiven MCP](/docs/tools/mcp-server.md), or using the extension manager in the Aiven Console. ## Install an extension[​](#install-an-extension "Direct link to Install an extension") To install an extension, run: ``` CREATE EXTENSION EXTENSION_NAME CASCADE; ``` ## Update an extension[​](#update-an-extension "Direct link to Update an extension") To upgrade an already-installed extension to the latest version, run: ``` ALTER EXTENSION EXTENSION_NAME UPDATE; ``` ## Delete an extension[​](#delete-an-extension "Direct link to Delete an extension") To delete an extension, run: ``` DROP EXTENSION EXTENSION_NAME; ``` ## Request an extension[​](#request-an-extension "Direct link to Request an extension") Aiven welcomes suggestions for additional extensions, and some extensions can be enabled on request. For any extension that's not on the [list of approved extensions](/docs/products/postgresql/reference/list-of-extensions.md), [open a support ticket](/docs/platform/howto/support.md#create-a-support-ticket) and include: * The extension you're requesting. * The database service and user database that need it. ## FAQ[​](#faq "Direct link to FAQ") * **Do you need extra configuration before installing `pg_stat_plans` or `pg_stat_monitor`?** Yes. Enable the matching [advanced configuration](/docs/products/postgresql/reference/advanced-params.md) parameter for your service, `pg_stat_plans_enable` or `pg_stat_monitor_enable`, before you install either extension. Enabling either parameter applies a service restart. * **Does a maintenance update also update your extensions?** No. User schemas and functions often rely on specific extension versions, so Aiven for PostgreSQL doesn't assume that every extension is safe to upgrade automatically. To test an extension upgrade before applying it to your live database, fork your service and run the upgrade on the copy. * **Can you install untrusted language extensions, such as `plpythonu`?** No. Aiven for PostgreSQL doesn't support *untrusted* language extensions because they would compromise the ability to guarantee the highest possible service level. Related pages * [Extensions on Aiven for PostgreSQL®](/docs/products/postgresql/reference/list-of-extensions.md) * [Extension versions per PostgreSQL release](/docs/products/postgresql/reference/list-of-extensions-for-each-version.md) * [Advanced parameters for Aiven for PostgreSQL®](/docs/products/postgresql/reference/advanced-params.md) * [Support](/docs/platform/howto/support.md) --- # Manage connection pooling [Connection pooling](/docs/products/postgresql/concepts/pg-connection-pooling.md) lets you maintain very large numbers of connections to a database while minimizing the consumption of server resources. note Connection pooling requires a startup plan or higher. ## Connection pooling tips[​](#connection-pooling-tips "Direct link to Connection pooling tips") You can connect directly to the PostgreSQL® server using the **Service URI** setting listed on the **Overview** page. However, this URI doesn't make use of the PgBouncer pooling. PgBouncer pools use a different port number than the regular PostgreSQL server port. The PgBouncer **Service URI** for a particular pool is in [Aiven Console](https://console.aiven.io/) > your service's page > **Connect** > **Connection pools**. You can also view the generic PgBouncer service URI for your pools on your service's **Overview** page > **Connection information** > **PgBouncer** tab. You can use both pooled and non-pooled connections at the same time. note If you have set a custom `search_path` for your database, this is not automatically set for your new connection pool. Remember to set it also for new connection pools when you create them. ## Manage connection pools[​](#manage-connection-pools "Direct link to Manage connection pools") To manage the connection pools: 1. Log in to [Aiven Console](https://console.aiven.io/) and select your Aiven for PostgreSQL service. 2. In the **Connect** section, click **Connection pools**. 3. In the **Connection pools** view, you can check the available connection pools and add or remove them. The settings available are as follows: * **Pool name**: Enter a name for your connection pool. This also becomes the `database` or `dbname` connection parameter for your pooled client connections. This parameter must be equal to the `Database` parameter. * **Database**: Choose the database to connect to. Each pool can only connect to a single database. * **Username**: Select the database username to use when connecting to the backend database. * **Pool Mode**: Select the [pooling mode](/docs/products/postgresql/concepts/pg-connection-pooling.md#pooling-modes). * **Pool Size**: Select how many PostgreSQL server connections this pool can use at a time. The value can be from 1 to 10000. note If you manage connection pools with the [Aiven Operator for Kubernetes®](/docs/tools/kubernetes.md), its `ConnectionPool` resource limits `poolSize` to a maximum of 1000, lower than the maximum available through the console, CLI, Terraform, or API. important The **Pool Size** parameter is NOT the maximum number of client connections of this pool. Each pool can handle from a minimum of 5000 client connections to a maximum defined by the lower threshold between: * 500 for each GB of RAM in the service * A total of 50000 client connections 4. To view the database connection settings for a pool, click **Actions** > **Info**. You can also manage connection pools with the [Aiven CLI](/docs/tools/cli/service/connection-pool.md), with Terraform, using the `aiven_connection_pool` resource, or with the [Aiven Operator for Kubernetes®](/docs/tools/kubernetes.md): ``` Loading... ``` ## Connection pools for replicas[​](#connection-pools-for-replicas "Direct link to Connection pools for replicas") For services with multiple nodes, whenever you define a connection pool, the same connection pool is created both for primary and standby servers. For standby servers, the connection pool URI is exactly the same as for the primary server, except that the host name has a `replica-` prefix. For example, if the primary connection URI is as follows: ``` postgres://avnadmin:password@pg-prod-myproject.aivencloud.com:20986/mypool?params ``` The replica connection pool URI is as follows: ``` postgres://avnadmin:password@replica-pg-prod-myproject.aivencloud.com:20986/mypool?params ``` --- # Manage Aiven for PostgreSQL® service users Create and manage service users in your Aiven for PostgreSQL® service to control access to its databases and tables. Service users only exist in the scope of the Aiven service. They are unique to the service and not shared with any other services. Every service has a default `avnadmin` user with full access to the service. ## Add a service user[​](#add-a-service-user "Direct link to Add a service user") * Aiven Console * Aiven CLI * Aiven API * Terraform 1. In your service, in the **Connect** section, click **Users**. 2. Click **Add service user** or **Create user**. 3. Enter a name for your service user. 4. Set up all the other configuration options. If a password is required, a random password is generated automatically. You can change it later. 5. Click **Add service user**. Run the [avn service user-create](/docs/tools/cli/service/user.md#avn-service-user-create) command: ``` avn service user-create SERVICE_NAME --username USERNAME ``` Replace the following: * `SERVICE_NAME`: the name of your Aiven for PostgreSQL service. * `USERNAME`: the name of the service user to create. Use the [ServiceUserCreate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUserCreate) endpoint: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/user \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{"username": "USERNAME"}' ``` Replace the placeholders with your project name, service name, bearer token, and the username to create. Use the [`aiven_pg_user` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/pg_user) to create and manage service users. Related pages * [Create a database](/docs/products/postgresql/howto/create-database.md) * [Connect to your service](/docs/products/postgresql/howto/list-code-samples.md) --- # Migrate PostgreSQL® databases to Aiven using aiven-db-migrate The `aiven-db-migrate` tool is an open source project available on [GitHub](https://github.com/aiven/aiven-db-migrate), and it is the preferred way to perform PostgreSQL® database migration. `aiven-db-migrate` performs a schema dump and migration first to ensure schema compatibility. It supports both logical replication, and using a dump and restore process. Logical replication is the default method which keeps the two databases synchronized until the replication is interrupted. Restrictions on logical replication Before you use the logical replication method, make sure you know and understand all the [restrictions](https://www.postgresql.org/docs/current/logical-replication-restrictions.html). If the preconditions for logical replication are not met for a database, the migration falls back to using `pg_dump`. note You can use logical replication when migrating from AWS RDS PostgreSQL® 10+ and [Google CloudSQL PostgreSQL](https://cloud.google.com/sql/docs/release-notes#August_30_2021). ## Requirements[​](#requirements "Direct link to Requirements") To perform a migration with `aiven-db-migrate`: * The source server needs to be publicly available or accessible via a virtual private cloud (VPC) peering connection between the private networks. * You have a user account with access to the destination cluster from an external IP, as configured in `pg_hba.conf` on the source cluster. To use the **logical replication** method, you'll need the following: * PostgreSQL® version is 10 or higher. * Sufficient access to the source cluster (either the `replication` permission or the `aiven-extras` extension installed). The extension allows you to perform publish/subscribe-style logical replication without a superuser account, and it is preinstalled on Aiven for PostgreSQL servers. See [Aiven Extras on GitHub](https://github.com/aiven/aiven-extras). * An available replication slot on the destination cluster for each database migrated from the source cluster. 1. If you don't have an Aiven for PostgreSQL database yet, run the following command to create a couple of PostgreSQL services via [Aiven CLI](/docs/tools/cli.md) substituting the parameters accordingly: ``` avn service create --project PROJECT_NAME -t pg -p DEST_PG_PLAN DEST_PG_NAME ``` 2. Enable the `aiven_extras` extension in the Aiven for PostgreSQL® target database as written in the [dedicated document](/docs/products/postgresql/concepts/dba-tasks-pg.md#aiven_extras_extension). 3. Set the `wal_level` to `logical` on source database. Check the following examples for the main managed databases: * [Amazon Aurora](/docs/products/postgresql/howto/logical-replication-aws-aurora.md) * [Amazon RDS](/docs/products/postgresql/howto/logical-replication-aws-rds.md) * [Google Cloud SQL](/docs/products/postgresql/howto/logical-replication-gcp-cloudsql.md) note Aiven for PostgreSQL has `wal_level` set to `logical` by default To review the current `wal_level`, run the following command on the source cluster via `psql` ``` show wal_level; ``` ## Variables[​](#pg_migrate_wal "Direct link to Variables") The following variables need to be substituted in the `aiven-db-migrate` calls | Variable | Description | | -------------- | ----------------------------------------------------------------------- | | `SRC_USERNAME` | Username for source PostgreSQL connection | | `SRC_PASSWORD` | Password for source PostgreSQL connection | | `SRC_HOSTNAME` | Hostname for source PostgreSQL connection | | `SRC_PORT` | Port for source PostgreSQL connection | | `DEST_PG_NAME` | Destination Aiven for PostgreSQL service name | | `DST_DBNAME` | Bootstrap database name for destination PostgreSQL connection | | `DB_TO_SKIP` | Comma separated list of database names for which to skip the migrations | ## Perform the migration with `aiven-db-migrate`[​](#perform-the-migration-with-aiven-db-migrate "Direct link to perform-the-migration-with-aiven-db-migrate") warning Running a logical replication migration twice on the same cluster creates duplicate data. Logical replication also has [limitations](https://www.postgresql.org/docs/current/logical-replication-restrictions.html) on what it can copy. ### Run `aiven-db-migrate` using the Aiven CLI[​](#run-aiven-db-migrate-using-the-aiven-cli "Direct link to run-aiven-db-migrate-using-the-aiven-cli") You can initiate a migration to an Aiven for PostgreSQL® service with the [Aiven CLI](/docs/tools/cli.md) and the following command, substituting the placeholders accordingly: ``` avn service update --project PROJECT_NAME -c migration.host=SRC_HOSTNAME \ -c migration.port=SRC_PORT \ -c migration.ssl=true \ -c migration.username=SRC_USERNAME \ -c migration.password=SRC_PASSWORD \ -c migration.dbname=DST_DBNAME \ -c migration.ignore_dbs=DB_TO_SKIP \ DEST_PG_NAME ``` note Using avn CLI shows limited status output, to troubleshoot failures run `aiven-db-migrate` [directly from Python](/docs/products/postgresql/howto/run-aiven-db-migrate-python.md). ### Display the migration status using the Aiven CLI[​](#display-the-migration-status-using-the-aiven-cli "Direct link to Display the migration status using the Aiven CLI") You can see the migration status using the [Aiven CLI](/docs/tools/cli.md) and the following call: ``` avn service migration-status --project PROJECT_NAME SERVICE_NAME ``` note There may be delay for migration status to update the current progress, keep running this command to see the most up-to-date status. The output is be similar to the following, which mentions that the `pg_dump` migration of the `defaultdb` database is `done` and the logical `replication` of the `has_aiven_extras` database is syncing: ``` -----Response Begin----- { "migration": { "error": null, "method": "", "status": "done" }, "migration_detail": [ { "dbname": "has_aiven_extras", "error": null, "method": "replication", "status": "syncing" }, { "dbname": "defaultdb", "error": null, "method": "pg_dump", "status": "done" } ] } -----Response End----- STATUS METHOD ERROR ====== ====== ===== done null ``` note The overall `method` field is left empty due to the mixed methods used to migrate each database. tip After the migration finishes, use an AI assistant connected to [Aiven MCP](/docs/tools/mcp-server.md) to check the migrated data. For example: > Compare the tables and row counts in `source-pg-service` and `destination-pg-service`. ### Stop the migration process using the Aiven CLI[​](#stop-the-migration-process-using-the-aiven-cli "Direct link to Stop the migration process using the Aiven CLI") Once the migration is finished, you can stop the related process using the [Aiven CLI](/docs/tools/cli.md). warning Make sure your migration process is in one of the following state when triggering the removal: * `done` if using `pg_dump` * `syncing` if using logical replication Otherwise, removing a migration configuration can leave the destination cluster in an inconsistent state. The migration process can be stopped with: ``` avn service update --project PROJECT_NAME --remove-option migration DEST_PG_NAME ``` This command removes all logical replication-related objects from both source and destination cluster. If using logical replication, the process stops it. It has no effect for the `pg_dump` method as it is a one-time operation. warning Don't stop the migration process while it is `running` state since both the logical replication and `pg-dump`/`pg-restore` methods are copying data from the source to the destination cluster. Once the migration is completed successfully, remove unused replication slots. The migration using `aiven-db-migrate` can also be [performed in Python](/docs/products/postgresql/howto/run-aiven-db-migrate-python.md) without requiring the Aiven CLI. --- # Migrate to a different cloud provider or region Any Aiven service can be relocated to a different cloud vendor or region. This is also valid for PostgreSQL® where the migration happens without downtime. Cloud provider/region migration features mean that you can relocate a service at any time, for example to meet specific latency requirements for a particular geography. To migrate a PostgreSQL service to a new cloud provider/region 1. In your service, click **Service settings** from the sidebar. 2. In the **Cloud and network** section, click **Actions** > **Change cloud**. 3. In the **Cloud** section , select a cloud provider and region, and click **Change**. The PostgreSQL cluster will enter the `REBALANCING` state, still serving queries from the old provider/region. note You can check the service's state at the top of the service's page, just below the service's name. New nodes will be added to the existing PostgreSQL cluster residing in the new provider/region and the data will be replicated to the new nodes. Once the new nodes are in sync, one of them will become the new primary node and all the nodes in the old provider/region will be decommissioned. After this phase the cluster enters in the `RUNNING` status, the PostgreSQL endpoint will not change. note You can check the nodes' availability at the top of the service's page, just below the service's name. tip To have consistent query time across the globe, consider [creating several read-only replicas across different cloud provider/regions](/docs/products/postgresql/howto/create-read-replica.md) --- # Migrate PostgreSQL® databases to Aiven using the Aiven Console Migrate PostgreSQL databases to the Aiven platform using the Aiven Console. note For the CLI method using `aiven-db-migrate`, see [Migrate to Aiven for PostgreSQL® with aiven-db-migrate](/docs/products/postgresql/howto/migrate-aiven-db-migrate.md). ## About migrating via console[​](#about-migrating-via-console "Direct link to About migrating via console") The console migration tool enables you to migrate PostgreSQL databases to managed PostgreSQL clusters in your Aiven organization. You can migrate the following: * Existing on-premise PostgreSQL databases * Cloud-hosted PostgreSQL databases * Managed PostgreSQL database clusters on Aiven. With the console migration tool, you can migrate your data using one of these methods: * [Continuous migration method](/docs/products/postgresql/howto/migrate-db-to-aiven-via-console.md#pg-continuous-migration) (default and recommended) * [One-time snapshot method](/docs/products/postgresql/howto/migrate-db-to-aiven-via-console.md#pg-dump-migration) (`pg_dump`). tip After completing the console migration workflow, use an AI assistant connected to [Aiven MCP](/docs/tools/mcp-server.md) to check the migrated data. For example: > Compare the tables and row counts in `source-pg-service` and `destination-pg-service`. ### Continuous migration[​](#pg-continuous-migration "Direct link to Continuous migration") The continuous migration method is used by default in the console. The continuous migration keeps the source database operational during the migration. This method uses [logical replication](https://www.postgresql.org/docs/current/logical-replication.html), which enables data transfer not only for the data that has already been there in the source database when triggering the migration but also for any data written to the source database during the migration. important Before you use the logical replication, make sure you know and understand all the restrictions it has. For details, see [Logical replication restrictions](https://www.postgresql.org/docs/current/logical-replication-restrictions.html). Using the continuous migration requires either superuser permissions or the `aiven_extras` extension installed on the source database. How to verify that you have superuser permissions. Use `psql` to run the `\du` command: ``` \du ``` Expected output ``` Role name | Attributes | Member of ----------+------------------------------------------------------------+----------------------------------------- _source_db | Superuser, Replication | {} example | Create role, Create DB, Replication, Bypass RLS | {pg_read_all_stats,pg_stat_scan_tables,pg_signal_backend} postgres | Superuser, Create role, Create DB, Replication, Bypass RLS | {} ``` Identify your role name in the `Role name` column and check if it has the `Superuser` attribute assigned in the `Attributes` column. If not, request it from your system administrator. No superuser permissions? Install `aiven_extras`. If you don't have superuser permissions, but you still want to use the continuous migration, you can install the `aiven_extras` extension on the source database using the following command: ``` CREATE EXTENSION `aiven_extras` CASCADE; ``` ### Dump migration[​](#pg-dump-migration "Direct link to Dump migration") `pg_dump` is a point-in-time snapshot. The data written to the source database during the migration process (after initiating the dump) is not migrated to the target database. When you start a dump migration, make sure no data is written to the source database by the time the dumping process is over. When you trigger the migration setup in the console and initial checks detect that your source database does not support the logical replication, you are notified about it via wizard. To continue with the migration, you can select the alternative `pg_dump` migration method in the wizard. No superuser permissions and no `aiven_extras`? Migrate using the dump method. Without superuser permissions or `aiven_extras` installed, you cannot use the logical replication and migrate in a continuous manner. In that case, you can migrate your database using the dump method if you have the following permissions: * Connect * Select on all tables in the database * Select on all the sequences in the database For the instruction on how to perform a dump, skip a few sections that follow and go straight to [Migrate a database](/docs/products/postgresql/howto/migrate-db-to-aiven-via-console.md#migrate-in-console). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") To use the default continuous migration method in the console: * Have the logical replication enabled on your source database either with superuser permissions or the `aiven_extras` extension. * The source database's hostname or IP address must be [accessible from the public Internet](/docs/platform/howto/public-access-in-vpc.md). * Collect the following source database's credentials and reference data: * Public hostname or connection string, or IP address used to connect to the database * Port used to connect to the database * Username (for a user with superuser permissions) * Password. * Firewalls protecting the source database and the target databases need to be open to allow the traffic and connection between the databases (update or disable the firewalls temporarily if needed). ## Pre-configure the source[​](#pre-configure-the-source "Direct link to Pre-configure the source") * Allow remote connections on the source database. Ensure your database allows all remote connections by using `psql` to run the following query: ``` SHOW listen_addresses; ``` If enabled, you can expect the following output (with `listen_addresses` set to `*`): ``` listen_addresses ----------- * (1 row) ``` If the command line returns something different, enable remote connections for your database with the following query: ``` ALTER SYSTEM SET listen_addresses = '*'; ``` * Change your IPv4 local connection to `0.0.0.0/0` to allow all incoming IP addresses. Find the `pg_hba.conf` configuration file using the following query: ``` SHOW hba_file; ``` Open `pg_hba.conf` in a text editor of your choice, for example, Visual Studio Code. ``` code pg_hba.conf ``` Under `IPv4 local connections`, find and replace the IP address with `0.0.0.0/0`. ``` # TYPE DATABASE USER ADDRESS METHOD # IPv4 local connections: host all all 0.0.0.0/0 md5 # IPv6 local connections: host all all ::/0 md5 ``` For more details on the configuration file's syntax, see [The pg\_hba.conf File](https://www.postgresql.org/docs/14/auth-pg-hba-conf.html). * Enable the logical replication. For cloud-hosted databases, the logical replication is usually enabled by default, while databases hosted on-premises can have the logical replication not enabled. Check that the logical replication is enabled using `psql` to run the following query: ``` SHOW wal_level; ``` Expected output if enabled ``` wal_level ----------- logical (1 row) ``` If the command prompt returns something different, enable the logical replication in your database by setting `wal_level` to `logical`: ``` ALTER SYSTEM SET wal_level = logical; ``` * Set the maximum number of replication slots to a value that is equal to or greater than the number of databases in the PostgreSQL server. Check the current status using the following query: ``` SHOW max_replication_slots; ``` You can expect the following output: ``` max_replication_slots ----------- (1 row) ``` If `number of slots` is smaller than the number of databases in your PostgreSQL server, modify it using the following query: ``` ALTER SYSTEM SET max_replication_slots = use_your_number; ``` where `use_your_number` stands for the number of databases in your server. * Restart your PostgreSQL server using the following command: ``` sudo service postgresql restart ``` ## Migrate a database[​](#migrate-in-console "Direct link to Migrate a database") 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. On the **Services** page, select the service where your target database is located. 3. From the sidebar on your service's page, select **Service settings**. 4. On the **Service settings** page, go to the **Service management** section, and select **Import database**. 5. Guided by the migration wizard, go through all the migration steps. ### Step 1: Configure[​](#step-1-configure "Direct link to Step 1: Configure") Get familiar with the guidelines provided in the **PostgreSQL migration configuration guide** window, make sure your configuration is in line with them, and select **Get started**. ### Step 2: Validation[​](#step-2-validation "Direct link to Step 2: Validation") 1. To establish a connection to your source database, enter required database details in the **Database connection and validation** window: * Hostname * Port * Database name * Username * Password 2. Select the **SSL encryption (recommended)** checkbox. 3. Optionally, exclude specific databases from the migration by entering their names (separated with spaces) into the **Exclude databases** field. 4. Select **Run check**. Cannot migrate the database using logical replication? If your connection test returns information that you cannot migrate the database using the logical replication due to the missing superuser permissions or `aiven_extras` extension, you can still migrate your data using the dump method. To start a dump, select checkbox **Start the migration using a one-time snapshot (dump method)**. ### Step 3: Migration[​](#step-3-migration "Direct link to Step 3: Migration") If all the checks pass with no error messages, you are ready to start the migration. Before you do that, be aware of its limitations and consequences. Impact on target databases It's recommended to migrate into an empty database. If you migrate into a populated database, colliding tables with primary keys are not affected, but tables without primary keys are appended. Check other limitations in [Logical replication restrictions](https://www.postgresql.org/docs/current/logical-replication-restrictions.html). Trigger the migration by selecting **Start migration** in the **Database migration** window. While the migration is in progress, you can take the following actions: * Let it proceed until completed by selecting **Close window**, which closes the wizard. You can come back to check the status at any time on the **Service settings** page > the **Service management** section > **Import database**. * Write to the target database. * Discontinue the migration by selecting **Stop migration**. Although the data already migrated is retained, you cannot restart the stopped process. To continue with the migration, start a new migration process from scratch. warning To avoid conflicts and replication issues while the migration is ongoing, take the following precautions: * Do not write to any tables in the target database that are being processed by the migration tool. * Do not change the replication configuration of the source database manually. Do not modify `wal_level` or reduce `max_replication_slots`. * Do not make database changes that can disrupt or prevent the connection between the source database and the target database. Do not change the listen address of the source database and do not modify or enable firewalls on the databases. Migration attempt failed? If you happen to get such a notification, investigate potential causes of the failure and try to fix the issues. When you are ready, trigger the migration again by selecting **Start over**. ### Step 4: Close[​](#step-4-close "Direct link to Step 4: Close") As soon as the wizard communicates the completion of the migration, check if there's also information about the replication mode being active. Replication mode active This information in the wizard means that your data has been transferred to Aiven, but some new data is still continuously being synced between the connected databases. * If there is no replication in progress, select **Close connection** in the migration wizard to finalize the migration process. As a result, on the **Service settings** page > the **Service management** section > **Import database**, you'll see the **Ready** tag. * If the replication mode is active, you can select **Keep replicating**. As a result, on the **Service settings** page > the **Service management** section > **Import database**, you'll see the **Syncing** tag, and you'll be able to see the status of the migration process by selecting **Status update**. You have successfully migrated your PostgreSQL database into you Aiven for PostgreSQL service. Related pages * [About aiven-db-migrate](/docs/products/postgresql/concepts/aiven-db-migrate.md) * [Migrate to Aiven for PostgreSQL® with aiven-db-migrate](/docs/products/postgresql/howto/migrate-aiven-db-migrate.md) * [Migrate to Aiven for PostgreSQL® with `pg_dump` and `pg_restore`](/docs/products/postgresql/howto/migrate-pg-dump-restore.md) * [Migrate between PostgreSQL® instances using aiven-db-migrate in Python](/docs/products/postgresql/howto/run-aiven-db-migrate-python.md) * [Migrate to Aiven for MySQL from an external MySQL](/docs/products/mysql/howto/migrate-from-external-mysql.md) --- # Migrate PostgreSQL® databases to Aiven using pg\_dump and pg\_restore Aiven for PostgreSQL® supports the same tools as a regular PostgreSQL database, so you can migrate using the standard `pg_dump` and `pg_restore` tools. tip We recommend to migrate your PostgreSQL® database to Aiven by using [aiven-db-migrate](/docs/products/postgresql/howto/migrate-aiven-db-migrate.md). The [`pg_dump`](https://www.postgresql.org/docs/current/app-pgdump.html) tool can be used to extract the data from your existing PostgreSQL database and [`pg_restore`](https://www.postgresql.org/docs/current/app-pgrestore.html) can then insert that data into your Aiven for PostgreSQL database. The duration of the process depends on the size of your existing database. During the migration no new data written to the database is included. You should turn off all write operations to your source database server before you run the `pg_dump`. tip You can keep the write operations enabled, and use the steps to try out the migration process first before the actual migration. This way, you will find out about the duration, and ensure everything works without downtime. ## Variables[​](#variables "Direct link to Variables") You can use the following variables in the code samples provided: | Variable | Description | | ------------------ | --------------------------------------------------------------------------------------- | | `SRC_SERVICE_URI` | Service URI for the source PostgreSQL connection | | `DUMP_FOLDER` | Local Folder used to store the source database dump files | | `DEST_PG_NAME` | Name of the destination Aiven for PostgreSQL service | | `DEST_PG_PLAN` | Aiven plan for the destination Aiven for PostgreSQL service | | `DEST_SERVICE_URI` | Service URI for the destination PostgreSQL connection, available from the Aiven Console | ## Perform the migration[​](#perform-the-migration "Direct link to Perform the migration") 1. If you don't have an Aiven for PostgreSQL database yet, run the following command to create a couple of PostgreSQL services via [Aiven CLI](/docs/tools/cli.md) substituting the parameters accordingly: ``` avn service create --project PROJECT_NAME -t pg -p DEST_PG_PLAN DEST_PG_NAME ``` tip Aiven for PostgreSQL allows you to switch between different service plans, but during the initial migration process using `pg_dump`, we recommend that you choose a service plan that is large enough for the task. This allows you to limit downtime during the migration process. Once migrated, you can scale the plan size up or down as needed. Aiven automatically creates a `defaultdb` database and `avnadmin` user account, which are used by default. 2. Run the `pg_dump` command substituting the `SRC_SERVICE_URI` with the service URI of your source PostgreSQL service, and `DUMP_FOLDER` with the folder where to store the dump in: ``` pg_dump -d 'SRC_SERVICE_URI' --jobs 4 --format directory -f DUMP_FOLDER ``` The `--jobs` option in this command instructs the operation to use 4 CPUs to dump the database. Depending on the number of CPUs you have available, you can use this option to adjust the performance to better suit your server. tip If you encounter problems with restoring your previous object ownerships to users that do not exist in your Aiven database, use the `--no-owner` option in the `pg_dump` command. You can create the ownership hierarchy after the data is migrated. 3. Run `pg_restore` to load the data into the new database: ``` pg_restore -d 'DEST_SERVICE_URI' --jobs 4 DUMP_FOLDER ``` note If you have more than one database to migrate, repeat the `pg_dump` and `pg_restore` steps for each database. 4. Switch the connection settings in your applications to use the new Aiven database once you have migrated all of your data. warning The user passwords are different from those on the server that you migrated from. In the **Connect** section, go to **Users** in the Aiven Console to check the new passwords. 5. Connect to the target database via `psql`: ``` psql 'DEST_SERVICE_URI' ``` 6. Run the `ANALYZE` command to apply proper database statistics for the newly loaded data: ``` newdb=> ANALYZE; ``` If you got this far, then all went well and your Aiven for PostgreSQL database is now ready to use. ## Handle `pg_restore` errors[​](#handle-pg_restore-errors "Direct link to handle-pg_restore-errors") When migrating PostgreSQL databases to Aiven via `pg_restore` you can encounter errors like: ``` could not execute query: ERROR: must be owner of extension ``` For example, the following `pg_restore` error appears quite commonly: ``` pg_restore: [archiver (db)] could not execute query: ERROR: must be owner of extension ``` This type of error is often related to the lack of superuser-level privileges blocking non-essential queries. A typical example is due to failing `COMMENT ON EXTENSION` queries trying to replace the documented comment string for an extension. In such cases, the errors are harmless and can be ignored. Alternatively, use the `--no-comments` parameter in `pg_restore` to skip these queries. tip `pg_restore` offers similar `--no-XXX` options to switch off other, often unnecessary restore queries. More information is available in the [PostgreSQL documentation](https://www.postgresql.org/docs/current/app-pgrestore.html). ### Poor performance after migration[​](#poor-performance-after-migration "Direct link to Poor performance after migration") Whenever you load data with the `pg_restore` or similar tools, it is recommended to run `ANALYZE` or `VACUUM ANALYZE` on your entire database to collect new statistics. The database will not have up-to-date statistics on the tables and indexes without these operations. In turn, this may lead to poor query plans and poor database performance. Generally, the Aiven platform automatically runs `ANALYZE` on your service after performing a major version upgrade to ensure the statistics are up-to-date. For more information about `ANALYZE`, you may see the official [SQL analyze](https://www.postgresql.org/docs/current/sql-analyze.html) documentation. --- # Migrate PostgreSQL® databases to Aiven using Bucardo The preferred approach to migrating a database to Aiven for PostgreSQL® is to use Aiven's open source migration tool ([About aiven-db-migrate](/docs/products/postgresql/concepts/aiven-db-migrate.md)). However, if you are running PostgreSQL 9.6 (or earlier) or do not have `superuser` access to your database to add replication slots, you can use the open source [Bucardo](https://bucardo.org) tool to allow replication to Aiven. **Requirements:** * An Aiven for PostgreSQL database * Your current database * A computer with Bucardo installed * Connectivity between Bucardo and the source and target databases ## Moving existing data[​](#moving-existing-data "Direct link to Moving existing data") To move existing data, you can follow the steps below and [update](https://bucardo.org/Bucardo/operations/onetimecopy) your `sync` job to use the `onetimecopy` and move existing data across. You can also use the standard `pg_dump` and `pg_restore` commands to fill the Aiven database and use Bucardo for syncing any changes to the source database and ensuring it remains up-to-date. ## Replicating changes[​](#replicating-changes "Direct link to Replicating changes") To migrate your data using Bucardo: 1. Install Bucardo using [the installation instructions](https://bucardo.org/Bucardo/installation/) on the Bucardo site. 2. Install the `aiven_extras` [extension](/docs/products/postgresql/concepts/dba-tasks-pg.md) to your current database. Bucardo requires the superuser role to set the `session_replication_role` parameter. Aiven uses the open source `aiven_extras` extension to allow you to run `superuser` commands as a different user, as direct `superuser` access is not provided for security reasons. 3. Open and edit the `Bucardo.pm` file with administrator privileges. The location of the file can vary according to your operating system, but you might find it in `/usr/local/share/perl5/5.32/Bucardo.pm`, for example. 1. Scroll down until you see a `disable_triggers` function, in line 5324 in [Bucardo.pm](https://github.com/bucardo/bucardo/blob/1ff4d32d1924f3437af3fbcc1a50c1a5b21d5f5c/Bucardo.pm). 2. In line 5359 in [Bucardo.pm](https://github.com/bucardo/bucardo/blob/1ff4d32d1924f3437af3fbcc1a50c1a5b21d5f5c/Bucardo.pm), change `SET session_replication_role = default` to the following: ``` $dbh->do(q{select aiven_extras.session_replication_role('replica');}); ``` 3. Scroll down to the `enable_triggers` function in line 5395 in [Bucardo.pm](https://github.com/bucardo/bucardo/blob/1ff4d32d1924f3437af3fbcc1a50c1a5b21d5f5c/Bucardo.pm). 4. On line 5428, change `SET session_replication_role = default` to the following: ``` $dbh->do(q{select aiven_extras.session_replication_role('origin');}); ``` 5. Save your changes and close the file. 4. Add your source and destination databases. For example: ``` bucardo add db srcdb dbhost=0.0.0.0 dbport=5432 dbname=all_your_base dbuser=$DBUSER dbpass=$DBPASS bucardo add db destdb dbhost=cg-pg-dev-sandbox.aivencloud.com dbport=21691 dbname=all_your_base dbuser=$DBUSER dbpass=$DBPASS ``` 5. Add the tables to replicate: ``` bucardo add table belong to us herd=$HERD db=srcdb ``` note You can set `$HERD` to any name you choose for the herd, which is used to set up the synchronization. 6. Dump and restore the database from your source to your Aiven service: ``` pg_dump --schema-only --no-owner all_your_base > base.sql psql "$AIVEN_DB_URL" < base.sql ``` You can restore the source data or provide only the schema, depending on the size of your current database. 7. Create the `dbgroup` for Bucardo: ``` bucardo add dbgroup src_to_dest srcdb:source destdb:target bucardo add sync sync_src_to_dest relgroup=$HERD db=srcdb,destdb (sudo) bucardo start bucardo status sync_src_to_dest ``` 8. Start Bucardo and run the `status` command. When `Current state` is `Good`, the data is flowing to your Aiven database. 9. Log in to the [Aiven Console](https://console.aiven.io), select your Aiven for PostgreSQL service from the **Services** list, and in the **Observe** section, click **Current queries** in your service's page. This shows you that the `bucardo` process is inserting data. 10. Once all your data is synchronized, switch the database connection for your applications to Aiven for PostgreSQL. --- # Monitor a database with Datadog [Database Monitoring with Datadog](https://www.datadoghq.com/product/database-monitoring/) enables you to capture key metrics on the Datadog platform for any Aiven for PostgreSQL® service with [Datadog Metrics](/docs/integrations/datadog/datadog-metrics.md) integration. Datadog Database Monitoring allows you to view query metrics and explain plans in a single place, with the ability to drill into precise execution details, along with query and host metrics correlation. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Apply any outstanding maintenance updates mentioning the Datadog integration. * Ensure the [Datadog Metrics integration](/docs/integrations/datadog/datadog-metrics.md) is enabled. * The [PostgreSQL extensions](/docs/products/postgresql/reference/list-of-extensions.md) - `pg_stat_statements` and `aiven_extras`, must be enabled by executing the following [CREATE EXTENSION](https://www.postgresql.org/docs/current/sql-createextension.html) SQL commands directly on the Aiven for PostgreSQL® database service. ``` CREATE EXTENSION pg_stat_statements; CREATE EXTENSION aiven_extras; ``` ## Enable monitoring[​](#enable-monitoring "Direct link to Enable monitoring") You can individually enable Datadog Database Monitoring for the specific [Datadog Metrics](/docs/integrations/datadog/datadog-metrics.md) integration for Aiven for PostgreSQL®, by configuring the `datadog_dbm_enabled` parameter. Repeat this action for every Datadog Metrics integration for Aiven for PostgreSQL®, which you plan to monitor. Using the `avn service integration-list` [Aiven CLI command](/docs/tools/cli/service/integration.md#avn_service_integration_list), you can obtain the Datadog Metric integration to monitor and enable the Datadog Database monitoring functionality by using the `datadog_dbm_enabled` configuration parameter. For example: * Find the UUID of the Datadog Metrics integration for a particular service: ``` avn service integration-list --project ``` * Enable the Datadog Database Monitoring for the Datadog Metrics integration with the following command, substituting the `` with the integration UUID retrieved at the previous step: ``` avn service integration-update --project --user-config '{"datadog_dbm_enabled": true}' ``` * Check if user-config `datadog_dbm_enabled` set correctly: ``` avn service integration-list \ --project \ --json | jq '.[] | select(.integration_type=="datadog").user_config' ``` `datadog_dbm_enabled` should be set to `true`: ``` { "datadog_dbm_enabled": true } ``` Related pages * Learn more about [Datadog and Aiven](/docs/integrations/datadog.md). * [Monitor PgBouncer with Datadog](/docs/products/postgresql/howto/monitor-pgbouncer-with-datadog.md). * [Collect relation and function metrics with Datadog](/docs/products/postgresql/howto/monitor-relation-function-metrics-datadog.md). * Learn more about [Database monitoring with Datadog](https://www.datadoghq.com/product/database-monitoring/) from the Datadog product page. --- # Monitor PgBouncer with Datadog for Aiven for PostgreSQL® Integrate PgBouncer with Datadog to track connection pool metrics and monitor application traffic on the Datadog platform. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Your service plan is Startup or higher. * You applied outstanding service maintenance updates affecting the Datadog Metrics integration. ## Enable monitoring[​](#enable-monitoring "Direct link to Enable monitoring") You can enable monitoring PgBouncer metrics for your Aiven for PostgreSQL service both [if the service already has a Datadog Metrics integration](#enable-monitoring-if-the-integration-exists) and [if it doesn't](#create-the-integration-and-enable-monitoring). ### Enable monitoring for an integrated service[​](#enable-monitoring-if-the-integration-exists "Direct link to Enable monitoring for an integrated service") Enable monitoring PgBouncer metrics for an Aiven for PostgreSQL service that already has a Datadog Metrics integration: 1. Obtain the `SERVICE_INTEGRATION_ID` of the Datadog Metrics integration for your Aiven for PostgreSQL service by running the [avn service integration-list](/docs/tools/cli/service/integration.md#avn_service_integration_list) command: ``` avn service integration-list SERVICE_NAME \ --project PROJECT_NAME ``` 2. To enable PgBouncer monitoring in Datadog, set up the `datadog_pgbouncer_enabled` parameter to `true`: ``` avn service integration-update SERVICE_INTEGRATION_ID \ --project PROJECT_NAME \ --user-config '{"datadog_pgbouncer_enabled": true}' ``` Replace SERVICE\_INTEGRATION\_ID with the service integration identifier acquired in the preceding step. ### Enable monitoring for a non-integrated service[​](#create-the-integration-and-enable-monitoring "Direct link to Enable monitoring for a non-integrated service") To enable monitoring PgBouncer metrics for an Aiven for PostgreSQL service that doesn't have a Datadog Metrics integration, [create the integration](/docs/tools/cli/service/integration.md#avn_service_integration_create) and enable monitoring by running: ``` avn service integration-create INTEGRATION_CREATE_PARAMETERS \ --user-config-json '{"datadog_pgbouncer_enabled": true}' ``` Replace INTEGRATION\_CREATE\_PARAMETERS with [the parameters required to create the Datadog Metrics integration](/docs/tools/cli/service/integration.md#avn_service_integration_create). ## Verify the changes[​](#verify-the-changes "Direct link to Verify the changes") Check that the `datadog_pgbouncer_enabled` user-config is set correctly: ``` avn service integration-list SERVICE_NAME \ --project PROJECT_NAME \ --json | jq '.[] | select(.integration_type=="datadog").user_config' ``` Expect the following output confirming that `datadog_pgbouncer_enabled` is set to `true`: ``` { "datadog_pgbouncer_enabled": true } ``` Related pages * [Database monitoring with Datadog](/docs/products/postgresql/howto/monitor-database-with-datadog.md) * [Access PgBouncer statistics](/docs/products/postgresql/howto/pgbouncer-stats.md) * [Datadog and Aiven](/docs/integrations/datadog.md) * [Create service integrations](/docs/platform/howto/create-service-integration.md) --- # Collect relation and function metrics with Datadog for Aiven for PostgreSQL® Configure the Datadog Metrics integration to collect per-table and per-index statistics, and per-function call statistics for PL/pgSQL functions, for your Aiven for PostgreSQL® service. Both options are off by default. Datadog bills relation and function metrics as custom metrics, and a single integration can produce many of them, so you opt in and choose the scope yourself. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * The [Datadog Metrics integration](/docs/integrations/datadog/datadog-metrics.md) is enabled for the service. * For function metrics only: the `track_functions` parameter is set to `pl` or `all` in your service's advanced configuration. Without it, PostgreSQL collects no function statistics, and `datadog_function_metrics_enabled` has no effect. ## Find the integration ID[​](#find-the-integration-id "Direct link to Find the integration ID") Both settings apply to a specific Datadog Metrics integration. Find its ID by running the [avn service integration-list](/docs/tools/cli/service/integration.md#avn_service_integration_list) command: ``` avn service integration-list --project PROJECT_NAME SERVICE_NAME ``` Use the `service_integration_id` value from the output as `INTEGRATION_ID` in the following commands. note If saving `datadog_pg_relations`, `datadog_function_metrics_enabled`, or `datadog_pg_dbname` fails, your service is pending a scheduled maintenance update. Apply the update, or wait for your next maintenance window, then try again. ## Choose the monitored database[​](#choose-the-monitored-database "Direct link to Choose the monitored database") The Datadog PostgreSQL check connects to the service's main database, so relation and function metrics come from that database only. To collect these metrics from a different database, set `datadog_pg_dbname`: ``` avn service integration-update --project PROJECT_NAME \ --user-config-json '{"datadog_pg_dbname": "DATABASE_NAME"}' \ INTEGRATION_ID ``` `datadog_pg_dbname` scopes relation and function metrics only. Database Monitoring collects query statistics from every database on the service regardless of this option. The name must be 1-63 characters, start with a letter, digit, or underscore, and otherwise contain only letters, digits, underscores, and hyphens. The database doesn't need to exist yet. If it doesn't, the Datadog agent logs a connection error until you create it, then reports metrics with no further configuration change needed. ## Collect relation metrics[​](#collect-relation-metrics "Direct link to Collect relation metrics") Set `datadog_pg_relations` to the relations you want metrics for. Each entry selects relations either by exact name or by regular expression: ``` avn service integration-update --project PROJECT_NAME \ --user-config-json '{ "datadog_pg_relations": [ {"relation_name": "orders"}, {"relation_name": "events", "schemas": ["public", "analytics"]}, {"relation_regex": "^metrics_.*", "relkind": ["r", "p"]} ] }' \ INTEGRATION_ID ``` This example collects metrics for the `orders` relation in all schemas, for the `events` relation in the `public` and `analytics` schemas, and for any relation matching `^metrics_.*`, with lock metrics limited to ordinary tables (`r`) and partitioned tables (`p`). Each entry accepts the following fields. | Field | Description | | ---------------- | ---------------------------------------------------------------------------------- | | `relation_name` | Name of a single relation, up to 63 characters. | | `relation_regex` | Regular expression matching relation names, up to 128 characters. | | `schemas` | Restricts the entry to these schemas, up to 8. Applies to all schemas when unset. | | `relkind` | Relation kinds that **lock** metrics cover. Applies to ordinary tables when unset. | Set exactly one of `relation_name` or `relation_regex` on each entry. Setting both, setting neither, or using a regular expression that doesn't compile fails validation when you save the configuration. `relation_name` and each entry in `schemas` must start with a letter, a digit, or an underscore, and can otherwise contain only letters, digits, underscores, and dollar signs. `relkind` takes PostgreSQL `pg_class` relation kinds and affects lock metrics only. Other relation metrics follow the name or regular expression match regardless of `relkind`. | Value | Relation kind | | ----- | ----------------- | | `r` | Ordinary table | | `i` | Index | | `S` | Sequence | | `t` | TOAST table | | `m` | Materialized view | | `c` | Composite type | | `f` | Foreign table | | `p` | Partitioned table | You can configure up to 32 entries in `datadog_pg_relations`. ## Collect function metrics[​](#collect-function-metrics "Direct link to Collect function metrics") Requires `track_functions` set to `pl` or `all` on the service, as described in [Prerequisites](#prerequisites). ``` avn service integration-update --project PROJECT_NAME \ --user-config-json '{"datadog_function_metrics_enabled": true}' \ INTEGRATION_ID ``` To confirm PostgreSQL is tracking functions before checking Datadog, connect to your database and run the following query: ``` SELECT funcname, calls FROM pg_stat_user_functions; ``` An empty result means `track_functions` is off, or the functions aren't written in PL/pgSQL. With `track_functions` set to `pl`, only PL/pgSQL functions are counted. Set it to `all` to also count SQL and C functions, at a higher overhead. ## Verify the configuration[​](#verify-the-configuration "Direct link to Verify the configuration") ``` avn service integration-list SERVICE_NAME \ --project PROJECT_NAME \ --json | jq '.[] | select(.integration_type=="datadog").user_config' ``` Updates merge into the existing configuration, so settings you don't mention are preserved. `datadog_pg_relations` is replaced as a whole and not appended to, so send the complete list every time you change it. ## Metrics collected[​](#metrics-collected "Direct link to Metrics collected") With relations configured, Datadog reports per-relation metrics including table size, index scans, index rows read and fetched, index blocks hit and read, live and dead row counts, time since the last autovacuum and autoanalyze operations, and lock counts. Each metric is tagged with the relation name. With function metrics enabled, Datadog reports call counts, and total and self execution time, for each PL/pgSQL function, tagged with the function name. Find these metrics in Datadog's Metrics Explorer under the `postgresql.` prefix. Related pages * [Database monitoring with Datadog](/docs/products/postgresql/howto/monitor-database-with-datadog.md) * [Monitor PgBouncer with Datadog](/docs/products/postgresql/howto/monitor-pgbouncer-with-datadog.md) * [Datadog and Aiven](/docs/integrations/datadog.md) * [PostgreSQL® metrics exposed in Grafana®](/docs/products/postgresql/reference/pg-metrics.md) --- # Monitor PostgreSQL® metrics with pgwatch2 [pgwatch2](https://github.com/cybertec-postgresql/pgwatch2) is an open source monitoring solution for PostgreSQL®, created by CYBERTEC and can be used to monitor instances of Aiven for PostgreSQL collecting key PostgreSQL metrics and also gathering data from a wide range of PostgreSQL extensions. note Aiven for PostgreSQL supports the most popular extensions, but, due to security implications, can't support all the available extensions. Check [Extensions on Aiven for PostgreSQL®](/docs/products/postgresql/reference/list-of-extensions.md) and [Extensions on Aiven for PostgreSQL®](/docs/products/postgresql/reference/list-of-extensions.md) for details on the supported PostgreSQL extensions on the Aiven platform. ## Prepare an Aiven for PostgreSQL instance for pgwatch2[​](#prepare-an-aiven-for-postgresql-instance-for-pgwatch2 "Direct link to Prepare an Aiven for PostgreSQL instance for pgwatch2") On the Aiven for PostgreSQL instance to be monitored with pgwatch2: 1. Create a user for pgwatch2: ``` CREATE USER pgwatch2 WITH PASSWORD 'password'; ``` 2. Limit the number of connections from pgwatch2 (optional, but recommended in the [pgwatch2 documentation](https://pgwatch2.readthedocs.io/en/latest/)): ``` ALTER ROLE pgwatch2 CONNECTION LIMIT 3; ``` 3. Allow pgwatch2 to read database statistics: ``` GRANT pg_read_all_stats TO pgwatch2; ``` 4. To collect data gathered from the `pg_stat_statements` extension, enable it: ``` CREATE EXTENSION IF NOT EXISTS pg_stat_statements; ``` 5. The [pgwatch2 documentation](https://pgwatch2.readthedocs.io/en/latest/) recommends to enable timing of database I/O calls by setting the PostgreSQL configuration parameter `track_io_timing` (see [Extensions on Aiven for PostgreSQL®](/docs/products/postgresql/reference/list-of-extensions.md)). warning According to the [PostgreSQL documentation](https://www.postgresql.org/docs/current/runtime-config-statistics.html), setting `track_io_timing = on` can cause significant overhead. ## Running pgwatch2[​](#running-pgwatch2 "Direct link to Running pgwatch2") pgwatch2 has multiple [installation options](https://pgwatch2.readthedocs.io/en/latest/installation_options.html) to choose from. For the sake of simplicity, the following example uses [ad-hoc mode](https://pgwatch2.readthedocs.io/en/latest/installation_options.html#ad-hoc-mode) with a Docker container: ``` docker run --rm -p 3000:3000 -p 8080:8080 \ -e PW2_ADHOC_CONN_STR='postgres://pgwatch2:password@HOST:PORT/defaultdb?sslmode=require' \ -e PW2_ADHOC_CONFIG='rds' --name pw2 cybertec/pgwatch2-postgres ``` This runs pgwatch2 with the container image provided by CYBERTEC. `PW2_ADHOC_CONN_STR` is set to the connection string of the PostgreSQL instance to be monitored, copied from the [Aiven web console](https://console.aiven.io/) replacing the username/password have been replaced by the ones specifically created for pgwatch2. See the [pgwatch2 documentation](https://pgwatch2.readthedocs.io/en/latest/) to decide on the best way to set up pgwatch2 in your environment. note pgwatch2 contains several dashboards that rely on extensions not available in Aiven for PostgreSQL, so it is to be expected that some dashboards are either empty or display error symbols. ![Screenshot of a pgwatch2 Dashboard](/docs/assets/images/pgwatch2-befa80ab77dfe69be53b2a55db8accdd.png) --- # Optimize Aiven for PostgreSQL® slow queries Aiven for PostgreSQL allows you to [identify slow queries](/docs/products/postgresql/howto/identify-pg-slow-queries.md) using the `pg_stat_statements` view. important You can also use [Aiven's AI capabilities](/docs/products/postgresql/howto/ai-insights.md) to identify and speed up slow queries. ## Limit the number of indexes[​](#limit-the-number-of-indexes "Direct link to Limit the number of indexes") Having many database indexes on a table can reduce write performance due to the overhead of maintaining them. ## Handle an increase in database connections[​](#handle-an-increase-in-database-connections "Direct link to Handle an increase in database connections") When your application code scales horizontally to accommodate high loads, you might find that you inadvertently reach the [connection limits](/docs/products/postgresql/reference/pg-connection-limits.md) for your service. Each connection in PostgreSQL runs in a separate process, and this makes them more expensive (compared to threads, for example) in terms of inter-process communication and memory usage, since each connection consumes a certain amount of RAM. In such cases, you can use the [connection pooling](/docs/products/postgresql/concepts/pg-connection-pooling.md), based on [PgBouncer](https://www.pgbouncer.org), to handle an increase in database connections. You can add and configure the connection pooling for your service in the **Connection pools** view in [Aiven Console](https://console.aiven.io/). ## Move read-only queries to standby nodes[​](#move-read-only-queries-to-standby-nodes "Direct link to Move read-only queries to standby nodes") If your Aiven for PostgreSQL® service has [standby nodes](/docs/products/postgresql/concepts/high-availability.md), you can reduce the effect of slow queries on the primary node by redirecting read-only queries to the additional [read-only](/docs/products/postgresql/howto/create-read-replica.md) nodes by directly connecting via the **read-only replica URL**. ## Move read-only queries to a remote read-only replica[​](#move-read-only-queries-to-a-remote-read-only-replica "Direct link to Move read-only queries to a remote read-only replica") You can also create a [remote read-only replica](/docs/products/postgresql/howto/create-read-replica.md) service in the same or a different cloud or region that you can use to reduce the query load on the primary service for read-only queries. Related pages * [AI DB Optimizer for Aiven for PostgreSQL®](/docs/products/postgresql/howto/ai-insights.md) * [Standalone query optimizer](/docs/tools/query-optimizer.md) * [Identify PostgreSQL® slow queries with \`pg\_stat\_statements](/docs/products/postgresql/howto/identify-pg-slow-queries.md) --- # Sample dataset for PostgreSQL®: Pagila Use a sample database that you can import in your Aiven for PostgreSQL® service. Pagila is a PostgreSQL® port of the [Sakila Sample Database](https://dev.mysql.com/doc/sakila/en/). The examples use one from `devrimgunduz`, [version 3.1.0](https://github.com/devrimgunduz/pagila). Sakila/Pagila is a database representing a DVD rental store, containing information about films, rental stores and rentals, where a customer rents a film from a store through its staff. With all this relational information, Pagila is a perfect fit to play around with PostgreSQL and the SQL language. ## Load Pagila to your Aiven for PostgreSQL service[​](#load-pagila-to-your-aiven-for-postgresql-service "Direct link to Load Pagila to your Aiven for PostgreSQL service") Before exploring the Pagila database, follow the create a service to start a PostgreSQL instance. 1. Download `pagila-schema.sql` and `pagila-data.sql` from the [devrimgunduz/pagila](https://github.com/devrimgunduz/pagila/tree/master) repository. tip You may use the following command on your terminal: ``` wget https://raw.githubusercontent.com/devrimgunduz/pagila/refs/heads/master/pagila-schema.sql wget https://raw.githubusercontent.com/devrimgunduz/pagila/refs/heads/master/pagila-data.sql ``` 2. Connect to the PostgreSQL instance using the following command. The `SERVICE_URI` value can be found in the Aiven Console dashboard. ``` psql 'SERVICE_URI' ``` 3. Within the `psql` shell, create a database named `pagila` and connect to it with the command below: ``` CREATE DATABASE pagila; \c pagila; ``` 4. Populate the database with the command below. This might take some time. ``` \i pagila-schema.sql; \i pagila-data.sql; ``` 5. Once the command finishes, make sure to reconnect to the database to access the imported data: ``` \c pagila; ``` ## Entity-relationship model diagram[​](#entity-relationship-model-diagram "Direct link to Entity-relationship model diagram") The image below shows an overview of the Pagila database tables and views, generated by [DBeaver](https://dbeaver.io). For example, the `film` table has string columns like `title` and `description`. It also relates to the table `language` with the columns `language_id` and `original_language_id`. With that information, you know that you can join both tables to get the language of each film, or to list all films for a specific language. ![A entity-relation model diagram for the Pagila databases, containing all the tables, fields and views.](/docs/assets/images/pagila-erm-99f1ecb38332114572b48c958aacf472.png) ## Sample queries[​](#sample_queries "Direct link to Sample queries") Let's explore the dataset with a few queries. All the queries results were limited by the first 10 items. List all the films by ordered by their length ``` select film_id, title, length from film order by length desc; ``` ``` | film_id | title | length | | ------- | ------------------ | ------ | | 426 | HOME PITY | 185 | | 690 | POND SEATTLE | 185 | | 609 | MUSCLE BRIGHT | 185 | | 991 | WORST BANGER | 185 | | 182 | CONTROL ANTHEM | 185 | | 141 | CHICAGO NORTH | 185 | | 349 | GANGS PRIDE | 185 | | 212 | DARN FORRESTER | 185 | | 817 | SOLDIERS EVOLUTION | 185 | | 872 | SWEET BROTHERHOOD | 185 | ``` List how many films there are in each film category ``` select category.name, count(category.name) category_count from category left join film_category on category.category_id = film_category.category_id left join film on film_category.film_id = film.film_id group by category.name order by category_count desc; ``` ``` | name | category_count | | ----------- | -------------- | | Sports | 74 | | Foreign | 73 | | Family | 69 | | Documentary | 68 | | Animation | 66 | | Action | 64 | | New | 63 | | Drama | 62 | | Sci-Fi | 61 | | Games | 61 | ``` Show the actors and actresses ordered by how many movies they are featured in ``` select actor.first_name, actor.last_name, count(actor.first_name) featured_count from actor left join film_actor on actor.actor_id = film_actor.actor_id group by actor.first_name, actor.last_name order by featured_count desc; ``` ``` | first_name | last_name | featured_count | | ---------- | --------- | -------------- | | SUSAN | DAVIS | 54 | | GINA | DEGENERES | 42 | | WALTER | TORN | 41 | | MARY | KEITEL | 40 | | MATTHEW | CARREY | 39 | | SANDRA | KILMER | 37 | | SCARLETT | DAMON | 36 | | VIVIEN | BASINGER | 35 | | VAL | BOLGER | 35 | | GROUCHO | DUNST | 35 | ``` Get a list of all active customers, ordered by their first name ``` select first_name, last_name from customer where active = 1 order by first_name asc; ``` ``` | first_name | last_name | | ---------- | --------- | | MARY | SMITH | | PATRICIA | JOHNSON | | LINDA | WILLIAMS | | BARBARA | JONES | | ELIZABETH | BROWN | | JENNIFER | DAVIS | | MARIA | MILLER | | SUSAN | WILSON | | MARGARET | MOORE | | DOROTHY | TAYLOR | ``` See who rented most DVDs - and how many times ``` select customer.first_name, customer.last_name, count(customer.first_name) rentals_count from customer left join rental on customer.customer_id = rental.customer_id group by customer.first_name, customer.last_name order by rentals_count desc; ``` ``` | first_name | last_name | rentals_count | | ---------- | --------- | ------------- | | ELEANOR | HUNT | 46 | | KARL | SEAL | 45 | | CLARA | SHAW | 42 | | MARCIA | DEAN | 42 | | TAMMY | SANDERS | 41 | | WESLEY | BULL | 40 | | SUE | PETERS | 40 | | MARION | SNYDER | 39 | | RHONDA | KENNEDY | 39 | | TIM | CARY | 39 | ``` ## Challenge[​](#challenge "Direct link to Challenge") After playing around with the sample queries, can you use SQL statements to answer some of these questions? 1. What is the total revenue of each rental store? See answer ``` select store.store_id, sum(payment.amount) as "total revenue" from store left join inventory on inventory.store_id = store.store_id left join rental on rental.inventory_id = inventory.inventory_id left join payment on payment.rental_id = rental.rental_id where payment.amount is not null group by store.store_id order by sum(payment.amount) desc; ``` ``` | store_id | total revenue | | -------- | ------------- | | 2 | 33726.77 | | 1 | 33689.74 | ``` 2. Can you list the top 5 film genres by their gross revenue? See answer ``` select category.name, film.title, sum(payment.amount) as "gross revenue" from film left join film_category on film_category.film_id = film.film_id left join category on film_category.category_id = category.category_id left join inventory on inventory.film_id = film.film_id left join rental on rental.inventory_id = inventory.inventory_id left join payment on payment.rental_id = rental.rental_id where payment.amount is not null group by category.name, film.title order by sum(payment.amount) desc limit 5; ``` ``` | name | title | gross revenue | | ----------- | ----------------- | ------------- | | Music | TELEGRAPH VOYAGE | 231.73 | | Documentary | WIFE TURN | 223.69 | | Comedy | ZORRO ARK | 214.69 | | Sci-Fi | GOODFELLAS SALUTE | 209.69 | | Sports | SATURDAY LAMBS | 204.72 | ``` 3. The `film.description` has the `text` type, allowing for [full text search](https://www.postgresql.org/docs/current/textsearch.html) queries, what will you search for? See answer ``` -- Select all descriptions with the words "documentary" and "robot" select film.title, film.description from film where to_tsvector(film.description) @@ to_tsquery('documentary & robot'); ``` ``` | title | description | | ---------------- | ------------------------------------------------------------------------------------------------------------------ | | CASPER DRAGONFLY | A Intrepid Documentary of a Boat And a Crocodile who must Chase a Robot in The Sahara Desert | | CHAINSAW UPTOWN | A Beautiful Documentary of a Boy And a Robot who must Discover a Squirrel in Australia | | CONTROL ANTHEM | A Fateful Documentary of a Robot And a Student who must Battle a Cat in A Monastery | | CROSSING DIVORCE | A Beautiful Documentary of a Dog And a Robot who must Redeem a Womanizer in Berlin | | KANE EXORCIST | A Epic Documentary of a Composer And a Robot who must Overcome a Car in Berlin | | RUNNER MADIGAN | A Thoughtful Documentary of a Crocodile And a Robot who must Outrace a Womanizer in The Outback | | SOUTH WAIT | A Amazing Documentary of a Car And a Robot who must Escape a Lumberjack in An Abandoned Amusement Park | | SWEDEN SHINING | A Taut Documentary of a Car And a Robot who must Conquer a Boy in The Canadian Rockies | | VIRGIN DAISY | A Awe-Inspiring Documentary of a Robot And a Mad Scientist who must Reach a Database Administrator in A Shark Tank | ``` ## Clean up[​](#clean-up "Direct link to Clean up") To clean up the environment and destroy the database, run the following commands: ``` \c defaultdb; DROP DATABASE pagila; ``` --- # Controlled maintenance updates in Aiven for PostgreSQL® [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Control when primary node switchover happens during Aiven for PostgreSQL® maintenance. important This feature is in [Limited availability](/docs/platform/concepts/service-and-feature-releases.md). ## Benefits and use cases[​](#benefits-and-use-cases "Direct link to Benefits and use cases") Use controlled maintenance updates to reduce impact from primary node promotion: * Align switchover with low-traffic periods. * Keep maintenance predictable for application teams. * Reduce risk during business-critical hours. * Use multiple short windows each day instead of one broad period. Typical use cases are the following: * A global application that needs different low-traffic windows each day. * A service that allows brief disruption only after batch processing ends. * A team that starts maintenance manually, then wants promotion in approved hours. ## How it works[​](#how-it-works "Direct link to How it works") 1. You define one or more daily switchover windows in `user_config.switchover_windows`. 2. Maintenance starts from your regular `maintenance` schedule or from a manual start. 3. Aiven waits for your configured switchover window before promoting the new primary. If nodes are not ready in the first available window, switchover waits for the next configured window. For detailed PostgreSQL failover behavior, see [Aiven for PostgreSQL upgrade and failover procedures](/docs/products/postgresql/concepts/upgrade-failover.md). ## Limitations[​](#limitations "Direct link to Limitations") Controlled switchover applies to the following: * Maintenance started automatically during the configured maintenance window * Maintenance started manually with [ServiceMaintenanceStart](https://api.aiven.io/doc/#tag/Service/operation/ServiceMaintenanceStart) Controlled switchover does not apply to the following: * Plan upgrades * Region migrations * Incident handling, such as node failure or network issues Window limits: * Define at least one window per day. * Define at most four windows per day. * Define at most 28 windows total. * Set each window length to at least 10 minutes. * Use UTC times. * Use `HH:MM:SS` for `start_time` and `end_time`. * Split a window that crosses midnight into two windows. ## Manage controlled switchover[​](#manage-controlled-switchover "Direct link to Manage controlled switchover") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before you configure controlled maintenance updates, ensure the following: * Your service is Aiven for PostgreSQL. * Aiven has enabled this [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) feature for your project. * Your service has a configured maintenance window. * You can update service configuration with [Aiven API endpoints](https://api.aiven.io/doc/#tag/Service) or [Aiven CLI service commands](/docs/tools/cli/service-cli.md). ### Enable and configure[​](#enable-and-configure "Direct link to Enable and configure") Follow these window rules: * Define at least one window per day. * Define at most four windows per day. * Define at most 28 windows total. * Set each window length to at least 10 minutes. * Use UTC times. * Use `HH:MM:SS` for `start_time` and `end_time`. * Split a window that crosses midnight into two windows. `start_time` and `end_time` are inclusive. Use one of the following tools to configure `user_config.switchover_windows`: * CLI * API To generate and apply daily windows with Aiven CLI, run: ``` WINDOWS_JSON="$(jq -c -n \ --argjson window '{"start_time": "09:00:00", "end_time": "10:00:00"}' \ --argjson weekdays '[ "monday", "tuesday", "wednesday", "thursday", "friday", "saturday", "sunday" ]' \ '$weekdays | map($window + {dow: .})' )" avn service update -c "switchover_windows=${WINDOWS_JSON}" SERVICE_NAME ``` Call the [ServiceUpdate endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) to set `user_config.switchover_windows`: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "switchover_windows": [ { "dow": "monday", "start_time": "22:00:00", "end_time": "22:15:00" }, { "dow": "tuesday", "start_time": "23:30:00", "end_time": "23:59:59" }, { "dow": "wednesday", "start_time": "00:00:00", "end_time": "00:30:00" } ] } }' ``` This example sets a Monday maintenance window and three switchover windows: ``` { "maintenance": { "dow": "monday", "time": "08:00:00" }, "user_config": { "switchover_windows": [ { "dow": "monday", "start_time": "22:00:00", "end_time": "22:15:00" }, { "dow": "tuesday", "start_time": "23:30:00", "end_time": "23:59:59" }, { "dow": "wednesday", "start_time": "00:00:00", "end_time": "00:30:00" } ... ] } } ``` If maintenance starts Monday at `08:00:00` UTC and nodes are ready before `22:00:00`, promotion happens in the Monday window. If nodes become ready after that window, promotion waits for the Tuesday or Wednesday window. If you update windows during ongoing maintenance, Aiven uses the latest configuration for the next eligible switchover time. ### Operate and monitor[​](#operate-and-monitor "Direct link to Operate and monitor") To start maintenance manually, use [`avn service maintenance-start`](/docs/tools/cli/service-cli.md#avn-service-maintenance-start). When controlled windows are configured, manual maintenance still follows those windows for promotion. To turn off controlled switchover and allow promotion at any time, set `switchover_windows` to an empty list: ``` { "user_config": { "switchover_windows": [] } } ``` This change takes effect soon after the update. Pending switchovers can proceed without waiting for a window. Check service state with: ``` avn service get SERVICE_NAME --json ``` In the response, inspect `maintenance.controlled_switchover`: ``` { "maintenance": { "controlled_switchover": { "enabled": true, "state": "COMPLETED", "scheduled_start_time": "2026-02-27T10:01:00.000000Z", "completed_at": "2026-02-27T10:02:00.000000Z", "termination_reason": "Controlled switchover" } } } ``` The `controlled_switchover` object includes the following fields: * `enabled`: Whether controlled switchover is configured for the service. * `state`: The current switchover state. See the state values that follow. * `scheduled_start_time`: The time when switchover is scheduled to start. Can be `null` for some states. * `completed_at`: The time when switchover completed. Can be `null` if switchover has not completed. * `termination_reason`: The reason for switchover completion, such as `old generation node timed out`. Can be `null`. State values: * `INACTIVE`: No active maintenance, or feature is not configured. * `PENDING`: Maintenance started, and new nodes are not ready. * `SCHEDULED`: New nodes are ready, and switchover is scheduled. * `RUNNING`: Switchover is in progress. * `COMPLETED`: Switchover finished within a controlled switchover window. * `COMPLETED_EARLY`: Switchover finished outside a controlled switchover window, for example when the old primary node failed during the wait. If the old service nodes time out or fail outside a configured window, Aiven promotes the new primary and completes the switchover early, without waiting for the next window. In this case, `state` is `COMPLETED_EARLY`. ## Best practices[​](#best-practices "Direct link to Best practices") * Use windows longer than 10 minutes to absorb timing variation. * Use at least 15-minute windows to reduce client impact around promotion. * Keep one predictable window per day before adding more windows. * Place windows outside business-critical traffic periods. * Treat midnight windows as two entries, one per day. * Monitor `state` and `scheduled_start_time` during maintenance events. Related pages * [Aiven for PostgreSQL upgrade and failover procedures](/docs/products/postgresql/concepts/upgrade-failover.md) * [Aiven CLI service commands](/docs/tools/cli/service-cli.md) --- # Detect and terminate long-running queries in Aiven for PostgreSQL® Aiven does not terminate any customer queries even if they run indefinitely, but long-running queries can cause issues by locking resources and therefore preventing database maintenance tasks. To identify and terminate such long-running queries, you can do it from either: * [Aiven Console](https://console.aiven.io) * [PostgreSQL® shell](https://www.postgresql.org/docs/current/app-psql.html) (`psql`) ## Terminate long running queries from the Aiven Console[​](#terminate-long-running-queries-from-the-aiven-console "Direct link to Terminate long running queries from the Aiven Console") 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. From the **Services** page, select your Aiven for PostgreSQL service. 3. In your service's page, in the **Observe** section, click **Current queries**. 4. In the **Current queries** page, you can check the query duration and select **Terminate** to stop any long-running queries. ## Detect and terminate long running queries with `psql`[​](#detect-and-terminate-long-running-queries-with-psql "Direct link to detect-and-terminate-long-running-queries-with-psql") You can [login to your service](/docs/products/postgresql/howto/connect-psql.md) by running on the terminal `psql `. Once connected, you can call the following function on the `psql` shell to terminate a query manually: ``` SELECT pg_terminate_backend(pid); ``` You can learn more about the `pg_terminate_backend()` function from the [official documentation](https://www.postgresql.org/docs/current/functions-admin.html). You can then use the following query to monitor currently running queries: ``` SELECT * FROM pg_stat_activity WHERE state <> 'idle'; ``` Client applications can use the `statement_timeout` session variable to voluntarily request the server to automatically cancel any query using the current connection that runs over a specified length of time. For example, the following would cancel any query that runs for more 15 seconds automatically: ``` SET statement_timeout = 15000 ``` You may check the [client connection defaults](https://www.postgresql.org/docs/current/runtime-config-client.html) documentation for more information on the available session variables. ## Database user error[​](#database-user-error "Direct link to Database user error") If you run the above command using a database user not being a member of the database you're connecting to, you will encounter the error: ``` ERROR: must be a member of the role whose process is being terminated or member of pg_signal_backend ``` You can check the roles assigned to each user with the following command: ``` SELECT r.rolname as username,r1.rolname as "role" FROM pg_catalog.pg_roles r JOIN pg_catalog.pg_auth_members m ON (m.member = r.oid) JOIN pg_roles r1 ON (m.roleid=r1.oid) WHERE r.rolcanlogin ORDER BY 1; ``` where you would see the following: ``` username | role ----------+--------------------- avnadmin | pg_read_all_stats avnadmin | pg_stat_scan_tables (3 rows) ``` To be able to check the database owner and grant the role, you can run the following: ``` \l ``` which you should see the role: ``` Name | Owner | -----------+----------+ testdb | testrole | ``` To resolve the permission issue, you may grant the user the appropriate role as per below: ``` grant testrole to avnadmin; ``` --- # Check the size of a database, a table or an index PostgreSQL® offers different commands and functions to get disk space usage for a database, a table, or an index. ## Get the size of a database[​](#get-the-size-of-a-database "Direct link to Get the size of a database") Retrieve the database size using either: * The `\l+ [ pattern ]` command * The `pg_database_size` function. Using the \l+ \[ pattern ] command ``` testdb2=> \l+ List of databases Name | Owner | Encoding | Collate | Ctype | Access privileges | Size | Tablespace | Description -----------+----------+----------+-------------+-------------+-----------------------+-----------+------------+------------------------------------ _aiven | postgres | UTF8 | en_US.UTF-8 | en_US.UTF-8 | =T/postgres +| No Access | pg_default | ... testdb2 | avnadmin | UTF8 | en_US.UTF-8 | en_US.UTF-8 | | 66 MB | pg_default | (6 rows) testdb2=> \l+ testdb2 List of databases Name | Owner | Encoding | Collate | Ctype | Access privileges | Size | Tablespace | Description ---------+----------+----------+-------------+-------------+-------------------+-------+------------+------------- testdb2 | avnadmin | UTF8 | en_US.UTF-8 | en_US.UTF-8 | | 66 MB | pg_default | (1 row)h ``` Using the pg\_database\_size function ``` testdb2=> select pg_database_size('testdb2'); pg_database_size ------------------ 68895523 (1 row) testdb2=> select pg_size_pretty(pg_database_size('testdb2')); pg_size_pretty ---------------- 66 MB (1 row) ``` Compare the outputs of the \l+ DB\_NAME command and the `pg_database_size` function The outputs for the testdb2 database size are the same for both methods. Since the `pg_database_size` function returns the database size in bytes, we use the `pg_size_pretty` function to retrieve an easy-to-read output. ## Get the size of a table[​](#get-the-size-of-a-table "Direct link to Get the size of a table") To get the table size, you can use either the `\dt+ [ pattern ]` command or the `pg_table_size` function. Using the \dt+ \[ pattern ] command ``` testdb2=> \dt+ mytable1 List of relations Schema | Name | Type | Owner | Persistence | Access method | Size | Description -------------+----------+-------+----------+-------------+---------------+-------+------------- test_schema | mytable1 | table | myowner | permanent | heap | 14 MB | (1 row) ``` Use the pg\_table\_size function ``` testdb2=> select pg_size_pretty(pg_table_size('mytable1')); pg_size_pretty ---------------- 14 MB (1 row) ``` ## Get the size of a table and its indices[​](#get-the-size-of-a-table-and-its-indices "Direct link to Get the size of a table and its indices") To get disk space usage for a table and its indexes, you can use the `pg_total_relation_size` function, which computes the total disk space used by the table, all its indices, and TOAST data: ``` testdb2=> select pg_size_pretty(pg_total_relation_size('mytable1')); pg_size_pretty ---------------- 15 MB (1 row) ``` warning It is not recommended to use the `pg_relation_size` function as it computes the disk space used by only one fork of the relation. To get the total size of all the relation's forks, use higher-level functions like `pg_total_relation_size` or `pg_table_size`. tip WAL files also contribute to the service disk usage. For more information, see [About PostgreSQL® disk usage](/docs/products/postgresql/concepts/pg-disk-usage.md) Related pages * [PostgreSQL interactive terminal](https://www.postgresql.org/docs/15/app-psql.html) * [Database Object Management Functions](https://www.postgresql.org/docs/current/functions-admin.html#FUNCTIONS-ADMIN-DBOBJECT) --- # Aiven for PostgreSQL® reads failover to the primary Enable automatic failover for your Aiven for PostgreSQL® read workloads to ensure uninterrupted access when standby nodes are unavailable. When you route read-only queries to standby nodes, a standby failure can make your replica URI temporarily unreachable until a new standby is provisioned and catches up. Reads failover to the primary automatically and temporarily redirects read-only traffic to the healthy primary node when all standbys are unavailable, helping you avoid downtime. ## Benefits[​](#benefits "Direct link to Benefits") * Improves availability for read workloads during standby outages. * Reduces operational effort; no app-side routing changes required. * Uses a single, stable connection endpoint for read traffic. ## How it works[​](#how-it-works "Direct link to How it works") * This feature doesn't create additional replicas; it redirects read traffic when replicas are unavailable. When enabled, your service exposes a dedicated HA replica DNS endpoint for read-only traffic. * Under normal conditions, this endpoint resolves to standby nodes. * If all standbys become unavailable, the endpoint automatically switches to the primary. * When standbys recover, the endpoint switches back to replicas. ## Enable the feature[​](#enable-the-feature "Direct link to Enable the feature") You can manage reads failover to the primary from the Aiven Console, CLI, API, or using Aiven Provider for Terraform. ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven for PostgreSQL service on a [Business or Premium plan](https://aiven.io/pricing?product=pg) (see how to change your plan) * Tool for managing the feature: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) * [Aiven Operator for Kubernetes®](/docs/tools/kubernetes.md) * During a failover to primary, read-only traffic is served by the primary. This means: * If your application expects eventual consistency from replicas, it will temporarily receive strong consistency from the primary. * If your application relies on read-after-write consistency, failover to the primary will maintain this guarantee, but switching back to replicas may reintroduce replication lag. * Ensure your application can tolerate these changes in consistency behavior during failover events. ### Use your preferred tool[​](#use-your-preferred-tool "Direct link to Use your preferred tool") * Console * CLI * API * Terraform * Kubernetes 1. In the [Aiven Console](https://console.aiven.io/), open your Aiven for PostgreSQL® service. 2. Go to service **Service settings** > **Advanced configuration**. 3. Click **Configure** > **Add configuration option**. 4. Use the search bar to find `enable_ha_replica_dns`, and set it to **Enabled**. 5. Click **Save configuration**. Run ``` aiven service update SERVICE_NAME -c enable_ha_replica_dns=true ``` Call the [ServiceUpdate endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) to set the `enable_ha_replica_dns` configuration to `true`: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "enable_ha_replica_dns": true } }' ``` Add or update your Terraform resource for the Aiven for PostgreSQL service: ``` resource "aiven_pg" "example" { project = "PROJECT_NAME" cloud_name = "CLOUD_REGION" plan = "PLAN_NAME" service_name = "SERVICE_NAME" user_config = { enable_ha_replica_dns = true # ...other config options } } ``` Add or update your Aiven Service custom resource manifest: ``` apiVersion: aiven.io/v1alpha1 kind: PostgreSQL metadata: name: SERVICE_NAME spec: project: PROJECT_NAME cloudName: CLOUD_REGION plan: PLAN_NAME userConfig: enable_ha_replica_dns: true # ...other config options ``` ## Use the HA-replica-DNS endpoint[​](#use-the-ha-replica-dns-endpoint "Direct link to Use the HA-replica-DNS endpoint") 1. After enabling, retrieve the replica connection URI from the console, CLI, or API. This URI will automatically redirect to the primary when replicas are unavailable and switch back once replicas are healthy. 2. Point your read-only clients to this URI to benefit from automatic failover without changing application logic. important Existing connections to replicas may fail during an outage. New connections using the HA replica DNS continue to succeed. To ensure application reliability, implement connection retry logic so your clients can reconnect automatically if a replica connection is interrupted. ## Disable the feature[​](#disable-the-feature "Direct link to Disable the feature") You can disable reads failover to the primary at any time. * Console * CLI * API * Terraform * Kubernetes 1. In the [Aiven Console](https://console.aiven.io/), open your Aiven for PostgreSQL® service. 2. Go to service **Service settings** > **Advanced configuration**. 3. Click **Configure** > **Add configuration option**. 4. Use the search bar to find `enable_ha_replica_dns`, and set it to **Disabled**. 5. Click **Save configuration**. Run ``` aiven service update SERVICE_NAME -c enable_ha_replica_dns=false ``` Call the [ServiceUpdate endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) to set the `enable_ha_replica_dns` configuration to `false`: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "enable_ha_replica_dns": false } }' ``` Add or update your Terraform resource for the Aiven for PostgreSQL service: ``` resource "aiven_pg" "example" { project = "PROJECT_NAME" cloud_name = "CLOUD_REGION" plan = "PLAN_NAME" service_name = "SERVICE_NAME" user_config = { enable_ha_replica_dns = false # ...other config options } } ``` Add or update your Aiven Service custom resource manifest: ``` apiVersion: aiven.io/v1alpha1 kind: PostgreSQL metadata: name: SERVICE_NAME spec: project: PROJECT_NAME cloudName: CLOUD_REGION plan: PLAN_NAME userConfig: enable_ha_replica_dns: false # ...other config options ``` --- # PG Studio for Aiven for PostgreSQL® [Early availability](/docs/platform/concepts/service-and-feature-releases.md) Aiven PG Studio lets you write, understand, and run SQL queries in the Aiven Console using natural language. It combines an SQL editor with an AI assistant that uses your database schema to generate and explain queries. note PG Studio and its AI features are on by default. To [turn off PG Studio or its AI features](/docs/products/postgresql/howto/pg-studio/security-connections.md#manage-pg-studio-and-ai-features), contact the Aiven support team. important PG Studio is an [Early availability](/docs/platform/concepts/service-and-feature-releases.md) feature. ## What PG Studio offers[​](#what-pg-studio-offers "Direct link to What PG Studio offers") PG Studio supports: * Writing SQL in plain English or any other language * Autocompleting SQL queries with the `Tab` key, based on PostgreSQL commands and your schema * Visualizing your database structure with an interactive schema map * Exploring tables in a **Tables** view with data preview * Exploring schemas and table relationships * Explaining queries and database objects * Running a single query or multiple selected queries at once, with live results in separate tabs * Executing write queries and data definition statements with built-in safety guardrails ## PG Studio components[​](#pg-studio-components "Direct link to PG Studio components") * **Schema visualization:** View your database structure as an interactive diagram showing tables, columns, and relationships. Open it from **Open schema map** or request it from the **AI Assistant** panel. Click the copy icon next to a table name to copy it to the clipboard. * **Tables view:** Browse tables in your selected schema and preview up to 100 rows. Open a table tab to start writing SQL. * **SQL editor:** Write and edit SQL across multiple tabs. Run a single statement or select multiple statements to execute them all at once, with each result shown in its own tab. Execute write operations with built-in safety guardrails. * **AI Assistant panel:** Describe what you need in natural language. The assistant generates SQL or explains queries, tables, and relationships using your database schema. ## Get started with PG Studio[​](#get-started-with-pg-studio "Direct link to Get started with PG Studio") ## [Get started](/docs/products/postgresql/howto/pg-studio/get-started.md) [Open PG Studio and run your first queries.](/docs/products/postgresql/howto/pg-studio/get-started.md) ## [Use AI Assistant](/docs/products/postgresql/howto/pg-studio/use-ai-assistant.md) [Generate and explain SQL queries with natural language.](/docs/products/postgresql/howto/pg-studio/use-ai-assistant.md) ## [Write and run queries](/docs/products/postgresql/howto/pg-studio/write-run-queries.md) [Use the SQL editor to write, edit, and execute queries.](/docs/products/postgresql/howto/pg-studio/write-run-queries.md) ## [Manage queries](/docs/products/postgresql/howto/pg-studio/manage-queries.md) [Save and organize your queries for reuse.](/docs/products/postgresql/howto/pg-studio/manage-queries.md) ## [Security and connections](/docs/products/postgresql/howto/pg-studio/security-connections.md) [Understand how PG Studio connects and protects your data.](/docs/products/postgresql/howto/pg-studio/security-connections.md) ## Related pages[​](#related-pages "Direct link to Related pages") * [AI Insights for Aiven for PostgreSQL](/docs/products/postgresql/howto/ai-insights.md) --- # Get started with PG Studio Open PG Studio and run your first queries. note PG Studio and its AI features are on by default, so no setup is needed. To turn them off, see [Manage PG Studio and AI features](/docs/products/postgresql/howto/pg-studio/security-connections.md#manage-pg-studio-and-ai-features). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") To use PG Studio, you need: * **Aiven permissions:** The `service:data:write` and `service:secrets:read` permissions at the organization, unit, or project level. These permissions are included in the **Admin**, **Developer**, and **Operator** roles. * **Network access:** Your IP address must be in the service's IP allowlist. PG Studio validates your browser's IP address, which must be allowed in the [service's IP filter configuration](/docs/platform/howto/restrict-access.md). If you get the `Access is not allowed from the IP address` error, add your IP address to the allowlist. ## Open PG Studio[​](#open-pg-studio "Direct link to Open PG Studio") 1. In the [Aiven Console](https://console.aiven.io/login), open your Aiven for PostgreSQL service. 2. Click **PG Studio**. 3. Click the source database and schema selectors. PG Studio opens a split view that shows the SQL editor and the **AI Assistant** panel. Use the editor selectors to change the database source and schema. If AI features are off for your organization, the **AI Assistant** panel does not appear. ## Run your first query[​](#run-your-first-query "Direct link to Run your first query") You can write SQL directly. If AI features are on, you can also use the **AI Assistant** to generate queries: ### Write SQL manually[​](#write-sql-manually "Direct link to Write SQL manually") 1. In the SQL editor, enter your query, for example: ``` SELECT * FROM users LIMIT 10; ``` 2. Click **Run**. 3. View the results in the results panel. ### Generate SQL with AI[​](#generate-sql-with-ai "Direct link to Generate SQL with AI") 1. In the **AI Assistant** panel, describe what you need, for example: **Show all users who signed up in the last 7 days**. 2. Review the generated SQL in the SQL editor. 3. Click **Run** to execute the query. ## Explore your schema[​](#explore-your-schema "Direct link to Explore your schema") 1. Click **Open schema map** to view your database structure as an interactive diagram. 2. Browse tables, columns, and relationships. 3. Ask schema questions in the **AI Assistant** panel, such as **How are the orders and customers tables related?**. ## Related pages[​](#related-pages "Direct link to Related pages") * [Use AI Assistant](/docs/products/postgresql/howto/pg-studio/use-ai-assistant.md) * [Write and run queries](/docs/products/postgresql/howto/pg-studio/write-run-queries.md) * [PG Studio overview](/docs/products/postgresql/howto/pg-studio.md) --- # Manage queries in PG Studio PG Studio lets you save useful SQL and revisit recently executed statements, so you can continue analysis without rewriting queries. ## Save queries[​](#save-queries "Direct link to Save queries") Save a query so you can open and run it again later. Saved queries are user-specific. Other users in the same Aiven for PostgreSQL service can't see your saved queries. To save a query: 1. In the SQL editor, write or generate your query. 2. Run the query and verify the result. 3. Click **Save**. PG Studio saves the query using the current tab name. The query appears in **Saved queries**. ## Access saved queries and query history[​](#access-saved-queries-and-query-history "Direct link to Access saved queries and query history") **Saved queries** shows both your explicitly saved queries and your recent query history, so you can return to any previous work. For write query history and fork-based rollback, see [View and revert write query history](/docs/products/postgresql/howto/pg-studio/write-run-queries.md#view-and-revert-write-query-history). To open a saved query or a recent query: 1. In the SQL editor, click **Saved queries**. 2. Select a query from the list to load it into the editor. 3. Click **Run** to execute the query. ## Rename saved queries[​](#rename-saved-queries "Direct link to Rename saved queries") To rename a saved query: 1. Click **Saved queries**. 2. Select a query from the list. 3. Double-click the query tab name, enter a new name, and press `Enter`. ## Delete saved queries[​](#delete-saved-queries "Direct link to Delete saved queries") To delete a saved query: 1. Click **Saved queries**. 2. Select a query from the list. 3. Click **Delete** to remove the query from your saved list. note Deleting a query removes it only from your saved queries list in the editor. It does not delete any database objects created earlier by running SQL in your database. ## Related pages[​](#related-pages "Direct link to Related pages") * [Write and run queries](/docs/products/postgresql/howto/pg-studio/write-run-queries.md) * [Get started with PG Studio](/docs/products/postgresql/howto/pg-studio/get-started.md) * [PG Studio overview](/docs/products/postgresql/howto/pg-studio.md) --- # Security and connections in PG Studio Learn how PG Studio connects to your database and ensures safe, controlled access. ## Database connection details[​](#database-connection-details "Direct link to Database connection details") PG Studio connects to your PostgreSQL service using: * **Database user:** The `avnadmin` user account, which has full read and write access to your databases. You can run SQL statements through PG Studio. * **Access scope:** Full database access with the same privileges as the `avnadmin` user. This is not limited to read-only access. ## Security safeguards[​](#security-safeguards "Direct link to Security safeguards") PG Studio ensures safe, controlled access: For AI query generation scope, see [How AI assistance works](/docs/products/postgresql/howto/pg-studio/use-ai-assistant.md#how-ai-assistance-works). * **Single-statement validation:** PG Studio allows only one SQL statement per execution. * **Automatic safety checks:** PG Studio validates all generated SQL for safety before execution. * **Restricted unsafe requests:** Requests for privilege escalation or malicious SQL are blocked. * **Timeouts and limits:** * Statement timeout: 30 seconds * Lock timeout: 10 seconds * Connection timeout: 10 seconds * Maximum result size: 5,000 rows * **Encrypted connections:** All database connections use SSL/TLS encryption. * **Rate limiting:** Two query executions every two seconds per user per service; one AI request every two seconds per user per service. * **Write query safeguards:** When you run a write query, PG Studio prompts you to confirm before executing. * **Fork testing option:** You can test write queries on a database fork instead of modifying live data directly. ## Network access requirements[​](#network-access-requirements "Direct link to Network access requirements") Your IP address must be in the service's IP allowlist. PG Studio validates your browser's IP address, which must be allowed in the [service's IP filter configuration](/docs/platform/howto/restrict-access.md). If you get the `Access is not allowed from the IP address` error, add your IP address to the allowlist. ## Required permissions[​](#required-permissions "Direct link to Required permissions") To use PG Studio, you need the `service:data:write` and `service:secrets:read` permissions at the organization, unit, or project level. These permissions are included in the **Admin**, **Developer**, and **Operator** roles. ## Manage PG Studio and AI features[​](#manage-pg-studio-and-ai-features "Direct link to Manage PG Studio and AI features") PG Studio and its AI features are on by default for all organizations and apply to all projects in your organization. Aiven manages these controls, so you cannot change them yourself. The two controls are independent, so you can turn either one off or on without affecting the other. To change either setting, contact the [Aiven support team](mailto:support@aiven.io). ## Related pages[​](#related-pages "Direct link to Related pages") * [Get started with PG Studio](/docs/products/postgresql/howto/pg-studio/get-started.md) * [Restrict access to services](/docs/platform/howto/restrict-access.md) * [PG Studio overview](/docs/products/postgresql/howto/pg-studio.md) --- # Use AI Assistant in PG Studio The AI Assistant in PG Studio helps you generate SQL queries and understand your database using natural language. note AI features are on by default. If they are off for your organization, the **AI Assistant** panel and the **Ask AI** action do not appear. Aiven manages this control. To change it, contact the Aiven support team. See [Manage PG Studio and AI features](/docs/products/postgresql/howto/pg-studio/security-connections.md#manage-pg-studio-and-ai-features). ## How AI assistance works[​](#how-ai-assistance-works "Direct link to How AI assistance works") The AI Assistant lets you get help with SQL queries and database questions in natural language. You can: * **Generate queries:** Describe what data you need, and the assistant creates the SQL query for you. * **Get explanations:** Ask about SQL syntax, query structure, or how specific database features work. * **Understand your schema:** Request information about table structures, relationships, and data organization. * **Troubleshoot queries:** Get help debugging errors, breaking down complex queries, or improving query performance. The AI Assistant uses your database schema as context for all responses, providing accurate and relevant suggestions specific to your data structure. It can suggest `SELECT`, data modification, and data definition statements when they match your request. For blocked request types, see [Security safeguards](/docs/products/postgresql/howto/pg-studio/security-connections.md#security-safeguards). ## Generate SQL with natural language[​](#generate-sql-with-natural-language "Direct link to Generate SQL with natural language") 1. In the **AI Assistant** panel, describe the query or result you need. 2. Review the generated SQL in the SQL editor. 3. Click **Run** to execute the query. Example queries to try: * Show monthly active users for the last 6 months * Find all orders with a total amount greater than 1000 dollars that are still pending * List customers who placed orders in the last 30 days * Find duplicate email addresses in the users table * List failed payment transactions from the last week with customer details ## Ask AI about your query[​](#ask-ai-about-your-query "Direct link to Ask AI about your query") Use **Ask AI** in the SQL editor to explain queries or create a modified version of your SQL. The AI response and any updated SQL appear in the **AI Assistant** panel. 1. Paste a query into the SQL editor. 2. Optional: Highlight the part of the query to focus on. 3. Click **Ask AI** in the SQL editor. 4. In the **AI Assistant** panel, enter your request. 5. Review the explanation or updated SQL in the **AI Assistant** panel. 6. Optional: Click **Run query** in the **AI Assistant** panel to move the updated SQL to the editor and execute it. When you highlight a snippet, the AI uses it as a focused context. You can ask for an explanation of that part or request changes to it, such as rewriting a filter or optimizing a subquery. ### What you can ask AI[​](#what-you-can-ask-ai "Direct link to What you can ask AI") * Explain what a specific SQL query does in plain language * Break down complex queries into simpler steps * Describe how `JOIN` operations combine data from multiple tables * Clarify the purpose of `WHERE` clauses and filters * Explain aggregate functions and `GROUP BY` operations * Rewrite or modify a highlighted part of a query * Optimize a selected subquery or filter condition ## How AI context works[​](#how-ai-context-works "Direct link to How AI context works") When you work in the SQL editor, click **Ask AI** for the current statement. PG Studio opens the **AI Assistant** panel, attaches SQL context, and focuses the input. PG Studio sends SQL context in two parts: * The full active SQL statement for context, including aliases, common table expressions (CTEs), and joins. * Any highlighted SQL snippet as a separate focus area. ## Explore your schema with AI[​](#explore-your-schema-with-ai "Direct link to Explore your schema with AI") 1. Ask a schema question in the **AI Assistant** panel, such as how tables relate or what a column stores. 2. Review the response or generated SQL. 3. Click **Open schema map** to browse tables and relationships visually. ## Related pages[​](#related-pages "Direct link to Related pages") * [Write and run queries](/docs/products/postgresql/howto/pg-studio/write-run-queries.md) * [Get started with PG Studio](/docs/products/postgresql/howto/pg-studio/get-started.md) * [PG Studio overview](/docs/products/postgresql/howto/pg-studio.md) --- # Write and run queries in PG Studio Use the SQL editor in PG Studio to write, edit, and execute queries. note PG Studio is on by default. If it is off for your organization, PG Studio shows a message to contact Aiven support and you cannot open the editor or run queries. Aiven manages this control. To change it, contact the Aiven support team. See [Manage PG Studio and AI features](/docs/products/postgresql/howto/pg-studio/security-connections.md#manage-pg-studio-and-ai-features). ## Write SQL manually[​](#write-sql-manually "Direct link to Write SQL manually") 1. In the SQL editor, enter your query. Use `Tab` to autocomplete Press the `Tab` key to autocomplete SQL keywords, table names, column names, and other database objects from your schema. When the autocomplete list is open, use the arrow keys to move between suggestions and press `Enter` or `Tab` to insert the selected suggestion. Press `Esc` to dismiss autocomplete. 2. Click **Run**, or press `Cmd+Enter` on macOS or `Ctrl+Enter` on Windows and Linux. 3. View the results in the results panel. The autocomplete feature helps you write queries faster by suggesting: * SQL keywords and clauses: `SELECT`, `FROM`, `WHERE`, `JOIN` * Table names from your selected schema * Column names from tables in your query * Database functions and operators ## Run a single query[​](#run-a-single-query "Direct link to Run a single query") note When the editor contains multiple SQL statements and you don't select any text, **Run** executes only the statement where your cursor is placed. 1. In the SQL editor, place your cursor inside the query you want to run, without selecting any text. 2. Click **Run**, or press `Cmd+Enter` on macOS or `Ctrl+Enter` on Windows and Linux. 3. View the results in the results panel. This is useful for: * Testing individual queries from a larger script * Iterating on a single query while keeping other statements in context * Testing each subquery independently before combining a complex query * Verifying data in a specific table without executing setup or cleanup statements ## Run multiple queries[​](#run-multiple-queries "Direct link to Run multiple queries") You can write several SQL statements in the editor and run them all at once. Select the statements you want to run. PG Studio runs them sequentially and displays the results in separate tabs. 1. In the SQL editor, write your SQL statements: ``` SELECT count(*) FROM orders; SELECT * FROM users LIMIT 10; ``` 2. Select all the statements you want to run. 3. Click **Run selected**, or press `Cmd+Enter` on macOS or `Ctrl+Enter` on Windows and Linux. 4. View the results in the results panel. When more than one statement runs, each result appears in its own tab labeled **Query 1**, **Query 2**, and so on: * A red error indicator on a tab means the statement encountered an error. If a statement fails, execution stops and remaining statements are not run. Results from statements that completed before the error are preserved. This is useful for: * Running a sequence of setup, query, and cleanup statements in one go * Executing a batch of related queries and comparing their results * Validating multiple statements before saving or sharing a script ## Execute queries[​](#execute-queries "Direct link to Execute queries") When you run a query, whether generated by the AI or written manually: * PG Studio automatically detects the query type and applies appropriate safeguards. * The editor executes it unless it is a write operation, in which case PG Studio prompts you to confirm before execution as a safety measure. * The results panel displays your query results. When running multiple statements, each result appears in its own tab. ## Run write queries[​](#run-write-queries "Direct link to Run write queries") You can execute data modification and data definition statements directly against your database. To run a write query: 1. In the SQL editor, enter your write query. 2. Click **Run**, or press `Cmd+Enter` on macOS or `Ctrl+Enter` on Windows and Linux. 3. PG Studio detects the write operation and shows **You are changing live data**. 4. Choose how to proceed: * **Run on production** runs the query directly against the connected database. * **Test on a fork** creates a fork so you can test the query without changing live data. 5. Optional: Select **Don't ask me again** to skip this confirmation for future write queries on this service. ## View and revert write query history[​](#view-and-revert-write-query-history "Direct link to View and revert write query history") PG Studio keeps a change log of write queries you run. You can review changes and create a fork from a point in time before a specific write query. To view the change log: 1. In the SQL editor, click **Change log**. This option appears after you run at least one write query. 2. In **Change log**, review each write query with its timestamp and query preview. To create a fork from a point before a specific write query: 1. Open **Change log**. 2. Select the write query you want to roll back to. 3. Click **Create fork**. PG Studio opens **Create fork** with a recovery timestamp set to the moment before that query ran. ## Query execution limits[​](#query-execution-limits "Direct link to Query execution limits") PG Studio applies these limits to query execution: * Statement timeout: 30 seconds * Lock timeout: 10 seconds * Connection timeout: 10 seconds * Maximum result size: 1,000 rows per statement * Maximum statements per run: 10 * **Rate limiting:** 10 query executions per 10 seconds per user per service * **AI rate limiting:** One AI request every two seconds per user per service ## Explore tables in the query editor[​](#explore-tables-in-the-query-editor "Direct link to Explore tables in the query editor") Use **Tables** in the query editor to browse your database tables and start to run SQL queries. The table list uses your selected source database and schema. To use the **Tables** view: 1. Open PG Studio. 2. Select **Tables** in the query editor. 3. Select a source database and schema. 4. Select a table from the list. When you open a table, PG Studio shows up to 100 rows. From **Tables**, you can open a table tab and use it as a starting point for querying that table. For column definitions and relationships, use **Open schema map**. ## Related pages[​](#related-pages "Direct link to Related pages") * [Use AI Assistant](/docs/products/postgresql/howto/pg-studio/use-ai-assistant.md) * [Manage queries](/docs/products/postgresql/howto/pg-studio/manage-queries.md) * [Security and connections](/docs/products/postgresql/howto/pg-studio/security-connections.md) * [PG Studio overview](/docs/products/postgresql/howto/pg-studio.md) --- # Access PgBouncer statistics for Aiven for PostgreSQL® PgBouncer is used at Aiven as a [connection pooler](/docs/products/postgresql/concepts/pg-connection-pooling.md) to lower the performance impact of opening new connections to Aiven for PostgreSQL®. Verify your password encryption method If you use PGBouncer connection pooling, [verify your password encryption method compatibility](/docs/products/postgresql/troubleshooting/pg-password-encryption-upgrade.md) to ensure successful connections. You may need to migrate to `SCRAM-SHA-256` to maintain compatibility as the MD5 password encryption will be deprecated in PostgreSQL 19. You can access PgBouncer statistics in two ways: * **Through integration endpoints**: This is the recommended option. PgBouncer statistics are exported as standard metrics that you can send to Datadog, Prometheus, or CloudWatch by using [integration endpoints](/docs/platform/howto/list-monitoring.md) and related service integrations. * **Direct database connection**: Connect to PgBouncer and run the `SHOW STATS` command to view statistics directly. ## Access statistics through integration endpoints[​](#access-statistics-through-integration-endpoints "Direct link to Access statistics through integration endpoints") Aiven exports PgBouncer `SHOW STATS` results as standard metrics. You can send them to external monitoring systems through integration endpoints and the related service integrations, without connecting to the database. The following statistics are available: * `total_xact_count` - Total number of SQL transactions pooled * `total_query_count` - Total number of SQL queries pooled * `total_received` - Total volume in bytes of network traffic received * `total_sent` - Total volume in bytes of network traffic sent * `total_xact_time` - Total transaction time in microseconds * `total_query_time` - Total query time in microseconds * `total_wait_time` - Total time in microseconds clients waited for a connection * `total_bind_count` - Total number of bind operations * `total_client_parse_count` - Total number of client-side parse operations * `total_server_assignment_count` - Total number of server assignment operations * `total_server_parse_count` - Total number of server-side parse operations * `avg_xact_count` - Average transactions per second * `avg_query_count` - Average queries per second * `avg_xact_time` - Average transaction time in microseconds * `avg_query_time` - Average query time in microseconds * `avg_wait_time` - Average wait time in microseconds * `avg_bind_count` - Average bind operations per second * `avg_client_parse_count` - Average client parse operations per second * `avg_server_assignment_count` - Average server assignments per second * `avg_server_parse_count` - Average server parse operations per second ### Metric format[​](#metric-format "Direct link to Metric format") PgBouncer metrics use the native format for each metrics integration. When you send PgBouncer metrics to Aiven for Metrics with an InfluxDB-compatible endpoint, Aiven exports them in InfluxDB line protocol format. In this case, metrics include the following tags: `cloud`, `db`, `host`, `instance`, `project`, `service`, and `service_type`. The `instance` tag distinguishes between metrics from different PgBouncer processes. For example, if a service runs two PgBouncer processes, their metrics have `instance` set to `pgbouncer_1` and `pgbouncer_2`. The following example shows the InfluxDB line protocol output: ``` pgbouncer,cloud=google-europe-west1,db=pool1,host=ae-pg-1,instance=pgbouncer_1,project=testproject,service=ae-pg,service_type=pg avg_query_time=383i,total_query_count=33i,total_wait_time=43703i,avg_wait_time=0i 1773830989000000000 ``` ### Enable metrics integrations[​](#enable-metrics-integrations "Direct link to Enable metrics integrations") To access PgBouncer metrics through integrations: * **Datadog**: Follow [Monitor PgBouncer with Datadog](/docs/products/postgresql/howto/monitor-pgbouncer-with-datadog.md) to enable PgBouncer monitoring. * **Other integrations**: Set up any [metrics integration](/docs/platform/howto/list-monitoring.md), such as Prometheus or CloudWatch, for your PostgreSQL service. PgBouncer metrics are automatically included. ## Access statistics through direct connection[​](#access-statistics-through-direct-connection "Direct link to Access statistics through direct connection") You can also connect directly to PgBouncer and run the `SHOW STATS` command to view statistics. note You have read-only access to PgBouncer statistics since PgBouncer pools are automatically managed by Aiven. ### Get PgBouncer URI[​](#extract-pgbouncer-uri "Direct link to Get PgBouncer URI") To get the PgBouncer URI, you can use either the [Aiven Console](https://console.aiven.io/) or the [Aiven CLI client](/docs/tools/cli.md). #### PgBouncer URI in the console[​](#pgbouncer-uri-in-the-console "Direct link to PgBouncer URI in the console") 1. Log in to the [Aiven Console](https://console.aiven.io/), and go to a desired organization, project, and service. 2. Click **Connection pools**, and find a desired pool. 3. Click **Actions** > **Info** > **Primary Connection URI**. #### PgBouncer URI with the Aiven CLI[​](#pgbouncer-uri-with-the-aiven-cli "Direct link to PgBouncer URI with the Aiven CLI") Use [jq](https://stedolan.github.io/jq/) to parse the JSON response. Execute the following command replacing `SERVICE_NAME` and `PROJECT_NAME` as needed: ``` avn service get SERVICE_NAME --project PROJECT_NAME --json | jq -r '.connection_info.pgbouncer' ``` Expect to receive an output similar to the following: ``` postgres://avnadmin:xxxxxxxxxxx@demo-pg-dev-advocates.aivencloud.com:13040/pgbouncer?sslmode=require ``` ### Connect to PgBouncer[​](#connect-to-pgbouncer "Direct link to Connect to PgBouncer") To connect to PgBouncer, use the [extracted URI](#extract-pgbouncer-uri): ``` psql 'EXTRACTED_PGBOUNCER_URI' ``` ### View statistics[​](#view-statistics "Direct link to View statistics") 1. Enable the expanded display by running: ``` pgbouncer=# \x ``` 2. Show the statistics by running: ``` pgbouncer=# SHOW STATS; ``` Depending on the load of your database, expect an output similar to the following: ``` database | total_xact_count | total_query_count | total_received | total_sent | total_xact_time | total_query_time | total_wait_time | avg_xact_count | avg_query_count | avg_xact_time | avg_query_time | avg_wait_time ----------+------------------+-------------------+----------------+------------+-----------------+------------------+-----------------+----------------+-----------------+---------------+----------------+--------------- pgbouncer | 1 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 (1 row) ``` tip Run `SHOW HELP` to see all `pgbouncer` commands. --- # Power on/off and delete your Aiven for PostgreSQL® service Power off your Aiven for PostgreSQL® service to release resources and save credits, power it back on when you need it, or delete it permanently. ## Power off a service[​](#power-off-a-service "Direct link to Power off a service") When you power off a service: * All virtual machines are removed from the public cloud. * The service configuration is stored on the Aiven Platform. * If there are no backups, all service data is lost. * If the service has time-based or point in time recovery backups, the backups remain on the Aiven Platform. Services powered off for more than 180 days are automatically deleted. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power off service**. To power off a service, run: ``` avn service update SERVICE_NAME --power-off ``` ## Power on a service[​](#power-on-a-service "Direct link to Power on a service") When you power on a service: * New virtual machines are created on the service's public cloud. * The service starts with the stored configuration parameters. * The latest time-based backup is restored. * Maintenance updates are automatically applied. * If a point in time recovery backup is available, the database transaction logs are replayed to recover the service data to a specific point in time. The restoration takes from a few minutes to a few hours, depending on the network bandwidth, the disk IOPS allocated to the service, and the size of the backup. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power on service**. To power on a service, run: ``` avn service update SERVICE_NAME --power-on ``` To see when the service is running, run: ``` avn service wait SERVICE_NAME ``` note When you power on an Aiven for PostgreSQL service, the latest backup is restored and the write-ahead log (WAL) is replayed to recover the data to the latest available point in time. note Static IP addresses are not removed when a service is powered off or deleted. They continue to generate the usual costs. To avoid these costs, [remove the static IP addresses](/docs/platform/concepts/static-ips.md). ## Delete a service[​](#delete-a-service "Direct link to Delete a service") * Console * CLI 1. In your project, click **Services**. 2. Open the service to delete, and click **Actions** > **Delete service**. To delete a service, run: ``` avn service terminate SERVICE_NAME ``` Related pages * [Aiven for PostgreSQL® backups](/docs/products/postgresql/concepts/pg-backups.md) * [Fork Aiven for PostgreSQL®](/docs/products/postgresql/howto/fork-service.md) --- # Prepare your Aiven for PostgreSQL® service for high load Prepare your Aiven for PostgreSQL® service for higher than usual traffic to avoid outages and keep performance stable. Prepare your services for higher than usual traffic to avoid service outages by doing the following: * **Subscribe to service notifications:** To receive notifications about service health and warnings when resources are low, you can [set service and project contacts](https://aiven.io/docs/platform/howto/technical-emails). You can also view the status of the Aiven Platform and get updates on incidents on the [status page](https://status.aiven.io/). Follow the RSS feed, subscribe to email or SMS notifications, or use the Slack integration to get notifications about incidents. * **Monitor your services:** [Monitor the health of your services](/docs/platform/howto/list-monitoring.md) using metrics, logs, alerts, and dashboards. * **Scale your services:** If you forecast a load that can't be handled by the service, you can scale up your service. * **Set the backup schedule:** To minimize the impact of the higher load during the backup process, schedule backups outside of peak traffic hours. * **Set the maintenance window:** Schedule maintenance updates outside of your peak traffic hours. * **Run load tests on service forks:** To test the impact of high traffic on a production service, fork the service and run your load test on the fork. Additionally, optimizing a service allows it to perform better under stress therefore avoiding the need of an upgrade. The more optimized a service is for your usage, the better you can weather spikes in traffic. Related pages * [Change the service plan](/docs/products/postgresql/howto/change-service-plan.md) * [Scale disk storage](/docs/products/postgresql/howto/scale-disk-storage.md) --- # Prevent PostgreSQL® full disk issues If your Aiven for PostgreSQL® service runs out of disk space, the service will start malfunctioning, which will preclude proper backup creation. To prevent this situation, Aiven automatically detects when your service is running out of free space and stops further write operations by setting the `default_transaction_read_only` parameter to `ON`. With this setting in place, clients trying to execute write operations will start facing errors like: ``` cannot execute CREATE TABLE in a read-only transaction. ``` To re-enable database writes, increase the available space, by either deleting data or upgrading to a larger plan. ## Increase free space by upgrading to a larger plan[​](#increase-free-space-by-upgrading-to-a-larger-plan "Direct link to Increase free space by upgrading to a larger plan") When upgrading to a larger plan, new nodes with bigger disk space capacity are created and replace the original nodes. Once the new nodes with increased disk capacity are up and running, the disk usage returns to below the critical level and the system automatically sets the `default_transaction_read_only` parameter to `OFF` allowing write operations again. You can upgrade the Aiven for PostgreSQL service plan via the [Aiven console](https://console.aiven.io/) or [Aiven CLI](/docs/tools/cli.md). To perform a plan upgrade via the [Aiven console](https://console.aiven.io/): 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Change**. 3. Select a plan that provides at least 125% of the current disk size and click **Change plan**. Once the new nodes with the increased disk capacity are up and running, the disk usage drops below the critical level and the system automatically sets the `default_transaction_read_only` parameter to `OFF` allowing write operations again. note You might want to temporarily upgrade your service to a bigger plan to enable all users or applications to delete data without strict time limits. In such cases, you may have to wait for the next service backup to complete before you can downgrade to a smaller plan. ## Increase free space by deleting data[​](#increase-free-space-by-deleting-data "Direct link to Increase free space by deleting data") To release space from a database, you can also delete data stored in it, but the database read-only mode also prevents this. You can enable deletions by either enabling writes for a specific session or for a limited amount of time over the full database. ### Enable database writes for a specific session[​](#enable-database-writes-for-a-specific-session "Direct link to Enable database writes for a specific session") To enable writes for a session, login to the required database and execute the following command: ``` SET default_transaction_read_only = OFF; ``` You can then delete data within your session. ### Enable database writes for a limited amount of time[​](#enable-database-writes-for-a-limited-amount-of-time "Direct link to Enable database writes for a limited amount of time") To enable any writes to the database for a limited amount of time, send the following `POST` request using [Aiven APIs](/docs/tools/api.md) and replacing the `PROJECT_NAME` and `SERVICE_NAME` placeholders: ``` https://api.aiven.io/v1/project//service//enable-writes ``` The above API call enables write operations in the target Aiven for PostgreSQL database for 15 minutes, allowing you to delete some data. --- # Restrict access to databases or tables in Aiven for PostgreSQL® You can restrict access to Aiven for PostgreSQL® databases and tables by setting up read-only permissions for specific user's roles. tip Use an AI assistant connected to [Aiven MCP](/docs/tools/mcp-server.md) to review PostgreSQL roles and configure read-only access. For example: > List the login-enabled roles and their grants on `my-pg-service`, and create a read-only role for the `public` schema. ## Set read-only access in a schema[​](#set-read-only-access-in-a-schema "Direct link to Set read-only access in a schema") 1. Modify default permissions for a user's role in a particular schema. ``` alter default privileges for role name_of_role in schema name_of_schema YOUR_GRANT_OR_REVOKE_PERMISSIONS ``` 2. Apply the new read-only access setting to your existing database objects that uses the affected schema. ``` grant select on all tables in schema name_of_schema to NAME_OF_READ_ONLY_ROLE ``` ## Set read-only access in a database[​](#set-read-only-access-in-a-database "Direct link to Set read-only access in a database") You can set up the read-only access for a specific user's role in a particular database. 1. Create a database which will be used as a template `create database ro__template...`. 2. For the new template database, set permissions and roles that you want as default ones in the template. 3. When creating a database, use `create database NAME with template = 'ro__template'`. --- # Rename your Aiven for PostgreSQL® service Change the name of your Aiven for PostgreSQL® service by forking it under a new name and deleting the original service. You cannot rename a service after creation. Instead, you can create a fork with the new name and delete the original service. ## Rename a service[​](#rename-a-service "Direct link to Rename a service") 1. Stop writing to the service. 2. Fork the service. 3. Add any integrations or SSO configurations that weren't copied. 4. Connect your clients to the new service. 5. Test the forked service. 6. Delete the original service. Related pages * [Fork Aiven for PostgreSQL®](/docs/products/postgresql/howto/fork-service.md) * [Power on/off and delete your Aiven for PostgreSQL® service](/docs/products/postgresql/howto/power-cycle-service.md) --- # Identify and repair issues with PostgreSQL® indexes with REINDEX PostgreSQL® indexes can become corrupted due to a variety of reasons including software bugs, hardware failures or unexpected duplicated data. `REINDEX` allows you to rebuild the index in such situations. ## Rebuild non-unique indexes[​](#rebuild-non-unique-indexes "Direct link to Rebuild non-unique indexes") You can rebuild corrupted indexes that do not have `UNIQUE` in their definition using the following command, that creates a new index replacing the old one: ``` REINDEX INDEX ; ``` warning Re-indexing applies locks to the table and may interfere with normal use of the database. In some cases, it can be useful to manually build a second index concurrently alongside the old index and remove the old index: ``` CREATE INDEX CONCURRENTLY foo_index_new ON table_a (...); DROP INDEX CONCURRENTLY foo_index_old; ALTER INDEX foo_index_new RENAME TO foo_index; ``` You can run the `REINDEX` command for: * all indexes of a table (`REINDEX TABLE`) * all indexes in the entire database (`REINDEX DATABASE`). For more information on the `REINDEX` command, see the [PostgreSQL documentation page](https://www.postgresql.org/docs/current/sql-reindex.html). ## Rebuild unique indexes[​](#rebuild-unique-indexes "Direct link to Rebuild unique indexes") A `UNIQUE` index works on top of one or more columns whose combination is unique in a table. In situations when the index is corrupted or disabled and duplicated physical rows appear in the table, breaking the uniqueness constraint of the index, then index rebuilding with `REINDEX` will fail. To solve such a problem, you'll first need to remove the duplicated rows from the table before attempting to rebuild the index. ### Identify conflicting duplicated rows[​](#identify-conflicting-duplicated-rows "Direct link to Identify conflicting duplicated rows") To identify conflicting duplicate rows, run a query that counts the number of rows for each combination of columns included in the index definition. For example, the following `route` table has a `unique_route_index` index defining unique rows based on the combination of the `source` and `destination` columns: ``` CREATE TABLE route( source TEXT, destination TEXT, description TEXT ); CREATE UNIQUE INDEX unique_route_index ON route (source, destination); ``` If the `unique_route_index` is corrupted, find duplicated rows in the `route` table by running: ``` SELECT source, destination, count FROM (SELECT source, destination, COUNT(*) AS count FROM route GROUP BY source, destination) AS foo WHERE count > 1; ``` This query groups the data by the same `source` and `destination` fields defined in the index, and filters any entries with more than one occurrence. The resulting rows identify the problematic entries, which must be resolved manually by deleting or merging the entries until no duplicates exist. Once duplicated entries are removed, you can use the `REINDEX` command to rebuild the index. --- # Monitor PostgreSQL® metrics with Grafana® As well as offering PostgreSQL-as-a-service, the Aiven platform gives you access to monitor the database. The metrics/dashboard integration in the Aiven console lets you send PostgreSQL® metrics to an external endpoint like Datadog or to create an integration and a [prebuilt dashboard](/docs/products/postgresql/reference/pg-metrics.md) in Aiven for Grafana®. Get detailed information about the metrics and dashboard sections in [PostgreSQL® metrics exposed in Grafana®](/docs/products/postgresql/reference/pg-metrics.md). ## Push PostgreSQL metrics to Aiven for Metrics or PostgreSQL[​](#push-postgresql-metrics-to-aiven-for-metrics-or-postgresql "Direct link to Push PostgreSQL metrics to Aiven for Metrics or PostgreSQL") To collect metrics about your PostgreSQL service you will need to configure a metrics integration and nominate somewhere to store the collected metrics. 1. In the **Overview** page of your Aiven for PostgreSQL service, go to **Manage integrations** and choose the **Store Metrics** option with **Store service metrics in a time-series database** as its description. 2. Choose either a new or existing Aiven for Metrics or PostgreSQL service. * A new service will ask you to select the cloud, region and plan to use. You should also give your service a name. The service overview page shows the nodes rebuilding, and indicates when they are ready. * If you're already using Aiven for Metrics or PostgreSQL, you can submit your PostgreSQL metrics to the existing service. warning Although you can send metrics of your PostgreSQL service to this very service, this is not recommended since it increases the load on the monitored system. This can also result in issues with the availability of the metrics in case of problems with the service. ## Provision and configure Grafana[​](#provision-and-configure-grafana "Direct link to Provision and configure Grafana") 1. Select the target Aiven for Metrics or PostgreSQL database service and go to its service page. Under **Manage integrations**, choose the **Grafana Metrics Dashboard** option to make the metrics available on that platform. 2. Choose either a new or existing Grafana service. * A new service will ask you to select the cloud, region and plan to use. You should also give your service a name. The service overview page shows the nodes rebuilding, and indicates when they are ready. * If you're already using Grafana on Aiven, you can integrate your Aiven for Metrics as an additional data source for that existing Grafana. 3. On the **Overview** page for your Aiven for Grafana service, select the **Service URI** link. The username and password for your Grafana service is also available on the service's **Overview** page. Now your Grafana service is connected to Aiven for Metrics as a data source and you can go ahead and visualize your PostgreSQL metrics. ## Open PostgreSQL metrics prebuilt dashboard[​](#open-postgresql-metrics-prebuilt-dashboard "Direct link to Open PostgreSQL metrics prebuilt dashboard") In Grafana, go to **Dashboards** and **Manage**, and double click the dashboard named after the metrics database. ![Screenshot of a Grafana Manage Dashboards panel](/docs/assets/images/metrics-dashboard-manage-5f346afa2b7865c4fd49c748a4650855.png) Browse the prebuilt dashboard or create your own monitoring views. More info about the dashboard and pushed metrics can be found at [PostgreSQL® metrics exposed in Grafana®](/docs/products/postgresql/reference/pg-metrics.md) ![Screenshot of the PostgreSQL Metrics Dashboard for Grafana](/docs/assets/images/metrics-dashboard-global-8de42b05a6e0a38714f5a299bd871816.png) --- # Restore PostgreSQL® from a backup Aiven for PostgreSQL® databases are automatically backed up and can be restored from a backup at any point in time within the **backup retention period**, which [varies by plan](/docs/products/postgresql/concepts/pg-backups.md). The restore is created by forking: a new PostgreSQL instance is created and content from the original database is restored into it. note Aiven for PostgreSQL doesn't allow a service to be rolled back to a backup in-place since it creates alternative timelines for the database, adding complexity for the user. To restore a PostgreSQL database: 1. In your service, in the **Backups** section, click **Backup management**. 2. Click **Fork & restore**. 3. Choose the backup to fork from. 4. Enter a name, and select the cloud and plan. 5. Click **Create fork**. Once the new service is running, you can change your application's connection settings to point to it. tip Forked services can also be very useful for testing purposes, allowing you to create a completely realistic, separate copy of the actual production database with its data. ## Manual restores[​](#manual-restores "Direct link to Manual restores") Manual restoration should only be necessary when data is accidentally corrupted by the pointing applications. Aiven automatically handles outages and software failures by replacing broken nodes with new ones that resume correctly from the point of failure. note The Hobbyist service plan does not support database forking, so you have to use an external tool, such as `pg_dump`, to perform a backup. To perform a manual backup, see [Create manual PostgreSQL® backups](/docs/products/postgresql/howto/create-manual-backups.md). --- # Migrate between PostgreSQL® instances using aiven-db-migrate in Python The `aiven-db-migrate` tool is an open source project available on [GitHub](https://github.com/aiven/aiven-db-migrate), useful to perform PostgreSQL migrations. It's written in Python and therefore you can start any migration by directly calling the correct module. The list of source and target database requirements is available in the [dedicated documentation](/docs/products/postgresql/howto/migrate-aiven-db-migrate.md). ## Variables[​](#variables "Direct link to Variables") The following variables need to be substituted in the `aiven-db-migrate` calls | Variable | Description | | -------------- | ------------------------------------------------------------- | | `SRC_USERNAME` | Username for source PostgreSQL connection | | `SRC_PASSWORD` | Password for source PostgreSQL connection | | `SRC_HOSTNAME` | Hostname for source PostgreSQL connection | | `SRC_PORT` | Port for source PostgreSQL connection | | `DST_USERNAME` | Username for destination PostgreSQL connection | | `DST_PASSWORD` | Password for destination PostgreSQL connection | | `DST_HOSTNAME` | Hostname for destination PostgreSQL connection | | `DST_PORT` | Port for destination PostgreSQL connection | | `DST_DBNAME` | Bootstrap database name for destination PostgreSQL connection | ## Execute the `aiven-db-migrate` in Python[​](#execute-the-aiven-db-migrate-in-python "Direct link to execute-the-aiven-db-migrate-in-python") You can run perform a PostgreSQL® migration using `aiven-db-migrate` by: * Cloning the `aiven-db-migrate` [GitHub repository](https://github.com/aiven/aiven-db-migrate) and follow the installation instructions * Executing the following script ``` python3 -m aiven_db_migrate.migrate pg -d \ -s postgres://SRC_USERNAME:SRC_PASSWORD@SRC_HOSTNAME:SRC_PORT \ -t postgres://DST_USERNAME:DST_PASSWORD@DST_HOSTNAME:DST_PORT/DST_DBNAME?sslmode=require \ --max-replication-lag 1 \ --stop-replication \ -f DB_TO_SKIP ``` The following flags can be set: * `-d` or `--debug`: enables debug logging, highly recommended when running migration to see details of the migration process * `-f` or `--filtered-db`: comma separated list of databases to filter out during migrations * `--stop-replication`: by default logical replication will be left running on both source and target database. If set to true, it requires also `--max-replication-lag >= 0` to wait until a fixed replication lag before stopping the replication * `--max-replication-lag`: max replication lag in bytes to wait for, by default no wait, this parameter is required when using `--stop-replication` * `--skip-table`: comma separated list of tables to filter out during migrations, useful when willing to skip extension tables like `spatial_ref_sys` from `postgis` extension. tip For live migration `--stop-replication` and `--max-replication-lag` flags are not needed to keep logical replication running. The full list of parameters and related descriptions is available by running the following command: ``` python3 -m aiven_db_migrate.migrate pg -h ``` --- # Scale disk storage for your Aiven for PostgreSQL® service Scale the disk storage of your Aiven for PostgreSQL® service up or down without disrupting the running service. /eol-for-major-versions#aiven-for-flinkAdding or removing disk storage does not disrupt the running service. You pay only for extra storage instead of upgrading compute resources. You can add extra storage when you create a service or after it is running. When you add storage to a running service, the Aiven Platform provisions the extra disk and adds it to the running instances. For a clustered service such as Aiven for Apache Kafka®, Aiven divides extra storage equally between the nodes. For a shared service, each node receives the full extra capacity. ## Limitations[​](#limitations "Direct link to Limitations") * Disk added for extra storage is slower than the original disk until the next maintenance update. The slower disk can reduce performance for I/O-intensive workloads. * Maximum storage depends on the plan, service type, and cloud provider. It can be up to five times the plan's base storage size. * Cloud providers limit how many times you can increase storage between maintenance updates. If you reach the limit, run a maintenance update to optimize performance. * You cannot add storage during a maintenance update. * Dynamic disk sizing (DDS) is not supported on custom service plans. Pricing If you add storage when you create a service, **Additional disk storage** shows an estimated monthly cost. The **Service summary** lists plan storage plus additional storage. The estimated monthly price includes the additional storage cost. If you add storage to a running service, the Aiven Console shows the cost of the additional storage and related backups. The same costs appear on your invoices. ## Add or remove storage[​](#add-or-remove-storage "Direct link to Add or remove storage") ### Add storage when you create a service[​](#add-storage-when-you-create-a-service "Direct link to Add storage when you create a service") To add storage while you create a service: 1. In **Additional disk storage**, set the size with the slider or enter a value in GB. 2. Review the estimated monthly cost. 3. In the **Service summary**, click **Create service**. Change additional storage later on the running service, or enable automatic disk scaling with Aiven Autoscaler. ### Change storage on a running service[​](#change-storage-on-a-running-service "Direct link to Change storage on a running service") You cannot add or remove storage when service nodes are in the rebuilding state, for example during a maintenance update or a service upgrade. If you are removing disk storage: * Make sure the data in your service does not exceed the allocated storage. If it does, you cannot remove the additional storage. * Plan for the time it takes to rebuild the service. The time depends on the service. - Console - CLI - Terraform 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Actions** > **Manage additional storage**. 3. Change the disk storage. note * The price shown for the additional storage includes backup costs. * You can only remove storage that you previously added using this feature. To downgrade further, you can change your service plan. 4. Click **Save Changes**. Use [Aiven CLI](/docs/tools/cli.md) to add or remove additional storage using [`avn service update`](/docs/tools/cli/service-cli.md#avn-cli-service-update) with the `--disk-space-gib` flag to specify the total disk space to provide to your service. For example, if your service has a 80-GiB disk and you would like to add an extra 10-GiB disk, use: ``` avn service update --disk-space-gib 90 --project PROJECT_NAME SERVICE_NAME ``` note * When you perform a horizontal service upgrade or downgrade, remember to include all additional disks the service uses. For example, when switching from `Startup-4` to `Business-4` or from `Business-4` to `Startup-4`, include all the additional disks available for this service. * When you fork an existing service, include all additional disks the service uses. Use the `additional_disk_space` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). If you added storage, the additional storage is available immediately. If you removed additional storage, the service nodes go through a rolling restart. Depending on the service type and configuration, there might be a short downtime for services with no HA capabilities. note Storage optimization is performed at the next maintenance update after a change to the storage size. Due to cloud provider limitations, there is a limit on how many times storage can be increased between two maintenance updates. When this limit is reached, perform a maintenance update for performance optimization. Plan increases to avoid reaching this limit. Related pages * [Disk autoscaler](/docs/products/postgresql/howto/disk-autoscaler.md) * [Change the service plan](/docs/products/postgresql/howto/change-service-plan.md) * [Memory and out-of-memory conditions](/docs/products/postgresql/concepts/pg-shared-buffers.md) --- # Set up logical replication to Aiven for PostgreSQL® Aiven for PostgreSQL® represents an ideal managed solution for a variety of use cases; remote production systems can be completely migrated to Aiven using different methods including [using Aiven-db-migrate](/docs/products/postgresql/howto/migrate-aiven-db-migrate.md) or the standard [dump and restore method](/docs/products/postgresql/howto/migrate-pg-dump-restore.md). Whether you are migrating or have another use case to keep an existing system in sync with an Aiven for PostgreSQL service, you can address that by setting up a **logical replica** and replicating tables from a self-managed PostgreSQL cluster to Aiven. note This content also works with AWS RDS PostgreSQL 10+ and [Google CloudSQL PostgreSQL](https://cloud.google.com/sql/docs/release-notes#August_30_2021). ## Variables[​](#variables "Direct link to Variables") These are the placeholders you will need to replace in the code sample: | Variable | Description | | -------------- | ------------------------------------------------ | | `SRC_HOST` | Hostname of the source PostgreSQL database | | `SRC_PORT` | Port of the source PostgreSQL database | | `SRC_DATABASE` | Database Name of the source PostgreSQL database | | `SRC_USER` | Username of the source PostgreSQL database | | `SRC_PASSWORD` | Password of the source PostgreSQL database | | `SRC_CONN_URI` | Connection URI of the source PostgreSQL database | ## Requirements[​](#requirements "Direct link to Requirements") * PostgreSQL version 10 or newer * Connection between the source cluster's PostgreSQL port and Aiven for PostgreSQL cluster * Access to an superuser role on the source cluster * `wal_level` setting to `logical` on the source cluster. To verify and change the `wal_level` setting, see [the instructions on setting this configuration](/docs/products/postgresql/howto/migrate-aiven-db-migrate.md#pg_migrate_wal). note If you are using an AWS RDS PostgreSQL cluster as source, the `rds.logical_replication` parameter must be set to `1` (`true`) in the parameter group. ## Set up the replication[​](#set-up-the-replication "Direct link to Set up the replication") To create a logical replication, there is no need to install any extensions on the source cluster, but a superuser account is required. tip The `aiven_extras` extension enables the creation of a publish/subscribe-style logical replication without a superuser account, and it is preinstalled on Aiven for PostgreSQL servers. For more information on `aiven_extras`, see the dedicated [GitHub repository](https://github.com/aiven/aiven-extras). The following example assumes the `aiven_extras` extension is not available in the source PostgreSQL database. This example assumes a source database called `origin_database` on a self-managed PostgreSQL cluster. The replication will mirror three tables, named `test_table`, `test_table_2`, and `test_table_3`, to the `defaultdb` database on Aiven for PostgreSQL. The process to setup the logical replication is the following: 1. On the source cluster, connect to the `origin_database` with `psql`. 2. Create the `PUBLICATION` entry, named `pub_source_tables`, for the test tables: ``` CREATE PUBLICATION pub_source_tables FOR TABLE test_table,test_table_2,test_table_3 WITH (publish='insert,update,delete'); ``` tip In PostgreSQL 10 and above, `PUBLICATION` entries define the tables to be replicated, which are in turn `SUBSCRIBED` to by the receiving database. When creating a publication entry, the `publish` parameter defines the operations to transfer. In this example, all the `INSERT`, `UPDATE`, or `DELETE` operations will be transferred. 3. PostgreSQL's logical replication doesn't copy table definitions, that can be extracted from the `origin_database` with `pg_dump` and included in a `origin-database-schema.sql` file with: ``` pg_dump --schema-only --no-publications \ SRC_CONN_URI \ -t test_table -t test_table_2 -t test_table_3 > origin-database-schema.sql ``` 4. Connect via `psql` to the destination Aiven for PostgreSQL database and create the new `aiven_extras` extension: ``` CREATE EXTENSION aiven_extras CASCADE; ``` 5. Create the table definitions in the Aiven for PostgreSQL destination database within `psql`: ``` \i origin-database-schema.sql ``` 6. Create a `SUBSCRIPTION` entry, named `dest_subscription`, in the Aiven for PostgreSQL destination database to start replicating changes from the source `pub_source_tables` publication: ``` SELECT * FROM aiven_extras.pg_create_subscription( 'dest_subscription', 'host=SRC_HOST password=SRC_PASSWORD port=SRC_PORT dbname=SRC_DATABASE user=SRC_USER', 'pub_source_tables', 'dest_slot', TRUE, TRUE); ``` 7. Verify that the subscription has been created successfully. As the `pg_subscription` catalog is superuser-only, you can use the `aiven_extras.pg_list_all_subscriptions()` function from the `aiven_extras` extension: ``` SELECT subdbid, subname, subowner, subenabled, subslotname FROM aiven_extras.pg_list_all_subscriptions(); subdbid | subname | subowner | subenabled | subslotname ---------+-------------------+----------+------------+------------- 16401 | dest_subscription | 10 | t | dest_slot (1 row) ``` 8. Verify the subscription status: ``` SELECT * FROM pg_stat_subscription; subid | subname | pid | relid | received_lsn | last_msg_send_time | last_msg_receipt_time | latest_end_lsn | latest_end_time -------+-------------------+-----+-------+--------------+-------------------------------+-------------------------------+----------------+------------------------------- 16444 | dest_subscription | 869 | | 0/C002360 | 2021-06-25 12:06:59.570865+00 | 2021-06-25 12:06:59.571295+00 | 0/C002360 | 2021-06-25 12:06:59.570865+00 (1 row) ``` 9. Verify the data is correctly copied over the Aiven for PostgreSQL target tables. ## Remove unused replication setup[​](#remove-unused-replication-setup "Direct link to Remove unused replication setup") It is important to remove unused replication setups since the underlying replication slots in PostgreSQL forces the server to keep all the data needed to replicate since the publication creation time. If the data stream has no readers, there will be an ever-growing amount of data on disk until it becomes full. To remove an unused subscription, essentially stopping the replication, run the following command in the Aiven for PostgreSQL target database: * When the source database is accessible ``` SELECT * FROM aiven_extras.pg_drop_subscription('dest_subscription'); ``` * When there is no access to the source database ``` SELECT * FROM aiven_extras.pg_drop_subscription('dest_subscription', FALSE); ``` Verify the replication removal with: ``` SELECT * FROM aiven_extras.pg_list_all_subscriptions(); subdbid | subname | subowner | subenabled | subconninfo | subslotname | subsynccommit | subpublications ---------+---------+----------+------------+-------------+-------------+---------------+----------------- (0 rows) ``` ## Manage inactive or lagging replication slots[​](#manage-inactive-or-lagging-replication-slots "Direct link to Manage inactive or lagging replication slots") Inactive or lagging replication can cause problems in a database, like an ever-increasing disk usage not associated to any growth of the amount of data in the database. Filling the disk causes the database instance to stop serving clients and a loss of service. 1. Assess the replication slots status via `psql`: ``` SELECT slot_name,restart_lsn FROM pg_replication_slots; ``` The command output is like: ``` slot_name │ restart_lsn ──────────────┼───────────── pghoard_local │ 6E/16000000 dest_slot | 5B/8B0 (2 rows) ``` 2. Compare the `restart_lsn` values between the replication slot in analysis (`dest_slot` in the above example) and `pghoard_local`: the hexadecimal difference between the them states how many write-ahead-logging (WAL) entries are waiting for the target `dest_slot` connector to catch up. note In the above example the difference is 0x6E - 0x5B = 19 entries 3. If, after assessing the lag, the `dest_slot` connector results lagging or inactive: * If the `dest_slot` connector is still in use, a recommended approach is to restart the process and verify if it solves the problem. You can disable and enable the associated subscription using `aiven_extras`: ``` SELECT * FROM aiven_extras.pg_alter_subscription_disable('dest_subscription'); SELECT * FROM aiven_extras.pg_alter_subscription_enable('dest_subscription'); ``` * If the `dest_slot` connector is no longer needed, run the following command to remove it: ``` SELECT pg_drop_replication_slot('dest_slot'); ``` 4. In both cases, after the next PostgreSQL checkpoint, the disk space that the WAL logs have reserved for the `dest_subscription` connector should be freed up. note The checkpoint occurs only when: * 15 minutes have elapsed (we use a `checkpoint_timeout` value of 900 seconds), or * 5% of disk write operations is reached (the `max_wal_size` value is set to 5% of the instance storage). For further information about WAL and checkpoints, read the [PostgreSQL documentation](https://www.postgresql.org/docs/current/wal-configuration.html). note The recreation of replication slots gets enabled automatically for services created or updated as of January 2021. Additional details are outlined in [our blog post](https://aiven.io/blog/aiven-for-pg-recreates-logical-replication-slots). Replication slots are recreated when a maintenance update is applied or a failover occurs (for multi-node clusters), but they are not recovered after major version upgrades. --- # Tag your Aiven for PostgreSQL® service Add key-value tags to your Aiven for PostgreSQL® service to organize services and track ownership, cost allocation, and governance. Use tags to add metadata to Aiven services to categorize them or run custom logic on them. Typical uses include: * Tagging for governance to deploy services with specific tags only. * Tagging for internal cost reporting, ownership, allocation, and accountability. A tag is a key/value pair: * **Key**: A case-sensitive string that starts with a letter and consists of letters, numbers, dashes, and underscores. The maximum length for a key is 64 characters. * **Value**: A string value limited to 64 UTF-8 characters. Within a service, the tag keys must be unique. * Console * Terraform 1. In the service, click **Service settings**. 2. In the **Service status** section, click **Actions** > **Add service tags**. 3. Enter a key and value for each tag. 4. Click **Save changes**. Use the `tag` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Power on/off and delete your Aiven for PostgreSQL® service](/docs/products/postgresql/howto/power-cycle-service.md) --- # Track restore progress for your Aiven for PostgreSQL® service Track the restore progress of individual nodes in your Aiven for PostgreSQL® service during node replacement, forking, or maintenance, using the Aiven API. You can track restore progress for individual nodes during service node replacement by using the Aiven API. For example, use this endpoint to monitor the restore progress of a forked service or when applying maintenance. The service object exposes restore progress under `node_states[].progress_updates`: * `service.node_states[]` contains per-node state entries. * When a node is restoring or catching up, its `state` is typically `syncing_data`. * When the state is `syncing_data`, the node may include `progress_updates` with one or more phase objects. * Other node states don't include restore progress data. note `progress_updates` may be missing or empty even when a node is in `syncing_data`. This can occur when a restore completes before detailed progress is reported or when the service does not emit detailed progress counters. ## API endpoints[​](#api-endpoints "Direct link to API endpoints") Restore progress fields are part of the standard service response payload. * Get a single service (recommended for polling): `GET /project/{project}/service/{service_name}` * List services in a project: `GET /project/{project}/service` - Request - Response ``` curl -H "Authorization: aivenv1 API_TOKEN" https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME ``` Replace the following placeholders: * `API_TOKEN`: Your Aiven API token. * `PROJECT_NAME`: Your Aiven project name. * `SERVICE_NAME`: The name of your service. ``` { "service": { ... "node_states": [ { "node_name": "...", "state": "syncing_data", "progress_updates": [ { "phase": "basebackup", "completed": false, "current": 3410567, "min": 0, "max": 7569280, "unit": "bytes_uncompressed" } ] } ], ... } } ``` ## Node states[​](#node-states "Direct link to Node states") Common values for `node_states[].state` include: * `setting_up_vm`: The virtual machine is being created or initialized. * `syncing_data`: The node is restoring data or catching up. * `running`: The node is operating normally. * `leaving`: The node is leaving the cluster. * `unknown`: A transient or error state. ## `progress_updates` data model[​](#progress_updates-data-model "Direct link to progress_updates-data-model") `progress_updates` is a list of phase objects. When present, phases appear in the following order: 1. `prepare` 2. `basebackup` 3. `stream` 4. `finalize` Each phase object includes the following fields: ``` { "completed": false, "current": 3410567, "max": 7569280, "min": 0, "phase": "basebackup", "unit": "bytes_uncompressed" } ``` ### Field semantics[​](#field-semantics "Direct link to Field semantics") * `phase`: String, required. The restore phase. Possible values: `prepare`, `basebackup`, `stream`, and `finalize`. * `completed`: Boolean, required. Whether the phase is complete. * `current`: Number or null, optional. The current progress value. This field can be missing or null. * `min`: Number or null, optional. The starting value for the phase. This field can be missing or null. * `max`: Number or null, optional. The expected total value for the phase. This value can be missing, null, or change while the restore is in progress. * `unit`: String or null, optional. The unit for `current`, `min`, and `max`. New unit values can appear over time. Important considerations * Treat `unit` as an opaque identifier. Unknown values can appear. * `max` may change while a restore is in progress. * Not all phases report numeric counters. Some services only indicate phase completion. ## Why `max` values can change[​](#why-max-values-can-change "Direct link to why-max-values-can-change") The `current`, `min`, and `max` values are best-effort progress indicators. They can be based on estimates or on system state that changes over time. Treat `max` as the latest known expected total, not as a fixed guarantee. Common reasons `max` can change include: * The restore process discovers additional work after it starts, such as files, segments, or objects that become visible only after metadata is read. * New data is added on the backend while the node is catching up, which moves the completion point forward. This is common during incremental catch-up phases. * Progress is calculated from system state, such as replication lag, rather than from a fixed work queue. As the system state changes, the value is recalculated. * The service switches restore strategies during the operation, for example from snapshot restore to replication catch-up, which changes what the counters represent. As a result: * Phase percentage can decrease even when the restore operates normally. * Remaining-time estimates based on `max` are unreliable. * Sudden changes in `max` are expected unless the node remains in `syncing_data` longer than expected. ## Restore phase meanings[​](#restore-phase-meanings "Direct link to Restore phase meanings") Phase names are standardized, but the underlying work and the meaning of the counters are service-specific. * `prepare`: Prepares the node for restore. * `basebackup`: Restores the full backup. * `stream`: Applies incremental changes, such as replication or log replay. * `finalize`: Completes final steps before serving traffic. Not all restores include every phase. ## Compute phase progress percentages[​](#compute-phase-progress-percentages "Direct link to Compute phase progress percentages") You cannot reliably compute overall restore progress. You can compute a phase-specific progress percentage when `min`, `max`, and `current` are present and `max != min`. ``` pct = round(((current - min) / (max - min)) * 100, 1) ``` When handling progress values: * If any of `min`, `max`, or `current` is null or missing, display `n/a`. * If `max == min`, treat the percentage as undefined. * Expect the percentage to decrease when `max` changes. * Clamp displayed values to the range `[0, 100]`. ## Polling guidance[​](#polling-guidance "Direct link to Polling guidance") Progress updates are best-effort and refresh every 10 seconds while a node is in `syncing_data`. Poll the service state every 10 to 30 seconds. More frequent polling does not provide additional detail. For each `node_states[]` entry: * If `state` is not `syncing_data`, no restore progress is available. * If `state` is `syncing_data`: * If `progress_updates` is missing or empty, the node is restoring without detailed progress data. * Otherwise, the current phase is the last phase where `completed` is `false`. Stop polling when all nodes reach the `running` state or when a stall is detected. ### Stall detection[​](#stall-detection "Direct link to Stall detection") The API does not provide per-phase timestamps. To detect stalls, use a time-based threshold, such as a node remaining in `syncing_data` longer than expected. Do not rely on counters or `max` values to estimate remaining time. Related pages * [Backups](/docs/products/postgresql/concepts/pg-backups.md) * [Fork your service](/docs/products/postgresql/howto/fork-service.md) --- # Perform a PostgreSQL® major version upgrade PostgreSQL® in-place upgrades allows to upgrade an instances to a new major version without needing to fork and redirect the traffic. The whole procedure usually takes 60 seconds or less for small databases. ## Before you begin[​](#before-you-begin "Direct link to Before you begin") ### Create a read-only replica[​](#create-a-read-only-replica "Direct link to Create a read-only replica") We recommend [creating a read-only replica](/docs/products/postgresql/howto/create-read-replica.md) before the upgrade: * Very large databases may take a long time to upgrade and will be unreadable during the upgrade. You can use a read-only replica service to keep the data readable during an upgrade. * A PostgreSQL upgrade has some risk of downtime and data loss if the node goes down before the system is back in a normal state. A read-only replica can help reduce this risk. ### Test upgrading on a fork[​](#test-upgrading-on-a-fork "Direct link to Test upgrading on a fork") We recommend to test the upgrade on a fork of the database to be upgraded. Testing on a fork provides the benefit of verifying the impact of the upgrade for the specific service without affecting the running service, mostly to: 1. Ensure that the upgrade succeeds and is performed quickly enough, which might not be the case if there are many databases or large objects. Smaller node sizes with a large dataset can run into OOM issues during the `pg_dump/pg_restore` phase of `pg_upgrade --link`. A fork will reveal this scenario. 2. Test query performance directly after upgrade under real world load, when no statistics are available and caches are cold. ### Check extension compatibility[​](#check-extension-compatibility "Direct link to Check extension compatibility") Before upgrading, confirm that the extensions installed on your service are available on the PostgreSQL major version you plan to upgrade to. The automatic `pg_upgrade --check` only catches upgrade-blocking issues at the PostgreSQL binary level. It doesn't check whether your extensions have a build for the target version. Extension availability varies by PostgreSQL major version. An extension that works on your current version might not yet be available on the version you're upgrading to. Choose a target version that supports all the extensions your service depends on. To check extension compatibility: 1. List the extensions installed on your service: ``` SELECT extname, extversion FROM pg_extension; ``` 2. Check the [list of extensions for each PostgreSQL version](/docs/products/postgresql/reference/list-of-extensions-for-each-version.md) and confirm that each extension you use is available for your specific target major version, not only available in general. 3. [Update extensions](/docs/products/postgresql/howto/manage-extensions.md#update-an-extension) to their latest available version where possible. 4. If an extension isn't yet available for your target version, wait until it is, or choose a different target version. 5. Review the release notes or changelog for each extension for behavior changes, deprecated features, or migration steps tied to the target version. 6. If you test the upgrade on a fork, confirm the extensions load correctly and that dependent queries or functions still work after the upgrade. note Some extensions are only compatible with certain PostgreSQL major versions. For TimescaleDB, check the [TimescaleDB compatibility matrix](https://www.tigerdata.com/docs/deploy/self-hosted/upgrades/major-upgrade#plan-your-upgrade-path) to confirm your TimescaleDB version supports the PostgreSQL major version you're upgrading to. ## Limitations[​](#limitations "Direct link to Limitations") ### Upgrade to major versions in sequence[​](#upgrade-to-major-versions-in-sequence "Direct link to Upgrade to major versions in sequence") It's not recommended to upgrade across multiple major versions in a single pass. For example, if you're on version 1.0 and need to be on 4.0, first upgrade to 2.0, next to 3.0, and finally to 4.0. Avoid updating from 1.0 directly to 4.0. ### No downgrade to major versions[​](#no-downgrade-to-major-versions "Direct link to No downgrade to major versions") Downgrading to an earlier major version is not supported. ### Conditions that block an upgrade[​](#conditions-that-block-an-upgrade "Direct link to Conditions that block an upgrade") The following conditions prevent an upgrade from completing: * The service is a read-only replica. Upgrade the primary service instead. * A version upgrade is already in progress on the service. * Maintenance updates are pending for the service cluster. * The service nodes are not all in a running state. ## Upgrade to a major version[​](#upgrade-to-a-major-version "Direct link to Upgrade to a major version") * Console * CLI * API * Terraform * Kubernetes To upgrade a PostgreSQL service: 1. Log in to [Aiven Console](https://console.aiven.io/), and select the instance to upgrade. 2. Select **Service settings** from the sidebar of your service's page. 3. Go to the **Service management** section, click **Actions** > **Upgrade version**. 4. In the **Upgrade Aiven for PostgreSQL Confirmation** window, select the version to upgrade to from the dropdown menu. note When you select the version, the system checks the compatibility of the upgrade. warning Upon clicking **Upgrade**: * The system applies the upgrade **immediately**. * The PostgreSQL instance can't be restored to the previous version. * Backups created with an earlier major version are no longer visible in the Aiven Console and cannot be used for operations such as Point In Time Recovery (PiTR). You can only use backups created after the major version upgrade for such purposes. 5. Select **Upgrade**. 1. An automatic check is executed to confirm whether an upgrade is possible (`pg_upgrade --check`). 2. If the service has more than one node, any standby nodes are shut down and removed, as replication can not be performed during the upgrade. 3. The primary node starts an in-place upgrade to the new major version. 4. After a successful upgrade the primary node becomes available for use. A new full backup is initiated. 5. After completion of the full backup, new standby nodes are created for services with more than one node. 6. If the service is a configured to have a [read-only replica service](/docs/products/postgresql/howto/create-read-replica.md), the replica service will now be upgraded to the same version using the very same process. Read-only replicas remain readable during the upgrade of the primary service, but will go offline for the upgrade at this point. 7. `ANALYZE` will be automatically run for all tables after the upgrade to refresh table statistics and optimize queries. Use the [Aiven CLI](/docs/tools/cli.md) to upgrade your PostgreSQL service. The upgrade requires two steps: run a compatibility check first, then apply the version change. The compatibility check result expires after one hour. 1. Run the upgrade compatibility check using [`avn service task-create`](/docs/tools/cli/service-cli.md#avn-service-task-create): ``` avn service task-create SERVICE_NAME \ --project PROJECT_NAME \ --operation upgrade_check \ --target-version TARGET_VERSION ``` Note the `task_id` in the output. 2. Poll the task until `task_status` is `DONE` using [`avn service task-get`](/docs/tools/cli/service-cli.md#avn-service-task-get): ``` avn service task-get SERVICE_NAME \ --project PROJECT_NAME \ --task-id TASK_ID ``` If `success` is `false`, resolve the reported issue before continuing. 3. Apply the version upgrade within one hour of the successful check using [`avn service update`](/docs/tools/cli/service-cli.md#avn-cli-service-update): ``` avn service update SERVICE_NAME \ --project PROJECT_NAME \ -c pg_version=TARGET_VERSION ``` Replace the following: * `SERVICE_NAME`: name of your Aiven for PostgreSQL service. * `PROJECT_NAME`: name of your Aiven project. * `TARGET_VERSION`: target PostgreSQL major version, for example `16`. * `TASK_ID`: task ID returned in step 1. Use the Aiven API to upgrade your PostgreSQL service. The upgrade requires two steps: run a compatibility check first, then apply the version change. The compatibility check result expires after one hour. 1. Run the upgrade compatibility check using [ServiceTaskCreate](https://api.aiven.io/doc/#tag/Service/operation/ServiceTaskCreate): ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/task \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{"task_type": "upgrade_check", "upgrade_check": {"target_version": "TARGET_VERSION"}}' ``` Note the `task_id` in the response. 2. Poll the task until `task_status` is `DONE` using [ServiceTaskGet](https://api.aiven.io/doc/#tag/Service/operation/ServiceTaskGet): ``` curl --request GET \ 'https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/task/TASK_ID' \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' ``` If `success` is `false`, resolve the reported issue before continuing. 3. Apply the version upgrade within one hour of the successful check using [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate): ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{"user_config": {"pg_version": "TARGET_VERSION"}}' ``` Replace the following: * `PROJECT_NAME`: name of your Aiven project. * `SERVICE_NAME`: name of your Aiven for PostgreSQL service. * `YOUR_BEARER_TOKEN`: your Aiven API authentication token. * `TARGET_VERSION`: target PostgreSQL major version, for example `16`. * `TASK_ID`: task ID returned in step 1. Use the [`aiven_pg`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/pg) resource to update [`pg_version`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/pg#pg_version-1) inside `pg_user_config` and apply the change: ``` resource "aiven_pg" "example_pg" { project = var.aiven_project_name cloud_name = "google-europe-west1" plan = "startup-4" service_name = "my-postgres-service" pg_user_config { pg_version = "TARGET_VERSION" } } ``` Replace `TARGET_VERSION` with the target PostgreSQL major version, for example `"16"`. The Aiven Terraform provider runs a pre-flight compatibility check before applying the upgrade. Use the [PostgreSQL](https://aiven.github.io/aiven-operator/resources/postgresql.html) resource to update [`pg_version`](https://aiven.github.io/aiven-operator/resources/postgresql.html#spec.userConfig.pg_version-property) inside `spec.userConfig` and apply the manifest: 1. Update `pg_version` in your manifest: ``` apiVersion: aiven.io/v1alpha1 kind: PostgreSQL metadata: name: my-pg-service spec: authSecretRef: name: aiven-token key: token project: PROJECT_NAME cloudName: google-europe-west1 plan: startup-4 userConfig: pg_version: "TARGET_VERSION" ``` 2. Apply the updated manifest: ``` kubectl apply -f my-pg-service.yaml ``` Replace `PROJECT_NAME` with your Aiven project name and `TARGET_VERSION` with the target PostgreSQL major version, for example `"16"`. The operator runs a compatibility check before applying the upgrade. note A full backup of a large database may take a long time to complete. It may take some time before the standby node becomes available, as they can only be launched when a backup taken from the new version is available. Related pages * [Upgrade and failover procedures](/docs/products/postgresql/concepts/upgrade-failover.md) * [Manage extensions](/docs/products/postgresql/howto/manage-extensions.md) --- # Troubleshoot PostGIS® upgrade issues Troubleshoot issues that block PostGIS® extension upgrades on Aiven for PostgreSQL® services, and complete the upgrade safely. ## Upgrade PostGIS® to 3.6 with topology columns[​](#upgrade-postgis-to-36-with-topology-columns "Direct link to Upgrade PostGIS® to 3.6 with topology columns") Upgrade PostGIS topology columns manually when an Aiven for PostgreSQL® service has `topology.topoelement` or `topology.topoelementarray` columns and a pre-flight check blocks the upgrade from a PostGIS version earlier than 3.6 to 3.6 or later. This procedure applies only to services that store topology columns, which is a small subset of PostGIS users. If your databases do not use topology columns, the standard upgrade process applies and no manual steps are required. ### Why the upgrade is blocked[​](#why-the-upgrade-is-blocked "Direct link to Why the upgrade is blocked") PostGIS 3.6.0 introduced a breaking change that affects user tables with `topology.topoelement` or `topology.topoelementarray` columns. In PostGIS 3.6.0, the base type of the `topology.topoelement` domain changed from `integer[]` to `bigint[]`. The upgrade script then adds a CHECK constraint that validates all existing data. When user tables store data in these columns, the upgrade fails: 1. Existing data was stored as `integer[]`, using 4-byte integers per element. 2. After the domain change, PostgreSQL reinterprets the same byte data as `bigint[]`, using 8-byte integers. 3. Array element access returns incorrect values. For example, `ARRAY[1,2]` stored as `integer[]` reads as `ARRAY[1,0]` when interpreted as `bigint[]`. 4. The CHECK constraint validation fails, which blocks the extension upgrade. The Aiven pre-flight check detects these columns and blocks the upgrade **before** any changes are made to your service, so the service stays in a consistent state. The error message lists the affected databases, schemas, tables, and columns, references upstream issue #5983, and directs you to Aiven support. Use this procedure to prepare your databases, with support assistance, before you retry the upgrade. When the upgrade fails, you might see: * `CheckViolation` errors during `ALTER EXTENSION postgis_topology UPDATE`. * Error messages about invalid topology element references. * An extension upgrade blocked with an error that mentions topology columns. ### Detect affected tables[​](#detect-affected-tables "Direct link to Detect affected tables") Run the following query in each database to identify tables with topology columns: ``` SELECT n.nspname AS schema_name, c.relname AS table_name, a.attname AS column_name, a.atttypid::pg_catalog.regtype::text AS col_type FROM pg_catalog.pg_attribute a JOIN pg_catalog.pg_class c ON c.oid = a.attrelid JOIN pg_catalog.pg_namespace n ON n.oid = c.relnamespace WHERE a.attnum > 0 AND NOT a.attisdropped AND c.relkind IN ('r', 'p') AND a.atttypid IN ( pg_catalog.to_regtype('topology.topoelement'), pg_catalog.to_regtype('topology.topoelementarray') ) ORDER BY 1, 2, 3; ``` If the query returns any rows, the database requires the manual upgrade procedure. Keep the output, because later steps reference the affected tables and columns. ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A maintenance window. Extension upgrades require exclusive locks on system catalogs. * A database backup. Take a backup and test restoration before you start. * The list of affected tables from the detection query. * The `avnadmin` user, or a role that owns the affected tables and can run `ALTER TABLE`. Aiven for PostgreSQL does not provide superuser access; the `avnadmin` user has the required privileges. ### Upgrade procedure[​](#upgrade-procedure "Direct link to Upgrade procedure") Complete the following steps in order. When you use the bulk options, keep the same session or connection across steps, because the bulk commands store state in a temporary table. #### Step 1: Detach topology columns[​](#step-1-detach-topology-columns "Direct link to Step 1: Detach topology columns") For each affected table, convert the column to its base array type: ``` -- For topology.topoelement columns: ALTER TABLE schema_name.table_name ALTER COLUMN column_name TYPE integer[]; -- For topology.topoelementarray columns: ALTER TABLE schema_name.table_name ALTER COLUMN column_name TYPE integer[][]; ``` This detaches the column from the `topology.topoelement` domain type and converts it to a plain PostgreSQL array, which prevents the CHECK constraint from firing during the extension upgrade. For example, if the detection query returned `public.parcels.topo_ref` with type `topology.topoelement`: ``` ALTER TABLE public.parcels ALTER COLUMN topo_ref TYPE integer[]; ``` To detach all affected columns in a single command, run the following block. It saves the original domain-typed columns into a temporary `topology_columns_to_reattach` table that Step 3 uses for reattachment: ``` DO $$ DECLARE rec RECORD; base_type text; BEGIN CREATE TEMP TABLE IF NOT EXISTS topology_columns_to_reattach ( schema_name text NOT NULL, table_name text NOT NULL, column_name text NOT NULL, col_type text NOT NULL ) ON COMMIT PRESERVE ROWS; TRUNCATE topology_columns_to_reattach; INSERT INTO topology_columns_to_reattach (schema_name, table_name, column_name, col_type) SELECT n.nspname AS schema_name, c.relname AS table_name, a.attname AS column_name, a.atttypid::pg_catalog.regtype::text AS col_type FROM pg_catalog.pg_attribute a JOIN pg_catalog.pg_class c ON c.oid = a.attrelid JOIN pg_catalog.pg_namespace n ON n.oid = c.relnamespace WHERE a.attnum > 0 AND NOT a.attisdropped AND c.relkind IN ('r', 'p') AND a.atttypid IN ( pg_catalog.to_regtype('topology.topoelement'), pg_catalog.to_regtype('topology.topoelementarray') ); FOR rec IN SELECT schema_name, table_name, column_name, col_type FROM topology_columns_to_reattach LOOP base_type := CASE rec.col_type WHEN 'topology.topoelement' THEN 'integer[]' WHEN 'topology.topoelementarray' THEN 'integer[][]' ELSE NULL END; IF base_type IS NOT NULL THEN EXECUTE format( 'ALTER TABLE %I.%I ALTER COLUMN %I TYPE %s', rec.schema_name, rec.table_name, rec.column_name, base_type ); END IF; END LOOP; END $$; ``` #### Step 2: Upgrade the PostGIS extensions[​](#step-2-upgrade-the-postgis-extensions "Direct link to Step 2: Upgrade the PostGIS extensions") Connect as the `avnadmin` user, which has the privileges to manage extensions, and upgrade all PostGIS extensions: ``` ALTER EXTENSION postgis UPDATE; ALTER EXTENSION postgis_topology UPDATE; -- If installed: ALTER EXTENSION postgis_tiger_geocoder UPDATE; ALTER EXTENSION postgis_raster UPDATE; -- Run PostGIS post-upgrade maintenance: SELECT public.postgis_extensions_upgrade(); ``` #### Step 3: Reattach topology columns[​](#step-3-reattach-topology-columns "Direct link to Step 3: Reattach topology columns") For each detached column, reattach it to the domain type with explicit casting: ``` -- For topology.topoelement columns: ALTER TABLE schema_name.table_name ALTER COLUMN column_name TYPE topology.topoelement USING column_name::bigint[]::topology.topoelement; -- For topology.topoelementarray columns: ALTER TABLE schema_name.table_name ALTER COLUMN column_name TYPE topology.topoelementarray USING column_name::bigint[][]::topology.topoelementarray; ``` The `USING` clause casts the data in two steps: 1. `integer[]` to `bigint[]`. This widening cast expands each element from 4 to 8 bytes and is safe for all values. 2. `bigint[]` to `topology.topoelement`. This attaches the domain and applies CHECK constraint validation. For example: ``` ALTER TABLE public.parcels ALTER COLUMN topo_ref TYPE topology.topoelement USING topo_ref::bigint[]::topology.topoelement; ``` To reattach all columns in a single command, run the following block. It uses the `topology_columns_to_reattach` table created in Step 1: ``` DO $$ DECLARE rec RECORD; domain_type text; cast_type text; BEGIN FOR rec IN SELECT schema_name, table_name, column_name, col_type FROM topology_columns_to_reattach LOOP domain_type := rec.col_type; cast_type := CASE rec.col_type WHEN 'topology.topoelement' THEN 'bigint[]' WHEN 'topology.topoelementarray' THEN 'bigint[][]' ELSE NULL END; IF cast_type IS NOT NULL THEN EXECUTE format( 'ALTER TABLE %I.%I ALTER COLUMN %I TYPE %s USING %I::%s::%s', rec.schema_name, rec.table_name, rec.column_name, domain_type, rec.column_name, cast_type, domain_type ); END IF; END LOOP; END $$; ``` #### Step 4: Repair TopoGeometry data[​](#step-4-repair-topogeometry-data "Direct link to Step 4: Repair TopoGeometry data") This step applies to PostGIS 3.6.1 or later. If you upgraded to PostGIS 3.6.1 or later, run the repair function to fix any corrupt TopoGeometry element arrays: ``` -- Check whether the repair function is available: SELECT proname FROM pg_proc WHERE proname = 'fixcorrupttopogeometrycolumn'; -- If available, repair all topology layers: SELECT topology.FixCorruptTopoGeometryColumn( schema_name, table_name, feature_column ) FROM topology.layer; ``` The `FixCorruptTopoGeometryColumn()` function scans the feature column and rebuilds the element arrays with correct topology references. PostGIS 3.6.1 introduced this function to repair TopoGeometry element arrays affected by the `integer[]` and `bigint[]` reinterpretation. note This step is optional for PostGIS 3.6.0, but required for PostGIS 3.6.1 or later if your data uses TopoGeometry types rather than plain topology elements. ### Verify the upgrade[​](#verify-the-upgrade "Direct link to Verify the upgrade") After you complete the procedure, verify the results. 1. Confirm that all columns are restored. Re-run the detection query. It returns the same rows as before, with columns typed as `topology.topoelement` or `topology.topoelementarray`: ``` SELECT n.nspname AS schema_name, c.relname AS table_name, a.attname AS column_name, a.atttypid::pg_catalog.regtype::text AS col_type FROM pg_catalog.pg_attribute a JOIN pg_catalog.pg_class c ON c.oid = a.attrelid JOIN pg_catalog.pg_namespace n ON n.oid = c.relnamespace WHERE a.attnum > 0 AND NOT a.attisdropped AND c.relkind IN ('r', 'p') AND a.atttypid IN ( pg_catalog.to_regtype('topology.topoelement'), pg_catalog.to_regtype('topology.topoelementarray') ) ORDER BY 1, 2, 3; ``` 2. Check data integrity. Sample the data to confirm that element arrays contain valid topology IDs: ``` SELECT * FROM schema_name.table_name WHERE column_name IS NOT NULL LIMIT 10; ``` 3. Validate the topology. If your data uses TopoGeometry types with registered topologies, run the following query. An empty result set indicates no topology errors: ``` SELECT * FROM topology.ValidateTopology('your_topology_name'); ``` ### Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") #### CheckViolation error during reattachment[​](#checkviolation-error-during-reattachment "Direct link to CheckViolation error during reattachment") **Symptom**: `ALTER TABLE ... TYPE topology.topoelement` fails with a CHECK constraint violation. **Cause**: The CHECK constraint in PostGIS 3.6 or later validates that element IDs exist in the topology tables. If your data references deleted topology elements, reattachment fails. **Solution**: The solution depends on your PostGIS version. * On PostGIS 3.6.1 or later, run `FixCorruptTopoGeometryColumn()` before you reattach the columns: ``` SELECT topology.FixCorruptTopoGeometryColumn( schema_name, table_name, feature_column ) FROM topology.layer WHERE schema_name = 'your_schema' AND table_name = 'your_table'; ``` * On PostGIS 3.6.0, repair or remove the invalid references: ``` -- Identify rows with invalid topology references: SELECT * FROM schema_name.table_name WHERE column_name IS NOT NULL AND NOT topology.IsValidTopoElement(column_name); -- Option A: Set invalid references to NULL: UPDATE schema_name.table_name SET column_name = NULL WHERE column_name IS NOT NULL AND NOT topology.IsValidTopoElement(column_name); -- Option B: Delete rows with invalid references, if acceptable: DELETE FROM schema_name.table_name WHERE column_name IS NOT NULL AND NOT topology.IsValidTopoElement(column_name); ``` #### Column left as integer\[][​](#column-left-as-integer "Direct link to Column left as integer\[]") **Symptom**: After a failed upgrade attempt, columns are still typed as `integer[]` instead of `topology.topoelement`. **Cause**: A previous detach or reattach attempt failed and left the columns detached. **Solution**: Follow Step 3 to reattach the columns. The procedure is idempotent, so running it on columns that are already detached restores them correctly. #### FixCorruptTopoGeometryColumn function does not exist[​](#fixcorrupttopogeometrycolumn-function-does-not-exist "Direct link to FixCorruptTopoGeometryColumn function does not exist") **Symptom**: `SELECT topology.FixCorruptTopoGeometryColumn(...)` fails with a function does not exist error. **Cause**: PostGIS 3.6.1 introduced this function. It is not available on PostGIS 3.6.0. **Solution**: Choose one of the following options: * Upgrade to PostGIS 3.6.1 or later. This is the recommended option. * Skip Step 4 if your data does not use TopoGeometry types. * Repair TopoGeometry element arrays manually using the PostGIS 3.6.0 API. #### Permission denied for table[​](#permission-denied-for-table "Direct link to Permission denied for table") **Symptom**: `ALTER TABLE ... ALTER COLUMN` fails with a permission denied error. **Cause**: The procedure requires ownership of the table, or the privileges of the `avnadmin` user. **Solution**: Run the commands as the table owner, the `avnadmin` user, or a role that has `ALTER` privilege on the table. Related pages * [Manage extensions](/docs/products/postgresql/howto/manage-extensions.md) * [Supported extensions](/docs/products/postgresql/reference/list-of-extensions.md) * [List of extensions for each version](/docs/products/postgresql/reference/list-of-extensions-for-each-version.md) * [PostGIS issue #5983](https://github.com/postgis/postgis/issues/5983) * [PostGIS 3.6.0 release notes](https://postgis.net/docs/release_notes.html) * [PostGIS FixCorruptTopoGeometryColumn() documentation](https://postgis.net/docs/FixCorruptTopoGeometryColumn.html) --- # Use the PostgreSQL® dblink extension `dblink` is a [PostgreSQL® extension](https://www.postgresql.org/docs/current/dblink.html) that allows you to connect to other PostgreSQL databases and to run arbitrary queries. With [Foreign Data Wrappers](https://www.postgresql.org/docs/current/postgres-fdw.html) (FDW) you can uniquely define a remote **foreign server** in order to access its data. The database connection details like hostnames are kept in a single place, and you only need to create once a **user mapping** storing remote connections credentials. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") To create a `dblink` foreign data wrapper you need the following information about the PostgreSQL remote server: * `TARGET_PG_HOST`: The remote database hostname * `TARGET_PG_PORT`: The remote database port * `TARGET_PG_USER`: The remote database user to connect * `TARGET_PG_PASSWORD`: The remote database password for the `TARGET_PG_USER` * `TARGET_PG_DATABASE_NAME`: The remote database name note If you're using Aiven for PostgreSQL as remote server, the above details are available in the [Aiven console](https://console.aiven.io) > the service's **Overview** page or via the `avn service get` command with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn_service_get). ## Enable `dblink` extension on Aiven for PostgreSQL[​](#enable-dblink-extension-on-aiven-for-postgresql "Direct link to enable-dblink-extension-on-aiven-for-postgresql") To enable the `dblink` extension on an Aiven for PostgreSQL service: 1. Connect to the database with the `avnadmin` user. The following shows how to do it with `psql`, the service URI can be found in the [Aiven console](https://console.aiven.io/) the service's **Overview** page: ``` psql "postgres://avnadmin:[AVNADMIN_PWD]@[PG_HOST]:[PG_PORT]/[PG_DB_NAME]?sslmode=require" ``` tip If you're using Aiven for PostgreSQL as remote server, you can connect to a service with the `avnadmin` user with the `avn service cli` command with the [Aiven CLI](/docs/tools/cli/service-cli.md#avn-service-cli). 2. Create the `dblink` extension: ``` CREATE EXTENSION dblink; ``` ## Create a foreign data wrapper using `dblink_fdw`[​](#create-a-foreign-data-wrapper-using-dblink_fdw "Direct link to create-a-foreign-data-wrapper-using-dblink_fdw") To create a foreign data wrapper using the `dblink_fwd`: 1. Connect to the database with the `avnadmin` user. The following shows how to do it with `psql`, the service URI can be found in the [Aiven console](https://console.aiven.io/) the service's **Overview** page: ``` psql "postgres://avnadmin:[AVNADMIN_PWD]@[PG_HOST]:[PG_PORT]/[PG_DB_NAME]?sslmode=require" ``` 2. Create a user `user1` that will access the `dblink`: ``` CREATE USER user1 PASSWORD 'secret1' ``` 3. Create a remote server definition, named `pg_remote`, using `dblink_fdw` and the target PostgreSQL connection details: ``` CREATE SERVER pg_remote FOREIGN DATA WRAPPER dblink_fdw OPTIONS ( host 'TARGET_PG_HOST', dbname 'TARGET_PG_DATABASE_NAME', port 'TARGET_PG_PORT' ); ``` 4. Create a user mapping for the `user1` to automatically authenticate as the `TARGET_PG_USER` when using the `dblink`: ``` CREATE USER MAPPING FOR user1 SERVER pg_remote OPTIONS ( user 'TARGET_PG_USER', password 'TARGET_PG_PASSWORD' ); ``` 5. Enable `user1` to use the remote PostgreSQL connection `pg_remote`: ``` GRANT USAGE ON FOREIGN SERVER pg_remote TO user1; ``` ## Query data using a foreign data wrapper[​](#query-data-using-a-foreign-data-wrapper "Direct link to Query data using a foreign data wrapper") To query a foreign data wrapper you must be a database user having the necessary grants to the remote server definition. We can use `user1` from the previous example. To query the remote table `inventory` defined in the target PostgreSQL database pointed by the `pg_remote` server definition: 1. Connect with the Aiven for PostgreSQL service with the database user (`user1`) having the necessary grants to the remote server definition. 2. Establish the `dblink` connection to the remote target: ``` SELECT dblink_connect('my_new_conn', 'pg_remote'); ``` 3. Execute the query passing the foreign server definition as parameter: ``` SELECT * FROM dblink('pg_remote','SELECT item_id FROM inventory') AS target_inventory(target_item_id int); ``` 4. Check the results: ``` target_item_id ---------------- 1 2 3 (3 rows) ``` --- # Collect audit logs in Aiven for PostgreSQL® Enable and configure the [Aiven for PostgreSQL® audit logging feature](/docs/products/postgresql/concepts/pg-audit-logging.md) on your service. Access and visualize your logs to monitor activities on your databases. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * PostgreSQL version 11 or higher * `avnadmin` superuser role * Dev tool of your choice to interact with the feature * [Aiven Console](https://console.aiven.io/) * [Aiven CLI client](/docs/tools/cli.md) * [psql](https://www.postgresql.org/docs/current/app-psql.html) for advanced configuration * Aiven for OpenSearch® service for accessing and visualizing your logs ## Enable audit logging[​](#enable-audit-logging "Direct link to Enable audit logging") Enable audit logging by setting the `pgaudit.feature_enabled` parameter to `true` in your service's advanced configuration. Using the Aiven [console](https://console.aiven.io/), [CLI](/docs/tools/cli.md), or [psql](https://www.postgresql.org/docs/current/app-psql.html). * Aiven Console * Aiven CLI * psql important In [Aiven Console](https://console.aiven.io/), you can enable audit logging at the service level only. To enable it on a database or for a user's role, use [psql](https://www.postgresql.org/docs/current/app-psql.html). 1. Log in to [Aiven Console](https://console.aiven.io/), and go to your organization > project > Aiven for PostgreSQL service. 2. On the **Overview** page of your service, select **Service settings** from the sidebar. 3. On the **Service settings** page, go to the **Advanced configuration** section and select **Configure**. 4. In the **Advanced configuration** window, select **Add configuration options**, add the `pgaudit.feature_enabled` parameter, set it to `true`, and select **Save configuration**. Use the [Aiven CLI client](/docs/tools/cli.md) to run the [avn service update](/docs/tools/cli/service-cli.md) command. Update your service by setting the `pgaudit.feature_enabled` parameter's value to `true`. ``` avn service update -c pgaudit.feature_enabled=true SERVICE_NAME ``` important By default, audit logging does not emit any audit records. To trigger a logging operation and start receiving audit records, configure audit logging parameters as detailed in [Configure audit logging](#configure-audit-logging). note psql allows for fine-grained enablement of audit logging: on a database, for a user's role, or for a database-role combination. #### Enable on a database[​](#enable-on-a-database "Direct link to Enable on a database") important If you use PostgreSQL 14 or earlier, [upgrade](/docs/products/postgresql/howto/upgrade.md) to PostgreSQL 15 or later to enable `pgaudit` on your database. 1. [Connect to your Aiven for PostgreSQL service](/docs/products/postgresql/howto/list-code-samples.md). 2. Run the following query: ``` SET pgaudit.log='ddl'; ``` #### Enable for a user's role[​](#enable-for-a-users-role "Direct link to Enable for a user's role") important If you use PostgreSQL 14 or earlier, [upgrade](/docs/products/postgresql/howto/upgrade.md) to PostgreSQL 15 or later to enable `pgaudit` for a user's role. 1. [Connect to your Aiven for PostgreSQL service](/docs/products/postgresql/howto/list-code-samples.md). 2. Run the following query: ``` ALTER ROLE ROLE_NAME SET pgaudit.log='ddl'; ``` ## Configure audit logging[​](#configure-audit-logging "Direct link to Configure audit logging") Configure audit logging by setting [its parameters](https://github.com/pgaudit/pgaudit/tree/6afeae52d8e4569235bf6088e983d95ec26f13b7#readme) in the [Aiven Console](https://console.aiven.io/), with the [Aiven CLI](/docs/tools/cli.md), or using [psql](https://www.postgresql.org/docs/current/app-psql.html). important * Advanced configuration of the audit logging feature requires using [psql](https://www.postgresql.org/docs/current/app-psql.html). * Any configuration changes take effect only on new connections. For information on all the audit logging configuration parameters, refer to [Settings](https://github.com/pgaudit/pgaudit/tree/6afeae52d8e4569235bf6088e983d95ec26f13b7). * Aiven Console * Aiven CLI * psql important In the [Aiven Console](https://console.aiven.io/), you can enable audit logging on a service only. To enable it on a database or for a user's role, use [psql](https://www.postgresql.org/docs/current/app-psql.html). 1. Log in to [Aiven Console](https://console.aiven.io/), and go to your organization > project > Aiven for PostgreSQL service. 2. On the **Overview** page of your service, select **Service settings** from the sidebar. 3. On the **Service settings** page, go to the **Advanced configuration** section and select **Configure**. 4. In the **Advanced configuration** window, select **Add configuration options**, find a desired parameter (all prefixed with `pgaudit.log`), set its value as needed, and select **Save configuration**. Use the [Aiven CLI client](/docs/tools/cli.md) to configure audit logging on your service by running the following command: ``` avn service update -c pgaudit.PARAMETER_NAME=PARAMETER_VALUE SERVICE_NAME ``` note psql allows for fine-grained configuration of audit logging: on a database, for a user's role, or for a database-role combination. #### Configure on a database[​](#configure-on-a-database "Direct link to Configure on a database") important If you use PostgreSQL 14 or earlier, [upgrade](/docs/products/postgresql/howto/upgrade.md) to PostgreSQL 15 or later to configure `pgaudit` on your database. 1. [Connect to your Aiven for PostgreSQL service](/docs/products/postgresql/howto/list-code-samples.md). 2. Run the following query: ``` SET pgaudit.PARAMETER_NAME='all'; ``` #### Configure for a user's role[​](#configure-for-a-users-role "Direct link to Configure for a user's role") important If you use PostgreSQL 14 or earlier, [upgrade](/docs/products/postgresql/howto/upgrade.md) to PostgreSQL 15 or later to configure `pgaudit` for a user's role. 1. [Connect to your Aiven for PostgreSQL service](/docs/products/postgresql/howto/list-code-samples.md). 2. Run the following query: ``` ALTER ROLE_NAME SET pgaudit.PARAMETER_NAME=PARAMETER_VALUE; ``` ## Configure session audit logging[​](#configure-session-audit-logging "Direct link to Configure session audit logging") Session audit logging allows recording detailed logs of all SQL statements and commands executed during a database session in the system's backend. important If you use PostgreSQL 14 or earlier, [upgrade](/docs/products/postgresql/howto/upgrade.md) to PostgreSQL 15 or later to use the session audit logging. To enable the session audit logging, run the following query: ``` ALTER DATABASE DATABASE_NAME SET pgaudit.log='ddl'; ``` Example ``` ALTER DATABASE defaultdb SET pgaudit.log='read,ddl'; ``` See also For more details on how to set up, configure, and use session audit logging, check [Session audit logging](https://github.com/pgaudit/pgaudit/tree/6afeae52d8e4569235bf6088e983d95ec26f13b7). ## Access your logs[​](#access-your-logs "Direct link to Access your logs") Choose one of the [tools or methods for accessing, monitoring, or analyzing your audit logs](/docs/products/postgresql/concepts/pg-audit-logging.md#collecting-and-visualizing-logs). Example: **Aiven for OpenSearch®** * Aiven Console * Aiven CLI * Aiven API Access your Aiven for PostgreSQL logs by [enabling OpenSearch log integration](/docs/products/opensearch/howto/opensearch-log-integration.md). Use the [Aiven CLI](/docs/tools/cli.md) to create the service integration. ``` avn service integration-create --project $PG_PROJECT \ -t logs \ -s $PG_SERVICE_NAME \ -d $OS_SERVICE_NAME ``` After the service integration is set up and propagated to the service configuration, the logs are available in Aiven for OpenSearch. Each log record emitted by audit logging is stored in Aiven for OpenSearch as a single message, which cannot be guaranteed for external integrations such as [Remote Syslog](/docs/integrations/rsyslog.md). Call the [ServiceIntegrationCreate](https://api.aiven.io/doc/#tag/Service_Integrations/operation/ServiceIntegrationCreate) endpoint passing the following parameters in the request body: * `integration_type`: `logs` * `source_service`: the name of an Aiven for PostgreSQL * `destination_service`: the name of an Aiven for OpenSearch service ``` curl --request POST \ --url https://api.aiven.io/v1/project/{project_name}/integration \ --header 'Authorization: Bearer REPLACE_WITH_YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "integration_type": "logs", "source_service": "REPLACE_WITH_POSTGRESQL_SERVICE_NAME", "destination_service": "REPLACE_WITH_OPENSEARCH_SERVICE_NAME", }' ``` ## Visualize your logs[​](#visualize-your-logs "Direct link to Visualize your logs") Choose one of the [tools or methods for visualizing your collected audit logs](/docs/products/postgresql/concepts/pg-audit-logging.md#collecting-and-visualizing-logs). Example: **[OpenSearch Dashboards](/docs/products/opensearch/dashboards/get-started.md)** 1. Integrate your Aiven for PostgreSQL with Aiven for OpenSearch by following the steps in [Access-your-logs](/docs/products/postgresql/howto/use-pg-audit-logging.md#access-your-logs). 2. Go to OpenSearch Dashboards. 3. Set up an **Index Pattern** in OpenSearch Dashboards to match your audit logs index. For guidance, see [Create index patterns in OpenSearch Dashboards](https://opensearch.org/docs/latest/dashboards/index-patterns/). The index name for audit logs typically starts with `avnlog-pg-`. 4. To filter the audit logs: 1. Go to **Discover**. 2. Select your audit logs index pattern. 3. In the filter tool, set the value of `AIVEN_AUDIT_FROM` to `pg`. 4. Apply the filter. 5. Preview and analyze the logs. ![PgAudit logs in OpenSearch Dashboards](/docs/assets/images/pgaudit-logs-in-os-dashboards-892214ca76faae406b619497f0b8c685.png) note If the index pattern in OpenSearch Dashboards had been configured before you enabled the service integration, the audit-specific AIVEN\_AUDIT\_FROM field is not available for filtering. Refresh the fields list for the index in OpenSearch Dashboards under **Stack Management** > **Index Patterns** > Your index pattern > **Refresh field list**. ## Disable audit logging[​](#disable-audit-logging "Direct link to Disable audit logging") Disable audit logging by setting the `pgaudit.feature_enabled` parameter to `false` in your service's advanced configuration. Use the Aiven [console](https://console.aiven.io/), [CLI](/docs/tools/cli.md), or [psql](https://www.postgresql.org/docs/current/app-psql.html). * Aiven Console * Aiven CLI * psql important In the [Aiven Console](https://console.aiven.io/), you can disable audit logging on a service only. To disable it on a database or for a user's role, use [psql](https://www.postgresql.org/docs/current/app-psql.html). 1. Log in to [Aiven Console](https://console.aiven.io/), and go to your organization > project > Aiven for PostgreSQL service. 2. On the **Overview** page of your service, select **Service settings** from the sidebar. 3. On the **Service settings** page, go to the **Advanced configuration** section and select **Configure**. 4. In the **Advanced configuration** window, select **Add configuration options**, add the `pgaudit.feature_enabled` parameter, set it to `false`, and select **Save configuration**. Use the [Aiven CLI client](/docs/tools/cli.md) to run the [avn service update](/docs/tools/cli/service-cli.md) command. Update your service by setting the `pgaudit.feature_enabled` parameter's value to `false`. ``` avn service update -c pgaudit.feature_enabled=false SERVICE_NAME ``` note psql allows you to disable audit logging on a few levels: database, user's role, or database-role combination. #### Disable on a database[​](#disable-on-a-database "Direct link to Disable on a database") important If you use PostgreSQL 14 or earlier, [upgrade](/docs/products/postgresql/howto/upgrade.md) to PostgreSQL 15 or later to disable `pgaudit` on your database. 1. [Connect to your Aiven for PostgreSQL service](/docs/products/postgresql/howto/list-code-samples.md). 2. Run the following query: ``` ALTER DATABASE DATABASE_NAME SET pgaudit.log='none'; ``` #### Disable for a user's role[​](#disable-for-a-users-role "Direct link to Disable for a user's role") important If you use PostgreSQL 14 or earlier, [upgrade](/docs/products/postgresql/howto/upgrade.md) to PostgreSQL 15 or later to disable `pgaudit` for a user's role. 1. [Connect to your Aiven for PostgreSQL service](/docs/products/postgresql/howto/list-code-samples.md). 2. Run the following query: ``` ALTER ROLE_NAME SET pgaudit.log='none'; ``` --- # Use the PostgreSQL® pg\_cron extension The `pg_cron` extension is a cron-based job scheduler for PostgreSQL (10 or higher) that runs inside the database. `pg_cron` can run multiple jobs in parallel, but it runs at most one instance of a job at a time. If a second run is supposed to start before the first one finishes, the second run is queued and started as soon as the first run completes. ## CRON syntax[​](#cron-syntax "Direct link to CRON syntax") The schedule uses the standard [cron syntax](https://en.wikipedia.org/wiki/Cron), where an asterisk (`*`) means *execute at every time interval*, and a specific number means *execute exclusively at this specific time*: ``` ┌───────────── min (0 - 59) │ ┌────────────── hour (0 - 23) │ │ ┌─────────────── day of month (1 - 31) │ │ │ ┌──────────────── month (1 - 12) │ │ │ │ ┌───────────────── day of week (0 - 6) (0 to 6 are Sunday to │ │ │ │ │ Saturday, or use names; 7 is also Sunday) │ │ │ │ │ │ │ │ │ │ * * * * * ``` You can also use `[1-59] seconds` to schedule a job based on an interval. ## Enable `pg_cron` for specific users[​](#enable-pg_cron-for-specific-users "Direct link to enable-pg_cron-for-specific-users") To use the `pg_cron` extension: 1. Connect to the database as `avnadmin` user and make sure to use the `defaultdb` database: ``` CREATE EXTENSION pg_cron; ``` 2. Optional: Grant usage permission to regular users: ``` GRANT USAGE ON SCHEMA cron TO janedoe; ``` ## Set up the cron job[​](#set-up-the-cron-job "Direct link to Set up the cron job") ### List all the jobs[​](#list-all-the-jobs "Direct link to List all the jobs") To view the full list of existing jobs, run: ``` SELECT * FROM cron.job; jobid | schedule | command | nodename | nodeport | database | username | active | jobname ------+-------------+--------------------------------+-----------+----------+----------+-----------+--------+------------------------- 106 | 29 03 * * * | vacuum freeze test_table | localhost | 8192 | database1| adminuser | t | database1 manual vacuum 1 | 59 23 * * * | vacuum freeze pgbench_accounts | localhost | 8192 | postgres | adminuser | t | manual vacuum (2 rows) ``` ### Schedule a job[​](#schedule-a-job "Direct link to Schedule a job") To schedule a new job, run: Vacuum every day at 10:00am (GMT) ``` SELECT cron.schedule('nightly-vacuum', '0 10 * * *', 'VACUUM'); ``` ### Unschedule a job[​](#unschedule-a-job "Direct link to Unschedule a job") To unschedule a job, you have two options: * By using the `jobname`: Unschedule jobs using jobname ``` SELECT cron.unschedule('nightly-vacuum' ); ``` * By using the `jobid`: Unschedule jobs using jobid ``` SELECT cron.unschedule(1); ``` ### View completed jobs[​](#view-completed-jobs "Direct link to View completed jobs") To list all completed job runs, run: ``` select * from cron.job_run_details order by start_time desc limit 5; +------------------------------------------------------------------------------------------------------------------------------------------------------------------+ ¦ jobid ¦ runid ¦ job_pid ¦ database ¦ username ¦ command ¦ status ¦ return_message ¦ start_time ¦ end_time ¦ +-------+-------+---------+----------+----------+-------------------+-----------+------------------+-------------------------------+-------------------------------¦ ¦ 10 ¦ 4328 ¦ 2610 ¦ postgres ¦ marco ¦ select process() ¦ succeeded ¦ SELECT 1 ¦ 2023-02-07 09:30:00.098164+01 ¦ 2023-02-07 09:30:00.130729+01 ¦ ¦ 10 ¦ 4327 ¦ 2609 ¦ postgres ¦ marco ¦ select process() ¦ succeeded ¦ SELECT 1 ¦ 2023-02-07 09:29:00.015168+01 ¦ 2023-02-07 09:29:00.832308+01 ¦ ¦ 10 ¦ 4321 ¦ 2603 ¦ postgres ¦ marco ¦ select process() ¦ succeeded ¦ SELECT 1 ¦ 2023-02-07 09:28:00.011965+01 ¦ 2023-02-07 09:28:01.420901+01 ¦ ¦ 10 ¦ 4320 ¦ 2602 ¦ postgres ¦ marco ¦ select process() ¦ failed ¦ server restarted ¦ 2023-02-07 09:27:00.011833+01 ¦ 2023-02-07 09:27:00.72121+01 ¦ ¦ 9 ¦ 4320 ¦ 2602 ¦ postgres ¦ marco ¦ select do_stuff() ¦ failed ¦ job canceled ¦ 2023-02-07 09:26:00.011833+01 ¦ 2023-02-07 09:26:00.22121+01 ¦ +------------------------------------------------------------------------------------------------------------------------------------------------------------------+ (10 rows) ``` --- # Use the PostgreSQL® pg\_repack extension [`pg_repack`](https://reorg.github.io/pg_repack/) is a PostgreSQL® extension that allows you to efficiently reorganize tables to remove any excess bloat the tables have accumulated. Reorganizing a table may take some time, but `pg_repack` tries to minimize the locks required to continue online operations. note Before you install the `pg_repack` extension, verify the version of the extension is supported on the PostgreSQL version you are using. See the [supported versions](https://reorg.github.io/pg_repack/). ## Variables[​](#variables "Direct link to Variables") The following variables need to be substituted when running the commands. | Variable | Description | | -------------- | ----------------------------------------------------- | | `HOSTNAME` | Hostname for PostgreSQL connection | | `PORT` | Port for PostgreSQL connection | | `DATABASENAME` | Database Name of your Aiven for PostgreSQL connection | | `TABLENAME` | Name of the table to reorganize | ## Use `pg_repack` extension[​](#use-pg_repack-extension "Direct link to use-pg_repack-extension") To use the `pg_repack` extension: 1. Connect to the database as `avnadmin` user, and run the following command to create the extension: ``` CREATE EXTENSION pg_repack; ``` 2. Run the `pg_repack` command on the table to reorganize it. ``` pg_repack -k -U avnadmin -h -p -d -t ``` note * Using `-k` skips the superuser checks in the client. This setting is useful when using `pg_repack` on platforms that support running it as non-superusers. * The target table must have a PRIMARY KEY, or at least a UNIQUE total index on a NOT NULL column. Related pages * [`pg_repack` documentation](https://reorg.github.io/pg_repack/) * [Install or update extension](/docs/products/postgresql/howto/manage-extensions.md) --- # Enable and use pgvector on Aiven for PostgreSQL® The [pgvector extension](/docs/products/postgresql/concepts/pgvector.md) allows you to perform the vector similarity search and use embedding techniques directly in Aiven for PostgreSQL. tip Use an AI assistant connected to [Aiven MCP](/docs/tools/mcp-server.md) to configure `pgvector` for an AI agent memory store, retrieval-augmented generation (RAG) pipeline, or semantic search. For example: > Enable `pgvector` on `my-pg-service`, create a table to store 1536-dimensional embeddings, and create an HNSW index for cosine similarity search. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven account * Running Aiven for PostgreSQL service * psql and a psql CLI client * Vector embeddings generated (for example, with the [OpenAI API](https://platform.openai.com/docs/api-reference/embeddings/create) client) ## Enable pgvector[​](#enable-pgvector "Direct link to Enable pgvector") Run the CREATE EXTENSION statement from a client such as psql connected to your service. This is needed for each database to perform the similarity search on. 1. [Connect to your Aiven for PostgreSQL service](/docs/products/postgresql/howto/list-code-samples.md) using, for example, the psql client (CLI). 2. Connect to your database. ``` \c database-name ``` 3. Run the CREATE EXTENSION statement. ``` CREATE EXTENSION vector; ``` ## Store embeddings[​](#store-embeddings "Direct link to Store embeddings") 1. Create a table to store the generated vector embeddings. Use the CREATE TABLE SQL command, adjusting the dimensions as needed. ``` CREATE TABLE items (id bigserial PRIMARY KEY, embedding vector(3)); ``` note As a result, the `items` table is created. The table includes the `embedding` column, which can store vectors with three dimensions. 2. Run the INSERT statement to store the embeddings generated with, for example, the [OpenAI API](https://platform.openai.com/docs/api-reference/embeddings/create) client. ``` INSERT INTO items (embedding) VALUES ('[1,2,3]'), ('[4,5,6]'); ``` note As a result, two new rows are inserted into the `items` table with the provided embeddings. ## Perform similarity search[​](#perform-similarity-search "Direct link to Perform similarity search") To calculate similarity, run the SELECT statements using the built-in vector operators. ``` SELECT * FROM items ORDER BY embedding <-> '[3,1,2]' LIMIT 5; ``` note As a result, the query computes the L2 distance between the selected vector and the vectors stored in the `items` table, arrange the results based on the calculated distance, and outputs its top five nearest neighbors (most similar items). Operators for calculating similarity * `<->` - Euclidean distance (L2 distance) * `<#>` - negative inner product * `<=>` - cosine distance ## Add indices[​](#add-indices "Direct link to Add indices") You can add an index on the vector column to use the *approximate* nearest neighbor search (instead of the default the *exact* nearest neighbor search). This can improve query performance with an ignorable cost on recall. Add an index is possible for all distance functions (L2 distance, cosine distance, inner product). To add an index, run a query similar to the following: ``` CREATE INDEX ON items USING ivfflat (embedding vector_l2_ops) WITH (lists = 100); ``` note As a result, the index is added to the `embedding` column for the L2 distance function. ## Disable pgvector[​](#disable-pgvector "Direct link to Disable pgvector") To stop the pgvector extension and remove it from a database, run the following SQL command: ``` DROP EXTENSION vector; ``` Related pages * [pgvector for AI-powered search in Aiven for PostgreSQL®](/docs/products/postgresql/concepts/pgvector.md) * [pgvector README on GitHub](https://github.com/pgvector/pgvector/blob/master/README.md) --- # Visualize PostgreSQL® data with Grafana® PostgreSQL® can hold a wide variety of types of data, and creating visualisations helps gather insights on top of raw figures. Aiven can set up the Grafana® and the integration between the two services for you. ## Integrate PostgreSQL and Grafana[​](#integrate-postgresql-and-grafana "Direct link to Integrate PostgreSQL and Grafana") 1. On the **Overview** page for your Aiven for PostgreSQL service, go to **Manage integrations** and choose the **Monitor Data in Grafana** option. 2. Choose either a new or existing Grafana service. * A new service will ask you to select the cloud, region and plan to use. You should also give your service a name. The service overview page shows the nodes rebuilding, and indicates when they are ready. * If you're already using Grafana on Aiven, you can integrate your PostgreSQL as a data source for that existing Grafana. 3. On the **Overview** page for your Aiven for Grafana service, select the **Service URI** link. The username and password for your Grafana service is also available on the service's **Overview** page. Now your Grafana service is connected to PostgreSQL as a data source and you can go ahead and visualise your data. ## Visualize PostgreSQL data in Grafana[​](#visualize-postgresql-data-in-grafana "Direct link to Visualize PostgreSQL data in Grafana") In Grafana, create a dashboard and add a panel to it. The datasource dropdown shows `--Grafana--` by default, but you will also see your PostgreSQL service. ![Screenshot of a Grafana panel showing PostgreSQL logo](data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAXwAAAAyCAYAAABfwmBiAAAAAXNSR0IArs4c6QAAADhlWElmTU0AKgAAAAgAAYdpAAQAAAABAAAAGgAAAAAAAqACAAQAAAABAAABfKADAAQAAAABAAAAMgAAAAClM6YSAAAcbklEQVR4Ae1dCWAN1/f+8rJLhIQsiF2K2Gpfa6t9qZ1SaqcbWlSValVpq7ai/qotfpSillZbFLVUtXZBK1W77BKRXfaX/zn3ZV7m5b1EXvISwT3Mmzt37jbfzP3uueeemViV9SiXgadA4mKjYUX/+D//ZO5EWBejBMUZjhKiS2cYp5yTe4mAREAi8DghoHmcGmu5tmYga5TThbKODWvJ/axhWnkkEZAISASKMwJPFeELmlczuzpcnO+SbJtEQCIgEbAAAk8V4RviZYLtRVRO8Ya55ZFEQCIgEXjcELB53Bqc7/Yyj5MpnrV8qwwKiHCmSV8YeKSdPt/YFmFG+8odYV++JazdasHasUwR1iyrkggUbwTSEyORfv8KkkNOIPnOYZONtXpqFm1joojddaSe6+KtSGJM/laZeU2iKCMLHQGNkwecG06GnVejQq9LViAReNwRSAk7j3i/FdAmhBtcylOj4QsFP1PLz0LAKCLrlAwVKwRKNn0btmXrID38HFKu/wBtfCCQ9qBYtVE2RiLwSBGwKQGNc0XY1ehHilFjcJ+JOfq2QZOKzIY/oJwtVtcvAZdiMMSofXQM0OADHgNMSo4nTKaWkZZDwL5yF0H2aWGnkHT2U2ij/5Nkbzl4ZUlPCgKkAHHf4D7CfYUVJO47aikS+mWyX1inhKi3gqMzhp+LR2yauhlFFVZp9JlBVUxRNULWYyYC9hXbiBypN3eLvUajgau7O2xtbc0sSSaXCDy5CKSmpCDq3j1otVpwX7Hxag7uO8l3DugvutAJ39vBCrOfcdRX6FvSGu/R8Qz/RH1c8QvIYaA43ROb0tVFc9Ljg8ULc65lyyIy/C4SExKKUzNlWyQCjxQBRycnlPHwROTdu+C+wqL0HaVhhUr4bL6ZVM0BLraGi6D9y9shJi0D/8al47eIVNGWTu628HbU4FRUGm3pSvssus+dxnM/a9GGyMLMQkDj4CrSW2Xa7G3t7CTZm4WgTPw0IMAKkDLrVfqK0neU67cI4fuW1KCkjZWeqJnoWavvTCSeneyVikdXsleCiE3N0KfjAeBUVCEuxpnidVNx+tbJgERAIiAReDIQyDfhK6aazh6GdlR/0tq9HTR6As8LTOpBobazdV6ymJ+moKRe0Pzmt1jmkAhIBCQCFkUg34Rfgcwv2cmeW8Y2elPC5hvegpK0+tOctrmrjZgdKJFs6nmUInn9UaIv65YISAQKE4F8E35cHol5V0gKVtxMIqLPmcg7u9tgFJl4mPx5EFjo64h3LL6oK6m8MB8kWbZEQCJQ/BHItx++f5wWuZE+n3vhVJzwxsmN7BmigxFpeOlcAhZc1XnuDKBFXR4EClOyfPFzHogKs35ZtkRAIiARKGoECsSqp6JS0SmbDZ8vII4WYZnAeVAwR9YHpAgvnc2Nnchv3xGnjscVgr++5TX9Bg0bo3qNGuJSr1+/hkt+5825bIO0uX3CISPD8oPTqHET0b1XH6z7ejX27/nZoC2P+wH7I3t7V4Cnp4e4lOioaISEhKBevbpwrVANITf8UbdpG2jT0xEXE43LfqcREBBIX+Aw9Cp73HGQ7ZcIKAgUiPDXB6agUwXjIub/m5gr2fu66Lx6ghO1CEo0JDEeJGZQ/tWNHTHA2xbrb+vcNpUGF6f9iNHj8fKY8aJJ2Uliw7qv8O26r81q7lcbvkO16j455rlz6xbGjhic4/n8nOD6ypR1R8WKlfOTvdjm4cGxXpMWaNFnDCKTdARe1yYNrWuVx55LYUhMTkXLVn3hWbokzl8PhrasFm3rdYDfvk24fO6kRUm/UZNm+O+KPxLi44stXrJhTwcCxmxtxnWfup+O0zGpaF42a6E2+IEWO4OMX6N1IWeeUdVsMbq6oasmu2TuDEjD/26mIuiBjvwP3k3DhoBk9K9iXSwJn0ly+qw5KOdVHru2b8FffxzTE/9GIvlWz7XDgEFD0Zr2ixbMw43rV/OGaqZmGXkvAvfvRxrluViAmYNRYU94BGv2jbsMxp9XI/RX2q1ZTXx1+CoeJKWIuKCIGLHv3coXR/1uICA8GZ1fnIjIoJu4G56VT19APgKlS5fGZ5+vwvFjRzF3luF3TfJR3CPJ4urqhv6Dh2Lvzz8iNCS4wG14aeQYXPn3X5w7faLAZckCzEOgQITPVdUuS9qTXZaWvvyirjOpm8Fkv7m9A3xLK0sGWeld7IDRtW0woLo1ZpxJwcEQ3UtXy/9LwbEajuTeCfLTV5f26MPT351Db3xa4ZWxI3A3LFQ0aOM6Xbsu+p0Db7u+34J5nyzG27Pfxyujh5vV6G2bN9BAss2sPDKxIQJl3FyRZFeaInXrQh6lnZGckoa0NGMz476TV9CvbV3s+P1vhNI3P8qULWMxwo+OjsaqzxfjxJ9/GDbwMTqq5lMTQ0eMQlhoCPb89EOBWz5q3Cs4S2QvCb/AUJpdQIEI39fVCqWcmLyzCPxgsPFbst91soMvvSwZm5KOnTfTcTBI1+k6e2swoJo1XOyoHCL2Ne1tMeMksIPSMMkfDEujfBqcDDfupGZfqYUyDB85Fp5e5fDauJf1ZM9FM8mzODuXRP2GDcmO74f3352ONes3k/Y/ARvJxGMJadH6OUydMQvxZB4Y//JQpKfrZlOTps7Ac+064J9LFzFvzky8+/48NGzcFF+uXI6RYyfA3dMTyclJ8Dt7Bh+9/y4eth7AWli75zuTDbwSHiQ8gP8/F/HRB7PA3+tQ5Jtvt8He3h5HDx1Er779wd+46dO1gzjd6rm2GDJsJGo8UxMpKcn479/L+ODdGaINSv7C3NvQd3YSMjV5rse1pCMqe7jCrVUJ/HP7Lv4L0H021t7OBgPbNUBFdxfUqFCGFAwtoqN1mr857Rv28hg8174DMmjdYP/ePfhl9y5xb9jU165jZ8TFxtC+Exo3a4EZU17TF833zdnZGZ/Me1/EDXnpZXTs1BXaDC1+/mGn0Kr5xJgJr8HVzQ3BQYHo1LUHIsLDsGHtV7jif1lflhIo6+6BOfM+wb49u9GtR2+4upbBpYvnsXThAv1957eV33lvLnx8aiHsbihOHj+Gth064a3XxyvFiP0A0ux79ukvwkNHjES58hXwzZdfiGOeyfYbNARubu44f/aUGNiUzHz/+/QfDJdSpfDbr3sJjx/gU7Mmxr86WZjLfOvUx0cLl2LOO1OVLHJfBAgUjPDdSLu3zSL74PgMI218YHUNfD2siOy1GHY4Ff73s9IzkS//Jw1butqitptO+3+vuTURPPnrk7lz520iM2VSUARg5KWKl0aNxY87vzcge87Hg0A/6hxdu/dEyZIumDbpFTEI7N/3izD3WIrwT5KmyCRSqXIVTH1nFhZ9PI/s/jXwQr+BIv6HHVvFZVQnrcytTFnMmvuR6OQP6LVrblfbDs9jzf82Y8LIYTle7oxZH6BLj14iX1R0FFwoH3fuHT/tx4v9eiIxUfcmtLd3RTCxsvbHkpqqGww4Lc9uWFJogOB6mzRrie92/owBvTqL+ML+iYuLwzOkSCjiXtoJIfdjEBwRC99KHgi5F4O4B8moV9kTt0IiEBIRhZj4JFR20CI83PAb4koZOe2ZuFrSQMzmDr43k6a+TYv4Plj62QIaBK1Rt34DBN65jVs3r6ERDcJtaGA+/vsR8Ro837cLfmdF0UzSPCjcj7wHxxIl6P7ORkkXF2zbvFHgX6VqNZqhpCGMFp4bN20htq7tWuhJXGmfq1sZ1KlXX2yslfNzwAvzHp5eeOetN0SyL9duQuWqVcW6gqOzExrTOoMp0VjbwJqugYWvRWOtC3fp3gszZn8gngV+tvoNHELX/AymvjGBBv8BeHP6THEdSUlJeHXyVDRr2RqbaRpsnZmfJsh0/QWiH1PNlXEPQSDfdDqnhTUWdbCBFRG+sgWrXqpS6u1STSPOzz+bZkD2yvlY4oih+1PBebmcUk7Am410D9XJsAzwVlykXoOGoiknSBtSy6uT38LmHT9hIBE+m3riiWxuXLsqkvz1x+9iz548eZXe/Qbhk8WfG2xTZ8zWZ585bYro5NzpyhPpzvt0sSCao4cP4u+LF/TpOBAWGoxu7VuhT7cO1NkniXy8BtGwURODdMpBe9JGmeyZvHkGMbh3V/Tr/jzuRYTDiTRRNmdll2tXr2Bwn+7o3qE1HB1LYO6Cz0SSzxd/ih4dW2NI3x6IIdNGKbJnsyZcFBJBNnjH9Fh9VfEPUhB0LxaX79xFREw8nBzIlkhyL+4BUtPpGSaiLlO6BO74nyNSNZ6l6gsyEXBxKYVDB/ZhxOC+GD6ojyD+Dp0NP0vL2Xbv2im+ZNhnwCBRSvdefUW9Wzd9i8pVqgqyZ9MPY9m7czskJMTjxeEj9TXyrGziqGEYNWwAViz9TMyo2j9vXI+S4cihA6I9vTo9h5DgICjPL2v3lapUgd/5M+K54Ht888Y1JZvBfvuWb7Fi2SIRt3nDOny1arkIv/n2u2CvJ24n399LF/xQ/9mGdI9LgWcF7CHF1/HykH7g55IHK/Zee2PCKPEM+v9zCTOnTjao62k84Fn4q5PegruHzpOMMeDwaJrR8TlLS74If1FHDcY0oiHajkwt6s3GmJxdSlAcpTkQQGlzECb95RdJm88sq7NPDgkfcTRr0izDR4+hh1v3l5dYs+9PC7QH9v4itHrnkiXBWn18fJzQeEaMGSfysMaXV6lYqTKatmhtsHXs3FWf/ca1/7Dv592CLNas3wSvchWEpvbx3Pf0aZTAN6tX6c0+586cFHZYPtel5wtKEoP98127i+MfaA3h9q0bIswa/Sbq7CxsUsoukyeOEdocxzcnTY5NO7zw/MuPO0VSDu/f85MIsyZcFHKfyMiBPrZma6NTHvi9i/R03TNoR3FJKbqFodthUXAnLUNLj6k3LSj5/XXE7Oa9+do4HCGz1uwPF2Dd5u1Co7a3dzAqh81v/v/8jbr1Gohz3Xv1Fhoy27Jbt+0g4sqX98bilavFxhFsIlSE78Od27fE4ekTf4q9j88zWEiLwgeOnRLbnkPHleS4fpX+bgAJk++R3/bDjoieZ35duvUUz86+n3T3hNPcvnWTd2BzkFIW71+bMk3Eq384DZdFY6S+rR6Zrq81a9XBn8eOimdg157fhILA3mo8KEgxRoAHvtiYWAwc8pIgeiZ7DqckJdOzcsk4QwFjbMzNP+ZZKwyqR3daZbdXynAR9nzlKHNPg4CVLdnvk7PFZzs8cCsDizPNQ2zP93UH/C3jKJGtpoIfij+RmFnMy6PHixCTfBcy57Dwgi0Lk786rYjMw893G9dj3y+7DVKyaUQtyxZ9TNP89ijtyguTEB4g3LEfJmdPnUBvMiOUK1feZFIfsrmzHD9mSHz79/6MKdPeETZ7JnR1XampOvLkfM1bteGdID01+Wg0/MzotBcRKOQfbqO/3wlUfLYvboZG4m50PBr5VEBEdIIgu2gy3yjiTzb9wR3q49BPWxEQGCzISjmXl/2yVV+R2eZZBJDZhjVlJycnsreb/nu732/ZSOauJcKsw+R77OghUQWbx1icyLxiZ6+bfcTGxCD6Pv1pzofIVnpeTv+lI/qUTLNa9iz3aNBlqVatGqpUqy7CprT6GHofYc0Xn4vz/HOBtPLS5KWjFiW/DZlkPD3LiVM0+RAzm5TkZGHj5xnFACKurrSG0I2Ui2NHDom1JXU5Mgxa00rGjm2biOSHC6JnTHgA4Dg+Z2kxi/Bd7Mnc0obuLC1smRJf4hBOoyb3A7e1aFHViuz4ROC5mEZFHlW5LsYKkqkqizTu5o3ror5v16+lKex5oX21pMWpuLhYLFn5pTi3k8he8dzhhVz23lmysjG5ZpqeMpu6gOioyIe6v/H0Pi4uRhA+hyPI5JIX4Y7IYu9AN8qE2NEiLIudrY50lCRpROpcD5s+WHtV7PjKeWXvVkZHdGlpqWTailWi9fs7RIpFJf4XzqJnMzJXUIVhkXFILxOIoU08UaVmLRy7dAupmaabiJgE2Gm0cCvlguFjX8Gvu7eTW+zDiZavg00YTPZs0mD7NcuiFf+XI+GzCy93ZDbR8aC0iZ4lFsV1lwfW9V/rniXGmreHid/5s2Se0a0DcFqfmrVFFnXelq3aivt3lIjXgdYHWDp06kx13RBhxU7Pi/L8DKulMa2/qOXqFd1CMT9Lr44ZoT9lTfZ+nsXwQrR4Z2T4IPGs/N83G8XakQvhy2QmxRABNenzmcIiey7bLMIvRSRsRTxgpXLD5ELUMqg+sPZMVswOejYGPUvae2/gxc00GGQpVlmJKFTH07DcUll/M8Ug3aM8+Puin6i+ZZu2gvB5zwuSbLPnTsKavUL2Sjt5AZNF8eJR4gu6HzR0BNj0o5Dwp0tXCnvtw8qtVKWqSHI3NNRkUtYE2SbdlbSyC5meR5yQNXcmKCb6nMie07GW27hpc4SH3cXIof056pFJeEQk7pw7jAbN+8PdJgn3bpwXBJ10MwSt61bB0Qs6sqtbzQuBQaHIKF0JTulk16e/GsQLlHkRJjDusNV9fNCt1wuoVbsOrY/obK9MgKbkzMm/hIbPNnDFbLaPTF7jXn0Dg4eNEHVfpun8xNenQGOlyTeOvJgeFBggBqRGTZsh/G6Y8LI6QF5Er0+ZDn6G+PlJoEXX59p3NNVUERcceEfs27RrT15gF4RZKSggADVohvLevI/JvPgTzRoHkDmvFQa90B0t6FnpTQu3NmQ6Cwi4QzNdZzEjVExqybSQy+tIvLB8+e9LOdb7NJ3gZ2jzxrWFfslm2fADyVutx7oM+Eeylm96m9KRvm3vQOcyhQl+wk7SRunbOsenEPH3ycCb7TLw9ZCsNJy0Sy3D8hb1MyxHKe9R7zf/by06d+shvHJ48Za9ZHixbvWKpUZkz/Z9NvPwy1iWFLahjp34miiSF754QdSLTDRjJ75uVE3jZs31cayZtydXS5azZ07p49WBsyf+EoeKLV451zdzoZEJJDf5k7xPWCpUrCjapKR1poFxIQ1K3M6iEtZwf9u7G3cOb8D6Re8imRZnA2iWdnrPRjLdBOqbkZqYiL2bvsDWz+dgzZKP80z2XAAT5ro1q4Qny/SZc9DzhX6CWPkcuyEqwukU+Xb9NyJ45NB+JYpcXxPw3oxpwquJiXo+ef6UpT/juOSz+fo06oBiUtOqylWf53B6erpYQB84ZBi1U4sln+rKYvfchfPnigF8+KhxYmBJStK9r5C9DD5mTx/W2Js2b4X5ny0TSaZPeRUhQUHgRf6Fy1YKkt9EfYNndTOnTUI0mYbY/v/pkhVkorIn9+Cl+jeNj/x2QLzdvWyVZfuFqbbLOEMErMp6lMt6Eg3P5Xo0qJEWbxG5e+tMyAZp9/sD478z1pDqUFW+tAVFWeHEraypaikaII5P0yK7Vr/DzwpTd5o1Jhm0Q30QG32fphCZFvXMabIy5RWxojm6NulOZ7WPy9Gd1sXxFJUjPpz9jp7k2SWNbfaKJs9k/+Eni4QNf+Lol9RNyTH81cYtZGOtIRY82TUvu7BZaPEnH2Htt98Llzr2snh7Mq3mk8fNohWrBfmMHjZQaHXsI8+eESzXaZE3kDStZtRh2dOGX/Fnrx0Wdp/kWcjWTRuE7ZUx+XHfYZGONdBzZ06gKi1Ws0bGnjtjR7xIHV1Hlr8e+Uu4ZXZqo9NoRYH0s3LNetSuUxdsBjpN2iy739Vv2ETY/3kmxIOjOVJ24F6RPGHvILH38vbGrcwFSXPK4bRMvM82qA+v5n1x5wFNV0laVHbC/vULERp6Vxzn94d91Jkc1eSen7L4zVYmaH5pKz/CJp3Vazfi69Ur6QW+rbQY6Km/Z0p5z9SqjWv/XUEFWjsIJTfPLbt+FjO7bh1aKUmM9uzWyc4IPDgpwl5Z7H3F151dStBahoODo35BX32ezTs2NnYmz6nTybB5CFSlNbgwGohZnHpsF/t7O3qIPf+YnnPqT+cc2H5eg+3ngYquGVgyOB0tdNwiMnRtQHbr1HRM225I+pdDrcBbdtk6MR10/43E15vHIssQvlHhBYhg0n171vuiUx3Yv5c+rfA7Ro6ZIEpkm31LIlD2xw8PC8NnCz7Me02Z2hp/24a37ML+1OzvzP7TTKYfzp4pkrD99hxp7GxK+XjxcuEKp+RlzYwHI55+szCJT5o4SoT5R1EQFY2RyWrki/3I+2KNGDCe79JdEBi7Zc7/YJYRcegLUgXefG08vfjzMXmetBeDCZ/iwWLHtu9I09NpiKrkRRrkAS2V1hfsbLOezYzEGCKyLBLLb4NCyQ/fEhIVRcqJhYRt8soArRTJprnlq78RM8M/aEbWgDzO+NniZyg3YZNQdsnNxMcDg3pwUOeVtnw1GkUXzreGr25iRbcM/Do9FS7Z7O6B9NxO22KDE9dNkzbnWzI0DS1rGE4yTl63QuB9Kyzbby326rryG7akhq+0IbePp7EZx1IvWyn1mbNXNPz578/C70d+I/NCLfKlp2/0mJg55FQu26Br1iZNkLRp9Ru2OaU3Fc/+5fzyTfa1DVNpc4qzpIbPdXh5eaL7iMnwj7GFe0k7aG8cx94fd+RU/WMXz5ivXrcJXyxbpH9TN/tF+Nathxmz5sKznJd4mevcmdNYMHd2vu9z9vLl8aNB4GEavkUIny+tjrcWv85MNXmVl4PIhHON3ORoz0Rep4IWrL0PbqE1Sj9tkw2+P5mlfRklyGdEYRC+0hTd55F9xCGbXRSzjnL+UezVhM8vvjzOYmnCZyxK0ecGupG9PeT2NXox6HDWG6CPM1Cy7U89Ag8j/HybdLIjG0trPvymrCmpWzUDdasak3v2tHO/Lxyyz16PpY+Z4IsDyVv6up7k8mLI9XXbBt3iqf51/yf5guW1SQQIAYsRvrc7kb3Kj95cdKeutcP3f1qsOeZW/8SlZ/c5WxtbhIUZL6Y9cRcrL0giIBHIEwIWY1haC8vVPz+31sQ+gCT73ADKx7nPF32Sj1wyi0RAIvAkI2Axwg+8R4yv0vCDIqywZIc9Tly2Rp0q6bTpTDr8bZ1xPQxt/fsvWKwZT/K9ktcmEZAISAQKhIDFmDYwQoOlu2xx+bY1gsI1+OdW1sIrn/tV9fatf4AGyyZlvXLLA4MUiYBEQCIgEShcBCxG+NzMJdvy9gGcbUfscIS0+l6tU3ErVINAGiCkSAQkAhIBiUDhImBRwjenqeFRGqz7RWr25mD2tKbVJkVB4+CKDJsSsKJPHvM7AY70Fmei6o3PpxUbed0SAQUB7hPK+zLcV1i476jlkRG+uhEyLBHIDYG06Buw82oCa+cK0EZfw316gawMfU/ItoLu0wi55ZXnJAJPCwJM9tw3WLivsHDfUYskfDUaMlwsEUi8uU8Qvm21Pkg+v1h86iHSzD9DWCwvTDZKIlBICHBfYeG+oxZpPFejIcPFEoHUkBNICTsLG6/mcGgyE5rS9F2gzClrsWywbJRE4FEgQH2C+wb3Ee4rycEnwH1HLVLDV6Mhw8UWgfiLX8PVvT6sPRrDkTYpEgGJQM4IZKQlIuHCaqMEFvuWjlHJxSyiML+lU8wu9Ylujn3ljrAv3xLWbrVg7Wj6zwg+0QDIi5MI5IBAemIk0u9fQTJp9cl3DptMJQmfYDH3e/gmkZSREgGJgESgmCMgbfjF/AbJ5kkEJAISAUshIAnfUkjKciQCEgGJQDFHQBJ+Mb9BsnkSAYmARMBSCEjCtxSSshyJgERAIlDMEZCEX8xvkGyeREAiIBGwFAL/D7oLbknOPmAtAAAAAElFTkSuQmCC) With your PostgreSQL service selected, the **Query** section will show the metrics from the database in its dropdown. tip If no metrics are shown, check that there is data in the database Once you are happy with your panel, give it a title and click "Save" in the top right hand corner. ![Screenshot of a Grafana panel](/docs/assets/images/view-data-postgresql-grafana-2cee1fb1a44f313bc1b5fcda9bd1f6c3.png) tip Grafana expects a Time column in the query, if you don't have any, use the function `NOW()` which generates the up-to-date timestamp. To get to know Grafana better, try the [Grafana Fundamentals](https://grafana.com/tutorials/grafana-fundamentals/?pg=docs) page on the Grafana project site. --- # Maintenance and lifecycle in Aiven for PostgreSQL® Keep your Aiven for PostgreSQL® service current by upgrading versions and applying maintenance updates. Related pages * [Upgrade the PostgreSQL version](/docs/products/postgresql/howto/upgrade.md) --- # Advanced parameters for Aiven for PostgreSQL® On creating a PostgreSQL® database, you can customize it using a series of the following advanced parameters: | Parameter | | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | []()[**additional\_backup\_regions**](#additional_backup_regions)`array`Additional Cloud Regions for Backup Replication | | []()[**migration**](#migration)`object,null`Migrate data from existing servermigration.host string Hostname or IP address of the server where to migrate data from migration.port integer - min: 1 - max: 65535 Port number of the server where to migrate data from migration.password string Password for authentication with the server where to migrate data from migration.ssl boolean - default: true The server where to migrate data from is secured with SSL migration.username string User name for authentication with the server where to migrate data from migration.dbname string Database name for bootstrapping the initial connection migration.ignore\_dbs string Comma-separated list of databases, which should be ignored during migration (supported by MySQL and PostgreSQL only at the moment) migration.ignore\_roles string Comma-separated list of database roles, which should be ignored during migration (supported by PostgreSQL only at the moment) migration.method string The migration method to be used (currently supported only by Redis, Dragonfly, MySQL and PostgreSQL service types) | | []()[**ip\_filter**](#ip_filter)`array`- default: `0.0.0.0/0,::/0`IP filterAllow incoming connections from CIDR address block, e.g. '10.20.0.0/16' | | []()[**static\_ips**](#static_ips)`boolean`Use static public IP addresses | | []()[**service\_log**](#service_log)`boolean,null`Service loggingStore logs for the service so that they are available in the HTTP API and console. | | []()[**enable\_ipv6**](#enable_ipv6)`boolean`Enable IPv6Register AAAA DNS records for the service, and allow IPv6 packets to service ports | | []()[**admin\_username**](#admin_username)`string,null`Custom username for admin user. This must be set only when a new service is being created. | | []()[**admin\_password**](#admin_password)`string,null`Custom password for admin user. Defaults to random string. This must be set only when a new service is being created. | | []()[**backup\_hour**](#backup_hour)`integer,null`- max: `23`The hour of day (in UTC) when backup for the service is started. New backup is only started if previous backup has already completed. | | []()[**backup\_minute**](#backup_minute)`integer,null`- max: `59`The minute of an hour when backup for the service is started. New backup is only started if previous backup has already completed. | | []()[**backup\_interval\_hours**](#backup_interval_hours)`integer,null`- min: `3`
- max: `24`Interval in hours between automatic backups. Minimum value is 3 hours. Must be a divisor of 24 (3, 4, 6, 8, 12, 24). (Applicable to ACU plans only) | | []()[**backup\_retention\_days**](#backup_retention_days)`integer,null`- min: `1`
- max: `30`Backup retention in daysNumber of days to retain automatic backups. Backups older than this value will be automatically deleted. (Applicable to ACU plans only) | | []()[**enable\_ha\_replica\_dns**](#enable_ha_replica_dns)`boolean`Enable HA replica DNSCreates a dedicated read-only DNS that automatically falls back to the primary if standby nodes are unavailable. It switches back when a standby recovers. | | []()[**node\_count**](#node_count)`integer`- min: `1`
- max: `100`Number of nodes for the service | | []()[**pgaudit**](#pgaudit)`object`System-wide settings for the pgaudit extension.pgaudit.feature\_enabled boolean Enable pgaudit extension. When enabled, pgaudit extension will be automatically installed.Otherwise, extension will be uninstalled but auditing configurations will be preserved. pgaudit.log array Specifies which classes of statements will be logged by session audit logging. pgaudit.log\_catalog boolean - default: true Log Catalog Specifies that session logging should be enabled in the case where all relations in a statement are in pg\_catalog. pgaudit.log\_client boolean Specifies whether log messages will be visible to a client process such as psql. pgaudit.log\_level string - default: log Specifies the log level that will be used for log entries. pgaudit.log\_max\_string\_length integer - min: -1 - max: 102400 - default: -1 Log Max String Length Crop parameters representation and whole statements if they exceed this threshold. A (default) value of -1 disable the truncation. pgaudit.log\_nested\_statements boolean - default: true This GUC allows to turn off logging nested statements, that is, statements that are executed as part of another ExecutorRun. pgaudit.log\_parameter boolean Log Parameter Specifies that audit logging should include the parameters that were passed with the statement. pgaudit.log\_parameter\_max\_size integer Log Parameter Max Size Specifies that parameter values longer than this setting (in bytes) should not be logged, but replaced with \. pgaudit.log\_relation boolean Specifies whether session audit logging should create a separate log entry for each relation (TABLE, VIEW, etc.) referenced in a SELECT or DML statement. pgaudit.log\_rows boolean Log Rows pgaudit.log\_statement boolean - default: true Specifies whether logging will include the statement text and parameters (if enabled). pgaudit.log\_statement\_once boolean Log Statement Once Specifies whether logging will include the statement text and parameters with the first log entry for a statement/substatement combination or with every entry. pgaudit.role string Specifies the master role to use for object audit logging. | | []()[**pglookout**](#pglookout)`object`- default: `[object Object]`PGLookout settingsSystem-wide settings for pglookout.pglookout.max\_failover\_replication\_time\_lag integer - min: 10 - max: 9223372036854776000 - default: 60 Max Failover Replication Time Lag Number of seconds of master unavailability before triggering database failover to standby | | []()[**pg\_service\_to\_fork\_from**](#pg_service_to_fork_from)`string,null`Name of the PG Service from which to fork (deprecated, use service\_to\_fork\_from). This has effect only when a new service is being created. | | []()[**service\_to\_fork\_from**](#service_to_fork_from)`string,null`Name of another service to fork from. This has effect only when a new service is being created. | | []()[**project\_to\_fork\_from**](#project_to_fork_from)`string,null`Name of another project to fork a service from. This has effect only when a new service is being created. | | []()[**synchronous\_replication**](#synchronous_replication)`string`Synchronous replication type. (deprecated, use synchronous\_commit instead). Note that the service plan also needs to support synchronous replication.This setting is deprecated. Use synchronous\_commit instead. Any change to this setting will automatically update synchronous\_commit. Setting the value to quorum changes synchronous\_commit to remote\_write, while setting it to off changes synchronous\_commit to off. | | []()[**pg\_read\_replica**](#pg_read_replica)`boolean,null`Should the service which is being forked be a read replica (deprecated, use read\_replica service integration instead).This setting is deprecated. Use read\_replica service integration instead. | | []()[**pg\_stat\_monitor\_enable**](#pg_stat_monitor_enable)`boolean`- Service restartEnable pg\_stat\_monitor extension if available for the current clusterEnable the pg\_stat\_monitor extension. Changing this parameter causes a service restart. When this extension is enabled, pg\_stat\_statements results for utility commands are unreliable | | []()[**pg\_stat\_plans\_enable**](#pg_stat_plans_enable)`boolean`- Service restartEnable pg\_stat\_plans extension if available for the current clusterEnable the pg\_stat\_plans extension. Changing this parameter causes a service restart. Tracks execution plans for SQL queries | | []()[**pg\_version**](#pg_version)`string,null`PostgreSQL major version | | []()[**pgbouncer**](#pgbouncer)`object`PGBouncer connection pooling settingsSystem-wide settings for pgbouncer.pgbouncer.autodb\_idle\_timeout integer - max: 86400 - default: 3600 If the automatically created database pools have been unused this many seconds, they are freed. If 0 then timeout is disabled. \[seconds] pgbouncer.autodb\_max\_db\_connections integer - max: 2147483647 Do not allow more than this many server connections per database (regardless of user). Setting it to 0 means unlimited. pgbouncer.autodb\_pool\_mode string - default: transaction PGBouncer pool mode pgbouncer.autodb\_pool\_size integer - max: 10000 If non-zero then create automatically a pool of that size per user when a pool doesn't exist. pgbouncer.ignore\_startup\_parameters array List of parameters to ignore when given in startup packet pgbouncer.max\_prepared\_statements integer - max: 3000 - default: 100 PgBouncer tracks protocol-level named prepared statements related commands sent by the client in transaction and statement pooling modes when max\_prepared\_statements is set to a non-zero value. Setting it to 0 disables prepared statements. max\_prepared\_statements defaults to 100, and its maximum is 3000. pgbouncer.min\_pool\_size integer - max: 10000 Add more server connections to pool if below this number. Improves behavior when usual load comes suddenly back after period of total inactivity. The value is effectively capped at the pool size. pgbouncer.server\_connect\_timeout number - min: 1 - max: 86400 If connection and login don’t finish in this amount of time, the connection will be closed. \[seconds] pgbouncer.server\_idle\_timeout integer - max: 86400 - default: 600 If a server connection has been idle more than this many seconds it will be dropped. If 0 then timeout is disabled. \[seconds] pgbouncer.server\_lifetime integer - min: 60 - max: 86400 - default: 3600 The pooler will close an unused server connection that has been connected longer than this. \[seconds] pgbouncer.server\_login\_retry number - min: 1 - max: 86400 If login to the server failed, because of failure to connect or from authentication, the pooler waits this much before retrying to connect. During the waiting interval, new clients trying to connect to the failing server will get an error immediately without another connection attempt. \[seconds] pgbouncer.server\_reset\_query\_always boolean Run server\_reset\_query (DISCARD ALL) in all pooling modes | | []()[**recovery\_target\_time**](#recovery_target_time)`string,null`Recovery target time when forking a service. This has effect only when a new service is being created. | | []()[**variant**](#variant)`string,null`Variant of the PostgreSQL service, may affect the features that are exposed by default | | []()[**private\_access**](#private_access)`object`Allow access to selected service ports from private networksprivate\_access.pg boolean Allow clients to connect to pg with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.pgbouncer boolean Allow clients to connect to pgbouncer with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.prometheus boolean Allow clients to connect to prometheus with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | []()[**privatelink\_access**](#privatelink_access)`object`Allow access to selected service components through Privatelinkprivatelink\_access.pg boolean Enable pg privatelink\_access.pgbouncer boolean Enable pgbouncer privatelink\_access.prometheus boolean Enable prometheus | | []()[**public\_access**](#public_access)`object`Allow access to selected service ports from the public Internetpublic\_access.pg boolean Allow clients to connect to pg from the public internet for service nodes that are in a project VPC or another type of private network public\_access.pgbouncer boolean Allow clients to connect to pgbouncer from the public internet for service nodes that are in a project VPC or another type of private network public\_access.prometheus boolean Allow clients to connect to prometheus from the public internet for service nodes that are in a project VPC or another type of private network | | []()[**pg**](#pg)`object`postgresql.conf configuration valuespg.autovacuum\_freeze\_max\_age integer - min: 200000000 - max: 1500000000 - Service restart autovacuum\_freeze\_max\_age Specifies the maximum age (in transactions) that a table's pg\_class.relfrozenxid field can attain before a VACUUM operation is forced to prevent transaction ID wraparound within the table. The system launches autovacuum processes to prevent wraparound even when autovacuum is otherwise disabled. Changing this parameter causes a service restart. pg.autovacuum\_max\_workers integer - min: 1 - max: 20 - Service restart autovacuum\_max\_workers Specifies the maximum number of autovacuum processes (other than the autovacuum launcher) that may be running at any one time. The default is 3. Changing this parameter causes a service restart. pg.autovacuum\_naptime integer - min: 1 - max: 86400 autovacuum\_naptime Specifies the minimum delay between autovacuum runs on any given database. The delay is measured in seconds. The default is 60. pg.autovacuum\_vacuum\_threshold integer - max: 2147483647 autovacuum\_vacuum\_threshold Specifies the minimum number of updated or deleted tuples needed to trigger a VACUUM in any one table. The default is 50. pg.autovacuum\_analyze\_threshold integer - max: 2147483647 autovacuum\_analyze\_threshold Specifies the minimum number of inserted, updated or deleted tuples needed to trigger an ANALYZE in any one table. The default is 50. pg.autovacuum\_vacuum\_scale\_factor number - max: 1 autovacuum\_vacuum\_scale\_factor Specifies a fraction of the table size to add to autovacuum\_vacuum\_threshold when deciding whether to trigger a VACUUM (e.g. 0.2 for 20% of the table size). The default is 0.2. pg.autovacuum\_analyze\_scale\_factor number - max: 1 autovacuum\_analyze\_scale\_factor Specifies a fraction of the table size to add to autovacuum\_analyze\_threshold when deciding whether to trigger an ANALYZE (e.g. 0.2 for 20% of the table size). The default is 0.2. pg.autovacuum\_vacuum\_cost\_delay integer - min: -1 - max: 100 autovacuum\_vacuum\_cost\_delay Specifies the cost delay value that will be used in automatic VACUUM operations. If -1 is specified, the regular vacuum\_cost\_delay value will be used. The default is 2 (upstream default). pg.autovacuum\_vacuum\_cost\_limit integer - min: -1 - max: 10000 autovacuum\_vacuum\_cost\_limit Specifies the cost limit value that will be used in automatic VACUUM operations. If -1 is specified, the regular vacuum\_cost\_limit value will be used. The default is -1 (upstream default). pg.bgwriter\_delay integer - min: 10 - max: 10000 bgwriter\_delay Specifies the delay between activity rounds for the background writer in milliseconds. The default is 200. pg.bgwriter\_flush\_after integer - max: 2048 Whenever more than bgwriter\_flush\_after bytes have been written by the background writer, attempt to force the OS to issue these writes to the underlying storage. Specified in kilobytes. Setting of 0 disables forced writeback. The default is 512. pg.bgwriter\_lru\_maxpages integer - max: 1073741823 bgwriter\_lru\_maxpages In each round, no more than this many buffers will be written by the background writer. Setting this to zero disables background writing. The default is 100. pg.bgwriter\_lru\_multiplier number - max: 10 The average recent need for new buffers is multiplied by bgwriter\_lru\_multiplier to arrive at an estimate of the number that will be needed during the next round, (up to bgwriter\_lru\_maxpages). 1.0 represents a “just in time” policy of writing exactly the number of buffers predicted to be needed. Larger values provide some cushion against spikes in demand, while smaller values intentionally leave writes to be done by server processes. The default is 2.0. pg.deadlock\_timeout integer - min: 500 - max: 1800000 deadlock\_timeout This is the amount of time, in milliseconds, to wait on a lock before checking to see if there is a deadlock condition. The default is 1000 (upstream default). pg.password\_encryption string,null password\_encryption Chooses the algorithm for encrypting passwords. pg.default\_toast\_compression string default\_toast\_compression Specifies the default TOAST compression method for values of compressible columns. The default is lz4. Only available for PostgreSQL 14+. pg.idle\_in\_transaction\_session\_timeout integer - max: 604800000 idle\_in\_transaction\_session\_timeout Time out sessions with open transactions after this number of milliseconds pg.io\_combine\_limit integer - min: 1 - max: 32 - default: 16 io\_combine\_limit EXPERIMENTAL: Controls the largest I/O size in operations that combine I/O in 8kB units. Version 17 and up only. pg.io\_max\_combine\_limit integer - min: 1 - max: 128 - default: 16 - Service restart io\_max\_combine\_limit EXPERIMENTAL: Controls the largest I/O size in operations that combine I/O in 8kB units, and silently limits the user-settable parameter io\_combine\_limit. Version 18 and up only. Changing this parameter causes a service restart. pg.io\_max\_concurrency integer - min: -1 - max: 1024 - default: -1 - Service restart io\_max\_concurrency EXPERIMENTAL: Controls the maximum number of I/O operations that one process can execute simultaneously. Version 18 and up only. Changing this parameter causes a service restart. pg.io\_method string - default: worker - Service restart io\_method EXPERIMENTAL: Controls the maximum number of I/O operations that one process can execute simultaneously. Version 18 and up only. Changing this parameter causes a service restart. pg.io\_workers integer - min: 1 - max: 32 io\_workers EXPERIMENTAL: Number of IO worker processes, for io\_method=worker. Version 18 and up only. pg.jit boolean Controls system-wide use of Just-in-Time Compilation (JIT). pg.log\_autovacuum\_min\_duration integer - min: -1 - max: 2147483647 log\_autovacuum\_min\_duration Causes each action executed by autovacuum to be logged if it ran for at least the specified number of milliseconds. Setting this to zero logs all autovacuum actions. Minus-one disables logging autovacuum actions. The default is 1000. pg.log\_error\_verbosity string log\_error\_verbosity Controls the amount of detail written in the server log for each message that is logged. pg.log\_line\_prefix string log\_line\_prefix Choose from one of the available log formats. pg.log\_min\_duration\_statement integer - min: -1 - max: 86400000 log\_min\_duration\_statement Log statements that take more than this number of milliseconds to run, -1 disables pg.log\_temp\_files integer - min: -1 - max: 2147483647 log\_temp\_files Log statements for each temporary file created larger than this number of kilobytes, -1 disables pg.max\_files\_per\_process integer - min: 1000 - max: 4096 - Service restart max\_files\_per\_process PostgreSQL maximum number of files that can be open per process. The default is 1000 (upstream default). Changing this parameter causes a service restart. pg.max\_prepared\_transactions integer - max: 10000 - Service restart max\_prepared\_transactions PostgreSQL maximum prepared transactions. The default is 0. Changing this parameter causes a service restart. pg.max\_pred\_locks\_per\_transaction integer - min: 64 - max: 5120 - Service restart max\_pred\_locks\_per\_transaction PostgreSQL maximum predicate locks per transaction. The default is 64 (upstream default). Changing this parameter causes a service restart. pg.max\_locks\_per\_transaction integer - min: 64 - max: 6400 - Service restart max\_locks\_per\_transaction PostgreSQL maximum locks per transaction. Changing this parameter causes a service restart. pg.max\_slot\_wal\_keep\_size integer - min: -1 - max: 2147483647 max\_slot\_wal\_keep\_size PostgreSQL maximum WAL size (MB) reserved for replication slots. If -1 is specified, replication slots may retain an unlimited amount of WAL files. The default is -1 (upstream default). wal\_keep\_size minimum WAL size setting takes precedence over this. pg.max\_stack\_depth integer - min: 2097152 - max: 6291456 max\_stack\_depth Maximum depth of the stack in bytes. The default is 2097152 (upstream default). pg.max\_standby\_archive\_delay integer - min: 1 - max: 43200000 max\_standby\_archive\_delay Max standby archive delay in milliseconds. The default is 30000 (upstream default). pg.max\_standby\_streaming\_delay integer - min: 1 - max: 43200000 max\_standby\_streaming\_delay Max standby streaming delay in milliseconds. The default is 30000 (upstream default). pg.max\_replication\_slots integer - min: 8 - max: 256 - Service restart max\_replication\_slots PostgreSQL maximum replication slots. The default is 20. Changing this parameter causes a service restart. pg.max\_logical\_replication\_workers integer - min: 4 - max: 256 - Service restart max\_logical\_replication\_workers PostgreSQL maximum logical replication workers (taken from the pool defined by max\_worker\_processes). The default is 4 (upstream default). Changing this parameter causes a service restart. pg.max\_parallel\_workers integer - max: 96 max\_parallel\_workers Sets the maximum number of workers that the system can support for parallel queries. The default is 8 (upstream default). pg.max\_parallel\_workers\_per\_gather integer - max: 96 max\_parallel\_workers\_per\_gather Sets the maximum number of workers that can be started by a single Gather or Gather Merge node. The default is 2 (upstream default). pg.max\_sync\_workers\_per\_subscription integer - min: 2 - max: 8 max\_sync\_workers\_per\_subscription Maximum number of synchronization workers per subscription. The default is 2. pg.max\_worker\_processes integer - min: 8 - max: 288 - Service restart max\_worker\_processes Sets the maximum number of background processes that the system can support. The default is 8. Changing this parameter causes a service restart. pg.pg\_partman\_bgw\.role string pg\_partman\_bgw\.role Controls which role to use for pg\_partman's scheduled background tasks. pg.pg\_partman\_bgw\.interval integer - min: 3600 - max: 604800 pg\_partman\_bgw\.interval Sets the time interval in seconds to run pg\_partman's scheduled tasks. The default is 3600. pg.pg\_stat\_monitor.pgsm\_max\_buckets integer - min: 1 - max: 10 - Service restart pg\_stat\_monitor.pgsm\_max\_buckets Sets the maximum number of buckets. Changing this parameter causes a service restart. Only available for PostgreSQL 13+. pg.pg\_stat\_monitor.pgsm\_enable\_query\_plan boolean pg\_stat\_monitor.pgsm\_enable\_query\_plan Enables or disables query plan monitoring. Only available for PostgreSQL 13+. pg.pg\_stat\_statements.track string pg\_stat\_statements.track Controls which statements are counted. Specify top to track top-level statements (those issued directly by clients), all to also track nested statements (such as statements invoked within functions), or none to disable statement statistics collection. The default is top. pg.pg\_stat\_plans.track string pg\_stat\_plans.track Controls which statements' plans are tracked. Specify top to track top-level statements (those issued directly by clients), all to also track nested statements (such as statements invoked within functions), or none to disable plan tracking. The default is top. pg.synchronous\_commit string synchronous\_commit Sets the current transaction's synchronization level. The default is off. This setting takes precedence over synchronous\_replication. pg.temp\_file\_limit integer - min: -1 - max: 2147483647 temp\_file\_limit PostgreSQL temporary file limit in KiB, -1 for unlimited pg.timezone string PostgreSQL service timezone pg.track\_activity\_query\_size integer - min: 1024 - max: 10240 - Service restart track\_activity\_query\_size Specifies the number of bytes reserved to track the currently executing command for each active session. Changing this parameter causes a service restart. pg.track\_commit\_timestamp string - Service restart track\_commit\_timestamp Record commit time of transactions. Changing this parameter causes a service restart. pg.track\_functions string track\_functions Enables tracking of function call counts and time used. pg.track\_io\_timing string track\_io\_timing Enables timing of database I/O calls. The default is off. When on, it will repeatedly query the operating system for the current time, which may cause significant overhead on some platforms. pg.max\_wal\_senders integer - min: 20 - max: 256 - Service restart max\_wal\_senders PostgreSQL maximum WAL senders. The default is 20. Changing this parameter causes a service restart. pg.wal\_sender\_timeout integer wal\_sender\_timeout Terminate replication connections that are inactive for longer than this amount of time, in milliseconds. Setting this value to zero disables the timeout. pg.wal\_writer\_delay integer - min: 10 - max: 200 wal\_writer\_delay WAL flush interval in milliseconds. The default is 200. Setting this parameter to a lower value may negatively impact performance. pg.max\_connections integer - min: 25 - max: 60000 - Service restart max\_connections Sets the PostgreSQL maximum number of concurrent connections to the database server. For services with a read replica, first increase the read replica's value. After the change is applied to the replica, you can increase the primary service's value. Changing this parameter causes a service restart. | | []()[**shared\_buffers\_percentage**](#shared_buffers_percentage)`number`- min: `20`
- max: `60`
- Service restartshared\_buffers\_percentagePercentage of total RAM that the database server uses for shared memory buffers. Valid range is 20-60 (float), which corresponds to 20% - 60%. This setting adjusts the shared\_buffers configuration value. Changing this parameter causes a service restart. | | []()[**switchover\_windows**](#switchover_windows)`array` | | []()[**timescaledb**](#timescaledb)`object`TimescaleDB extension configuration valuesSystem-wide settings for the timescaledb extensiontimescaledb.max\_background\_workers integer - min: 1 - max: 4096 - default: 16 - Service restart timescaledb.max\_background\_workers The number of background workers for timescaledb operations. You should configure this setting to the sum of your number of databases and the total number of concurrent background workers you want running at any given point in time. Changing this parameter causes a service restart. | | []()[**work\_mem**](#work_mem)`integer`- min: `1`
- max: `1024`work\_memSets the maximum amount of memory to be used by a query operation (such as a sort or hash table) before writing to temporary disk files, in MB. The default is 1MB + 0.075% of total RAM (up to 32MB). | --- # Keep-alive connections parameters PostgreSQL® keep-alive connection parameters are useful to manage Idle connections. The following is a reference to the default Aiven for PostgreSQL® parameters on the server side, and what keep-alive parameters can be used at the client side. ## Keep-alive server side parameters[​](#keep-alive-server-side-parameters "Direct link to Keep-alive server side parameters") Currently, the following default keep-alive timeouts are used on the [server-side](https://www.postgresql.org/docs/current/runtime-config-connection.html#RUNTIME-CONFIG-CONNECTION-SETTINGS): | Parameter (server) | Value | Description | | ------------------------- | ----- | ------------------------------------------------------------------------------------------------------------------------------------------- | | `tcp_keepalives_idle` | 180 | Specifies the amount of time with no network activity after which the operating system should send a TCP `keepalive` message to the client. | | `tcp_keepalives_count` | 6 | Specifies the number of TCP `keepalive` messages that can be lost before the server's connection to the client is considered dead. | | `tcp_keepalives_interval` | 10 | Specifies the amount of time after which a TCP `keepalive` message that has not been acknowledged by the client should be retransmitted. | ## Keep-alive client side parameters[​](#keep-alive-client-side-parameters "Direct link to Keep-alive client side parameters") The [client-side](https://www.postgresql.org/docs/current/libpq-connect.html#LIBPQ-KEEPALIVES) keep-alive parameters can be set to whatever values you want. | Parameter (client) | Description | | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | | `keepalives` | Controls whether client-side TCP `keepalives` are used. The default value is 1, meaning on. | | `keepalives_idle` | Controls the number of seconds of inactivity after which TCP should send a `keepalive` message to the server. | | `keepalive_count` | Controls the number of TCP `keepalives` that can be lost before the client's connection to the server is considered dead. | | `keepalives_interval` | Controls the number of seconds after which a TCP `keepalive` message that is not acknowledged by the server should be retransmitted. | Even though TCP connections usually stay open for extended periods of time, you should also make sure that your applications can reconnect, since TCP connections are liable to break at times. Also, when reconnecting you should make sure that your client always resolves the DNS address on connection, since the underlying address will change during automatic failover when a primary node fails. --- # Extensions on Aiven for PostgreSQL® PostgreSQL® extensions allow you to extend the functionality by adding more capabilities to your Aiven for PostgreSQL. important Some extensions: * Have dependencies and need to be created in a predetermined order. * Require resetting the client connection before they are fully available. ## How it works[​](#how-it-works "Direct link to How it works") ### List extensions[​](#list-extensions "Direct link to List extensions") To list the extensions and their details, such as extension version numbers, run the following command in your Aiven for PostgreSQL server: ``` SELECT * FROM pg_available_extensions; ``` note Superuser-only and untrusted extensions listed in `pg_available_extensions` cannot be installed with a few exceptions. * To list all superuser-only and untrusted extensions, run ``` SELECT name, version FROM pg_available_extension_versions WHERE superuser AND NOT trusted ``` * To list superuser-only and untrusted extensions **available for installation**, run ``` SELECT * FROM pg_available_extension_versions WHERE name = ANY(string_to_array(current_setting('extwlist.extensions'), ',')); ``` ### Install and manage extensions[​](#install-and-manage-extensions "Direct link to Install and manage extensions") See [Manage Aiven for PostgreSQL extensions](/docs/products/postgresql/howto/manage-extensions.md) for the instructions. ## Supported extensions[​](#supported-extensions "Direct link to Supported extensions") ### Auditing[​](#auditing "Direct link to Auditing") * [pgaudit](https://github.com/pgaudit/pgaudit) provides session and object audit logging often required for compliance with government, financial, or ISO certifications. * [tcn](https://www.postgresql.org/docs/current/tcn.html). Triggered change notifications. ### Connectivity[​](#connectivity "Direct link to Connectivity") * [dblink](https://www.postgresql.org/docs/current/contrib-dblink-function.html). Connect to other PostgreSQL databases from within a database. * [postgres\_fdw](https://www.postgresql.org/docs/current/postgres-fdw.html). Foreign-data wrapper for remote PostgreSQL servers. ### Data types[​](#data-types "Direct link to Data types") * [citext](https://www.postgresql.org/docs/current/citext.html). Data type for case-insensitive character strings. * [cube](https://www.postgresql.org/docs/current/cube.html). Data type for multidimensional cubes. * [hll](https://github.com/citusdata/postgresql-hll). Type for storing `hyperloglog` data. * [hstore](https://www.postgresql.org/docs/current/hstore.html). Data type for storing sets of (key, value) pairs. * [ip4r](https://github.com/RhodiumToad/ip4r). IPv4 and IPv6 address range data types with indexing support. * [isn](https://www.postgresql.org/docs/current/isn.html). Data types for international product numbering standards. * [ltree](https://www.postgresql.org/docs/current/ltree.html). Data type for hierarchical tree-like structures. * [seg](https://www.postgresql.org/docs/current/seg.html). Data type for representing line segments or floating-point intervals. * [timescaledb](https://github.com/timescale/timescaledb). Enables scalable inserts and complex queries for time-series data. * [unit](https://github.com/df7cb/postgresql-unit). SI units extension. * [uuid-ossp](https://www.postgresql.org/docs/current/uuid-ossp.html). Generate universally unique identifiers (UUIDs). ### Geographical features[​](#geographical-features "Direct link to Geographical features") * [address\_standardizer](https://postgis.net/docs/standardize_address.html). Used to parse an address into constituent elements. Generally used to support geocoding address normalization step. * [address\_standardizer\_data\_us](https://postgis.net/docs/standardize_address.html). `Address standardizer` US dataset example. * [earthdistance](https://www.postgresql.org/docs/current/earthdistance.html). Calculate great-circle distances on the surface of the Earth. * [h3](https://github.com/zachasme/h3-pg). PostgreSQL bindings for H3, a hierarchical hexagonal geospatial indexing system. * [h3\_postgis](https://github.com/zachasme/h3-pg). H3 PostGIS integration. * [pgrouting](https://github.com/pgRouting/pgrouting). Extends the PostGIS/PostgreSQL geospatial database to provide geospatial routing and other network analysis functionality. * [postgis](https://postgis.net/). PostGIS geometry and geography spatial types and functions. * [postgis\_legacy](https://postgis.net/). Legacy functions for PostGIS. * [postgis\_raster](https://postgis.net/docs/RT_reference.html). PostGIS raster types and functions. * [postgis\_sfcgal](https://postgis.net/docs/reference_sfcgal.html). PostGIS SFCGAL functions. * [postgis\_tiger\_geocoder](https://postgis.net/docs/Extras.html#Tiger_Geocoder). PostGIS tiger geocoder and reverse geocoder. * [postgis\_topology](https://postgis.net/docs/Topology.html). PostGIS topology spatial types and functions. ### Machine learning (ML) and artificial intelligence (AI)[​](#machine-learning-ml-and-artificial-intelligence-ai "Direct link to Machine learning (ML) and artificial intelligence (AI)") * [pgvector](https://github.com/pgvector/pgvector) designed for vector similarity search for PostgreSQL. * [pgvectorscale](https://github.com/timescale/pgvectorscale) complements [pgvector](https://github.com/pgvector/pgvector) as a vector data extension for PostgreSQL. `PG16 and newer` important Supported `pgvectorscale` versions: * PG16: pgvectorscale-0.6.0 * PG17: pgvectorscale-0.6.0 Read about all [pgvectorscale releases](https://github.com/timescale/pgvectorscale/releases). ### Procedural language[​](#procedural-language "Direct link to Procedural language") * [plperl](https://www.postgresql.org/docs/current/plperl.html). PL/Perl procedural language. * [plpgsql](https://www.postgresql.org/docs/current/plpgsql.html). PL/pgSQL procedural language. ### Search and text handling[​](#search-and-text-handling "Direct link to Search and text handling") * [bloom](https://www.postgresql.org/docs/current/bloom.html). Bloom access method - signature file based index. * [btree\_gin](https://www.postgresql.org/docs/current/btree-gin.html). Support for indexing common data types in GIN. * [btree\_gist](https://www.postgresql.org/docs/current/btree-gist.html). Support for indexing common data types in GiST. * [dict\_int](https://www.postgresql.org/docs/current/dict-int.html). Text search dictionary template for integers. * [fuzzystrmatch](https://www.postgresql.org/docs/current/fuzzystrmatch.html). Determine similarities and distance between strings. * [pg\_similarity](https://github.com/eulerto/pg_similarity). Support similarity queries. * [pg\_trgm](https://www.postgresql.org/docs/current/pgtrgm.html). Text similarity measurement and index searching based on trigrams. * [pgcrypto](https://www.postgresql.org/docs/current/pgcrypto.html). Cryptographic functions. * [rum](https://github.com/postgrespro/rum). RUM index access method. * [unaccent](https://www.postgresql.org/docs/current/unaccent.html). Text search dictionary that removes accents. ### Utilities[​](#utilities "Direct link to Utilities") * [aiven\_extras](https://github.com/aiven/aiven-extras). This extension is meant for use in environments where you want non-superusers to be able to use certain database features. * [bool\_plperl](https://www.postgresql.org/docs/current/plperl-funcs.html). Transform between `bool` and `plperl`. * [intagg](https://www.postgresql.org/docs/current/intagg.html). Integer aggregator and enumerator (obsolete). * [intarray](https://www.postgresql.org/docs/current/intarray.html). Functions, operators, and index support for 1-D arrays of integers. * [jsonb\_plperl](https://www.postgresql.org/docs/current/datatype-json.html). Transform between `jsonb` and `plperl`. * [lo](https://www.postgresql.org/docs/current/lo.html). Large Object maintenance. * [pageinspect](https://www.postgresql.org/docs/current/pageinspect.html). Inspect the contents of database pages at a low level. * [pg\_buffercache](https://www.postgresql.org/docs/current/pgbuffercache.html). Examine the shared buffer cache. * [pg\_cron](https://github.com/citusdata/pg_cron). Job scheduler for PostgreSQL. * [pg\_partman](https://github.com/pgpartman/pg_partman). Extension to manage partitioned tables by time or ID. * [pg\_prewarm](https://www.postgresql.org/docs/current/pgprewarm.html). Prewarm relation data. * [pg\_repack](https://pgxn.org/dist/pg_repack/1.4.6/). Reorganize tables in PostgreSQL databases with minimal locks. * [pg\_stat\_statements](https://www.postgresql.org/docs/current/pgstatstatements.html). Track planning and execution statistics of all SQL statements executed. * [pgrowlocks](https://www.postgresql.org/docs/current/pgrowlocks.html). Show row-level locking information. * [pgstattuple](https://www.postgresql.org/docs/current/pgstattuple.html). Show tuple-level statistics. * [postgresql\_anonymizer](https://postgresql-anonymizer.readthedocs.io/en/latest/). Mask or replace personally identifiable information (PII) or commercially sensitive data from a PostgreSQL database. `PG15 and newer` The user who [installs this extension](/docs/products/postgresql/howto/manage-extensions.md#install-an-extension) (preferably `avnadmin`) becomes its **owner**. * Only the **avnadmin** user can update extension settings, such as `anon.salt`, even if they are not the extension owner. * Only the extension **owner** can read internal tables and settings, even if they are not the `avnadmin` user. Use the `anon.current_setting(SETTING_NAME)` function to access extension settings. * [sslinfo](https://www.postgresql.org/docs/current/sslinfo.html). Information about SSL certificates. * [tablefunc](https://www.postgresql.org/docs/current/tablefunc.html). Functions that manipulate whole tables, including `crosstab`. * [tsm\_system\_rows](https://www.postgresql.org/docs/current/tsm-system-rows.html). TABLESAMPLE method which accepts number of rows as a limit. * [tsm\_system\_time](https://www.postgresql.org/docs/current/tsm-system-time.html). TABLESAMPLE method which accepts time in milliseconds as a limit. --- # Extension versions per PostgreSQL release Extension availability and versions in Aiven for PostgreSQL® vary by PostgreSQL major version, with each extension having one default version per PostgreSQL release. See [Manage Aiven for PostgreSQL extensions](/docs/products/postgresql/howto/manage-extensions.md) for setup and configuration instructions. ## PostgreSQL 19 extensions[​](#postgresql-19-extensions "Direct link to PostgreSQL 19 extensions") | Extension name | Default version | Supported versions | | -------------------- | --------------- | ---------------------------------------------------- | | bloom | 1.0 | 1.0 | | bool\_plperl | 1.0 | 1.0 | | btree\_gin | 1.4 | 1.0, 1.1, 1.2, 1.3, 1.4 | | btree\_gist | 1.9 | 1.9 | | citext | 1.8 | 1.4, 1.5, 1.6, 1.7, 1.8 | | cube | 1.5 | 1.2, 1.3, 1.4, 1.5 | | dblink | 1.2 | 1.2 | | dict\_int | 1.0 | 1.0 | | earthdistance | 1.2 | 1.1, 1.2 | | fuzzystrmatch | 1.2 | 1.1, 1.2 | | hstore | 1.8 | 1.4, 1.5, 1.6, 1.7, 1.8 | | intagg | 1.1 | 1.1 | | intarray | 1.5 | 1.2, 1.3, 1.4, 1.5 | | isn | 1.3 | 1.1, 1.2, 1.3 | | jsonb\_plperl | 1.0 | 1.0 | | lo | 1.2 | 1.1, 1.2 | | ltree | 1.3 | 1.1, 1.2, 1.3 | | pg\_buffercache | 1.7 | 1.2, 1.3, 1.4, 1.5, 1.6, 1.7 | | pg\_prewarm | 1.2 | 1.1, 1.2 | | pg\_stat\_plans | 2.1 | 2.0, 2.1 | | pg\_stat\_statements | 1.13 | 1.10, 1.11, 1.12, 1.13, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9 | | pg\_trgm | 1.6 | 1.3, 1.4, 1.5, 1.6 | | pgcrypto | 1.4 | 1.3, 1.4 | | pgrowlocks | 1.2 | 1.2 | | pgstattuple | 1.5 | 1.4, 1.5 | | plperl | 1.0 | 1.0 | | plpgsql | 1.0 | 1.0 | | postgres\_fdw | 1.3 | 1.0, 1.1, 1.2, 1.3 | | seg | 1.4 | 1.1, 1.2, 1.3, 1.4 | | sslinfo | 1.2 | 1.2 | | tablefunc | 1.0 | 1.0 | | tcn | 1.0 | 1.0 | | tsm\_system\_rows | 1.0 | 1.0 | | tsm\_system\_time | 1.0 | 1.0 | | unaccent | 1.1 | 1.1 | | uuid-ossp | 1.1 | 1.1 | | vector | 0.8.6 | 0.8.6 | ## PostgreSQL 18 extensions[​](#postgresql-18-extensions "Direct link to PostgreSQL 18 extensions") | Extension name | Default version | Supported versions | | ------------------------------- | --------------- | ---------------------------------------------------------------- | | address\_standardizer | 3.6.4 | 3.5.7, 3.6.4 | | address\_standardizer\_data\_us | 3.6.4 | 3.5.7, 3.6.4 | | aiven\_extras | 1.1.22 | 1.1.22 | | anon | 3.1.3 | 3.1.3 | | bloom | 1.0 | 1.0 | | bool\_plperl | 1.0 | 1.0 | | btree\_gin | 1.3 | 1.0, 1.1, 1.2, 1.3 | | btree\_gist | 1.8 | 1.2, 1.3, 1.4, 1.5, 1.6, 1.7, 1.8 | | citext | 1.8 | 1.4, 1.5, 1.6, 1.7, 1.8 | | cube | 1.5 | 1.2, 1.3, 1.4, 1.5 | | dblink | 1.2 | 1.2 | | dict\_int | 1.0 | 1.0 | | earthdistance | 1.2 | 1.1, 1.2 | | fuzzystrmatch | 1.2 | 1.1, 1.2 | | h3 | 4.5.0 | 4.5.0 | | h3\_postgis | 4.5.0 | 4.5.0 | | hll | 2.20 | 2.10, 2.11, 2.12, 2.13, 2.14, 2.15, 2.16, 2.17, 2.18, 2.19, 2.20 | | hstore | 1.8 | 1.4, 1.5, 1.6, 1.7, 1.8 | | intagg | 1.1 | 1.1 | | intarray | 1.5 | 1.2, 1.3, 1.4, 1.5 | | ip4r | 2.4 | 2.4 | | isn | 1.3 | 1.1, 1.2, 1.3 | | jsonb\_plperl | 1.0 | 1.0 | | lo | 1.2 | 1.1, 1.2 | | ltree | 1.3 | 1.1, 1.2, 1.3 | | pg\_buffercache | 1.6 | 1.2, 1.3, 1.4, 1.5, 1.6 | | pg\_cron | 1.6 | 1.0, 1.1, 1.2, 1.3, 1.4, 1.4-1, 1.5, 1.6 | | pg\_partman | 5.2.4 | 5.2.4 | | pg\_prewarm | 1.2 | 1.1, 1.2 | | pg\_repack | 1.5.2 | 1.5.2 | | pg\_stat\_monitor | 2.3 | 2.0, 2.1, 2.2, 2.3 | | pg\_stat\_plans | 2.1 | 2.0, 2.1 | | pg\_stat\_statements | 1.12 | 1.10, 1.11, 1.12, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9 | | pg\_trgm | 1.6 | 1.3, 1.4, 1.5, 1.6 | | pgaudit | 18.0 | 18.0 | | pgcrypto | 1.4 | 1.3, 1.4 | | pgrouting | 4.0.1 | 4.0.1 | | pgrowlocks | 1.2 | 1.2 | | pgstattuple | 1.5 | 1.4, 1.5 | | plperl | 1.0 | 1.0 | | plpgsql | 1.0 | 1.0 | | postgis | 3.6.4 | 3.5.7, 3.6.4 | | postgis\_raster | 3.6.4 | 3.5.7, 3.6.4 | | postgis\_sfcgal | 3.6.4 | 3.5.7, 3.6.4 | | postgis\_tiger\_geocoder | 3.6.4 | 3.5.7, 3.6.4 | | postgis\_topology | 3.6.4 | 3.5.7, 3.6.4 | | postgres\_fdw | 1.2 | 1.0, 1.1, 1.2 | | rum | 1.3 | 1.3 | | seg | 1.4 | 1.1, 1.2, 1.3, 1.4 | | sslinfo | 1.2 | 1.2 | | tablefunc | 1.0 | 1.0 | | tcn | 1.0 | 1.0 | | timescaledb | 2.30.0 | 2.23.1, 2.24.0, 2.25.2, 2.26.4, 2.27.2, 2.28.3, 2.29.2, 2.30.0 | | tsm\_system\_rows | 1.0 | 1.0 | | tsm\_system\_time | 1.0 | 1.0 | | unaccent | 1.1 | 1.1 | | unit | 7 | 3, 4, 5, 6, 7 | | uuid-ossp | 1.1 | 1.1 | | vector | 0.8.6 | 0.8.6 | | vectorscale | 0.9.0 | 0.9.0 | ## PostgreSQL 17 extensions[​](#postgresql-17-extensions "Direct link to PostgreSQL 17 extensions") | Extension name | Default version | Supported versions | | ------------------------------- | --------------- | -------------------------------------------------------------------------------------------------------------- | | address\_standardizer | 3.6.4 | 3.3.10, 3.5.7, 3.6.4 | | address\_standardizer\_data\_us | 3.6.4 | 3.3.10, 3.5.7, 3.6.4 | | aiven\_extras | 1.1.22 | 1.1.22 | | anon | 3.1.3 | 3.1.3 | | bloom | 1.0 | 1.0 | | bool\_plperl | 1.0 | 1.0 | | btree\_gin | 1.3 | 1.0, 1.1, 1.2, 1.3 | | btree\_gist | 1.7 | 1.2, 1.3, 1.4, 1.5, 1.6, 1.7 | | citext | 1.6 | 1.4, 1.5, 1.6 | | cube | 1.5 | 1.2, 1.3, 1.4, 1.5 | | dblink | 1.2 | 1.2 | | dict\_int | 1.0 | 1.0 | | earthdistance | 1.2 | 1.1, 1.1\_aiven, 1.2 | | fuzzystrmatch | 1.2 | 1.1, 1.2 | | h3 | 4.5.0 | 4.5.0 | | h3\_postgis | 4.5.0 | 4.5.0 | | hll | 2.20 | 2.10, 2.11, 2.12, 2.13, 2.14, 2.15, 2.16, 2.17, 2.18, 2.19, 2.20 | | hstore | 1.8 | 1.4, 1.5, 1.6, 1.7, 1.8 | | intagg | 1.1 | 1.1 | | intarray | 1.5 | 1.2, 1.3, 1.4, 1.5 | | ip4r | 2.4 | 2.4 | | isn | 1.2 | 1.1, 1.2 | | jsonb\_plperl | 1.0 | 1.0 | | lo | 1.1 | 1.1 | | ltree | 1.3 | 1.1, 1.2, 1.3 | | pg\_buffercache | 1.5 | 1.2, 1.3, 1.4, 1.5 | | pg\_cron | 1.6 | 1.0, 1.1, 1.2, 1.3, 1.4, 1.4-1, 1.5, 1.6 | | pg\_partman | 4.7.4 | 4.7.4 | | pg\_prewarm | 1.2 | 1.1, 1.2 | | pg\_repack | 1.5.1 | 1.5.1 | | pg\_similarity | 1.0 | 1.0 | | pg\_stat\_monitor | 2.1 | 2.0, 2.1 | | pg\_stat\_statements | 1.11 | 1.10, 1.11, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9 | | pg\_trgm | 1.6 | 1.3, 1.4, 1.5, 1.6 | | pgaudit | 17.0 | 17.0 | | pgcrypto | 1.3 | 1.3 | | pgrouting | 4.0.1 | 3.1.3, 4.0.1 | | pgrowlocks | 1.2 | 1.2 | | pgstattuple | 1.5 | 1.4, 1.5 | | plperl | 1.0 | 1.0 | | plpgsql | 1.0 | 1.0 | | postgis | 3.6.4 | 3.3.10, 3.5.7, 3.6.4 | | postgis\_raster | 3.6.4 | 3.3.10, 3.5.7, 3.6.4 | | postgis\_sfcgal | 3.6.4 | 3.3.10, 3.5.7, 3.6.4 | | postgis\_tiger\_geocoder | 3.6.4 | 3.3.10, 3.5.7, 3.6.4 | | postgis\_topology | 3.6.4 | 3.3.10, 3.5.7, 3.6.4 | | postgres\_fdw | 1.1 | 1.0, 1.1 | | rum | 1.3 | 1.3 | | seg | 1.4 | 1.1, 1.2, 1.3, 1.4 | | sslinfo | 1.2 | 1.2 | | tablefunc | 1.0 | 1.0 | | tcn | 1.0 | 1.0 | | timescaledb | 2.30.0 | 2.17.2, 2.18.2, 2.19.3, 2.20.3, 2.21.4, 2.22.1, 2.23.1, 2.24.0, 2.25.2, 2.26.4, 2.27.2, 2.28.3, 2.29.2, 2.30.0 | | tsm\_system\_rows | 1.0 | 1.0 | | tsm\_system\_time | 1.0 | 1.0 | | unaccent | 1.1 | 1.1 | | unit | 7 | 3, 4, 5, 6, 7 | | uuid-ossp | 1.1 | 1.1 | | vector | 0.8.6 | 0.8.6 | | vectorscale | 0.9.0 | 0.6.0, 0.9.0 | ## PostgreSQL 16 extensions[​](#postgresql-16-extensions "Direct link to PostgreSQL 16 extensions") | Extension name | Default version | Supported versions | | ------------------------------- | --------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | | address\_standardizer | 3.5.7 | 3.3.10, 3.5.7 | | address\_standardizer\_data\_us | 3.5.7 | 3.3.10, 3.5.7 | | aiven\_extras | 1.1.22 | 1.1.22 | | anon | 3.1.3 | 3.1.3 | | bloom | 1.0 | 1.0 | | bool\_plperl | 1.0 | 1.0 | | btree\_gin | 1.3 | 1.0, 1.1, 1.2, 1.3 | | btree\_gist | 1.7 | 1.2, 1.3, 1.4, 1.5, 1.6, 1.7 | | citext | 1.6 | 1.4, 1.5, 1.6 | | cube | 1.5 | 1.2, 1.3, 1.4, 1.5 | | dblink | 1.2 | 1.2 | | dict\_int | 1.0 | 1.0 | | earthdistance | 1.2 | 1.1, 1.1\_aiven, 1.2 | | fuzzystrmatch | 1.2 | 1.1, 1.2 | | h3 | 4.5.0 | 4.5.0 | | h3\_postgis | 4.5.0 | 4.5.0 | | hll | 2.20 | 2.10, 2.11, 2.12, 2.13, 2.14, 2.15, 2.16, 2.17, 2.18, 2.19, 2.20 | | hstore | 1.8 | 1.4, 1.5, 1.6, 1.7, 1.8 | | intagg | 1.1 | 1.1 | | intarray | 1.5 | 1.2, 1.3, 1.4, 1.5 | | ip4r | 2.4 | 2.4 | | isn | 1.2 | 1.1, 1.2 | | jsonb\_plperl | 1.0 | 1.0 | | lo | 1.1 | 1.1 | | ltree | 1.2 | 1.1, 1.2 | | pg\_buffercache | 1.4 | 1.2, 1.3, 1.4 | | pg\_cron | 1.6 | 1.0, 1.1, 1.2, 1.3, 1.4, 1.4-1, 1.5, 1.6 | | pg\_partman | 4.7.4 | 4.7.4 | | pg\_prewarm | 1.2 | 1.1, 1.2 | | pg\_repack | 1.5.0 | 1.5.0 | | pg\_similarity | 1.0 | 1.0 | | pg\_stat\_monitor | 1.0 | 1.0 | | pg\_stat\_statements | 1.10 | 1.10, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9 | | pg\_trgm | 1.6 | 1.3, 1.4, 1.5, 1.6 | | pgaudit | 16.0 | 16.0 | | pgcrypto | 1.3 | 1.3 | | pgrouting | 4.0.1 | 3.1.3, 4.0.1 | | pgrowlocks | 1.2 | 1.2 | | pgstattuple | 1.5 | 1.4, 1.5 | | plperl | 1.0 | 1.0 | | plpgsql | 1.0 | 1.0 | | postgis | 3.5.7 | 3.3.10, 3.5.7 | | postgis\_raster | 3.5.7 | 3.3.10, 3.5.7 | | postgis\_sfcgal | 3.5.7 | 3.3.10, 3.5.7 | | postgis\_tiger\_geocoder | 3.5.7 | 3.3.10, 3.5.7 | | postgis\_topology | 3.5.7 | 3.3.10, 3.5.7 | | postgres\_fdw | 1.1 | 1.0, 1.1 | | rum | 1.3 | 1.3 | | seg | 1.4 | 1.1, 1.2, 1.3, 1.4 | | sslinfo | 1.2 | 1.2 | | tablefunc | 1.0 | 1.0 | | tcn | 1.0 | 1.0 | | timescaledb | 2.30.0 | 2.13.1, 2.14.2, 2.15.3, 2.16.1, 2.17.2, 2.18.2, 2.19.3, 2.20.3, 2.21.4, 2.22.1, 2.23.1, 2.24.0, 2.25.2, 2.26.4, 2.27.2, 2.28.3, 2.29.2, 2.30.0 | | tsm\_system\_rows | 1.0 | 1.0 | | tsm\_system\_time | 1.0 | 1.0 | | unaccent | 1.1 | 1.1 | | unit | 7 | 3, 4, 5, 6, 7 | | uuid-ossp | 1.1 | 1.1 | | vector | 0.8.6 | 0.8.6 | | vectorscale | 0.6.0 | 0.6.0 | ## PostgreSQL 15 extensions[​](#postgresql-15-extensions "Direct link to PostgreSQL 15 extensions") | Extension name | Default version | Supported versions | | ------------------------------- | --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | address\_standardizer | 3.3.10 | 3.2.10, 3.3.10 | | address\_standardizer\_data\_us | 3.3.10 | 3.2.10, 3.3.10 | | aiven\_extras | 1.1.22 | 1.1.22 | | anon | 3.1.3 | 3.1.3 | | bloom | 1.0 | 1.0 | | bool\_plperl | 1.0 | 1.0 | | btree\_gin | 1.3 | 1.0, 1.1, 1.2, 1.3 | | btree\_gist | 1.7 | 1.2, 1.3, 1.4, 1.5, 1.6, 1.7 | | citext | 1.6 | 1.4, 1.5, 1.6 | | cube | 1.5 | 1.2, 1.3, 1.4, 1.5 | | dblink | 1.2 | 1.2 | | dict\_int | 1.0 | 1.0 | | earthdistance | 1.2 | 1.1, 1.1\_aiven, 1.2 | | fuzzystrmatch | 1.1 | 1.1 | | h3 | 4.5.0 | 4.5.0 | | h3\_postgis | 4.5.0 | 4.5.0 | | hll | 2.20 | 2.10, 2.11, 2.12, 2.13, 2.14, 2.15, 2.16, 2.17, 2.18, 2.19, 2.20 | | hstore | 1.8 | 1.4, 1.5, 1.6, 1.7, 1.8 | | intagg | 1.1 | 1.1 | | intarray | 1.5 | 1.2, 1.3, 1.4, 1.5 | | ip4r | 2.4 | 2.4 | | isn | 1.2 | 1.1, 1.2 | | jsonb\_plperl | 1.0 | 1.0 | | lo | 1.1 | 1.1 | | ltree | 1.2 | 1.1, 1.2 | | pg\_buffercache | 1.3 | 1.2, 1.3 | | pg\_cron | 1.6 | 1.0, 1.1, 1.2, 1.3, 1.4, 1.4-1, 1.5, 1.6 | | pg\_partman | 4.7.4 | 4.7.4 | | pg\_prewarm | 1.2 | 1.1, 1.2 | | pg\_repack | 1.4.7 | 1.4.7 | | pg\_similarity | 1.0 | 1.0 | | pg\_stat\_monitor | 1.0 | 1.0 | | pg\_stat\_statements | 1.10 | 1.10, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9 | | pg\_trgm | 1.6 | 1.3, 1.4, 1.5, 1.6 | | pgaudit | 1.7 | 1.7 | | pgcrypto | 1.3 | 1.3 | | pgrouting | 4.0.1 | 3.1.3, 4.0.1 | | pgrowlocks | 1.2 | 1.2 | | pgstattuple | 1.5 | 1.4, 1.5 | | plperl | 1.0 | 1.0 | | plpgsql | 1.0 | 1.0 | | postgis | 3.3.10 | 3.2.10, 3.3.10 | | postgis\_legacy | 3.3 | 3.3 | | postgis\_raster | 3.3.10 | 3.2.10, 3.3.10 | | postgis\_sfcgal | 3.3.10 | 3.2.10, 3.3.10 | | postgis\_tiger\_geocoder | 3.3.10 | 3.2.10, 3.3.10 | | postgis\_topology | 3.3.10 | 3.2.10, 3.3.10 | | postgres\_fdw | 1.1 | 1.0, 1.1 | | rum | 1.3 | 1.3 | | seg | 1.4 | 1.1, 1.2, 1.3, 1.4 | | sslinfo | 1.2 | 1.2 | | tablefunc | 1.0 | 1.0 | | tcn | 1.0 | 1.0 | | timescaledb | 2.28.3 | 2.10.3, 2.11.2, 2.12.2, 2.13.1, 2.14.2, 2.15.3, 2.16.1, 2.17.2, 2.18.2, 2.19.3, 2.20.3, 2.21.4, 2.22.1, 2.23.1, 2.24.0, 2.25.2, 2.26.4, 2.27.2, 2.28.3, 2.9.3 | | tsm\_system\_rows | 1.0 | 1.0 | | tsm\_system\_time | 1.0 | 1.0 | | unaccent | 1.1 | 1.1 | | unit | 7 | 3, 4, 5, 6, 7 | | uuid-ossp | 1.1 | 1.1 | | vector | 0.8.6 | 0.8.6 | ## PostgreSQL 14 extensions[​](#postgresql-14-extensions "Direct link to PostgreSQL 14 extensions") | Extension name | Default version | Supported versions | | ------------------------------- | --------------- | ----------------------------------------------------------------------------------------------------------------- | | address\_standardizer | 3.3.10 | 3.1.13, 3.2.10, 3.3.10 | | address\_standardizer\_data\_us | 3.3.10 | 3.1.13, 3.2.10, 3.3.10 | | aiven\_extras | 1.1.22 | 1.1.22 | | bloom | 1.0 | 1.0 | | bool\_plperl | 1.0 | 1.0 | | btree\_gin | 1.3 | 1.0, 1.1, 1.2, 1.3 | | btree\_gist | 1.6 | 1.2, 1.3, 1.4, 1.5, 1.6 | | citext | 1.6 | 1.4, 1.5, 1.6 | | cube | 1.5 | 1.2, 1.3, 1.4, 1.5 | | dblink | 1.2 | 1.2 | | dict\_int | 1.0 | 1.0 | | earthdistance | 1.2 | 1.1, 1.1\_aiven, 1.2 | | fuzzystrmatch | 1.1 | 1.1 | | h3 | 4.5.0 | 4.5.0 | | h3\_postgis | 4.5.0 | 4.5.0 | | hll | 2.20 | 2.10, 2.11, 2.12, 2.13, 2.14, 2.15, 2.16, 2.17, 2.18, 2.19, 2.20 | | hstore | 1.8 | 1.4, 1.5, 1.6, 1.7, 1.8 | | intagg | 1.1 | 1.1 | | intarray | 1.5 | 1.2, 1.3, 1.4, 1.5 | | ip4r | 2.4 | 2.4 | | isn | 1.2 | 1.1, 1.2 | | jsonb\_plperl | 1.0 | 1.0 | | lo | 1.1 | 1.1 | | ltree | 1.2 | 1.1, 1.2 | | pg\_buffercache | 1.3 | 1.2, 1.3 | | pg\_cron | 1.6 | 1.0, 1.1, 1.2, 1.3, 1.4, 1.4-1, 1.5, 1.6 | | pg\_partman | 4.7.4 | 4.7.4 | | pg\_prewarm | 1.2 | 1.1, 1.2 | | pg\_repack | 1.4.7 | 1.4.7 | | pg\_similarity | 1.0 | 1.0 | | pg\_stat\_monitor | 1.0 | 1.0 | | pg\_stat\_statements | 1.9 | 1.4, 1.5, 1.6, 1.7, 1.8, 1.9 | | pg\_trgm | 1.6 | 1.3, 1.4, 1.5, 1.6 | | pgaudit | 1.6.1 | 1.6.1 | | pgcrypto | 1.3 | 1.3 | | pgrouting | 4.0.1 | 3.1.3, 4.0.1 | | pgrowlocks | 1.2 | 1.2 | | pgstattuple | 1.5 | 1.4, 1.5 | | plperl | 1.0 | 1.0 | | plpgsql | 1.0 | 1.0 | | postgis | 3.3.10 | 3.1.13, 3.2.10, 3.3.10 | | postgis\_legacy | 3.3 | 3.3 | | postgis\_raster | 3.3.10 | 3.1.13, 3.2.10, 3.3.10 | | postgis\_sfcgal | 3.3.10 | 3.1.13, 3.2.10, 3.3.10 | | postgis\_tiger\_geocoder | 3.3.10 | 3.1.13, 3.2.10, 3.3.10 | | postgis\_topology | 3.3.10 | 3.1.13, 3.2.10, 3.3.10 | | postgres\_fdw | 1.1 | 1.0, 1.1 | | rum | 1.3 | 1.3 | | seg | 1.4 | 1.1, 1.2, 1.3, 1.4 | | sslinfo | 1.2 | 1.2 | | tablefunc | 1.0 | 1.0 | | tcn | 1.0 | 1.0 | | timescaledb | 2.19.3 | 2.10.3, 2.11.2, 2.12.2, 2.13.1, 2.14.2, 2.15.3, 2.16.1, 2.17.2, 2.18.2, 2.19.3, 2.5.2, 2.6.1, 2.7.2, 2.8.1, 2.9.3 | | tsm\_system\_rows | 1.0 | 1.0 | | tsm\_system\_time | 1.0 | 1.0 | | unaccent | 1.1 | 1.1 | | unit | 7 | 3, 4, 5, 6, 7 | | uuid-ossp | 1.1 | 1.1 | | vector | 0.8.6 | 0.8.6 | ## PostgreSQL 13 extensions[​](#postgresql-13-extensions "Direct link to PostgreSQL 13 extensions") | Extension name | Default version | Supported versions | | ------------------------------- | --------------- | ------------------------------------------------------------------------------------------------------ | | address\_standardizer | 3.3.10 | 3.0.7, 3.1.13, 3.2.10, 3.3.10 | | address\_standardizer\_data\_us | 3.3.10 | 3.0.7, 3.1.13, 3.2.10, 3.3.10 | | aiven\_extras | 1.1.22 | 1.1.22 | | bloom | 1.0 | 1.0 | | bool\_plperl | 1.0 | 1.0 | | btree\_gin | 1.3 | 1.0, 1.1, 1.2, 1.3 | | btree\_gist | 1.5 | 1.2, 1.3, 1.4, 1.5 | | citext | 1.6 | 1.4, 1.5, 1.6 | | cube | 1.4 | 1.2, 1.3, 1.4 | | dblink | 1.2 | 1.2 | | dict\_int | 1.0 | 1.0 | | earthdistance | 1.1 | 1.1 | | fuzzystrmatch | 1.1 | 1.1 | | hll | 2.20 | 2.10, 2.11, 2.12, 2.13, 2.14, 2.15, 2.16, 2.17, 2.18, 2.19, 2.20 | | hstore | 1.7 | 1.4, 1.5, 1.6, 1.7 | | intagg | 1.1 | 1.1 | | intarray | 1.3 | 1.2, 1.3 | | ip4r | 2.4 | 2.4 | | isn | 1.2 | 1.1, 1.2 | | jsonb\_plperl | 1.0 | 1.0 | | lo | 1.1 | 1.1 | | ltree | 1.2 | 1.1, 1.2 | | pg\_buffercache | 1.3 | 1.2, 1.3 | | pg\_cron | 1.6 | 1.0, 1.1, 1.2, 1.3, 1.4, 1.4-1, 1.5, 1.6 | | pg\_partman | 4.7.4 | 4.7.4 | | pg\_prewarm | 1.2 | 1.1, 1.2 | | pg\_repack | 1.4.7 | 1.4.7 | | pg\_similarity | 1.0 | 1.0 | | pg\_stat\_monitor | 1.0 | 1.0 | | pg\_stat\_statements | 1.8 | 1.4, 1.5, 1.6, 1.7, 1.8 | | pg\_trgm | 1.5 | 1.3, 1.4, 1.5 | | pgaudit | 1.5.1 | 1.5.1 | | pgcrypto | 1.3 | 1.3 | | pgrouting | 4.0.1 | 3.1.3, 4.0.1 | | pgrowlocks | 1.2 | 1.2 | | pgstattuple | 1.5 | 1.4, 1.5 | | plperl | 1.0 | 1.0 | | plpgsql | 1.0 | 1.0 | | postgis | 3.3.10 | 3.0.7, 3.1.13, 3.2.10, 3.3.10 | | postgis\_legacy | 3.3 | 3.3 | | postgis\_raster | 3.3.10 | 3.0.7, 3.1.13, 3.2.10, 3.3.10 | | postgis\_sfcgal | 3.3.10 | 3.0.7, 3.1.13, 3.2.10, 3.3.10 | | postgis\_tiger\_geocoder | 3.3.10 | 3.0.7, 3.1.13, 3.2.10, 3.3.10 | | postgis\_topology | 3.3.10 | 3.0.7, 3.1.13, 3.2.10, 3.3.10 | | postgres\_fdw | 1.0 | 1.0 | | rum | 1.3 | 1.3 | | seg | 1.3 | 1.1, 1.2, 1.3 | | sslinfo | 1.2 | 1.2 | | tablefunc | 1.0 | 1.0 | | tcn | 1.0 | 1.0 | | timescaledb | 2.15.3 | 2.10.3, 2.11.2, 2.12.2, 2.13.1, 2.14.2, 2.15.3, 2.2.1, 2.3.1, 2.4.2, 2.5.2, 2.6.1, 2.7.2, 2.8.1, 2.9.3 | | tsm\_system\_rows | 1.0 | 1.0 | | tsm\_system\_time | 1.0 | 1.0 | | unaccent | 1.1 | 1.1 | | unit | 7 | 3, 4, 5, 6, 7 | | uuid-ossp | 1.1 | 1.1 | | vector | 0.8.6 | 0.8.6 | --- # Supported log formats Aiven for PostgreSQL® supports setting different log formats which are compatible with popular log analysis tools like `pgbadger` or `pganalyze`. You can customise this functionality by navigating to your PostgreSQL® service on the [Aiven Console](https://console.aiven.io/). From the sidebar on your service's page, select **Service settings**. On the **Service settings** page, go to the **Advanced configuration** section, and select **Configure** > **Add configuration options**. Next, you can select the `pg.log_line_prefix` parameter and a desired format based on a pre-fixed list. The supported log formats are available below with an example of the output: * `'pid=%p,user=%u,db=%d,app=%a,client=%h '` ``` [pg-user-test-1]2023-01-11T23:58:46.010530[postgresql-14][14-1] pid=625,user=postgres,db=defaultdb,app=[unknown],client=[local] LOG: connection authorized: user=postgres database=defaultdb application_name=aiven-pruned [pg-user-test-1]2023-01-11T23:58:46.019705[postgresql-14][15-1] pid=625,user=postgres,db=defaultdb,app=aiven-pruned,client=[local] LOG: disconnection: session time: 0:00:00.010 user=postgres database=defaultdb host=[local] ``` * `'%t [%p]: [%l-1] user=%u,db=%d,app=%a,client=%h '` ``` [pg-user-test-1]2023-01-11T23:59:46.592609[postgresql-14][16-1] 2023-01-11 23:59:46 GMT [949]: [2-1] user=postgres,db=defaultdb,app=[unknown],client=[local] LOG: connection authorized: user=postgres database=defaultdb application_name=aiven-pruned [pg-user-test-1]2023-01-11T23:59:46.602035[postgresql-14][17-1] 2023-01-11 23:59:46 GMT [949]: [3-1] user=postgres,db=defaultdb,app=aiven-pruned,client=[local] LOG: disconnection: session time: 0:00:00.010 user=postgres database=defaultdb host=[local] ``` * `'%m [%p] %q[user=%u,db=%d,app=%a] '` ``` [pg-user-test-1]2023-01-12T00:00:57.839867[postgresql-14][18-1] 2023-01-12 00:00:57.839 GMT [1323] [user=postgres,db=defaultdb,app=[unknown]] LOG: connection authorized: user=postgres database=defaultdb application_name=aiven-pruned [pg-user-test-1]2023-01-12T00:00:57.849223[postgresql-14][19-1] 2023-01-12 00:00:57.849 GMT [1323] [user=postgres,db=defaultdb,app=aiven-pruned] LOG: disconnection: session time: 0:00:00.010 user=postgres database=defaultdb host=[local] ``` After selecting one of the available log formats from the drop down menu, select **Save configuration** to have the change take effect. Once the setting has been enabled, you can go to the logs tab on your service page to check if the log format has been successfully changed. At the moment, the formats available are known to be compatible with majority of the log analysis tools. For additional information on how to check the service logs, you can visit our [access service logs](/docs/platform/howto/list-monitoring.md) documentation. --- # Connection limits per plan for Aiven for PostgreSQL® Find the default `max_connections` value for each Aiven for PostgreSQL® plan, and learn how to change it for your service. By default, Aiven for PostgreSQL® instances limit the number of allowed connections to make sure that the database is able to serve them all. ## `max_connections` defaults[​](#max_connections-defaults "Direct link to max_connections-defaults") Default values of the `max_connections` setting vary according to the service plan: | Plan | Max connections | | ------------------------------------- | --------------- | | Developer | 15 | | Free | 20 | | Hobbyist | 25 | | Startup/Business/Premium-4 | 100 | | Startup/Business/Premium-8 | 200 | | Startup/Business/Premium-16 | 400 | | Startup/Business/Premium-32 | 800 | | Startup/Business/Premium-64 and above | 1000 | note Aiven can utilize any number of the connections for managing the service. Aiven for PostgreSQL doesn't apply a fixed per-GiB connection formula (for example, 100 connections per GiB of RAM). The plan-based defaults in the preceding table are the only default values, and there's no separate `max_connection_limit` setting. `max_connections` is the only parameter that controls the total number of connections for your service. tip During a connection-exhaustion incident, use an AI assistant connected to [Aiven MCP](/docs/tools/mcp-server.md) to check current connection usage against the configured limit. For example: > Show the current connection count on `my-pg-service`, grouped by role and state, and compare the total with `max_connections`. ## Increase or decrease `max_connections`[​](#increase-or-decrease-max_connections "Direct link to increase-or-decrease-max_connections") To increase or decrease the number of allowed connections for your service, set the [`max_connections`](/docs/products/postgresql/reference/advanced-params.md#pg_max_connections) parameter, which accepts a value from 25 to 60000. note Changing `max_connections` causes a service restart. If your service has a read replica, increase the replica's value first. After that change is applied, increase the primary service's value. * Console * CLI * Terraform * API 1. Log in to [Aiven Console](https://console.aiven.io/), and go to your organization > project > Aiven for PostgreSQL service. 2. On the **Overview** page of your service, select **Service settings** from the sidebar. 3. On the **Service settings** page, go to the **Advanced configuration** section, and select **Configure**. 4. Select **Add configuration options**, add the `max_connections` parameter, set the value, and select **Save configuration**. Run the [service update](/docs/tools/cli/service-cli.md#avn-cli-service-update) command: ``` avn service update SERVICE_NAME -c pg.max_connections=VALUE ``` Set `max_connections` in the `pg_user_config.pg` block of the [`aiven_pg`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/pg) resource: ``` resource "aiven_pg" "example" { # ... pg_user_config { pg { max_connections = 1000 } } } ``` Call the [ServiceUpdate endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate): ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ -H 'Authorization: Bearer BEARER_TOKEN' \ -H 'content-type: application/json' \ --data '{ "user_config": { "pg": { "max_connections": 1000 } } }' ``` ## Use connection pooling[​](#use-connection-pooling "Direct link to Use connection pooling") When several clients or client threads are connecting to the database, Aiven recommends using [connection pooling](/docs/products/postgresql/concepts/pg-connection-pooling.md) to limit the number of actual backend connections. Connection pooling is available in all Aiven for PostgreSQL Startup, Business, and Premium plans, and can be [configured in the console](/docs/products/postgresql/howto/manage-pool.md). Related pages * [Advanced parameters for Aiven for PostgreSQL](/docs/products/postgresql/reference/advanced-params.md) * [Aiven for PostgreSQL connection pooling](/docs/products/postgresql/concepts/pg-connection-pooling.md) * [Manage a connection pool](/docs/products/postgresql/howto/manage-pool.md) * [Change the service plan](/docs/products/postgresql/howto/change-service-plan.md) --- # PostgreSQL® metrics exposed in Grafana® The metrics/dashboard integration in the Aiven console enables you to push PostgreSQL® metrics to an external endpoint like Datadog or to create an integration and a prebuilt dashboard in Aiven for Grafana®. For more information on enabling the integration, see [Monitor PostgreSQL® metrics with Grafana®](/docs/products/postgresql/howto/report-metrics-grafana.md). ## General info about default dashboards[​](#general-info-about-default-dashboards "Direct link to General info about default dashboards") A few key points about the default dashboards pre-created by Aiven in Grafana: 1. The PostgreSQL dashboards show all tables and indexes for all logical databases since Aiven cannot determine tables or indexes relevance. 2. Some metrics are gathered but not shown in the default dashboard, you can access all available metrics by creating new dashboards. 3. New dashboards can be created to show any metrics or use any filtering criteria. The default dashboard can be used as a template to make the process easier. warning When creating new dashboards, do not prefix the names with **"Aiven"** because they may be removed or replaced. The "Aiven" prefix is used to identify Aiven's system-managed dashboards. This also applies to the default dashboard, for which any direct editing to it can be lost. ## PostgreSQL metrics prebuilt dashboard[​](#postgresql-metrics-prebuilt-dashboard "Direct link to PostgreSQL metrics prebuilt dashboard") The PostgreSQL default dashboard is split into several sections under two main categories: Generic and PostgreSQL. **Generic** metrics are not specific to the type of service running on the node and mostly related to CPU, memory, disk, and network. **PostgreSQL** metrics are specific for the service. ## General metrics[​](#general-metrics "Direct link to General metrics") ### Overview[​](#overview "Direct link to Overview") This section shows a high-level overview of the service node health. Major issues with the service are often visible directly in this section. note In the Overview section, the figures for services with multiple nodes are averages of all nodes that belong to the service. For some metrics, such as disk space, this typically does not matter since it's equal across all the nodes. For other metrics, especially when related to load concentrated only on the primary node, high values can be dampened by the average. Node-specific values are shown in the **system metrics** section. ![Grafana Dashboard for PostgreSQL Overview Section](/docs/assets/images/metrics-dashboard-overview-bf7f2d60a90a0032ee9dbebb2b42b97a.png) The following metrics are shown: | Parameter Name | Parameter Definition | Additional Notes | | ------------------ | ----------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- | | `Uptime` | The time the service has been up and running. | | | `Load average` | The number of processes that would want to run. | If the `Load average` figure is higher than the number of CPUs on the nodes, the service can be **under-provisioned**. | | `Memory available` | Memory not allocated by running processes. | | | `Disk free` | Amount of unused disk space. | | ### System metrics[​](#system-metrics "Direct link to System metrics") This section shows a more detailed listing of various generic system-related metrics. ![Grafana Dashboard for PostgreSQL System Metrics Section](/docs/assets/images/metrics-dashboard-system-47b39c3b18671b67cea1356f6c40678e.png) The following metrics are shown: | Parameter Name | Parameter Definition | Additional Notes | | ------------------------------------- | ------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `CPU` | `System`, `user`, `iowait`, and `interrupt` request (IRQ) CPU usage. | A high `iowait` is an indication that the system is writing or reading too much data to or from disk. | | `Load average` | The number of processes that would want to run. | The `Load average` figure is higher than the number of CPUs on the nodes, the service might be **under-provisioned**. | | `Memory available` | The amount of memory not allocated by running processes. | | | `Memory unused` | The amount of memory not allocated by running processes or used for buffer caches. | | | `Context switches` | The number of switches from one process or thread to another. | | | `Interrupts` | The number of interrupts per second. | | | `Processes` | The number of processes that are actively doing something. | Processes that are mostly idle are not included. | | `Disk free` | The current amount of remaining disk space. | Aiven suggest to actively monitor this value and associate it with an alert. The database will stop working correctly if it runs out of disk space. | | `Disk i/o` | The number of bytes read and written per second on each of the nodes. | | | `Data disk usage` | The amount of disk space that is in use on the service's data disk. | | | `CPU iowait` | The percentage of CPU time spent waiting for the disk to become available for read and write operations | Aiven suggest to create an alert that is triggered when `iowait` goes beyond a certain threshold for an extended time. This gives you an opportunity to respond when the database starts to slow down from too many read and write operations. | | `Network` | The number of inbound and outbound bytes per second for a node. | | | `Network (sum of all nodes)` | The same as the `Network` graph, but values are not grouped by service node. | | | `TCP connections` | The number of open TCP connections, grouped by node. | | | `TCP socket state total on all nodes` | The number of TCP connections across all service nodes, grouped by the TCP connection state. | | ## PostgreSQL-specific metrics[​](#postgresql-specific-metrics "Direct link to PostgreSQL-specific metrics") For most metrics, the metric name identifies the internal PostgreSQL statistics view. See the [PostgreSQL documentation](https://www.postgresql.org/docs/current/monitoring-stats.html) for more detailed explanations of the various metric values. Metrics that are currently recorded but not shown in the default dashboard include `postgresql.pg_stat_bgwriter` and `postgresql.pg_class` metrics as a whole, as well as some individual values from other metrics. ### PostgreSQL overview[​](#postgresql-overview "Direct link to PostgreSQL overview") The metrics in the PostgreSQL overview section are grouped by logical database. In addition, some metrics are grouped by host. ![Grafana Dashboard for PostgreSQL database Overview Section](/docs/assets/images/metrics-dashboard-pg-overview-eb416e7b7cbc14eed586a92080962d89.png) | Parameter Name | Parameter Definition | Additional Notes | | ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `Database size` | The size of the files associated with a logical database | Some potentially large files that are not included in this value. Most notably, the write-ahead log (WAL) is not included in the size of the logical databases as it is not tied to any specific logical database. | | `Connections` | The number of open connections to the database | Each connection puts a large burden on the PostgreSQL server and this number should typically be [fairly small](/docs/products/postgresql/reference/pg-connection-limits.md). Use connection pooling to [reduce the number of connections](/docs/products/postgresql/concepts/pg-connection-pooling.md) to the actual database server. | | `Oldest running query age` | The age of the oldest running query | Typical queries run in milliseconds, and having queries that run for minutes often indicates an issue. | | `Oldest connection age` | The age of the oldest connection. | Old open connections with open transactions are a problem, because they prevent `VACUUM` from performing correctly, resulting in bloat and performance degradation. | | `Commits / sec` | The number of commits per second | | | `Rollbacks / sec` | The number of rollbacks per second | | | `Disk block reads / sec` | The number of 8 kB disk blocks that PostgreSQL reads per second, excluding reads that were satisfied by the buffer cache. | The read operations may have been satisfied by the operating system's file system cache. | | `Buffer cache disk block reads / sec` | The number of 8 kB disk blocks that PostgreSQL reads per second that were already in buffer cache. | | | `Temp files created / min` | The number of temporary files that PostgreSQL created per minute. | Temporary files are usually created when a query requests a large result set that can't fit in memory and needs to be sorted or when a query joins large result sets. A high number of temporary files or temporary file bytes may indicate that you should increase the working memory setting. | | `Temp file bytes written / sec` | The number of bytes written to temporary files per second | This value should be kept at reasonable levels to avoid the server becoming IO-bound from having to write so much data to temporary files. | | `Deadlocks / min` | The number of deadlocks per minute. | Deadlocks occur when different transactions obtain row-level locks for two or more of the same rows in a different order. You can resolve deadlock situations by retrying the transactions on the client side, but deadlocks can create significant bottlenecks and high counts are something that you should investigate. | ### PostgreSQL indexes[​](#postgresql-indexes "Direct link to PostgreSQL indexes") This section contains graphs related to the size and use of **indexes**. Since the default dashboard contains all indexes in all logical databases, it is convoluted for complex databases. tip You might want to make a copy of the default dashboard and add additional constraints for the graphs to filter out uninteresting indexes. For example, for the size graph, you might want to include only indexes that are above `X` megabytes in size. ![Grafana Dashboard for PostgreSQL database Indexes Section](/docs/assets/images/metrics-dashboard-pg-indexes-79ae953c1edbf0abb0113b57c375aa1d.png) | Parameter Name | Parameter Definition | Additional Notes | | --------------------------- | ---------------------------------------------------------- | ---------------- | | `Index size` | The size of indexes on disk | | | `Index scans / sec` | The number of scans per second per index | | | `Index tuple reads / sec` | The number or tuples read from an index during index scans | | | `Index tuple fetches / sec` | The number of table rows fetched during index scans | | ## Tables[​](#tables "Direct link to Tables") This section contains graphs related to the size and use of **tables**. As with indexes, the graph will be convoluted for complex databases, and you may want to make a copy of the dashboard to add additional filters that exclude uninteresting tables. ![Grafana Dashboard for PostgreSQL database Indexes Section](/docs/assets/images/metrics-dashboard-pg-tables-b7cae5ff9d025775c13a886b8230c24f.png) | Parameter Name | Parameter Definition | Additional Notes | | ----------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `Table size` | The size of tables, excluding indexes and [TOAST data](https://www.postgresql.org/docs/current/storage-toast.html) | | | `Table size total` | The total size of tables, including indexes and [TOAST data](https://www.postgresql.org/docs/current/storage-toast.html) | | | `Table seq scans / sec` | The number of sequential scans per table per second | For small tables, sequential scans may be the best way of accessing the table data and having a lot of sequential scans may be normal, but for larger tables, sequential scans should be very rare. | | `Table tuple inserts / sec` | The number of tuples inserted per second | | | `Table tuple updates / sec` | The number of tuples updated per second | | | `Table tuple deletions / sec` | The number of tuples deleted per second | | | `Table dead tuples` | The number of rows that have become un-referenced due to an update or deletion for the same row, and uncommitted transactions older than the update or delete operation are no longer running. The rows will be marked reusable during the next `VACUUM`. | High values may indicate that vacuuming is not aggressive enough. Consider adjusting its configuration to make it run more often, because frequent vacuums reduce table bloat and make the system work better. The `n_live_tup` value is available and can be used to create graphs that show tables with high ratios of dead and live tuples. | | `Table modifications since analyze` | The number of inserts, updates, or deletions since the last `ANALYZE` operation | A high number for this parameter means that the query planner may end up creating bad query plans because it is operating on obsolete data. Vacuuming also performs `ANALYZE`, and you may want to adjust your vacuum settings if you see slow queries and high table modification counts for the related tables. | ### PostgreSQL vacuum and analyse[​](#postgresql-vacuum-and-analyse "Direct link to PostgreSQL vacuum and analyse") This section contains graphs related to **vacuum** and **analyze** operations. The graphs are grouped by table and, for complex databases, you probably want to add additional filter criteria to only show results where values are outside the expected range. ![Grafana Dashboard for PostgreSQL database Vacuum and Analyse Section](/docs/assets/images/metrics-dashboard-pg-vacuum-bb74dc713dde4ec0c95286d07f404965.png) | Parameter Name | Parameter Definition | Additional Notes | | ---------------------- | ----------------------------------------------------------------- | ---------------- | | `Last vacuum age` | Time since the last manual vacuum operation for a table | | | `Last autovacuum age` | Time since the last automatic vacuum operation for a table | | | `Last analyze age` | Time since the last manual analyze operation for a table | | | `Last autoanalyze age` | Time since last automatic analyze operation for a table | | | `Maint ops / min` | The number of vacuum and analyze operations per table, per minute | | ### PostgreSQL miscellaneous[​](#postgresql-miscellaneous "Direct link to PostgreSQL miscellaneous") This section contains PostgreSQL metrics graphs that are not covered by the previous sections. ![Grafana Dashboard for PostgreSQL database Miscellaneous Section](/docs/assets/images/metrics-dashboard-pg-miscellaneous-f3bb26a309467a534c9b504e3b30b9d6.png) | Parameter Name | Parameter Definition | Additional Notes | | ------------------------ | ------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `Xact replay lag` | The replication lag between primary and standby nodes | | | `Replication bytes diff` | The replication lag in bytes. This is the total diff across all replication clients. | To differentiate between different standby nodes, you can additionally group by the `client_addr` tag. This graph shows a difference based on `write_lsn`; `flush_lsn` is also available. | | `Unfrozen transactions` | The number of transactions that have not been frozen as well as the freeze limit | In very busy systems, the number of transactions that have not been frozen by vacuum operations may rise rapidly and you should monitor this value to ensure the freeze limit is not reached. Reaching the limit causes the system to stop working. If the `txns` values get close to the freeze limit, vacuum settings need to be made more aggressive, and you must resolve any problems that prevent vacuum operations from completing, such as long-running open transactions. | --- # Resource capability of Aiven for PostgreSQL® plans When creating or updating an Aiven service, the plan that you choose will drive the specific resources (CPU, memory, disk IOPS, etc.) powering your service. Aiven is a cloud data platform, so the underlying instance types are chosen appropriately for the type of service. Elements like local NVMe SSDs, sufficient memory for the expected workloads, fast access to backup storage, ability to encrypt the disks, contribute to the choice of the instance types to use on the selected cloud platforms. In addition, particular instance types are sometimes not available in a specific cloud region. There is no one-size-fits-all for choosing the optimal instance type, so Aiven takes all of these criteria into account to select the right instance type for a given service. To know how much your service can handle, you can benchmark it with your specific workload, in a representative setup. It will be affected by more than just the instance type: network throughput, latency to your applications, number of connections active at the time, type of TLS encryption in use, and a whole range of things specific to the cloud environment can contribute to a service's expected performance. The best way to know how your application will work is to benchmark it in your setup. You can move between plans, scaling up and down, without any downtime in order to try different sizes. note Aiven also uses [industry-standard benchmarks](https://aiven.io/blog/aiven-for-postgresql-13-performance-on-gcp-aws-and-azure-benchmark) to offer guidance about the right plan for your workload. --- # Terminology for PostgreSQL® * **Primary node**: The PostgreSQL® primary node is the main server node that processes SQL queries, makes the necessary changes to the database files on the disk, and returns the results to the client application. * **Standby node**: PostgreSQL standby nodes (also called replicas) replicate the changes from the primary node and try to maintain an up-to-date copy of the same database files that exists on the primary node. * **Write-Ahead Log (WAL)**: The WAL is a log file storing all the database transactions in a sequential manner. * **pglookout**: [pglookout](https://github.com/aiven/pglookout) is a PostgreSQL replication monitoring and failover daemon. * **PGHoard**: [PGHoard](https://github.com/aiven/pghoard) is a PostgreSQL backup daemon and restore tooling that stores backup data in cloud object stores --- # Use of deprecated TLS versions TLS versions `TLSv1` and `TLSv1.1` are considered insecure, and are no longer supported in Aiven for PostgreSQL® deployments. Older services (and forks of older services) may still allow connections using these TLS versions. Support for these versions is deprecated and will be removed in the future. We recommend updating clients, or configuring them to only use `TLSv1.2` and above. Refer to the documentation of your PostgreSQL clients. To check the TLS versions clients are connecting with, you can query the `pg_stat_activity` table joined with `pg_stat_ssl`: ``` SELECT datname, pid, usesysid, usename, application_name, client_addr, ssl, version, cipher, backend_start FROM pg_stat_activity JOIN pg_stat_ssl USING (pid) WHERE client_addr IS NOT NULL; ``` ``` datname │ pid │ usesysid │ usename │ application_name │ client_addr │ ssl │ version │ cipher │ backend_start ──────────┼─────────┼──────────┼──────────┼──────────────────┼────────────────┼─────┼─────────┼────────────────────────┼─────────────────────────────── defaultdb │ 2172508 │ 16412 │ avnadmin │ psql │ 192.178.0.1 │ t │ TLSv1.3 │ TLS_AES_256_GCM_SHA384 │ 2022-09-12 12:39:12.644646+00 ``` Connections logs are also available that contain this information, for example: ``` [11-1] pid=2460224,user=test-user,db=test-db,app=[unknown],client=192.18.0.1 LOG: connection authorized: user=test-user database=test-db SSL enabled (protocol=TLSv1.1, cipher=AES256-SHA, bits=256, compression=off) ``` --- # Aiven for PostgreSQL® version lifecycle Learn how Aiven manages Aiven for PostgreSQL® version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. ## Aiven version support and upstream EOL[​](#aiven-version-support-and-upstream-eol "Direct link to Aiven version support and upstream EOL") Aiven aims to follow the EOL schedule set by the original authors and maintainers of the open source software (the upstream projects). Once the upstream project retires a specific version, they do not receive security updates and critical bug fixes anymore by the maintainers. Outdated services don't offer the level of protection you need, so Aiven follows the upstream project's EOL schedule to ensure that Aiven services are always running on supported versions. ## Service version numbering[​](#service-version-numbering "Direct link to Service version numbering") Aiven services inherit the upstream project's software versioning scheme. Depending on the service, a major version can be either a single digit or in the format `major.minor`. The exact version of the service is visible in the [Aiven Console](https://console.aiven.io/) when the service is running. ## Service version EOL policy[​](#service-version-eol-policy "Direct link to Service version EOL policy") Aiven sets an EOL date for each major version of the service. This policy covers both running and powered-off services on affected versions. ## EOL notifications[​](#eol-notifications "Direct link to EOL notifications") When Aiven sets the EOL date for a service major version: * You receive an email notification along with instructions on the next steps. * The [Aiven Console](https://console.aiven.io/) shows an EOL alert for affected services. * You receive email reminders monthly. * In the month of the EOL date, you receive weekly reminders. ## EOL best practices[​](#eol-best-practices "Direct link to EOL best practices") * Use service forking to test the version upgrade before upgrading your production services. * Upgrade to the supported version before the EOL date. This gives you time to test compatibility, resolve any issues, and plan the upgrade on your schedule. After the EOL date: * If the service is powered on, it's automatically upgraded to the latest version when possible, or to another supported version. note If it's not possible to upgrade a powered-on service to a supported version, the service is powered off and ultimately deleted. * If the service is powered off, it's deleted. ## Version EOL dates[​](#version-eol-dates "Direct link to Version EOL dates") Aiven for PostgreSQL® major versions reach EOL on the same date as the upstream open source project's EOL. | Version | Aiven EOL | Service creation supported until | Service creation supported from | | ------- | ---------- | -------------------------------- | ------------------------------- | | 9.5 | 2021-04-15 | 2021-01-26 | 2015-12-22 | | 9.6 | 2021-11-11 | 2021-05-11 | 2016-09-29 | | 10 | 2022-11-10 | 2022-05-10 | 2017-01-14 | | 11 | 2023-11-09 | 2023-05-09 | 2017-03-06 | | 12 | 2024-11-14 | 2024-05-14 | 2019-11-18 | | 13 | 2025-11-13 | 2025-05-13 | 2021-02-15 | | 14 | 2026-11-12 | 2026-05-12 | 2021-11-11 | | 15 | 2027-11-11 | 2027-05-12 | 2022-12-12 | | 16 | 2028-11-09 | 2028-05-09 | 2024-01-08 | | 17 | 2029-11-08 | 2029-05-08 | 2024-12-09 | | 18 | 2030-11-07 | 2030-05-07 | 2025-09-25 | Related pages * [Perform a PostgreSQL® major version upgrade](/docs/products/postgresql/howto/upgrade.md) --- # Scaling and performance in Aiven for PostgreSQL® Tune performance and manage disk usage to keep your Aiven for PostgreSQL® service running efficiently as it grows. Related pages * [Shared buffers](/docs/products/postgresql/concepts/pg-shared-buffers.md) * [Disk usage](/docs/products/postgresql/concepts/pg-disk-usage.md) * [Prevent a full disk](/docs/products/postgresql/howto/prevent-full-disk.md) --- # Verify the Aiven for PostgreSQL® password encryption method Verify that your Aiven for PostgreSQL® connections use `scram-sha-256` password encryption. Aiven for PostgreSQL defaults to `scram-sha-256` password encryption for enhanced security, replacing the MD5 method. In some configurations, you might need to enforce this setting manually. [Check when configuration changes are required](/docs/products/postgresql/troubleshooting/pg-password-encryption-upgrade.md#check-when-to-update-your-configuration) and, if so, update your configuration to enable `scram-sha-256`. important PostgreSQL 19 will no longer support the MD5 password encryption, making the `scram-sha-256` password encryption mandatory. ## Check when to update your configuration[​](#check-when-to-update-your-configuration "Direct link to Check when to update your configuration") * **No changes needed** if your Aiven for PostgreSQL services have: * **No** PgBouncer connection pools tied to specific database users. * All database users managed by Aiven. * **Configuration updates required** if your Aiven for PostgreSQL services have: * PgBouncer connection pools tied to specific database users. * Database users **not** managed by Aiven. When configuration updates are required, review the [`scram-sha-256` compatibility guidelines](/docs/products/postgresql/troubleshooting/pg-password-encryption-upgrade.md#ensure-scram-sha-256-compatibility) and follow the appropriate setup steps based on your configuration. ## Ensure scram-sha-256 compatibility[​](#ensure-scram-sha-256-compatibility "Direct link to Ensure scram-sha-256 compatibility") ### Ensure app connections to PgBouncer connection pools[​](#ensure-app-connections-to-pgbouncer-connection-pools "Direct link to Ensure app connections to PgBouncer connection pools") When a connection pool is configured with a specific username, an attempt to connect using another role after `scram-sha-256` is enforced fails with a `permission denied` error. This is due to the challenge-response authentication flow initiated by the PostgreSQL client and proxied by PgBouncer to PostgreSQL. 1. Check which connection pools have specific usernames by running the [`avn service connection-pool-list`](/docs/tools/cli/service/connection-pool.md) command: ``` avn service connection-pool-list --project PROJECT_NAME SERVICE_NAME ``` Example output: ``` POOL_NAME DATABASE USERNAME POOL_MODE POOL_SIZE =============== ============ ======== =========== ========= my_pool defaultdb pool_usr session 20 general_pool defaultdb transaction 15 ``` 2. Review the `USERNAME` column to identify potential issues: * **Pools with usernames** (`my_pool` with `pool_usr`) may experience authentication issues with `scram-sha-256`. * **Pools without usernames** (`general_pool`) are compatible with `scram-sha-256`. 3. For pools with specific usernames, check your application's connection string `postgresql://pool_usr:password@service-host:port/my_pool` to verify the username matches exactly: * Connection string username: `pool_usr` * Pool configuration username: `pool_usr` 4. If the usernames don't match, connect your application to a pool with a matching username or migrate the pool using one of the following methods: * Remove the username from the pool: ``` avn service connection-pool-update \ --project PROJECT_NAME SERVICE_NAME my_pool \ --username="" ``` * [Re-hash the pool user's password](/docs/products/postgresql/troubleshooting/pg-password-encryption-upgrade.md#re-hash-database-user-passwords). * Update your application to use a different compatible pool without specific username requirements: ``` postgresql://any_user:password@service-host:port/general_pool ``` ### Update service's `user_config`[​](#update-services-user_config "Direct link to update-services-user_config") Update the password encryption value in your service's `user_config`: ``` { "pg": { "password_encryption": "scram-sha-256" } } ``` This enables hashing and authenticating new managed users' passwords using `scram-sha-256`. important While this maintains the MD5 compatibility, [re-hash the passwords](/docs/products/postgresql/troubleshooting/pg-password-encryption-upgrade.md#re-hash-database-user-passwords) at your earlier convenience. ### Re-hash database user passwords[​](#re-hash-database-user-passwords "Direct link to Re-hash database user passwords") Re-hash existing passwords supported by MD5 to use the `scram-sha-256` encryption: ``` ALTER ROLE ROLE_NAME PASSWORD 'ROLE_PASSWORD'; ``` ## Troubleshoot connection issues[​](#troubleshoot-connection-issues "Direct link to Troubleshoot connection issues") If you experience authentication failures: * **Check client library support**: Ensure your PostgreSQL client supports `scram-sha-256`. * **Review connection logs**: Look for authentication method mismatches. --- # Troubleshoot connection pooling issues in Aiven for PostgreSQL® Discover the PgBouncer connection pooler and learn how to cope with some specific connection pooling issues. Verify your password encryption method If you use PGBouncer connection pooling, [verify your password encryption method compatibility](/docs/products/postgresql/troubleshooting/pg-password-encryption-upgrade.md) to ensure successful connections. You may need to migrate to `SCRAM-SHA-256` to maintain compatibility as the MD5 password encryption will be deprecated in PostgreSQL 19. ## About connection pooling with PgBouncer[​](#about-connection-pooling-with-pgbouncer "Direct link to About connection pooling with PgBouncer") PgBouncer is a lightweight connection pooler for PostgreSQL® with low memory requirements (2 kB per connection by default). PgBouncer offers several methods when rotating connections: * **Session pooling:** This is the most permissive method. When a client connects, it gets assigned with a server connection that is maintained as long as the client stays connected. When the client disconnects, the server connection is put back into the pool. This mode supports all PostgreSQL features. * **Transaction pooling:** A server connection is assigned to a client only during a transaction. When PgBouncer notices that the transaction is over, the server connection is put back into the pool. warning This mode breaks a few session-based features of PostgreSQL. Use it only when the application cooperates without using the features that break. For incompatible features, see [PostgreSQL feature map for pooling modes](https://www.pgbouncer.org/features). * **Statement pooling:** This is the most restrictive method, which disallows multi-statement transactions. This is meant to enforce the `autocommit` mode on the client and is mostly targeted at PL/Proxy. ## Handling connection pooling issues[​](#handling-connection-pooling-issues "Direct link to Handling connection pooling issues") A high CPU utilization while using the PgBouncer pooling may indicate a usage anti-pattern with a suboptimal pooling method selection or frequent reconnect operations. SSL handshakes are expensive resource-wise with asymmetric cryptography adding overhead. After a negotiation, relatively efficient symmetric ciphers are used. If clients in the application pool frequently disconnect between queries, this negates part of the benefit of the pooler and adds additional overhead. For most applications with a large pool of clients, the transaction pooling allows the application pool to maintain their connections, which helps avoid the overhead of new connection requests. For the setup and configurations of PgBouncer, refer to [Connection pooling](/docs/products/postgresql/concepts/pg-connection-pooling.md). --- # Troubleshoot out-of-shared-memory errors Identify and resolve the `out of shared memory` issue caused by stuck sessions. ## Symptoms[​](#symptoms "Direct link to Symptoms") If your Aiven for PostgreSQL® service becomes unavailable and you see repeated `FATAL: out of shared memory` messages in the logs, it may indicate a blocking session, particularly one that is in an `idle in transaction` state. Your Aiven for PostgreSQL logs might show messages similar to the following: ``` 2024-09-11T18:31:19.653257+0000 postgresql-13: pid=1031851,user=_db,db=_db FATAL: out of shared memory pgbouncer_internal: login failed: FATAL: out of shared memory ``` This may prevent superuser connections. ## Identify the problem[​](#identify-the-problem "Direct link to Identify the problem") Inspect the metric for `Oldest query age` / `Longest running query` in Grafana under the PostgreSQL dashboard. Look for the **Query Statistics** panel or a similar section. A value like `16.7 hours` indicates an open transaction that’s been running too long and likely causing memory exhaustion. ## Causes[​](#causes "Direct link to Causes") A common reason for the out-of-shared-memory issue is a session stuck in the `idle in transaction` state. Such a session may hold on to locks and memory resources indefinitely, eventually leading to memory exhaustion. There are two typical causes of this problem: * App-level exception handling 1. The application opens a transaction with `BEGIN`. 2. The application executes a query. 3. The application throws an unhandled exception while processing the result. 4. The application thread hangs but keeps the connection open. Aiven for PostgreSQL backend waits indefinitely, holding the transaction open. * Missing cleanup logic Applications may not implement proper transaction timeouts or cleanup routines. ## Prevent the problem[​](#prevent-the-problem "Direct link to Prevent the problem") * Set `idle_in_transaction_session_timeout` in Aiven for PostgreSQL to automatically terminate sessions stuck in this state. Default is `24 hours`. You may reduce it to `5 minutes` or so in your user config. * Implement connection timeouts and error handling in your application logic, monitor long-running queries and transactions using PostgreSQL metrics. --- # Aiven Runtime overview Aiven Runtime lets you deploy and run containerized applications directly within your existing Aiven project infrastructure. This means you can host your applications where your data is. Aiven automatically handles the underlying networking, and container orchestration, enabling developers to focus on application logic rather than DevOps overhead. ## Key features[​](#key-features "Direct link to Key features") Aiven Runtime offers the following capabilities to streamline your development lifecycle: * **Seamless integration of compute and data resources**: Run applications where your Aiven data services are to reduce complexity, latency, and egress costs. * **Familiar tools**: Deploy applications directly from your Git repositories to decrease time-to-production. * **Secure networking**: Applications run within your trusted network boundary. ## Use cases[​](#use-cases "Direct link to Use cases") Aiven Runtime is ideal for filtering data streams in real time, building anomaly detection, creating admin dashboards, running LLMs and AI Agents securely, or shipping other internal tools in a secure and scalable way. --- # Change cloud for Aiven Runtime You can change the cloud provider or region of an Aiven Runtime application. 1. In your project, click **Runtime**. 2. Open your application. 3. In the **Cloud and network** section, click **Change cloud**. 4. Select a cloud provider and region. 5. Optional: Select a different plan. 6. Click **Change**. The app is rebuilt in the new cloud provider and region. This process may take a few moments. --- # Connect or configure a GitHub account Connect your GitHub account to deploy applications from your GitHub repositories. Required roles or permissions:`role:organization:admin` GitHub permissions:To connect a GitHub organization account, you must be an [organization owner](https://docs.github.com/en/organizations/managing-peoples-access-to-your-organization-with-roles/roles-in-an-organization#organization-owners). You can also connect a personal GitHub account. You cannot connect the same GitHub organization or personal account to more than one Aiven organization. note When you connect a GitHub account to your Aiven organization, all users in that organization can select that account in Aiven Runtime. ## Connect a GitHub account[​](#connect-a-github-account "Direct link to Connect a GitHub account") To connect your GitHub account, install the Aiven Platform app on GitHub: 1. In your project, click **Runtime**. 2. Click **Deploy application**. 3. If you did not connect an account before, click **Connect GitHub account**. If you have connected other accounts, click **Connect another account**. 4. Click **GitHub**. 5. On the tab that opens, select a GitHub account. 6. Select **All repositories** or choose specific repositories. All users in the Aiven organization can view and deploy from the connected repositories. 7. Click **Install & Authorize**. 8. To confirm, click **Connect**. In the Aiven Console tab, you can select your account and connected repositories to [deploy your application](/docs/products/runtime/deploy-apps.md). ## Configure or uninstall the Aiven Platform app on GitHub[​](#configure-or-uninstall-the-aiven-platform-app-on-github "Direct link to Configure or uninstall the Aiven Platform app on GitHub") 1. In your project, click **Runtime**. 2. Click **Deploy application**. 3. Click **Connect another account**. 4. Click **GitHub**. 5. On the tab that opens, click **Configure** on a GitHub account. You can change the connected repositories, suspend the installation, or uninstall the Aiven Platform app. ## Troubleshooting[​](#troubleshooting "Direct link to Troubleshooting") If you have issues connecting your GitHub account, [uninstall the Aiven Platform app from GitHub](https://docs.github.com/en/enterprise-cloud@latest/apps/using-github-apps/reviewing-and-modifying-installed-github-apps) and try again. --- # Connect services to Aiven Runtime Connect your deployed application to [Aiven services](/docs/products/services.md). You can connect an existing Aiven for Apache Kafka®, Aiven for PostgreSQL®, Aiven for OpenSearch®, or Aiven for Valkey™ service. You can also define integrations when you create your application by using [Compose files](/docs/products/runtime/manifest-files/compose-files.md). note An Aiven Runtime application cannot be integrated with another Runtime application. ## Connect an Aiven service[​](#connect-an-aiven-service "Direct link to Connect an Aiven service") * Console * CLI * API 1. In your project, click **Runtime**. 2. Open your application. 3. In the **Connected services** section, click **Connect service**. 4. Select the service to connect. 5. Click **Connect**. Use the `avn service integration-create` command. For example, to integrate a PostgreSQL service with your application, run: ``` avn service integration-create \ --project PROJECT_NAME \ --integration-type application_service_credential \ --source-service SERVICE_NAME \ --dest-service APPLICATION_NAME \ --user-config-json '{ "service_type": "pg", "exposed_values": { "connection_string": { "environment_variable_key": "DATABASE_URL" } } }' ``` Where: * `PROJECT_NAME` is the name of your Aiven project. * `source-service` is the name of the data service to connect. * `dest-service` is the name of your application. * `service_type` is the type of data service. For example, `pg` for PostgreSQL. * `environment_variable_key` is the environment variable your application reads for the connection URI. For other services, view the list of [default variables](/docs/products/runtime/secrets-and-variables.md#default-environment-variables). Use the `POST /v1/project/{project}/integration` endpoint. For example, to integrate an existing PostgreSQL service with an application: ``` curl -sS -X POST "https://api.aiven.io/v1/project/PROJECT_NAME/integration" \ -H "Authorization: Bearer $AIVEN_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "integration_type": "application_service_credential", "source_service": "prod-pg", "dest_service": "web-app", "user_config": { "service_type": "pg", "exposed_values": { "connection_string": { "environment_variable_key": "DATABASE_URL" } } } }' ``` Where: * `PROJECT_NAME` is the name of your Aiven project. * `source_service` is the name of the data service to integrate with your application. * `dest_service` is the name of your application. * `service_type` is the type of data service, for example `pg` for PostgreSQL. * `environment_variable_key` is the environment variable your application reads for the connection URI. For other services, view the list of [default variables](/docs/products/runtime/secrets-and-variables.md#default-environment-variables). ## Connect a Karapace schema registry[​](#connect-a-karapace-schema-registry "Direct link to Connect a Karapace schema registry") To connect services that are integrated with your application to a Karapace schema registry: * Connect the application to the Aiven for Apache Kafka® service. * Add the schema registry connection details as environment variables. ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven for Apache Kafka® service with the [Karapace schema registry enabled](/docs/products/kafka/karapace/howto/enable-karapace.md). * [The connection details](/docs/products/kafka/howto/use-schema-registry-in-java.md#get-connection-details) for the schema registry. ### Connect a schema registry during application creation[​](#connect-a-schema-registry-during-application-creation "Direct link to Connect a schema registry during application creation") * Console * CLI * API 1. In your project, click **Runtime**. 2. Click **Deploy application**. 3. Select or connect your **GitHub account**. 4. Select your **Account**, **Repository**, and **Branch**. 5. Click **Next**. 6. Select your manifest file and click **Scan**. Aiven Runtime automatically detects what applications and services are needed. 7. On the Kafka service, click **Swap with an existing service**. 8. Select the Kafka service you created and click **Apply**. 9. To configure the integration with the schema registry, click **Configure** and add the connection details as environment variables. 10. To deploy the application, click **Deploy**. When you create the application, [connect the Kafka service](#connect-an-aiven-service) and include the schema registry details in `application.environment_variables`. For example: ``` avn service create example-application \ --project example-project \ --service-type application \ --plan startup-50-1024 \ --cloud aws-eu-west-1 \ --user-config-json '{ "application": { "source": { "repository_url": "REPOSITORY_URL", "branch": "main", "build_path": "./", "containerfile_path": "Dockerfile" }, "environment_variables": [ { "key": "SCHEMA_REGISTRY_URL", "value": "SCHEMA_REGISTRY_URI", "kind": "variable" }, { "key": "SCHEMA_REGISTRY_USER", "value": "SCHEMA_REGISTRY_USER", "kind": "variable" }, { "key": "SCHEMA_REGISTRY_PASSWORD", "value": "SCHEMA_REGISTRY_PASSWORD", "kind": "secret" } ] } }' ``` Where: `SCHEMA_REGISTRY_URI`, `SCHEMA_REGISTRY_USER`, and `SCHEMA_REGISTRY_PASSWORD` are the service URI, user, and password from the Kafka service Schema Registry connection information. When you create the application, [connect the Kafka service](#connect-an-aiven-service) and include the schema registry details in `user_config.application.environment_variables`. For example: ``` curl -sS -X POST "https://api.aiven.io/v1/project/example-project/service" \ -H "Authorization: Bearer $AIVEN_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "service_name": "web-app", "service_type": "application", "plan": "startup-50-1024", "cloud": "aws-eu-west-1", "user_config": { "application": { "source": { "repository_url": "REPOSITORY_URL", "branch": "main", "build_path": "./", "containerfile_path": "Dockerfile" }, "environment_variables": [ { "key": "SCHEMA_REGISTRY_URL", "value": "SCHEMA_REGISTRY_URI", "kind": "variable" }, { "key": "SCHEMA_REGISTRY_USER", "value": "SCHEMA_REGISTRY_USER", "kind": "variable" }, { "key": "SCHEMA_REGISTRY_PASSWORD", "value": "SCHEMA_REGISTRY_PASSWORD", "kind": "secret" } ] } }, "service_integrations": [ { "integration_type": "application_service_credential", "source_service": "KAFKA_SERVICE_NAME", "user_config": { "service_type": "kafka", "exposed_values": { "bootstrap_servers": { "environment_variable_key": "KAFKA_BOOTSTRAP_SERVER" }, "security_protocol": { "environment_variable_key": "KAFKA_SECURITY_PROTOCOL" }, "access_key": { "environment_variable_key": "KAFKA_ACCESS_KEY" }, "access_cert": { "environment_variable_key": "KAFKA_ACCESS_CERT" }, "ca_cert": { "environment_variable_key": "KAFKA_CA_CERT" } } } } ] }' ``` Where: * `KAFKA_SERVICE_NAME` is the connected Kafka service with Karapace enabled. * `SCHEMA_REGISTRY_URI`, `SCHEMA_REGISTRY_USER`, and `SCHEMA_REGISTRY_PASSWORD` are the service URI, user, and password from the Kafka service Schema Registry connection information. ### Connect a schema registry to an existing application[​](#connect-a-schema-registry-to-an-existing-application "Direct link to Connect a schema registry to an existing application") * Console * CLI * API 1. In your project, click **Runtime**. 2. Open your application. 3. In the **Environment variables** section, click **Edit**. 4. On the **Variables** tab, add the connection details as environment variables. 5. Click **Save**. Use the `avn service update` command. warning This replaces the application's environment variables. To keep the existing variables, include them in the `environment_variables` list. To view a list of the existing environment variables, run `avn service get APPLICATION_NAME`. For example: ``` avn service update example-application \ --project example-project \ -c 'application.environment_variables=[ { "key": "SCHEMA_REGISTRY_URL", "value": "SCHEMA_REGISTRY_URI", "kind": "variable" }, { "key": "SCHEMA_REGISTRY_USER", "value": "SCHEMA_REGISTRY_USER", "kind": "variable" }, { "key": "SCHEMA_REGISTRY_PASSWORD", "value": "SCHEMA_REGISTRY_PASSWORD", "kind": "secret" } ]' ``` Where: `SCHEMA_REGISTRY_URI`, `SCHEMA_REGISTRY_USER`, and `SCHEMA_REGISTRY_PASSWORD` are the service URI, user, and password from the Kafka service Schema Registry connection information. Use the `PUT /v1/project/{project}/service/{service}` endpoint. warning This replaces the application's environment variables. To keep the existing variables, include them in the `environment_variables` list. To view a list of the existing environment variables, call `GET /v1/project/{project}/service/{service}`. For example: ``` curl -sS -X PUT "https://api.aiven.io/v1/project/PROJECT_NAME/service/example-application" \ -H "Authorization: Bearer $AIVEN_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "user_config": { "application": { "environment_variables": [ { "key": "SCHEMA_REGISTRY_URL", "value": "SCHEMA_REGISTRY_URI", "kind": "variable" }, { "key": "SCHEMA_REGISTRY_USER", "value": "SCHEMA_REGISTRY_USER", "kind": "variable" }, { "key": "SCHEMA_REGISTRY_PASSWORD", "value": "SCHEMA_REGISTRY_PASSWORD", "kind": "secret" } ] } } }' ``` Where: `SCHEMA_REGISTRY_URI`, `SCHEMA_REGISTRY_USER`, and `SCHEMA_REGISTRY_PASSWORD` are the service URI, user, and password from the Kafka service Schema Registry connection information. ## Disconnect an Aiven service[​](#disconnect-an-aiven-service "Direct link to Disconnect an Aiven service") * Console * CLI * API 1. In your project, click **Runtime**. 2. Open your application. 3. In the **Connected services** section, find the service to disconnect. 4. Click **Actions** > **Disconnect service**. 5. Click **Disconnect** to confirm. 1) Get the integration ID for the connected service using the `service integration-list` command: ``` avn service integration-list APPLICATION_NAME --project PROJECT_NAME ``` 2) To remove the integration, run: ``` avn service integration-remove APPLICATION_NAME SERVICE_INTEGRATION_ID --project PROJECT_NAME ``` 1. List integrations for the application and copy the `service_integration_id` for the `application_service_credential` integration to remove: ``` curl -sS -X GET \ "https://api.aiven.io/v1/project/PROJECT_NAME/service/APPLICATION_NAME/integration" \ -H "Authorization: Bearer $AIVEN_TOKEN" ``` 2. Delete the integration: ``` curl -sS -X DELETE \ "https://api.aiven.io/v1/project/PROJECT_NAME/integration/SERVICE_INTEGRATION_ID" \ -H "Authorization: Bearer $AIVEN_TOKEN" ``` ## Apply database schema changes[​](#apply-database-schema-changes "Direct link to Apply database schema changes") Aiven Runtime does not automatically support pre-deploy commands or one-off task execution. To run database schema migrations, you can do one of the following: * **Run migrations at container startup**: You can update the `CMD` or entrypoint of your Containerfile or Dockerfile so that the database schema changes are applied every time the container starts up. * **Run migrations in CI/CD before deploying**: If you use a CI/CD pipeline, you can run migrations as a pipeline step before [deployment](/docs/products/runtime/deploy-apps.md#redeploy-an-application). --- # Connect a custom domain to an Aiven Runtime Connect a custom domain to an Aiven Runtime application using Cloudflare. Cloudflare receives traffic for your custom domain at its edge, and a Cloudflare Worker forwards each request to the Aiven-generated application hostname. This approach adds an extra network hop and a Worker subrequest before traffic reaches Aiven. This can increase latency slightly. Worker usage and pricing depend on your Cloudflare plan. ## Limitations and security considerations[​](#limitations-and-security-considerations "Direct link to Limitations and security considerations") This approach fronts your app's existing public URL. It doesn't replace the URL, and you cannot rely on the Worker as a security boundary. * **The application URL stays publicly reachable**: Traffic reaches your container with the public `*.aiven.app` host. Anyone who knows that URL can bypass your custom domain and any Cloudflare-layer protections you configure on it. * **The app must be host-aware**: Requests reach your app with the `Host` set to the Aiven hostname. To avoid leaking that hostname in redirects, links, and cookies, configure your app to trust `X-Forwarded-Host` and `X-Forwarded-Proto`. Only the `Location` response header is rewritten for you. * **Streaming and WebSockets need validation**: Server-sent events usually work with this proxy pattern. WebSockets need explicit upgrade handling in the Worker. Test both paths before production use. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven Runtime application with a publicly accessible URL. * A domain that is managed by Cloudflare and uses Cloudflare nameservers. * The domain record is set to **Proxied** in Cloudflare. * Universal SSL covers your root domain and first-level subdomains. Use Advanced Certificate Manager for deeper subdomains. ## Connect a custom domain managed by Cloudflare[​](#connect-a-custom-domain-managed-by-cloudflare "Direct link to Connect a custom domain managed by Cloudflare") 1. To get the Aiven Runtime application URL, in the Aiven Console, click **Runtime** and open your application. 2. In the **Connection information** section, copy the **Application URL**. 3. In Cloudflare, open your domain and click **DNS Records**. 4. Click **Add Record**. 5. For the **Type**, select **CNAME**. 6. Enter the **Name**. 7. For the **Target**, enter the application URL without `https://`. 8. Set the **Proxy status** to **Proxied**, and **TTL** to **Auto**. 9. Click **Back to Domains**, and in the sidebar, click **Compute** > **Workers & Pages**. 10. Click **Create application** > **Worker**. You can also use an existing Worker or deploy with [Wrangler](https://developers.cloudflare.com/workers/wrangler/). note Each exposed port has its own Aiven hostname. Front each one with its own Worker/route, or map hostnames to upstreams within a single Worker. 11. Configure the Worker. The following example forwards requests to an Aiven Runtime application host while preserving the original request method, body, and headers: ``` const AIVEN_HOST = "AIVEN_APP_HOSTNAME"; export default { async fetch(request, env, ctx) { const incomingUrl = new URL(request.url); const upstreamUrl = new URL(request.url); upstreamUrl.protocol = "https:"; upstreamUrl.hostname = AIVEN_HOST; upstreamUrl.port = ""; const headers = new Headers(request.headers); headers.delete("Host"); // Host/SNI are taken from the fetch URL automatically headers.set("X-Forwarded-Host", incomingUrl.host); headers.set("X-Forwarded-Proto", "https"); const init = { method: request.method, headers, redirect: "manual", }; if (request.method !== "GET" && request.method !== "HEAD") { init.body = request.body; } const response = await fetch(upstreamUrl.toString(), init); const responseHeaders = new Headers(response.headers); rewriteLocationHeader(responseHeaders, incomingUrl); return new Response(response.body, { status: response.status, statusText: response.statusText, headers: responseHeaders, }); }, }; function rewriteLocationHeader(headers, incomingUrl) { const location = headers.get("Location"); if (!location) { return; } const publicOrigin = `${incomingUrl.protocol}//${incomingUrl.host}`; const upstreamOrigin = `https://${AIVEN_HOST}`; if (location.startsWith(upstreamOrigin)) { headers.set("Location", location.replace(upstreamOrigin, publicOrigin)); } } ``` 12. Click **Deploy**. 13. Open your domain and click **Workers Routes**. 14. Click **Add route**. 15. Add the **Route** using the pattern that matches the hostname you configured: * **Root domain**: `example.com/*` * **`www` subdomain**: `www.example.com/*` * **Other subdomain**: `app.example.com/*` 16. Select the **Worker** and click **Save**. 17. Open the custom domain in a browser to confirm the application loads. --- # Deploy an application Build and deploy applications using Aiven Runtime from source code in a GitHub repository. Required roles or permissions:`role:organization:admin` to connect a GitHub account. `project:services:write`, `role:project:manager`, or `role:project:admin` to deploy applications. GitHub permissions:To connect a GitHub organization account, you must be an [organization owner](https://docs.github.com/en/organizations/managing-peoples-access-to-your-organization-with-roles/roles-in-an-organization#organization-owners). You can also connect a personal GitHub account. note When you connect a GitHub account to your Aiven organization, all users in that organization can select that account in Aiven Runtime. You cannot use Compose files to deploy applications through the Aiven API or Aiven MCP. Use [Containerfiles or Dockerfiles](/docs/products/runtime/manifest-files/containerfiles.md) instead. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Console * CLI * API - A GitHub account * The [Aiven CLI installed](/docs/tools/cli.md) * [An Aiven token](/docs/platform/concepts/authentication-tokens.md) * A [connected GitHub account](/docs/products/runtime/connect-github-account.md) - [An Aiven token](/docs/platform/concepts/authentication-tokens.md) - A [connected GitHub account](/docs/products/runtime/connect-github-account.md) ## Deploy an application[​](#deploy-an-application "Direct link to Deploy an application") * Console * CLI * API important When you connect a GitHub account to your Aiven organization, all users in that organization can select that account in Aiven Runtime. 1. In your project, click **Runtime**. 2. Click **Deploy application**. 3. Select or connect your **GitHub account**. 4. Select your **Account**, **Repository**, and **Branch**. 5. Click **Next**. 6. Select your manifest file and click **Scan**. Aiven Runtime automatically detects what applications and services are needed. 7. To change the configuration of an application, click . To change the configuration of a service integration, click **Configure**. 8. To deploy the application and create the services, click **Deploy**. 1) To choose a project, run: ``` avn project switch PROJECT_NAME ``` Where `PROJECT_NAME` is the name of your Aiven project. 2) Optional: Create data services for the app to use with the `avn service create` command. The following example creates a PostgreSQL service: ``` avn service create example-postgres \ --project PROJECT_NAME \ -t pg \ --cloud aws-eu-west-1 \ --plan startup-4 ``` 3) Get your `VCS_INTEGRATION_ID` from the Aiven API. This is Aiven's ID for the GitHub Aiven App installation linked to your organization when you [connected your GitHub account](/docs/products/runtime/connect-github-account.md). To get your ID, run: ``` curl -sS \ "https://api.aiven.io/v1/organization/ORGANIZATION_ID/application/vcs-integrations" \ -H "Authorization: Bearer $AIVEN_TOKEN" ``` Where `ORGANIZATION_ID` is the [Aiven organization ID](/docs/platform/reference/get-resource-IDs.md) the GitHub account is connected to. 4) Get the ID of the connected repository from the Aiven API. To get the `REMOTE_REPOSITORY_ID`, run the following command using the `VCS_INTEGRATION_ID`: ``` curl -sS \ "https://api.aiven.io/v1/organization/ORGANIZATION_ID/application/vcs-integrations/VCS_INTEGRATION_ID/repositories" \ -H "Authorization: Bearer $AIVEN_TOKEN" ``` 5) To create the application, run the following: ``` avn service create example-app \ --project PROJECT_NAME \ -t application \ --cloud aws-eu-west-1 \ --plan startup-50-1024 \ -c application.source.vcs_integration_id=VCS_INTEGRATION_ID \ -c application.source.remote_repository_id=REMOTE_REPOSITORY_ID \ -c application.source.repository_url=REPOSITORY_URL \ -c application.source.branch=BRANCH_NAME \ -c application.source.build_path=. \ -c application.source.containerfile_path=Dockerfile \ -c 'application.ports=[{"name":"http","port":8080,"protocol":"HTTP"}]' \ ``` Where: * `VCS_INTEGRATION_ID` is the GitHub Aiven app ID. * `REMOTE_REPOSITORY_ID` is the ID of the connected repository. * `REPOSITORY_URL` is the URL of the connected repository. * `BRANCH_NAME` is the branch to deploy. To use a project VPC, add `--project-vpc-id VPC_ID`. 6) Optional: Integrate your data services with the app. For example, to integrate the PostgreSQL service with the app, run: ``` avn service integration-create \ --project PROJECT_NAME \ -t application_service_credential \ -s example-postgres \ -d example-app \ --user-config-json '{"service_type":"pg","exposed_values":{"connection_string":{"environment_variable_key":"DATABASE_URL"}}}' ``` tip To check the status of your services or applications, run `avn service wait SERVICE_NAME --project PROJECT_NAME`. 1. Optional: Create data services to integrate with your application using the `POST /v1/project/{project}/service` endpoint. For example, the following creates Aiven for PostgreSQL® service: ``` curl -sS -X POST "https://api.aiven.io/v1/project/example-project/service" \ -H "Authorization: Bearer $AIVEN_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "service_name": "example-postgres-service", "service_type": "pg", "plan": "startup-4", "cloud": "aws-eu-west-1" }' ``` 2. To create the application, use the `POST/v1/project/{project}/service` endpoint. The following example deploys an application, sets environment variables, and integrates the app with an existing PostgreSQL service: ``` curl -sS -X POST \ "https://api.aiven.io/v1/project/PROJECT_NAME/service" \ -H "Authorization: Bearer $AIVEN_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "service_name": "example-app", "service_type": "application", "cloud": "aws-eu-west-1", "plan": "startup-50-1024", "user_config": { "application": { "source": { "repository_url": "REPOSITORY_URL", "branch": "BRANCH_NAME", "build_path": "./", "containerfile_path": "Dockerfile" }, "ports": [ { "name": "http", "port": 8080, "protocol": "HTTP" } ], "environment_variables": [ { "key": "LOG_LEVEL", "value": "INFO", "kind": "variable" }, { "key": "API_KEY", "value": "secret", "kind": "secret" } ] } }, "service_integrations": [ { "integration_type": "application_service_credential", "source_service": "example-postgres-service", "user_config": { "service_type": "pg", "exposed_values": { "connection_string": { "environment_variable_key": "DATABASE_URL" } } } } ] }' ``` Where: * `PROJECT_NAME` is the name of your Aiven project. * `REPOSITORY_URL` is the URL of the connected repository. * `BRANCH_NAME` is the branch to deploy. * `containerfile_path`: Use the repository-relative path for your Dockerfile or Containerfile. For example, `./Dockerfile` or `./api/Dockerfile.prod`. * `build_path` is the build context and defaults to `./.`. If you set `build_path` and omit `containerfile_path`, Aiven searches that directory for a Dockerfile/Containerfile. * `source_service` is the name of the service to integrate with the application. To use a project VPC, add `"project_vpc_id": "VPC_ID"`. ## Redeploy an application[​](#redeploy-an-application "Direct link to Redeploy an application") When you redeploy an application, Aiven deploys the latest commit from the selected branch. 1. In your project, click **Runtime**. 2. Open your application. 3. On the **Overview** page, click **Actions** > **Deploy latest commit**. --- # Change branch You can change your application's branch at any time. Changing the branch triggers a redeployment of the selected branch's latest commit. To change the branch: 1. In your project, click **Runtime**. 2. Open your application. 3. In the **Deployment configuration** section, click **Edit**. 4. Edit the branch and click **Save changes**. --- # Create Compose files for Aiven Runtime Aiven Runtime scans your repository for Compose files, such as [Docker Compose files](https://docs.docker.com/compose/), to detect applications, identify supported data services, and create integrations. Compose files must be in YAML format and follow the [Compose specification](https://compose-spec.io). Aiven recognizes Compose files with the following file naming conventions: | File type | Supported file naming formats | | --------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Compose files | - `docker-compose.yml`
- `docker-compose.yaml`
- `compose.yml`
- `compose.yaml`
- `compose.aiven.yaml`
- `compose.existing-db.yaml` | | Environment-specific and override files | - `docker-compose.override.yml`
- `compose.override.yaml`
- `docker-compose.ENVIRONMENT.yml` For example: `docker-compose.prod.yml`
- `compose.ENVIRONMENT.yaml` For example: `compose.dev.yaml` | Aiven automatically analyzes Compose files to detect the applications to build and the Aiven services to create. note You cannot use Compose files to deploy applications through the Aiven API or Aiven MCP. Use [Containerfiles or Dockerfiles](/docs/products/runtime/manifest-files/containerfiles.md) instead. ## Create a Compose file[​](#create-a-compose-file "Direct link to Create a Compose file") Use the following guidelines to create your Compose files for Aiven Runtime. More information on formatting Compose files is available in the [Compose specification](https://github.com/compose-spec/compose-spec/blob/main/spec.md) and in the [Docker Compose file reference](https://docs.docker.com/reference/compose-file). ### Service integrations[​](#service-integrations "Direct link to Service integrations") Aiven Runtime automatically detects and creates the following data services based on Docker image names: Aiven for Apache Kafka®, Aiven for PostgreSQL®, Aiven for Valkey™, and Aiven for OpenSearch®. Runtime uses variable names in the Compose file that point to the data services. If there are no variable names, it uses [default environment variable names](/docs/products/runtime/secrets-and-variables.md#integrated-service-environment-variables). Aiven integrates the data services listed in the `depends_on` property. You define the service type and tags with the `image` property, for example: ``` services: web-app: build: . depends_on: # Services to integrate with - postgres-db - valkey-cache environment: - DATABASE_URL=postgresql://DB_USER:DB_PASSWORD@postgres-db:5432/DB_NAME postgres-db: image: postgres:15 # PostgreSQL service valkey-cache: image: valkey/valkey:7.2 # Valkey service ``` note Service names must: * Consist only of lowercase letters a-z, numbers 0-9, and `-` * Begin with a lowercase letter * Be between 1 and 64 characters in length Aiven Runtime only recognizes specific, standard images for each service type. When a data service is recognized as an Aiven service, any specified version number is ignored. The latest stable version is used. You can change the version in the Aiven Console before or after deploying the application. note An Aiven Runtime application cannot be integrated with another Runtime application. #### Kafka[​](#kafka "Direct link to Kafka") Runtime recognizes the following images for Kafka: * `apache/kafka` * `confluentinc/cp-kafka` * `bitnami/kafka` The following is an example for the official Apache Kafka image: ``` services: message-broker: image: apache/kafka:3.9 ``` The following is an example for the Confluent Platform Kafka image: ``` kafka-confluent: image: confluentinc/cp-kafka:latest ``` #### PostgreSQL[​](#postgresql "Direct link to PostgreSQL") Runtime recognizes the following PostgreSQL images: * `postgres` * `postgresql` * `bitnami/postgresql` The following is an example for the official PostgreSQL image: ``` services: database: image: postgres:15 ``` #### Valkey[​](#valkey "Direct link to Valkey") Runtime recognizes the following Valkey images: * `valkey` * `valkey/valkey` * `bitnami/valkey` * `redis` * `bitnami/redis` The following is an example for the official Valkey image: ``` services: cache: image: valkey/valkey:7.2 ``` #### OpenSearch[​](#opensearch "Direct link to OpenSearch") Aiven Runtime only recognizes the official `opensearchproject/opensearch` image, for example: ``` services: search: image: opensearchproject/opensearch:2.11 ``` It can also include registry prefixes such as `docker.io/opensearchproject/opensearch` and `ghcr.io/opensearchproject/opensearch`. #### Custom builds and unsupported images[​](#custom-builds-and-unsupported-images "Direct link to Custom builds and unsupported images") Aiven Runtime doesn't support custom builds, non-standard images, and some image distributions. If you need to use custom images locally, you can use a separate Compose file for Aiven Runtime. For example, the following Compose file includes a custom PostgreSQL build and a non-standard Redis image: ``` # compose.yaml services: db: build: ./postgres-with-my-extensions # Custom build cache: image: my-redis-fork:7.0 # Non-standard image api: build: ./app ``` To deploy this setup on Aiven Runtime without editing your main Compose file, create a `compose.aiven.yaml` file with the following: ``` services: db: image: postgres:15 # Standard image Aiven recognizes cache: image: valkey:7.2 # Standard image Aiven recognizes api: build: ./app ``` #### Incorrect service detection[​](#incorrect-service-detection "Direct link to Incorrect service detection") There are cases where Runtime incorrectly identifies a service as an Aiven data service because it contains a string that matches a recognized image name. For example, a Compose file with an image name that includes `valkey` is recognized as an Aiven for Valkey™ service: ``` services: admin-app: image: my-valkey-admin-app ``` To deploy this setup on Aiven Runtime, create a separate `compose.aiven.yaml` file and a Containerfile or Dockerfile to describe the application. For the Valkey case, the following is an example of the Valkey admin app in a `compose.aiven.yaml` file: ``` services: admin-app: build: context: my-valkey-admin-app dockerfile: Dockerfile ``` An example Containerfile for this case is: ``` FROM my-valkey-admin-app EXPOSE 3000 CMD ["my-valkey-admin-app"] ``` ### Environment variables[​](#environment-variables "Direct link to Environment variables") You can use a list or dictionary format for environment variables. When you set environment variables in both Containerfiles and Compose files, Aiven merges them with the following priority: 1. Containerfile or Dockerfile environment variables are merged first, preserving their order. 2. Compose file environment variables override any Containerfile or Dockerfile variables with the same key. 3. Variables that are only in the Compose file are added at the end. Keys managed by service integrations are excluded. #### List format example[​](#list-format-example "Direct link to List format example") ``` services: app: environment: - DATABASE_URL=postgresql://user:pass@postgres-db:5432/mydb - VALKEY_URL=valkey://valkey-cache:6379 - NODE_ENV=production ``` #### Dictionary format example[​](#dictionary-format-example "Direct link to Dictionary format example") ``` services: app: environment: DATABASE_URL: postgresql://user:pass@postgres-db:5432/mydb VALKEY_URL: valkey://valkey-cache:6379 NODE_ENV: production ``` ## Example Compose files[​](#example-compose-files "Direct link to Example Compose files") ### Simple web application with PostgreSQL[​](#simple-web-application-with-postgresql "Direct link to Simple web application with PostgreSQL") The following example uses a Docker Compose file and a Dockerfile to configure a basic web application that is integrated with a PostgreSQL database. The `docker-compose.yml` file defines the application and the PostgreSQL service, along with the environment variables for integration: ``` version: '3.8' services: # Application service web-app: build: . ports: - "3000:3000" depends_on: - postgres-db environment: # Aiven will detect this integration and provide credentials - DATABASE_URL=postgresql://${POSTGRES_USER}:${POSTGRES_PASSWORD}@postgres-db:5432/${POSTGRES_DB} - NODE_ENV=production # Aiven PostgreSQL service (automatically detected) postgres-db: image: postgres:15 environment: POSTGRES_DB: ${POSTGRES_DB} POSTGRES_USER: ${POSTGRES_USER} POSTGRES_PASSWORD: ${POSTGRES_PASSWORD} ``` The Dockerfile defines how to build the web application: ``` FROM node:18-alpine WORKDIR /app # Copy package files COPY package*.json ./ # Install dependencies RUN npm ci --only=production # Copy application code COPY . . # Expose port EXPOSE 3000 # Start the application CMD ["npm", "start"] ``` ### Web application with Kafka[​](#web-application-with-kafka "Direct link to Web application with Kafka") The following example uses a Compose file to configure a web application and integrate it with a Kafka broker using SASL authentication. The file defines the application and the Kafka service, along with the environment variables for integration: ``` version: '3.8' services: # Application service web-app: build: . ports: - "8000:8000" depends_on: - kafka-broker environment: # Aiven will detect this integration and provide credentials - KAFKA_BOOTSTRAP_SERVERS=${KAFKA_BOOTSTRAP_SERVERS:-kafka-broker:9092} - KAFKA_SECURITY_PROTOCOL=${KAFKA_SECURITY_PROTOCOL:-SASL_PLAINTEXT} - KAFKA_SASL_MECHANISM=${KAFKA_SASL_MECHANISM:-PLAIN} - KAFKA_SASL_USERNAME=${KAFKA_SASL_USERNAME:-appuser} - KAFKA_SASL_PASSWORD=${KAFKA_SASL_PASSWORD:-appsecret} # Aiven Kafka broker service (automatically detected) kafka-broker: image: apache/kafka:3.9 hostname: kafka-broker ports: - "9092:9092" environment: KAFKA_NODE_ID: 1 KAFKA_LISTENER_SECURITY_PROTOCOL_MAP: "CONTROLLER:PLAINTEXT,SASL_PLAINTEXT:SASL_PLAINTEXT" KAFKA_ADVERTISED_LISTENERS: "SASL_PLAINTEXT://kafka-broker:9092" KAFKA_PROCESS_ROLES: "broker,controller" KAFKA_CONTROLLER_QUORUM_VOTERS: "1@kafka-broker:29093" KAFKA_LISTENERS: "CONTROLLER://:29093,SASL_PLAINTEXT://:9092" KAFKA_INTER_BROKER_LISTENER_NAME: "SASL_PLAINTEXT" KAFKA_CONTROLLER_LISTENER_NAMES: "CONTROLLER" CLUSTER_ID: "4L6g3nShT-eMCtK--X86sw" KAFKA_OFFSETS_TOPIC_REPLICATION_FACTOR: 1 KAFKA_TRANSACTION_STATE_LOG_MIN_ISR: 1 KAFKA_TRANSACTION_STATE_LOG_REPLICATION_FACTOR: 1 KAFKA_LOG_DIRS: "/tmp/kraft-combined-logs" KAFKA_SASL_ENABLED_MECHANISMS: "PLAIN" KAFKA_SASL_MECHANISM_INTER_BROKER_PROTOCOL: "PLAIN" KAFKA_INTER_BROKER_PROTOCOL_VERSION: "3.5" ``` ### Multi-service application[​](#multi-service-application "Direct link to Multi-service application") The following example Compose file defines a more complex application with a React frontend, a backend API, a background worker, and integrations with PostgreSQL, Valkey, and OpenSearch services. ``` services: # React frontend frontend: build: context: ./frontend dockerfile: Dockerfile ports: - "3000:3000" environment: - REACT_APP_API_URL=http://localhost:8000 # Backend API backend-api: build: context: ./backend dockerfile: Dockerfile ports: - "8000:8000" depends_on: - postgres-main - valkey-sessions - search-engine environment: # Primary database connection - DATABASE_URL=postgresql://app:password@postgres-main:5432/maindb # Session storage - VALKEY_URL=valkey://valkey-sessions:6379 # Search functionality - OPENSEARCH_URI=https://admin:password@search-engine:9200 - JWT_SECRET=your-jwt-secret - NODE_ENV=production # Background job processor worker: build: context: ./backend dockerfile: Dockerfile.worker depends_on: - postgres-main - valkey-sessions environment: - DATABASE_URL=postgresql://app:password@postgres-main:5432/maindb - VALKEY_URL=valkey://valkey-sessions:6379 - WORKER_MODE=true # Aiven PostgreSQL - Main database postgres-main: image: postgres:15 environment: POSTGRES_DB: maindb POSTGRES_USER: app POSTGRES_PASSWORD: password volumes: - postgres_data:/var/lib/postgresql/data # Aiven for Valkey - Session store and job queue valkey-sessions: image: valkey/valkey:7.2 # Aiven for OpenSearch - Full-text search search-engine: image: opensearchproject/opensearch:2.11 environment: discovery.type: single-node plugins.security.disabled: true "OPENSEARCH_JAVA_OPTS=-Xms512m -Xmx512m" volumes: - opensearch_data:/usr/share/opensearch/data volumes: postgres_data: opensearch_data: ``` Related pages * [Docker Compose Quickstart](https://docs.docker.com/compose/gettingstarted) * [Manage secrets securely in Docker Compose](https://docs.docker.com/compose/how-tos/use-secrets/) * [A Developer's Guide to Aiven Apps](https://aiven.io/blog/developers-guide-to-aiven-apps) * [Deploying Apache Kafka® Streams next to your data](https://aiven.io/blog/apache-kafka-streams-next-to-your-data) * [A Practical Guide to Deploying LMM-Powered Apps with CLIP and pgvector](https://aiven.io/blog/aiven-apps-deploy-lmm-demo) --- # Create Containerfiles and Dockerfiles for Aiven Runtime Aiven Runtime automatically detects and analyzes Containerfiles and Dockerfiles in your repository to configure applications. It recognizes Containerfiles and Dockerfiles in the following formats: * `Containerfile` * `Dockerfile` * `Containerfile.*` * `Dockerfile.*` * `*.containerfile` * `*.dockerfile` Aiven automatically analyzes these files to parse instructions for things like port detection and environment variables. ## Create a Containerfile or Dockerfile[​](#create-a-containerfile-or-dockerfile "Direct link to Create a Containerfile or Dockerfile") Containerfiles must start with a `FROM` instruction. This can come after optional parser directives, comments, and global `ARG` instructions. ## Port Detection[​](#port-detection "Direct link to Port Detection") Port numbers are extracted from the `EXPOSE` instruction. The following are example instructions for exposing ports 8000 and 9090: ``` EXPOSE 8080 EXPOSE 9090/tcp ``` ## Environment Variables[​](#environment-variables "Direct link to Environment Variables") You can define environment variables using the `ENV` instruction. For example: ``` ENV NODE_ENV=production ENV DEBUG=true ``` When you set environment variables in both Containerfiles and Compose files, Aiven merges them with the following priority: 1. Containerfile or Dockerfile environment variables are merged first, preserving their order. 2. Compose file environment variables override any Containerfile or Dockerfile variables with the same key. 3. Variables that are only in the Compose file are added at the end. Keys managed by service integrations are excluded. ## Other commands and best practices[​](#other-commands-and-best-practices "Direct link to Other commands and best practices") Details on commands and recommendations for creating Dockerfiles are available in Docker's documentation for [building best practices](https://docs.docker.com/build/building/best-practices/). ## Examples[​](#examples "Direct link to Examples") ### Single application example[​](#single-application-example "Direct link to Single application example") The following example defines a simple web application using a Containerfile: ``` FROM node:18-alpine WORKDIR /app COPY package*.json ./ RUN npm install COPY . . EXPOSE 3000 # Port ENV NODE_ENV=production # Environment variable CMD ["npm", "start"] ``` ### Multi-stage build examples[​](#multi-stage-build-examples "Direct link to Multi-stage build examples") #### Node.js application with Nginx[​](#nodejs-application-with-nginx "Direct link to Node.js application with Nginx") The following example demonstrates a multi-stage build that compiles a Node.js application and serves it with Nginx, exposing both HTTP and HTTPS ports: ``` FROM node:18-alpine AS builder WORKDIR /app COPY package*.json ./ RUN npm install COPY . . RUN npm run build FROM nginx:alpine COPY --from=builder /app/dist /usr/share/nginx/html EXPOSE 80 # HTTP port EXPOSE 443 # HTTPS port ENV NGINX_PORT=80 ``` #### Java application with custom JRE[​](#java-application-with-custom-jre "Direct link to Java application with custom JRE") The following is from the [anomaly detection example](https://github.com/Aiven-Labs/anomaly-detection-example) and demonstrates a multi-stage build that compiles a Java application and serves it with a custom JRE: ``` # --- Set ARG values for use in the following build stages # This value specifies the app to build ARG APP_NAME="AnomalyDetectorApp" # --- First stage: Get the app, work out its dependencies, create a JRE FROM gradle:9.3.0-jdk25-noble AS builder WORKDIR /app # ARG values do not persist over FROM boundaries # Explicitly reference it again to make it available ARG APP_NAME ENV APP_NAME="AnomalyDetectorApp" # Copy the gradle build environment over RUN mkdir app COPY app ./app/ RUN mkdir gradle COPY gradle ./gradle/ COPY settings.gradle ./ # Copy the run scripts for stage 2 COPY run.sh ./ COPY setup_auth.sh ./ # Start by building the app as a fat (uber) JAR # This gives us a smaller executable in stage 2 RUN gradle clean ${APP_NAME}UberJar --no-daemon ENV FAT_JAR_NAME=${APP_NAME}-uber.jar RUN cp app/build/libs/$FAT_JAR_NAME ./ # Unpack the contents of the fat JAR RUN mkdir temp && cd temp && jar xf ../$FAT_JAR_NAME # Identify the dependencies to get from the external Java environment RUN jdeps --print-module-deps \ --ignore-missing-deps \ --recursive \ --multi-release 17 \ --class-path="./temp/BOOT-INF/lib/*" \ --module-path="./temp/BOOT-INF/lib/*" \ ./$FAT_JAR_NAME > modules.txt # Assemble custom JRE with only those things in it RUN $JAVA_HOME/bin/jlink \ --verbose \ --add-modules $(cat modules.txt) \ --strip-debug \ --no-man-pages \ --no-header-files \ --compress=zip-6 \ --output ./custom-jre # ---------------------------------------------------------------------------- # --- Second stage: Run the actual image # Use the smallest base image possible (alpine) FROM debian:bookworm-slim WORKDIR /app # Install openssl (for run.sh) and RocksDB library (for Kafka Streams) RUN apt-get update \ && apt-get install -y librocksdb7.8 RUN apt-get autoremove -y \ && apt-get clean -y \ && apt-get autoclean -y \ && rm -rf /var/lib/apt/lists/* # Get the ARG value to make it available in this stage ARG APP_NAME # Set an ENV value to that value ENV APP_NAME=$APP_NAME # Copy the custom JRE and application artifacts from the builder stage COPY --from=builder /app/custom-jre /usr/lib/jvm/custom-jre COPY --from=builder /app/$APP_NAME-uber.jar ./ COPY --from=builder /app/setup_auth.sh ./ COPY --from=builder /app/run.sh ./ ENV JAVA_HOME="/usr/lib/jvm/custom-jre" ENV PATH="$JAVA_HOME/bin:$PATH" # Copy the entrypoint script and make it executable COPY run.sh ./ RUN chmod +x ./run.sh RUN chmod +x ./setup_auth.sh # Set the custom entrypoint CMD [ "./run.sh" ] ``` --- # Manifest files for Aiven Runtime Aiven Runtime uses container manifests to understand how to build and deploy your applications. You can define applications using two types of container manifests that work together to create complete solutions. * **[Compose files](/docs/products/runtime/manifest-files/compose-files.md)** define multi-service solutions that can reference and orchestrate multiple Containerfiles along with data services. * **[Containerfiles and Dockerfiles](/docs/products/runtime/manifest-files/containerfiles.md)** define how to build a single application. Aiven scans your repository for manifest files using standard naming conventions and parses instructions and configuration details like ports and environment variables. For integrating applications with Aiven data services, Aiven uses the information in Compose files to create and connect the services. --- # Manage ports for Aiven Runtime To make your application available on public networks, you can configure it to listen on ports for HTTP/S traffic. Public ports allow traffic between your application and clients on the internet such as browsers. You cannot use the following TCP destination ports for outbound connections from your application: * 23 * 25 * 119 * 135 * 137 * 138 * 139 * 179 * 445 * 465 * 631 The domain name for your application is in the **Connection information** section for the application. ## Add ports to an application[​](#add-ports-to-an-application "Direct link to Add ports to an application") To expose ports for an existing application: 1. In your project, click **Runtime**. 2. Open your application. 3. In the **Connection information** section, click **Edit ports**. 4. Click **Add port**. 5. Enter port number and name. 6. Click **Save**. ## Change or remove exposed ports[​](#change-or-remove-exposed-ports "Direct link to Change or remove exposed ports") 1. In your project, click **Runtime**. 2. Open your application. 3. In the **Connection information** section, click **Edit ports**. 4. Edit the ports. To delete a port, click . 5. Click **Save**. --- # Power off Aiven Runtime applications You can power an Aiven Runtime application on or off at any time. Powering off applications doesn't affect the connected services. You can power off services individually. Applications that are powered off for more than 180 days are automatically deleted. ## Power off an application[​](#power-off-an-application "Direct link to Power off an application") 1. In your project, click **Runtime**. 2. Open your application. 3. Click **Actions** > **Power off app**. ## Power on an application[​](#power-on-an-application "Direct link to Power on an application") 1. In your project, click **Runtime**. 2. Open your application. 3. Click **Actions** > **Power on app**. When the application finishes rebuilding, its status is **Running**. This process can take a few moments. --- # Change application plan for Aiven Runtime Adjust the plan of your applications at any time to scale them and optimize costs. When you change an application plan, the currently running commit is redeployed. Service plans for the connected services do not change. You can change the service plans for each service separately. To change your application's plan: 1. In your project, click **Runtime**. 2. Open your application. 3. In the **Application plan usage** section, click **Change plan**. 4. Select a tier, cloud, and plan. 5. Click **Change**. --- # Manage secrets and environment variables for Aiven Runtime Environment variables and secrets let you configure your application at runtime instead of embedding settings and sensitive information into your code. You can use them to pass information like API keys and database connection details to the application. This keeps sensitive data safe and makes it easy to adjust how your application behaves in different setups. Aiven Runtime also automatically exposes connection details as environment variables for connected data services. ## Manage secrets and environment variables for an application[​](#manage-secrets-and-environment-variables-for-an-application "Direct link to Manage secrets and environment variables for an application") When you edit secrets and environment variables, Aiven redeploys your application with the new configuration. It deploys the same commit from your Git branch that was deployed previously. To deploy the latest commit, you can manually [redeploy your app](/docs/products/runtime/deploy-apps.md#redeploy-an-application). 1. In your project, click **Runtime**. 2. Open your application. 3. On the **Overview** page, go to **Environment variables**. 4. Click **Edit**. 5. To add a secret, on the **Secrets** tab, click **Add secret**. To add an environment variable, on the **Variables** tab, click **Add variable**. 6. Click **Save**. ## Integrated service environment variables[​](#integrated-service-environment-variables "Direct link to Integrated service environment variables") To create an application, you can select a Compose, Containerfile, or Dockerfile manifest to scan. For Compose files, Aiven detects supported data service images, and suggests Aiven services and integrations. If an environment variable in the Compose file points to one of those detected data services, Aiven uses that variable name. Otherwise, if you listed the service in the `depends_on` property of the Compose file, a default environment variable name is suggested. The environment variables are required for the integrations, but you can customize the variable names. ### Default environment variables[​](#default-environment-variables "Direct link to Default environment variables") The following environment variables are added by default: | Service | Key | Value | | ----------------------- | ------------------------- | ---------------------------------- | | Aiven for PostgreSQL® | `DATABASE_URL` | The complete service URI. | | Aiven for Valkey™ | `VALKEY_URL` | The complete service URI. | | Aiven for OpenSearch® | `OPENSEARCH_URI` | The complete service URI. | | Aiven for Apache Kafka® | `KAFKA_BOOTSTRAP_SERVERS` | The service connection address. | | Aiven for Apache Kafka® | `KAFKA_SECURITY_PROTOCOL` | Set to `SSL`. | | Aiven for Apache Kafka® | `KAFKA_ACCESS_KEY` | The access key. | | Aiven for Apache Kafka® | `KAFKA_ACCESS_CERT` | The access certificate. | | Aiven for Apache Kafka® | `KAFKA_CA_CERT` | The trusted CA certificate bundle. | ### Credential formats[​](#credential-formats "Direct link to Credential formats") Each service type supports one credential format. PostgreSQL, Valkey, and OpenSearch provide a complete connection URI. Separate host, port, username, and password variables are not supported. Kafka credentials use the five separate variables listed in the table. Review integration suggestions before deployment. If your application expects a different format, update it to consume the supplied connection URI or Kafka variables. For example, if an application expects separate `PGHOST` and `PGPASSWORD` variables, update it to use `DATABASE_URL` instead. --- # Services Deploy fully managed and scalable open source data technologies as individual services and advanced data pipelines in minutes. ## Streaming[​](#streaming "Direct link to Streaming") [Aiven for Apache Kafka®](/docs/products/kafka.md) [Build your streaming data pipelines.](/docs/products/kafka.md) [Aiven for Apache Flink®](/docs/products/flink.md) [Control your event-driven applications and streaming analytics needs.](/docs/products/flink.md) ## Databases[​](#databases "Direct link to Databases") [Aiven for PostgreSQL®](/docs/products/postgresql.md) [The object-relational database with exentions and Aiven's AI capabilities.](/docs/products/postgresql.md) [Aiven for OpenSearch®](/docs/products/opensearch.md) [Explore and visualize your data with dashboard and plugins.](/docs/products/opensearch.md) [Aiven for ClickHouse®](/docs/products/clickhouse.md) [The cloud data warehouse to generate real-time analytical data.](/docs/products/clickhouse.md) [Aiven for Valkey™](/docs/products/valkey.md) [An in-memory NoSQL database with a small footprint and high performance.](/docs/products/valkey.md) [Aiven for MySQL®](/docs/products/mysql.md) [The relational database with all the integrations you need.](/docs/products/mysql.md) [Aiven for Dragonfly](/docs/products/dragonfly.md) [A scalable in-memory data store for high-performance.](/docs/products/dragonfly.md) [Aiven for Metrics](/docs/products/metrics.md) [Fully managed Thanos metrics – a cost-effective, open source Prometheus solution.](/docs/products/metrics.md) [Aiven for Grafana®](/docs/products/grafana.md) [Create dashboards and observe your data.](/docs/products/grafana.md) ## Managed apps[​](#managed-apps "Direct link to Managed apps") [Aiven for DataHub](/docs/products/datahub.md) [Unified governance tool for data discovery, documentation, and lineage.](/docs/products/datahub.md) --- # Aiven for Valkey™ Aiven for Valkey™ is a fully managed in-memory NoSQL database service that offers high performance, scalability, and security. Deployable in the cloud of your choice, it helps you store and access data efficiently. Developed under the Linux Foundation, Valkey™ is an open-source fork of Redis® designed to provide a seamless and reliable alternative to Redis OSS. Aiven for Valkey ensures full compatibility with Redis OSS 7.2.4, making it easy for users to transition their existing applications without disruption. With Aiven for Valkey, you can leverage the power of this in-memory database to improve the performance of your applications by setting up high-performance data caching. Additionally, it can be integrated seamlessly into your observability stack for purposes such as logging and monitoring. Aiven for Valkey supports a wide range of data structures, including strings, hashes, lists, sets, sorted sets with range queries, bitmaps, hyperloglogs, geospatial indexes, and streams. ## Key features and benefits[​](#key-features-and-benefits "Direct link to Key features and benefits") Aiven for Valkey has many features that make it easy and stress-free to use: * **Open source**: Valkey is licensed under the permissive BSD-3 license, ensuring open-source availability and freedom from restrictive licensing changes. * **Redis compatible**: Fully compatible with Redis OSS 7.2.4, providing a seamless transition for users with existing Redis applications. * **High performance**: As an in-memory NoSQL database, Valkey offers fast data retrieval with low latency, ideal for applications requiring real-time data processing. * **Managed service**: Fully managed service, so you don't have to worry about setup, management, or updates. Aiven provides tools and integrations to help you use this in-memory data store in your data pipelines. * **Fast and easy deployment**: Launch a production-ready service within minutes. Choose from multiple public clouds across numerous global regions, using high-performance clusters with optimally selected instance types and storage options. * **Integration with data infrastructure**: Aiven ensures secure network connectivity using VPC peering, PrivateLink, or TransitGateway technologies. Aiven integrates with various observability tooling, including Datadog, Prometheus, and Jolokia, or you can use Aiven's observability tools for improved monitoring and logging. * **DevOps-friendly management and development**: Manage your Valkey solution using [Aiven Console](https://console.aiven.io/), [Aiven CLI](https://github.com/aiven/aiven-client), or [Aiven Provider for Terraform](/docs/tools/terraform.md). Features like scaling, forking, and upgrading your Aiven for Valkey cluster are simple and efficient. Compatible with open-source software, it integrates with your existing applications and facilitates cloud and regional migrations. * **Backups and disaster recovery**: Automatic and configurable backups ensure data safety. Backups are performed every 24 hours, with retention periods varying by service plan. * **Migration support**: Supports data migration from external Redis® implementations and Valkey databases to Aiven for Valkey with minimal downtime. ## Ways to use Aiven for Valkey[​](#ways-to-use-aiven-for-valkey "Direct link to Ways to use Aiven for Valkey") * Use Aiven for Valkey and its in-memory data store technology as a supplementary data system alongside primary databases like PostgreSQL®. * Ideal for transient data, caching values for quick access, and data that can be reestablished, such as session data. While this service is not inherently a persistent storage solution, it can be configured for persistence. * Perfect for high-performance applications requiring fast data access, real-time analytics, session management, and distributed caching scenarios. Related pages * [Get started](/docs/products/valkey/get-started.md) * [Valkey GitHub repository](https://github.com/valkey-io/valkey) * [Valkey documentation](https://valkey.io/docs/) * [Aiven.io](https://aiven.io/valkey) * [Redis to Valkey migration guide](https://valkey.io/topics/migration/) --- # Backups and migration in Aiven for Valkey™ Back up and migrate your Aiven for Valkey™ service data. Related pages * [Aiven for Valkey™ service backups](/docs/products/valkey/howto/configure-backups.md) * [Migrate to Aiven for Valkey](/docs/products/valkey/howto/migrate-redis-aiven-via-console.md) --- # High availability in Aiven for Valkey™ Explore high availability with Aiven for Valkey™ across multiple plans. Gain insights into service continuity and understand the approach to handling failures. Non-clustered Aiven for Valkey services get automatic failover to a standby node; see [Failure handling](#failure-handling) for details. If your service uses [clustering](/docs/products/valkey/concepts/valkey-cluster.md), see [Failover](/docs/products/valkey/concepts/valkey-cluster.md#failover) instead. Compare Aiven for Valkey™ plans in the table below. Each plan offers unique configurations and features. Choose what best fits your needs. | Plan | Node configuration | High availability & Backup features | Backup history | | ------------ | ---------------------------------------- | ----------------------------------------------------------------------------------------------------- | -------------------------------------------------- | | **Hobbyist** | Single-node | Limited availability. No automatic failover. | Single backup only for disaster recovery. N/A | | **Startup** | Single-node | Limited availability. No automatic failover. | Automatic backups to a remote location. | | **Business** | Two-node (primary + standby) | High availability with automatic failover to a standby node if the primary fails. | Automatic backups to a remote location. | | **Premium** | Three-node (primary + standby + standby) | Enhanced high availability with automatic failover among multiple standby nodes if the primary fails. | Automatic backups to a remote location. | | **Custom** | Custom configurations | Custom high availability and failover features based on user requirements. | Custom backup features based on user requirements. | Refer to the [Plans & Pricing](https://aiven.io/pricing?product=redis) page for more information. ## Failure handling[​](#failure-handling "Direct link to Failure handling") * **Minor failures**: Aiven automatically handles minor failures, such as service process crashes or temporary loss of network access, without any significant changes to the service deployment. In all plans, the service automatically restores regular operation by restarting crashed processes or restoring network access when available. * **Severe failures**: In case of severe hardware or software problems, such as losing an entire node, more drastic recovery measures are required. Aiven's monitoring infrastructure automatically detects a failing node when it reports problems with its self-diagnostics or stops communicating altogether. The monitoring infrastructure then schedules the creation of a new replacement node. note In case of database failover, the **Service URI** remains unchanged, and only the IP address is updated to point to the new primary node. ## Highly available business, premium, and custom service plans[​](#highly-available-business-premium-and-custom-service-plans "Direct link to Highly available business, premium, and custom service plans") If a standby Valkey node fails, the primary node continues running normally, serving client applications without interruption. Once the replacement standby node is ready and synchronized with the primary, it begins real-time replication until the system stabilizes. When the failed node is a Valkey primary, the combined information from the Aiven monitoring infrastructure and the standby node is used to make a failover decision. The standby node is promoted as the new primary and immediately serves clients. A new replacement node is automatically scheduled and becomes the new standby node. If both the primary and standby nodes fail simultaneously, new nodes are created automatically to take their place as the new primary and standby. However, this process involves some degree of data loss since the primary node is restored from the most recent backup available. Therefore, any writes made to the database since the last backup is lost. note The amount of time it takes to replace a failed node depends mainly on the used **cloud region** and the **amount of data** that needs to be restored. However, in services with two or more nodes, such as business, premium, and custom plans, the surviving nodes continue serving clients even while recreating the other node. This is automatic and requires no administrator intervention. ## Single-node hobbyist and startup service plans[​](#single-node-hobbyist-and-startup-service-plans "Direct link to Single-node hobbyist and startup service plans") Losing the only node in the service triggers an automatic process of creating a new replacement node. The new node then restores its state from the latest available backup and resumes serving customers. During the restore operation, the service is unavailable since there is only one node providing the service. Any write operations made after the last backup are lost. Related pages * [Aiven for Valkey clustering](/docs/products/valkey/concepts/valkey-cluster.md) * [Get started with Aiven for Valkey](/docs/products/valkey/get-started.md) * [Read replica](/docs/products/valkey/concepts/read-replica.md) --- # Lua scripts with Aiven for Valkey™ Learn how to leverage the built-in support for Lua scripting in Aiven for Valkey™. Aiven for Valkey has inbuilt support for running Lua scripts to perform various actions directly on the Valkey server. Scripting is typically controlled using the `EVAL`, `EVALSHA` and `SCRIPT LOAD` commands. For all newly created Aiven for Valkey instances, `EVAL`, `EVALSHA` and `SCRIPT LOAD` commands are enabled by default. note Any outage caused by customer usage, including custom scripts, is not covered by the service level agreement (SLA). For more information about Redis scripting, check [Redis documentation](https://redis.io/commands/eval). --- # Memory management and persistence in Aiven for Valkey™ Learn how Aiven for Valkey™ addresses the challenges of high memory usage and high change rate. Discover how it implements robust memory management and persistence strategies. Aiven for Valkey™ functions primarily as a database cache. Data fetched from a database is stored in the caching system. Subsequent queries with the same parameters first check this cache, bypassing the need for a repeat database query. This efficiency can lead to challenges such as increased memory usage and frequent data changes, which Aiven for Valkey is specifically designed to manage. ## Data eviction policy in Aiven for Valkey[​](#data-eviction-policy-in-aiven-for-valkey "Direct link to Data eviction policy in Aiven for Valkey") Data eviction policy is one of the most important caching settings and it is available in the Aiven Console. Aiven for Valkey sets the `maxmemory` config, which determines the maximum amount of data that can be stored. The data eviction policy specifies what happens when this limit is reached. By default, all Aiven for Valkey services have the eviction policy set to *No eviction*. If you continue storing data without removing anything, write operations fail once the maximum memory is reached. When data is consumed at a rate similar to how it is written, it is acceptable to use the current eviction policy. However, for other use cases, it is better to use the `allkeys-lru` eviction policy. This policy starts dropping the old keys based on the least recently used strategy when `maxmemory` is reached. Another way to handle the situation is to drop random keys. note If you continue to write data, you will eventually reach the `maxmemory` limit, regardless of the eviction policy you use. ## High memory and high change rate behavior[​](#high-memory-and-high-change-rate-behavior "Direct link to High memory and high change rate behavior") For all new Aiven for Valkey services, the `maxmemory` setting is configured to **70% of available RAM** (minus management overhead) plus 5% for the replication backlog for multi-node services. This configuration limits the memory usage to below 100%, accommodating operations that require additional memory: * When a new Valkey replica connects to the primary, the service process on the primary node forks and creates an RDB snapshot, which is streamed to the new Valkey replica node. * A similar forking process occurs when the state of the Valkey service is persisted to disk, which for Aiven for Valkey happens **every 10 minutes by default**. tip * To reduce snapshot frequency, set `frequent_snapshots=false`. * To disable persistence, set `valkey_persistence=off`. note When a fork occurs, all memory pages of the new process are identical to the parent and do not consume extra memory. However, any changes in the parent process cause the memory configurations to diverge, increasing actual memory allocation. **Example** If the forked process took 4 minutes to write the RDB snapshot to disk, with new data written at 5 megabytes per second, the system memory usage can increase by approximately 1.2 gigabytes during that time. The duration of backup and replication operations depends on the total amount of memory in use. As the size of the plan increases, the memory usage also increases, which can cause a memory divergence. Therefore, the amount of memory reserved for completing these operations without using a swap is directly proportional to the total memory available. If memory usage exceeds the available memory to the extent that a swap is required, the node may become unresponsive and require replacement. In scenarios involving on-disk persistence, the child process attempts to dump all its memory onto the disk. If the parent process undergoes significant changes and uses up available RAM, it may force the child process to swap unfinished memory pages to disk. This transition makes the child process more I/O bound and further slows its operation. Increased divergence from the parent process results in more frequent swapping, further slowing the process. In such cases, the node is limited to writing and reading from swap memory. The rate at which data can be written is influenced by the size of the data values and the specifics of the Aiven plan. tip Writing at about 5 megabytes per second is manageable in most situations. However, attempting to write at about 50 megabytes per second is likely to lead to failures, particularly when memory usage approaches the allowed maximum or during the initialization of a new node. ## Initial synchronization[​](#initial-synchronization "Direct link to Initial synchronization") During system upgrades or in a high-availability setup in case of node failure, a new Valkey replica node needs to be synchronized with the current primary. The new node starts with an empty state, connects to the primary, and requests a full copy of its current state. After receiving this copy, the new node begins following the replication stream from the primary to achieve complete synchronization. Once initial synchronization is complete, the new node begins to follow the replication stream from the primary. The value of the replica client output buffer limit is set to 20% of the total memory Valkey is allowed to use. For a 28 gigabyte plan, this is around 4 gibibytes. Under memory pressure conditions, the value is set to 10%. If the volume of changes during the initial sync exceeds the replica client output buffer limit, the primary node logs `Client [...] scheduled to be closed ASAP for overcoming of output buffer limits.` The replica is disconnected, the buffer is cleared, and the replication process starts from scratch. tip Writing new changes at 5 megabytes per second, totaling 1.8 gigabytes over 6 minutes, allows the new node to start up successfully. A higher constant rate of change can cause the synchronization to fail unless the replica client output buffer limit is increased. ## Mitigation[​](#mitigation "Direct link to Mitigation") Aiven does not impose a rate limit on traffic for Valkey services because limiting only relevant write operations would require a specialized proxy, and restricting all traffic can negatively impact non-write operations. Rate limiting might be considered in the future, but for now, it is advisable to manage your workloads carefully and keep write rates moderate to prevent node failures due to excessive memory usage or challenges in initializing new nodes. note If you frequently need to write large volumes of data, contact Aiven support to discuss service configuration options that can accommodate your needs. ## Memory fragmentation and active defragmentation[​](#memory-fragmentation-and-active-defragmentation "Direct link to Memory fragmentation and active defragmentation") As keys are written, updated, and deleted, the memory allocator can end up with data spread across sparsely used memory pages. This fragmentation increases the memory a service uses beyond the size of the data it stores, and it tends to grow on long-running services with a high rate of change. To reduce fragmentation, enable the **Active memory defragmentation** (`valkey_activedefrag`) advanced configuration option. When enabled, Valkey relocates objects off sparsely used memory pages and returns the freed memory to the operating system. note Active memory defragmentation is disabled by default. It runs on the main thread and consumes CPU, so it can increase latency under load. For the full list of configuration options, including `valkey_activedefrag`, see [Advanced parameters for Aiven for Valkey™](/docs/products/valkey/reference/advanced-params.md). ## Service memory limits[​](#service-memory-limits "Direct link to Service memory limits") The practical memory limit will always be less than the service physical memory limit. **All services are subject to operating overhead:** * A small amount of memory is required by the operating system kernel to manage system resources, including networking functions and disk cache. * Aiven's cloud data platform requires memory to monitor availability, provide metrics, logging and manage backups. A server or node's **usable memory** can be calculated as: `usable memory = RAM - overhead` Where: * `overhead` is 350 MiB (≈ 0.34 GiB). Services may utilize optional components, service integrations, connection pooling, or plug-ins, which are not included in overhead calculations. If a service is overcommitted, the operating system, management layer, backups or availability monitoring, may fail status checks or operations due to resource contention. In severe instances, the node may fail completely with an out-of-memory condition. ## Out of memory conditions[​](#out-of-memory-conditions "Direct link to Out of memory conditions") Many processes request more memory from the kernel than they will ever use or need. In these cases, the kernel overallocates memory. This allows it to satisfy multiple processes requesting more memory than is available, which is not used or is freed by the time any other process actually needs it. However, if enough processes start using all their allocated memory simultaneously there may not be enough physical memory available and an `Out Of Memory` (`OOM`) condition occurs. warning This situation is critical and must be resolved immediately. The solution that the Linux kernel employs is to invoke the `Out of Memory Killer` (or `OOM Killer`). This reviews all running processes and kills one or more of them to free up system memory and keep the system running. The `OOM Killer` selects process to kill based on an `oom_score`; a calculation that balances how much memory the process is using with how long the process has been running. Processes that have been running for a long time are less likely to be killed. Subprocesses are summed with parent processes in terms of memory usage, so a process which forks many subprocesses, but itself does not use a lot of memory, may still be killed. In most instances, the hosted data service, or a child process, will have the highest memory footprint and be a prime candidate for termination when the OOM Killer inspects the running processes. Aiven's cloud data platform leverages kernel namespaces or containers to isolate processes from each other. Isolation has several benefits, including: * A smaller footprint for security‑related concerns * A smaller blast radius for failure * Greater control of system resources Left unchecked, the `OOM Killer` may opt to kill the primary service. This is undesirable as unclean termination of the primary service can lead to data loss, inconsistency, or corrupted backups. Further, if Aiven's management platform detects that the primary service is unavailable for , the service will be marked as down and a failover will occur. To mitigate this scenario, namespaces are used, some with additional memory limits, in combination with an `oom_score_adjust` on the primary process, to coax the `OOM Killer` into selection of less critical processes. This will still result in a service restart, but in a more controlled process, where the database is shut down, rather than killed; exposure to data loss is limited and recovery is faster when the service restarts, often avoiding failover. warning Out of Memory conditions can still lead to unexpected behavior, including data unavailable or data loss conditions. ## Avoid running low on memory[​](#avoid-running-low-on-memory "Direct link to Avoid running low on memory") The OOM killer only runs when the system is critically low on memory. To prevent it from running, either reduce your memory usage or increase the available memory. For most databases, the service memory footprint can often be reduced by: * Reducing concurrency or implementing connection pooling * Tuning queries to limit result sets * Tuning indexes for query load * Dropping unused objects from storage In cases where the working set no longer fits into memory, consider scaling your service. Related pages * [Change the service plan](/docs/products/valkey/howto/change-service-plan.md) * [Advanced parameters for Aiven for Valkey™](/docs/products/valkey/reference/advanced-params.md) --- # Aiven for Valkey™ read replica [Early availability](/docs/platform/concepts/service-and-feature-releases.md) Aiven for Valkey™ read replica replicates data from a primary service to a replica service across different DNS zones, clouds, or regions, enhancing data availability and supporting disaster recovery. ## Features and benefits[​](#features-and-benefits "Direct link to Features and benefits") * **High availability:** By replicating data across services in different locations, Aiven for Valkey read replica ensures your application remains available even if one service experiences downtime. * **Disaster recovery:** The replica service is automatically promoted to primary in case of failure in the original primary service, ensuring business continuity. * **Geographical distribution:** Valkey read replica supports all clouds and regions, reducing latency by keeping data closer to users. ## How read replica works[​](#how-read-replica-works "Direct link to How read replica works") Aiven for Valkey uses an **active-passive replication model**, where only data added to the primary service is replicated to the replica. This ensures high availability and data resilience. **Replication process:** New data added to the primary service is asynchronously copied to the replica service. The primary service handles all read and write operations to ensure performance and data consistency under normal conditions. **Handling failures:** If the primary service fails, Aiven’s automatic failover mechanism promotes the replica service to primary, ensuring continuous availability. To resume normal operations, update your application's connection settings to point to the new primary service. ## Limitations[​](#limitations "Direct link to Limitations") * Read replicas are supported only within the same service type. Cross-service replication is not available. * You can create a maximum of 5 read replicas per primary service. * Read replicas are only supported on the **Startup** plan. Related pages [Create an Aiven for Valkey™ read replica and promote it to primary](/docs/products/valkey/howto/create-valkey-read-replica.md) --- # Aiven for Valkey™ clustering [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Aiven for Valkey™ clustering provides a managed, scalable solution for distributed in-memory data storage with built-in high availability and automatic failover capabilities. Valkey clustering distributes your data across multiple nodes (shards) to handle larger datasets and higher traffic loads than a single-node deployment can support. Each shard contains a portion of your data, and the cluster automatically routes requests to the appropriate shard. ## Key features[​](#key-features "Direct link to Key features") ### High availability[​](#high-availability "Direct link to High availability") * **Automatic failover**: If a primary node fails, a replica is automatically promoted to maintain service availability. * **Minimal downtime**: Designed to handle both expected maintenance and unexpected failures with minimal service interruption. * **Read replicas**: Each shard includes at least one read replica for redundancy and improved read performance. ### Scalability[​](#scalability "Direct link to Scalability") * **Flexible sizing**: Supports various instance sizes, including smaller 4 GB RAM instances for cost optimization. ### Compatibility[​](#compatibility "Direct link to Compatibility") * **Cluster-enabled mode**: Fully compatible with existing Valkey and Redis cluster-aware client libraries. * **Standard protocols**: If your application currently uses a client for Valkey standalone mode, switch to a cluster-aware client to enable compatibility with Aiven for Valkey clustering. ## Architecture overview[​](#architecture-overview "Direct link to Architecture overview") ![Aiven for Valkey™ service architecture](/docs/assets/images/valkey-cluster-b409e2037adced378a5c2f42fbdb4386.png) ### Multi-shard deployment[​](#multi-shard-deployment "Direct link to Multi-shard deployment") The typical cluster deployment consists of three primary nodes, each with at least one replica, providing true high availability and scalability. * **Distributed data**: Data is automatically partitioned across multiple shards. * **Independent replicas**: Each shard has its own set of replicas for redundancy. * **Load distribution**: Requests are distributed across shards based on data location. ### Single-shard deployment[​](#single-shard-deployment "Direct link to Single-shard deployment") While Aiven for Valkey supports single-node clusters, this configuration is functionally equivalent to a standalone Valkey instance and is not the primary use case for clustering. * **Initial configuration**: Starts with one primary node and 0 - 2 read replicas * **Use case**: Ideal for smaller datasets or applications with moderate traffic * **High availability**: Automatic failover to replicas if the primary fails ## Benefits[​](#benefits "Direct link to Benefits") ### Performance[​](#performance "Direct link to Performance") * **Higher throughput**: Distribute read and write operations across multiple nodes. * **Read scaling**: Multiple replicas per shard increase read capacity. ### Reliability[​](#reliability "Direct link to Reliability") * **Fault tolerance**: Adding replicas for each shard at service creation ensures your service remains available even if individual nodes fail. * **Automatic recovery**: Failed nodes are automatically replaced and synchronized. * **Data protection**: Multiple copies of your data across different nodes. ### Operational simplicity[​](#operational-simplicity "Direct link to Operational simplicity") * **Managed service**: Aiven handles cluster setup, maintenance, and scaling. * **Automated operations**: Node discovery, failover, and resharding happen automatically. * **Monitoring included**: Built-in metrics for performance and health monitoring ## Use cases[​](#use-cases "Direct link to Use cases") ### High-traffic applications[​](#high-traffic-applications "Direct link to High-traffic applications") * Applications requiring more throughput than a single node can provide * Systems with high read/write ratios that benefit from multiple replicas * Services needing guaranteed uptime despite hardware failures ### Large datasets[​](#large-datasets "Direct link to Large datasets") * Data that exceeds the memory capacity of a single node * Applications requiring data partitioning for performance optimization * Systems that need to scale storage capacity ### Mission-critical systems[​](#mission-critical-systems "Direct link to Mission-critical systems") * Applications requiring high availability and automatic failover * Services that cannot tolerate single points of failure * Systems with strict uptime requirements ## How it works[​](#how-it-works "Direct link to How it works") ### Plan your deployment[​](#plan-your-deployment "Direct link to Plan your deployment") 1. **Assess your requirements**: Determine your data size, traffic patterns, and availability needs. 2. **Choose your configuration**: Start with a single shard for smaller workloads or multiple shards for larger datasets. 3. **Select instance sizes**: Choose appropriate memory and compute resources for your workload. ### Create a clustered service[​](#create-a-clustered-service "Direct link to Create a clustered service") To enable clustering in Aiven for Valkey, choose a multi-node cluster plan when creating your service. tip For high availability and improved read scalability, **add replicas** to each service shard during service creation. This allows you to fully leverage the benefits of clustering from the start. ### Configure a client[​](#configure-a-client "Direct link to Configure a client") * Ensure your application uses a cluster-aware Valkey/Redis client library. If your application currently uses a client for Valkey standalone mode, switch to a cluster-aware client to enable compatibility with Aiven for Valkey clustering. * Configure your client to discover and connect to cluster nodes automatically. * Test failover behavior to ensure your application handles node changes gracefully. ## Failover[​](#failover "Direct link to Failover") When a primary node stops responding, Aiven for Valkey promotes one of its replicas to primary so the shard keeps accepting writes. This differs from failover in [non-clustered Aiven for Valkey services](/docs/products/valkey/concepts/high-availability.md). A non-clustered service holds your whole dataset on a single primary, so a primary failure pauses every key until a standby takes over. A clustered service splits your dataset across multiple shards, each with its own primary, so only the keys in the affected shard's slot range are unavailable during a failover; the other shards keep serving traffic. How Aiven promotes the replacement primary for the affected shard depends on how many shards the cluster has: * **Clusters with three or more shards, where every shard has a replica**: The surviving primaries vote and promote a replica of the affected shard automatically, following Valkey's built-in cluster election. Aiven only steps in if this election doesn't complete. * **Clusters with one or two shards, or a shard with no replica**: Promoting a replica through Valkey's election requires a majority vote among primaries, and a cluster with fewer than three shards can't reach that majority once one primary is unreachable. In this case, Aiven promotes the replica that's furthest ahead in replication, meaning the one with the least data loss, once it confirms the failed primary is no longer part of the service. If the shard has no replica, Aiven provisions a new node instead and restores its data from the most recent backup. Failover isn't instant, so while a shard has uncovered hash slots, commands for keys in that shard's slot range fail; other shards keep serving their own keys without interruption. Add at least one replica to every shard to reduce how long a failure affects that shard. You can inspect cluster health at any time by running `CLUSTER NODES`. ## Resharding[​](#resharding "Direct link to Resharding") Aiven for Valkey distributes data across primary nodes using hash slots. When the number of primary nodes in your cluster changes, Aiven reshards the cluster automatically. Resharding redistributes the hash slots, and the keys they hold, across the available primary nodes to keep the slots evenly balanced across shards. Resharding runs as part of a service plan change that adds or removes primary nodes. Aiven manages the entire process: * **Slot redistribution**: Aiven divides the ranges of hash slots owned by each primary node and reassigns them across the updated set of primary nodes. * **Key migration**: Keys move together with their slots while the cluster stays available to clients. * **No manual slot management**: You cannot move individual slots or assign them to specific nodes. Aiven controls slot placement to keep the cluster balanced and consistent. The `MIGRATE` command that resharding uses to move keys between nodes stays disabled for direct use. To inspect the slot layout, run `CLUSTER NODES` on any Valkey node in the cluster. It shows the current slot distribution across the cluster nodes, so you can also use it to follow the progress of a resharding operation. When you scale in a cluster, meaning you reduce the number of primary nodes, the same dataset needs to fit into fewer nodes. If your dataset size exceeds the reduced memory capacity, Valkey starts evicting keys to free up space. warning Before you scale in a cluster, set your eviction policy to `allkeys-lru`, `allkeys-lfu`, or `allkeys-random`. Aiven requires one of these eviction policies for scale-in operations. For more information, see [Memory management](/docs/products/valkey/concepts/memory-usage.md). ## Backup and restore[​](#backup-and-restore "Direct link to Backup and restore") Aiven for Valkey automatically backs up your clustered service. Each primary node backs up the data for the hash slots it owns, and Aiven stores these backups in a remote location. Backups run independently for each primary and need no coordination from your application. To restore a cluster, Aiven combines the stored backups with the recorded hash slot layout, so your data returns to the same slot distribution. The cluster must keep the same number of primary nodes for a restore to succeed. note Cluster backups are not point-in-time recovery (PITR). Because each primary node is backed up independently, backups are not consistent across shards. A restored cluster reflects each primary's data as of its own backup, not a single moment in time across the whole cluster. Design your application to tolerate this if you rely on a restore. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * Valkey clustering is in [limited availability (LA)](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-). * Valkey clustering is supported for new services only. * Performance factors * Network latency between shards can affect cross-shard operations. * Resharding operations may temporarily impact performance. * Client library choice can affect cluster performance and behavior. Related pages * [Get started with Aiven for Valkey](/docs/products/valkey/get-started.md) * [Scaling and performance](/docs/products/valkey/scaling-performance.md) * [High availability](/docs/products/valkey/concepts/high-availability.md) * [Read replica](/docs/products/valkey/concepts/read-replica.md) --- # Aiven for Valkey™ free tier Use Aiven for Valkey™ for free. You don't need a credit card to sign up and you can use it indefinitely free of charge. ## Features and limitations[​](#features-and-limitations "Direct link to Features and limitations") Free Valkey™ services include: * A single node * 1 CPU per virtual machine * 1 GB RAM * `maxmemory` set to 50% * Monitoring for metrics and logs * Backups There are some limitations of the free tier: * Cannot create the service in a VPC * No static IPs * No integrations * No forking * No support services * Only one service of each service type in your [organization](/docs/platform/concepts/orgs-units-projects.md) * Not covered under Aiven's 99.99% SLA Free services do not have any time limitations. However, Aiven reserves the right to: * Power off free services with no initial usage within the first few hours after the service is running. You can power them back on at any time. * Power off free services with no continuative activity on the service. A notification is sent before the service is powered off. You can power them back on at any time. * Shut down services if Aiven believes they violate the [acceptable use policy](https://aiven.io/terms). * Change the cloud provider, region, or configuration at any time. ## Upgrade or downgrade a free service[​](#upgrade-or-downgrade-a-free-service "Direct link to Upgrade or downgrade a free service") You can upgrade your free service to a paid plan at any time by adding a payment method to the project's billing group. To upgrade a free service: 1. Go to the service **Overview** page. 2. In the **Service plan usage** section, click **Upgrade plan**. The upgrade happens immediately; however, it can take up to 3 hours for Basic tier support to be available. You can also downgrade a paid plan to the free tier as long as: * The data you have in that trial or paid service fits into the smaller instance size. * The free tier is available in the same cloud as the paid plan. --- # Get started with Aiven for Valkey™ Begin your journey with Aiven for Valkey™, the versatile in-memory data store offering high-performance capabilities for caching, message queues, and efficient data storage solutions. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Console * CLI * Terraform - Access to the [Aiven Console](https://console.aiven.io) * [Aiven CLI](https://github.com/aiven/aiven-client#installation) installed * [A personal token](https://docs.aiven.io/docs/platform/howto/create_authentication_token.html) - [Terraform installed](https://www.terraform.io/downloads) - A [personal token](/docs/platform/howto/create_authentication_token.md) ## Create a service[​](#create-a-service "Direct link to Create a service") * Console * CLI * Terraform 1. In your project, click **Services**. 2. Click **Create service**. 3. Select **Valkey**. 4. Select a **Service tier**. 5. Select a **Cloud**. note You cannot choose a cloud provider or a specific cloud region on the Free tier. 6. Select a **Plan**. note The plans available can vary between cloud providers and regions for the same service. 7. In the **Service details**, enter a name for your service. 8. Optional: Add service tags. 9. In the **Service summary**, click **Create service**. The status of the service is **Rebuilding** during its creation. When the status is **Running**, you can start using the service. This typically takes a couple of minutes and can vary between cloud providers and regions. 1. Determine the service specifications, including plan, cloud provider, region, and project name. 2. Run the following command to create a Valkey service named `demo-valkey`: ``` avn service create demo-valkey \ --service-type valkey \ --cloud CLOUD_AND_REGION \ --plan PLAN \ --project PROJECT_NAME ``` Parameters: * `avn service create demo-valkey`: Command to create new Aiven service named `demo-valkey`. * `--service-type valkey`: Specifies the service type as Aiven for Valkey. * `--cloud CLOUD_AND_REGION`: Specifies the cloud provider and region for deployment. * `--plan PLAN`: Specifies the service plan or tier. * `--project PROJECT_NAME`: Specifies the project where the service will be created. To see: * A full list of default flags, run `avn service create -h` * Type-specific options, run `avn service types -v` The following example files are also available in the [Aiven Terraform Provider repository](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/valkey) on GitHub. 1. Create a file named `provider.tf` and add the following: ``` Loading... ``` 2. Create a file named `service.tf` and add the following: ``` Loading... ``` 3. Create a file named `variables.tf` and add the following: ``` Loading... ``` 4. Create the `terraform.tfvars` file and add the values for your token and project name. To apply your Terraform configuration: 1. Initialize Terraform by running: ``` terraform init ``` The output is similar to the following: ``` Initializing the backend... Initializing provider plugins... - Finding aiven/aiven versions matching ">= 4.0.0, < 5.0.0"... - Installing aiven/aiven v4.9.2... - Installed aiven/aiven v4.9.2 ... Terraform has been successfully initialized! ... ``` 2. To create an execution plan and preview the changes, run: ``` terraform plan ``` 3. To deploy your changes, run: ``` terraform apply --auto-approve ``` ## Connect to Aiven for Valkey[​](#connect-to-aiven-for-valkey "Direct link to Connect to Aiven for Valkey") Learn how to connect to Aiven for Valkey using different programming languages or through `valkey-cli`: * [valkey-cli](/docs/products/valkey/howto/connect-valkey-cli.md) * [Go](/docs/products/valkey/howto/connect-go.md) * [Node](/docs/products/valkey/howto/connect-node.md) * [PHP](/docs/products/valkey/howto/connect-php.md) * [Python](/docs/products/valkey/howto/connect-python.md) * [Java](/docs/products/valkey/howto/connect-java.md) --- # Back up your Aiven for Valkey™ service to another region Copy your Aiven for Valkey™ service backups to a secondary region for disaster recovery. In addition to the primary service backup, you can have a secondary backup in an alternative location. important This feature is in [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-). Contact your account team to enable it. Backup to another region (BTAR) is a disaster recovery feature that allows backup files to be copied from the service's primary backup region to an additional (secondary) region. BTAR can bolster data resilience and helps improve data protection against disasters in the primary backup region. When the primary region is down, BTAR allows forking the service from an additional copy of the backup residing in a secondary region. ## Limitations[​](#limitations "Direct link to Limitations") * The cloud provider for your additional backup region must match the cloud provider for your service and the primary backup. * Secondary backup can only be restored in the region where it was stored. For a service that has the backup to another region (BTAR) feature enabled, you can check the service backup status, change the backup region, monitor the replication lag, fork and restore using the cross-region backup, or migrate to another cloud or region. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * At least one Aiven service with BTAR enabled * Access to the [Aiven Console](https://console.aiven.io/) * [Aiven API](/docs/tools/api.md) * [Aiven CLI](/docs/tools/cli.md) ## Change a backup region[​](#change-a-backup-region "Direct link to Change a backup region") 1. Log in to the [Aiven Console](https://console.aiven.io/) and go to your project. 2. On the **Services** page, select an Aiven service on which you'd like to enable BTAR. 3. On your service page, **Backups**. 4. On the **Backups** page, click **Actions** > **Edit secondary backup location**. 5. In the **Edit secondary backup location** window, use the **Backup location** menu to select a region for your additional backup. Confirm your choice by selecting **Save**. important You can change the backup region once in 24 hours. ## Monitor a service with BTAR[​](#monitor-a-service-with-btar "Direct link to Monitor a service with BTAR") There are a few things you may want to check for your Aiven service in the context of BTAR: * What is the status of a secondary backup? * Does your service have a backup in another region? * What is the target region of the secondary backup? * What is the replication lag between data availability in the primary region and the secondary region? ### Check BTAR status[​](#check-btar-status "Direct link to Check BTAR status") To see the availability, the status, and the target region of a secondary (BTAR) backup in the [Aiven Console](https://console.aiven.io/), go to your service page > **Backups** > **Secondary backup location**. ### Determine replication lag[​](#determine-replication-lag "Direct link to Determine replication lag") Determine the target region and the replication lag for a secondary (BTAR) backup of your service, call the [ServiceBackupToAnotherRegionReport](https://api.aiven.io/doc/#tag/Service/operation/ServiceBackupToAnotherRegionReport) endpoint. Configure the call as follows: 1. Enter `YOUR-PROJECT-NAME` and `YOUR-SERVICE-NAME` into the URL. 2. Specify `DESIRED-TIME-PERIOD` depending on the time period you need the metrics for: select one of the following values for the `period` key: `hour`, `day`, `week`, `month`, or `year`. ``` curl --request POST \ --url https://api.aiven.io/v1/project/YOUR-PROJECT-NAME/service/YOUR-SERVICE-NAME/backup_to_another_region/report \ --header 'Authorization: Bearer YOUR-BEARER-TOKEN' \ --header 'content-type: application/json' \ --data '{"period":"DESIRED-TIME-PERIOD"}' ``` As output, you get metrics including replication lags at specific points in time. ## Fork and restore a service with BTAR[​](#fork-and-restore "Direct link to Fork and restore a service with BTAR") You can use the [Aiven Console](/docs/tools/aiven-console.md), [API](/docs/tools/api.md), or [CLI](/docs/tools/cli.md) to recover your service from a backup in another region. To restore your service using BTAR, create a fork of the original service in the region where the secondary backup resides. note When you **fork & restore** from the secondary backup, your new fork service is created in the cloud and region where the secondary backup is located. The fork service gets the same plan that the primary service uses. Backups of the fork service are located in the region where this new service is hosted. * Aiven Console * Aiven CLI * Aiven API 1. Open the [Aiven Console](https://console.aiven.io/) and go to your service homepage. 2. Click **Backups**. 3. On the **Backups** page, select **Fork & restore**. 4. In the **New database fork** window: 1. Set **Backup location** to either **Primary location** or **Secondary location**. 2. Set **Backup version** to one of the following: * **Latest transaction** * **Point in time**: Set it up to no earlier than the time of taking the oldest replicated base backup. 3. Specify a name for the new fork service. 4. Select **Create fork**. Run the [avn service create](/docs/tools/cli/service-cli.md#avn-cli-service-create) command with the `--service-to-fork-from` option and the `--recovery-target-time`option. Set `--recovery-target-time` to no earlier than the time of taking the oldest replicated base backup. ``` avn service create FORK_SERVICE_NAME \ --plan SERVICE_PLAN \ --project PROJECT_NAME \ --service-type SERVICE_TYPE \ --cloud SECONDARY_BACKUP_REGION \ --recovery-target-time "YYYY-MM-DDTHH:MM:SS+00:00" \ --service-to-fork-from PRIMARY_SERVICE_NAME ``` Replace the following with meaningful data: * `FORK_SERVICE_NAME` * `SERVICE_PLAN` * `PROJECT_NAME` * `SERVICE_TYPE` * `SECONDARY_BACKUP_REGION` * `PRIMARY_SERVICE_NAME` * `YYYY-MM-DDTHH:MM:SS+00:00` Use the [ServiceCreate](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) API to create a fork service. When constructing the API request, add the `user_config` object to the request body and nest the `service_to_fork_from` field and the `recovery_target_time` field inside. Set `recovery_target_time` to no earlier than the time of taking the oldest replicated base backup. ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "cloud": "SECONDARY_BACKUP_REGION", "plan": "SERVICE_PLAN", "service_name": "FORK_SERVICE_NAME", "service_type": "SERVICE_TYPE", "user_config": { "service_to_fork_from": "PRIMARY_SERVICE_NAME", "recovery_target_time": "YYYY-MM-DDTHH:MM:SS+00:00" } }' ``` Replace the following with meaningful data: * `FORK_SERVICE_NAME` * `SERVICE_PLAN` * `PROJECT_NAME` * `SERVICE_TYPE` * `SECONDARY_BACKUP_REGION` * `PRIMARY_SERVICE_NAME` * `YYYY-MM-DDTHH:MM:SS+00:00` ## Migrate a service with BTAR[​](#migrate-a-service-with-btar "Direct link to Migrate a service with BTAR") You can migrate a service with BTAR the same way you [migrate a service with a regular backup](/docs/platform/howto/migrate-services-cloud-region.md). note When you migrate your service, locations of service backups, both primary and secondary ones, do not change. ## Delete a cross-region backup[​](#delete-a-cross-region-backup "Direct link to Delete a cross-region backup") Delete an additional service backup created in a region different from your primary backup region. You can delete a cross-region backup using the [Aiven Console](/docs/tools/aiven-console.md), [API](/docs/tools/api.md), or [CLI](/docs/tools/cli.md). When you delete the additional cross-region backup, you still have the default backup located in the primary, service-hosting region. * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/). 2. From the **Services** view, select an Aiven service on which you'd like to disable BTAR. 3. On your service's page, click **Backups**. 4. On the **Backups** page, click **Actions** > **Secondary backup location**. 5. In the **Edit secondary backup location** window, select **Disable**. Your additional service backup is no longer visible on your service's **Backups** page in the **Secondary backup location** column. To remove secondary backups for your service, use the [avn service update](/docs/tools/cli/service-cli.md) command to remove all target region names from the `additional_backup_regions` array. ``` avn service update your-sevice-name \ -c additional_backup_regions=\[\] ``` To remove secondary backups for your service, update the service configuration. Use the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint to remove all target regions names from the `additional_backup_regions` array. ``` curl --request PUT \ --url https://api.aiven.io/v1/project/YOUR_PROJECT_NAME/service/YOUR_SERVICE_NAME \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "additional_backup_regions": [] } }' ``` Related pages * [Backups](/docs/products/valkey/howto/configure-backups.md) * [Track restore progress](/docs/products/valkey/howto/track-restore-progress.md) --- # Benchmark Aiven for Valkey™ performance Aiven for Valkey™ uses `memtier_benchmark`, a command-line tool by Redis, for load generation and performance evaluation of NoSQL key-value databases. warning `redis-benchmark` is not supported to work with Aiven services, since `CONFIG` command is not allowed to run. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An Aiven for Valkey service running. * `memtier_benchmark` installed. To install the tool, download the source code from [GitHub](https://github.com/RedisLabs/memtier_benchmark) and see the [README](https://github.com/RedisLabs/memtier_benchmark/blob/master/README.md) to install all the dependencies. Next, build and install the tool. note The `Testing` section within the [README](https://github.com/RedisLabs/memtier_benchmark/blob/master/README.md) is optional. ## Run benchmark[​](#run-benchmark "Direct link to Run benchmark") Before using `memtier_benchmark`, explore its capabilities with `mentier_benchmark -h` or this [Redis article](https://redis.com/blog/memtier_benchmark-a-high-throughput-benchmarking-tool-for-redis-memcached/). Substitute the following variables in the commands. The **Overview** page of your Aiven for Valkey service contains this information. | Variable | Description | | ---------- | ---------------------------------------- | | `USERNAME` | User name of Aiven for Valkey connection | | `PASSWORD` | Password of Aiven for Valkey connection | | `HOST` | Hostname for Valkey connection | | `PORT` | Port for Valkey connection | The following is a sample command from the [Redis blog](https://redis.com/blog/benchmark-shared-vs-dedicated-redis-instances/). This command executes `10000 (-n 10000)` SET & GET operations with a `1:1 ratio (--ratio 1:1)`. It launches `4 threads (-t 4)`, with each thread opening `25 connections (-c 25)`. The tool performs `10 iterations (-x 10)` of each run to collect meaningful aggregate averages. ``` memtier_benchmark -a 'USERNAME:PASSWORD' -s 'HOST' -p 'PORT' --tls --tls-skip-verify -t 4 -n 10000 --ratio 1:1 -c 25 -x 10 -d 100 --key-pattern S:S ``` The output provides detailed metrics on operations per second, latency, and throughput for each test run. Below is an example output for a single benchmark cycle: ``` Writing results to stdout [RUN #1] Preparing benchmark client... [RUN #1] Launching threads now... [RUN #1 100%, 216 secs] 0 threads: 1000000 ops, 1996 (avg: 4621) ops/sec, 277.88KB/sec (avg: 642.15KB/sec), 50.54 (avg: 21.63) msec latency <<<< many lines for RUN #2 to RUN #9 [RUN #10] Preparing benchmark client... [RUN #10] Launching threads now... [RUN #10 100%, 224 secs] 0 threads: 1000000 ops, 4116 (avg: 4444) ops/sec, 572.83KB/sec (avg: 617.53KB/sec), 24.40 (avg: 22.49) msec latency 4 Threads 25 Connections per thread 10000 Requests per client BEST RUN RESULTS ============================================================================================================================ Type Ops/sec Hits/sec Misses/sec Avg. Latency p50 Latency p99 Latency p99.9 Latency KB/sec ---------------------------------------------------------------------------------------------------------------------------- Sets 2404.11 --- --- 21.62541 20.60700 48.89500 112.63900 339.90 Gets 2404.11 2404.11 0.00 21.62707 20.60700 49.15100 105.98300 328.16 Waits 0.00 --- --- --- --- --- --- --- Totals 4808.23 2404.11 0.00 21.62624 20.60700 49.15100 111.10300 668.06 Request Latency Distribution Type <= msec Percent ------------------------------------------------------------------------ SET 11.327 0.000 <<<< many lines GET 66.559 100.000 --- WAIT 0.000 100.000 WORST RUN RESULTS ============================================================================================================================ Type Ops/sec Hits/sec Misses/sec Avg. Latency p50 Latency p99 Latency p99.9 Latency KB/sec ---------------------------------------------------------------------------------------------------------------------------- Sets 2249.10 --- --- 22.94219 21.63100 47.87100 109.56700 317.98 Gets 2249.10 2249.10 0.00 22.94561 21.63100 47.87100 109.05500 307.00 Waits 0.00 --- --- --- --- --- --- --- Totals 4498.20 2249.10 0.00 22.94390 21.63100 47.87100 109.56700 624.99 Request Latency Distribution Type <= msec Percent ------------------------------------------------------------------------ SET 10.047 0.000 <<<< many lines GET 191.487 100.000 --- WAIT 0.000 100.000 AGGREGATED AVERAGE RESULTS (10 runs) ============================================================================================================================ Type Ops/sec Hits/sec Misses/sec Avg. Latency p50 Latency p99 Latency p99.9 Latency KB/sec ---------------------------------------------------------------------------------------------------------------------------- Sets 2312.01 --- --- 22.42681 21.24700 47.35900 101.88700 326.88 Gets 2312.01 2312.01 0.00 22.42914 21.24700 47.35900 101.88700 315.59 Waits 0.00 --- --- --- --- --- --- --- Totals 4624.02 2312.01 0.00 22.42798 21.24700 47.35900 101.88700 642.47 Request Latency Distribution Type <= msec Percent ------------------------------------------------------------------------ SET 9.791 0.000 <<<< many lines GET 712.703 100.000 --- WAIT 0.000 100.000 ``` This demonstrates the performance data obtainable with `memtier_benchmark`. The initial sections present data from the `10` runs. The following sections present the `BEST RUN`, `WORST RUN` and `AGGREGATED AVERAGE` results as well as the `Request Latency Distribution` of the operations. Running this command on various Aiven for Valkey services or the same service under different conditions allows for effective performance comparisons note Aiven has `rate limit` on services. By default it's `200` new connections per 0.5 second per CPU core. Also be aware of the connection limit depending on memory size as explained in [Estimate maximum number of connection](/docs/products/valkey/howto/estimate-max-number-of-connections.md). Aiven enforces a `rate limit` on services. By default, it's set to `200` new connections per 0.5 seconds per CPU core. Additionally, consider the connection limit based on memory size as explained in [Estimate maximum number of connection](/docs/products/valkey/howto/estimate-max-number-of-connections.md). --- # Change the cloud or region for your Aiven for Valkey™ service Move your Aiven for Valkey™ service to a different cloud provider or region. 1. In your service, click **Service settings** from the sidebar. 2. In the **Cloud and network** section, click **Actions** > **Change cloud**. 3. In the **Cloud** section , select a cloud provider and region, and click **Change**. Your service starts a migration to the new location and remains available during the process. When the migration completes, the service continues running in the new cloud or region. Related pages * [Fork your service](/docs/products/valkey/howto/fork-service.md) * [Migrate to another cloud or region](/docs/platform/howto/migrate-services-cloud-region.md) --- # Change the plan for your Aiven for Valkey™ service Change the service plan for your Aiven for Valkey™ service to scale resources up or down and optimize costs. Adjust the plan of your services at any time to scale your services as needed and optimize costs. If you can't find a suitable plan, you can [request a custom plan](/docs/platform/concepts/service-pricing.md). tip If you plan to upgrade your service plan, do it immediately after a full backup. This reduces the amount of incremental changes that need to be applied on top of the base backup, which speeds up the upgrade itself. important * When changing a service plan, reserve an additional 25% of disk space. This requirement applies to upgrades and downgrades. * Downgrading to a plan with fewer VMs is supported for most services, including Aiven for Apache Kafka®, Aiven for PostgreSQL®, Aiven for OpenSearch®, Aiven for ClickHouse®, Aiven for MySQL®, Aiven for Metrics, and Aiven for Valkey™. * Changing a service plan triggers a node recycle, service rebuilding, and any pending maintenance updates. - Console - Terraform - CLI 1. In your service, click **Service settings**. 2. In the **Service plan** section, click **Change plan**. 3. Select a plan that provides at least 125% of the current disk size and click **Change plan**. Update the `plan` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). To change a service plan in the Aiven CLI, use the [`avn service update --plan `](/docs/tools/cli/service-cli.md#avn-cli-service-update) command. Your service's state becomes **Rebuilding** and remains accessible. When the state switches to **Running**, your new service plan is active. Related pages * [Prepare for high load](/docs/products/valkey/howto/prepare-for-high-load.md) --- # Configure ACL permissions in Aiven for Valkey™ Aiven for Valkey™ uses [access control lists (ACLs)](https://redis.io/docs/management/security/acl/) to manage the usage of commands and keys based on specific username and password combinations. Direct use of [ACL commands](https://redis.io/commands/acl-list/) is restricted to ensure the reliability of replication, configuration management, and disaster recovery backups for the default user. However, you can create custom ACLs using either the [Aiven Console](https://console.aiven.io/) or [Aiven CLI](/docs/tools/cli.md). ## Create user and configure ACLs[​](#create-user-and-configure-acls "Direct link to Create user and configure ACLs") * Console * CLI To create a user and configure ACLs using the Aiven Console: 1. Log in to [Aiven Console](https://console.aiven.io/), select your project, and select your Aiven for Valkey service. 2. Click **Users** from the left sidebar. 3. Click **Create user**, and provide the following details: * **Username:** Enter a username for the user. * **Categories:** Define the command categories accessible to the user. For example, use the prefix `+@all` or a similar convention to grant users access to all categories. Separate each category entry with a single space. * **Commands:** List the commands the user can execute, separating each command by a single space. For example, input `+set -get` to grant the user permission to execute the SET command and deny access to the GET command. * **Channels:** Specify the Pub/Sub channels the user can access, separating each with a space. * **Keys:** Define the keys the user can interact with. For example, specify keys like `user:123` or `product:456`, or `order:789` to grant the user access to interact with these specific keys in Aiven for Valkey. 4. Once you have defined the ACL permissions for the user, click **Save**. To create a user and configure ACLs using the Aiven CLI: 1. Ensure the [CLI tool](/docs/tools/cli.md) is set up and configured. 2. Use the following command to create a user named `mynewuser` with specific ACLs: ``` avn service user-create \ --project myproject \ --service myservicename \ --username mynewuser \ --redis-acl-keys 'mykeys.*' \ --redis-acl-commands '+get' \ --redis-acl-categories '' ``` 3. Test the ACL settings by connecting to the service using the new username: ``` valkey-cli \ --user mynewuser \ --pass ... \ --tls \ -h myservice-myproject.aivencloud.com \ -p 12719 myservice-myproject.aivencloud.com:12719> get mykeys.hello (nil) myservice-myproject.aivencloud.com:12719> set mykeys.hello world (error) NOPERM this user has no permissions to run the 'set' command or its subcommand ``` ## User management[​](#user-management "Direct link to User management") Manage users of your Aiven for Valkey service directly from the Aiven Console. ### Reset password[​](#reset-password "Direct link to Reset password") 1. Click **Users** from the left sidebar. 2. Find the user whose password needs to be reset and Click **Actions** > **Reset password**. 3. Confirm the password reset by clicking **Reset** on the confirmation dialog. ### Edit ACL rules[​](#edit-acl-rules "Direct link to Edit ACL rules") 1. Click **Users** from the left sidebar. 2. Find the user whose ACL rules require editing and Click **Actions** > **Edit ACL rules **. 3. Make the necessary changes to the ACL rules on the **Edit access control** dialog. 4. Click **Save**. ### Duplicate user[​](#duplicate-user "Direct link to Duplicate user") 1. Click **Users** from the left sidebar. 2. Locate the user you wish to duplicate and click **Actions** > **Duplicate user**. 3. Enter a name for the new user in the **Duplicate user** dialog. 4. Click **Add user**. ### Delete user[​](#delete-user "Direct link to Delete user") 1. Click **Users** from the left sidebar. 2. Find the user you intend to delete and click **Actions** > **Delete user**. 3. Confirm the deletion by clicking **Delete** on the confirmation dialog. --- # Aiven for Valkey™ service backups Learn how backups work for your Aiven for Valkey™ service and set the time when automatic backups are taken. ## How backups work[​](#how-backups-work "Direct link to How backups work") Backups are stored in the object storage of the cloud region where the service is created, for example, S3 for AWS or Google Cloud Storage for Google Cloud. note If you change a service's cloud provider or an availability zone, its backups are not migrated from their original location. Whenever a service is powered on from a powered-off state, the latest available backup is automatically restored. Backups are automatically deleted 30 days after the service's deletion date. ## Configure the backup time[​](#configure-the-backup-time "Direct link to Configure the backup time") To edit the backup schedule for your service: * Console * Aiven API * Aiven CLI * Terraform 1. In your service, **Backups**. 2. Click **Actions** > **Configure backup settings**. 3. Click **Add configuration options**. 4. Add `backup_hour` and `backup_minute`, and set their values. 5. Click **Save configuration**. Call the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint, and add the following properties to the `user_config` object: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{ "user_config": { "backup_hour": BACKUP_HOUR, "backup_minute": BACKUP_MINUTE } }' ``` Replace the following: * `SERVICE_NAME`: the name of your service. * `PROJECT_NAME`: the name of your project. * `BACKUP_HOUR`: the hour when the service backup starts. Accepted values are integers between `0` and `23`. * `BACKUP_MINUTE`: the minute when the service backup starts. Accepted values are integers between `0` and `59`. Run the [avn service update](/docs/tools/cli/service-cli.md#avn-cli-service-update) command, and add the following properties to the `user_config` object: ``` avn service update SERVICE_NAME \ --project PROJECT_NAME \ --user-config '{ "backup_hour": BACKUP_HOUR, "backup_minute": BACKUP_MINUTE }' ``` Replace the following: * `SERVICE_NAME`: the name of your service. * `PROJECT_NAME`: the name of your project. * `BACKUP_HOUR`: the hour when the service backup starts. Accepted values are integers between `0` and `23`. * `BACKUP_MINUTE`: the minute when the service backup starts. Accepted values are integers between `0` and `59`. Use the `backup_hour` and `backup_minute` attributes in [your service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs) to set the start time for backups. If a backup was recently made, it can take another backup cycle before the new backup time takes effect. note When `backup_hour` is set, the backup frequency changes from 12 hours to 24 hours. Setting `backup_hour` has no effect while `valkey_persistence` is set to `off`. Related pages * [Fork Aiven for Valkey™](/docs/products/valkey/howto/fork-service.md) * [High availability](/docs/products/valkey/concepts/high-availability.md) --- # Connect to Aiven for Valkey™ with Go Establish a connection to the Aiven for Valkey™ service using Go. This example demonstrates how to connect to Aiven for Valkey from Go using the `go-valkey/valkey` library, designed to interact with the Valkey protocol. ## Variables[​](#variables "Direct link to Variables") Replace placeholders in the code sample with values from your service overview page: | Variable | Description | | ------------- | ----------------------------------------------- | | `SERVICE_URI` | URI for the Aiven for Valkey service connection | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Install the `valkey-go` library using the following command: ``` go get github.com/valkey-io/valkey-go ``` ## Set up and run[​](#set-up-and-run "Direct link to Set up and run") 1. Create a file named `main.go` and insert the code below, substituting the placeholder with your Aiven for Valkey URI: ``` package main import ( "context" "fmt" "github.com/valkey-io/valkey-go" ) var ctx = context.Background() func main() { valkeyURI := "SERVICE_URI" opts, err = valkey.ParseURL(valkeyURI) if err != nil { return panic(err) } client, err := valkey.NewClient(opts) if err != nil { panic(err) } defer client.Close() err = client.Do(ctx, client.B().Set().Key("key").Value("hello world").Nx().Build()).Error() if err != nil { panic(err) } value, err := client.Do(ctx, client.B().Get().Key("key").Build()).ToString() if err != nil { panic(err) } fmt.Println("The value of key is:", value) } ``` This code creates a key named `key` with the value `hello world` without an expiration. It then retrieves this key from the Valkey service and outputs its value. 2. Run the script using the following command: ``` go run main.go ``` A successful connection displays: ``` The value of key is: hello world ``` Related pages * For additional information, see [valkey-go github repository](https://github.com/valkey-io/valkey-go). --- # Connect to Aiven for Valkey™ with Java Establish a connection to your Aiven for Valkey™ service using Java and the `jedis` library. ## Variables[​](#variables "Direct link to Variables") Replace placeholders in the code sample with values from your service overview page: | Variable | Description | | ------------- | ----------------------------------------------- | | `SERVICE_URI` | URI for the Aiven for Valkey service connection | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") 1. [Install Maven](https://maven.apache.org/install.html). 2. Alternatively, if you choose not to install Maven, manually download the dependencies from the [Maven Central Repository](https://search.maven.org) and place them in the `lib` folder. 3. If you have Maven installed, run the following commands in the `lib` folder to download `jedis` and its dependencies: ``` mvn org.apache.maven.plugins:maven-dependency-plugin:2.8:get -Dartifact=redis.clients:jedis:4.1.1:jar -Ddest=lib/jedis-4.1.1.jar \ && mvn org.apache.maven.plugins:maven-dependency-plugin:2.8:get -Dartifact=org.apache.commons:commons-pool2:2.11.1:jar -Ddest=lib/commons-pool2-2.11.1.jar \ && mvn org.apache.maven.plugins:maven-dependency-plugin:2.8:get -Dartifact=org.slf4j:slf4j-api:1.7.35:jar -Ddest=lib/slf4j-api-1.7.35.jar \ && mvn org.apache.maven.plugins:maven-dependency-plugin:2.8:get -Dartifact=com.google.code.gson:gson:2.8.9:jar -Ddest=lib/gson-2.8.9.jar ``` ## Set up and run[​](#set-up-and-run "Direct link to Set up and run") 1. Create a file named `ValkeyExample.java` and insert the code below, substituting the placeholder with your Aiven for Valkey URI: ``` import redis.clients.jedis.JedisPooled; public class ValkeyExample { public static void main(String[] args) { if (args.length != 1) { throw new IllegalArgumentException("Expected only one argument service URI"); } else { JedisPooled jedisPooled = new JedisPooled(args[0]); jedisPooled.set("key", "hello world"); System.out.println("The value of key is: " + jedisPooled.get("key")); } } } ``` This code connects to Aiven for Valkey, sets a `key` named key with the value `hello world` (without expiration), then retrieves and prints the value of this key. 2. To compile and run the script, replace **SERVICE\_URI** with your actual service URI: ``` javac -cp "lib/*:." ValkeyExample.java && java -cp "lib/*:." ValkeyExample SERVICE_URI ``` A successful connection displays: ``` The value of key is: hello world ``` --- # Connect to Aiven for Valkey™ with NodeJS Connect to the Aiven for Valkey™ service using NodeJS with the `ioredis` library. ## Variables[​](#variables "Direct link to Variables") Replace placeholders in the code sample with values from your service overview page: | Variable | Description | | ------------- | ----------------------------------------------- | | `SERVICE_URI` | URI for the Aiven for Valkey service connection | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Install the `ioredis` library using the following command: ``` npm install --save ioredis ``` ## Set up and run[​](#set-up-and-run "Direct link to Set up and run") 1. Create a file named `index.js` and insert the code below, substituting the placeholder with your Aiven for Valkey URI: ``` const Valkey = require('ioredis'); const serviceUri = 'SERVICE_URI'; const valkey = new Valkey(serviceUri); valkey.set('key', 'hello world'); valkey.get('key').then(function (result) { console.log(`The value of key is: ${result}`); valkey.disconnect(); }); ``` This code creates a key named `key` with the value `hello world` without an expiration. It then retrieves this key from the Valkey service and outputs its value. 2. Run the script using the following command: ``` node index.js ``` A successful connection displays: ``` The value of key is: hello world ``` --- # Connect to Aiven for Valkey™ with PHP Connect to the Aiven for Valkey™ database using PHP, making use of the `predis` library. ## Variables[​](#variables "Direct link to Variables") Replace placeholders in the code sample with values from your service overview page: | Variable | Description | | ------------- | ----------------------------------------------- | | `SERVICE_URI` | URI for the Aiven for Valkey service connection | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Install the `predis` library using the following command: ``` composer require predis/predis ``` ## Set up and run[​](#set-up-and-run "Direct link to Set up and run") 1. Create a file named `index.php` and insert the code below, substituting the placeholder with your Aiven for Valkey URI: ``` set('key', 'hello world'); $value = $client->get('key'); echo "The value of key is: {$value}"; ``` This code creates a key named `key` with the value `hello world` without an expiration. It then retrieves this key from the Valkey service and outputs its value. 2. Run the script using the following command: ``` php index.php ``` A successful connection displays: ``` The value of key is: hello world ``` --- # Connect to Aiven for Valkey™ with Python Connect to the Aiven for Valkey™ service using Python with the `valkey-py` library. `valkey-py` is a Python interface specifically designed for the Valkey key-value store. ## Variables[​](#variables "Direct link to Variables") Replace placeholders in the code sample with values from your service overview page: | Variable | Description | | ------------- | ----------------------------------------------- | | `SERVICE_URI` | URI for the Aiven for Valkey service connection | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Install the `valkey-py` library with optional performance enhancements from `hiredis` by running the following commands: ``` pip install valkey pip install "valkey[hiredis]" ``` ## Set up and run[​](#set-up-and-run "Direct link to Set up and run") 1. Create a file named `main.py` and insert the code below, substituting the placeholder with your Aiven for Valkey URI: ``` import valkey def main(): valkey_uri = 'VALKEY_SERVICE_URI' # Replace 'VALKEY_SERVICE_URI' with your actual Valkey service URI valkey_client = valkey.from_url(valkey_uri) valkey_client.set('key', 'hello world') key = valkey_client.get('key').decode('utf-8') print('The value of key is:', key) if __name__ == '__main__': main() ``` This code creates a key named `key` with the value `hello world` without an expiration. It then retrieves this key from the Valkey service and outputs its value. 2. Run the script using the following command: ``` python main.py ``` note Use `python3` instead of `python` on systems where it defaults to Python 2. Successful execution results in the following output: ``` The value of key is: hello world ``` ## Related Pages[​](#related-pages "Direct link to Related Pages") * For additional information about `valkey-py`, see the [valkey-py GitHub Repository](https://github.com/valkey-io/valkey-py). --- # Connect to Aiven for Valkey™ Connect to the Aiven for Valkey™ service using various programming languages or tools. ## [valkey-cli](/docs/products/valkey/howto/connect-valkey-cli.md) [Learn how to establish a connection to an Aiven for Valkey™ service using the valkey-cli.](/docs/products/valkey/howto/connect-valkey-cli.md) ## [Go](/docs/products/valkey/howto/connect-go.md) [Establish a connection to the Aiven for Valkey™ service using Go. This example demonstrates how to connect to Aiven for Valkey from Go using the go-valkey/valkey library, designed to interact with the Valkey protocol.](/docs/products/valkey/howto/connect-go.md) ## [NodeJS](/docs/products/valkey/howto/connect-node.md) [Connect to the Aiven for Valkey™ service using NodeJS with the ioredis library.](/docs/products/valkey/howto/connect-node.md) ## [PHP](/docs/products/valkey/howto/connect-php.md) [Connect to the Aiven for Valkey™ database using PHP, making use of the predis library.](/docs/products/valkey/howto/connect-php.md) ## [Python](/docs/products/valkey/howto/connect-python.md) [Connect to the Aiven for Valkey™ service using Python with the valkey-py library. valkey-py is a Python interface specifically designed for the Valkey key-value store.](/docs/products/valkey/howto/connect-python.md) ## [Java](/docs/products/valkey/howto/connect-java.md) [Establish a connection to your Aiven for Valkey™ service using Java and the jedis library.](/docs/products/valkey/howto/connect-java.md) ## [Estimate max connections](/docs/products/valkey/howto/estimate-max-number-of-connections.md) [The number of simultaneous connections for Aiven for Valkey™ depends on the total available memory on the server.](/docs/products/valkey/howto/estimate-max-number-of-connections.md) ## [Connection issues](/docs/products/valkey/troubleshooting/troubleshoot-connection-issues.md) [Learn troubleshooting techniques for your Aiven for Valkey™ service and resolve common connection issues.](/docs/products/valkey/troubleshooting/troubleshoot-connection-issues.md) --- # Connect to Aiven for Valkey™ with valkey-cli Learn how to establish a connection to an Aiven for Valkey™ service using the `valkey-cli`. ## Variables[​](#variables "Direct link to Variables") Replace the following placeholders in the code sample with actual values from your service overview page: | Variable | Description | | ------------- | ----------------------------------------------- | | `SERVICE_URI` | URI for the Aiven for Valkey service connection | ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Ensure that the valkey-cli client is installed. You can install it as part of the [Valkey server installation](https://valkey.io/topics/installation/). ## Setup and run[​](#setup-and-run "Direct link to Setup and run") Run the following command from a terminal window to connect: ``` valkey-cli -u SERVICE_URI ``` This command initiates a connection to the Aiven for Valkey service. To verify the connection, execute: ``` INFO ``` This command displays all server parameters, ensuring the connection is active: ``` # Server redis_version:7.2.4 server_name:valkey valkey_version:7.2.7 redis_git_sha1:c0b10003 redis_git_dirty:0 redis_build_id:e63142036e093656 redis_mode:standalone ... ``` To set a key, execute the following command: ``` SET mykey mykeyvalue123 ``` Following successful execution, a confirmation message `OK` appears. To retrieve the key value, execute the following command: ``` GET mykey ``` This will display the value of `mykey`, in this case,`"mykeyvalue123"`. --- # Controlled upgrade pipelines for your Aiven for Valkey™ service [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Link Aiven for Valkey™ services in an upgrade pipeline to test maintenance updates in a development or staging environment before they reach production. Control when your Aiven managed services receive maintenance updates and test maintenance updates in development or staging environments before they reach production. important Controlled upgrade pipeline is a [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-) feature. [Contact Aiven](https://aiven.io/contact) to request access. Aiven performs automatic service maintenance for security fixes, minor software updates, and other platform changes. The controlled upgrade pipeline feature lets you link services of the same type in an ordered sequence to control when each service receives updates. After a maintenance update upgrades a service at the initial pipeline step, you validate that service version before the update proceeds to the service at the next pipeline step. Validating means approving the new version as safe to roll out to the next service. Validation can be manual or automatic after a configurable delay. ## Why use controlled upgrade pipelines[​](#why-use-controlled-upgrade-pipelines "Direct link to Why use controlled upgrade pipelines") Controlled upgrade pipelines prevent production incidents caused by automatic updates reaching production before teams can test the new version in a lower environment. They give you full oversight of the update process: * **Risk mitigation**: Prevents unexpected maintenance updates from breaking your production environment by ensuring they are tested in a non-production setting first. * **Stability**: Keeps destination services (such as production) on a known-good version until you, or the automatic timer, confirm the new version is safe. * **Process control**: Allows platform teams to standardize their deployment and maintenance lifecycle across environments. ## About controlled upgrade pipelines[​](#about-controlled-upgrade-pipelines "Direct link to About controlled upgrade pipelines") ### Upgrade steps[​](#upgrade-steps "Direct link to Upgrade steps") An upgrade step is a pair of services linked by an upgrade constraint: * **Source service**: The service that receives maintenance updates first * **Destination service**: The service that waits for validation before receiving updates Each destination service can have only one source service. A source service can have multiple destination services. ### Upgrade pipelines[​](#upgrade-pipelines "Direct link to Upgrade pipelines") An upgrade pipeline is a chain of upgrade steps that spans multiple environments. For example: * Single chain: development → staging → production * Multiple destinations: development → production-eu and development → production-na ## How validation works[​](#how-validation-works "Direct link to How validation works") When a maintenance update upgrades your source service: 1. The source service receives the update first. 2. Test the updated source service to verify it works as expected. 3. Validate the update manually using the API or CLI, or wait for automatic validation after the configured delay. The default delay is 7 days. 4. After validation, the destination service becomes eligible for the same maintenance update. 5. The destination service receives the update during its next maintenance window. If one source service has multiple destination services, one validation for the source service applies to all connected destination services. ### Validation and maintenance windows[​](#validation-and-maintenance-windows "Direct link to Validation and maintenance windows") Validation and the maintenance window control different things: * **Validation** controls *what* version the destination service upgrades to. * The maintenance window controls *when* the upgrade happens. After you validate an update, or automatic validation applies, the destination service receives the validated version during its next scheduled maintenance window. Validation does not trigger an immediate upgrade outside the maintenance window. Upgrade pipelines add a constraint on what is installed during a maintenance update; they do not change when maintenance runs. Nodes in the destination service maintain the validated version until a newer version is validated, either when you validate it manually or when automatic validation applies after the configured delay. When a node is recycled, it uses the same validated version, not the latest available version. When you create a step, the destination service keeps the newest version that is already validated at that moment. If the destination service is already applying maintenance during step creation, the in-progress target version becomes the initial validated version. warning A powered-off source service cannot receive maintenance updates, so you cannot validate it. If you power off services earlier in the chain, the destination service upgrades regardless. For example, in a development → staging → production chain, if both development and staging are powered off, production upgrades without testing and validation in the earlier environments. Keep services in the chain powered on to preserve the protection that upgrade pipelines provide. ## Limitations and considerations[​](#limitations-and-considerations "Direct link to Limitations and considerations") * **Same service type**: You can only link services of the same type. For example, two Aiven for PostgreSQL services. * **Chain length**: The default maximum chain depth is 3 services, which is 2 steps. If you need a longer chain, [contact Aiven](https://aiven.io/contact). * **No cycles**: You cannot create circular dependencies between services. * **Emergency overrides**: Aiven can apply critical security or stability fixes to a destination service before explicit validation. * **Supported services**: This feature supports all Aiven service types except Aiven for Apache Flink® and Aiven for MySQL. * **Automatic maintenance updates only**: Pipelines apply to automatic maintenance updates, such as minor service version updates and node image updates. Major version upgrades, for example Aiven for PostgreSQL® 15 to 16, require manual action and are not promoted automatically through the pipeline. * **No permanent blocking**: You cannot prevent an update indefinitely. Automatic validation applies after the configured delay, up to the maximum delay. * **No validation rollback**: You cannot undo a validation after it is recorded. ## Use controlled upgrade pipelines[​](#use-controlled-upgrade-pipelines "Direct link to Use controlled upgrade pipelines") ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") To use controlled upgrade pipelines, you need the following: * The feature enabled by Aiven ([Limited availability](/docs/platform/concepts/service-and-feature-releases.md)) * Dev tool of your choice: * [Aiven CLI](/docs/tools/cli.md) Install the latest version of the Aiven CLI to access the `upgrade-pipeline` commands. * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) * Set `PROVIDER_AIVEN_ENABLE_BETA=true` before running Terraform. * See the [resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) for full schema, import format, and lifecycle behavior. * [Aiven Operator for Kubernetes](/docs/tools/kubernetes.md) Install the operator and create an Aiven token secret named `aiven-token` that the operator uses to authenticate against the Aiven API. * Write access to the source and destination projects * At least two services of the same type (for example, two Aiven for PostgreSQL® services) * Services can be in different projects in the same organization ### Set up an upgrade pipeline[​](#set-up-an-upgrade-pipeline "Direct link to Set up an upgrade pipeline") Use the Aiven CLI or API to create upgrade steps between your services. note The `upgrade-pipeline` CLI commands require Aiven CLI version 4.x or later. Command names and parameters may change before general availability. #### Create an upgrade step[​](#create-an-upgrade-step "Direct link to Create an upgrade step") Create a step to link a source service and a destination service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ [--source-project SOURCE_PROJECT] SOURCE_SERVICE \ [--destination-project DESTINATION_PROJECT] DESTINATION_SERVICE \ [--auto-validation-delay-days DAYS] ``` **Options** * `--organization-id` is required. * `--source-project` and `--destination-project` are optional. If you omit either project option, Aiven CLI uses the current default project set with `avn project switch`. * `--auto-validation-delay-days` is optional. Defaults to 7 days if not specified. ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "SOURCE_PROJECT_NAME", "source_service_name": "SOURCE_SERVICE_NAME", "destination_project_name": "DESTINATION_PROJECT_NAME", "destination_service_name": "DESTINATION_SERVICE_NAME", "auto_validation_delay_days": 7 }' ``` Use the [`aiven_upgrade_step`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) resource: ``` resource "aiven_upgrade_step" "example" { organization_id = "ORGANIZATION_ID" source_project_name = "SOURCE_PROJECT_NAME" source_service_name = "SOURCE_SERVICE_NAME" destination_project_name = "DESTINATION_PROJECT_NAME" destination_service_name = "DESTINATION_SERVICE_NAME" auto_validation_delay_days = 7 } ``` Apply an `UpgradePipelineStep` manifest with `kubectl`: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: upgrade-step-sample spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: SOURCE_PROJECT_NAME sourceServiceName: SOURCE_SERVICE_NAME destinationProjectName: DESTINATION_PROJECT_NAME destinationServiceName: DESTINATION_SERVICE_NAME autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable after the resource is created. Parameters: * `source_project_name`: Name of the project containing the source service * `source_service_name`: Name of the source service * `destination_project_name`: Name of the project containing the destination service * `destination_service_name`: Name of the destination service * `auto_validation_delay_days`: Optional. Number of days before automatic validation. The value must be at least `1`. The default is 7 days. The maximum delay you can configure is 30 days. #### List upgrade steps[​](#list-upgrade-steps "Direct link to List upgrade steps") View all upgrade steps you have access to: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step list --organization-id ORGANIZATION_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) To list managed upgrade steps, use: ``` terraform state list 'aiven_upgrade_step.*' terraform state show 'aiven_upgrade_step.example' ``` List `UpgradePipelineStep` resources in the current namespace: ``` kubectl get upgradepipelinesteps ``` #### View a specific step[​](#view-a-specific-step "Direct link to View a specific step") Get details about a specific upgrade step, including the last validation: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step get \ --organization-id ORGANIZATION_ID \ STEP_ID ``` ``` curl https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` terraform state show aiven_upgrade_step.example ``` Show the manifest and full status, including `id`, `conditions`, and `lastValidation`: ``` kubectl describe upgradepipelinestep RESOURCE_NAME kubectl get upgradepipelinestep RESOURCE_NAME -o yaml ``` The step details include `last_validation` values such as `validated_at`, `validated_by_user`, and `comment` when validation exists (available through the API). ### Validate an upgrade[​](#validate-an-upgrade "Direct link to Validate an upgrade") After testing your source service with the new update, validate the version to allow the destination service to receive the same update. #### Manual validation[​](#manual-validation "Direct link to Manual validation") Validate the current version of your source service: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step validate-for-service \ --project SOURCE_PROJECT \ SERVICE_NAME \ [--comment "COMMENT"] ``` `--comment` is optional. Use it to record a note about the validation, for example `"Tested and verified in development"`. ``` curl -X POST https://api.aiven.io/v1/project/SOURCE_PROJECT/service/SOURCE_SERVICE/upgrade-validation \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "comment": "Tested and verified in development" }' ``` Terraform manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. The operator manages upgrade steps, but validation is done through the API or CLI. Use the **CLI** or **API** tab to validate and optionally add a comment. #### Automatic validation[​](#automatic-validation "Direct link to Automatic validation") If you do not manually validate an update, the system automatically validates the source service version after the configured delay. Auto-validation starts from when the source service receives the update. ### Manage upgrade steps[​](#manage-upgrade-steps "Direct link to Manage upgrade steps") #### Update a step[​](#update-a-step "Direct link to Update a step") Modify the automatic validation delay for an existing step: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step update \ --organization-id ORGANIZATION_ID \ --auto-validation-delay-days 14 \ STEP_ID ``` ``` curl -X PATCH https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "auto_validation_delay_days": 14 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` resource "aiven_upgrade_step" "example" { # ...required fields... auto_validation_delay_days = 14 # Updated from 7 to 14 } ``` Apply the changes: ``` terraform plan terraform apply ``` Edit `autoValidationDelayDays` in your manifest and re-apply: ``` spec: autoValidationDelayDays: 14 ``` ``` kubectl apply -f upgrade-step.yaml ``` The `organizationId`, `sourceProjectName`, `sourceServiceName`, `destinationProjectName`, and `destinationServiceName` fields are immutable. To change them, delete the resource and create a new one. #### Delete a step[​](#delete-a-step "Direct link to Delete a step") Remove an upgrade step to allow the destination service to receive updates independently: * CLI * API * Terraform * Kubernetes ``` avn upgrade-pipeline step delete --organization-id ORGANIZATION_ID STEP_ID ``` Find `STEP_ID` from the upgrade step list command. ``` curl -X DELETE https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps/STEP_ID \ -H "Authorization: Bearer TOKEN" ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) Remove the resource from configuration and apply, or destroy it directly: ``` terraform apply terraform destroy -target=aiven_upgrade_step.example ``` Delete the `UpgradePipelineStep` resource: ``` kubectl delete upgradepipelinestep RESOURCE_NAME ``` Deleting a step removes all associated validations. ### Example: Three-environment pipeline[​](#example-three-environment-pipeline "Direct link to Example: Three-environment pipeline") Create a pipeline that promotes updates from development to staging to production: * CLI * API * Terraform * Kubernetes 1. Create a step from development to staging: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project dev-project \ --destination-project staging-project \ --auto-validation-delay-days 3 \ pg-dev pg-staging ``` 2. Create a step from staging to production: ``` avn upgrade-pipeline step create \ --organization-id ORGANIZATION_ID \ --source-project staging-project \ --destination-project prod-project \ --auto-validation-delay-days 7 \ pg-staging pg-prod ``` 1) Create a step from development to staging: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "dev-project", "source_service_name": "pg-dev", "destination_project_name": "staging-project", "destination_service_name": "pg-staging", "auto_validation_delay_days": 3 }' ``` 2) Create a step from staging to production: ``` curl -X POST https://api.aiven.io/v1/organization/ORGANIZATION_ID/upgrade-pipeline/steps \ -H "Authorization: Bearer TOKEN" \ -H "Content-Type: application/json" \ -d '{ "source_project_name": "staging-project", "source_service_name": "pg-staging", "destination_project_name": "prod-project", "destination_service_name": "pg-prod", "auto_validation_delay_days": 7 }' ``` **Reference**: [`aiven_upgrade_step` resource documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/upgrade_step) ``` # Step 1: Development → Staging resource "aiven_upgrade_step" "dev_to_staging" { organization_id = "ORGANIZATION_ID" source_project_name = "dev-project" source_service_name = "pg-dev" destination_project_name = "staging-project" destination_service_name = "pg-staging" auto_validation_delay_days = 3 } # Step 2: Staging → Production resource "aiven_upgrade_step" "staging_to_prod" { organization_id = "ORGANIZATION_ID" source_project_name = "staging-project" source_service_name = "pg-staging" destination_project_name = "prod-project" destination_service_name = "pg-prod" auto_validation_delay_days = 7 } ``` Apply the configuration: ``` export PROVIDER_AIVEN_ENABLE_BETA=true terraform init terraform plan terraform apply ``` Define both steps in a single manifest and apply it: ``` apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: dev-to-staging spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: dev-project sourceServiceName: pg-dev destinationProjectName: staging-project destinationServiceName: pg-staging autoValidationDelayDays: 3 --- apiVersion: aiven.io/v1alpha1 kind: UpgradePipelineStep metadata: name: staging-to-prod spec: authSecretRef: name: aiven-token key: token organizationId: ORGANIZATION_ID sourceProjectName: staging-project sourceServiceName: pg-staging destinationProjectName: prod-project destinationServiceName: pg-prod autoValidationDelayDays: 7 ``` ``` kubectl apply -f upgrade-pipeline.yaml ``` When a maintenance update arrives: 1. The development service receives the update. 2. After testing, validate the development version or wait 3 days for auto-validation. 3. The staging service receives the update during its next maintenance window. 4. After testing, validate the staging version or wait 7 days for auto-validation. 5. The production service receives the update during its next maintenance window. Related pages * [Maintenance and updates for your Aiven for Valkey™ service](/docs/products/valkey/howto/maintenance-updates.md) * [Change the service plan](/docs/products/valkey/howto/change-service-plan.md) * [Service and feature releases](/docs/platform/concepts/service-and-feature-releases.md) * [Aiven CLI](/docs/tools/cli.md) --- # Create read replica in Aiven for Valkey™ [Early availability](/docs/platform/concepts/service-and-feature-releases.md) [Aiven for Valkey™ read replica](/docs/products/valkey/concepts/read-replica.md) enables data replication from a primary to a replica service, improving performance and increasing redundancy for high availability and disaster recovery. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Aiven for Valkey service * Aiven API token * [Aiven CLI tool](https://github.com/aiven/aiven-client) ## Limitations[​](#limitations "Direct link to Limitations") * You can create a maximum of **5 read replicas** per primary service. * Read replicas are supported only on the **Startup** plan. ## Create a read replica[​](#create-a-read-replica "Direct link to Create a read replica") * Aiven Console * Aiven CLI * Aiven API 1. On the **Overview** page of your service, go to the **Read replica** section. 2. Click **Create replica**. 3. Enter a name for the replica. 4. Select a **Cloud**. 5. Select a **Plan**. 6. Click **Create**. The read replica is listed in your project services and can be accessed from the **Read replica** section on the service's **Overview** page. The service type is identified by chips labeled **Primary** or **Replica** at the top. To create a read replica using the Aiven CLI, use the following command with placeholders: ``` avn service create -t valkey -p startup-4 --project PROJECT_NAME \ --read-replica-for PRIMARY_SERVICE_NAME REPLICA_SERVICE_NAME ``` Parameters: * `PROJECT_NAME`: Your project name. * `CLOUD_NAME`: Cloud provider name. * `PRIMARY_SERVICE_NAME`: Primary service name. * `REPLICA_SERVICE_NAME`: Replica service name. To create a read replica for your Aiven for Valkey service via API, use the following command: ``` curl -X POST https://api.aiven.io/v1/project/PROJECT_NAME/service \ -H "Content-Type: application/json" \ -H "Authorization: bearer YOUR_AUTH_TOKEN" \ -d '{ "cloud": "CLOUD_NAME", "plan": "startup-4", "service_name": "REPLICA_SERVICE_NAME", "service_type": "valkey", "project": "PROJECT_NAME", "service_integrations": [ { "integration_type": "read_replica", "source_service_name": "PRIMARY_SERVICE_NAME" } ] }' ``` Parameters: * `PROJECT_NAME`: Your project name. * `YOUR_AUTH_TOKEN`: Your API authentication token. * `CLOUD_NAME`: Cloud provider name. * `REPLICA_SERVICE_NAME`: Replica service name. * `PRIMARY_SERVICE_NAME`: Primary service name. ## Promote a replica to a standalone primary[​](#promote-a-replica-to-a-standalone-primary "Direct link to Promote a replica to a standalone primary") You can promote your read replica to a primary service. After promotion, the replica becomes an independent primary service, disconnected from the original. * Aiven Console * Aiven CLI * Aiven API 1. Log in to the [Aiven Console](https://console.aiven.io/) and select your Aiven for Valkey service. 2. On the **Overview** page, scroll to the **Read replica** section. 3. Click the replica service to promote. 4. On the **Replica** service's **Overview** page, click **Promote to primary**. 5. In the confirmation dialog, click **Promote**. To promote the replica to primary using the Aiven CLI: 1. Retrieve the service integration ID with the `integration-list` command: ``` avn service integration-list --project ``` Parameters: * `--project `: Your project name. * ``: Your replica service name. 2. Delete the service integration with the following command: note Deleting the service integration breaks the replication link, promoting the replica to a standalone primary service. ``` avn service integration-delete --project PROJECT_NAME ``` Parameters: * `PROJECT_NAME`: Your project name. * ``: The integration ID obtained in the previous step. To promote the replica to primary using the Aiven API: 1. Get the service integration ID via an API call. ``` curl -s \ -H "Authorization: Bearer " \ "https://api.aiven.io/v1/project/PROJECT_NAME/service/REPLICA_SERVICE_NAME/integration" \ | jq -r '.service_integrations[] | select(.source_service_name=="PRIMARY_SERVICE_NAME").service_integration_id' ``` Parameters: * `Authorization: Bearer `: Your API authentication [token](/docs/platform/concepts/authentication-tokens.md). * `PROJECT_NAME`: Your project name. * `REPLICA_SERVICE_NAME`: Your replica service name. * `PRIMARY_SERVICE_NAME`: Your primary service name. * `service_integration_id`: Extracts the integration ID. 2. Delete the service integration using the obtained integration ID. note Deleting the service integration breaks the replication link, promoting the replica to a standalone primary service. ``` curl -X DELETE \ -H "Authorization: Bearer " \ "https://api.aiven.io/v1/project/PROJECT_NAME/integration/INTEGRATION_ID" ``` Parameters: * `YOUR_AUTH_TOKEN`: Your API authentication [token](/docs/platform/concepts/authentication-tokens.md). * `PROJECT_NAME`: Your project name. * `INTEGRATION_ID`: The integration ID obtained in the previous step. Related pages * [Aiven for Valkey read replica](/docs/products/valkey/concepts/read-replica.md) --- # Estimate the maximum number of connections for Aiven for Valkey™ The number of simultaneous connections for Aiven for Valkey™ depends on the total available memory on the server. You can use the following to estimate: ``` max_number_of_connections = 4 * m ``` where `m` represents the memory in megabytes. With at least 10,000 connections available, even on the smallest servers. For example, on a server with 4 GB memory (4,096 MB), the estimated simultaneous connections are: ``` 4 * 4096 = 16384 connections ``` note Make sure to convert the memory figure `m` to megabytes. This number is an estimate based on the available memory, so it varies between plans and cloud providers. To see the exact maximum connections allowed for your specific service, use the [valkey-cli](/docs/products/valkey/howto/connect-valkey-cli.md) with the `info` command as follows: ``` echo "info" | valkey-cli -u VALKEY_URI | grep maxclients ``` --- # Fork your Aiven for Valkey™ service Fork your Aiven for Valkey™ service to create an independent copy for testing, debugging, or development without affecting the original service. Fork an Aiven service to create a complete copy of it from its latest backup. Forked services are independent and don't share resources with or increase the load on the original service. Common use cases for forking include: * Creating a snapshot to analyze an issue. * Creating a development copy of your production environment. * Testing upgrades before applying them to production services. * Creating an instance in a different cloud provider, region, or with a different plan. * Renaming a service. During the forking process, the fork might initially have only one node while backups are being taken. The other nodes appear after the backup process is complete. When you fork a service, its configuration, data, and service users are copied to the new service. ## Limitations[​](#limitations "Direct link to Limitations") * You can only fork services that have at least one backup. * Service integrations are not copied to the fork. * Cross-project forking is supported only within the same organization. ## Fork a service[​](#fork-a-service "Direct link to Fork a service") * Console * CLI * API * Terraform 1. In your service, in the **Backups** section, click **Backup management**. 2. Click **Fork & restore**. 3. Choose the backup to fork from. 4. Enter a name, and select the cloud and plan. 5. Click **Create fork**. Use the [create service command](/docs/tools/cli/service-cli.md#avn-cli-service-create) with: * `--service-to-fork-from`: the name of the service to use as the source. * `--project-to-fork-from`: to fork a service in a different project, set this to the project name the source service is in. Use the [`ServiceCreate` endpoint](https://api.aiven.io/doc/#tag/Service/operation/ServiceCreate) and in the `user_config` property set: * `service_to_fork_from`: the name of the source service. * `project_to_fork_from`: to fork a service in a different project, set this to the name of the project the source service is in. Use the `service_to_fork_from` attribute in the user config of your service resource. To fork a service in a different project, set the `project_to_fork_from` attribute. More information on the service resources and their configuration options is available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Aiven for Valkey™ service backups](/docs/products/valkey/howto/configure-backups.md) * [Rename your Aiven for Valkey™ service](/docs/products/valkey/howto/rename-service.md) --- # Maintenance and updates for your Aiven for Valkey™ service Manage maintenance updates and set the maintenance window for your Aiven for Valkey™ service. ## Maintenance updates[​](#maintenance-updates "Direct link to Maintenance updates") Aiven applies some maintenance updates automatically. The following are the types of updates: * **Mandatory updates:** Security updates, quarterly patch releases, and platform updates that affect reliability or stability of the service nodes. * **Optional updates:** All other updates are initially optional. After six months, they become mandatory and are applied in the next week’s maintenance window. * **Periodic infrastructure updates:** Scheduled automatically for services with nodes active for 180 days and more. These updates are mandatory for all services, except those with maintenance turned off. Critical security updates are applied during the next available maintenance window. For other updates, Aiven gives you at least seven days' notice. Maintenance updates are also automatically applied during service upgrades. To view pending updates: * Console * CLI * API 1. In your service, click **Service settings**. 2. Go to the **Service management** section. Use the [`avn service get`](/docs/tools/cli/service-cli.md#avn_service_get) command. Use the [`service`](https://api.aiven.io/doc/#tag/Service/operation/ServiceGet) endpoint. ## Maintenance window[​](#maintenance-window "Direct link to Maintenance window") The maintenance window is the time period when Aiven can automatically apply maintenance updates to a service. When an update becomes available, Aiven schedules it for the next available maintenance window for each service. The update runs in the first window after it becomes available, and can begin any time after the start time. For example, if a service has a maintenance window of Monday 12:00 UTC, and an update becomes available on Tuesday, the update will be applied on the following Monday. During maintenance, Aiven might restart or replace service nodes. This can cause brief connection interruptions, but services are designed to minimize downtime. Aiven performs maintenance in a rolling-forward style, creating new nodes alongside existing ones and retiring the old nodes after the upgrade completes. Major service upgrades are triggered manually. A manually triggered upgrade starts immediately, regardless of the maintenance window. important You cannot control the order in which services are updated. Each service updates according to its own configured maintenance window, and there is no guaranteed way to control the update sequence. Manual updates and maintenance window adjustments only help for non-critical updates. ## Set the maintenance window[​](#set-the-maintenance-window "Direct link to Set the maintenance window") To set the maintenance window for your service: * Console * Terraform 1. In the Aiven Console, open your service. 2. In the **Maintenance** section, click **Actions** > **Change maintenance window**. 3. Set the day and time. 4. Click **Save changes**. Use the `maintenance_window_dow` and `maintenance_window_time` attributes in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). ## Certificate rotation[​](#certificate-rotation "Direct link to Certificate rotation") If your Aiven for Valkey™ service still uses the project CA certificate rather than the default browser-recognized certificate, Aiven periodically rotates this CA certificate. This rotation uses the same maintenance process described in [Maintenance updates](#maintenance-updates), applied during your service's maintenance window, to update your service to trust and use the new certificate. Update any client that trusts the project CA certificate to trust the new certificate before the rotation completes, otherwise your client can't verify the server certificate and the connection fails. For details on the certificate bundle and rotation process, see [TLS/SSL certificates](/docs/platform/concepts/tls-ssl-certificates.md#certificate-rotation). Related pages * [Version upgrades](/docs/products/valkey/howto/valkey-version-upgrade.md) * [Change the service plan](/docs/products/valkey/howto/change-service-plan.md) * [TLS/SSL certificates](/docs/platform/concepts/tls-ssl-certificates.md) * [Manage SSL connectivity in Aiven for Valkey™](/docs/products/valkey/howto/manage-ssl-connectivity.md) --- # Manage Aiven for Valkey™ service users Create and manage service users in your Aiven for Valkey™ service to control access to your data. Service users only exist in the scope of the Aiven service. They are unique to the service and not shared with any other services. Every service has a default `avnadmin` user with full access to the service. ## Add a service user[​](#add-a-service-user "Direct link to Add a service user") * Aiven Console * Aiven CLI * Aiven API * Terraform 1. In your service, **Users**. 2. Click **Add service user** or **Create user**. 3. Enter a name for your service user. 4. Set up all the other configuration options. If a password is required, a random password is generated automatically. You can change it later. 5. Click **Add service user**. Run the [avn service user-create](/docs/tools/cli/service/user.md#avn-service-user-create) command: ``` avn service user-create SERVICE_NAME --username USERNAME ``` Replace the following: * `SERVICE_NAME`: the name of your Aiven for Valkey service. * `USERNAME`: the name of the service user to create. Use the [ServiceUserCreate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUserCreate) endpoint: ``` curl --request POST \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME/user \ --header 'Authorization: Bearer YOUR_BEARER_TOKEN' \ --header 'content-type: application/json' \ --data '{"username": "USERNAME"}' ``` Replace the placeholders with your project name, service name, bearer token, and the username to create. Use the [`aiven_valkey_user` resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/valkey_user) to create and manage service users. To control what each user can access, configure [ACL permissions](/docs/products/valkey/howto/configure-acl-permissions.md). Related pages * [Configure ACL permissions](/docs/products/valkey/howto/configure-acl-permissions.md) * [Connect to your service](/docs/products/valkey/howto/connect-services.md) --- # Manage SSL connectivity in Aiven for Valkey™ Manage SSL connectivity for your Aiven for Valkey™ service by enabling secure connections and configuring stunnel for clients without SSL support. ## Client support for SSL-encrypted connections[​](#client-support-for-ssl-encrypted-connections "Direct link to Client support for SSL-encrypted connections") ### Default support[​](#default-support "Direct link to Default support") Aiven for Valkey uses SSL encrypted connections by default. This is indicated by the `valkeys://` (with double `s`) prefix in the `Service URI` on the [Aiven Console](https://console.aiven.io/). Aiven for Valkey services use a browser-recognized (Let's Encrypt) certificate by default, so no CA certificate download is required. Services created before this certificate mode was enabled still use the Aiven project CA certificate. There's no self-service option to use the project CA certificate for a service that uses a browser-recognized certificate. To request this, [open a support ticket](/docs/platform/howto/support.md). For details, see [TLS/SSL certificates](/docs/platform/concepts/tls-ssl-certificates.md#certificate-requirements). Valkey-compatible CLI tools support SSL connections. You can connect directly to your service using: ``` valkey-cli -u valkeys://username:password@host:port ``` Alternatively, you can use the third-party [Redli tool](https://github.com/IBM-Cloud/redli): ``` redli -u valkeys://username:password@host:port ``` Not every Redis client supports SSL-encrypted connections. In such cases, disabling or bypassing SSL is possible but **not recommended**. You can use one of the following options to achieve this. ## Set up `stunnel` process[​](#set-up-stunnel-process "Direct link to set-up-stunnel-process") Set up a `stunnel` process on the client to manage SSL settings on the database side while hiding it from the client. Use the following `stunnel` configuration, for example `stunnel.conf`, to set up a `stunnel` process. ``` client = yes foreground = yes debug = info delay = yes [redis] accept = 127.0.0.1:6380 connect = myredis.testproject.aivencloud.com:28173 TIMEOUTclose = 0 ; Only needed for services that use the Aiven project CA certificate. Most ; environments trust Let's Encrypt certificates by default, without a CAfile. ; CAfile = /path/to/optional/project/cacert/that/you/can/download/from/aiven/console ``` For details about the global options of the stunnel configuration, see the [Stunnel Global Options](https://www.stunnel.org/static/stunnel.html#GLOBAL-OPTIONS). More details about setting up such a process are available on the [Stunnel website page](https://www.stunnel.org/index.html). For `service-level option`, the following parameters are configured: * `accept => *[host:]port*`: Accept connections on the specified address. * `connect => *[host:]port*`: Connect to a remote address. * `TIMEOUTclose => *seconds*`: Time to wait for `close_notify`. note Adjust settings according to your service. On the **Overview** page, the **Connection information** section lists your Host and Port to configure the connection. HAProxy terminates SSL connections before forwarding them to Aiven for Valkey. The HAProxy connection timeout is set to 300 seconds (5 minutes) by default, matching the `valkey_timeout` value. This prevents idle connections from staying open too long. You can adjust the `valkey_timeout` value in the service's **Advanced configuration** section in the [Aiven Console](https://console.aiven.io). ## Allow plain-text connections[​](#allow-plain-text-connections "Direct link to Allow plain-text connections") As an alternative to SSL, you can enable plain-text connections. * Aiven Console * Aiven CLI 1. In the service's **Overview** page, click **Service settings**. 2. Go to the **Advanced configuration** section. 3. Set `valkey_ssl` to `false`. To disable SSL on an existing Aiven for Valkey service, use the following command, replacing `SERVICE_NAME `with your service name: ``` avn service update SERVICE_NAME -c "valkey_ssl=false" ``` warning Enabling plain-text connections compromises the security of your Aiven for Valkey service. Disabling SSL allows potential eavesdroppers to access sensitive credentials and data. After these changes, the `Service URI` will change to a new URL, starting with the `valkey://` prefix (removing the extra 's'). This indicates a direct, non-SSL connection to the Aiven for Valkey service. --- # Migrate Valkey™ databases to Aiven for Valkey™ Migrate your Valkey™ databases to Aiven for Valkey™ using the Aiven Console migration tool. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting the migration process, ensure the following: * Target [Aiven for Valkey service](/docs/products/valkey/get-started.md) that doesn't use [Valkey clustering](/docs/products/valkey/concepts/valkey-cluster.md), which isn't supported as a migration target * Source database details: * **Hostname or connection string**: The public hostname, connection string, or IP address used to connect to the database * **Port**: The port used to connect to the database * **Username**: The username with sufficient permissions to access the data * **Password**: The password used to connect to the database * Firewall rules updated or temporarily disabled to allow traffic between source and target databases * Source Valkey service secured with SSL * Publicly accessible source Valkey service or one with a VPC peering connection between private networks. You'll need the VPC ID and cloud name. note The migration does not include service user accounts or commands in progress. ## Database migration steps[​](#database-migration-steps "Direct link to Database migration steps") 1. Log in to the [Aiven Console](https://console.aiven.io/) and select the Aiven for Valkey service as your migration target. 2. Go to **Service settings** from the sidebar. 3. On the **Service settings** page, go to **Service management** > **Actions** > **Migrate database**. 4. Follow the wizard to guide you through the migration process. ### Step 1: Configure[​](#step-1-configure "Direct link to Step 1: Configure") In the migration wizard, review the prerequisites and click **Get started** to begin. ### Step 2: Validate[​](#step-2-validate "Direct link to Step 2: Validate") In the migration screen, enter the connection details for the migration source: * Hostname * Port * Username * Password Select **SSL encryption recommended** for a secure connection, and click **Run check**. The [Aiven Console](https://console.aiven.io/) validates the database configurations. If any errors occur, follow the on-screen instructions to resolve them and rerun the check. ### Step 3: Migrate[​](#step-3-migrate "Direct link to Step 3: Migrate") Once validation is complete, click **Start migration** to begin migrating data to Aiven for Valkey. ### Step 4: Replicate[​](#step-4-replicate "Direct link to Step 4: Replicate") While the migration is in progress: * You can close the migration wizard and monitor the progress later from the **Overview** page. * To stop the migration, click **Stop migration** in the migration progress window. Data already transferred to Aiven for Valkey is preserved. To prevent conflicts during replication: * Do not create or delete databases on the source service. * Avoid network or configuration changes that might disrupt the connection between source and target databases, such as firewall modifications. If the migration fails, resolve the issue and click **Start over**. ### Step 5: Close and complete the migration[​](#step-5-close-and-complete-the-migration "Direct link to Step 5: Close and complete the migration") After the migration, select one of the following: * **Stop replication**: If no further synchronization is needed, and you are ready to switch to Aiven for Valkey after testing. * **Keep replicating**: If continuous data synchronization is needed. Avoid system updates or configuration changes during active replication to prevent unintended migrations. note When replication is active, Aiven for Valkey ensures your data stays in sync by continuously synchronizing new writes from the source database. --- # Migrate from Aiven for Dragonfly® to Aiven for Valkey™ Migrate your data from Aiven for Dragonfly® to Aiven for Valkey™ using a manual key-by-key migration or a scan-based tool. Use this approach when the built-in console migration wizard is not available for your source service type. note The Aiven Console migration wizard supports migration **into** Dragonfly from Valkey or Caching, but not from Dragonfly to Valkey. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting: * A running [Aiven for Valkey](/docs/products/valkey/get-started.md) target service. * A running Aiven for Dragonfly source service. * [`valkey-cli`](https://valkey.io/topics/cli/) installed locally. * Connection details for both services—available on the **Overview** page of each service in the [Aiven Console](https://console.aiven.io/). tip To get connection details quickly from the CLI, run: ``` avn service get --project -v ``` ## Compatibility considerations[​](#compatibility-considerations "Direct link to Compatibility considerations") Dragonfly and Valkey are both Redis®-compatible, but differences exist in supported commands and data type behavior. Before migrating: * Review the [Dragonfly command compatibility list](https://www.dragonflydb.io/docs/command-reference/compatibility) and identify any commands or features your application uses that differ in Valkey. * Valkey does not support all Dragonfly-specific extensions. If your data uses `JSON`, `Search`, or `TimeSeries` modules, verify that the target Valkey service has equivalent modules enabled or plan a data transformation step. * Test your application against Valkey in a staging environment before cutting over production traffic. ## Step 1: Prepare for migration[​](#step-1-prepare-for-migration "Direct link to Step 1: Prepare for migration") Before migrating, review your data for keys or data types that may require transformation. ### Handle module-specific data types[​](#handle-module-specific-data-types "Direct link to Handle module-specific data types") If your Dragonfly service uses module data types (for example, `ReJSON-RL` for JSON keys), those keys cannot be loaded directly into a Valkey service that does not have the equivalent module enabled. Options: * **Enable the module on Valkey**: In the [Aiven Console](https://console.aiven.io/), go to your Valkey service **Service settings** > **Advanced configuration** and enable the required module before loading data. * **Transform the data**: Extract module-type values and re-insert them using native Valkey data types (for example, store JSON as a plain string with `SET`). * **Skip the keys**: If the data is not critical, exclude those keys from migration. ### Filter keys (optional)[​](#filter-keys-optional "Direct link to Filter keys (optional)") To exclude specific keys or key patterns from the migration, use a script to scan and pipe only the keys you need. See [Step 2](#step-2-migrate-data-to-aiven-for-valkey) for the scan-based approach. ## Step 2: Migrate data to Aiven for Valkey[​](#step-2-migrate-data-to-aiven-for-valkey "Direct link to Step 2: Migrate data to Aiven for Valkey") Choose one of the following methods based on your use case. ### Method A: Key-by-key migration using SCAN and DUMP/RESTORE[​](#method-a-key-by-key-migration-using-scan-and-dumprestore "Direct link to Method A: Key-by-key migration using SCAN and DUMP/RESTORE") Use this approach to migrate all or a subset of keys. Run the migration in batches so that each key isn't a separate network round trip. Issuing `PTTL`, `DUMP`, and `RESTORE` one key at a time is slow for large datasets. Instead, use a client that pipelines commands, such as the Python [`valkey`](https://pypi.org/project/valkey/) library, which batches reads from the source and writes to the target: ``` import valkey BATCH_SIZE = 500 src = valkey.Valkey( host="", port=, password="", ssl=True, ) dst = valkey.Valkey( host="", port=, password="", ssl=True, ) cursor = 0 while True: cursor, keys = src.scan(cursor=cursor, count=BATCH_SIZE) if keys: # Pipeline the PTTL and DUMP reads from the source. read = src.pipeline(transaction=False) for key in keys: read.pttl(key) read.dump(key) results = read.execute() # Pipeline the RESTORE writes to the target. write = dst.pipeline(transaction=False) for i, key in enumerate(keys): ttl, payload = results[2 * i], results[2 * i + 1] if payload is None: continue write.restore(key, ttl if ttl and ttl > 0 else 0, payload, replace=True) write.execute() if cursor == 0: break ``` Install the client with `pip install valkey` before running the script. Increase `BATCH_SIZE` to trade memory for throughput. note `DUMP` and `RESTORE` use a serialization format that is Redis-version-specific. If the Dragonfly and Valkey versions differ significantly, some keys may fail to restore. Test with a small key set first. ### Method B: Use RedisShake for bulk migration[​](#method-b-use-redisshake-for-bulk-migration "Direct link to Method B: Use RedisShake for bulk migration") [RedisShake](https://github.com/tair-opensource/RedisShake) is a vendor-neutral, actively maintained tool for migrating data between Redis-compatible services. Use the `scan_reader` mode, which iterates over all keys using SCAN commands. note Aiven for Dragonfly does not support the replication commands (`SYNC`, `PSYNC`) required by RedisShake's `sync_reader` mode. Always use `scan_reader` when migrating from Aiven for Dragonfly. 1. Download the latest RedisShake release for your platform from the [RedisShake releases page](https://github.com/tair-opensource/RedisShake/releases) and extract it. Alternatively, run it with Docker: ``` docker pull ghcr.io/tair-opensource/redisshake:latest ``` 2. Create a `shake.toml` configuration file that defines the Dragonfly source and the Valkey target. Both Aiven services require TLS: ``` [scan_reader] address = ":" password = "" tls = true count = 500 # keys per SCAN iteration; raise for throughput [redis_writer] address = ":" password = "" tls = true ``` 3. Start the migration: ``` ./redis-shake shake.toml ``` RedisShake scans all keys from Dragonfly and writes them to Valkey. It stops automatically when the scan is complete. note `scan_reader` performs a one-time bulk copy and does not stream ongoing writes. Stop writes to your Dragonfly service before starting the migration to avoid missing new data written during the scan. ## Step 3: Verify the migration[​](#step-3-verify-the-migration "Direct link to Step 3: Verify the migration") After loading data into Valkey, verify that the migration is complete: 1. Compare the key count on both services: ``` # On Dragonfly valkey-cli -h -p \ --tls --no-auth-warning -a \ DBSIZE # On Valkey valkey-cli -h -p \ --tls --no-auth-warning -a \ DBSIZE ``` 2. Spot-check a sample of keys to confirm values, types, and time-to-live (TTL) values transferred correctly: ``` valkey-cli -h -p \ --tls --no-auth-warning -a \ TYPE valkey-cli -h -p \ --tls --no-auth-warning -a \ TTL ``` 3. Run your application's integration tests against the Valkey service before switching production traffic. ## Step 4: Cut over to Aiven for Valkey[​](#step-4-cut-over-to-aiven-for-valkey "Direct link to Step 4: Cut over to Aiven for Valkey") When the data is verified and your application is tested: 1. Stop writes to the Dragonfly service, or route them to Valkey first. 2. If using Method A (SCAN/DUMP/RESTORE script), run it one final time to capture any remaining keys, then stop it. 3. Update your application connection strings to point to the Aiven for Valkey service. 4. Monitor your application for errors after cutover. tip Keep the Aiven for Dragonfly service running for a short period after cutover as a fallback. Power off or delete it once you are confident the migration is stable. Related pages * [Get started with Aiven for Valkey™](/docs/products/valkey/get-started.md) * [Migrate from Redis®\* to Aiven for Valkey™ via console](/docs/products/valkey/howto/migrate-redis-aiven-via-console.md) * [Migrate Valkey™ databases to Aiven for Valkey™](/docs/products/valkey/howto/migrate-caching-valkey-to-aiven-for-valkey.md) * [Aiven for Dragonfly® overview](/docs/products/dragonfly.md) --- # Migrate from Redis®\* to Aiven for Valkey™ using the CLI Move your data from a source, standalone Redis®\* data store to an Aiven-managed Valkey™ service. The migration process first attempts to use the `replication` method, and if it fails, it switches to `scan`. Create an Aiven for Valkey service and migrate data from AWS ElastiCache Redis. The Aiven project name is `test`, and the service name for the target Aiven for Valkey is `valkey`. Limitations * Migrating from **Google Cloud Memorystore** for Redis is not supported. * Migrating into a target service that uses [Valkey clustering](/docs/products/valkey/concepts/valkey-cluster.md) is not supported. * Source Redis version must be equal to or lower than: * Redis version 7.2 * Target Aiven for Valkey version ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * A target [Aiven for Valkey](/docs/products/valkey/get-started.md) service. * The hostname, port, and password of the source Redis service. * The source Redis service secured with SSL, which is the default for migration. * Publicly accessible source Redis service or a service with a VPC peering between the private networks. The migration process requires VPC ID and the cloud name. note AWS ElastiCache for Redis instances cannot have public IP addresses and require project VPC and peering connection. ## Create a service and perform the migration[​](#create-a-service-and-perform-the-migration "Direct link to Create a service and perform the migration") 1. Check the Aiven configuration options and connection details * To view Aiven configuration options, enter: ``` avn service types -v ... Service type 'valkey' options: ... Remove migration => --remove-option migration Hostname or IP address of the server where to migrate data from => -c migration.host= Password for authentication with the server where to migrate data from => -c migration.password= Port number of the server where to migrate data from => -c migration.port= The server where to migrate data from is secured with SSL => -c migration.ssl= (default=True) User name for authentication with the server where to migrate data from => -c migration.username= ``` * For the VPC information, enter: ``` avn vpc list --project test PROJECT_VPC_ID CLOUD_NAME ... ==================================== ============= 40ddf681-0e89-4bce-bd89-25e246047731 aws-eu-west-1 ``` note Note the hostname, port, and password of the source Redis service, as well as the VPC ID and cloud name. You need these details to complete the migration. 2. Create the Aiven for Valkey service and start the migration. If you do not have a service already, create one with: ``` avn service create --project test -t valkey -p hobbyist --cloud aws-eu-west-1 --project-vpc-id 40ddf681-0e89-4bce-bd89-25e246047731 -c migration.host="master.jappja-redis.kdrxxz.euw1.cache.amazonaws.com" -c migration.port=6379 -c migration.password= valkey ``` tip You can skip specifying the project-vpc-id and cloud if the source Redis server is publicly accessible. 3. Check the migration status: ``` avn service migration-status --project test valkey STATUS METHOD ERROR ====== ====== ===== done scan null ``` note Status can be one of `done`, `failed` or `running`. In case of failure, the error contains the error message: ``` avn service migration-status --project test valkey STATUS METHOD ERROR ====== ====== ================ failed scan invalid password ``` ## Migrate to an existing Aiven for Valkey service[​](#migrate-to-an-existing-aiven-for-valkey-service "Direct link to Migrate to an existing Aiven for Valkey service") To update an existing service, run: ``` avn service update --project test -c migration.host="master.jappja-redis.kdrxxz.euw1.cache.amazonaws.com" -c migration.port=6379 -c migration.password= redis ``` ## Remove migration from configuration[​](#remove-migration-from-configuration "Direct link to Remove migration from configuration") Migration is one-time operation. Once completed and the status is `done`, you cannot restart the same migration. To perform the migration again, first remove the existing configuration and reconfigure the settings to initiate a new migration: ``` avn service update --project test --remove-option migration valkey ``` --- # Migrate from Redis®\* to Aiven for Valkey™ using Aiven Console Migrate your Redis®\* databases, whether on-premise or cloud-hosted, to Aiven for Valkey™, using Aiven Console's guided wizard. Limitations * Migrating from **Google Cloud Memorystore** for Redis is not supported. * Migrating into a target service that uses [Valkey clustering](/docs/products/valkey/concepts/valkey-cluster.md) is not supported. * Source Redis version must be equal to or lower than: * Redis version 7.2 * Target Aiven for Valkey version ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Before starting the migration process, ensure you have the following: * A target Aiven for Valkey service. To create one, see [Get started with Aiven for Valkey](/docs/products/valkey/get-started.md). * Source database information: * **Hostname or connection string:** This is the public hostname, connection string, or IP address used to connect to the database. Refer to [accessible from the public Internet](/docs/platform/howto/public-access-in-vpc.md). * **Port:** The port used to connect to the database. * **Username:** The username used to connect to the database. Ensure this user has sufficient permissions to access the data to migrate. * **Password:** The password used to connect to the database. * To enable traffic and connection between the source and target databases, ensure that you update or disable the firewalls that protect them. If necessary, you can temporarily disable the firewalls. * A source Redis®\* service that is secured with SSL is a default migration requirement. * A publicly accessible source Redis®\* service or a service with a VPC peering connection between private networks. The VPC ID and cloud name are required for the migration process. note AWS ElastiCache for Redis®\* instances cannot have public IP addresses and require project VPC and peering connection. ## Migrate a Redis database[​](#migrate-a-redis-database "Direct link to Migrate a Redis database") To migrate a Redis database to Aiven for Valkey service: 1. Log in to the [Aiven Console](https://console.aiven.io/), and select the target Aiven for Valkey service for migrating the Redis® database. 2. Click **Service settings** on the sidebar. 3. Scroll to the **Service management** section, and click **Actions** > **Import database**. to initiate the import process. 4. Follow the wizard to guide you through the database migration process. ### Step 1: Configure[​](#step-1-configure "Direct link to Step 1: Configure") Read through the guidelines on the Redis migration wizard and click **Get started** to proceed with the database migration. ![Screenshot of the database migration wizard](/docs/assets/images/redis-db-migration-get-started-bc53202b88d9c1978b25aa4a06338896.png) ### Step 2: Validation[​](#step-2-validation "Direct link to Step 2: Validation") 1. On the **Database connection and validation** screen, enter the following information to establish a connection to your source database: * **Hostname:** This is the public hostname, connection string, or IP address used to connect to the database. * **Port:** The port used to connect to the database. * **Username:** The username used to connect to the database. * **Password:** The password used to connect to the database. 2. Select the **SSL encryption recommended** checkbox. 3. Click the **Run check** to validate the connection. If the check returns any warning, resolve the issues before proceeding with the migration process. ![Connect to database](/docs/assets/images/redis-migration-validation-a876b2c81ffbd97abd223b547cf4318b.png) ### Step 3: Migration[​](#step-3-migration "Direct link to Step 3: Migration") On the **Database migration** screen, click **Start Migration** to begin the migration. ![Start database migration](/docs/assets/images/redis-start-migration-cad0d8697fda279794abde94fafefdda.png) Live replication during migration While the migration is connected and running, changes in the source are continuously replicated to Aiven for Valkey. If a key is created, modified, or deleted in the source, the same change is reflected on the Aiven service. While the migration is in progress, you can * Close the wizard by clicking **Close window**. To check the migration status anytime,return to the wizard from the service's overview page. * Continue writing to the target database. * Stop the migration by clicking **Stop migration**. Data that has already been migrated is retained, and you can initiate a new migration at any time Stopping and starting over If you choose to stop the migration, this action immediately halts the replication of your data. However, any data that has already been migrated to Aiven is retained. You can initiate a new migration later, and this process overwrites any previously migrated databases. Migration attempt failed If so, investigate possible causes of the failure and address the identified issues. Once you have resolved the underlying problems, initiate the migration by clicking **Start over**. ### Step 4: Close[​](#step-4-close "Direct link to Step 4: Close") When the wizard informs you about the completion of the migration, you can choose one of the following options: * Click **Close connection** to disconnect the databases and stop the replication process if it is still active. * Click **Keep replicating** to keep the connection open and the replication mode active for continuous synchronization of any subsequent additions to the connected databases. ![Close database connection](/docs/assets/images/redis-migration-complete-35baa1e5b8c3902ef1a06436398d5f5f.png) Verify sync and plan cutover * **Initial sync completion**: The Valkey replica exposes `master_sync_in_progress`. When it is `0`, the initial sync (RDB load) is complete and the migration is considered done. * **Cutover when replication continues**: If the source is still receiving writes, perform the final sync verification to ensure all data is replicated: 1. Stop writes to the source. 2. Run `INFO REPLICATION` on both the source and the Aiven replica. 3. Compare the values of `master_repl_offset` on the source and the `slave_repl_offset` on the replica. If `slave_repl_offset` ≥ `master_repl_offset`, you can safely close the connection. Related pages * [Get started with Aiven for Valkey](/docs/products/valkey/get-started.md) --- # Power on/off and delete your Aiven for Valkey™ service Power off your Aiven for Valkey™ service to release resources and save credits, power it back on when you need it, or delete it permanently. ## Power off a service[​](#power-off-a-service "Direct link to Power off a service") When you power off a service: * All virtual machines are removed from the public cloud. * The service configuration is stored on the Aiven Platform. * If there are no backups, all service data is lost. * If the service has time-based or point in time recovery backups, the backups remain on the Aiven Platform. Services powered off for more than 180 days are automatically deleted. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power off service**. To power off a service, run: ``` avn service update SERVICE_NAME --power-off ``` ## Power on a service[​](#power-on-a-service "Direct link to Power on a service") When you power on a service: * New virtual machines are created on the service's public cloud. * The service starts with the stored configuration parameters. * The latest time-based backup is restored. * Maintenance updates are automatically applied. * If a point in time recovery backup is available, the database transaction logs are replayed to recover the service data to a specific point in time. The restoration takes from a few minutes to a few hours, depending on the network bandwidth, the disk IOPS allocated to the service, and the size of the backup. * Console * CLI 1. In your project, click **Services**. 2. Select the service to open the **Overview** page. 3. Click **Actions** > **Power on service**. To power on a service, run: ``` avn service update SERVICE_NAME --power-on ``` To see when the service is running, run: ``` avn service wait SERVICE_NAME ``` note Aiven for Valkey stores data in memory. When you power off the service, in-memory data is kept only if a backup is available, and it is restored from the latest backup when you power the service back on. note Static IP addresses are not removed when a service is powered off or deleted. They continue to generate the usual costs. To avoid these costs, [remove the static IP addresses](/docs/platform/concepts/static-ips.md). ## Delete a service[​](#delete-a-service "Direct link to Delete a service") * Console * CLI 1. In your project, click **Services**. 2. Open the service to delete, and click **Actions** > **Delete service**. To delete a service, run: ``` avn service terminate SERVICE_NAME ``` Related pages * [Aiven for Valkey™ service backups](/docs/products/valkey/howto/configure-backups.md) * [Fork Aiven for Valkey™](/docs/products/valkey/howto/fork-service.md) --- # Prepare your Aiven for Valkey™ service for high load Prepare your Aiven for Valkey™ service for higher than usual traffic to avoid outages and keep performance stable. Prepare your services for higher than usual traffic to avoid service outages by doing the following: * **Subscribe to service notifications:** To receive notifications about service health and warnings when resources are low, you can [set service and project contacts](https://aiven.io/docs/platform/howto/technical-emails). You can also view the status of the Aiven Platform and get updates on incidents on the [status page](https://status.aiven.io/). Follow the RSS feed, subscribe to email or SMS notifications, or use the Slack integration to get notifications about incidents. * **Monitor your services:** [Monitor the health of your services](/docs/platform/howto/list-monitoring.md) using metrics, logs, alerts, and dashboards. * **Scale your services:** If you forecast a load that can't be handled by the service, you can scale up your service. * **Set the backup schedule:** To minimize the impact of the higher load during the backup process, schedule backups outside of peak traffic hours. * **Set the maintenance window:** Schedule maintenance updates outside of your peak traffic hours. * **Run load tests on service forks:** To test the impact of high traffic on a production service, fork the service and run your load test on the fork. Additionally, optimizing a service allows it to perform better under stress therefore avoiding the need of an upgrade. The more optimized a service is for your usage, the better you can weather spikes in traffic. Related pages * [Change the service plan](/docs/products/valkey/howto/change-service-plan.md) --- # Rename your Aiven for Valkey™ service Change the name of your Aiven for Valkey™ service by forking it under a new name and deleting the original service. You cannot rename a service after creation. Instead, you can create a fork with the new name and delete the original service. ## Rename a service[​](#rename-a-service "Direct link to Rename a service") 1. Stop writing to the service. 2. Fork the service. 3. Add any integrations or SSO configurations that weren't copied. 4. Connect your clients to the new service. 5. Test the forked service. 6. Delete the original service. Related pages * [Fork Aiven for Valkey™](/docs/products/valkey/howto/fork-service.md) * [Power on/off and delete your Aiven for Valkey™ service](/docs/products/valkey/howto/power-cycle-service.md) --- # Tag your Aiven for Valkey™ service Add key-value tags to your Aiven for Valkey™ service to organize services and track ownership, cost allocation, and governance. Use tags to add metadata to Aiven services to categorize them or run custom logic on them. Typical uses include: * Tagging for governance to deploy services with specific tags only. * Tagging for internal cost reporting, ownership, allocation, and accountability. A tag is a key/value pair: * **Key**: A case-sensitive string that starts with a letter and consists of letters, numbers, dashes, and underscores. The maximum length for a key is 64 characters. * **Value**: A string value limited to 64 UTF-8 characters. Within a service, the tag keys must be unique. * Console * Terraform 1. In the service, click **Service settings**. 2. In the **Service status** section, click **Actions** > **Add service tags**. 3. Enter a key and value for each tag. 4. Click **Save changes**. Use the `tag` attribute in [your Aiven service resource](https://registry.terraform.io/providers/aiven/aiven/latest/docs). Related pages * [Fork Aiven for Valkey™](/docs/products/valkey/howto/fork-service.md) --- # Track restore progress for your Aiven for Valkey™ service Track the restore progress of individual nodes in your Aiven for Valkey™ service during node replacement, forking, or maintenance, using the Aiven API. You can track restore progress for individual nodes during service node replacement by using the Aiven API. For example, use this endpoint to monitor the restore progress of a forked service or when applying maintenance. The service object exposes restore progress under `node_states[].progress_updates`: * `service.node_states[]` contains per-node state entries. * When a node is restoring or catching up, its `state` is typically `syncing_data`. * When the state is `syncing_data`, the node may include `progress_updates` with one or more phase objects. * Other node states don't include restore progress data. note `progress_updates` may be missing or empty even when a node is in `syncing_data`. This can occur when a restore completes before detailed progress is reported or when the service does not emit detailed progress counters. ## API endpoints[​](#api-endpoints "Direct link to API endpoints") Restore progress fields are part of the standard service response payload. * Get a single service (recommended for polling): `GET /project/{project}/service/{service_name}` * List services in a project: `GET /project/{project}/service` - Request - Response ``` curl -H "Authorization: aivenv1 API_TOKEN" https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME ``` Replace the following placeholders: * `API_TOKEN`: Your Aiven API token. * `PROJECT_NAME`: Your Aiven project name. * `SERVICE_NAME`: The name of your service. ``` { "service": { ... "node_states": [ { "node_name": "...", "state": "syncing_data", "progress_updates": [ { "phase": "basebackup", "completed": false, "current": 3410567, "min": 0, "max": 7569280, "unit": "bytes_uncompressed" } ] } ], ... } } ``` ## Node states[​](#node-states "Direct link to Node states") Common values for `node_states[].state` include: * `setting_up_vm`: The virtual machine is being created or initialized. * `syncing_data`: The node is restoring data or catching up. * `running`: The node is operating normally. * `leaving`: The node is leaving the cluster. * `unknown`: A transient or error state. ## `progress_updates` data model[​](#progress_updates-data-model "Direct link to progress_updates-data-model") `progress_updates` is a list of phase objects. When present, phases appear in the following order: 1. `prepare` 2. `basebackup` 3. `stream` 4. `finalize` Each phase object includes the following fields: ``` { "completed": false, "current": 3410567, "max": 7569280, "min": 0, "phase": "basebackup", "unit": "bytes_uncompressed" } ``` ### Field semantics[​](#field-semantics "Direct link to Field semantics") * `phase`: String, required. The restore phase. Possible values: `prepare`, `basebackup`, `stream`, and `finalize`. * `completed`: Boolean, required. Whether the phase is complete. * `current`: Number or null, optional. The current progress value. This field can be missing or null. * `min`: Number or null, optional. The starting value for the phase. This field can be missing or null. * `max`: Number or null, optional. The expected total value for the phase. This value can be missing, null, or change while the restore is in progress. * `unit`: String or null, optional. The unit for `current`, `min`, and `max`. New unit values can appear over time. Important considerations * Treat `unit` as an opaque identifier. Unknown values can appear. * `max` may change while a restore is in progress. * Not all phases report numeric counters. Some services only indicate phase completion. ## Why `max` values can change[​](#why-max-values-can-change "Direct link to why-max-values-can-change") The `current`, `min`, and `max` values are best-effort progress indicators. They can be based on estimates or on system state that changes over time. Treat `max` as the latest known expected total, not as a fixed guarantee. Common reasons `max` can change include: * The restore process discovers additional work after it starts, such as files, segments, or objects that become visible only after metadata is read. * New data is added on the backend while the node is catching up, which moves the completion point forward. This is common during incremental catch-up phases. * Progress is calculated from system state, such as replication lag, rather than from a fixed work queue. As the system state changes, the value is recalculated. * The service switches restore strategies during the operation, for example from snapshot restore to replication catch-up, which changes what the counters represent. As a result: * Phase percentage can decrease even when the restore operates normally. * Remaining-time estimates based on `max` are unreliable. * Sudden changes in `max` are expected unless the node remains in `syncing_data` longer than expected. ## Restore phase meanings[​](#restore-phase-meanings "Direct link to Restore phase meanings") Phase names are standardized, but the underlying work and the meaning of the counters are service-specific. * `prepare`: Prepares the node for restore. * `basebackup`: Restores the full backup. * `stream`: Applies incremental changes, such as replication or log replay. * `finalize`: Completes final steps before serving traffic. Not all restores include every phase. ## Compute phase progress percentages[​](#compute-phase-progress-percentages "Direct link to Compute phase progress percentages") You cannot reliably compute overall restore progress. You can compute a phase-specific progress percentage when `min`, `max`, and `current` are present and `max != min`. ``` pct = round(((current - min) / (max - min)) * 100, 1) ``` When handling progress values: * If any of `min`, `max`, or `current` is null or missing, display `n/a`. * If `max == min`, treat the percentage as undefined. * Expect the percentage to decrease when `max` changes. * Clamp displayed values to the range `[0, 100]`. ## Polling guidance[​](#polling-guidance "Direct link to Polling guidance") Progress updates are best-effort and refresh every 10 seconds while a node is in `syncing_data`. Poll the service state every 10 to 30 seconds. More frequent polling does not provide additional detail. For each `node_states[]` entry: * If `state` is not `syncing_data`, no restore progress is available. * If `state` is `syncing_data`: * If `progress_updates` is missing or empty, the node is restoring without detailed progress data. * Otherwise, the current phase is the last phase where `completed` is `false`. Stop polling when all nodes reach the `running` state or when a stall is detected. ### Stall detection[​](#stall-detection "Direct link to Stall detection") The API does not provide per-phase timestamps. To detect stalls, use a time-based threshold, such as a node remaining in `syncing_data` longer than expected. Do not rely on counters or `max` values to estimate remaining time. Related pages * [Backups](/docs/products/valkey/howto/configure-backups.md) * [Fork your service](/docs/products/valkey/howto/fork-service.md) --- # Manage Aiven for Valkey™ versions Aiven for Valkey™ supports multiple versions of Valkey running concurrently in the platform. Choose a version that best fits your needs and upgrade your service when ready. ## Supported Valkey versions[​](#supported-valkey-versions "Direct link to Supported Valkey versions") From version 9.0, Aiven for Valkey supports two major upstream Valkey versions at a time. These are the two latest major versions that are stable on the Aiven Platform. You can select either version when you create a service or upgrade an existing service. If you do not select a version, the default is the latest stable version on the Aiven Platform. See the supported versions in the [Aiven for Valkey version reference](/docs/platform/reference/eol-for-major-versions.md#aiven-for-valkey). ## Before you upgrade[​](#before-you-upgrade "Direct link to Before you upgrade") ### Available or upcoming upgrades[​](#available-or-upcoming-upgrades "Direct link to Available or upcoming upgrades") Track upgrades for your service via: * [Aiven Console](https://console.aiven.io): Service **Overview** page > **Maintenance** section > List of available mandatory and optional upgrades * Email: notifications for automated upgrades ### Downgrade restriction[​](#downgrade-restriction "Direct link to Downgrade restriction") Downgrading to a previous version is not supported due to potential data format incompatibilities. Always test upgrades in a non-production environment first. To revert to a previous version: 1. Create a service with the desired version. 2. Restore data from a backup taken before the upgrade. 3. Update your application connection strings. ### Prerequisites for upgrade[​](#prerequisites-for-upgrade "Direct link to Prerequisites for upgrade") Before upgrading your service: * **Test in development**: Test the upgrade in a development environment first. * **Backup your data**: Ensure you have recent backups. Backups are automatic, but verify they exist. To upgrade your service version, check that: * Your Aiven for Valkey service is running. * Target version to upgrade to is [available for manual upgrade](/docs/products/valkey/howto/valkey-version-upgrade.md#available-or-upcoming-upgrades). * You can use one of the following tools to upgrade: * [Aiven Console](https://console.aiven.io/) * [Aiven CLI](/docs/tools/cli.md) * [Aiven API](/docs/tools/api.md) * [Aiven Provider for Terraform](/docs/tools/terraform.md) * [Aiven Operator for Kubernetes®](/docs/tools/kubernetes.md) ## Upgrade your service[​](#upgrade-your-service "Direct link to Upgrade your service") * Console * CLI * API * Terraform * Kubernetes 1. In the [Aiven Console](https://console.aiven.io/), go to your Valkey service. 2. On the **Overview** page, go to the **Maintenance** section. 3. Click **Actions** > **Upgrade version**. 4. Select a version to upgrade to. warning When you click **Upgrade**: * The system applies the upgrade immediately. * You cannot downgrade the service to a previous version. 5. Click **Upgrade**. Upgrade the service version using the [avn service update](https://aiven.io/docs/tools/cli/service-cli#avn-cli-service-update) command: ``` avn service update SERVICE_NAME -c valkey_version="N.N" ``` Parameters: * `SERVICE_NAME`: Name of your service * `N.N`: Target service version to upgrade to, for example `9.0` Call the [ServiceUpdate](https://api.aiven.io/doc/#tag/Service/operation/ServiceUpdate) endpoint to set `valkey_version`: ``` curl --request PUT \ --url https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME \ --header 'Authorization: Bearer BEARER_TOKEN' \ --header 'Content-Type: application/json' \ --data '{ "user_config": { "valkey_version": "N.N" } }' ``` Parameters: * `PROJECT_NAME`: Name of your project * `SERVICE_NAME`: Name of your service * `BEARER_TOKEN`: Your API authentication token * `N.N`: Target service version to upgrade to, for example `9.0` Use the [`aiven_valkey`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/valkey) resource to set [`valkey_version`](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/valkey#valkey_version-1): ``` resource "aiven_valkey" "example" { project = var.PROJECT_NAME cloud_name = "CLOUD_NAME" plan = "PLAN_NAME" service_name = "SERVICE_NAME" valkey_user_config { valkey_version = "N.N" } } ``` Parameters: * `PROJECT_NAME`: Name of your project * `CLOUD_NAME`: Cloud region identifier * `PLAN_NAME`: Service plan * `SERVICE_NAME`: Name of your service * `N.N`: Target service version to upgrade to, for example `9.0` Use the [Valkey](https://aiven.github.io/aiven-operator/resources/valkey.html) resource to set [`valkey_version`](https://aiven.github.io/aiven-operator/resources/valkey.html#spec.userConfig.valkey_version-property): ``` apiVersion: aiven.io/v1alpha1 kind: Valkey metadata: name: SERVICE_NAME spec: authSecretRef: name: aiven-token key: token connInfoSecretTarget: name: valkey-connection project: PROJECT_NAME cloudName: CLOUD_NAME plan: PLAN_NAME userConfig: valkey_version: "N.N" ``` Apply the updated configuration: ``` kubectl apply -f valkey-service.yaml ``` Parameters: * `PROJECT_NAME`: Name of your project * `SERVICE_NAME`: Name of your service * `CLOUD_NAME`: Cloud region identifier * `PLAN_NAME`: Service plan * `N.N`: Target service version to upgrade to, for example `9.0` ## Version selection for new services[​](#version-selection-for-new-services "Direct link to Version selection for new services") When creating an Aiven for Valkey service: * **Default version**: The latest stable version on the Aiven Platform is the default version. * **Explicit selection**: You can specify a version using the `valkey_version` parameter. * **Version availability**: Only versions in `available` state can be selected. Example (CLI): ``` avn service create SERVICE_NAME \ --service-type valkey \ --plan PLAN_NAME \ --cloud CLOUD_NAME \ -c valkey_version="N.N" ``` Parameters: * `SERVICE_NAME`: Name of your service * `PLAN_NAME`: Service plan * `CLOUD_NAME`: Cloud region identifier * `N.N`: Service version, for example `8.1` Related pages --- # Maintenance and lifecycle in Aiven for Valkey™ Keep your Aiven for Valkey™ service current by upgrading versions and applying maintenance updates. Related pages * [Upgrade the Valkey version](/docs/products/valkey/howto/valkey-version-upgrade.md) --- # Advanced parameters for Aiven for Valkey™ See the configuration options available for Aiven for Valkey™: | Parameter | | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | []()[**additional\_backup\_regions**](#additional_backup_regions)`array`Additional Cloud Regions for Backup Replication | | []()[**backup\_hour**](#backup_hour)`integer,null`- max: `23`The hour of day (in UTC) when backup for the service is started. New backup is only started if previous backup has already completed. | | []()[**backup\_minute**](#backup_minute)`integer,null`- max: `59`The minute of an hour when backup for the service is started. New backup is only started if previous backup has already completed. | | []()[**ip\_filter**](#ip_filter)`array`- default: `0.0.0.0/0,::/0`IP filterAllow incoming connections from CIDR address block, e.g. '10.20.0.0/16' | | []()[**service\_log**](#service_log)`boolean,null`Service loggingStore logs for the service so that they are available in the HTTP API and console. | | []()[**static\_ips**](#static_ips)`boolean`Use static public IP addresses | | []()[**migration**](#migration)`object,null`Migrate data from existing servermigration.host string Hostname or IP address of the server where to migrate data from migration.port integer - min: 1 - max: 65535 Port number of the server where to migrate data from migration.password string Password for authentication with the server where to migrate data from migration.ssl boolean - default: true The server where to migrate data from is secured with SSL migration.username string User name for authentication with the server where to migrate data from migration.dbname string Database name for bootstrapping the initial connection migration.ignore\_dbs string Comma-separated list of databases, which should be ignored during migration (supported by MySQL and PostgreSQL only at the moment) migration.ignore\_roles string Comma-separated list of database roles, which should be ignored during migration (supported by PostgreSQL only at the moment) migration.method string The migration method to be used (currently supported only by Redis, Dragonfly, MySQL and PostgreSQL service types) | | []()[**private\_access**](#private_access)`object`Allow access to selected service ports from private networksprivate\_access.prometheus boolean Allow clients to connect to prometheus with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations private\_access.valkey boolean Allow clients to connect to valkey with a DNS name that always resolves to the service's private IP addresses. Only available in certain network locations | | []()[**privatelink\_access**](#privatelink_access)`object`Allow access to selected service components through Privatelinkprivatelink\_access.prometheus boolean Enable prometheus privatelink\_access.valkey boolean Enable valkey | | []()[**public\_access**](#public_access)`object`Allow access to selected service ports from the public Internetpublic\_access.prometheus boolean Allow clients to connect to prometheus from the public internet for service nodes that are in a project VPC or another type of private network public\_access.valkey boolean Allow clients to connect to valkey from the public internet for service nodes that are in a project VPC or another type of private network | | []()[**recovery\_basebackup\_name**](#recovery_basebackup_name)`string`Name of the basebackup to restore in forked service | | []()[**valkey\_maxmemory\_policy**](#valkey_maxmemory_policy)`string,null`- default: `noeviction`Valkey maxmemory-policy | | []()[**valkey\_pubsub\_client\_output\_buffer\_limit**](#valkey_pubsub_client_output_buffer_limit)`integer`- min: `32`
- max: `262144`Set output buffer limit for pub / sub clients in MB. The value is the hard limit, the soft limit is 1/4 of the hard limit. When setting the limit, be mindful of the available memory in the selected service plan. | | []()[**valkey\_number\_of\_databases**](#valkey_number_of_databases)`integer`- min: `1`
- max: `128`Set number of Valkey databases. Changing this will cause a restart of the Valkey service. | | []()[**valkey\_io\_threads**](#valkey_io_threads)`integer`- min: `1`
- max: `256`Set Valkey IO thread count. Changing this will cause a restart of the Valkey service. | | []()[**valkey\_lfu\_log\_factor**](#valkey_lfu_log_factor)`integer`- max: `100`
- default: `10`Counter logarithm factor for volatile-lfu and allkeys-lfu maxmemory-policies | | []()[**valkey\_lfu\_decay\_time**](#valkey_lfu_decay_time)`integer`- min: `1`
- max: `120`
- default: `1`LFU maxmemory-policy counter decay time in minutes | | []()[**valkey\_ssl**](#valkey_ssl)`boolean`- default: `true`Require SSL to access Valkey | | []()[**valkey\_timeout**](#valkey_timeout)`integer`- max: `2073600`
- default: `300`Valkey idle connection timeout in seconds | | []()[**valkey\_notify\_keyspace\_events**](#valkey_notify_keyspace_events)`string`Set notify-keyspace-events option | | []()[**valkey\_persistence**](#valkey_persistence)`string`When persistence is 'rdb', Valkey does RDB dumps each 10 minutes if any key is changed. Also RDB dumps are done according to backup schedule for backup purposes. When persistence is 'off', no RDB dumps and backups are done, so data can be lost at any moment if service is restarted for any reason, or if service is powered off. Also service can't be forked. | | []()[**frequent\_snapshots**](#frequent_snapshots)`boolean`- default: `true`When enabled, Valkey will create frequent local RDB snapshots. When disabled, Valkey will only take RDB snapshots when a backup is created, based on the backup schedule. This setting is ignored when `valkey_persistence` is set to `off`. | | []()[**valkey\_active\_expire\_effort**](#valkey_active_expire_effort)`integer`- min: `1`
- max: `10`
- default: `1`Active expire effortValkey reclaims expired keys both when accessed and in the background. The background process scans for expired keys to free memory. Increasing the active-expire-effort setting (default 1, max 10) uses more CPU to reclaim expired keys faster, reducing memory usage but potentially increasing latency. | | []()[**valkey\_activedefrag**](#valkey_activedefrag)`boolean`Enable active memory defragmentation. When enabled, Valkey relocates objects off sparsely-used memory pages to reduce fragmentation and return memory to the operating system. Defragmentation runs on the main thread and consumes CPU, so it may increase latency under load. | | []()[**valkey\_active\_defrag\_ignore\_bytes**](#valkey_active_defrag_ignore_bytes)`integer`- min: `1048576`
- max: `1073741824`
- default: `104857600`Active defrag minimum fragmentation wasteMinimum amount of fragmentation waste, in bytes, before active defragmentation starts. Only takes effect when `valkey_activedefrag` is enabled. | | []()[**valkey\_active\_defrag\_threshold\_lower**](#valkey_active_defrag_threshold_lower)`integer`- min: `1`
- max: `100`
- default: `10`Minimum percentage of fragmentation before active defragmentation starts. Only takes effect when `valkey_activedefrag` is enabled. | | []()[**valkey\_acl\_channels\_default**](#valkey_acl_channels_default)`string`Default ACL for pub/sub channels used when a Valkey user is createdDetermines default pub/sub channels' ACL for new users if ACL is not supplied. When this option is not defined, all\_channels is assumed to keep backward compatibility. This option doesn't affect Valkey configuration acl-pubsub-default. | | []()[**valkey\_version**](#valkey_version)`string,null`Valkey major version | | []()[**service\_to\_fork\_from**](#service_to_fork_from)`string,null`Name of another service to fork from. This has effect only when a new service is being created. | | []()[**project\_to\_fork\_from**](#project_to_fork_from)`string,null`Name of another project to fork a service from. This has effect only when a new service is being created. | | []()[**enable\_ipv6**](#enable_ipv6)`boolean`Enable IPv6Register AAAA DNS records for the service, and allow IPv6 packets to service ports | --- # Restricted commands in Aiven for Valkey™ For optimal performance, stability, and security, Aiven for Valkey™ disables specific commands. Commands you cannot use in Aiven for Valkey™ are the following: * `bgrewriteaof`: Initiates a background append-only file rewrite. * `cluster`: Manages Valkey cluster commands. * `command`: Provides details about all Valkey commands. * `debug`: Contains sub-commands for debugging Valkey. * `failover`: Manages manual failover of a master to a replica. * `migrate`: Atomically transfers a key from a Valkey instance to another one. * `slaveof`: Makes the server a replica of another instance, or promotes it as master. * `acl`: Manages Valkey Access Control Lists. * `bgsave`: Creates a snapshot of the dataset into a dump file. * `config`: Alters the configuration of a running Valkey server. * `lastsave`: Returns the UNIX timestamp of the last successful save to disk. * `monitor`: Streams back every command processed by the Valkey server. * `replicaof`: Makes the server a replica of another instance. * `save`: Synchronously saves the dataset to disk. * `shutdown`: Synchronously saves the dataset to disk and shuts down the server. --- # Aiven for Valkey™ metrics available via Prometheus Monitor and optimize your Aiven for Valkey™ service with metrics available via Prometheus. These metrics help track cluster health, replication status, and overall performance. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Enable Prometheus integration](/docs/platform/howto/integrations/prometheus-metrics.md). * Note the Prometheus **username** and **password** in the **Integration endpoints** section of the [Aiven Console](https://console.aiven.io/). ## Access Prometheus metrics[​](#access-prometheus-metrics "Direct link to Access Prometheus metrics") * Browser * CLI 1. Open your service's **Overview** page in the [Aiven Console](https://console.aiven.io/). 2. In the **Connection information** section, click the **Prometheus** tab. 3. Copy the **Service URI**. 4. Paste the Service URI into your browser's address bar. 5. When prompted, enter your Prometheus credentials. 6. Click **Login**. To retrieve metrics, run the following `curl` command: ``` curl --user 'USERNAME:PASSWORD' PROMETHEUS_URL/metrics ``` Replace `USERNAME:PASSWORD` with your Prometheus credentials and `PROMETHEUS_URL` with the Service URI from the **Connection information** section. ## Host metrics[​](#host-metrics "Direct link to Host metrics") Host metrics provide insights into system-level performance, including CPU, memory, disk, and network usage. ### CPU utilization[​](#cpu-utilization "Direct link to CPU utilization") CPU utilization metrics offer insights into CPU usage. These metrics include time spent on different processes, system load, and overall uptime. | Metric | Description | | ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | | `cpu_usage_guest` | CPU time spent running a virtual CPU for guest operating systems | | `cpu_usage_guest_nice` | CPU time running low-priority virtual CPUs for guest operating systems; interrupted by higher-priority tasks and measured in hundredths of a second | | `cpu_usage_idle` | Time the CPU spends doing nothing | | `cpu_usage_iowait` | Time waiting for I/O to complete | | `cpu_usage_irq` | Time servicing interrupts | | `cpu_usage_nice` | Time running user-niced processes | | `cpu_usage_softirq` | Time servicing softirqs | | `cpu_usage_steal` | Time spent in other operating systems when running in a virtualized environment | | `cpu_usage_system` | Time spent running system processes | | `cpu_usage_user` | Time spent running user processes | | `system_load1` | System load average for the last minute | | `system_load15` | System load average for the last 15 minutes | | `system_load5` | System load average for the last 5 minutes | | `system_n_cpus` | Number of CPU cores available | | `system_n_users` | Number of users logged in | | `system_uptime` | Time for which the system has been up and running | ### Disk space utilization[​](#disk-space-utilization "Direct link to Disk space utilization") Disk space utilization metrics provide a snapshot of disk usage. These metrics include information about free and used disk space, as well as `inode` usage and total disk capacity. | Metric | Description | | ------------------- | ----------------------------- | | `disk_free` | Amount of free disk space | | `disk_inodes_free` | Number of free inodes | | `disk_inodes_total` | Total number of inodes | | `disk_inodes_used` | Number of used inodes | | `disk_total` | Total disk space | | `disk_used` | Amount of used disk space | | `disk_used_percent` | Percentage of disk space used | ### Disk input and output[​](#disk-input-and-output "Direct link to Disk input and output") Metrics such as `diskio_io_time` and `diskio_iops_in_progress` provide insights into disk I/O operations. These metrics cover read/write operations, the duration of these operations, and the number of bytes read/written. | Metric | Description | | ------------------------- | ------------------------------------------------------------------------------- | | `diskio_io_time` | Total time spent on I/O operations | | `diskio_iops_in_progress` | Number of I/O operations currently in progress | | `diskio_merged_reads` | Number of read operations that were merged | | `diskio_merged_writes` | Number of write operations that were merged | | `diskio_read_bytes` | Total bytes read from disk | | `diskio_read_time` | Total time spent on read operations | | `diskio_reads` | Total number of read operations | | `diskio_weighted_io_time` | Weighted time spent on I/O operations, considering their duration and intensity | | `diskio_write_bytes` | Total bytes written to disk | | `diskio_write_time` | Total time spent on write operations | | `diskio_writes` | Total number of write operations | ### Generic memory[​](#generic-memory "Direct link to Generic memory") The following metrics, including `mem_active` and `mem_available`, provide insights into your system's memory usage. | Metric | Description | | ----------------------- | ------------------------------------------------------------- | | `mem_active` | Amount of actively used memory | | `mem_available` | Amount of available memory | | `mem_available_percent` | Percentage of available memory | | `mem_buffered` | Amount of memory used for buffering I/O | | `mem_cached` | Amount of memory used for caching | | `mem_commit_limit` | Maximum amount of memory that can be committed | | `mem_committed_as` | Total amount of committed memory | | `mem_dirty` | Amount of memory waiting to be written to disk | | `mem_free` | Amount of free memory | | `mem_high_free` | Amount of free memory in the high memory zone | | `mem_high_total` | Total amount of memory in the high memory zone | | `mem_huge_pages_free` | Number of free huge pages | | `mem_huge_page_size` | Size of huge pages | | `mem_huge_pages_total` | Total number of huge pages | | `mem_inactive` | Amount of inactive memory | | `mem_low_free` | Amount of free memory in the low memory zone | | `mem_low_total` | Total amount of memory in the low memory zone | | `mem_mapped` | Amount of memory mapped into the process's address space | | `mem_page_tables` | Amount of memory used by page tables | | `mem_shared` | Amount of memory shared between processes | | `mem_slab` | Amount of memory used by the kernel for data structure caches | | `mem_swap_cached` | Amount of swap memory cached | | `mem_swap_free` | Amount of free swap memory | | `mem_swap_total` | Total amount of swap memory | | `mem_total` | Total amount of memory | | `mem_used` | Amount of used memory | | `mem_used_percent` | Percentage of used memory | | `mem_vmalloc_chunk` | Largest contiguous block of vmalloc memory available | | `mem_vmalloc_total` | Total amount of vmalloc memory | | `mem_vmalloc_used` | Amount of used vmalloc memory | | `mem_wired` | Amount of wired memory | | `mem_write_back` | Amount of memory being written back to disk | | `mem_write_back_tmp` | Amount of temporary memory being written back to disk | ### Network[​](#network "Direct link to Network") The following metrics, including `net_bytes_recv` and `net_packets_sent`, provide insights into your system's network operations. | Metric | Description | | ----------------------------- | ----------------------------------------------------------------- | | `net_bytes_recv` | Total bytes received on the network interfaces | | `net_bytes_sent` | Total bytes sent on the network interfaces | | `net_drop_in` | Incoming packets dropped | | `net_drop_out` | Outgoing packets dropped | | `net_err_in` | Incoming packets with errors | | `net_err_out` | Outgoing packets with errors | | `net_icmp_inaddrmaskreps` | Number of ICMP address mask replies received | | `net_icmp_inaddrmasks` | Number of ICMP address mask requests received | | `net_icmp_incsumerrors` | Number of ICMP checksum errors | | `net_icmp_indestunreachs` | Number of ICMP destination unreachable messages received | | `net_icmp_inechoreps` | Number of ICMP echo replies received | | `net_icmp_inechos` | Number of ICMP echo requests received | | `net_icmp_inerrors` | Number of ICMP messages received with errors | | `net_icmp_inmsgs` | Total number of ICMP messages received | | `net_icmp_inparmprobs` | Number of ICMP parameter problem messages received | | `net_icmp_inredirects` | Number of ICMP redirect messages received | | `net_icmp_insrcquenchs` | Number of ICMP source quench messages received | | `net_icmp_intimeexcds` | Number of ICMP time exceeded messages received | | `net_icmp_intimestampreps` | Number of ICMP timestamp reply messages received | | `net_icmp_intimestamps` | Number of ICMP timestamp request messages received | | `net_icmpmsg_intype3` | Number of ICMP type 3 (destination unreachable) messages received | | `net_icmpmsg_intype8` | Number of ICMP type 8 (echo request) messages received | | `net_icmpmsg_outtype0` | Number of ICMP type 0 (echo reply) messages sent | | `net_icmpmsg_outtype3` | Number of ICMP type 3 (destination unreachable) messages sent | | `net_icmp_outaddrmaskreps` | Number of ICMP address mask reply messages sent | | `net_icmp_outaddrmasks` | Number of ICMP address mask request messages sent | | `net_icmp_outdestunreachs` | Number of ICMP destination unreachable messages sent | | `net_icmp_outechoreps` | Number of ICMP echo reply messages sent | | `net_icmp_outechos` | Number of ICMP echo request messages sent | | `net_icmp_outerrors` | Number of ICMP messages sent with errors | | `net_icmp_outmsgs` | Total number of ICMP messages sent | | `net_icmp_outparmprobs` | Number of ICMP parameter problem messages sent | | `net_icmp_outredirects` | Number of ICMP redirect messages sent | | `net_icmp_outsrcquenchs` | Number of ICMP source quench messages sent | | `net_icmp_outtimeexcds` | Number of ICMP time exceeded messages sent | | `net_icmp_outtimestampreps` | Number of ICMP timestamp reply messages sent | | `net_icmp_outtimestamps` | Number of ICMP timestamp request messages sent | | `net_icmp_outratelimitglobal` | Number of globally rate-limited ICMP messages sent | | `net_icmp_outratelimithost` | Number of ICMP messages rate-limited per host | | `net_ip_defaultttl` | Default time-to-live for IP packets | | `net_ip_forwarding` | Indicates if IP forwarding is enabled | | `net_ip_forwdatagrams` | Number of forwarded IP datagrams | | `net_ip_fragcreates` | Number of IP fragments created | | `net_ip_fragfails` | Number of failed IP fragmentations | | `net_ip_fragoks` | Number of successful IP fragmentations | | `net_ip_inaddrerrors` | Number of incoming IP packets with address errors | | `net_ip_indelivers` | Number of incoming IP packets delivered to higher layers | | `net_ip_indiscards` | Number of incoming IP packets discarded | | `net_ip_inhdrerrors` | Number of incoming IP packets with header errors | | `net_ip_inreceives` | Total number of incoming IP packets received | | `net_ip_inunknownprotos` | Number of incoming IP packets with unknown protocols | | `net_ip_outdiscards` | Number of outgoing IP packets discarded | | `net_ip_outnoroutes` | Number of outgoing IP packets with no route available | | `net_ip_outrequests` | Total number of outgoing IP packets requested to be sent | | `net_ip_outtransmits` | Number of IP packets transmitted successfully | | `net_ip_reasmfails` | Number of failed IP reassembly attempts | | `net_ip_reasmoks` | Number of successful IP reassembly attempts | | `net_ip_reasmreqds` | Number of IP fragments received needing reassembly | | `net_ip_reasmtimeout` | Number of IP reassembly timeouts | | `net_packets_recv` | Total number of packets received on the network interfaces | | `net_packets_sent` | Total number of packets sent on the network interfaces | | `netstat_tcp_close` | Number of TCP connections in the CLOSE state | | `netstat_tcp_close_wait` | Number of TCP connections in the CLOSE\_WAIT state | | `netstat_tcp_closing` | Number of TCP connections in the CLOSING state | | `netstat_tcp_established` | Number of TCP connections in the ESTABLISHED state | | `netstat_tcp_fin_wait1` | Number of TCP connections in the FIN\_WAIT\_1 state | | `netstat_tcp_fin_wait2` | Number of TCP connections in the FIN\_WAIT\_2 state | | `netstat_tcp_last_ack` | Number of TCP connections in the LAST\_ACK state | | `netstat_tcp_listen` | Number of TCP connections in the LISTEN state | | `netstat_tcp_none` | Number of TCP connections in the NONE state | | `netstat_tcp_syn_recv` | Number of TCP connections in the SYN\_RECV state | | `netstat_tcp_syn_sent` | Number of TCP connections in the SYN\_SENT state | | `netstat_tcp_time_wait` | Number of TCP connections in the TIME\_WAIT state | | `netstat_udp_socket` | Number of UDP sockets | | `net_tcp_activeopens` | Number of active TCP open connections | | `net_tcp_attemptfails` | Number of failed TCP connection attempts | | `net_tcp_currestab` | Number of currently established TCP connections | | `net_tcp_estabresets` | Number of established TCP connections reset | | `net_tcp_incsumerrors` | Number of TCP checksum errors in incoming packets | | `net_tcp_inerrs` | Number of incoming TCP packets with errors | | `net_tcp_insegs` | Number of TCP segments received | | `net_tcp_maxconn` | Maximum number of TCP connections supported | | `net_tcp_outrsts` | Number of TCP reset packets sent | | `net_tcp_outsegs` | Number of TCP segments sent | | `net_tcp_passiveopens` | Number of passive TCP open connections | | `net_tcp_retranssegs` | Number of TCP segments retransmitted | | `net_tcp_rtoalgorithm` | TCP retransmission timeout algorithm | | `net_tcp_rtomax` | Maximum TCP retransmission timeout | | `net_tcp_rtomin` | Minimum TCP retransmission timeout | | `net_udp_ignoredmulti` | Number of UDP multicast packets ignored | | `net_udp_incsumerrors` | Number of UDP checksum errors in incoming packets | | `net_udp_indatagrams` | Number of UDP datagrams received | | `net_udp_inerrors` | Number of incoming UDP packets with errors | | `net_udp_memerrors` | Number of UDP packets dropped due to memory errors | | `net_udplite_ignoredmulti` | Number of UDP-Lite multicast packets ignored | | `net_udplite_incsumerrors` | Number of UDP-Lite checksum errors in incoming packets | | `net_udplite_indatagrams` | Number of UDP-Lite datagrams received | | `net_udplite_inerrors` | Number of incoming UDP-Lite packets with errors | | `net_udplite_memerrors` | Number of UDP-L | ### Kernel[​](#kernel "Direct link to Kernel") The metrics listed below, such as `kernel_boot_time` and `kernel_context_switches`, provide insights into the operations of your system's kernel. | Metric | Description | | ------------------------- | ----------------------------------------------------------- | | `kernel_boot_time` | Time at which the system was last booted | | `kernel_context_switches` | Number of context switches that have occurred in the kernel | | `kernel_entropy_avail` | Amount of available entropy in the kernel's entropy pool | | `kernel_interrupts` | Number of interrupts that have occurred | | `kernel_processes_forked` | Number of processes that have been forked | ### Process[​](#process "Direct link to Process") Metrics such as `processes_running` and `processes_zombies` provide insights into the management of the system's processes. | Metric | Description | | ------------------------- | ------------------------------------------------------------------------ | | `processes_blocked` | Number of processes that are blocked | | `processes_dead` | Number of processes that have terminated | | `processes_idle` | Number of processes that are idle | | `processes_paging` | Number of processes that are paging | | `processes_running` | Number of processes currently running | | `processes_sleeping` | Number of processes that are sleeping | | `processes_stopped` | Number of processes that are stopped | | `processes_total` | Total number of processes | | `processes_total_threads` | Total number of threads across all processes | | `processes_unknown` | Number of processes in an unknown state | | `processes_zombies` | Number of zombie processes (terminated but not reaped by parent process) | ### Swap usage[​](#swap-usage "Direct link to Swap usage") Metrics such as `swap_free` and `swap_used` provide insights into the usage of the system's swap memory. | Metric | Description | | ------------------- | ----------------------------------- | | `swap_free` | Amount of free swap memory | | `swap_in` | Amount of data swapped in from disk | | `swap_out` | Amount of data swapped out to disk | | `swap_total` | Total amount of swap memory | | `swap_used` | Amount of used swap memory | | `swap_used_percent` | Percentage of swap memory used | ## Valkey-specific metrics[​](#valkey-specific-metrics "Direct link to Valkey-specific metrics") [Valkey-specific metrics](https://github.com/influxdata/telegraf/blob/master/plugins/inputs/redis/README.md#metrics) provide insights into the performance and health of your Aiven for Valkey service. --- # Supported Valkey™ modules Aiven for Valkey™ includes pre-enabled modules that extend core Valkey functionality with additional data types and commands. ## Valkey Bloom[​](#valkey-bloom "Direct link to Valkey Bloom") Valkey Bloom adds a probabilistic Bloom filter data type to Valkey. A Bloom filter is a space-efficient data structure that tests set membership: it can tell you that an element is *possibly in a set* or *definitely not in a set*, using a fraction of the memory needed to store the elements themselves. Valkey Bloom is available on Aiven for Valkey services running Valkey 9 and later. ### Configuration[​](#configuration "Direct link to Configuration") * Valkey Bloom is enabled by default. * No configuration is required. * Valkey Bloom cannot be disabled. ### Capabilities[​](#capabilities "Direct link to Capabilities") Valkey Bloom allows you to: * Create Bloom filters with a custom capacity and false positive rate * Add single or multiple elements to a filter * Check whether one or more elements might exist in a filter * Scale filters automatically as more elements are added * Persist filters alongside the rest of your data Bloom filters suit high-traffic workloads where you need to prevent cache penetration, deduplicate streams, or filter out known items at low memory cost. You interact with them through the `BF.*` command family, such as `BF.ADD`, `BF.EXISTS`, `BF.MADD`, `BF.MEXISTS`, and `BF.RESERVE`. For complete documentation on Valkey Bloom commands and usage, see the [Valkey Bloom documentation](https://valkey.io/topics/bloomfilters/). ## Valkey JSON[​](#valkey-json "Direct link to Valkey JSON") Valkey JSON provides native JSON document storage and manipulation capabilities within Valkey. ### Configuration[​](#configuration-1 "Direct link to Configuration") * Valkey JSON is enabled by default. * No configuration is required. * Valkey JSON cannot be disabled. ### Capabilities[​](#capabilities-1 "Direct link to Capabilities") Valkey JSON allows you to: * Store JSON documents as values * Query JSON documents using JSONPath * Make atomic updates to JSON elements * Index and search JSON data For complete documentation on Valkey JSON commands and usage, see the [Valkey JSON documentation](https://valkey.io/topics/valkey-json/). ## Valkey Search[​](#valkey-search "Direct link to Valkey Search") Valkey Search is a high-performance search engine module that supports vector search, full-text search, numeric filtering, and tag filtering. It enables indexing data stored in Valkey Hash or Valkey JSON data types and querying it with low latency. Valkey Search is available on Aiven for Valkey services running Valkey 9 and later. ### Configuration[​](#configuration-2 "Direct link to Configuration") * Valkey Search is enabled by default. * No configuration is required. * Valkey Search cannot be disabled. ### Capabilities[​](#capabilities-2 "Direct link to Capabilities") Valkey Search allows you to: * Create indexes over Valkey Hash and Valkey JSON data * Run vector similarity searches using Approximate Nearest Neighbor (HNSW) or exact K-Nearest Neighbor (KNN) algorithms * Apply numeric, tag, and full-text filters in hybrid queries * Aggregate search results using `FT.AGGREGATE` * Monitor index status and backfill progress using `FT.INFO` ### Supported commands[​](#supported-commands "Direct link to Supported commands") The following commands are available with Valkey Search: | Command | Description | | -------------- | ----------------------------------- | | `FT.CREATE` | Create an index | | `FT.DROPINDEX` | Delete an index | | `FT.INFO` | Return index details and statistics | | `FT._LIST` | List all indexes | | `FT.SEARCH` | Search an index | | `FT.AGGREGATE` | Aggregate search results | For the full command reference, see [Valkey Search commands](https://valkey.io/commands/#search). ### Cluster mode[​](#cluster-mode "Direct link to Cluster mode") Valkey Search is supported on Aiven for Valkey cluster plans. ### Further reading[​](#further-reading "Direct link to Further reading") For more information about Valkey Search, see the following upstream resources: * [Valkey Search overview](https://valkey.io/topics/search/) * [Search query syntax](https://valkey.io/topics/search-query/) * [Search expressions](https://valkey.io/topics/search-expressions/) * [Search data formats](https://valkey.io/topics/search-data-formats/) * [Search monitoring](https://valkey.io/topics/search-monitoring/) Related pages * [Valkey data types](https://valkey.io/topics/data-types/) * [Valkey commands reference](https://valkey.io/commands/) --- # Aiven for Valkey™ version lifecycle Learn how Aiven manages Aiven for Valkey™ version support, end of life (EOL) dates, and what happens to your service after a version reaches EOL. ## Aiven version support and upstream EOL[​](#aiven-version-support-and-upstream-eol "Direct link to Aiven version support and upstream EOL") Aiven aims to follow the EOL schedule set by the original authors and maintainers of the open source software (the upstream projects). Once the upstream project retires a specific version, they do not receive security updates and critical bug fixes anymore by the maintainers. Outdated services don't offer the level of protection you need, so Aiven follows the upstream project's EOL schedule to ensure that Aiven services are always running on supported versions. ## Service version numbering[​](#service-version-numbering "Direct link to Service version numbering") Aiven services inherit the upstream project's software versioning scheme. Depending on the service, a major version can be either a single digit or in the format `major.minor`. The exact version of the service is visible in the [Aiven Console](https://console.aiven.io/) when the service is running. ## Service version EOL policy[​](#service-version-eol-policy "Direct link to Service version EOL policy") Aiven sets an EOL date for each major version of the service. This policy covers both running and powered-off services on affected versions. ## EOL notifications[​](#eol-notifications "Direct link to EOL notifications") When Aiven sets the EOL date for a service major version: * You receive an email notification along with instructions on the next steps. * The [Aiven Console](https://console.aiven.io/) shows an EOL alert for affected services. * You receive email reminders monthly. * In the month of the EOL date, you receive weekly reminders. ## EOL best practices[​](#eol-best-practices "Direct link to EOL best practices") * Use service forking to test the version upgrade before upgrading your production services. * Upgrade to the supported version before the EOL date. This gives you time to test compatibility, resolve any issues, and plan the upgrade on your schedule. After the EOL date: * If the service is powered on, it's automatically upgraded to the latest version when possible, or to another supported version. note If it's not possible to upgrade a powered-on service to a supported version, the service is powered off and ultimately deleted. * If the service is powered off, it's deleted. ## Version EOL dates[​](#version-eol-dates "Direct link to Version EOL dates") | Version | Aiven EOL | Service creation supported until | Service creation supported from | | ------- | --------------- | -------------------------------- | ------------------------------- | | 8.1.x | To be announced | To be announced | 2025-11-18 | | 9.0.x | 2026-08-31 | 2026-08-31 | 2026-03-09 | | 9.1.x | To be announced | To be announced | 2026-07-15 | Related pages * [Manage Aiven for Valkey™ versions](/docs/products/valkey/howto/valkey-version-upgrade.md) --- # Scaling and performance in Aiven for Valkey™ Scale your Aiven for Valkey™ service vertically or horizontally, and tune performance and memory to keep it running efficiently as it grows. * **Vertical scaling**: Change your service plan to move to a node with more memory and compute. * **Horizontal scaling**: Use [clustering](/docs/products/valkey/concepts/valkey-cluster.md) to shard your data across multiple nodes, increasing capacity and throughput beyond what a single node can provide. Related pages * [Aiven for Valkey clustering](/docs/products/valkey/concepts/valkey-cluster.md) * [Memory usage](/docs/products/valkey/concepts/memory-usage.md) * [Overcommit memory warning](/docs/products/valkey/troubleshooting/warning-overcommit_memory.md) --- # Troubleshoot Aiven for Valkey™ connection issues Learn troubleshooting techniques for your Aiven for Valkey™ service and resolve common connection issues. important By default, [Aiven for Valkey uses SSL connections](/docs/products/valkey/howto/manage-ssl-connectivity.md), and these connections are closed automatically after 12 hours. This is not a parameter that can be changed. Aiven also sets the `valkey_timeout` advanced parameter to 300 seconds by default. ## Some Valkey connections are closed intermittently[​](#some-valkey-connections-are-closed-intermittently "Direct link to Some Valkey connections are closed intermittently") When experiencing connection issues with your Aiven for Valkey service, some common things to check: * Some Redis®\* clients do not support SSL connections. It is recommended to check the documentation for the Redis®\* client being used to ensure SSL connections are supported. * If you notice older connections terminating, check the value configured for the [`valkey_timeout` advanced parameter](/docs/products/valkey/reference/advanced-params.md). This parameter controls the timeout value for idle connections. Once the timeout is reached, the connection is terminated. ## Methods for troubleshooting connections[​](#methods-for-troubleshooting-connections "Direct link to Methods for troubleshooting connections") A great way to troubleshoot connection issues is to arrange for a packet capture to take place. This can be achieved with tools like [Tcpdump](https://www.tcpdump.org/) and [Wireshark](https://www.wireshark.org/). This allows you to see if connections are making it outside your network to the Aiven for Valkey instance. Another tool you can use to help diagnose connection issues is the Socket Statistics CLI tool which dumps socket statistics. --- # Handle the overcommit memory warning When starting an Aiven for Valkey™ service in the [Aiven Console](https://console.aiven.io/), you may notice on **Logs** the following **warning** `overcommit_memory`: ``` # WARNING overcommit_memory is set to 0! Background save may fail under low memory condition. To fix this issue add 'vm.overcommit_memory = 1' to /etc/sysctl.conf and reboot or run the command 'sysctl vm.overcommit_memory=1' for this to take effect. ``` This warning can be safely ignored as Aiven for Valkey ensures that the available memory never drops low enough to hit this particular failure case. --- # Aiven dev tools You can interact with the Aiven platform with various interfaces and tools that best suit your workflow. [Aiven Terraform Provider](/docs/tools/terraform.md) [Automate infrastructure provisioning and management on the Aiven Platform.](/docs/tools/terraform.md) [Aiven Kubernetes Operator](/docs/tools/kubernetes.md) [Create and manage Aiven services directly within your Kubernetes clusters.](/docs/tools/kubernetes.md) [Aiven API](/docs/tools/api.md) [Programmatically interact with and manage your Aiven infrastructure.](/docs/tools/api.md) [Aiven CLI](/docs/tools/cli.md) [Manage your Aiven services through the command-line interface.](/docs/tools/cli.md) [SQL query optimizer](/docs/tools/query-optimizer.md) [Use AI to optimize your queries.](/docs/tools/query-optimizer.md) [Aiven MCP](/docs/tools/mcp-server.md) [Manage Aiven services and access documentation from AI-powered coding assistants.](/docs/tools/mcp-server.md) [Managed Agents](/docs/tools/agents.md) [Create and run agents on the Aiven Platform.](/docs/tools/agents.md) --- # Managed Agents [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Managed Agents lets you create and run AI agents on the Aiven Platform. Agents can use connected tools to gather information, investigate issues, and perform tasks across Aiven and other systems. You define what an agent does, choose the AI model it uses, and give it access to the tools it needs. You can interact with an agent in a chat or configure scheduled tasks to run automatically. Managed Agents is built on open-source agent technology. Aiven manages the infrastructure required to run your agents, so you don't need to deploy or maintain the underlying infrastructure. note Managed Agents is in [limited availability](/docs/platform/concepts/service-and-feature-releases.md#limited-availability-). You need access for each project. In the project, click **Agents** > **Request access**. After you have access, click **Enable agents**. To run Managed Agents in a project VPC, click **Enable in VPC**. ## Example uses[​](#example-uses "Direct link to Example uses") Use agents for tasks that involve gathering information, analyzing it, and taking actions through connected tools. Run these tasks on demand or [on a schedule](/docs/tools/agents/schedule-agent.md). For example, you can create an agent to: * Review Aiven for PostgreSQL® logs, summarize the findings, and send an update to Slack. * Check consumer lag in Aiven for Apache Kafka® and create a Jira issue when the lag requires action. * Investigate an incident using information from multiple systems and summarize the findings. * Run recurring operational checks and report the results to your team. What an agent can do depends on its system instructions, AI model, and available tools and integrations. ## How Managed Agents works[​](#how-managed-agents-works "Direct link to How Managed Agents works") You configure an agent with: * **System instructions** that define what the agent does and how it behaves. * **An AI model** that processes the agent's instructions and requests. * **Tools and integrations** that give the agent access to information and external systems. * **Schedules** that let the agent perform recurring tasks automatically. You can also start a chat with an agent to give it a task or ask follow-up questions. ## Tools and integrations[​](#tools-and-integrations "Direct link to Tools and integrations") Agents can use built-in tools such as Web Fetch and Web Search. You can also connect an agent to Aiven through [Aiven MCP](/docs/tools/mcp-server.md) or to other MCP integrations, such as Slack, GitHub, and Jira. You choose which tools each agent can use. When you connect Aiven MCP, you grant access to services in the current project and assign an MCP role. You can grant any role up to your own. Aiven creates a scoped token automatically. ## Run agents interactively or on a schedule[​](#run-agents-interactively-or-on-a-schedule "Direct link to Run agents interactively or on a schedule") You can use an agent in two ways: * [Chat with an agent](/docs/tools/agents/chat-with-agent.md) to send requests when needed. * [Schedule an agent](/docs/tools/agents/schedule-agent.md) to run tasks automatically at a specified time or interval. For example, you can use chat to investigate an incident as it happens, or create a daily schedule that asks the agent to summarize service health. ## Next steps[​](#next-steps "Direct link to Next steps") * [Create an agent](/docs/tools/agents/create-agent.md) * [Chat with an agent](/docs/tools/agents/chat-with-agent.md) * [Schedule an agent](/docs/tools/agents/schedule-agent.md) * [Manage an agent](/docs/tools/agents/manage-agent.md) * [Manage integrations](/docs/tools/agents/manage-integrations.md) Related pages * [AI tools on Aiven](/docs/ai-features.md) * [Aiven MCP](/docs/tools/mcp-server.md) --- # Chat with an agent [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Use chat to send tasks or questions to an agent and continue the conversation with follow-up messages. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") An existing agent. To create one, see [Create an agent](/docs/tools/agents/create-agent.md). ## Start a chat[​](#start-a-chat "Direct link to Start a chat") 1. In the Aiven Console, open your project. 2. Click **Agents**. 3. Click the agent. The agent opens in a new chat. 4. Enter a message and send it. You can continue the conversation by sending additional messages. To start a separate conversation, click **New chat**. ## View chat history[​](#view-chat-history "Direct link to View chat history") 1. Click **Chat history**. 2. Click a conversation to open it. Use chat when you want to interact with the agent on demand. To automate recurring tasks, see [Schedule an agent](/docs/tools/agents/schedule-agent.md). Related pages * [Managed Agents](/docs/tools/agents.md) * [Create an agent](/docs/tools/agents/create-agent.md) * [Schedule an agent](/docs/tools/agents/schedule-agent.md) --- # Create an agent [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Create an agent by describing a task or by configuring the agent manually. When you describe a task, Aiven generates a configuration that you can review and test before you create the agent. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Managed Agents enabled for the project. If you have not requested access or enabled Managed Agents, see [Managed Agents](/docs/tools/agents.md). ## Create an agent by describing a task[​](#create-an-agent-by-describing-a-task "Direct link to Create an agent by describing a task") 1. In the Aiven Console, open your project. 2. Click **Agents** > **Create agent**. ### Describe the task[​](#describe-the-task "Direct link to Describe the task") 1. In **Describe what you want this agent to do**, enter the task. For example: > Check the health of the PostgreSQL services in this project. Summarize any issues you find and highlight anything that needs attention. To start from a suggestion, select an option under **Or start from a suggestion**, such as **Slow-query digest**, **Incident RCA**, **PR reviewer**, **Sprint status sync**, or **Deep researcher**. Some suggestions use other MCP integrations, such as Slack or Jira. 2. Click **Create agent**. Do not close this window. Aiven prepares a draft. The agent is not saved until you click **Create agent** again. If you leave, the draft and any test runs are lost. Aiven analyzes the task and prepares the configuration. This can take a few minutes. During this process, Aiven: * Determines the agent's goal * Generates **System instructions** * Selects MCP tools and permissions * Creates the **Task prompt** ### Review the generated configuration[​](#review-the-generated-configuration "Direct link to Review the generated configuration") When the configuration is ready, review it before you create the agent. You can change: * AI model * System instructions * Task prompt * Built-in tools and integrations * MCP role and available tools **System instructions** define how the agent behaves. The **Task prompt** defines the task the agent performs when it runs. If a required integration shows **Not connected**, click **set it up now**. **Manage integrations** opens. You can also click **+ Add other integrations**. For more information, see [Manage integrations](/docs/tools/agents/manage-integrations.md). ### Test the agent[​](#test-the-agent "Direct link to Test the agent") 1. Optional: Click **Run test** to review the output. 2. If you change the configuration, click **Save changes and Run test**. **Create agent** stays unavailable until you save and run the test again. ### Create the agent[​](#create-the-agent "Direct link to Create the agent") 1. Click **Create agent**. 2. For **Schedule**, select **On demand** or a recurring schedule. If you select **Custom**, set the cadence and any extra options, such as **Time** and **Time zone**. 3. Click **Create**. The agent appears on the **Agents** page with the status **Active**. The agent **Overview** shows **Agent details** and **System instructions**. ## Configure an agent manually[​](#configure-an-agent-manually "Direct link to Configure an agent manually") 1. In the Aiven Console, open your project. 2. Click **Agents** > **Create agent** > **Configure manually**. Manual configuration has three steps: **Agent details**, **Integrations**, and **Schedule**. ### Configure agent details[​](#configure-agent-details "Direct link to Configure agent details") 1. Enter a **Name** for the agent. 2. Select the **AI model**. 3. Enter **System instructions** that describe what the agent does, how it responds, and anything it avoids. 4. Click **Continue**. ### Select integrations[​](#select-integrations "Direct link to Select integrations") 1. Under **Built-in tools and integrations**, select the tools the agent can use: * **Web Fetch** * **Web Search** * **Aiven MCP** Connected integrations also appear in this list. 2. Optional: Click **adding or creating MCP integrations**. For more information, see [Manage integrations](/docs/tools/agents/manage-integrations.md). 3. Click **Continue**. ### Configure the schedule[​](#configure-the-schedule "Direct link to Configure the schedule") 1. For **How often should this agent run?**, select **On demand** or a recurring schedule. If you select **Custom**, set the cadence and any extra options, such as **Time** and **Time zone**. 2. For a scheduled agent, enter a **Task prompt**. Aiven sends this message to the agent on each scheduled run. For more information, see [Schedule an agent](/docs/tools/agents/schedule-agent.md). 3. Click **Create agent**. The agent appears on the **Agents** page with the status **Active**. ## Next steps[​](#next-steps "Direct link to Next steps") * [Chat with an agent](/docs/tools/agents/chat-with-agent.md) * [Schedule an agent](/docs/tools/agents/schedule-agent.md) * [Manage an agent](/docs/tools/agents/manage-agent.md) * [Manage integrations](/docs/tools/agents/manage-integrations.md) Related pages * [Managed Agents](/docs/tools/agents.md) --- # Manage an agent [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Update an existing agent, including its system instructions, AI model, and tools. To send a task or question, see [Chat with an agent](/docs/tools/agents/chat-with-agent.md). To run a task on a schedule, see [Schedule an agent](/docs/tools/agents/schedule-agent.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") An existing agent. To create one, see [Create an agent](/docs/tools/agents/create-agent.md). ## Open an agent[​](#open-an-agent "Direct link to Open an agent") 1. In the Aiven Console, open your project. 2. Click **Agents**. 3. Click the agent. The agent opens in a new chat. To chat or open a previous conversation, see [Chat with an agent](/docs/tools/agents/chat-with-agent.md). ## View agent details[​](#view-agent-details "Direct link to View agent details") Click **Overview**. **Overview** shows **Agent details** and **System instructions**. **Agent details** includes the agent name and AI model. ### Edit the agent name or model[​](#edit-the-agent-name-or-model "Direct link to Edit the agent name or model") 1. In **Agent details**, click **Edit**. 2. Change the agent name or the **AI model**. 3. Click **Save**. ### Edit system instructions[​](#edit-system-instructions "Direct link to Edit system instructions") 1. In **System instructions**, click **Edit**. 2. Update the instructions. 3. Click **Save**. ## Change tools and integrations[​](#change-tools-and-integrations "Direct link to Change tools and integrations") Click **Integrations** to change the tools the agent can use, including built-in tools and Aiven MCP. For more information, see [Manage integrations](/docs/tools/agents/manage-integrations.md). ## Manage schedules[​](#manage-schedules "Direct link to Manage schedules") Click **Agent schedules** to view or create scheduled tasks for the agent. For more information, see [Schedule an agent](/docs/tools/agents/schedule-agent.md). Related pages * [Managed Agents](/docs/tools/agents.md) * [Create an agent](/docs/tools/agents/create-agent.md) * [Chat with an agent](/docs/tools/agents/chat-with-agent.md) --- # Manage integrations [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Choose which tools and integrations an agent can use, including built-in tools and [Aiven MCP](/docs/tools/mcp-server.md). You can also connect other Model Context Protocol (MCP) integrations. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") Managed Agents enabled for the project. If you have not requested access or enabled Managed Agents, see [Managed Agents](/docs/tools/agents.md). ## Configure tools and integrations[​](#configure-tools-and-integrations "Direct link to Configure tools and integrations") 1. In the Aiven Console, open your project. 2. Select **Agents**, then select the agent. 3. Select **Integrations**. 4. Select the tools and integrations the agent can use. 5. Select **Save changes**. Built-in tools include: * **Web Fetch** * **Web Search** * **Aiven MCP** To add or manage integrations, select **Advanced integration settings**. ## Connect an integration[​](#connect-an-integration "Direct link to Connect an integration") In **Advanced integration settings**, connected integrations appear at the top of the page. The **Integrations catalog** lists other integrations you can connect. To connect an integration: 1. Find the integration in the **Integrations catalog**. 2. Select **Connect**. 3. Enter the required details and credentials. 4. Select **Connect**. To connect an integration that is not available in the catalog, select **Add custom integration**. ## Set the Aiven MCP role and tools[​](#set-the-aiven-mcp-role-and-tools "Direct link to Set the Aiven MCP role and tools") When you enable **Aiven MCP**, Aiven creates a scoped token automatically. You can assign an MCP role up to your own project permissions. The following roles are available: | Role | Access | | ------------- | ---------------------------------------------------------------------------------------- | | **Read-only** | View services, configuration, logs, and metrics. No changes. | | **Developer** | Manage databases, topics, connectors, and run queries. Cannot create or delete services. | | **Operator** | Full access to all services in the project, including creating and deleting services. | 1. Select an **MCP role**. 2. Under **Available tools**, select the tool groups the agent can use. 3. Select **Save changes**. Related pages * [Managed Agents](/docs/tools/agents.md) * [Aiven MCP](/docs/tools/mcp-server.md) * [Create an agent](/docs/tools/agents/create-agent.md) * [Manage an agent](/docs/tools/agents/manage-agent.md) --- # Schedule an agent [Limited availability](/docs/platform/concepts/service-and-feature-releases.md) Create scheduled tasks so an agent runs automatically at a specified time or interval. Use a schedule for recurring work, such as a daily summary of service health. To interact with an agent on demand, see [Chat with an agent](/docs/tools/agents/chat-with-agent.md). You can also set **On demand** or a schedule when you [create an agent](/docs/tools/agents/create-agent.md). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") An existing agent. To create one, see [Create an agent](/docs/tools/agents/create-agent.md). ## Create a schedule[​](#create-a-schedule "Direct link to Create a schedule") 1. In the Aiven Console, open your project. 2. Click **Agents**. 3. Click the agent. 4. Click **Agent schedules**. 5. Click **Create schedule**. 6. Enter a **Name**. 7. In **Task**, enter the work the agent performs on each run. For example: > Summarize yesterday's signups and flag anomalies. 8. Select the **Cadence**. If the cadence runs at a specific time, set **Time** and **Time zone**. 9. Click **Create schedule**. The Console shows when the schedule runs. For example: **Runs daily at 09:00 Europe/Berlin**. **System instructions** define how the agent behaves. **Task** defines the work this schedule performs. ## Edit a schedule[​](#edit-a-schedule "Direct link to Edit a schedule") 1. On **Agent schedules**, find the schedule to edit. 2. Click **Actions** > **Edit**. 3. Change the **Name**, **Task**, **Cadence**, **Time**, or **Time zone**. ## Enable or disable a schedule[​](#enable-or-disable-a-schedule "Direct link to Enable or disable a schedule") The **Status** column shows **Enabled** or **Disabled**. A disabled schedule does not run. To enable a schedule: 1. On **Agent schedules**, find the schedule to enable. 2. Click **Actions** > **Enable**. To disable a schedule: 1. On **Agent schedules**, find the schedule to disable. 2. Click **Actions** > **Disable**. ## Run a schedule[​](#run-a-schedule "Direct link to Run a schedule") 1. On **Agent schedules**, find the schedule to run. 2. Click **Actions** > **Run now**. The agent opens in a chat and runs the **Task** from the schedule. You can send follow-up messages in the same conversation. ## View schedule runs[​](#view-schedule-runs "Direct link to View schedule runs") 1. On **Agent schedules**, find the schedule. 2. Click **Actions** > **View runs**. ## Delete a schedule[​](#delete-a-schedule "Direct link to Delete a schedule") 1. On **Agent schedules**, find the schedule to delete. 2. Click **Actions** > **Delete**. Related pages * [Managed Agents](/docs/tools/agents.md) * [Create an agent](/docs/tools/agents/create-agent.md) * [Chat with an agent](/docs/tools/agents/chat-with-agent.md) --- # Aiven Console overview In the [Aiven Console](https://console.aiven.io) you can create and manage Aiven services, update your user profile, manage settings across organizations and projects, set up billing groups, view invoices, and more. tip Use an AI assistant connected to [Aiven MCP](/docs/tools/mcp-server.md) to create, update, and view details for Aiven services from clients such as Cursor and Claude Code. ## User profile[​](#user-profile "Direct link to User profile") To view your personal information, authentication settings, and organizations you belong to, click the label **User information** profile icon in the top right. The user profile is also the place where you can get [referral links](/docs/platform/reference/referrals.md), enable [feature previews](/docs/platform/howto/feature-preview.md) to test upcoming features, create tokens, and configure other personal settings. ### Name and email[​](#name-and-email "Direct link to Name and email") You can [update your name and other personal details](/docs/platform/howto/edit-user-profile.md) in your user profile. note You cannot edit your email address. Instead, you can [migrate your Aiven resources to another email address](/docs/platform/howto/change-your-email-address.md) within specific projects. ### User authentication[​](#user-authentication "Direct link to User authentication") On the **Authentication methods** tab of the **User profile**, you can manage your password and authentication settings, including: * [Adding authentication methods](/docs/platform/howto/add-authentication-method.md) * [Managing two-factor authentication](/docs/platform/howto/user-2fa.md) ### Tokens[​](#tokens "Direct link to Tokens") On the **Tokens** page, you can generate or revoke [personal tokens](/docs/platform/concepts/authentication-tokens.md). ## Organization and organizational unit settings[​](#orgs-units-settings "Direct link to Organization and organizational unit settings") The [organization or organizational unit](/docs/platform/concepts/orgs-units-projects.md) that you are currently working with is displayed at the top of the page. You can switch to another organization or organizational unit by clicking the name to open the drop-down menu. If you don't have an organization, click **Create organization** to [create your first organization](/docs/tools/aiven-console/howto/create-orgs-and-units.md). note We strongly recommend creating an organization. It makes managing your projects much easier and comes with many additional features, such as groups, billing groups, and SAML authentication. Organization and organizational unit settings are available on the **Admin** page where you can: * [Manage your groups](/docs/platform/howto/manage-groups.md) * Create new projects under an organization or organizational unit * Configure [authentication policies for an organization](/docs/platform/howto/set-authentication-policies.md) * View logs of activity such as the adding or removing of users, changing authentication methods, and more * Rename or delete an organization or organizational unit ## Projects and services[​](#projects-and-services "Direct link to Projects and services") To navigate between different projects or view all projects click the **Projects** drop-down menu. This menu shows only the projects within the organization or organizational unit that you are currently working in. Selecting a project opens the **Services** page with a list of all services in that project. You can view the status of the services and create new services. On the **Services** page you can also access the [integration endpoints](/docs/platform/concepts/service-integration.md), VPCs, project logs, project permissions, and project settings. ## Billing groups[​](#billing-groups "Direct link to Billing groups") Billing groups let you use billing details across multiple projects and generate a consolidated invoice. Click **Billing** to see and [manage your billing groups](/docs/platform/howto/use-billing-groups.md) and [payment cards](/docs/platform/howto/manage-payment-card.md). --- # Create organizations and organizational units Organizations and organizational units help you group projects and apply common settings like authentication and access. When you sign up for Aiven, an organization is automatically created for you. You can add organizational units in your organization to group related projects and create custom [hierarchical organizations](/docs/platform/concepts/orgs-units-projects.md). ## Create an organizational unit[​](#create-an-organizational-unit "Direct link to Create an organizational unit") You can create an organizational unit within an organization to group your projects by, for example, your departments or environments. Only one level of nesting is supported. This means that you can't create organizational units within other units. * Console * Terraform 1. In the organization, click **Admin**. 2. Click **Organization**. 3. In the **Organizational units** section, click **Create organizational unit**. 4. Enter a name for the unit. 5. Click **Create organizational unit**. ``` Loading... ``` More information on this resource and its configuration options are available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/organizational_unit). ## Create an organization[​](#create-an-organization "Direct link to Create an organization") You can create only one organization for your user account. To separate and organize your projects and services, create organizational units. important You can only verify a domain in one organization, meaning you can't set up SAML authentication, user provisioning with SCIM, or managed users for the same domain in another organization. Additionally, support and commitment contracts cannot be shared across organizations. When you create another organization you also have to manually configure all settings for the new organization such as: * billing groups * authentication policies * users and groups * roles and permissions - Console - Terraform You cannot create an organization while logged in with an [identity provider](/docs/platform/howto/list-identity-providers.md). To create an organization, log in to the Aiven Console using another authentication method. 1. Click **User information** > **Organizations**. 2. Click **Create organization**. 3. Enter a name for the organization. 4. Optional: Select projects to assign to this organization. 5. Click **Create organization**. note You cannot create an organization with a token that you created when you were logged in using an [identity provider](/docs/platform/howto/list-identity-providers.md). ``` Loading... ``` More information on this resource and its configuration options are available in the [Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs/resources/organization). --- # Aiven API Use the Aiven API to programmatically access and automate tasks in the Aiven platform. Common use cases for the Aiven API: * Use with continuous integration to create services during test runs. * Integrate with other parts of your existing automation setup to complete complex tasks. * Deploy and tear down development or demo platforms on a schedule. * Scale your disks based on specific events. tip Want to manage Aiven services from an AI assistant? Use [Aiven MCP](/docs/tools/mcp-server.md) to create, update, and inspect services in natural language from clients such as Cursor and Claude Code. ## Get started with Aiven API[​](#get-started-with-aiven-api "Direct link to Get started with Aiven API") Use the [Postman workspace](https://www.postman.com/aiven-apis/workspace/aiven/overview) to try the Aiven API. 1. [Create a token](/docs/platform/howto/create_authentication_token.md). 2. Fork the Postman **collection** and **environment**. [![Run In Postman](https://run.pstmn.io/button.svg)](https://app.getpostman.com/run-collection/32630186-ad894eed-9287-4a30-8534-d2b4422a254a?action=collection%2Ffork\&source=rip_markdown\&collection-url=entityId%3D32630186-ad894eed-9287-4a30-8534-d2b4422a254a%26entityType%3Dcollection%26workspaceId%3D85dcbab4-8e52-4836-839f-9157852fef73#?env%5BAiven%20environment%5D=W3sia2V5IjoiYXV0aFRva2VuIiwidmFsdWUiOiIiLCJlbmFibGVkIjp0cnVlLCJ0eXBlIjoic2VjcmV0Iiwic2Vzc2lvblZhbHVlIjoiIiwic2Vzc2lvbkluZGV4IjowfSx7ImtleSI6ImJhc2VVcmwiLCJ2YWx1ZSI6Imh0dHBzOi8vYXBpLmFpdmVuLmlvL3YxIiwiZW5hYmxlZCI6dHJ1ZSwidHlwZSI6ImRlZmF1bHQiLCJzZXNzaW9uVmFsdWUiOiJodHRwczovL2FwaS5haXZlbi5pby92MSIsInNlc3Npb25JbmRleCI6MX0seyJrZXkiOiJhY2NvdW50X2lkIiwidmFsdWUiOiJpZCIsImVuYWJsZWQiOnRydWUsInNlc3Npb25WYWx1ZSI6ImlkIiwic2Vzc2lvbkluZGV4IjoyfSx7ImtleSI6InRlYW1faWQiLCJ2YWx1ZSI6ImlkIiwiZW5hYmxlZCI6dHJ1ZSwic2Vzc2lvblZhbHVlIjoiaWQiLCJzZXNzaW9uSW5kZXgiOjN9LHsia2V5IjoicHJvamVjdCIsInZhbHVlIjoiaWQiLCJlbmFibGVkIjp0cnVlLCJzZXNzaW9uVmFsdWUiOiJpZCIsInNlc3Npb25JbmRleCI6NH0seyJrZXkiOiJiaWxsaW5nX2dyb3VwX2lkIiwidmFsdWUiOiJpZCIsImVuYWJsZWQiOnRydWUsInNlc3Npb25WYWx1ZSI6ImlkIiwic2Vzc2lvbkluZGV4Ijo1fSx7ImtleSI6InRlbmFudCIsInZhbHVlIjoidGVuYW50IiwiZW5hYmxlZCI6dHJ1ZSwic2Vzc2lvblZhbHVlIjoidGVuYW50Iiwic2Vzc2lvbkluZGV4Ijo2fSx7ImtleSI6Im9yZ2FuaXphdGlvbl9pZCIsInZhbHVlIjoiaWQiLCJlbmFibGVkIjp0cnVlLCJzZXNzaW9uVmFsdWUiOiJpZCIsInNlc3Npb25JbmRleCI6N30seyJrZXkiOiJkb21haW5faWQiLCJ2YWx1ZSI6ImlkIiwiZW5hYmxlZCI6dHJ1ZSwic2Vzc2lvblZhbHVlIjoiaWQiLCJzZXNzaW9uSW5kZXgiOjh9LHsia2V5IjoidXNlcl9ncm91cF9pZCIsInZhbHVlIjoiaWQiLCJlbmFibGVkIjp0cnVlLCJzZXNzaW9uVmFsdWUiOiJpZCIsInNlc3Npb25JbmRleCI6OX0seyJrZXkiOiJzZXJ2aWNlX25hbWUiLCJ2YWx1ZSI6Im5hbWUiLCJlbmFibGVkIjp0cnVlLCJzZXNzaW9uVmFsdWUiOiJuYW1lIiwic2Vzc2lvbkluZGV4IjoxMH1d) 3. Insert the token in the `authToken` in the Postman environment. 4. See the [API documentation](https://api.aiven.io/doc/). 5. Send your requests via Postman. ## API examples[​](#api-examples "Direct link to API examples") ### List your projects[​](#list-your-projects "Direct link to List your projects") * Request * Response ``` curl -H "Authorization: aivenv1 TOKEN" https://api.aiven.io/v1/project ``` Where `TOKEN` is your token. ``` { "project_membership": { "my-best-demo": "admin", "aiven-sandbox": "admin" }, "project_memberships": { "my-best-demo": [ "admin" ], "aiven-sandbox": [ "admin" ] }, "projects": [ { "account_id": "a225dad8d3c4", "account_name": "Aiven Accounts", "address_lines": [], "available_credits": "0.00", "billing_address": "", "billing_currency": "USD", "billing_emails": [], "billing_extra_text": null, "billing_group_id": "588a8e63-fda7-4ff7-9bff-577debfee604", "billing_group_name": "Billing", "card_info": null, "city": "", "company": "", "country": "", "country_code": "", "default_cloud": "google-europe-north1", "end_of_life_extension": {}, "estimated_balance": "4.11", "estimated_balance_local": "4.11", "payment_method": "no_payment_expected", "project_name": "my-best-demo", "state": "", "tags": {}, "tech_emails": [], "tenant_id": "aiven", "trial_expiration_time": null, "vat_id": "", "zip_code": "" }, { //... } ] } ``` ### List cloud regions[​](#list-cloud-regions "Direct link to List cloud regions") * Request * Response ``` curl https://api.aiven.io/v1/clouds ``` This endpoint does not require authorization. If you aren't authenticated, it returns the standard set of cloud regions. ``` { "clouds": [ { "cloud_description": "Africa, South Africa - Amazon Web Services: Cape Town", "cloud_name": "aws-af-south-1", "geo_latitude": -33.92, "geo_longitude": 18.42, "geo_region": "africa" }, { "cloud_description": "Africa, South Africa - Azure: South Africa North", "cloud_name": "azure-south-africa-north", "geo_latitude": -26.198, "geo_longitude": 28.03, "geo_region": "africa" }, } ``` Related pages * [Personal tokens](/docs/platform/concepts/authentication-tokens.md) * [Aiven Provider for Terraform](https://registry.terraform.io/providers/aiven/aiven/latest/docs) * [API reference docs](https://api.aiven.io/doc/) * [Aiven MCP](/docs/tools/mcp-server.md): Manage Aiven services in natural language. --- # Secret redaction in Aiven API Service user passwords, secret service `user_config` fields, and integration endpoint secrets are redacted in API responses by default. Redacted values are shown as ``. For integrations that send config back to Aiven, leaving this placeholder prevents accidental overwriting of the secret. To read secrets on a `GET` request: * Send `include_secrets=true` as a query parameter. * Use a [role or permission](/docs/platform/concepts/permissions.md) that can read that secret type. For service secrets, calls without the required permissions receive a `403 Forbidden` response. For integration endpoint secrets, calls return redacted values. The following endpoints can return secrets in plaintext: * `GET /project/PROJECT/service/SERVICE_NAME/user/SERVICE_USERNAME` * `GET /project/PROJECT/service/SERVICE_NAME` * `GET /project/PROJECT/service` * `GET /project/PROJECT/integration_endpoint/INTEGRATION_ENDPOINT_ID` Only `GET` endpoints can reveal secrets. Most write endpoints redact secrets. One exception is the `POST /project/PROJECT/service/SERVICE_NAME/user` endpoint, which returns newly generated credentials in plaintext. --- # Aiven CLI The Aiven command line interface (CLI) lets you use the Aiven platform and services in a scriptable way through the API. tip Want to manage Aiven services from an AI assistant? Use [Aiven MCP](/docs/tools/mcp-server.md) to create, update, and inspect services in natural language from clients such as Cursor and Claude Code. ## Install the Aiven CLI[​](#install-the-aiven-cli "Direct link to Install the Aiven CLI") 1. The `avn` utility is a [Python package](https://pypi.org/project/aiven-client/): * pip * Homebrew ``` pip install aiven-client ``` ``` brew install aiven-client ``` 2. To check your installation, run: ``` avn --version ``` ## Authenticate with the Aiven CLI[​](#authenticate-with-the-aiven-cli "Direct link to Authenticate with the Aiven CLI") You can authenticate using your password or a [token](/docs/platform/concepts/authentication-tokens.md). * With a password * With a token 1. To log in with your email, run: ``` avn user login EMAIL_ADDRESS ``` 2. When prompted, enter your password. 1) [Create a token](/docs/platform/howto/create_authentication_token.md). 2) To authenticate with a token, run: ``` avn user login EMAIL_ADDRESS --token ``` note If you are registered on Aiven through the AWS or GCP marketplace, use the `--tenant` option. For example: ``` avn user login EMAIL_ADDRESS --tenant aws ``` ## Configure the output format[​](#configure-the-output-format "Direct link to Configure the output format") To get information in JSON format, use the `--json` switch with any command. Related pages * [Learn how to use the Aiven CLI](https://aiven.io/blog/aiven-cmdline) for common tasks. * Watch the [how to get started tutorial](https://www.youtube.com/watch?v=nf3PPn5w6K8). * Go to the [aiven-client repository on GitHub](https://github.com/aiven/aiven-client). * [Aiven MCP](/docs/tools/mcp-server.md): Manage Aiven services in natural language. --- # avn byoc Set up and manage your [custom clouds](/docs/platform/concepts/byoc.md) using the Aiven client and `avn byoc` commands. ## Manage a custom cloud[​](#manage-a-custom-cloud "Direct link to Manage a custom cloud") ### `avn byoc create`[​](#avn-byoc-create "Direct link to avn-byoc-create") [Creates a custom cloud](/docs/platform/howto/byoc/create-cloud/create-custom-cloud.md) in an organization. | Parameter | Required | Information | | -------------------- | -------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `--organization-id` | Yes | Identifier of an organization where to create the custom cloud | | `--deployment-model` | Yes | Determines the [deployment model](/docs/platform/concepts/byoc.md#byoc-architecture), for example `standard` (the default deployment model with a private workload network) | | `--cloud-provider` | Yes | Cloud provider to be used for running the custom cloud, for example`aws` (Amazon Web Services) | | `--cloud-region` | Yes | Cloud region where to create the custom cloud, for example `eu-west-1` | | `--reserved-cidr` | Yes | IP address range of the VPC to be created in your cloud account for Aiven services hosted on a custom cloud | | `--display-name` | Yes | Name of the custom cloud | ### `avn byoc delete`[​](#avn-byoc-delete "Direct link to avn-byoc-delete") Deletes a custom cloud from an organization. | Parameter | Required | Information | | ------------------- | -------- | -------------------------------------------------------------- | | `--organization-id` | Yes | Identifier of an organization where to delete the custom cloud | | `--byoc-id` | Yes | Identifier of the custom cloud to be deleted | ### `avn byoc list`[​](#avn-byoc-list "Direct link to avn-byoc-list") Returns a list of all the custom clouds in an organization. | Parameter | Required | Information | | ------------------- | -------- | ----------------------------- | | `--organization-id` | Yes | Identifier of an organization | ### `avn byoc provision`[​](#avn-byoc-provision "Direct link to avn-byoc-provision") Provisions resources for a custom cloud. | Parameter | Required | Information | | ----------------------------------------------- | ----------- | --------------------------------------------------------------------------------------------------------- | | `--organization-id` | Yes | Identifier of an organization where to modify the custom cloud | | `--byoc-id` | Yes | Identifier of the custom cloud to be modified | | `--aws-iam-role-arn` | Yes (AWS) | Identifier of the role created when running the infrastructure template in your AWS account | | `--google-privilege-bearing-service-account-id` | Yes (GCP) | Identifier of the service account created when running the infrastructure template in your Google account | | `--azure-subscription-id` | Yes (Azure) | Azure subscription ID where the BYOC infrastructure is deployed | | `--azure-tenant-id` | Yes (Azure) | Entra ID tenant ID of the directory where the Aiven CCE enterprise application is installed | ### `avn byoc update`[​](#avn-byoc-update "Direct link to avn-byoc-update") Modifies a custom cloud configuration. | Parameter | Required | Information | | -------------------- | -------- | --------------------------------------------------------------------------------------------------------------------- | | `--organization-id` | Yes | Identifier of an organization where to modify the custom cloud | | `--byoc-id` | Yes | Identifier of the custom cloud to be modified | | `--deployment-model` | No | Private or public network architecture model defining how resources are arranged and connected to your cloud provider | | `--cloud-provider` | No | Cloud provider running the custom cloud | | `--cloud-region` | No | Cloud region where the custom cloud runs | | `--reserved-cidr` | No | IP address range of the VPC in your cloud account for Aiven services created in the custom cloud | | `--display-name` | No | Name of the custom cloud | | `--contact-email` | No | Customer contact as `email="EMAIL",real_name="NAME",role="ROLE"`. Repeatable. Replaces the full contact list. | ## Tag a custom cloud[​](#tag-a-custom-cloud "Direct link to Tag a custom cloud") Custom cloud tags are key-value pairs that you can attach to your custom cloud for resource categorization. They propagate to resources on the Aiven platform and in your own cloud infrastructure. Custom cloud tags are cascaded to bastion nodes and disks in private [deployment models](https://aiven.io/docs/platform/concepts/byoc#byoc-architecture). ### `avn byoc tags list`[​](#avn-byoc-tags-list "Direct link to avn-byoc-tags-list") Returns infrastructure tags attached to a custom cloud. **Syntax** ``` avn byoc tags list \ --organization-id "ORGANIZATION_IDENTIFIER" \ --byoc-id "CUSTOM_CLOUD_IDENTIFIER" ``` **Parameters** | Parameter | Required | Information | | ------------------- | -------- | -------------------------------------------------------------- | | `--organization-id` | Yes | Identifier of an organization where to modify the custom cloud | | `--byoc-id` | Yes | Identifier of the custom cloud to be modified | ### `avn byoc tags update`[​](#avn-byoc-tags-update "Direct link to avn-byoc-tags-update") Adds, updates, or removes infrastructure tags on a custom cloud. **Syntax** ``` avn byoc tags update \ --organization-id "ORGANIZATION_IDENTIFIER" \ --byoc-id "CUSTOM_CLOUD_IDENTIFIER" \ --add-tag TAG_KEY_A=TAG_VALUE_A \ --remove-tag TAG_KEY_B ``` **Parameters** | Parameter | Required | Information | | ------------------- | -------- | ----------------------------------------------------------------------------- | | `--organization-id` | Yes | Identifier of an organization where to modify the custom cloud | | `--byoc-id` | Yes | Identifier of the custom cloud to be modified | | `--add-tag` | No | Adds or updates key-value pairs on a custom cloud for resource categorization | | `--remove-tag` | No | Deletes key-value pairs attached to a custom cloud | ### `avn byoc tags replace`[​](#avn-byoc-tags-replace "Direct link to avn-byoc-tags-replace") Replaces all existing tags with new ones. **Syntax** ``` avn byoc tags replace \ --organization-id "ORGANIZATION_IDENTIFIER" \ --byoc-id "CUSTOM_CLOUD_IDENTIFIER" \ --tag TAG_KEY_A=TAG_VALUE_A ``` **Parameters** | Parameter | Required | Information | | ------------------- | -------- | ------------------------------------------------------------------------------------ | | `--organization-id` | Yes | Identifier of an organization where to modify the custom cloud | | `--byoc-id` | Yes | Identifier of the custom cloud to be modified | | `--tag` | Yes | Key-value pair that replaces all existing key-value pairs attached to a custom cloud | ## Manage custom cloud permissions[​](#manage-custom-cloud-permissions "Direct link to Manage custom cloud permissions") ### `avn byoc cloud permissions add`[​](#avn-byoc-cloud-permissions-add "Direct link to avn-byoc-cloud-permissions-add") Adds new permissions to use the custom cloud in projects or accounts (organizational units) while keeping any existing permissions in place. | Parameter | Required | Information | | ------------------- | -------- | ------------------------------------------------------------------------------------------ | | `--organization-id` | Yes | Identifier of an organization where the custom cloud is located | | `--byoc-id` | Yes | Identifier of the custom cloud for which to grant the new permissions | | `--account` | No | Identifier of your account (organizational unit) for which the new permissions are granted | | `--project` | No | Name of the project for which to grant permissions | ### `avn byoc cloud permissions get`[​](#avn-byoc-cloud-permissions-get "Direct link to avn-byoc-cloud-permissions-get") Retrieves permissions to a custom cloud. | Parameter | Required | Information | | ------------------- | -------- | ------------------------------------------------------------------------ | | `--organization-id` | Yes | Identifier of an organization where the custom cloud is located | | `--byoc-id` | Yes | Identifier of the custom cloud for which to retrieve permissions details | ### `avn byoc cloud permissions remove`[​](#avn-byoc-cloud-permissions-remove "Direct link to avn-byoc-cloud-permissions-remove") Revokes permissions to a custom cloud. | Parameter | Required | Information | | ------------------- | -------- | -------------------------------------------------------------------------------------- | | `--organization-id` | Yes | Identifier of an organization where the custom cloud is located | | `--byoc-id` | Yes | Identifier of the custom cloud for which to revoke permissions | | `--account` | No | Identifier of your account (organizational unit) for which the permissions are revoked | | `--project` | No | Name of the project for which to revoke permissions | ### `avn byoc cloud permissions set`[​](#avn-byoc-cloud-permissions-set "Direct link to avn-byoc-cloud-permissions-set") Replaces all permissions there may be for using the custom cloud in projects or accounts (organizational units). After you run this command successfully, there are no permissions other than the ones you've just granted using this command. | Parameter | Required | Information | | ------------------- | -------- | --------------------------------------------------------------------------------------- | | `--organization-id` | Yes | Identifier of an organization where the custom cloud is located | | `--byoc-id` | Yes | Identifier of the custom cloud for which to set permissions | | `--account` | No | Identifier of your account (organizational unit) for which the permissions are replaced | | `--project` | No | Name of the project for which to set permissions | ## Manage an infrastructure template[​](#manage-an-infrastructure-template "Direct link to Manage an infrastructure template") Manage a custom cloud Terraform infrastructure template. ### `avn byoc template terraform get-template`[​](#avn-byoc-template-terraform-get-template "Direct link to avn-byoc-template-terraform-get-template") Downloads a custom cloud Terraform template. | Parameter | Required | Information | | ------------------- | -------- | ------------------------------------------------------------------------- | | `--organization-id` | Yes | Identifier of an organization where the custom cloud is located | | `--byoc-id` | Yes | Identifier of the custom cloud for which to download a Terraform template | ### `avn byoc template terraform get-vars`[​](#avn-byoc-template-terraform-get-vars "Direct link to avn-byoc-template-terraform-get-vars") Downloads a custom cloud Terraform variables file. | Parameter | Required | Information | | ------------------- | -------- | -------------------------------------------------------------------- | | `--organization-id` | Yes | Identifier of an organization where the custom cloud is located | | `--byoc-id` | Yes | Identifier of the custom cloud for which to download a variable file | --- # avn cloud The `avn cloud` command allows you to list the clouds available in a given project. ## List cloud region details[​](#list-cloud-region-details "Direct link to List cloud region details") Commands for listing cloud regions to be used when creating or moving instances with `avn` commands. ### `avn cloud list`[​](#avn-cloud-list "Direct link to avn-cloud-list") Lists cloud regions with related geographical region, latitude and longitude. | Parameter | Information | | ----------- | -------------------------------- | | `--project` | The project to fetch details for | **Example:** Show the clouds available to the currently selected project. ``` avn cloud list ``` **Example:** Show the clouds available to a named project. ``` avn cloud list --project my-project ``` A reference of the cloud regions is available in the [dedicated document](/docs/platform/reference/list_of_clouds.md). --- # avn credits Full list of commands for `avn credits`. ## Aiven credits[​](#aiven-credits "Direct link to Aiven credits") All commands for managing Aiven credits. ### `avn credits claim`[​](#avn-credits-claim "Direct link to avn-credits-claim") Add an Aiven credit code to a project. | Parameter | Information | | ----------- | ------------------------------------ | | `code` | Credit code to claim | | `--project` | The project to claim the credits for | **Example:** Add a credit code to the currently selected project. ``` avn credits claim "credit-code-123" ``` **Example:** Add a credit code to a named project. ``` avn credits claim "credit-code-123" --project my-project ``` ### `avn credits list`[​](#avn-credits-list "Direct link to avn-credits-list") List the credit codes associated with a project. | Parameter | Information | | ----------- | ----------------------------------- | | `--project` | The project to list the credits for | **Example:** List all credit codes associated with the currently selected project. ``` avn credits list ``` **Example:** List all credit codes associated with a named project. ``` avn credits list --project my-project ``` --- # avn events The `avn events` command is an audit log of things that have happened in a particular project, including: * timestamp of when the event occurred * which user performed the action (email address of the user, or "Aiven Automation" where it was an automatic change) * the type of event, such as `service_create`, `project_vpc_create` or `service_master_promotion` (this is not an exhaustive list and new event types are added from time to time * the name of the service affected * a more detailed description of the event An example of events output (note that the newest event is shown first): ``` TIME ACTOR EVENT_TYPE SERVICE_NAME EVENT_DESC ==================== ======================= ======================== ============== ============================================================================================== 2021-08-10T13:38:23Z my_user@aiven.io service_update demo-pg Reset service user password 2021-08-10T13:38:22Z my_user@aiven.io service_update demo-kafka Reset service user password 2021-08-10T13:37:21Z Aiven Automation service_master_promotion demo-pg Promoted demo-pg-1 to be the new master in service demo-pg. 2021-08-10T13:35:39Z my_user@aiven.io service_create demo-pg Created 'pg' service 'demo-pg' with plan 'business-4' in cloud 'google-europe-west3' 2021-08-10T13:35:22Z my_user@aiven.io service_create demo-kafka Created 'kafka' service 'demo-kafka' with plan 'business-4' in cloud 'google-europe-west3' ``` ## Project events[​](#project-events "Direct link to Project events") Information about the events that have occurred in a project. ### `avn events`[​](#avn-events "Direct link to avn-events") Lists instance or integration creation, deletion or modification events. | Parameter | Information | | ----------- | -------------------------------- | | `--project` | The project to fetch details for | **Example:** Show the recent events of the currently selected project. ``` avn events ``` **Example:** Show the most recent 10 events of a named project. ``` avn events -n 10 --project my-project ``` --- # avn mirrormaker Full list of commands for `avn mirrormaker`. ## Create and manage Aiven for Apache Kafka® MirrorMaker 2 replication flows[​](#create-and-manage-aiven-for-apache-kafka-mirrormaker-2-replication-flows "Direct link to Create and manage Aiven for Apache Kafka® MirrorMaker 2 replication flows") Commands for managing Aiven for Apache Kafka® MirrorMaker 2 replication flows. ### `avn mirrormaker replication-flow create`[​](#avn-mirrormaker-replication-flow-create "Direct link to avn-mirrormaker-replication-flow-create") Creates a new Aiven for Apache Kafka® MirrorMaker 2 replication flow. warning Before creating a replication flow, an [integration](/docs/tools/cli/service/integration.md#avn_service_integration_create) needs to be created between the Aiven for Apache Kafka MirrorMaker 2 service and each of the source and the target services. for example, An integration with alias `kafka-target-alias` between an Aiven for Apache Kafka service named `kafka-target` and an Aiven for Apache Kafka MirrorMaker 2 named `kafka-mm` can be created with: ``` avn service integration-create \ -s kafka-target \ -d kafka-mm \ -t kafka_mirrormaker \ -c cluster_alias=kafka-target-alias ``` At most **one** replication flow can be build between any two Aiven for Apache Kafka services. | Parameter | Information | | ------------------------- | ------------------------------------------------------------------------------------------------------ | | `service_name` | The Aiven for Apache Kafka MirrorMaker 2 service where to create the replication flow | | `--source-cluster` | The Aiven for Apache Kafka service to be used as source for replication | | `--target-cluster` | The Aiven for Apache Kafka service to be used as target for replication | | `replication_flow_config` | JSON string or path (preceded by `@`) to a JSON configuration file for the replication flow definition | **Example:** In the service `kafka-mm` create a replication flow from an Aiven for Apache Kafka service with integration alias `kafka-source-alias` to a service named `kafka-target-alias` with the following settings: * include all topics with name starting with `my-src-topic` (topic name patterns can be defined using [Java patterns](https://docs.oracle.com/javase/7/docs/api/java/util/regex/Pattern)) * exclude all topics with name ending with `not-include` * `DefaultReplicationPolicy` as replication policy class * enable MirrorMaker 2 heartbeats * enable synching of consumer groups offset every `60` seconds ``` avn mirrormaker replication-flow create kafka-mm \ --source-cluster kafka-source-alias \ --target-cluster kafka-target-alias \ ' { "emit_heartbeats_enabled": true, "enabled": true, "replication_policy_class": "org.apache.kafka.connect.mirror.DefaultReplicationPolicy", "source_cluster": "kafka-source-alias", "sync_group_offsets_enabled": true, "sync_group_offsets_interval_seconds": 60, "target_cluster": "kafka-target-alias", "topics": [ "my-src-topic.*" ], "topics.blacklist": [ ".*not-include" ] } ' ``` ### `avn mirrormaker replication-flow delete`[​](#avn-mirrormaker-replication-flow-delete "Direct link to avn-mirrormaker-replication-flow-delete") Deletes an existing Aiven for Apache Kafka® MirrorMaker 2 replication flow. | Parameter | Information | | ------------------ | ------------------------------------------------------------------------------------- | | `service_name` | The Aiven for Apache Kafka MirrorMaker 2 service where to delete the replication flow | | `--source-cluster` | The Aiven for Apache Kafka service to be used as source for replication | | `--target-cluster` | The Aiven for Apache Kafka service to be used as target for replication | **Example:** In the service `kafka-mm` delete the replication flow from an Aiven for Apache Kafka service with integration alias `kafka-source-alias` to the service named `kafka-target-alias`. ``` avn mirrormaker replication-flow delete kafka-mm \ --source-cluster kafka-source-alias \ --target-cluster kafka-target-alias ``` ### `avn mirrormaker replication-flow get`[​](#avn-mirrormaker-replication-flow-get "Direct link to avn-mirrormaker-replication-flow-get") Retrieves the configuration details of an existing Aiven for Apache Kafka® MirrorMaker 2 replication flow. | Parameter | Information | | ------------------ | ------------------------------------------------------------------------------------------ | | `service_name` | The Aiven for Apache Kafka MirrorMaker 2 service where to get the replication flow details | | `--source-cluster` | The Aiven for Apache Kafka service to be used as source for replication | | `--target-cluster` | The Aiven for Apache Kafka service to be used as target for replication | **Example:** In the service `kafka-mm` retrieve the details of the replication flow from an Aiven for Apache Kafka service with integration alias `kafka-source-alias` to the service named `kafka-target-alias`. ``` avn mirrormaker replication-flow get kafka-mm \ --source-cluster kafka-source-alias \ --target-cluster kafka-target-alias ``` An example of the `avn mirrormaker replication-flow get` command output: ``` { "emit_heartbeats_enabled": true, "enabled": true, "replication_policy_class": "org.apache.kafka.connect.mirror.DefaultReplicationPolicy", "source_cluster": "kafka-source-alias", "sync_group_offsets_enabled": true, "sync_group_offsets_interval_seconds": 60, "target_cluster": "kafka-target-alias", "topics": [ "my-src-topic.*" ], "topics.blacklist": [ ".*not-include" ] } ``` ### `avn mirrormaker replication-flow list`[​](#avn-mirrormaker-replication-flow-list "Direct link to avn-mirrormaker-replication-flow-list") Lists the configuration details for all replication flows defined in an existing Aiven for Apache Kafka® MirrorMaker 2 service. | Parameter | Information | | -------------- | ----------------------------------------------------------------------------------- | | `service_name` | The Aiven for Apache Kafka MirrorMaker 2 service where to list the replication flow | **Example:** List the configuration details for all replication flows defined in an existing Aiven for Apache Kafka MirrorMaker 2 named `kafka-mm`. ``` avn mirrormaker replication-flow list kafka-mm ``` An example of the `avn mirrormaker replication-flow list` command output: ``` [ { "emit_heartbeats_enabled": true, "enabled": true, "replication_policy_class": "org.apache.kafka.connect.mirror.DefaultReplicationPolicy", "source_cluster": "kafka-source-alias", "sync_group_offsets_enabled": true, "sync_group_offsets_interval_seconds": 60, "target_cluster": "kafka-target-alias", "topics": [ "my-src-topic.*" ], "topics.blacklist": [ ".*not-include" ] } ] ``` ### `avn mirrormaker replication-flow update`[​](#avn-mirrormaker-replication-flow-update "Direct link to avn-mirrormaker-replication-flow-update") Updates an existing Aiven for Apache Kafka® MirrorMaker 2 replication flow. | Parameter | Information | | ------------------------- | ------------------------------------------------------------------------------------------------------ | | `service_name` | The Aiven for Apache Kafka MirrorMaker 2 service where to update the replication flow | | `--source-cluster` | The Aiven for Apache Kafka service to be used as source for replication | | `--target-cluster` | The Aiven for Apache Kafka service to be used as target for replication | | `replication_flow_config` | JSON string or path (preceded by `@`) to a JSON configuration file for the replication flow definition | **Example:** In the service `kafka-mm` update the replication flow from an Aiven for Apache Kafka service with integration alias `kafka-source-alias` to a service named `kafka-target-alias` with the settings contained in a file named `replication-flow.json`. ``` avn mirrormaker replication-flow update kafka-mm \ --source-cluster kafka-source-alias \ --target-cluster kafka-target-alias \ @replication-flow.json ``` --- # avn service Full list of commands for `avn service`. ## Manage service details[​](#manage-service-details "Direct link to Manage service details") Commands for managing Aiven services via `avn` commands. ### `avn service acl`[​](#avn-service-acl "Direct link to avn-service-acl") Manages the Aiven for Apache Kafka® ACL entries. More information on `acl-add`, `acl-delete` and `acl-list` can be found in [the dedicated page](/docs/tools/cli/service/acl.md). ### `avn service backup-list`[​](#avn-service-backup-list "Direct link to avn-service-backup-list") Retrieves the list of backups for a certain service. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** Retrieve the list of backups for the service `grafana-25c408a5`. ``` avn service backup-list grafana-25c408a5 ``` An example of `service backup-list` output: ``` BACKUP_NAME BACKUP_TIME DATA_SIZE STORAGE_LOCATION ============================== ==================== ========= =================== grafana-20220614T140308137245Z 2022-06-14T14:03:08Z 774144 google-europe-west3 ``` ### `avn service ca get`[​](#avn_service_ca_get "Direct link to avn_service_ca_get") Retrieves the project CA that the selected service belongs to. | Parameter | Information | | ------------------- | ------------------------------------------------------ | | `service_name` | The name of the service | | `--target-filepath` | The file path used to store the CA certificate locally | **Example:** Retrieve the CA certificate for the project where the service named `kafka-doc` belongs and store it under `/tmp/ca.pem`. ``` avn service ca get kafka-doc --target-filepath /tmp/ca.pem ``` ### `avn service cli`[​](#avn-service-cli "Direct link to avn-service-cli") Opens the appropriate interactive shell, such as `psql` or `valkey-cli`, to the given service. Supported only for Aiven for PostgreSQL®, Aiven for Valkey™ and Aiven for Caching. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** Open a new `psql` shell connecting to an Aiven for PostgreSQL® service named `pg-doc`. ``` avn service cli pg-doc ``` ### `avn service connection-info`[​](#avn-service-connection-info "Direct link to avn-service-connection-info") Retrieves the connection information for Aiven for Apache Kafka®, Aiven for PostgreSQL® and Aiven for Caching in a variety of formats. More information on `connection-info` can be found in [the dedicated page](/docs/tools/cli/service/connection-info.md). ### `avn service connection-pool`[​](#avn-service-connection-pool "Direct link to avn-service-connection-pool") Manages the [PgBouncer connection pool](/docs/products/postgresql/concepts/pg-connection-pooling.md) for a given PostgreSQL® service. More information on `connection-pool-add`, `connection-pool-delete`, `connection-pool-list` and `connection-pool-update` can be found in [the dedicated page](/docs/tools/cli/service/connection-pool.md). ### `avn service connector`[​](#avn-service-connector "Direct link to avn-service-connector") Set of commands for managing Aiven for Apache Kafka® Connect connectors. More information on `connector available`, `connector create`, `connector delete`, `connector list`, `connector pause`, `connector restart`, `connector restart-task`, `connector resume`, `connector schema`, `connector status` and `connector update` can be found in the [dedicated page](/docs/tools/cli/service/connector.md). ### `avn service create`[​](#avn-cli-service-create "Direct link to avn-cli-service-create") Creates a new service. | Parameter | Information | | --------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--service-type` | The type of service; the [service types command](/docs/tools/cli/service-cli.md#avn-cli-service-type) has the available values | | `--plan` | Aiven subscription plan name; check [`avn_service_plan`](/docs/tools/cli/service-cli.md#avn-service-plan) for more information | | `--cloud` | The cloud region name; check [avn-cloud-list](/docs/tools/cli/cloud.md#avn-cloud-list) for more information | | `--disk-space-gib` | Total amount of disk space for data storage (GiB) | | `--no-fail-if-exists` | The create command will not fail if a service with the same name already exists | | `--project-vpc-id` | Id of the project VPC where to include the created service. The cloud of the project's VPC must match the service's cloud | | `--no-project-vpc` | Stops the service to be included in the project VPC even if one is available in the selected cloud | | `--enable-termination-protection` | Enables termination protection for the service | | `-c KEY=VALUE` | Any additional configuration settings for your service; check our documentation for more information, or use the [service types command](/docs/tools/cli/service-cli.md#avn-cli-service-type) which has a verbose mode that shows all options. | **Example:** Create new Aiven for Kafka® service named `kafka-demo` in the region `google-europe-west3` with: * the `business-4` plan * Kafka Connect enabled * 600 GiB of total storage capacity ``` avn service create kafka-demo \ --service-type kafka \ --cloud google-europe-west3 \ --plan business-4 \ -c kafka_connect=true \ --disk-space-gib 600 ``` ### `avn service credentials-reset`[​](#avn-service-credentials-reset "Direct link to avn-service-credentials-reset") Resets the service credentials. More information on user password change is provided in the [dedicated page](/docs/tools/cli/service/user.md). | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** Reset the credentials of a service named `kafka-demo`. ``` avn service credentials-reset kafka-demo ``` ### `avn service current-queries`[​](#avn-service-current-queries "Direct link to avn-service-current-queries") List current service connections/queries for an Aiven for PostgreSQL®, Aiven for MySQL or Aiven for Caching service. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** List the queries running for a service named `pg-demo`. ``` avn service current-queries pg-demo ``` ### `avn service database`[​](#avn-service-database "Direct link to avn-service-database") Manages databases within an Aiven for PostgreSQL®, or Aiven for MySQL. More information on `database-add`, `database-delete` and `database-list` can be found in [the dedicated page](/docs/tools/cli/service/database.md). ### `avn service es-acl`[​](#avn-service-es-acl "Direct link to avn-service-es-acl") Manages rules to OpenSearch® ACL and extended ACL configuration. More information on `es-acl-add`, `es-acl-del`, `es-acl-disable`, `es-acl-enable`, `es-acl-extended-disable`, `es-acl-extended-enable` and `es-acl-extended-list` can be found in [the dedicated page](/docs/tools/cli/service/es-acl.md). ### `avn service flink`[​](#avn-service-flink "Direct link to avn-service-flink") Manages Aiven for Apache Flink® tables and jobs. More info on `flink create-application`, `flink list-applications`, `flink get-application`, `flink update-application`, `flink delete-application`, `flink create-application-version`, `flink validate-application-version`, `flink get-application-version`, `flink delete-application-version`, `flink list-application-deployments`, `flink get-application-deployment`, `flink create-application-deployment`, `flink delete-application-deployment`, `flink stop-application-deployment`, `flink cancel-application-deployment` can be found in [the dedicated page](/docs/tools/cli/service/flink.md). ### `avn service get`[​](#avn_service_get "Direct link to avn_service_get") Retrieves a single service details. | Parameter | Information | | -------------- | --------------------------- | | `service_name` | The name of the service | | `--format` | Format of the output string | **Example:** Retrieve the `pg-demo` service details in the `'{service_name} {service_uri}'` format. ``` avn service get pg-demo --format '{service_name} {service_uri}' ``` **Example:** Retrieve the `pg-demo` full service details in JSON format. ``` avn service get pg-demo --json ``` ### `avn service index`[​](#avn-service-index "Direct link to avn-service-index") Manages OpenSearch® service indexes. [the dedicated page](/docs/tools/cli/service/service-index.md). ### `avn service integration`[​](#avn-service-integration "Direct link to avn-service-integration") Manages Aiven internal and external services integrations. More information on `integration-delete`, `integration-endpoint-create`, `integration-endpoint-delete`, `integration-endpoint-list`, `integration-endpoint-types-list`, `integration-endpoint-update`, `integration-list`, `integration-types-list` and `integration-update` can be found in [the dedicated page](/docs/tools/cli/service/integration.md). ### `avn service kafka-acl`[​](#avn-service-kafka-acl "Direct link to avn-service-kafka-acl") Manages the Apache Kafka® native ACL entries. More information on `kafka-acl-add`, `kafka-acl-delete` and `kafka-acl-list` can be found in [the dedicated page](/docs/tools/cli/service/kafka-acl.md). ### `avn service list`[​](#avn-service-list "Direct link to avn-service-list") Lists services within an Aiven project. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** Retrieve all the services running in the currently selected project. ``` avn service list ``` An example of `service list` output: ``` SERVICE_NAME SERVICE_TYPE STATE CLOUD_NAME PLAN CREATE_TIME UPDATE_TIME ================== ============ ======= =================== =========== ==================== ==================== os-24a6d6db opensearch RUNNING google-europe-west3 business-4 2021-09-27T10:18:04Z 2021-09-27T10:23:31Z kafka-2134 kafka RUNNING google-europe-west3 business-4 2021-09-27T08:48:35Z 2021-09-27T11:20:55Z mysql-12f7628c mysql RUNNING google-europe-west3 business-4 2021-09-27T10:18:09Z 2021-09-27T10:23:02Z pg-123456 pg RUNNING google-europe-west3 business-4 2021-09-27T07:41:04Z 2021-09-27T10:56:19Z ``` **Example:** Retrieve all the services with name `demo-pg` running in the project named `mytestproject`. ``` avn service list demo-pg --project mytestproject ``` ### `avn service logs`[​](#avn-service-logs "Direct link to avn-service-logs") Retrieves the selected service logs. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** Retrieve the logs for the service named `pg-demo`. ``` avn service logs pg-demo ``` note Use `avn service logs SERVICE_NAME -f` sparingly. For continuous log monitoring, set up a [log integration](/docs/platform/concepts/service-integration.md). ### `avn service maintenance-start`[​](#avn-service-maintenance-start "Direct link to avn-service-maintenance-start") Starts the service maintenance updates. warning Maintenance updates do not typically cause any noticeable impact on the service in use but may sometimes cause a short period of lower performance or downtime which will not exceed 1 hour. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** Start the maintenance updates for the service named `pg-demo`. ``` avn service maintenance-start pg-demo ``` note If there are no updates available, the command will show a `service is up to date, maintenance not required` message. ### `avn service metrics`[​](#avn-service-metrics "Direct link to avn-service-metrics") Retrieves the metrics for a defined service in Google chart compatible format. The list of service metrics includes: * `cpu_usage`: CPU usage percentage * `disk_usage`: Disk space usage percentage * `disk_ioread`: Disk reads IOPS * `disk_iowrites`: Disk writes IOPS * `load_average`: 5 min CPU load average * `mem_usage`: Memory usage percentage * `net_receive`: Network traffic received in bytes/s * `net_send`: Network traffic transmitted in bytes/s | Parameter | Information | | -------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--period` | The time period to retrieve the metrics for (possible values `hour`, `day`, `week`, `month`, `year`); the time period is relative to the current date and time, for example, `hour` will retrieve metrics for the last hour. | note The **granularity** of retrieved data changes based on the `--period` flag: * `hour`: 30 seconds * `day`: 5 minutes * `week`: 30 minutes * `month`: 3 hours * `year`: 1 day **Example:** Retrieve the daily metrics for the service named `pg-demo`. ``` avn service metrics pg-demo --period day ``` ### `avn service migration-status`[​](#avn-cli-service-migration-status "Direct link to avn-cli-service-migration-status") Get migration status ### `avn service plans`[​](#avn-service-plan "Direct link to avn-service-plan") Lists the service plans available in a selected project for a defined service type. | Parameter | Information | | ---------------- | --------------------------------------------------------------------------------------------------------------------------- | | `--service-type` | The type of service, check [avn-cli-service-type](/docs/tools/cli/service-cli.md#avn-cli-service-type) for more information | | `--cloud` | The cloud region | | `--monthly` | To show the monthly price estimate | **Example:** List the service plans available for a PostgreSQL® service in the `google-europe-west3` region. ``` avn service plans --service-type pg --cloud google-europe-west3 ``` An example of `service plans` output: ``` pg:hobbyist $0.034/h Hobbyist (1 CPU, 2 GB RAM, 8 GB disk) pg:startup-4 $0.136/h Startup-4 (1 CPU, 4 GB RAM, 80 GB disk) pg:startup-8 $0.267/h Startup-8 (2 CPU, 8 GB RAM, 175 GB disk) ... pg:premium-360 $36.027/h Premium-360 (96 CPU, 384 GB RAM, 3000 GB disk) 3-node high availability set pg:premium-512 $43.836/h Premium-512 (128 CPU, 512 GB RAM, 3000 GB disk) 3-node high availability set pg:premium-896 $72.329/h Premium-896 (224 CPU, 896 GB RAM, 3000 GB disk) 3-node high availability set ``` ### `avn service privatelink`[​](#avn-service-privatelink "Direct link to avn-service-privatelink") Manages Aiven privatelink connections for AWS and Azure. More information on `privatelink availability`, `privatelink aws` and `privatelink azure` can be found in [the dedicated page](/docs/tools/cli/service/privatelink.md). ### `avn service queries`[​](#avn-service-queries "Direct link to avn-service-queries") Lists the service connections/queries statistics for an Aiven for PostgreSQL® or Aiven for MySQL. The list of queries data points retrievable includes: * the `public.pg_stat_statements` columns (see the [documentation for these statistics columns](https://www.postgresql.org/docs/current/pgstatstatements.html)) for Aiven for PostgreSQL services. * the `performance_schema.events_statements_summary_by_digest` (refer to [documentation on the events information from the performance schema](https://dev.mysql.com/doc/refman/8.4/en/performance-schema-statement-summary-tables.html)) for Aiven for MySQL services. A description of the retrieved columns for Aiven for PostgreSQL can be found in the dedicated [PostgreSQL documentation](https://www.postgresql.org/docs/current/pgstatstatements.html). | Parameter | Information | | -------------- | ---------------------------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--format` | The format string for output defining the query metrics to retrieve, for example, `'{calls} {total_time}'` | **Example:** List the queries for an Aiven for PostgreSQL service named `pg-demo` including the query blurb, number of calls and both total and mean execution time. ``` avn service queries pg-demo --format '{query},{calls},{total_time},{mean_time}' ``` ### `avn service queries-reset`[​](#avn-service-queries-reset "Direct link to avn-service-queries-reset") Resets service connections/queries statistics for an Aiven for PostgreSQL® or Aiven for MySQL service. Resetting query statistics can be useful to measure database behaviour in a precise point in time or after a change has been deployed. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** Reset the queries for a service named `pg-demo`. ``` avn service queries-reset pg-demo ``` ### `avn service schema`[​](#avn-service-schema "Direct link to avn-service-schema") Service Schema commands ### `avn service schema-registry-acl`[​](#avn-service-schema-registry-acl "Direct link to avn-service-schema-registry-acl") Manages [Aiven for Apache Kafka® Karapace schema registry authorization](/docs/products/kafka/karapace/concepts/schema-registry-authorization.md). More information on `schema-registry-acl-add`, `schema-registry-acl-delete`, `schema-registry-acl-list` can be found in [the dedicated page](/docs/tools/cli/service/schema-registry-acl.md). ### `avn service sstableloader`[​](#avn-service-sstableloader "Direct link to avn-service-sstableloader") Service `sstableloader` commands ### `avn service tags`[​](#avn-service-tags "Direct link to avn-service-tags") Manage service tags. More information on `tags list`, `tags replace` and `tags update` can be found in [the dedicated page](/docs/tools/cli/service/tags.md). ### `avn service task-create`[​](#avn-service-task-create "Direct link to avn-service-task-create") Create a service task | Parameter | Information | | ---------------------- | ---------------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--project` | Project name (defaults to `None`) | | `--operation` | Task operation (default: `upgrade_check`, possible values: `migration_check`, `upgrade_check`) | | `--target-version` | Upgrade target version (used for PostgreSQL) | | `--source-service-uri` | Migration: source URI for migration | | `--ignore-dbs` | Migration: comma-separated list of databases to be ignored (MySQL only) | | `--format` | Format string for output, for example, `{name} {retention_hours}` | | `--json` | Raw JSON output | **Example:** Create a migration task to migrate a MySQL database to Aiven to the service `mysql` in project `myproj` ``` avn service task-create --operation migration_check --source-service-uri mysql://user:password@host:port/databasename --project myproj mysql ``` An example `avn service task-create` output: ``` TASK_TYPE SUCCESS TASK_ID ===================== ======= ==================================== mysql_migration_check null e2df7736-66c5-4696-b6c9-d33a0fc4cbed ``` ### `avn service task-get`[​](#avn-service-task-get "Direct link to avn-service-task-get") Get details for a single task for your service | Parameter | Information | | -------------- | ----------------------------------------------------------------- | | `service_name` | The name of the service | | `--project` | Project name (defaults to `None`) | | `--task-id` | The task ID to check | | `--format` | Format string for output, for example, `{name} {retention_hours}` | | `--json` | Raw JSON output | **Example:** Check the status of your migration task with id `e2df7736-66c5-4696-b6c9-d33a0fc4cbed` for the service named `mysql` in the `myproj` project ``` avn service task-get --task-id e2df7736-66c5-4696-b6c9-d33a0fc4cbed --project myproj mysql ``` Example output ``` TASK_TYPE SUCCESS TASK_ID RESULT ===================== ======= ==================================== ==================================================================================== mysql_migration_check true e2df7736-66c5-4696-b6c9-d33a0fc4cbed All pre-checks passed successfully, preferred migration method will be [Replication] ``` ### `avn service terminate`[​](#avn-cli-service-terminate "Direct link to avn-cli-service-terminate") Permanently deletes a service. warning The `terminate` command deletes the service and the associated data. The data is not recoverable. To temporarily shut down the service use the [service update command](/docs/tools/cli/service-cli.md#avn-cli-service-update): `avn service update SERVICE_NAME --power-off` | Parameter | Information | | -------------- | ----------------------------------------------- | | `service_name` | The name of the service | | `--force` | Force the action without requiring confirmation | **Example:** Terminate the service named `demo-pg`. ``` avn service terminate demo-pg ``` note To avoid accidental service deletion, enable the termination protection during service [creation](/docs/tools/cli/service-cli.md#avn-cli-service-create) or [update](/docs/tools/cli/service-cli.md#avn-cli-service-update) by using the `--enable-termination-protection` flag ### `avn service topic`[​](#avn-service-topic "Direct link to avn-service-topic") Manages Aiven for Apache Kafka® topics. More information on `topic-create`, `topic-delete`, `topic-list` and `topic-update` can be found in [the dedicated page](/docs/tools/cli/service/topic.md). ### `avn service types`[​](#avn-cli-service-type "Direct link to avn-cli-service-type") Lists the Aiven service types available in a project. **Example:** Retrieve all the services types available in the currently selected project. ``` avn service types ``` An example of `service types` output: ``` SERVICE_TYPE DESCRIPTION ================= =================================================================================== elasticsearch Elasticsearch - Search & Analyze Data in Real Time grafana Grafana - Metrics Dashboard kafka Kafka - High-Throughput Distributed Messaging System kafka_connect Kafka Connect - Kafka Connect service kafka_mirrormaker Kafka MirrorMaker - Kafka MirrorMaker service mysql MySQL - Relational Database Management System opensearch OpenSearch - Search & Analyze Data in Real Time, derived from Elasticsearch v7.10.2 pg PostgreSQL - Object-Relational Database Management System caching Caching - Compatible with legacy Redis® OSS ``` The service types command in verbose mode also shows all the configuration options for each type of service: ``` avn service types -v ``` You might find it helpful to pipe the output to `less` since there are a large number of options available and the command output is long. ### `avn service update`[​](#avn-cli-service-update "Direct link to avn-cli-service-update") Updates the settings for an Aiven service. | Parameter | Information | | ---------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--cloud` | The name of the cloud region where to deploy the service. See [avn-cloud-list](/docs/tools/cli/cloud.md#avn-cloud-list). | | `-c KEY=VALUE` | Apply a configuration setting. Run `avn service types -v` to view available values. | | `--disk-space-gib` | Total amount of disk space for data storage (GiB) | | `--plan` | Aiven subscription plan name. See [`avn_service_plan`](/docs/tools/cli/service-cli.md#avn-service-plan). | | `--power-on` | Power on the service | | `--power-off` | Power off the service | | `--maintenance-dow` | Set the automatic maintenance window's day of the week (possible values `monday`, `tuesday`, `wednesday`, `thursday`, `friday`, `saturday`, `sunday`, `never`) | | `--maintenance-time` | Set the automatic maintenance window's start time (`HH:MM:SS`) | | `--enable-termination-protection` | Enable termination protection | | `--disable-termination-protection` | Disable termination protection | | `--project-vpc-id` | The ID of the project VPC to use for the service. The cloud of the project's VPC must match the service's cloud. | | `--no-project-vpc` | The service will not use any VPC | | `--force` | Force the action without requiring confirmation | **Example:** Update the service named `demo-pg`, move it to `azure-germany-north` region and enable termination protection. ``` avn service update demo-pg \ --cloud azure-germany-north \ --enable-termination-protection ``` **Example:** Update the service named `big-service` to scale it down to the `Business-4` plan. ``` avn service update big-service \ --plan business-4 ``` **Example:** Update the service named `secure-database` to only accept connections from the range `10.0.1.0/24` and the IP `10.25.10.12`. ``` avn service update secure-database \ -c ip_filter=10.0.1.0/24,10.25.10.1/32 ``` note There is no whitespace between the IP addresses and comma in the command. **Example:** Update the Kafka version of the service named `kafka-service`. ``` avn service update \ kafka-service -c kafka_version=X.X ``` note This also works for other service types. To see a full list of configuration parameters, have a look at `avn service types -v` ### `avn service user`[​](#avn-service-user "Direct link to avn-service-user") Manages Aiven users and credentials. More information on `user-create`, `user-creds-acknowledge`, `user-creds-download`, `user-delete`, `user-get`, `user-kafka-java-creds`, `user-list`, `user-password-reset` and `user-set-access-control` can be found in [the dedicated page](/docs/tools/cli/service/user.md). ### `avn service versions`[​](#avn-service-versions "Direct link to avn-service-versions") For each service, lists the versions available together with: * `STATE`: if the version is `available` or `unavailable` * `AVAILABILITY_START_TIME` and `AVAILABILITY_END_TIME`: Period in which the specific version is available * `AIVEN_END_OF_LIFE_TIME`: Aiven deprecation date for the specific version * `UPSTREAM_END_OF_LIFE_TIME`: Upstream deprecation date for the specific version * `TERMINATION_TIME`: Termination time of the active instances * `END_OF_LIFE_HELP_ARTICLE_URL`: URL to "End of Life" documentation **Example:** List all service versions. ``` avn service versions ``` An example of `service versions` output: ``` SERVICE_TYPE MAJOR_VERSION STATE AVAILABILITY_START_TIME AVAILABILITY_END_TIME AIVEN_END_OF_LIFE_TIME UPSTREAM_END_OF_LIFE_TIME TERMINATION_TIME END_OF_LIFE_HELP_ARTICLE_URL ============= ============= =========== ======================= ===================== ====================== ========================= ================ ==================================================================================================== OpenSearch 7 unavailable 2020-08-27T00:00:00Z 2021-09-23T00:00:00Z 2022-03-23T00:00:00Z null null https://help.aiven.io/en/articles/5424825 OpenSearch 7.10 unavailable 2021-02-22T00:00:00Z 2021-09-23T00:00:00Z 2022-03-23T00:00:00Z null null https://help.aiven.io/en/articles/5424825 OpenSearch 7.9 unavailable 2020-08-27T00:00:00Z 2021-09-23T00:00:00Z 2022-03-23T00:00:00Z null null https://help.aiven.io/en/articles/5424825 kafka 2.3 unavailable 2019-09-05T00:00:00Z 2021-08-13T00:00:00Z 2021-08-13T00:00:00Z null null https://help.aiven.io/en/articles/4472730-eol-instructions-for-aiven-for-kafka kafka 2.4 unavailable 2019-10-21T00:00:00Z 2021-08-13T00:00:00Z 2021-08-13T00:00:00Z null null https://help.aiven.io/en/articles/4472730-eol-instructions-for-aiven-for-kafka ... pg 12 available 2019-11-18T00:00:00Z 2024-05-14T00:00:00Z 2024-11-14T00:00:00Z 2024-11-14T00:00:00Z null https://help.aiven.io/en/articles/2461799-how-to-perform-a-postgresql-in-place-major-version-upgrade pg 13 available 2021-02-15T00:00:00Z 2025-05-13T00:00:00Z 2025-11-13T00:00:00Z 2025-11-13T00:00:00Z null https://help.aiven.io/en/articles/2461799-how-to-perform-a-postgresql-in-place-major-version-upgrade pg 9.6 unavailable 2016-09-29T00:00:00Z 2021-05-11T00:00:00Z 2021-11-11T00:00:00Z 2021-11-11T00:00:00Z null https://help.aiven.io/en/articles/2461799-how-to-perform-a-postgresql-in-place-major-version-upgrade ``` ### `avn service wait`[​](#avn-service-wait "Direct link to avn-service-wait") Waits for the service to reach the `RUNNING` state | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** Wait for the service named `pg-doc` to reach the `RUNNING` state. ``` avn service wait pg-doc ``` *** *Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries.* --- # avn service acl Full list of commands for `avn service acl`. ## Manage Aiven ACL[​](#manage-aiven-acl "Direct link to Manage Aiven ACL") The `avn service acl` command manages access control lists (ACLs) in Aiven for Apache Kafka®. ACLs define permissions for accessing topics and controlling user access. They support wildcard patterns (`*` and `?`) for both topics and usernames. Supported permissions are `read`, `write`, and `readwrite`. ### `avn service acl-add`[​](#avn-service-acl-add "Direct link to avn-service-acl-add") Add an Aiven for Apache Kafka® ACL entry. | Parameter | Information | | -------------- | ------------------------------------------------------------------- | | `service_name` | Name of the service | | `--permission` | Permission type: possible values are `read`, `write` or `readwrite` | | `--topic` | Topic name pattern: accepts `*` and `?` as wildcard characters | | `--username` | Username pattern: accepts `*` and `?` as wildcard characters | **Example:** Add an ACL for usernames ending with `userA` to have `readwrite` access to topics starting with `topic2020` in service `kafka-doc`. ``` avn service acl-add kafka-doc --username *userA --permission readwrite --topic topic2020* ``` ### `avn service acl-delete`[​](#avn-service-acl-delete "Direct link to avn-service-acl-delete") Delete an Aiven for Apache Kafka® ACL entry. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | Name of the service | | `acl_id` | ID of the ACL to delete | **Example:** Delete the ACL with ID `acl3604f96c74a` from the Aiven for Apache Kafka service `kafka-doc`. ``` avn service acl-delete kafka-doc acl3604f96c74a ``` ### `avn service acl-list`[​](#avn-service-acl-list "Direct link to avn-service-acl-list") List Aiven for Apache Kafka® ACL entries. | Parameter | Information | | -------------- | ------------------- | | `service_name` | Name of the service | **Example:** List ACLs defined for service `kafka-doc`. ``` avn service acl-list kafka-doc ``` Example output of `avn service acl-list`: ``` ID USERNAME TOPIC PERMISSION ============== ======== ========= ========== default * * admin acl3604f96c74a Jon orders readwrite acl3604fa706cb Frida invoices* write ``` ## Related page[​](#related-page "Direct link to Related page") For managing Kafka-native ACLs, see [`avn service kafka-acl`](/docs/tools/cli/service/kafka-acl.md). --- # avn service connection-info Full list of commands for `avn service connection-info`. ## Retrieve connection information[​](#avn_cli_service_connection_info_kcat "Direct link to Retrieve connection information") ### `avn service connection-info kafkacat`[​](#avn-service-connection-info-kafkacat "Direct link to avn-service-connection-info-kafkacat") Retrieves the `kcat` command necessary to connect to an Aiven for Apache Kafka® service and produce/consume messages to topics, learn more in [Use `kcat` with Aiven for Apache Kafka®](/docs/products/kafka/howto/kcat.md). | Parameter | Information | | ------------------------------- | ------------------------------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--route` | The type of route to use to connect to the service. Possible values are `dynamic`, `privatelink` and `public` | | `--privatelink-connection-id` | The ID of the privatelink to use | | `--kafka-authentication-method` | The Aiven for Apache Kafka® authentication method. Possible values are `certificate` and `sasl` | | `--username` | The username used to connect if using `sasl` authentication method | | `--ca` | The path to the CA certificate file | | `--client-cert` | The path to the client certificate file | | `--client-key` | The path to the client key file | | `--write` | Save the certificate and key files if not existing | | `--overwrite` | Save (or overwrite if already existing) the certificate and key files | **Example:** Retrieve the `kcat` command to connect to an Aiven for Apache Kafka service named `demo-kafka` with SSL authentication (`certificate`), download the certificates necessary for the connection: ``` avn service connection-info kafkacat demo-kafka --write ``` An example of `service connection-info kafkacat` output: ``` kafkacat -b demo-kafka-dev-advocates.aivencloud.com:13041 -X security.protocol=SSL -X ssl.ca.location=ca.pem -X ssl.key.location=service.key -X ssl.certificate.location=service.crt ``` warning The command output uses the old `kafkacat` naming. To be able to execute `kcat` commands, replace `kafkacat` with `kcat`. ### `avn service connection-info pg string`[​](#avn-service-connection-info-pg-string "Direct link to avn-service-connection-info-pg-string") Retrieves the connection parameters for a certain Aiven for PostgreSQL® service. | Parameter | Information | | ----------------------------- | -------------------------------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--route` | The type of route to use to connect to the service. Possible values are `dynamic`, `privatelink` and `public`. | | `--usage` | The database connection usage. Possible values are `primary` and `replica` | | `--privatelink-connection-id` | The ID of the privatelink to use | | `--username` | The username used to connect if using `sasl` authentication method | | `--dbname` | The database name to use to connect | | `--sslmode` | The `sslmode` to use. Possible values are `require`, `verify-ca`, `verify-full`, `disable`, `allow`, `prefer`. | **Example:** Retrieve the connection parameters for an Aiven for PostgreSQL® service named `demo-pg`: ``` avn service connection-info pg string demo-pg ``` An example of `avn service connection-info pg string` output: ``` host='demo-pg-dev-project.aivencloud.com' port='13039' user=avnadmin dbname='defaultdb' ``` ### `avn service connection-info pg uri`[​](#avn-service-connection-info-pg-uri "Direct link to avn-service-connection-info-pg-uri") Retrieves the connection URI for an Aiven for PostgreSQL® service. | Parameter | Information | | ----------------------------- | ------------------------------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--route` | The type of route to use to connect to the service. Possible values are `dynamic`, `privatelink` and `public` | | `--usage` | The database connection usage. Possible values are `primary` and `replica` | | `--privatelink-connection-id` | The ID of the privatelink to use | | `--username` | The username used to connect if using `sasl` authentication method | | `--dbname` | The database name to use to connect | | `--sslmode` | The `sslmode` to use. Possible values are `require`, `verify-ca`, `verify-full`, `disable`, `allow`, `prefer` | **Example:** Retrieve the connection URI for an Aiven for PostgreSQL® service named `demo-pg`: ``` avn service connection-info pg uri demo-pg ``` An example of `avn service connection-info pg uri` output: ``` postgres://avnadmin:XXXXXXXXXX@demo-pg-dev-project.aivencloud.com:13039/defaultdb?sslmode=require ``` ### `avn service connection-info psql`[​](#avn-service-connection-info-psql "Direct link to avn-service-connection-info-psql") Retrieves the `psql` command needed to connect to an Aiven for PostgreSQL® service. | Parameter | Information | | ----------------------------- | ------------------------------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--route` | The type of route to use to connect to the service. Possible values are `dynamic`, `privatelink` and `public` | | `--usage` | The database connection usage. Possible values are `primary` and `replica` | | `--privatelink-connection-id` | The Id of the privatelink to use | | `--username` | The username used to connect if using `sasl` authentication method | | `--dbname` | The database name to use to connect | | `--sslmode` | The `sslmode` to use. Possible values are `require`, `verify-ca`, `verify-full`, `disable`, `allow`, `prefer` | **Example:** Retrieve the `psql` command needed to connect to an Aiven for PostgreSQL® service named `demo-pg`: ``` avn service connection-info psql demo-pg ``` An example of `avn service connection-info psql` output: ``` psql postgres://avnadmin:XXXXXXXXXXXX@demo-pg-dev-advocates.aivencloud.com:13039/defaultdb?sslmode=require ``` ### `avn service connection-info redis uri`[​](#avn-service-connection-info-redis-uri "Direct link to avn-service-connection-info-redis-uri") Retrieves the connection URI needed to connect to an Aiven for Caching service. | Parameter | Information | | ----------------------------- | ------------------------------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--route` | The type of route to use to connect to the service. Possible values are `dynamic`, `privatelink` and `public` | | `--usage` | The database connection usage. Possible values are `primary` and `replica` | | `--privatelink-connection-id` | The ID of the privatelink to use | | `--username` | The username used to connect if using `sasl` authentication method | | `--db` | The database name to use to connect | **Example:** Retrieve the connection URI needed to connect to an Aiven for Caching service named `demo-redis`: ``` avn service connection-info redis uri demo-redis ``` An example of `avn service connection-info redis uri` output: ``` rediss://default:XXXXXXXXXX@demo-redis-dev-project.aivencloud.com:13040 ``` --- # avn service connection-pool Full list of commands for `avn service connection-pool`. ## Manage PgBouncer connection pools[​](#manage-pgbouncer-connection-pools "Direct link to Manage PgBouncer connection pools") ### `avn service connection-pool-create`[​](#avn-service-connection-pool-create "Direct link to avn-service-connection-pool-create") Creates a new [PgBouncer connection pool](/docs/products/postgresql/concepts/pg-connection-pooling.md) for a given PostgreSQL® service. | Parameter | Information | | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--pool-name` | The name of the connection pool | | `--dbname` | The name of the database | | `--username` | The database username to use for the connection pool | | `--pool-size` | Size of the connection pool in number of connections | | `--pool-mode` | The [pool mode](/docs/products/postgresql/concepts/pg-connection-pooling.md#pooling-modes). Possible values are `transaction`, `session` and `statement` | **Example:** In the service `demo-pg` Create a connection pool named `cp-analytics-it` for the database `it-analytics` with: * username `avnadmin` * pool-size of `10` connections * `transaction` pool-mode ``` avn service connection-pool-create demo-pg \ --pool-name cp-analytics-it \ --dbname analytics-it \ --username avnadmin \ --pool-size 10 \ --pool-mode transaction ``` ### `avn service connection-pool-delete`[​](#avn-service-connection-pool-delete "Direct link to avn-service-connection-pool-delete") Deletes a [PgBouncer connection pool](/docs/products/postgresql/concepts/pg-connection-pooling.md) for a given PostgreSQL® service. | Parameter | Information | | -------------- | ------------------------------- | | `service_name` | The name of the service | | `--pool-name` | The name of the connection pool | **Example:** In the service `demo-pg` delete a connection pool named `cp-analytics-it`. ``` avn service connection-pool-delete demo-pg \ --pool-name cp-analytics-it ``` ### `avn service connection-pool-list`[​](#avn-service-connection-pool-list "Direct link to avn-service-connection-pool-list") Lists the [PgBouncer connection pool](/docs/products/postgresql/concepts/pg-connection-pooling.md) for a given PostgreSQL® service. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** List the connection pools available in the service `demo-pg`. ``` avn service connection-pool-list demo-pg ``` An example of `avn service connection-pool-list` output: ``` POOL_NAME DATABASE USERNAME POOL_MODE POOL_SIZE =============== ============ ======== =========== ========= cp-analytics-it analytics-it avnadmin transaction 10 cp-sales sales-it test-usr session 20 ``` ### `avn service connection-pool-update`[​](#avn-service-connection-pool-update "Direct link to avn-service-connection-pool-update") Updates a [PgBouncer connection pool](/docs/products/postgresql/concepts/pg-connection-pooling.md) for a given PostgreSQL® service. | Parameter | Information | | -------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--pool-name` | The name of the connection pool | | `--dbname` | The name of the database | | `--username` | The database username to use for the connection pool | | `--pool-size` | Size of the connection pool in number of connections | | `--pool-mode` | The [pool mode](/docs/products/postgresql/concepts/pg-connection-pooling.md#pooling-modes). Possible values are `transaction`, `session` and `statement` | **Example:** In the service `demo-pg` update the connection pool named `cp-analytics-it` for the database `it-analytics` with: * username `avnadmin` * pool-size of `20` connections * `session` pool-mode ``` avn service connection-pool-update demo-pg \ --pool-name cp-analytics-it \ --dbname analytics-it \ --username avnadmin \ --pool-size 20 \ --pool-mode session ``` --- # avn service connector Full list of commands for `avn service connector`. ## Manage Apache Kafka® Connect connectors details[​](#manage-apache-kafka-connect-connectors-details "Direct link to Manage Apache Kafka® Connect connectors details") Commands for managing Aiven for Apache Kafka® Connect connectors via `avn` commands. ### `avn service connector available`[​](#avn_cli_service_connector_available "Direct link to avn_cli_service_connector_available") Lists Apache Kafka® Connect connector plugins available in a given Aiven for Apache Kafka® service. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the Service | **Example:** List the Kafka Connect connector plugins available for the service `kafka-demo`. ``` avn service connector available kafka-demo ``` ### `avn service connector create`[​](#avn_service_connector_create "Direct link to avn_service_connector_create") Creates a new Apache Kafka® Connect connector in a given Aiven for Apache Kafka® service. | Parameter | Information | | ------------------ | ------------------------------------------------------------------------------------------ | | `service_name` | The name of the Service | | `connector_config` | JSON string or path (preceded by `@`) to a Kafka Connect connector JSON configuration file | **Example:** Create a JDBC source Kafka Connect connector in the service `kafka-demo` passing the JSON configuration string. ``` avn service connector create kafka-demo '{ "name": "pg-bulk-invoices-source", "connector.class": "io.aiven.connect.jdbc.JdbcSourceConnector", "connection.url": "jdbc:postgresql://demo-pg-myinventedprojectname.aivencloud.com:13039/defaultdb?sslmode=require", "connection.user": "avnadmin", "connection.password": "verysecurepassword123", "table.whitelist": "invoices", "mode": "bulk", "poll.interval.ms": "10000", "topic.prefix": "pg_source_" }' ``` ### `avn service connector delete`[​](#avn-service-connector-delete "Direct link to avn-service-connector-delete") Deletes an Apache Kafka® Connect connector in a given Aiven for Apache Kafka® service. | Parameter | Information | | -------------- | ---------------------------- | | `service_name` | The name of the Service | | `connector` | Kafka Connect connector name | **Example:** Delete the Kafka Connect connector named `pg-bulk-invoices-source` in the service `kafka-demo`. ``` avn service connector delete kafka-demo pg-bulk-invoices-source ``` ### `avn service connector list`[​](#avn-service-connector-list "Direct link to avn-service-connector-list") Lists Apache Kafka® Connect connectors in a given Aiven for Apache Kafka® service. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the Service | **Example:** List all Kafka Connect connectors in the service `kafka-demo`. ``` avn service connector list kafka-demo ``` An example of `avn service connector list` output: ``` { "connectors": [ { "config": { "connection.password": "verysecurepassword123", "connection.url": "jdbc:postgresql://demo-test-myinventedprojectname.aivencloud.com:13039/defaultdb?sslmode=require", "connection.user": "avnadmin", "connector.class": "io.aiven.connect.jdbc.JdbcSourceConnector", "mode": "bulk", "name": "pg-bulk-invoices-source", "poll.interval.ms": "10000", "table.whitelist": "invoices", "topic.prefix": "pg_source_" }, "name": "pg-bulk-invoices-source", "plugin": { "author": "Aiven", "class": "io.aiven.connect.jdbc.JdbcSourceConnector", "docURL": "https://github.com/aiven/aiven-kafka-connect-jdbc/blob/master/docs/source-connector.md", "title": "JDBC Source", "type": "source", "version": "6.6.0" }, "tasks": [] } ] } ``` ### `avn service connector pause`[​](#avn-service-connector-pause "Direct link to avn-service-connector-pause") Pauses an Apache Kafka® Connect connector in a given Aiven for Apache Kafka® service. | Parameter | Information | | -------------- | ---------------------------- | | `service_name` | The name of the Service | | `connector` | Kafka Connect connector name | **Example:** Pause the Kafka Connect connector named `pg-bulk-invoices-source` in the service `kafka-demo`. ``` avn service connector pause kafka-demo pg-bulk-invoices-source ``` ### `avn service connector restart`[​](#avn-service-connector-restart "Direct link to avn-service-connector-restart") Restarts an Apache Kafka® Connect connector in a given Aiven for Apache Kafka® service. | Parameter | Information | | -------------- | ---------------------------- | | `service_name` | The name of the Service | | `connector` | Kafka Connect connector name | **Example:** Restart the Kafka Connect connector named `pg-bulk-invoices-source` in the service `kafka-demo`. ``` avn service connector restart kafka-demo pg-bulk-invoices-source ``` ### `avn service connector restart-task`[​](#avn-service-connector-restart-task "Direct link to avn-service-connector-restart-task") Restarts an Apache Kafka® Connect connector task in a given Aiven for Apache Kafka® service. | Parameter | Information | | -------------- | ------------------------------- | | `service_name` | The name of the Service | | `connector` | Kafka Connect connector name | | `task` | Kafka Connect connector task id | **Example:** Restart the task with id `0` in the Kafka Connect connector named `pg-bulk-invoices-source` belonging to the service `kafka-demo`. ``` avn service connector restart-task kafka-demo pg-bulk-invoices-source 0 ``` ### `avn service connector resume`[​](#avn-service-connector-resume "Direct link to avn-service-connector-resume") Resumes an Apache Kafka® Connect connector in a given Aiven for Apache Kafka® service. | Parameter | Information | | -------------- | ---------------------------- | | `service_name` | The name of the Service | | `connector` | Kafka Connect connector name | **Example:** Resume the Kafka Connect connector named `pg-bulk-invoices-source` belonging to the service `kafka-demo`. ``` avn service connector resume kafka-demo pg-bulk-invoices-source ``` ### `avn service connector schema`[​](#avn-service-connector-schema "Direct link to avn-service-connector-schema") Retrieves the configuration information for an Apache Kafka® Connect connector plugin in a given Aiven for Apache Kafka® service. | Parameter | Information | | -------------- | ----------------------------------------- | | `service_name` | The name of the Service | | `connector` | Kafka Connect connector plugin class name | **Example:** Retrieve the schema for the Kafka Connect plugin with class `io.debezium.connector.sqlserver.SqlServerConnector` belonging to the service `kafka-demo`. ``` avn service connector schema kafka-demo io.debezium.connector.sqlserver.SqlServerConnector ``` ### `avn service connector status`[​](#avn-service-connector-status "Direct link to avn-service-connector-status") Gets an Apache Kafka® Connect connector status in a given Aiven for Apache Kafka service. | Parameter | Information | | -------------- | ---------------------------- | | `service_name` | The name of the Service | | `connector` | Kafka Connect connector name | **Example:** Check the status of a Kafka Connect connector named `pg-bulk-invoices-source` belonging to the service `kafka-demo`. ``` avn service connector status kafka-demo pg-bulk-invoices-source ``` An example of `avn service connector status` output: ``` { "status": { "state": "RUNNING", "tasks": [ { "id": 0, "state": "RUNNING", "trace": "" } ] } } ``` ### `avn service connector stop`[​](#avn-service-connector-stop "Direct link to avn-service-connector-stop") Stops an Apache Kafka® Connect connector in a given Aiven for Apache Kafka® service. | Parameter | Information | | -------------- | ---------------------------- | | `service_name` | The name of the Service | | `connector` | Kafka Connect connector name | **Example:** Stop the Kafka Connect connector named `pg-bulk-invoices-source` in the service `kafka-demo`. ``` avn service connector stop kafka-demo pg-bulk-invoices-source ``` ### `avn service connector update`[​](#avn-service-connector-update "Direct link to avn-service-connector-update") Updates an Apache Kafka® Connect connector in a given Aiven for Apache Kafka® service. | Parameter | Information | | ------------------ | ------------------------------------------------------------------------------------------ | | `service_name` | The name of the Service | | `connector` | Kafka Connect connector name | | `connector_config` | JSON string or path (preceded by `@`) to a Kafka Connect connector JSON configuration file | **Example:** Update a the JDBC source Kafka Connect connector named `pg-bulk-invoices-source` in the service `kafka-demo` with the JSON configuration string contained in the file `kafka-connect-config.json`. ``` avn service connector update kafka-demo pg-bulk-invoices-source @kafka-connect-config.json ``` --- # avn service database Full list of commands for `avn service database`. ## Manage databases[​](#manage-databases "Direct link to Manage databases") ### `avn service database-create`[​](#avn-service-database-create "Direct link to avn-service-database-create") Creates a database within an Aiven for PostgreSQL® or Aiven for MySQL. | Parameter | Information | | -------------- | ------------------------ | | `service_name` | The name of the service | | `--dbname` | The name of the database | **Example:** Create a database named `analytics-it` within the service named `pg-demo`. ``` avn service database-create pg-demo --dbname analytics-it ``` ### `avn service database-delete`[​](#avn-service-database-delete "Direct link to avn-service-database-delete") Removes a specific database within an Aiven for PostgreSQL® or Aiven for MySQL. | Parameter | Information | | -------------- | ------------------------ | | `service_name` | The name of the service | | `--dbname` | The name of the database | **Example:** Delete the database named `analytics-it` within the service named `pg-demo` ``` avn service database-delete pg-demo --dbname analytics-it ``` ### `avn service database-list`[​](#avn-service-database-list "Direct link to avn-service-database-list") Lists the service databases available in an Aiven for PostgreSQL® or Aiven for MySQL. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** List the service databases within the service named `pg-demo` ``` avn service database-list pg-demo ``` --- # avn service es-acl Full list of commands for `avn service es-acl`. ## Manage OpenSearch® access control lists[​](#manage-opensearch-access-control-lists "Direct link to Manage OpenSearch® access control lists") ### `avn service es-acl-add`[​](#avn-service-es-acl-add "Direct link to avn-service-es-acl-add") Add rules to OpenSearch ACL configuration ### `avn service es-acl-del`[​](#avn-service-es-acl-del "Direct link to avn-service-es-acl-del") Delete rules from OpenSearch ACL configuration ### `avn service es-acl-disable`[​](#avn-service-es-acl-disable "Direct link to avn-service-es-acl-disable") Disable OpenSearch ACL configuration ### `avn service es-acl-enable`[​](#avn-service-es-acl-enable "Direct link to avn-service-es-acl-enable") Enable OpenSearch ACL configuration ### `avn service es-acl-extended-disable`[​](#avn-service-es-acl-extended-disable "Direct link to avn-service-es-acl-extended-disable") Disable OpenSearch Extended ACL ### `avn service es-acl-extended-enable`[​](#avn-service-es-acl-extended-enable "Direct link to avn-service-es-acl-extended-enable") Enable OpenSearch Extended ACL ### `avn service es-acl-list`[​](#avn-service-es-acl-list "Direct link to avn-service-es-acl-list") List OpenSearch ACL configuration --- # avn service flink Full list of commands for `avn service flink`. warning The Aiven for Apache Flink® CLI commands have been updated, and to execute them, you must use `aiven-client` version `2.18.0`. ## Manage Aiven for Apache Flink® applications[​](#manage-aiven-for-apache-flink-applications "Direct link to Manage Aiven for Apache Flink® applications") ### `avn service flink create-application`[​](<#avn service flink create-application> "Direct link to avn service flink create-application") Create an Aiven for the Apache Flink® application in the specified service and project. | Parameter | Information | | ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `project` | The name of the project | | `service_name` | The name of the service | | `application_properties` | Application properties definition for Aiven for Flink application, either as a JSON string or a file path (prefixed with '@') containing the JSON configuration | The `application_properties` parameter should contain the following common properties in JSON format: | Parameter | Information | | --------------------- | ---------------------------------------- | | `name` | The name of the application | | `application_version` | Optional: The version of the application | **Example:** Creates an Aiven for Apache Flink application named `DemoApp` in the service `flink-democli` and project `my-project`. ``` avn service flink create-application flink-democli \ --project my-project \ "{\"name\":\"DemoApp\"}" ``` An example of `avn service flink create-application` output: ``` { "application_versions": [], "created_at": "2023-02-08T07:37:25.165996Z", "created_by": "wilma@example.com", "id": "2b29f4aa-a496-4fca-8575-23544415606e", "name": "DemoApp", "updated_at": "2023-02-08T07:37:25.165996Z", "updated_by": "wilma@example.com" } ``` ### `avn service flink list-applications`[​](#avn-service-flink-list-applications "Direct link to avn-service-flink-list-applications") Lists all the Aiven for Apache Flink® applications in a specified project and service. | Parameter | Information | | -------------- | ----------------------- | | `project` | The name of the project | | `service_name` | The name of the service | **Example:** Lists all the Aiven for Flink applications for the service `flink-democli` in the project `my-project`. ``` avn service flink list-applications flink-democli \ --project my-project ``` An example of `avn service flink list-applications` output: ``` { "applications": [ { "created_at": "2023-02-08T07:37:25.165996Z", "created_by": "wilma@example.com", "id": "2b29f4aa-a496-4fca-8575-23544415606e", "name": "DemoApp", "updated_at": "2023-02-08T07:37:25.165996Z", "updated_by": "wilma@example.com" } ] } ``` ### `avn service flink get-application`[​](#avn-service-flink-get-application "Direct link to avn-service-flink-get-application") Retrieves the information about the Aiven for Flink® applications in a specified project and service. | Parameter | Information | | ---------------- | ------------------------------------------------------------------------ | | `project` | The name of the project | | `service_name` | The name of the service | | `application-id` | The ID of the Aiven for Flink application to retrieve information about. | **Example:** Retrieves information about Aiven for Flink® application with application-id `2b29f4aa-a496-4fca-8575-23544415606e` for service `flink-democli` and project `my-project` ``` avn service flink get-application flink-democli \ --project my-project \ --application-id 2b29f4aa-a496-4fca-8575-23544415606e ``` An example of `avn service flink list-applications` output: ``` { "application_versions": [], "created_at": "2023-02-08T07:37:25.165996Z", "created_by": "wilma@example.com", "id": "2b29f4aa-a496-4fca-8575-23544415606e", "name": "DemoApp", "updated_at": "2023-02-08T07:37:25.165996Z", "updated_by": "wilma@example.com" } ``` ### `avn service flink update-application`[​](#avn-service-flink-update-application "Direct link to avn-service-flink-update-application") Update an Aiven for Flink® application in a specified project and service. | Parameter | Information | | ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `project` | The name of the project | | `service_name` | The name of the service | | `application-id` | The ID of the Aiven for Flink application to update | | `application-properties` | Application properties definition for Aiven for Flink® application, either as a JSON string or a file path (prefixed with '@') containing the JSON configuration | The `application_properties` parameter should contain the following common properties in JSON format | Parameter | Information | | --------- | --------------------------- | | `name` | The name of the application | **Example:** Updates the name of the Aiven for Flink application from `Demo` to `DemoApp` for application-id `986b2d5f-7eda-480c-bcb3-0f903a866222` in the service `flink-democli` and project `my-project`. ``` avn service flink update-application flink-democli \ --project my-project \ --application-id 986b2d5f-7eda-480c-bcb3-0f903a866222 \ "{\"name\":\"DemoApp\"}" ``` ### `avn service flink delete-application`[​](#avn--service-flink-delete-application "Direct link to avn--service-flink-delete-application") Delete an Aiven for Flink® application in a specified project and service. | Parameter | Information | | ---------------- | --------------------------------------------------- | | `project` | The name of the project | | `service_name` | The name of the service | | `application-id` | The ID of the Aiven for Flink application to delete | **Example:** Deletes the Aiven for Flink application with application-id `64192db8-d073-4e28-956b-82c71b016e3e` for the service `flink-democli` in the project `my-project`. ``` avn service flink delete-application flink-democli \ --project my-project \ --application-id 64192db8-d073-4e28-956b-82c71b016e3e ``` ### `avn service flink create-application-version`[​](#avn-service-flink-create-application-version "Direct link to avn-service-flink-create-application-version") Create an Aiven for Flink® application version in a specified project and service. warning Before creating an application, [create integrations](/docs/products/flink/howto/create-integration.md) between Aiven for Apache Flink and the source/sinks data services. As of now you can define integration with: * Aiven for Apache Kafka® as source/sink * Aiven for Apache PostgreSQL® as source/sink * Aiven for OpenSearch® as sink Sinking data using the [Slack connector](/docs/products/flink/howto/slack-connector.md), doesn't need an integration. **Example**: to create an integration between an Aiven for Apache Flink service named `flink-democli` and an Aiven for Apache Kafka service named `demo-kafka` you can use the following command: ``` avn service integration-create \ --integration-type flink \ --dest-service flink-democli \ --source-service demo-kafka ``` All the available command integration options can be found in the [dedicated document](/docs/tools/cli/service/integration.md#avn_service_integration_create) | Parameter | Information | | -------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `project` | The name of the project | | `service_name` | The name of the service | | `application-id` | The ID of the Aiven for Flink application to create a version | | `application_version_properties` | Application version properties definition for Aiven for Flink® application, either as a JSON string or a file path (prefixed with '@') containing the JSON configuration | The `application_version_properties` parameter should contain the following common properties in JSON format: | Parameter | Information | | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `sinks` | An array of objects that contains the table creation statements creation statements of the sinks | | `create_table` | A string that defines the CREATE TABLE statement of the sink including the integration ID. The integration ID can be found with the [integration-list](/docs/tools/cli/service/integration.md#avn_service_integration_list) command | | `source` | An array of objects that contains the table creation statements of the source | | `create_table` | A string that defines the CREATE TABLE statement of the source including the integration ID. The integration ID can be found with the [integration-list](/docs/tools/cli/service/integration.md#avn_service_integration_list) command | | `statement` | The transformation SQL statement of the application | **Example:** Creates a new Aiven for Flink application version for application-id `986b2d5f-7eda-480c-bcb3-0f903a866222` with the following details: * **Source**: a table, named `special_orders` coming from an Apache Kafka® topic named `special_orders_topic` using the integration with id `4ec23427-9e9f-4827-90fa-ea9e38c31bc3` and the following columns: ``` id INT, name VARCHAR, topping VARCHAR ``` * **Sink**: a table, called `pizza_orders`, writing to an Apache Kafka® topic named `pizza_orders_topic` using the integration with id `4ec23427-9e9f-4827-90fa-ea9e38c31bc3` and the following columns: ``` id INT, name VARCHAR, topping VARCHAR ``` * **SQL statement**: ``` INSERT INTO special_orders SELECT id, name, c.topping FROM pizza_orders CROSS JOIN UNNEST(pizzas) b CROSS JOIN UNNEST(b.additionalToppings) AS c(topping) WHERE c.topping IN ('🍍 pineapple', '🍓 strawberry','🍌 banana') ``` ``` avn service flink create-application-version flink-democli \ --project my-project \ --application-id 986b2d5f-7eda-480c-bcb3-0f903a866222 \ """{ \"sources\": [ { \"create_table\": \"CREATE TABLE special_orders ( \ id INT, \ name VARCHAR, \ topping VARCHAR \ ) \ WITH ( \ 'connector' = 'kafka', \ 'properties.bootstrap.servers' = '', \ 'scan.startup.mode' = 'earliest-offset', \ 'value.fields-include' = 'ALL', \ 'topic' = 'special_orders_topic', \ 'value.format' = 'json' \ )\", \"integration_id\": \"4ec23427-9e9f-4827-90fa-ea9e38c31bc3\" } ], \"sinks\": [ { \"create_table\": \"CREATE TABLE pizza_orders ( \ id INT, \ shop VARCHAR, \ name VARCHAR, \ phoneNumber VARCHAR, \ address VARCHAR, \ pizzas ARRAY)>) \ WITH ( \ 'connector' = 'kafka', \ 'properties.bootstrap.servers' = '', \ 'scan.startup.mode' = 'earliest-offset', \ 'topic' = 'pizza_orders_topic', \ 'value.format' = 'json' \ )\", \"integration_id\": \"4ec23427-9e9f-4827-90fa-ea9e38c31bc3\" } ], \"statement\": \"INSERT INTO special_orders \ SELECT id, \ name, \ c.topping \ FROM pizza_orders \ CROSS JOIN UNNEST(pizzas) b \ CROSS JOIN UNNEST(b.additionalToppings) AS c(topping) \ WHERE c.topping IN ('🍍 pineapple', '🍓 strawberry','🍌 banana')\" }""" ``` ### `avn service flink validate-application-version`[​](#avn-service-flink-validate-application-version "Direct link to avn-service-flink-validate-application-version") Validates the Aiven for Flink® application version in a specified project and service. warning Before creating an application, [create integrations](/docs/products/flink/howto/create-integration.md) between Aiven for Apache Flink and the source/sinks data services. As of now you can define integration with: * Aiven for Apache Kafka® as source/sink * Aiven for Apache PostgreSQL® as source/sink * Aiven for OpenSearch® as sink Sinking data using the [Slack connector](/docs/products/flink/howto/slack-connector.md), doesn't need an integration. **Example**: to create an integration between an Aiven for Apache Flink service named `flink-democli` and an Aiven for Apache Kafka service named `demo-kafka` you can use the following command: ``` avn service integration-create \ --integration-type flink \ --dest-service flink-democli \ --source-service demo-kafka ``` All the available command integration options can be found in the [dedicated document](/docs/tools/cli/service/integration.md#avn_service_integration_create) | Parameter | Information | | -------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `project` | The name of the project | | `service_name` | The name of the service | | `application-id` | The ID of the Aiven for Flink application to create a version | | `application_version_properties` | Application version properties definition for Aiven for Flink application, either as a JSON string or a file path (prefixed with '@') containing the JSON configuration | The `application_version_properties` parameter should contain the following common properties in JSON format | Parameter | Information | | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `sinks` | An array of objects that contains the table creation statements creation statements of the sinks | | `create_table` | A string that defines the CREATE TABLE statement of the sink including the integration ID. The integration ID can be found with the [integration-list](/docs/tools/cli/service/integration.md#avn_service_integration_list) command | | `source` | An array of objects that contains the table creation statements of the source | | `create_table` | A string that defines the CREATE TABLE statement of the source including the integration ID. The integration ID can be found with the [integration-list](/docs/tools/cli/service/integration.md#avn_service_integration_list) command | | `statement` | The transformation SQL statement of the application | **Example:** Validates the Aiven for Flink application version for the application-id `986b2d5f-7eda-480c-bcb3-0f903a866222`. ``` avn service flink validate-application-version flink-democli \ --project my-project \ --application-id 986b2d5f-7eda-480c-bcb3-0f903a866222 \ """{ \"sources\": [ { \"create_table\": \"CREATE TABLE special_orders ( \ id INT, \ name VARCHAR, \ topping VARCHAR \ ) \ WITH ( \ 'connector' = 'kafka', \ 'properties.bootstrap.servers' = '', \ 'scan.startup.mode' = 'earliest-offset', \ 'value.fields-include' = 'ALL', \ 'topic' = 'special_orders_topic', \ 'value.format' = 'json' \ )\", \"integration_id\": \"4ec23427-9e9f-4827-90fa-ea9e38c31bc3\" } ], \"sinks\": [ { \"create_table\": \"CREATE TABLE pizza_orders ( \ id INT, \ shop VARCHAR, \ name VARCHAR, \ phoneNumber VARCHAR, \ address VARCHAR, \ pizzas ARRAY)>) \ WITH ( \ 'connector' = 'kafka', \ 'properties.bootstrap.servers' = '', \ 'scan.startup.mode' = 'earliest-offset', \ 'topic' = 'pizza_orders_topic', \ 'value.format' = 'json' \ )\", \"integration_id\": \"4ec23427-9e9f-4827-90fa-ea9e38c31bc3\" } ], \"statement\": \"INSERT INTO special_orders \ SELECT id, \ name, \ c.topping \ FROM pizza_orders \ CROSS JOIN UNNEST(pizzas) b \ CROSS JOIN UNNEST(b.additionalToppings) AS c(topping) \ WHERE c.topping IN ('🍍 pineapple', '🍓 strawberry','🍌 banana')\" }""" ``` ### `avn service flink get-application-version`[​](#avn-service-flink-get-application-version "Direct link to avn-service-flink-get-application-version") Retrieves information about a specific version of an Aiven for Flink® application in a specified project and service. | Parameter | Information | | ------------------------ | ------------------------------------------------------------------------------- | | `project` | The name of the project | | `service_name` | The name of the service | | `application-id` | The ID of the Aiven for Flink application | | `application-version-id` | The ID of the Aiven for Flink application version to retrieve information about | **Example:** Retrieves the information specific to the Aiven for Flink® application for the service `flink-demo-cli` and project `my-project` with: * Application id: `986b2d5f-7eda-480c-bcb3-0f903a866222` * Application version id: `7a1c6266-64da-4f6f-a8b0-75207f997c8d` ``` avn service flink get-application-version flink-democli \ --project my-project \ --application-id 986b2d5f-7eda-480c-bcb3-0f903a866222 \ --application-version-id 7a1c6266-64da-4f6f-a8b0-75207f997c8d ``` ### `avn service flink delete-application-version`[​](#avn-service-flink-delete-application-version "Direct link to avn-service-flink-delete-application-version") Deletes a version of the Aiven for Flink® application in a specified project and service. | Parameter | Information | | ------------------------ | ----------------------------------------------------------- | | `project` | The name of the project | | `service_name` | The name of the service | | `application-id` | The ID of the Aiven for Flink application | | `application-version-id` | The ID of the Aiven for Flink application version to delete | **Example:** Delete the Aiven for Flink application version for service `flink-demo-cli` and project `my-project` with: * Application id: `986b2d5f-7eda-480c-bcb3-0f903a866222` * Application version id: `7a1c6266-64da-4f6f-a8b0-75207f997c8d` ``` avn service flink delete-application-version flink-democli \ --project my-project \ --application-id 986b2d5f-7eda-480c-bcb3-0f903a866222 \ --application-version-id 7a1c6266-64da-4f6f-a8b0-75207f997c8d ``` ### `avn service flink list-application-deployments`[​](#avn-service-flink-list-application-deployments "Direct link to avn-service-flink-list-application-deployments") Lists all the Aiven for Flink® application deployments in a specified project and service. | Parameter | Information | | ---------------- | ----------------------------------------- | | `project` | The name of the project | | `service_name` | The name of the service | | `application-id` | The ID of the Aiven for Flink application | **Example:** Lists all the Aiven for Flink application deployments for application-id `f171af72-fdf0-442c-947c-7f6a0efa83ad` for the service `flink-democli`, in the project `my-project`. ``` avn service flink list-application-deployments flink-democli \ --project my-project \ --application-id f171af72-fdf0-442c-947c-7f6a0efa83ad ``` ### `avn service flink get-application-deployment`[​](#avn-service-flink-get-application-deployment "Direct link to avn-service-flink-get-application-deployment") Retrieves information about an Aiven for Flink® application deployment in a specified project and service. | Parameter | Information | | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `project` | The name of the project | | `service_name` | The name of the service | | `application-id` | The ID of the Aiven for Flink application | | `deployment-id` | The ID of the Aiven for Flink application deployment. This ID can be obtained from the output of the `avn service flink list-application-deployments` command | **Example:** Retrieves the details of the Aiven for Flink application deployment for the application-id `f171af72-fdf0-442c-947c-7f6a0efa83ad`, deployment-id `bee0b5cb-01e7-49e6-bddb-a750caed4229` for the service `flink-democli`, in the project `my-project`. ``` avn service flink get-application-deployment flink-democli \ --project my-project \ --application-id f171af72-fdf0-442c-947c-7f6a0efa83ad \ --deployment-id bee0b5cb-01e7-49e6-bddb-a750caed4229 ``` ### `avn service flink create-application-deployment`[​](#avn-service-flink-create-application-deployment "Direct link to avn-service-flink-create-application-deployment") Creates a new Aiven for Flink® application deployment in a specified project and service. | Parameter | Information | | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `project` | The name of the project | | `service_name` | The name of the service | | `application-id` | The ID of the Aiven for Flink application | | `deployment_properties` | The deployment properties definition for Aiven for Flink application, either as a JSON string or a file path (prefixed with '@') containing the JSON configuration | The `deployment_properties` parameter should contain the following common properties in JSON format | Parameter | Information | | -------------------- | ----------------------------------------------------------- | | `parallelism` | The number of parallel instance for the task | | `restart_enabled` | Specifies whether a Flink Job is restarted in case it fails | | `starting_savepoint` | (Optional) The savepoint from where you want to deploy. | | `version_id` | The ID of the application version. | **Example:** Create an Aiven for Flink application deployment for the application id `986b2d5f-7eda-480c-bcb3-0f903a866222`. ``` avn service flink create-application-deployment flink-democli \ --project my-project \ --application-id 986b2d5f-7eda-480c-bcb3-0f903a866222 \ "{\"parallelism\": 1,\"restart_enabled\": true, \"version_id\": \"7a1c6266-64da-4f6f-a8b0-75207f997c8d\"}" ``` ### `avn service flink delete-application-deployment`[​](#avn-service-flink-delete-application-deployment "Direct link to avn-service-flink-delete-application-deployment") Deletes an Aiven for Flink® application deployment in a specified project and service. | Parameter | Information | | ---------------- | --------------------------------------------------------------- | | `project` | The name of the project | | `service_name` | The name of the service | | `application-id` | The ID of the Aiven for Flink® application | | `deployment-id` | The ID of the Aiven for Flink® application deployment to delete | **Example:** Deletes the Aiven for Flink application deployment with application-id `f171af72-fdf0-442c-947c-7f6a0efa83ad` and deployment-id `6d5e2c03-2235-44a5-ab8f-c544a4de04ef`. ``` avn service flink delete-application-deployment flink-democli \ --project my-project \ --application-id f171af72-fdf0-442c-947c-7f6a0efa83ad \ --deployment-id 6d5e2c03-2235-44a5-ab8f-c544a4de04ef ``` ### `avn service flink stop-application-deployment`[​](#avn-service-flink-stop-application-deployment "Direct link to avn-service-flink-stop-application-deployment") Stops a running Aiven for Flink® application deployment in a specified project and service. | Parameter | Information | | ---------------- | ------------------------------------------------------------ | | `project` | The name of the project | | `service_name` | The name of the service | | `application-id` | The ID of the Aiven for Flink application | | `deployment-id` | The ID of the Aiven for Flink application deployment to stop | **Example:** Stops the Aiven for Flink application deployment with application-id `f171af72-fdf0-442c-947c-7f6a0efa83ad` and deployment-id `6d5e2c03-2235-44a5-ab8f-c544a4de04ef`. ``` avn service flink stop-application-deployment flink-democli \ --project my-project \ --application-id f171af72-fdf0-442c-947c-7f6a0efa83ad \ --deployment-id 6d5e2c03-2235-44a5-ab8f-c544a4de04ef ``` ### `avn service flink cancel-application-deployments`[​](#avn-service-flink-cancel-application-deployments "Direct link to avn-service-flink-cancel-application-deployments") Cancels an Aiven for Flink® application deployment in a specified project and service. | Parameter | Information | | ---------------- | -------------------------------------------------------------- | | `project` | The name of the project | | `service_name` | The name of the service | | `application-id` | The ID of the Aiven for Flink application | | `deployment-id` | The ID of the Aiven for Flink application deployment to cancel | **Example:** Cancels the Aiven for Flink application deployment with application-id `f171af72-fdf0-442c-947c-7f6a0efa83ad` and deployment-id `6d5e2c03-2235-44a5-ab8f-c544a4de04ef`. ``` avn service flink cancel-application-deployments flink-democli \ --project my-project \ --application-id f171af72-fdf0-442c-947c-7f6a0efa83ad \ --deployment-id 6d5e2c03-2235-44a5-ab8f-c544a4de04ef ``` --- # avn service integration A full list of commands for `avn service integration`. ## Manage Aiven internal and external integrations[​](#manage-aiven-internal-and-external-integrations "Direct link to Manage Aiven internal and external integrations") ### `avn service integration-create`[​](#avn_service_integration_create "Direct link to avn_service_integration_create") Creates a new service integration. | Parameter | Information | | ---------------------- | -------------------------------------------------------------------------------------------- | | `--integration-type` | The [integration type](/docs/tools/cli/service/integration.md#avn_service_integration_types) | | `--source-service` | The integration source service | | `--dest-service` | The integration destination service | | `--source-endpoint-id` | The integration source endpoint ID | | `--dest-endpoint-id` | The integration destination endpoint ID | | `--user-config-json` | The integration parameters as JSON string or path to file preceded by `@` | | `-c KEY=VALUE` | The custom configuration settings. | tip Endpoint IDs are used when creating an integration with external services. To get an integration endpoint ID use the [dedicated endpoint list](/docs/tools/cli/service/integration.md#avn_service_integration_endpoint_list) command. note Both the `--user-config-json` and `-c` flags provide a way to customise the service integration using different methods. Only one of the flags are allowed per command. When using both in the same command, an error is shown: ``` ERROR command failed: UserError: -c (user config) and --user-config-json parameters cannot be used at the same time ``` **Example:** Create a `kafka_logs` service integration to send the logs of the service named `demo-pg` to an Aiven for Kafka service named `demo-kafka` in the topic `test_log`. ``` avn service integration-create \ --integration-type kafka_logs \ --source-service demo-pg \ --dest-service demo-kafka \ -c 'kafka_topic=test_log' ``` ### `avn service integration-delete`[​](#avn-service-integration-delete "Direct link to avn-service-integration-delete") Deletes a service integration. | Parameter | Information | | ---------------- | ----------------------------------- | | `integration-id` | The ID of the integration to delete | **Example:** Delete the integration with id `8e752fa9-a0c1-4332-892b-f1757390d53f`. ``` avn service integration-delete 8e752fa9-a0c1-4332-892b-f1757390d53f ``` ### `avn service integration-endpoint-create`[​](#avn_service_integration_endpoint_create "Direct link to avn_service_integration_endpoint_create") Creates an external service integration endpoint. | Parameter | Information | | -------------------- | ---------------------------------------------------------------------------------------------------------- | | `--endpoint-name` | The name of the endpoint | | `--endpoint-type` | The [endpoint type](/docs/tools/cli/service/integration.md#avn%20service%20integration%20endpoint%20types) | | `--user-config-json` | The endpoint configuration in JSON format or as path to a file preceded by `@` | | `-c KEY=VALUE` | The custom configuration settings. | **Example:** Create an external Apache Kafka® endpoint named `demo-ext-kafka`. ``` avn service integration-endpoint-create --endpoint-name demo-ext-kafka \ --endpoint-type external_kafka \ --user-config-json '{"bootstrap_servers":"servertest:123","security_protocol":"PLAINTEXT"}' ``` note For more examples of creating external Apache Kafka® endpoints, see [Integrate Aiven for Apache Flink® with Apache Kafka®](/docs/products/flink/howto/ext-kafka-flink-integration.md#step-4-create-an-external-apache-kafka-endpoint). **Example:** Create an external Loggly endpoint named `Loggly-ext`. ``` avn service integration-endpoint-create \ --endpoint-name Loggly-ext \ -d loggly -t rsyslog \ -c server=logs-01.loggly.com \ -c port=6514 \ -c format=rfc5424 \ -c tls=true \ -c sd='TOKEN@NNNNN TAG="tag-of-your-choice"' \ -c ca='loggly-tls-cert' ``` ### `avn service integration-endpoint-delete`[​](#avn-service-integration-endpoint-delete "Direct link to avn-service-integration-endpoint-delete") Deletes a service integration endpoint. | Parameter | Information | | ------------- | -------------------------------- | | `endpoint-id` | The ID of the endpoint to delete | **Example:** Delete the endpoint with ID `97590813-4a58-4c0c-91fd-eef0f074873b`. ``` avn service integration-endpoint-delete 97590813-4a58-4c0c-91fd-eef0f074873b ``` ### `avn service integration-endpoint-list`[​](#avn_service_integration_endpoint_list "Direct link to avn_service_integration_endpoint_list") Lists all service integration endpoints available in a selected project. **Example:** Lists all service integration endpoints available in the selected project. ``` avn service integration-endpoint-list ``` An example of `avn service integration-endpoint-list` output: ``` ENDPOINT_ID ENDPOINT_NAME ENDPOINT_TYPE ==================================== ================ ============== 97590813-4a58-4c0c-91fd-eef0f074873b datadog instance datadog 821e0144-1503-42db-aa9f-b4aa34c4af6b demo-ext-kafka external_kafka ``` ### `avn service integration-endpoint-types-list`[​](<#avn service integration endpoint types> "Direct link to avn service integration endpoint types") Lists all available integration endpoint types for given project. **Example:** Lists all service integration endpoint types available in the selected project. ``` avn service integration-endpoint-types-list ``` An example of `avn service integration-endpoint-types-list` output: ``` TITLE ENDPOINT_TYPE SERVICE_TYPES =========================================== =============================== ===================================================================================================================================================================================================================== Send service metrics to Datadog datadog elasticsearch, kafka, kafka_connect, kafka_mirrormaker, mysql, pg, redis Send service logs to AWS CloudWatch external_aws_cloudwatch_logs alerta, alertmanager, clickhouse, elasticsearch, flink, grafana, kafka, kafka_connect, kafka_mirrormaker, mysql, opensearch, pg, redis, sw Send service metrics to AWS CloudWatch external_aws_cloudwatch_metrics elasticsearch, kafka, kafka_connect, kafka_mirrormaker, mysql, pg, redis Send service logs to external Elasticsearch external_elasticsearch_logs alerta, alertmanager, clickhouse, elasticsearch, flink, grafana, kafka, kafka_connect, kafka_mirrormaker, mysql, opensearch, pg, redis, sw Send service logs to Google Cloud Logging external_google_cloud_logging alerta, alertmanager, clickhouse, elasticsearch, flink, grafana, kafka, kafka_connect, kafka_mirrormaker, mysql, opensearch, pg, redis, sw Integrate external Kafka cluster external_kafka alerta, alertmanager, clickhouse, elasticsearch, flink, grafana, kafka, kafka_connect, kafka_mirrormaker, kafka_mirrormaker, mysql, opensearch, pg, redis, sw Integrate external Schema Registry external_schema_registry kafka Access JMX metrics via Jolokia jolokia kafka, kafka_connect, kafka_mirrormaker Send service metrics to Prometheus prometheus elasticsearch, kafka, kafka_connect, kafka_mirrormaker, mysql, pg, redis Send service logs to remote syslog rsyslog alerta, alertmanager, clickhouse, elasticsearch, flink, grafana, kafka, kafka_connect, kafka_mirrormaker, mysql, opensearch, pg, redis, sw Send service metrics to SignalFX signalfx kafka ``` ### `avn service integration-endpoint-update`[​](#avn-service-integration-endpoint-update "Direct link to avn-service-integration-endpoint-update") Updates a service integration endpoint. | Parameter | Information | | -------------------- | ------------------------------------------------------------------------------ | | `endpoint-id` | The ID of the endpoint | | `--user-config-json` | The endpoint configuration in JSON format or as path to a file preceded by `@` | | `-c KEY=VALUE` | The custom configuration settings. | **Example:** Update an external Apache Kafka® endpoint with id `821e0144-1503-42db-aa9f-b4aa34c4af6b`. ``` avn service integration-endpoint-update 821e0144-1503-42db-aa9f-b4aa34c4af6b \ --user-config-json '{"bootstrap_servers":"servertestABC:123","security_protocol":"PLAINTEXT"}' ``` ### `avn service integration-list`[​](#avn_service_integration_list "Direct link to avn_service_integration_list") Lists the integrations defined for a selected service. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** List all integrations for the service named `demo-pg`. ``` avn service integration-list demo-pg ``` An example of `account service integration-list` output: ``` SERVICE_INTEGRATION_ID SOURCE DEST INTEGRATION_TYPE ENABLED ACTIVE DESCRIPTION ==================================== ============ ========== ================ ======= ====== ============================================================ 0e431dab-175a-4029-b417-d74a6437af1a demo-grafana demo-pg dashboard true true Provide a datasource for Grafana service (integration not enabled) demo-grafana demo-pg datasource false false Provide a datasource for Grafana service (without dashboard) (integration not enabled) demo-kafka demo-pg metrics false false Receive service metrics from service 8e752fa9-a0c1-4332-892b-f1757390d53f demo-pg demo-kafka kafka_logs true true Send logs to Kafka (integration not enabled) demo-pg demo-pg metrics false false Send service metrics to Aiven for Metrics or PostgreSQL service ``` ### `avn service integration-types-list`[​](#avn_service_integration_types "Direct link to avn_service_integration_types") Lists all available integration types for given project. **Example:** List all integration types for the currently selected project. ``` avn service integration-types-list ``` An example of `account service integration-types-list` output: ``` INTEGRATION_TYPE DEST_DESCRIPTION DEST_SERVICE_TYPE SOURCE_DESCRIPTION SOURCE_SERVICE_TYPES =============================== ==================================================================== =============================== ========================================================== ================================================================================================================================================================================================== datadog Receive service metrics from service datadog Send service metrics to Datadog endpoint elasticsearch, kafka, kafka_connect, kafka_mirrormaker, mysql, pg, redis datasource Provide a datasource for Grafana service (without dashboard) elasticsearch Grafana datasource grafana datasource Provide a datasource for Kafka Connect service alerta Kafka Connect datasource kafka, kafka_connect datasource Provide a datasource for PostgreSQL service pg PostgreSQL datasource pg datasource Provide a datasource for Elasticsearch service elasticsearch Elasticsearch datasource elasticsearch ... schema_registry_proxy Proxy Schema Registry requests kafka external_schema_registry signalfx Receive service metrics from service signalfx Send service metrics to SignalFX kafka ``` ### `avn service integration-update`[​](#avn_service_integration_update "Direct link to avn_service_integration_update") Updates an existing service integration. | Parameter | Information | | -------------------- | --------------------------------------------------------------------------- | | `integration_id` | The ID of integration | | `--user-config-json` | The integration parameters as JSON string or path to file (preceded by `@`) | | `-c KEY=VALUE` | The custom configuration settings. | **Example:** Update the service integration with ID `8e752fa9-a0c1-4332-892b-f1757390d53f` changing the Aiven for Kafka topic storing the logs to `test_pg_log`. ``` avn service integration-update 8e752fa9-a0c1-4332-892b-f1757390d53f \ -c 'kafka_topic=test_pg_log' ``` *Elasticsearch is a trademark of Elasticsearch B.V., registered in the U.S. and in other countries.* --- # avn service kafka-acl Full list of commands for `avn service kafka-acl`. ## Manage Kafka-native ACLs[​](#manage-kafka-native-acls "Direct link to Manage Kafka-native ACLs") The `avn service kafka-acl` command manages **Kafka-native access control lists (ACLs)** in Aiven for Apache Kafka®. Kafka-native ACLs define advanced, resource-level permissions for accessing resources such as topics, consumer groups, clusters, and transactional IDs. They support fine-grained access control with both `ALLOW` and `DENY` rules, and wildcard patterns (`*` and `?`) for resources and usernames. ### `avn service kafka-acl-add`[​](#avn-service-kafka-acl-add "Direct link to avn-service-kafka-acl-add") Add a Kafka-native ACL entry. | Parameter | Information | | ------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `service_name` | Name of the service | | `--principal` | Principal for the ACL, in the form `User:` | | `--topic` | Topic resource for the ACL | | `--group` | Consumer group resource for the ACL | | `--cluster` | Cluster resource for the ACL | | `--transactional-id` | `TransactionalId` resource for the ACL | | `--operation` | Operation type: possible values are `Describe`, `DescribeConfigs`,
`Alter`, `IdempotentWrite`, `Read`, `Delete`, `Create`, `ClusterAction`,
`All`, `Write`, `AlterConfigs`, `CreateTokens`, `DescribeTokens` | | `--host` | Host for the ACL, where `*` matches all hosts (default: `*`) | | `--resource-pattern-type` | Resource pattern type, either `LITERAL` or `PREFIXED` (default: `LITERAL`) | | ! `--deny` | Create a `DENY` rule (default: `ALLOW`) | **Example:** Add a Kafka-native ACL for user `userA` to `Read` on topics with names starting with `topic2020` in service `kafka-doc`. ``` avn service kafka-acl-add kafka-doc \ --principal User:userA \ --operation Read \ --topic topic2020 \ --resource-pattern-type PREFIXED ``` ### `avn service kafka-acl-delete`[​](#avn-service-kafka-acl-delete "Direct link to avn-service-kafka-acl-delete") Delete a Kafka-native ACL entry. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | Name of the service | | `acl_id` | ID of the ACL to delete | **Example:** Delete a Kafka-native ACL with ID `acl3604f96c74a` on service `kafka-doc`. ``` avn service kafka-acl-delete kafka-doc acl3604f96c74a ``` ### `avn service kafka-acl-list`[​](#avn-service-kafka-acl-list "Direct link to avn-service-kafka-acl-list") List Kafka-native ACL entries. | Parameter | Information | | -------------- | ------------------- | | `service_name` | Name of the service | **Example:** List Kafka-native ACLs defined for service `kafka-doc`. ``` avn service kafka-acl-list kafka-doc ``` Example output of `avn service kafka-acl-list`: ``` ID PERMISSION_TYPE PRINCIPAL OPERATION RESOURCE_TYPE PATTERN_TYPE RESOURCE_NAME HOST ============== =============== ========== ========= ============= ============ ============= ==== acl4f9ed69c8aa ALLOW User:John Write Topic LITERAL orders * acl4f9ed6e6371 ALLOW User:Frida Write Topic PREFIXED invoices * ``` ## Related page[​](#related-page "Direct link to Related page") For managing Aiven ACLs, see [`avn service acl`](/docs/tools/cli/service/acl.md). --- # avn service privatelink Full list of commands for `avn service privatelink`. ## Manage Aiven privatelink service for AWS and Azure[​](#manage-aiven-privatelink-service-for-aws-and-azure "Direct link to Manage Aiven privatelink service for AWS and Azure") ### `avn service privatelink availability`[​](#avn_service_privatelink_availability "Direct link to avn_service_privatelink_availability") Lists PrivateLink cloud availability and prices. | Parameter | Information | | ----------- | -------------------------------- | | `--project` | The project to fetch details for | | `--format` | Format of the output string | **Example:** Lists PrivateLink cloud availability and prices. ``` avn service privatelink availability ``` ``` CLOUD_NAME PRICE_USD =============================== ========= aws-ca-central-1 0.0600 aws-eu-central-1 0.0600 aws-us-east-1 0.0600 azure-canadacentral 0.0600 azure-eastus 0.0600 azure-france-central 0.0600 azure-germany-north 0.0600 azure-india-central 0.0600 azure-westus 0.0600 ``` ### `avn service privatelink aws connection list`[​](#avn_service_privatelink_aws_connection_list "Direct link to avn_service_privatelink_aws_connection_list") Lists AWS PrivateLink connection information for a service. | Parameter | Information | | -------------- | -------------------------------- | | `service_name` | The name of the service | | `--project` | The project to fetch details for | **Example:** List AWS PrivateLink connection information for the `kafka-12a3b4c5` service. ``` avn service privatelink aws connection list kafka-12a3b4c5 ``` An example of output: ``` { "dns_name": "vpce-0123456789abc1345-qfhrjbis.vpce-svc-0abcdef0123456789.us-east-1.vpce.amazonaws.com", "privatelink_connection_id": "plc39413abcdef", "state": "active", "vpc_endpoint_id": "vpce-0123456789abc1345" } ``` ### `avn service privatelink aws create`[​](#avn_service_privatelink_aws_create "Direct link to avn_service_privatelink_aws_create") Creates an AWS PrivateLink for a service. To add multiple principals, repeat `\--principal` parameter. | Parameter | Information | | -------------- | ------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--project` | The project to fetch details for | | `--principal` | ARN that is allowed to connect (example: `arn:aws:iam::123456789012:user/cloud_user`) | | `--format` | Format of the output string | **Example:** Create an AWS PrivateLink for the `kafka-12a3b4c5` service. ``` avn service privatelink aws create --principal 'arn:aws:iam::123456789012:user/cloud_user' --principal 'arn:aws:iam::987654321098:user/cloud_user' kafka-12a3b4c5 ``` An example of output: ``` AWS_SERVICE_ID AWS_SERVICE_NAME PRINCIPALS STATE ============== ================ ==================================================================================== ======== null null arn:aws:iam::123456789012:user/cloud_user, arn:aws:iam::987654321098:user/cloud_user creating ``` ### `avn service privatelink aws delete`[​](#avn_service_privatelink_aws_delete "Direct link to avn_service_privatelink_aws_delete") Deletes an AWS PrivateLink defined for a service. | Parameter | Information | | -------------- | -------------------------------- | | `service_name` | The name of the service | | `--project` | The project to fetch details for | | `--format` | Format of the output string | **Example:** Delete the AWS PrivateLink for the `kafka-12a3b4c5` service. ``` avn service privatelink aws delete kafka-12a3b4c5 ``` An example of output: ``` AWS_SERVICE_ID AWS_SERVICE_NAME PRINCIPALS STATE ========================== ======================================================= ========================================= ======== vpce-svc-1234567890abc1234 com.amazonaws.vpce.us-east-1.vpce-svc-1234567890abc1234 arn:aws:iam::123456789012:user/cloud_user deleting ``` tip The deletion can take some time to complete. You can check the status by running `avn service privatelink aws get`. ### `avn service privatelink aws get`[​](#avn_service_privatelink_aws_get "Direct link to avn_service_privatelink_aws_get") Lists AWS PrivateLink information for a service. | Parameter | Information | | -------------- | -------------------------------- | | `service_name` | The name of the service | | `--project` | The project to fetch details for | | `--format` | Format of the output string | **Example:** List AWS PrivateLink information for the `kafka-12a3b4c5` service. ``` avn service privatelink aws get kafka-12a3b4c5 ``` An example of output: ``` AWS_SERVICE_ID AWS_SERVICE_NAME PRINCIPALS STATE ========================== ======================================================= ========================================= ====== vpce-svc-1234567890abc1234 com.amazonaws.vpce.us-east-1.vpce-svc-1234567890abc1234 arn:aws:iam::123456789012:user/cloud_user active ``` ### `avn service privatelink aws update`[​](#avn_service_privatelink_aws_update "Direct link to avn_service_privatelink_aws_update") Updates AWS PrivateLink principals for a service. To update multiple principals, repeat `\--principal` parameter. | Parameter | Information | | -------------- | ------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--project` | The project to fetch details for | | `--principal` | ARN that is allowed to connect (example: `arn:aws:iam::123456789012:user/cloud_user`) | | `--format` | Format of the output string | **Example:** Update AWS principals for the `kafka-12a3b4c5` service. ``` avn service privatelink aws update \ --principal 'arn:aws:iam::123456789012:user/cloud_user' \ kafka-12a3b4c5 ``` An example of output: ``` AWS_SERVICE_ID AWS_SERVICE_NAME PRINCIPALS STATE ========================== ======================================================= ========================================= ====== vpce-svc-1234567890abc1234 com.amazonaws.vpce.us-east-1.vpce-svc-1234567890abc1234 arn:aws:iam::123456789012:user/cloud_user active ``` ### `avn service privatelink aws refresh`[​](#avn_service_privatelink_aws_refresh "Direct link to avn_service_privatelink_aws_refresh") Refreshes incoming AWS PrivateLink endpoint connections. | Parameter | Information | | -------------- | -------------------------------- | | `service_name` | The name of the service | | `--project` | The project to fetch details for | | `--format` | Format of the output string | **Example:** Refresh incoming AWS PrivateLink endpoint connections for the `kafka-12a3b4c5` service. ``` avn service privatelink aws refresh kafka-12a3b4c5 ``` ### `avn service privatelink azure connection approve`[​](#avn_service_privatelink_azure_connection_approve "Direct link to avn_service_privatelink_azure_connection_approve") Approves a pending Azure Private Link connection endpoint. | Parameter | Information | | --------------------------- | ----------------------------------- | | `service_name` | The name of the service | | `privatelink_connection_id` | The Aiven privatelink connection ID | | `--project` | The project to fetch details for | | `--format` | Format of the output string | **Example:** Approve the Azure Private Link `plc12345abcdef` connection for the `kafka-12a3b4c5` service. ``` avn service privatelink azure connection approve kafka-12a3b4c5 plc12345abcdef ``` An example of output: ``` PRIVATE_ENDPOINT_ID PRIVATELINK_CONNECTION_ID STATE USER_IP_ADDRESS ======================================================================================================================================== ========================= ============= =============== /subscriptions/12345678-90ab-cdef-0987-6543210abcde/resourceGroups/group-eastus/providers/Microsoft.Network/privateEndpoints/pl-endpoint plc12345abcdef user-approved null ``` ### `avn service privatelink azure connection list`[​](#avn_service_privatelink_azure_connection_list "Direct link to avn_service_privatelink_azure_connection_list") Lists Azure Private Link connection information for a service. | Parameter | Information | | -------------- | -------------------------------- | | `service_name` | The name of the service | | `--project` | The project to fetch details for | | `--format` | Format of the output string | **Example:** List Azure Private Link connection information for the `kafka-12a3b4c5` service. ``` avn service privatelink azure connection list kafka-12a3b4c5 ``` An example of output: ``` PRIVATELINK_CONNECTION_ID PRIVATE_ENDPOINT_ID STATE USER_IP_ADDRESS ========================= ======================================================================================================================================== ===================== =============== plc12345abcdef /subscriptions/12345678-90ab-cdef-0987-6543210abcde/resourceGroups/group-eastus/providers/Microsoft.Network/privateEndpoints/pl-endpoint pending-user-approval null ``` ### `avn service privatelink azure connection update`[​](#avn_service_privatelink_azure_connection_update "Direct link to avn_service_privatelink_azure_connection_update") Updates an Azure Private Link connection with the Private IP address of the private endpoint's Network interface. | Parameter | Information | | --------------------------- | ----------------------------------------------------------- | | `service_name` | The name of the service | | `privatelink_connection_id` | The Aiven PrivateLink connection ID | | `--endpoint-ip-address` | (Private) IP address of Azure endpoint in user subscription | | `--project` | The project to fetch details for | | `--format` | Format of the output string | **Example:** In the `kafka-12a3b4c5` service, update the IP of the Azure Private Link connection `plc12345abcdef` to `10.19.1.4`. ``` avn service privatelink azure connection update \ --endpoint-ip-address 10.19.1.4 \ kafka-12a3b4c5 \ plc12345abcdef ``` An example of output: ``` PRIVATE_ENDPOINT_ID PRIVATELINK_CONNECTION_ID STATE USER_IP_ADDRESS ======================================================================================================================================== ========================= ====== =============== /subscriptions/12345678-90ab-cdef-0987-6543210abcde/resourceGroups/group-eastus/providers/Microsoft.Network/privateEndpoints/pl-endpoint plc12345abcdef active 10.19.1.4 ``` ### `avn service privatelink azure create`[​](#avn_service_privatelink_azure_create "Direct link to avn_service_privatelink_azure_create") Creates an Azure Private Link for a service. | Parameter | Information | | ------------------------ | ----------------------------------------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--project` | The project to fetch details for | | `--user-subscription-id` | Azure subscription IDs allowed to connect to the Private Link service (example: `12345678-90ab-cdef-0987-6543210abcde`) | | `--format` | Format of the output string | **Example:** Create an Azure Private Link for the `kafka-12a3b4c5` service. ``` avn service privatelink azure create \ --user-subscription-id \ 12345678-90ab-cdef-0987-6543210abcde \ kafka-12a3b4c5 ``` An example of output: ``` AZURE_SERVICE_ALIAS AZURE_SERVICE_ID STATE USER_SUBSCRIPTION_IDS =================== ================ ======== ==================================== null null creating 12345678-90ab-cdef-0987-6543210abcde ``` ### `avn service privatelink azure delete`[​](#avn_service_privatelink_azure_delete "Direct link to avn_service_privatelink_azure_delete") Deletes an Azure Private Link defined for a service. | Parameter | Information | | -------------- | -------------------------------- | | `service_name` | The name of the service | | `--project` | The project to fetch details for | | `--format` | Format of the output string | **Example:** Delete Azure Private Link for the `kafka-12a3b4c5` service. ``` avn service privatelink azure delete kafka-12a3b4c5 ``` An example of output: ``` AZURE_SERVICE_ALIAS AZURE_SERVICE_ID STATE USER_SUBSCRIPTION_IDS ============================================================================================ ========================================================================================================================================================================================= ======== ==================================== aivenprod-ss123456789ab.12345678-90ab-cdef-9876-543210abcdef.eastus.azure.privatelinkservice /subscriptions/12345678-90ab-cdef-1234-567890abcdef/resourceGroups/aivenprod-12345678-90ab-cdef-1234-567890abcdef/providers/Microsoft.Network/privateLinkServices/aivenprod-ss123456789ab deleting 12345678-90ab-cdef-0987-6543210abcde ``` ### `avn service privatelink azure get`[​](#avn_service_privatelink_azure_get "Direct link to avn_service_privatelink_azure_get") Lists Azure Private Link information for a service. | Parameter | Information | | -------------- | -------------------------------- | | `service_name` | The name of the service | | `--project` | The project to fetch details for | | `--format` | Format of the output string | **Example:** List Azure Private Link information for the `kafka-12a3b4c5` service. ``` avn service privatelink azure get kafka-12a3b4c5 ``` An example of output: ``` AZURE_SERVICE_ALIAS AZURE_SERVICE_ID STATE USER_SUBSCRIPTION_IDS ============================================================================================ ========================================================================================================================================================================================= ====== ==================================== aivenprod-ss123456789ab.12345678-90ab-cdef-9876-543210abcdef.eastus.azure.privatelinkservice /subscriptions/12345678-90ab-cdef-1234-567890abcdef/resourceGroups/aivenprod-12345678-90ab-cdef-1234-567890abcdef/providers/Microsoft.Network/privateLinkServices/aivenprod-ss123456789ab active 12345678-90ab-cdef-0987-6543210abcde ``` ### `avn service privatelink azure refresh`[​](#avn_service_privatelink_azure_refresh "Direct link to avn_service_privatelink_azure_refresh") Refreshes incoming Azure Private Link endpoint connections. | Parameter | Information | | -------------- | -------------------------------- | | `service_name` | The name of the service | | `--project` | The project to fetch details for | | `--format` | Format of the output string | **Example:** Refresh incoming Azure Private Link endpoint connections for the `kafka-12a3b4c5` service. ``` avn service privatelink azure refresh kafka-12a3b4c5 ``` --- # avn service quota Full list of commands for `avn service quota`. ## Manage Kafka service quotas[​](#manage-kafka-service-quotas "Direct link to Manage Kafka service quotas") The `avn service quota` command manages quotas for Aiven for Apache Kafka® services. Quotas limit network throughput and CPU usage for producers and consumers. This prevents individual clients from overloading the cluster. You can scope quotas to a specific user, a client ID, or both. For an overview of how quotas work, see [Quotas in Aiven for Apache Kafka®](/docs/products/kafka/concepts/kafka-quotas.md). ### `avn service quota create`[​](#avn-service-quota-create "Direct link to avn-service-quota-create") Create a quota for an Aiven for Apache Kafka service. At least one of `--client-id` or `--user` is required to identify the quota subject. At least one quota parameter (`--consumer-byte-rate`, `--producer-byte-rate`, or `--request-percentage`) is also required. | Parameter | Information | | ---------------------- | ------------------------------------------------------------------------------------------------------- | | `service_name` | Name of the service | | `--client-id` | Client ID to scope the quota to | | `--user` | Username to scope the quota to | | `--consumer-byte-rate` | Maximum bytes per second that consumer clients with this quota can read from the cluster (0—1073741824) | | `--producer-byte-rate` | Maximum bytes per second that producer clients with this quota can write to the cluster (0—1073741824) | | `--request-percentage` | Maximum percentage of CPU time for request handler I/O and network threads per broker (0—100) | note To apply a quota to all users or all client IDs, use the keyword `default` as the value for `--user` or `--client-id`. **Example:** Set a 1 MiB/s producer and consumer throttle for user `alice` on service `kafka-doc`. ``` avn service quota create kafka-doc \ --user alice \ --consumer-byte-rate 1048576 \ --producer-byte-rate 1048576 ``` **Example:** Set a 25% CPU throttle for client ID `analytics-consumer` on service `kafka-doc`. ``` avn service quota create kafka-doc \ --client-id analytics-consumer \ --request-percentage 25 ``` **Example:** Set a default quota for all users on service `kafka-doc`. ``` avn service quota create kafka-doc \ --user default \ --consumer-byte-rate 5242880 ``` ### `avn service quota list`[​](#avn-service-quota-list "Direct link to avn-service-quota-list") List all quotas defined for an Aiven for Apache Kafka service. | Parameter | Information | | -------------- | ------------------- | | `service_name` | Name of the service | **Example:** List all quotas for service `kafka-doc`. ``` avn service quota list kafka-doc ``` Example output: ``` CLIENT-ID USER CONSUMER_BYTE_RATE PRODUCER_BYTE_RATE REQUEST_PERCENTAGE ==================== ===== ================== ================== ================== analytics-consumer 1048576 1048576 25 alice 524288 524288 ``` ### `avn service quota describe`[​](#avn-service-quota-describe "Direct link to avn-service-quota-describe") Describe a specific quota on an Aiven for Apache Kafka service. At least one of `--client-id` or `--user` is required. | Parameter | Information | | -------------- | ---------------------------------- | | `service_name` | Name of the service | | `--client-id` | Client ID of the quota to describe | | `--user` | Username of the quota to describe | **Example:** Describe the quota for user `alice` on service `kafka-doc`. ``` avn service quota describe kafka-doc --user alice ``` **Example:** Describe the quota scoped to both a user and a client ID. ``` avn service quota describe kafka-doc \ --user alice \ --client-id analytics-consumer ``` ### `avn service quota delete`[​](#avn-service-quota-delete "Direct link to avn-service-quota-delete") Delete a quota from an Aiven for Apache Kafka service. At least one of `--client-id` or `--user` is required. | Parameter | Information | | -------------- | -------------------------------- | | `service_name` | Name of the service | | `--client-id` | Client ID of the quota to delete | | `--user` | Username of the quota to delete | **Example:** Delete the quota for user `alice` on service `kafka-doc`. ``` avn service quota delete kafka-doc --user alice ``` ## Related pages[​](#related-pages "Direct link to Related pages") * [Quotas in Aiven for Apache Kafka®](/docs/products/kafka/concepts/kafka-quotas.md) * [Manage quotas](/docs/products/kafka/howto/manage-quotas.md) --- # avn service schema-registry-acl Full list of commands for `avn service schema-registry-acl`. ## Manage Karapace schema registry access control lists for Apache Kafka®[​](#manage-karapace-schema-registry-access-control-lists-for-apache-kafka "Direct link to Manage Karapace schema registry access control lists for Apache Kafka®") Using the following commands you can manage [Karapace schema registry authorization](/docs/products/kafka/karapace/concepts/schema-registry-authorization.md) for your Aiven for Apache Kafka® service via the `avn` commands. ### `avn service schema-registry-acl-add`[​](#avn-service-schema-registry-acl-add "Direct link to avn-service-schema-registry-acl-add") You can add a Karapace schema registry ACL entry by using the command: ``` avn service schema-registry-acl-add ``` Where: | Parameter | Information | | -------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--permission` | The permission type:- `schema_registry_read`
- `schema_registry_write` | | `--resource` | The resource format can be `Config:` or `Subject:`. For more information, see [Schema registry ACL definitions](/docs/products/kafka/karapace/concepts/acl-definition.md). | | `--username` | The name of a service user | **Example** The following example shows you how to add an ACL entry to grant a user (`user_1`) read options (`schema_registry_read`) to subject `s1`. Replace the placeholders `PROJECT_NAME` and `APACHE_KAFKA_SERVICE_NAME` with the name of the project and the Aiven for Apache Kafka® service. ``` avn service schema-registry-acl-add kafka-doc \ --username 'user_1' \ --permission schema_registry_read \ --resource 'Subject:s1' ``` note You cannot edit a Karapace schema registry ACL entry. Create a new entry and delete the older entry. ### `avn service schema-registry-acl-delete`[​](#avn-service-schema-registry-acl-delete "Direct link to avn-service-schema-registry-acl-delete") You can delete a Karapace schema registry ACL entry using the command: ``` avn service schema-registry-acl-delete ``` Where: | Parameter | Information | | -------------- | ---------------------------------------------------- | | `service_name` | The name of the service | | `acl_id` | The ID of the Karapace schema registry ACL to delete | **Example:** The following example deletes the Karapace schema registry ACL with ID `acl3604f96c74a` on the Aiven for Apache Kafka® instance named `kafka-doc`. ``` avn service schema-registry-acl-delete kafka-doc acl3604f96c74a ``` ### `avn service schema-registry-acl-list`[​](#avn-service-schema-registry-acl-list "Direct link to avn-service-schema-registry-acl-list") List all Karapace schema registry ACL entries defined: ``` avn service schema-registry-acl-list ``` Where: | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** The following example lists the ACLs defined for an Aiven for Apache Kafka® service named `kafka-doc`. ``` avn service schema-registry-acl-list kafka-doc ``` The command output is: ``` ID USERNAME RESOURCE PERMISSION ======================== ======== =============== ===================== default-sr-admin-config avnadmin Config: schema_registry_write default-sr-admin-subject avnadmin Subject:* schema_registry_write acl12345678901 userAB* Subject:s123* schema_registry_write ``` --- # avn service index Full list of commands for `avn service index`. ## Manage OpenSearch® indexes[​](#manage-opensearch-indexes "Direct link to Manage OpenSearch® indexes") ### `avn service index-delete`[​](#avn-service-index-delete "Direct link to avn-service-index-delete") Deletes OpenSearch service indexes. ### `avn service index-list`[​](#avn-service-index-list "Direct link to avn-service-index-list") Lists OpenSearch service indexes. --- # avn service tags Full list of commands for `avn service tags`. ## Manage service tags[​](#manage-service-tags "Direct link to Manage service tags") ### `avn service tags list`[​](#avn-service-tags-list "Direct link to avn-service-tags-list") Retrieves the tags associated with an Aiven service. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** Retrieve the tags associated with the service named `kafka-demo`. ``` avn service tags list kafka-demo ``` ``` KEY VALUE ======= =================== team frontend scope userclicks-tracking ``` **Example:** Retrieve the tags associated with the service named `kafka-demo` in JSON format. ``` avn service tags list kafka-demo --json ``` ``` [ { "key": "team", "value": "frontend" }, { "key": "scope", "value": "userclicks-tracking" } ] ``` ### `avn service tags replace`[​](#avn-service-tags-replace "Direct link to avn-service-tags-replace") Replaces a tag associated with an Aiven service, deleting the any old entry first. | Parameter | Information | | -------------- | ----------------------------------------------------- | | `service_name` | The name of the service | | `--tag` | The service tag to replace, in the format `KEY=VALUE` | **Example:** in the `demo-kafka` Aiven service, replace the tag with key `scope` to the value `userclicks` ``` avn service tags replace demo-kafka \ --tag scope=userclicks ``` ### `avn service tags update`[​](#avn-service-tags-update "Direct link to avn-service-tags-update") Update tags associated with an Aiven service. | Parameter | Information | | -------------- | ------------------------------------------------- | | `service_name` | The name of the service | | `--add-tag` | The service tag to add, in the format `KEY=VALUE` | | `--remove-tag` | The service tag key to remove | **Example:** in the `demo-kafka` Aiven service, modify the following: * add the tag with key `scope` and value `userclicks` * add the tag with key `bu` and value `emea` * remove the tag with key `team` ``` avn service tags update demo-kafka \ --add-tag scope=userclicks \ --add-tag bu=emea \ --remove-tag team=frontend ``` --- # avn service topic Full list of commands for `avn service topic`. ## Manage Aiven for Apache Kafka® topics[​](#avn_cli_service_topic_create "Direct link to Manage Aiven for Apache Kafka® topics") ### `avn service topic-create`[​](#avn-service-topic-create "Direct link to avn-service-topic-create") Creates a new Kafka topic on the specified Aiven for Apache Kafka service. | Parameter | Information | | ----------------------- | -------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `topic` | The name of the topic | | `--partitions` | The number of topic partitions | | `--replication` | The topic replication factor | | `--min-insync-replicas` | The minimum required nodes In Sync Replicas (ISR) for the topic/partition (default: 1) | | `--retention` | The retention period in hours (default: unlimited) | | `--retention-bytes` | The retention limit in bytes (default: unlimited) | | `--cleanup-policy` | The topic cleanup policy; can be either `delete` or `compact`. | | `--tag KEY[=VALUE]` | Topic tagging | **Example:** Create a topic named `invoices` in the `demo-kafka` service with: * `3` partitions * `2` as replication factor * 2 hours of retention time * `BU=FINANCE` tag ``` avn service topic-create demo-kafka invoices \ --partitions 3 \ --replication 2 \ --retention 2 \ --tag BU=FINANCE ``` ### `avn service topic-delete`[​](#avn-cli-delete-topic "Direct link to avn-cli-delete-topic") Deletes a Kafka topic on the specified Aiven for Apache Kafka service. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | | `topic` | The name of the topic | **Example:** Delete the topic named `invoices` in the `demo-kafka` service. ``` avn service topic-delete demo-kafka invoices ``` ### `avn service topic-get`[​](#avn-service-topic-get "Direct link to avn-service-topic-get") Retrieves Kafka topic on the specified Aiven for Apache Kafka service. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | | `topic` | The name of the topic | **Example:** Retrieve the information about a topic named `invoices` in the `demo-kafka` service. ``` avn service topic-get demo-kafka invoices ``` An example of `avn service topic-get` output: ``` PARTITION ISR SIZE EARLIEST_OFFSET LATEST_OFFSET GROUPS ========= === ==== =============== ============= ====== 0 2 0 0 0 0 1 2 null 0 0 0 2 2 0 0 0 0 (No consumer groups) ``` ### `avn service topic-list`[​](#avn-service-topic-list "Direct link to avn-service-topic-list") Lists Kafka topics on the specified Aiven for Apache Kafka service together with the following information: * partitions * replication * min in-sync replicas * retention bytes * retention hours * cleanup policy * tags | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** Retrieve list of topics available in the `demo-kafka` service. ``` avn service topic-list demo-kafka ``` An example of `avn service topic-get` output: ``` TOPIC_NAME PARTITIONS REPLICATION MIN_INSYNC_REPLICAS RETENTION_BYTES RETENTION_HOURS CLEANUP_POLICY TAGS ========== ========== =========== =================== =============== =============== ============== ========== bills 3 2 1 -1 unlimited delete invoices 3 2 1 -1 168 delete BU=FINANCE orders 2 3 1 -1 unlimited delete ``` ### `avn service topic-update`[​](#avn-cli-topic-update "Direct link to avn-cli-topic-update") Updates a Kafka topic on the specified Aiven for Apache Kafka service. | Parameter | Information | | ----------------------- | -------------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `topic` | The name of the topic | | `--partitions` | The number of topic partitions | | `--replication` | The topic replication factor | | `--min-insync-replicas` | The minimum required nodes In Sync Replicas (ISR) for the topic/partition (default: 1) | | `--retention` | The retention period in hours (default: unlimited) | | `--retention-bytes` | The retention limit in bytes (default: unlimited) | | `--cleanup-policy` | The topic cleanup policy; can be either `delete` or `compact`. | | `--tag KEY[=VALUE]` | Topic tagging | | `--untag KEY` | Topic tag to remove | **Example:** Update the topic named `invoices` in the `demo-kafka` service. Set `4` partitions and `3` as replication factor. Remove the `BU` tag and add a new `CC=FINANCE_DE` tag. ``` avn service topic-update demo-kafka invoices \ --partitions 4 \ --replication 3 \ --tag CC=FINANCE_DE \ --untag BU ``` --- # avn service user Full list of commands for `avn service user`. ## Manage Aiven users and credentials[​](#manage-aiven-users-and-credentials "Direct link to Manage Aiven users and credentials") ### `avn service user-create`[​](#avn-service-user-create "Direct link to avn-service-user-create") Creates a new user for the selected service. | Parameter | Information | | ------------------------ | ------------------------------------------------------------------------------- | | `service_name` | The name of the service | | `--username` | The new username to be created | | `--m3-group` | The name of the group the user belongs to (for Aiven for Metrics services only) | | `--redis-acl-keys` | The ACL rules for keys (Aiven for Caching services only) | | `--redis-acl-commands` | The ACL rules for commands (Aiven for Caching services only) | | `--redis-acl-categories` | The ACL rules for categories (Aiven for Caching services only) | | `--redis-acl-channels` | The ACL rules for channels (Aiven for Caching services only) | **Example:** Create new user named `janedoe` for a service named `pg-demo`. ``` avn service user-create pg-demo --username janedoe ``` ### `avn service user-creds-acknowledge`[​](#avn_service_user_creds_acknowledge "Direct link to avn_service_user_creds_acknowledge") Acknowledges the usage of the [renewed SSL certificate](/docs/products/kafka/howto/renew-ssl-certs.md) for a specific service user. | Parameter | Information | | -------------- | --------------------------------------------------- | | `service_name` | The name of the service | | `--username` | The username for which to download the certificates | **Example:** Acknowledge the usage of the new SSL certificate for the user `janedoe` belonging to a service named `kafka-demo`. ``` avn service user-creds-acknowledge kafka-demo --username janedoe ``` ### `avn service user-creds-download`[​](#avn_service_user_creds_download "Direct link to avn_service_user_creds_download") Downloads the SSL certificate, key and CA certificate for the selected service. | Parameter | Information | | -------------- | ------------------------------------------------------ | | `service_name` | The name of the service | | `--username` | The username for which to download the certificates | | `-d` | The target directory where certificates will be stored | **Example:** Download the SSL certificate, key and CA certificate in a folder named `/tmp/certs` for the user `janedoe` belonging to a service named `kafka-demo`. ``` avn service user-creds-download kafka-demo --username janedoe -d /tmp/certs ``` ### `avn service user-delete`[​](#avn-service-user-delete "Direct link to avn-service-user-delete") Delete a service in a given Aiven service. | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | | `--username` | The username to delete | **Example:** Delete the user `janedoe` defined in a service named `kafka-demo`. ``` avn service user-delete kafka-demo --username janedoe ``` ### `avn service user-get`[​](#avn-service-user-get "Direct link to avn-service-user-get") Retrieves the details for a single user in a given Aiven service. | Parameter | Information | | -------------- | ---------------------------------------------- | | `service_name` | The name of the service | | `--username` | The username for which to retrieve the details | **Example:** Retrieve the details for the user `janedoe` defined for a service named `kafka-demo`. ``` avn service user-get kafka-demo --username janedoe ``` tip Use the `--json` parameter to retrieve all the service specific information for a specific user. ### `avn service user-kafka-java-creds`[​](#avn_service_user_kafka_java_creds "Direct link to avn_service_user_kafka_java_creds") Download user certificate/key/CA certificate and create a Java keystore/truststore/properties from them Downloads the SSL certificate, key and CA certificate and creates a Java keystore and truststore for the selected service. | Parameter | Information | | -------------- | --------------------------------------------------------------- | | `service_name` | The name of the service | | `--username` | The username for which to download the certificates | | `-d` | The target directory where certificates will be stored | | `--password` | The Java keystore and truststore password (default: `changeit`) | **Example:** Download the SSL certificate, key and CA certificate in a folder named `/tmp/certs` for the user `janedoe` belonging to a service named `kafka-demo`. Secure the Java keystore and truststore with the password `safePassword123`. ``` avn service user-kafka-java-creds kafka-demo --username janedoe -d /tmp/certs --password safePassword123 ``` ### `avn service user-list`[​](#avn-service-user-list "Direct link to avn-service-user-list") Lists the users defined for the selected service, and the related type (`primary` or `normal`). | Parameter | Information | | -------------- | ----------------------- | | `service_name` | The name of the service | **Example:** List the users defined for a service named `pg-doc`. ``` avn service user-list pg-doc ``` An example of `account service user-list` output: ``` USERNAME TYPE ========= ======= analytics normal avnadmin primary ``` ### `avn service user-password-reset`[​](#avn-service-user-password-reset "Direct link to avn-service-user-password-reset") Resets or changes the service user password. | Parameter | Information | | ---------------- | --------------------------------------- | | `service_name` | The name of the service | | `--username` | The username to change the password for | | `--new-password` | The new password for the user | **Example:** Change the password for the `avnadmin` user of the service named `pg-doc` to `VerySecurePwd123`. ``` avn service user-password-reset pg-doc --username avnadmin --new-password VerySecurePwd123 ``` ### `avn service user-set-access-control`[​](#avn-service-user-set-access-control "Direct link to avn-service-user-set-access-control") Set Caching service user access control --- # avn user Manage users and personal tokens with the `avn user` commands. ## Manage users[​](#manage-users "Direct link to Manage users") ### `avn user info`[​](#avn-user-info "Direct link to avn-user-info") Retrieves the current user information such as: * Username * Real name * State (`active` or `inactive`) * Token validity start date * Associated projects * Authentication method **Example:** Retrieve the information for the currently logged user. ``` avn user info ``` An example of user information: ``` USER REAL_NAME STATE TOKEN_VALIDITY_BEGIN PROJECTS AUTH ==================== ========= ====== ================================ ============================= ======== john.doe@example.com John Doe active 2021-08-18T09:24:10.298796+00:00 dev-sandbox, prod-environment password ``` ### `avn user login`[​](#avn-user-login "Direct link to avn-user-login") Logs the user in. | Parameter | Information | | --------- | ----------------------------------------- | | `email` | The email associated to the user | | `--token` | Logs in the user with a pre-created token | **Example:** Log the `john.doe@example.com` user in, and prompt for password. ``` avn user login john.doe@example.com ``` The user will be prompted to insert the password. **Example:** Log the `john.doe@example.com` user in, using a personal token. ``` avn user login john.doe@example.com --token ``` The user will be prompted to insert the personal token. ### `avn user logout`[​](#avn-user-logout "Direct link to avn-user-logout") Logs the user out. **Example:** Log the user out. ``` avn user logout ``` ## Manage tokens[​](#manage-tokens "Direct link to Manage tokens") Commands for managing a user's tokens. ### `avn user access-token create`[​](#avn-user-access-token-create "Direct link to avn-user-access-token-create") Creates a token for the logged in user. | Parameter | Information | | -------------------- | -------------------------------------------------------------------------------------------- | | `--description` | Description of how the token will be used | | `--max-age-seconds` | Maximum age of the token in seconds, if any, after which it will expire (30 days by default) | | `--extend-when-used` | Extend token's expiry time when used (only applicable if token is set to expire) | **Example:** Create a token. ``` avn user access-token create --description "To be used with Python Notebooks" ``` **Example:** Create a token expiring every hour if not used. ``` avn user access-token create \ --description "To be used with Python Notebooks" \ --max-age-seconds 3600 \ --extend-when-used ``` The output will be similar to the following: ``` EXPIRY_TIME DESCRIPTION MAX_AGE_SECONDS EXTEND_WHEN_USED FULL_TOKEN ==================== ================================ =============== ================ =============================== 2021-08-16T16:26:10Z To be used with python notebooks 3600 true 6JsKDclT3OMQd1V2Fl2...RaraBPg== ``` ### `avn user access-token list`[​](#avn-user-access-token-list "Direct link to avn-user-access-token-list") Retrieves the information for all the active tokens: * Expiration time * Token prefix * Description * Token's max age in seconds * Extended when used flag * Last used time * Last IP address * Last user agent **Example:** Retrieve the information for the logged-in user. ``` avn user access-token list ``` An example of user information: ``` EXPIRY_TIME TOKEN_PREFIX DESCRIPTION MAX_AGE_SECONDS EXTEND_WHEN_USED LAST_USED_TIME LAST_IP LAST_USER_AGENT ==================== ============ ================================ =============== ================ ==================== =========== =================== 2021-09-15T15:29:14Z XCJ3+bgWywIh Test token 2592000 true 2021-08-16T15:29:14Z 192.168.1.1 aiven-client/2.12.0 2021-08-16T16:26:10Z 6JsKDclT3OMQ To be used with Python Notebooks 3600 true null null null ``` ### `avn user access-token revoke`[​](#avn-user-access-token-revoke "Direct link to avn-user-access-token-revoke") Revokes a token. | Parameter | Information | | -------------- | -------------------------------------------------------------- | | `token_prefix` | The full token or token prefix identifying the token to revoke | **Example:** Revoke the token starting with `6JsKDclT3OMQ`. ``` avn user access-token revoke "6JsKDclT3OMQ" ``` ### `avn user access-token update`[​](#avn-user-access-token-update "Direct link to avn-user-access-token-update") Updates the description of a token. | Parameter | Information | | --------------- | -------------------------------------------------------------- | | `token_prefix` | The full token or token prefix identifying the token to update | | `--description` | Description of how the token will be used | **Example:** Update the description of the token starting with `6JsKDclT3OMQ`. ``` avn user access-token update "6JsKDclT3OMQ" --description "To be used with Jupyter Notebooks" ``` ### `avn user tokens-expire`[​](<#avncli user-tokens-expire> "Direct link to avncli user-tokens-expire") Expires all the tokens associated with the user expired. **Example:** ``` avn user tokens-expire ``` ### `avn user access-token`[​](#avn-user-access-token "Direct link to avn-user-access-token") Set of commands for managing a user's tokens. ### `avn user create`[​](#avn-user-create "Direct link to avn-user-create") Creates a new user. | Parameter | Information | | ------------- | -------------------------------- | | `email` | The email associated to the user | | `--real-name` | The user's real name | **Example:** Create a user. ``` avn user create john.doe@example.com ``` **Example:** Create a user specifying the real name. ``` avn user create john.doe@example.com --real-name "John Doe" ``` --- # avn vpc The list of commands for project VPCs (`avn vpc`) and organization VPCs (`avn organization vpc`) ## Manage VPCs[​](#manage-vpcs "Direct link to Manage VPCs") ### Create VPCs[​](#create-vpcs "Direct link to Create VPCs") * Project VPC * Organization VPC #### Command: `avn vpc create`[​](#command-avn-vpc-create "Direct link to command-avn-vpc-create") | Parameter | Information | | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `--project` | The project where to create the VPC | | `--cloud` | The cloud and region where to host the VPC. See the list of available cloud regions using the [`avn cloud list`](/docs/tools/cli/cloud.md#avn-cloud-list) command. | | `--network-cidr` | The network range in the Aiven project VPC in CIDR format (a.b.c.d/e) (required) | **Example:** Create a VPC in the `aws-us-west-1` cloud region with network range `10.1.2.0/24`: ``` avn vpc create \ --cloud aws-us-west-1 \ --network-cidr 10.1.2.0/24 ``` The command output is similar to: ``` PROJECT_VPC_ID STATE CLOUD_NAME NETWORK_CIDR ==================================== ======== ============= ============ 123abc45-1234-abcd-1234-123abc456def APPROVED aws-us-west-1 10.1.2.0/24 ``` #### Command: `avn organization vpc create`[​](#command-avn-organization-vpc-create "Direct link to command-avn-organization-vpc-create") | Parameter | Information | | ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `--organization-id` | The organization where to create the VPC | | `--cloud` | The cloud and region where to host the VPC. See the list of available cloud regions using the [`avn cloud list`](/docs/tools/cli/cloud.md#avn-cloud-list) command. | | `--network-cidr` | The network range in the Aiven organization VPC in CIDR format (a.b.c.d/e) (required) | **Example:** Create a VPC in the `aws-us-west-1` cloud region with network range `10.1.2.0/24`: ``` avn organization vpc create \ --organization-id org123abc456de \ --cloud aws-us-west-1 \ --network-cidr 10.1.2.0/24 ``` The command output is similar to: ``` CLOUDS CREATE_TIME ORGANIZATION_ID ORGANIZATION_VPC_ID PEERING_CONNECTIONS PENDING_BUILD_ONLY_PEERING_CONNECTIONS STATE UPDATE_TIME ============================================================== ==================== =============== ==================================== =================== ====================================== ======== ==================== {"cloud_name": "aws-us-west-1", "network_cidr": "10.1.2.0/24"} YYYY-MM-DDTHH:MM:SSZ org123abc456de 123abc45-1234-abcd-1234-123abc456def null APPROVED YYYY-MM-DDTHH:MM:SSZ ``` ### Get VPCs[​](#get-vpcs "Direct link to Get VPCs") #### Command: `avn organization vpc get`[​](#command-avn-organization-vpc-get "Direct link to command-avn-organization-vpc-get") | Parameter | Information | | ----------------------- | ---------------------------------------------------------- | | `--organization-id` | The ID of the organization where the organization VPC runs | | `--organization-vpc-id` | The ID of the organization VPC to fetch details for | **Example:** Retrieve information about the organization VPC with ID `abcd1234-abcd-1234-abcd-abcd1234` in organization `org123abc`: ``` avn organization vpc get \ --organization-id org123abc \ --organization-vpc-id abcd1234-abcd-1234-abcd-abcd1234 ``` The command output is similar to: ``` ORGANIZATION_VPC_ID CLOUDS STATE ================================ =============================================================== ====== abcd1234-abcd-1234-abcd-abcd1234 {"cloud_name": "cloud-region-n", "network_cidr": "NN.N.N.N/NN"} ACTIVE ``` ### Delete VPCs[​](#delete-vpcs "Direct link to Delete VPCs") * Project VPC * Organization VPC #### Command: `avn vpc delete`[​](#command-avn-vpc-delete "Direct link to command-avn-vpc-delete") | Parameter | Information | | ------------------ | -------------------------------------------------------------------------------- | | `--project` | The project to use when a project isn't specified for an `avn` command | | `--project-vpc-id` | The project VPC ID. To get the list of VPC IDs execute `avn vpc list` (required) | **Example:** Delete the VPC with id `abcd1234-abcd-1234-abcd-abcd1234`: ``` avn vpc delete \ --project-vpc-id abcd1234-abcd-1234-abcd-abcd1234 ``` The command output is similar to: ``` PROJECT_VPC_ID STATE CLOUD_NAME NETWORK_CIDR ================================ ======== ============= ============ abcd1234-abcd-1234-abcd-abcd1234 DELETING aws-us-west-1 10.1.2.0/24 ``` #### Command: `avn organization vpc delete`[​](#command-avn-organization-vpc-delete "Direct link to command-avn-organization-vpc-delete") | Parameter | Information | | ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `--organization-id` | The ID of the organization hosting the VPC to be deleted (required) | | `--organization-vpc-id` | The ID of the organization VPC to be deleted (required). To get the list of VPC IDs, run [avn organization vpc list](/docs/tools/cli/vpc.md#command-avn-organization-vpc-list). | **Example:** Delete the VPC with id `abcd1234-abcd-1234-abcd-abcd1234` in organization `org123abc`: ``` avn organization vpc delete \ --organization-id org123abc \ --organization-vpc-id abcd1234-abcd-1234-abcd-abcd1234 ``` The command output is similar to: ``` CLOUDS CREATE_TIME ORGANIZATION_ID ORGANIZATION_VPC_ID PEERING_CONNECTIONS PENDING_BUILD_ONLY_PEERING_CONNECTIONS STATE UPDATE_TIME ==================================================================== ==================== =============== ================================ =================== ====================================== ======== ==================== {"cloud_name": "provider-region-n", "network_cidr": "NNN.NN.N.N/NN"} YYYY-MM-DDTHH:MM:SSZ org123abc abcd1234-abcd-1234-abcd-abcd1234 null DELETING YYYY-MM-DDTHH:MM:SSZ ``` ### List VPCs[​](#list-vpcs "Direct link to List VPCs") * Project VPC * Organization VPC #### Command: `avn vpc list`[​](#command-avn-vpc-list "Direct link to command-avn-vpc-list") | Parameter | Information | | ----------- | ---------------------------------------------------------------------- | | `--project` | The project to use when a project isn't specified for an `avn` command | | `--json` | Retrieve the output in JSON format | | `--verbose` | Retrieve the verbose output | **Example:** List all project's VPCs: ``` avn vpc list ``` The command output is similar to: ``` PROJECT_VPC_ID CLOUD_NAME NETWORK_CIDR STATE ==================================== ================== ============= ====== b132dfbf-b035-4cf5-8b15-b7cd6a68aqqd aws-us-east-1 10.2.1.0/24 ACTIVE c36a0a6a-6cfb-4718-93ce-ec043ae94qq5 aws-us-west-2 10.13.4.0/24 ACTIVE d7a984bf-6ebf-4503-bbbd-e7950c49bqqb azure-eastus 10.213.2.0/24 ACTIVE f99601f3-4b00-44d6-b4d9-6f16e9f55qq8 google-us-central1 10.1.13.0/24 ACTIVE 8af49368-3125-48a8-b94e-3d1a3d601qqf google-us-east1 10.50.8.0/24 ACTIVE 6ba650ce-cc08-4e0a-a386-5a354c327qq6 google-us-east4 10.1.17.0/24 ACTIVE c4bc3a59-87da-4dce-9243-c197edb43qq2 google-us-west3 10.1.13.0/24 ACTIVE ``` #### Command: `avn organization vpc list`[​](#command-avn-organization-vpc-list "Direct link to command-avn-organization-vpc-list") | Parameter | Information | | ------------------- | --------------------------------------------------------------- | | `--organization-id` | The ID of the organization hosting VPCs to be listed (required) | | `--json` | Retrieve the output in JSON format | | `--verbose` | Retrieve the verbose output | **Example:** List all organization VPCs for an organization: ``` avn organization vpc list \ --organization-id org123abc ``` The command output is similar to: ``` ORGANIZATION_VPC_ID CLOUD_NAME NETWORK_CIDR STATE ==================================== ================== ============= ====== b132dfbf-b035-4cf5-8b15-b7cd6a68aqqd aws-us-east-1 10.2.1.0/24 ACTIVE c36a0a6a-6cfb-4718-93ce-ec043ae94qq5 aws-us-west-2 10.13.4.0/24 ACTIVE d7a984bf-6ebf-4503-bbbd-e7950c49bqqb azure-eastus 10.213.2.0/24 ACTIVE f99601f3-4b00-44d6-b4d9-6f16e9f55qq8 google-us-central1 10.1.13.0/24 ACTIVE 8af49368-3125-48a8-b94e-3d1a3d601qqf google-us-east1 10.50.8.0/24 ACTIVE 6ba650ce-cc08-4e0a-a386-5a354c327qq6 google-us-east4 10.1.17.0/24 ACTIVE c4bc3a59-87da-4dce-9243-c197edb43qq2 google-us-west3 10.1.13.0/24 ACTIVE ``` ## Manage VPC peering connections[​](#manage-vpc-peering-connections "Direct link to Manage VPC peering connections") ### Create peering connections[​](#create-peering-connections "Direct link to Create peering connections") * Project VPC * Organization VPC #### Command: `avn vpc peering-connection create`[​](#command-avn-vpc-peering-connection-create "Direct link to command-avn-vpc-peering-connection-create") | Parameter | Information | | -------------------------- | ------------------------------------------------------------------------------------- | | `--project` | The project to use when a project isn't specified for an `avn` command | | `--project-vpc-id` | Aiven project VPC ID. To get the list of VPC IDs execute `avn vpc list` (required) | | `--peer-cloud-account` | AWS account ID, Google project ID, or Azure subscription ID (required) | | `--peer-vpc` | AWS VPC ID, Google VPC network name, or Azure VNet name (required) | | `--peer-region` | AWS region of peer VPC, if different than the region defined in the Aiven project VPC | | `--peer-resource-group` | Azure resource group name (required for Azure) | | `--peer-azure-app-id` | Azure app object ID (required for Azure) | | `--peer-azure-tenant-id` | Azure AD tenant ID (required for Azure) | | `--user-peer-network-cidr` | User-defined peer network IP range for routing/firewall | **Example:** Create a peering connection for AWS. ``` avn vpc peering-connection create \ --project-vpc-id b032dfbf-b035-4cf5-8b15-b7cd6a68aqqd \ --peer-cloud-account 012345678901 \ --peer-vpc vpc-abcdef01234567890 ``` The command output is: ``` CREATE_TIME PEER_AZURE_APP_ID PEER_AZURE_TENANT_ID PEER_CLOUD_ACCOUNT PEER_RESOURCE_GROUP PEER_VPC STATE STATE_INFO UPDATE_TIME USER_PEER_NETWORK_CIDRS VPC_PEERING_CONNECTION_TYPE ==================== ================= ==================== ================== =================== ===================== ======== ========== ==================== ======================= =========================== 2022-06-15T14:50:54Z null null 012345678901 null vpc-abcdef01234567890 APPROVED null 2022-06-15T14:50:54Z ``` #### Command: `avn organization vpc peering-connection create`[​](#command-avn-organization-vpc-peering-connection-create "Direct link to command-avn-organization-vpc-peering-connection-create") | Parameter | Information | | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `--organization-id` | The ID of the Aiven organization hosting the VPC to be peered (required) | | `--organization-vpc-id` | The ID of the Aiven organization VPC to be peered (required). To get the list of VPC IDs, run [avn organization vpc list](/docs/tools/cli/vpc.md#command-avn-organization-vpc-list). | | `--peer-cloud-account` | AWS account ID, Google project ID, Azure subscription ID, or the `upcloud` string for UpCloud (required) | | `--peer-vpc` | AWS VPC ID, Google VPC network name, Azure VNet name, or UpCloud private network UUID (required) | | `--peer-region` | AWS region of peer VPC, if different than the region defined in the Aiven organization VPC | | `--peer-resource-group` | Azure resource group name (required for Azure) | | `--peer-azure-app-id` | Azure app object ID (required for Azure) | | `--peer-azure-tenant-id` | Azure AD tenant ID (required for Azure) | | `--user-peer-network-cidr` | User-defined peer network IP range for routing/firewall | **Example:** Create a peering connection for UpCloud: ``` avn organization vpc peering-connection create \ --organization-id org123abc456de \ --organization-vpc-id 123abc45-abcd-1234-abcd-123abc456def \ --peer-cloud-account upcloud \ --peer-vpc abcd1234-abcd-1234-abcd-abcd1234abcd ``` The command output is similar to: ``` PEER_CLOUD_ACCOUNT PEER_RESOURCE_GROUP PEER_VPC PEER_REGION STATE ================== =================== ==================================== =========== ======== upcloud null abcd1234-abcd-1234-abcd-abcd1234abcd null APPROVED ``` ### Delete peering connections[​](#delete-peering-connections "Direct link to Delete peering connections") * Project VPC * Organization VPC #### Command: `avn vpc peering-connection delete`[​](#command-avn-vpc-peering-connection-delete "Direct link to command-avn-vpc-peering-connection-delete") | Parameter | Information | | ----------------------- | ------------------------------------------------------------------------------------- | | `--project` | The project to use when a project isn't specified for an `avn` command | | `--project-vpc-id` | Aiven project VPC ID. To get the list of VPC IDs execute `avn vpc list` (required) | | `--peer-cloud-account` | AWS account ID, Google project ID, or Azure subscription ID (required) | | `--peer-vpc` | AWS VPC ID, Google VPC network name, or Azure VNet name (required) | | `--peer-region` | AWS region of peer VPC, if different than the region defined in the Aiven project VPC | | `--peer-resource-group` | Azure resource group name (required for Azure) | **Example:** Delete the VPC peering connection between the `b032dfbf-b035-4cf5-8b15-b7cd6a68aqqd` Aiven VPC and the `vpc-abcdef01234567890` AWS VPC. ``` avn vpc peering-connection delete \ --project-vpc-id b032dfbf-b035-4cf5-8b15-b7cd6a68aqqd \ --peer-cloud-account 012345678901 \ --peer-vpc vpc-abcdef01234567890 ``` The command output is: ``` CREATE_TIME PEER_AZURE_APP_ID PEER_AZURE_TENANT_ID PEER_CLOUD_ACCOUNT PEER_REGION PEER_RESOURCE_GROUP PEER_VPC STATE STATE_INFO UPDATE_TIME USER_PEER_NETWORK_CIDRS VPC_PEERING_CONNECTION_TYPE ==================== ================= ==================== ================== =========== =================== ===================== ======== ========== ==================== ======================= =========================== 2022-06-15T14:50:54Z null null 012345678901 us-east-1 null vpc-abcdef01234567890 DELETING null 2022-06-15T15:02:12Z ``` #### Command: `avn organization vpc peering-connection delete`[​](#command-avn-organization-vpc-peering-connection-delete "Direct link to command-avn-organization-vpc-peering-connection-delete") | Parameter | Information | | ------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `--organization-id` | The ID of the Aiven organization hosting the peered VPC | | `--organization-vpc-id` | The ID of the Aiven organization VPC with the peering to be deleted (required). To get the list of VPC IDs, run [avn organization vpc list](/docs/tools/cli/vpc.md#command-avn-organization-vpc-list). | | `--peering-connection-id` | The ID of the peering connection to be deleted (required). To get the list of peering connection IDs, run [avn organization vpc peering-connection list](/docs/tools/cli/vpc.md#command-avn-organization-vpc-peering-connection-list). | **Example:** Delete the VPC peering connection between the `123abc45-abcd-1234-abcd-123abc456def` Aiven VPC and the `abc123ab-1234-abcd-1234-456def123abc` UpCloud network: ``` avn organization vpc peering-connection delete \ --organization-id org123abc456de \ --organization-vpc-id 123abc45-abcd-1234-abcd-123abc456def \ --peering-connection ``` The command output is similar to: ``` CREATE_TIME PEER_AZURE_APP_ID PEER_AZURE_TENANT_ID PEER_CLOUD_ACCOUNT PEER_REGION PEER_RESOURCE_GROUP PEER_VPC STATE STATE_INFO UPDATE_TIME USER_PEER_NETWORK_CIDRS VPC_PEERING_CONNECTION_TYPE ==================== ================= ==================== ================== =========== =================== ===================== ======== ========== ==================== ======================= =========================== 2022-06-15T14:50:54Z null null 012345678901 us-east-1 null vpc-abcdef01234567890 DELETING null 2022-06-15T15:02:12Z ``` ### Get peering connections[​](#get-peering-connections "Direct link to Get peering connections") #### Command: `avn vpc peering-connection get`[​](#command-avn-vpc-peering-connection-get "Direct link to command-avn-vpc-peering-connection-get") | Parameter | Information | | ---------------------- | ---------------------------------------------------------------------------------- | | `--project` | The project to use when a project isn't specified for an `avn` command | | `--project-vpc-id` | Aiven project VPC ID. To get the list of VPC IDs execute `avn vpc list` (required) | | `--peer-cloud-account` | AWS account ID, Google project ID, or Azure subscription ID (required) | | `--peer-vpc` | AWS VPC ID, Google VPC network name, or Azure VNet name (required) | | `--json` | Retrieve the output in JSON format | | `--verbose` | Retrieve the verbose output | **Example:** Fetch VPC peering connection details. ``` avn vpc peering-connection get \ --project-vpc-id b032dfbf-b035-4cf5-8b15-b7cd6a68aabd \ --peer-cloud-account 012345678901 \ --peer-vpc vpc-abcdef01234567890 ``` The command output is: ``` State: ACTIVE Message: Peering connection active AWS_VPC_PEERING_CONNECTION_ID TYPE ============================= ================================= pcx-abcdef01234567890 aws-vpc-peering-connection-active ``` ### List peering connections[​](#list-peering-connections "Direct link to List peering connections") * Project VPC * Organization VPC #### Command: `avn vpc peering-connection list`[​](#command-avn-vpc-peering-connection-list "Direct link to command-avn-vpc-peering-connection-list") | Parameter | Information | | ------------------ | ---------------------------------------------------------------------------------- | | `--project` | The project to use when a project isn't specified for an `avn` command | | `--project-vpc-id` | Aiven project VPC ID. To get the list of VPC IDs execute `avn vpc list` (required) | **Example:** List VPC peering connections for the VPC with id `b032dfbf-b035-4cf5-8b15-b7cd6a68aabd`. ``` avn vpc peering-connection list --project-vpc-id b032dfbf-b035-4cf5-8b15-b7cd6a68aabd ``` The command output is: ``` PEER_CLOUD_ACCOUNT PEER_RESOURCE_GROUP PEER_VPC PEER_REGION STATE ================== =================== ===================== =========== ====== 012345678901 null vpc-abcdef01234567890 us-east-1 ACTIVE ``` #### Command: `avn organization vpc peering-connection list`[​](#command-avn-organization-vpc-peering-connection-list "Direct link to command-avn-organization-vpc-peering-connection-list") | Parameter | Information | | ----------------------- | ------------------------------------------------------------------------------------------- | | `--organization-id` | The organization where the peered VPC resides | | `--organization-vpc-id` | The ID of the peered VPC obtainable with the `avn organization vpc list` command (required) | **Example:** List VPC peering connections for the VPC with id `b032dfbf-b035-4cf5-8b15-b7cd6a68aabd` in the `org123abc456de` organization. ``` avn organization vpc peering-connection list \ --organization-id org123abc456de \ --organization-vpc-id b032dfbf-b035-4cf5-8b15-b7cd6a68aabd ``` The command output is similar to: ``` PEERING_CONNECTION_ID PEER_CLOUD_ACCOUNT PEER_RESOURCE_GROUP PEER_VPC PEER_REGION STATE ==================================== ==================================== =================== ======== =========== ============ 123abc45-abcd-1234-abcd-123abc456def 123abc45-1234-abcd-1234-123abc456def test_resource_group test_net null PENDING_PEER ``` --- # Monitor Aiven documentation changes with GitHub Actions Set up automated monitoring of the Aiven documentation to track changes and get notifications when content is updated. To monitor the [Aiven documentation](https://aiven.io/docs) for changes, set up an automated GitHub Actions workflow in your own personal or company account. ## Benefits of automated monitoring[​](#benefits-of-automated-monitoring "Direct link to Benefits of automated monitoring") By setting this up, you gain three major advantages over checking the site manually: * **Diff history**: Because the script commits the new version to your repository, you can click the **Commits** tab in GitHub to see a line-by-line comparison of what was added or removed. * **Zero noise**: You only get a Slack ping when a functional change is made to the document map. * **LLM ready**: If you use AI to manage your Aiven services, you can point your AI tool (like an **MCP server**) at your `current_llms.txt` to ensure it always has the most recent documentation context. ## Set up automated monitoring[​](#set-up-automated-monitoring "Direct link to Set up automated monitoring") ### Create a monitoring repository[​](#create-a-monitoring-repository "Direct link to Create a monitoring repository") You need a home repository for your monitor. This repository stores the script and a history of the `llms.txt` file so you can see exactly what changed over time. 1. Log into GitHub and create a repository, for example, `aiven-docs-monitor`. 2. Set this to **Private** if you don't want others to see your monitoring activity. ### Set up notifications[​](#set-up-notifications "Direct link to Set up notifications") Rather than checking GitHub manually for changes, set up notifications to receive automatic alerts when changes are detected. #### Set up Slack notifications[​](#set-up-slack-notifications "Direct link to Set up Slack notifications") Use a Slack webhook: 1. Create an **Incoming Webhook** in your Slack workspace. 2. In your GitHub repository, go to **Settings** > **Secrets and variables** > **Actions**. 3. Click **New repository secret** and name it `SLACK_WEBHOOK_URL`. Paste your webhook link as the value. #### Set up email notifications[​](#set-up-email-notifications "Direct link to Set up email notifications") Configure SMTP credentials: 1. **Set up an email account for sending notifications**: * Use an existing Gmail account, or create a new one specifically for notifications. * For Gmail, you'll need to create an App Password rather than using your regular password. 2. **Create the App Password for Gmail**: * Go to your Google Account settings. * Navigate to **Security** > **2-Step Verification** > **App passwords**. * Select **Mail** and **Other (Custom name)**. * Enter `Aiven Docs Monitor` as the name. * Copy the generated 16-character password. 3. **Configure GitHub secrets**: * In your GitHub repository, go to **Settings** > **Secrets and variables** > **Actions**. * Click **New repository secret** and add: * `SMTP_USERNAME`: Your Gmail address (for example, `notifications@example.com`) * `SMTP_PASSWORD`: The app password you generated 4. **Update the recipient email**: * In the automation script, change `john.doe@example.com` to your email address. * You can add multiple recipients by separating emails with commas. Alternative SMTP providers You can use other email providers instead of Gmail. Update the `server_address` and `server_port` in the script accordingly: * **Outlook.com**: `smtp-mail.outlook.com`, port `587` * **Yahoo**: `smtp.mail.yahoo.com`, port `587` * **Custom SMTP**: Contact your email provider for the correct settings ### Create the automation script[​](#create-the-automation-script "Direct link to Create the automation script") Inside your new repository, create a folder path: `.github/workflows/`. Inside that folder, create a file named `monitor.yml`. Add the following code: ``` name: Monitor Aiven docs on: schedule: - cron: '0 9 * * *' # Runs daily at 9:00 AM UTC workflow_dispatch: # Allows you to run it manually to test permissions: contents: write jobs: check-changes: runs-on: ubuntu-latest steps: - name: Checkout repository uses: actions/checkout@v4 - name: Fetch latest llms.txt run: | curl -s https://aiven.io/docs/llms.txt -o latest_llms.txt - name: Compare and notify id: compare run: | if [ -f current_llms.txt ]; then if ! cmp -s current_llms.txt latest_llms.txt; then echo "changed=true" >> $GITHUB_OUTPUT else echo "changed=false" >> $GITHUB_OUTPUT fi else cp latest_llms.txt current_llms.txt echo "changed=false" >> $GITHUB_OUTPUT fi - name: Send slack alert if: steps.compare.outputs.changed == 'true' env: SLACK_WEBHOOK: ${{ secrets.SLACK_WEBHOOK_URL }} run: | curl -X POST -H 'Content-type: application/json' \ --data "{\"text\":\"🔔 *Aiven Docs Update:* Changes detected in llms.txt. View live: https://aiven.io/docs/llms.txt\"}" \ $SLACK_WEBHOOK - name: Send email alert if: steps.compare.outputs.changed == 'true' uses: dawidd6/action-send-mail@v3 with: server_address: smtp.gmail.com server_port: 587 username: ${{ secrets.SMTP_USERNAME }} password: ${{ secrets.SMTP_PASSWORD }} subject: 🔔 Aiven docs update detected to: john.doe@example.com from: ${{ secrets.SMTP_USERNAME }} body: | Hello, Changes have been detected in the Aiven documentation llms.txt file. 🔍 What changed: The llms.txt file has been updated 📅 Detection time: ${{ github.run_id }} - ${{ github.run_number }} 🌐 View live document: https://aiven.io/docs/llms.txt 📊 Workflow run: ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }} This is an automated notification from the Aiven docs monitor. Best regards, Aiven docs monitor - name: Save changes to history if: steps.compare.outputs.changed == 'true' run: | git config user.name "Aiven Monitor" git config user.email "monitor@example.com" mv latest_llms.txt current_llms.txt git add current_llms.txt git commit -m "Detect changes in Aiven llms.txt" git push ``` tip You can expand this script to monitor specific product folders, like only PostgreSQL® or Valkey™. ## Test the monitoring setup[​](#test-the-monitoring-setup "Direct link to Test the monitoring setup") Once you save the file, go to the **Actions** tab in your GitHub repository, select **Monitor Aiven Docs**, and click **Run workflow**. This performs the first capture of the file. --- # Aiven Operator for Kubernetes® Manage Aiven infrastructure with [Aiven Operator for Kubernetes®](https://github.com/aiven/aiven-operator/) by using [Custom Resource Definitions (CRD)](https://kubernetes.io/docs/tasks/extend-kubernetes/custom-resources/custom-resource-definitions/). The [Aiven Kubernetes Operator documentation](https://aiven.github.io/aiven-operator/index.html) includes an API reference and example usage for the resources. ## Get started[​](#get-started "Direct link to Get started") Take your first steps by configuring the Aiven Operator and deploying a PostgreSQL® database. ### Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Sign up for Aiven](https://console.aiven.io/signup). * [Install the Aiven Operator](https://aiven.github.io/aiven-operator/installation/helm.html). * Have admin access to a Kubernetes cluster where you can run the operator. * Create a [personal token](/docs/platform/howto/create_authentication_token.md). * [Create a Kubernetes Secret](https://aiven.github.io/aiven-operator/authentication.html). ### Deploy Aiven for PostgreSQL®[​](#deploy-aiven-for-postgresql "Direct link to Deploy Aiven for PostgreSQL®") This example creates an Aiven for PostgreSQL service using the operator's custom resource: 1. Create a file named `pg-sample.yaml` and add the following: ``` apiVersion: aiven.io/v1alpha1 kind: PostgreSQL metadata: name: pg-sample spec: # Gets the token from the `aiven-token` secret authSecretRef: name: aiven-token key: token # Outputs the PostgreSQL connection information to the `pg-connection` secret connInfoSecretTarget: name: pg-connection project: PROJECT-NAME cloudName: google-europe-west1 plan: startup-4 maintenanceWindowDow: friday maintenanceWindowTime: 23:00:00 # PostgreSQL configuration userConfig: pg_version: '16' ``` Where `PROJECT-NAME` is the name of the Aiven project to create the service in. 2. To apply the resource, run: ``` kubectl apply -f pg-sample.yaml ``` 3. Verify the status of your service by running: ``` kubectl get postgresqls.aiven.io pg-sample ``` Once the `STATE` field is `RUNNING`, your service is ready for use. The connection information is automatically created by the operator within a Kubernetes Secret named `pg-connection`. For PostgreSQL, the connection information is the service URI. ### Create a connection pool[​](#create-a-connection-pool "Direct link to Create a connection pool") This example creates a [PgBouncer connection pool](/docs/products/postgresql/concepts/pg-connection-pooling.md) for the `pg-sample` service. 1. Create a file named `pg-pool.yaml` and add the following: ``` apiVersion: aiven.io/v1alpha1 kind: ConnectionPool metadata: name: pg-connection-pool spec: authSecretRef: name: aiven-token key: token # Outputs the pool connection information to the `pg-connection-pool-secret` secret connInfoSecretTarget: name: pg-connection-pool-secret project: PROJECT-NAME serviceName: pg-sample databaseName: defaultdb username: avnadmin poolMode: transaction poolSize: 25 ``` Where `PROJECT-NAME` is the name of the Aiven project, `defaultdb` is an existing database on the `pg-sample` service, and `avnadmin` is an existing service user. note `poolSize` can be from 1 to 1000 with the Aiven Operator for Kubernetes, lower than the maximum of 10000 available through the console, CLI, Terraform, or API. 2. To apply the resource, run: ``` kubectl apply -f pg-pool.yaml ``` 3. Verify the status of your connection pool by running: ``` kubectl get connectionpools.aiven.io pg-connection-pool ``` ### Use the service[​](#use-the-service "Direct link to Use the service") When the service is running, you can deploy a pod to test the connection to PostgreSQL from Kubernetes. 1. Create a file named `pod-psql.yaml` with the following: ``` apiVersion: v1 kind: Pod metadata: name: psql-test-connection spec: restartPolicy: Never containers: - image: postgres:11-alpine name: postgres command: ['psql', '$(DATABASE_URI)', '-c', 'SELECT version();'] # The pg-connection secret becomes environment variables envFrom: - secretRef: name: pg-connection ``` 2. Apply the resource to create the pod and test the connection by running: ``` kubectl apply -f pod-psql.yaml ``` 3. To check the logs, run: ``` kubectl logs psql-test-connection ``` ### Clean up[​](#clean-up "Direct link to Clean up") 1. To destroy the resources, run: ``` kubectl delete pod psql-test-connection kubectl delete connectionpools.aiven.io pg-connection-pool kubectl delete postgresqls.aiven.io pg-sample ``` 2. To remove the operator and `cert-manager`, run: ``` helm uninstall aiven-operator helm uninstall aiven-operator-crds kubectl delete -f https://github.com/jetstack/cert-manager/releases/latest/download/cert-manager.yaml ``` ## Related links[​](#related-links "Direct link to Related links") * [Aiven Operator for Kubernetes repository](https://github.com/aiven/aiven-operator/) * [Aiven Kubernetes Operator examples](https://aiven.github.io/aiven-operator/resources/project.html) * [Kubernetes Basics](https://kubernetes.io/docs/tutorials/kubernetes-basics/) --- # Aiven MCP Create and manage Aiven services from AI assistants such as Cursor and Claude Code, including PostgreSQL®, Apache Kafka®, plans, metrics, logs, and service configuration. Use [read-only mode](#read-only-mode) to limit MCP tools to viewing services and other resources. You can enable it in your client or as an organization policy. You can also limit tools to specific services to keep the assistant focused. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * An [Aiven account](https://console.aiven.io/signup?campaign=mcp) * An MCP-compatible client, such as Cursor, Claude Code, Claude Desktop, VS Code, or Gemini CLI * MCP access enabled for your organization by an organization admin. In the Aiven Console, click **Admin** > **Authentication** > **Allow MCP connections**. Authentication uses OAuth 2.0 with PKCE. When you authenticate a client for the first time, your browser opens so you can sign in and authorize MCP access. This is usually a one-time setup step per client. note Aiven owns the `aiven.live` domain and operates the hosted MCP server at `mcp.aiven.live`. The server uses the Aiven API at `api.aiven.io` for authorization. ## Install the Aiven MCP[​](#configure-aiven-mcp "Direct link to Install the Aiven MCP") * Claude Code * Cursor * Claude Desktop * VS Code * Gemini CLI * Other clients * Local installation 1. Open a terminal. 2. Choose your options, then run the generated command: Services \[x]All\[ ]PostgreSQL\[ ]Kafka\[ ]Integrations AccessFull access (full) Read and write. The assistant can create, modify, and delete. ⚠ With full access, the assistant can create, modify, and delete services and data. [Review security and responsibility](#security-and-responsibility). Advanced settings Development\[ ]Allow connection credentials (development only) MarketplaceNone (Aiven) Select your marketplace only if you subscribed to Aiven through AWS, Azure, or Google Cloud Marketplace. This helps you sign in to the correct console. ``` claude mcp add --transport http aiven "https://mcp.aiven.live/mcp" ``` 3. Run `claude` to start Claude Code. 4. Run `/mcp`, select **aiven**, and authenticate in your browser when prompted. 5. Test the connection with a prompt such as `List my Aiven projects.` and approve the tool execution if prompted. For more information, see the [Claude Code MCP documentation](https://code.claude.com/docs/en/mcp). For one-click installation: [Add to Cursor](cursor://anysphere.cursor-deeplink/mcp/install?name=aiven-mcp\&config=eyJ1cmwiOiJodHRwczovL21jcC5haXZlbi5saXZlL21jcCIsInR5cGUiOiJodHRwIn0=) To add the server manually: 1. In your project root, create or edit the `.cursor/mcp.json` file. 2. Choose your options below, then copy the generated configuration into the file: Services \[x]All\[ ]PostgreSQL\[ ]Kafka\[ ]Integrations AccessFull access (full) Read and write. The assistant can create, modify, and delete. ⚠ With full access, the assistant can create, modify, and delete services and data. [Review security and responsibility](#security-and-responsibility). Advanced settings Development\[ ]Allow connection credentials (development only) MarketplaceNone (Aiven) Select your marketplace only if you subscribed to Aiven through AWS, Azure, or Google Cloud Marketplace. This helps you sign in to the correct console. ``` { "mcpServers": { "aiven": { "type": "http", "url": "https://mcp.aiven.live/mcp" } } } ``` 3. Save the file. 4. Restart Cursor. 5. Open **Settings** > **Tools & MCP**. 6. Select **aiven** and click **Connect**. In Cursor Chat (**Cmd+L** / **Ctrl+L**), test the connection with a prompt such as `List my Aiven projects.` and click **Allow** if prompted. For more information, see the [Cursor MCP documentation](https://cursor.com/docs/mcp). 1. Open the Claude Desktop configuration file. If it does not exist, create it: * **macOS:** `~/Library/Application Support/Claude/claude_desktop_config.json` * **Windows:** `%APPDATA%\Claude\claude_desktop_config.json` 2. Choose your options, then add the generated configuration to the file: Services \[x]All\[ ]PostgreSQL\[ ]Kafka\[ ]Integrations AccessFull access (full) Read and write. The assistant can create, modify, and delete. ⚠ With full access, the assistant can create, modify, and delete services and data. [Review security and responsibility](#security-and-responsibility). Advanced settings Development\[ ]Allow connection credentials (development only) MarketplaceNone (Aiven) Select your marketplace only if you subscribed to Aiven through AWS, Azure, or Google Cloud Marketplace. This helps you sign in to the correct console. ``` { "mcpServers": { "aiven": { "command": "npx", "args": [ "-y", "mcp-remote", "https://mcp.aiven.live/mcp" ] } } } ``` 3. Save the file. 4. Restart Claude Desktop. 5. In a new conversation, test the connection with a prompt such as `List my Aiven projects.` and click **Allow** if prompted. For more information, see the [Claude Desktop MCP documentation](https://modelcontextprotocol.io/docs/develop/connect-local-servers). note Requires VS Code 1.102 or later with the GitHub Copilot extension installed and enabled. 1. Open your workspace in VS Code. 2. In the workspace root, create a `.vscode` directory. 3. In the `.vscode` directory, create or edit the `mcp.json` file. 4. Choose your options, then add the generated configuration to the file: Services \[x]All\[ ]PostgreSQL\[ ]Kafka\[ ]Integrations AccessFull access (full) Read and write. The assistant can create, modify, and delete. ⚠ With full access, the assistant can create, modify, and delete services and data. [Review security and responsibility](#security-and-responsibility). Advanced settings Development\[ ]Allow connection credentials (development only) MarketplaceNone (Aiven) Select your marketplace only if you subscribed to Aiven through AWS, Azure, or Google Cloud Marketplace. This helps you sign in to the correct console. ``` { "servers": { "aiven": { "type": "http", "url": "https://mcp.aiven.live/mcp" } } } ``` 5. Save the file. 6. Reload VS Code. 7. Open the Command Palette, run **MCP: List Servers**, and confirm that **aiven** appears. 8. In Copilot Chat, test the connection with a prompt such as `List my Aiven projects.` and click **Allow** if prompted. For more information, see the [VS Code MCP documentation](https://code.visualstudio.com/docs/copilot/customization/mcp-servers). 1. Create or edit the `~/.gemini/settings.json` file. 2. Choose your options, then add the generated configuration to the file: Services \[x]All\[ ]PostgreSQL\[ ]Kafka\[ ]Integrations AccessFull access (full) Read and write. The assistant can create, modify, and delete. ⚠ With full access, the assistant can create, modify, and delete services and data. [Review security and responsibility](#security-and-responsibility). Advanced settings Development\[ ]Allow connection credentials (development only) MarketplaceNone (Aiven) Select your marketplace only if you subscribed to Aiven through AWS, Azure, or Google Cloud Marketplace. This helps you sign in to the correct console. ``` { "mcpServers": { "aiven": { "httpUrl": "https://mcp.aiven.live/mcp" } } } ``` 3. Save the file. 4. Run `gemini` to start the CLI. 5. Run `/mcp auth aiven` to sign in through your browser. 6. Test the connection with a prompt such as `List my Aiven projects.` and approve the tool execution if prompted. 1) Open your MCP client configuration. 2) Choose your options, then add the generated configuration to your client: Services \[x]All\[ ]PostgreSQL\[ ]Kafka\[ ]Integrations AccessFull access (full) Read and write. The assistant can create, modify, and delete. ⚠ With full access, the assistant can create, modify, and delete services and data. [Review security and responsibility](#security-and-responsibility). Advanced settings Development\[ ]Allow connection credentials (development only) MarketplaceNone (Aiven) Select your marketplace only if you subscribed to Aiven through AWS, Azure, or Google Cloud Marketplace. This helps you sign in to the correct console. ``` { "mcpServers": { "aiven": { "url": "https://mcp.aiven.live/mcp", "oauth": { "oauthScopes": [ "projects", "services", "accounts:read" ] } } } } ``` Most clients use a configuration similar to the preceding example. 3) Save the file and restart your client. 4) In your AI assistant, test the connection with a prompt such as `List my Aiven projects.` and approve the tool execution if prompted. Some clients require a transport type, such as `"type": "http"`. If the configuration fails, see your client documentation. Run the Aiven MCP server locally with `npx` instead of using the hosted server. You need an [Aiven API token](/docs/platform/howto/create_authentication_token.md) to authenticate requests to Aiven. Set `AIVEN_READ_ONLY="true"` to enable read-only mode. Set `AIVEN_ALLOW_SECRETS="true"` to let the AI agent access PostgreSQL and Kafka connection credentials. Use this option only in development environments for non-production services. For more information, see [Security and responsibility](#security-and-responsibility). * Claude Code * Cursor * Claude Desktop * VS Code * Other clients 1. Open a terminal. 2. Run the following command, replacing `your-token-here` with your Aiven API token: ``` claude mcp add --scope user aiven-mcp -e AIVEN_TOKEN=your-token-here -e AIVEN_READ_ONLY=false -e AIVEN_ALLOW_SECRETS=false -- npx -y mcp-aiven ``` 3. Run `claude` to start Claude Code. 4. Run `/mcp` to verify that the server is registered. 1) In your project root, create or edit the `.cursor/mcp.json` file. 2) Add the following configuration, replacing `your-token-here` with your Aiven API token: ``` { "mcpServers": { "aiven-mcp": { "command": "npx", "args": ["-y", "mcp-aiven"], "env": { "AIVEN_TOKEN": "your-token-here", "AIVEN_READ_ONLY": "false", "AIVEN_ALLOW_SECRETS": "false" } } } } ``` 3) Save the file and restart Cursor. 1. Open the Claude Desktop configuration file. 2. Add the following configuration, replacing `your-token-here` with your Aiven API token: ``` { "mcpServers": { "aiven-mcp": { "command": "npx", "args": ["-y", "mcp-aiven"], "env": { "AIVEN_TOKEN": "your-token-here", "AIVEN_READ_ONLY": "false", "AIVEN_ALLOW_SECRETS": "false" } } } } ``` 3. Save the file and restart Claude Desktop. 1) In the `.vscode` directory, create or edit the `mcp.json` file. 2) Add the following configuration, replacing `your-token-here` with your Aiven API token: ``` { "servers": { "aiven-mcp": { "command": "npx", "args": ["-y", "mcp-aiven"], "env": { "AIVEN_TOKEN": "your-token-here", "AIVEN_READ_ONLY": "false", "AIVEN_ALLOW_SECRETS": "false" } } } } ``` 3) Save the file and reload VS Code. 1. Open your MCP client configuration. 2. Add the following configuration, replacing `your-token-here` with your Aiven API token: ``` { "mcpServers": { "aiven-mcp": { "command": "npx", "args": ["-y", "mcp-aiven"], "env": { "AIVEN_TOKEN": "your-token-here" } } } } ``` 3. Save the file and restart your client. ## What Aiven MCP can do[​](#what-aiven-mcp-can-do "Direct link to What Aiven MCP can do") After you connect to Aiven MCP, you can work with Aiven resources in natural language. For example, you can do the following: * **View resources**: List projects, services, and integrations, or check the status, plan, and cloud region of a service. * **Manage services**: Create, update, and delete services such as PostgreSQL® and Apache Kafka®. You can also change service plans and configuration. [Read-only mode](#read-only-mode) restricts MCP tools to read-only operations, such as viewing services and other resources. * **Inspect and troubleshoot services**: View service metrics, logs, and configuration to investigate issues. * **Use Aiven documentation**: Ask questions and get answers based on the Aiven documentation. ## Read-only mode[​](#read-only-mode "Direct link to Read-only mode") Read-only mode restricts MCP tools to viewing services and other resources. You cannot create, modify, or delete resources. Read-only mode applies at the organization level and the client level. ### Organization level[​](#organization-level "Direct link to Organization level") Organization admins can restrict MCP connections to read-only operations for all users in the organization. 1. In the organization, click **Admin** > **Authentication**. 2. Select **Restrict MCP connections to read-only operations**. Users cannot override this setting in their client configuration. Write operations fail even if a user configures their client without read-only mode. For more information, see [Set authentication policies for organization users](/docs/platform/howto/set-authentication-policies.md). ### Client level[​](#client-level "Direct link to Client level") You can enable read-only mode for your client in the [installation configuration](#configure-aiven-mcp). The setting applies only to that client and does not change the organization policy. It cannot override an organization-level restriction. ## Security and responsibility[​](#security-and-responsibility "Direct link to Security and responsibility") important MCP tools can perform destructive operations on your Aiven services, including creating, modifying, and deleting services, databases, topics, and data. AI agents run operations from natural language prompts, which can be misinterpreted. Using the Aiven MCP server can result in damage to or loss of data. Aiven secures the MCP server and data in transit. Your selected AI agent provider determines how the agent uses your data, including whether it uses that data for training. Review the provider's terms before you enable the integration. Under the [shared responsibility model](https://aiven.io/responsibility-matrix), security and compliance for MCP usage are shared between Aiven and your organization. Aiven secures the platform and API. You are responsible for the following: * **Decide on MCP access.** Evaluate whether to enable MCP in your organization and the associated risks. * **Control access.** Scope API tokens to the minimum permissions needed and rotate them regularly. * **Review AI agent actions.** Review actions before they run, especially write or delete operations on production resources. * **Limit write access.** Use [read-only mode](#read-only-mode) at the organization level for all users, or at the client level for an individual client. * **Keep credentials off in production.** The **Allow connection credentials** option (`allow_secrets=true`) returns PostgreSQL and Kafka connection credentials, including URIs, passwords, and certificates, to the AI agent. Use it only for development with non-production services that do not hold sensitive data. The Aiven MCP server can also connect directly to PostgreSQL services running in a private VPC. --- # Standalone SQL query optimizer [Early availability](/docs/platform/concepts/service-and-feature-releases.md) Use Aiven's AI-powered **SQL query optimizer** for PostgreSQL® and MySQL® to get query optimization recommendations for an ad-hoc query. important If you are running a PostgreSQL or MySQL service, Aiven automatically suggests optimizations for slow queries from the **AI insights** menu entry. Also see [AI Database Optimizer for PG](/docs/products/postgresql/howto/ai-insights.md) and [AI Database Optimizer for MySQL](/docs/products/mysql/howto/ai-insights.md). To optimize a query: 1. Click **Tools** > **SQL query optimizer**. 2. Click **Optimize a query**. 3. Select your database type and version. 4. Paste your query and click **Next**. 5. Optional: 1. Provide your table structure and statistics by running the query provided in the UI. 2. Paste it in the **Query output** field. 6. Click **Optimize**. The optimization report shows the optimized query and potential optimal indexes. To learn more about the recommendations, click **Optimization details**. Frequently asked questions **Does Aiven AI Optimizer mask/obfuscate my queries?** Yes, Aiven AI Optimizer provides a non-intrusive solution to optimize your database performance without compromising sensitive data access. It achieves this by gathering information on schema structure, database statistics, and other signals to detect potential performance problems and offer optimization recommendations, without requiring credentials or access to the actual data in the database. To address the possibility of slow query logs containing sensitive data, Aiven offers data masking capabilities that replace sensitive parameters within queries with question marks (`?`). Data masking is enabled by default. note The masking option is not available for the Standalone SQL query optimizer yet. Related pages * [Optimizing queries in Aiven for PostgreSQL®](/docs/products/postgresql/howto/ai-insights.md) * [Optimizing queries in Aiven for MySQL®](/docs/products/mysql/howto/ai-insights.md) --- # Aiven Provider for Terraform Use the [Aiven Provider for Terraform](https://registry.terraform.io/providers/aiven/aiven/latest/docs) to provision and manage your Aiven infrastructure. tip You can also use an AI assistant connected to [Aiven MCP](/docs/tools/mcp-server.md) to create, update, and view details for Aiven services from clients such as Cursor and Claude Code. ## Get started[​](#get-started "Direct link to Get started") The Aiven Platform uses [organizations, organizational units, and projects](https://aiven.io/docs/platform/concepts/orgs-units-projects) to organize services. This example shows you how to use the Aiven Provider for Terraform to create an organization with two organizational units, and add projects to those units. The following example file is also available in the [Aiven Terraform Provider repository](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/clickhouse) on GitHub. 1. [Sign up for Aiven](https://console.aiven.io/signup?utm_source=github\&utm_medium=organic\&utm_campaign=devportal\&utm_content=repo). 2. [Download and install Terraform](https://www.terraform.io/downloads). 3. [Create a token](/docs/platform/howto/create_authentication_token.md). 4. Create a file named `main.tf` and add the following: ``` Loading... ``` 5. Create a file named `variables.tf` and add the following: ``` Loading... ``` 6. Create the `terraform.tfvars` file and assign values to the variables for the token and the project names. To apply your Terraform configuration: 1. Initialize Terraform by running: ``` terraform init ``` The output is similar to the following: ``` Initializing the backend... Initializing provider plugins... - Finding aiven/aiven versions matching ">= 4.0.0, < 5.0.0"... - Installing aiven/aiven v4.9.2... - Installed aiven/aiven v4.9.2 ... Terraform has been successfully initialized! ... ``` 2. To create an execution plan and preview the changes, run: ``` terraform plan ``` 3. To deploy your changes, run: ``` terraform apply --auto-approve ``` ## Next steps[​](#next-steps "Direct link to Next steps") * Follow an [example to set up your own organization](https://github.com/aiven/terraform-provider-aiven/tree/main/examples/get-started) with a user group and permissions. * Try one of the other [examples](https://github.com/aiven/terraform-provider-aiven/tree/main/examples) to learn how to create a service or integration using the Aiven Terraform Provider. * Get details about all the available resources and data sources in the [Aiven Provider for Terraform documentation](https://registry.terraform.io/providers/aiven/aiven/latest/docs). * Learn how to manage Aiven services with AI assistance using [Aiven MCP](/docs/tools/mcp-server.md). --- # Use OpenTofu with Aiven Provider for Terraform [OpenTofu](https://opentofu.org/) is an open source infrastructure-as-code tool that you can use to configure your Aiven infrastructure. Set up your first Aiven project and service using this example to get started with OpenTofu. If you already use the [Aiven Provider for Terraform](/docs/tools/terraform.md), you can [migrate your Terraform resources to OpenTofu](https://opentofu.org/docs/intro/migration/). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [Sign up for Aiven](https://console.aiven.io/signup?utm_source=github\&utm_medium=organic\&utm_campaign=devportal\&utm_content=repo) * [Install OpenTofu](https://opentofu.org/docs/intro/install/) * [Create a token](/docs/platform/howto/create_authentication_token.md) ## Configure your project and services[​](#configure-your-project-and-services "Direct link to Configure your project and services") Set up the OpenTofu project in an empty folder: 1. Create a file, `provider.tf`, and add the following code to declare a dependency on the Aiven Provider, specifying the [version](https://registry.terraform.io/providers/aiven/aiven/latest). ``` terraform { required_providers { aiven = { source = "aiven/aiven" version = ">=4.0.0, < 5.0.0" } } } provider "aiven" { api_token = var.aiven_token } ``` 2. Create a file named `project.tf` and add the following code to create an Aiven project in your organization: ``` # Get information about your organization data "aiven_organization" "main" { name = "ORGANIZATION_NAME" } # Create a new project in your organization resource "aiven_project" "example_project" { project = "ORGANIZATION_NAME-example-project" parent_id = data.aiven_organization.main.id } ``` Where `ORGANIZATION_NAME` is your [Aiven organization](/docs/platform/concepts/orgs-units-projects.md) name. 3. Create a file named `service.tf` and add the following code to define the configuration of an [Aiven for PostgreSQL®](/docs/products/postgresql.md) service: ``` resource "aiven_pg" "pg" { project = aiven_project.example_project.project cloud_name = "google-europe-west1" plan = "startup-4" service_name = "example-pg" maintenance_window_dow = "monday" maintenance_window_time = "10:00:00" pg_user_config { pg { idle_in_transaction_session_timeout = 900 log_min_duration_statement = -1 } } } ``` 4. Create a file named `variables.tf` and add the following code to declare the Aiven token variable: ``` variable "aiven_token" { description = "Aiven token" type = string } ``` 5. Create a file named `terraform.tfvars` with the following code to store the token value: ``` aiven_token = "AIVEN_TOKEN" ``` Where `AIVEN_TOKEN` is your token. ## Plan and apply the configuration[​](#plan-and-apply "Direct link to Plan and apply the configuration") 1. The `init` command prepares the working directly for use with OpenTofu. Use it to automatically find, download, and install the necessary Aiven Provider plugins: ``` tofu init ``` 2. Run the `plan` command to create an execution plan and preview the changes. This shows you what resources OpenTofu will create or modify: ``` tofu plan ``` The output is similar to the following: ``` OpenTofu used the selected providers to generate the following execution plan. ... Plan: 2 to add, 0 to change, 0 to destroy. ``` 3. To create the resources, run: ``` tofu apply --auto-approve ``` The output is similar to the following: ``` Apply complete! Resources: 2 added, 0 changed, 0 destroyed. ``` You can also see the PostgreSQL service in the [Aiven Console](https://console.aiven.io). ## Clean up[​](#clean-up "Direct link to Clean up") To delete the project, service, and data: 1. Create a destroy plan and preview the changes to your infrastructure with the following command: ``` tofu plan -destroy ``` 2. To delete the resources and all data, run: ``` tofu destroy ``` 3. Enter **yes** to confirm. The output is similar to the following: ``` Do you really want to destroy all resources? OpenTofu will destroy all your managed infrastructure, as shown above. There is no undo. Only 'yes' will be accepted to confirm. Enter a value: yes ... Destroy complete! Resources: 2 destroyed. ``` Related pages * Try OpenTofu with [another sample project](https://github.com/aiven/terraform-provider-aiven/blob/main/sample_project/sample.tf). ---