Enable ML Commons for Aiven for OpenSearch®
Configure ML Commons cluster settings for your Aiven for OpenSearch® service, and deploy a pretrained or externally hosted model.
For background on what ML Commons supports in Aiven for OpenSearch, see ML Commons for Aiven for OpenSearch.
Configure ML Commons cluster settings
Configure the ML Commons cluster settings as
advanced parameters on your Aiven
for OpenSearch service. The following example sets ml_commons_only_run_on_ml_node to
false, which is only needed on a plan without dedicated ML nodes.
- Console
- CLI
- API
- Log in to the Aiven Console.
- Click Services, then select your Aiven for OpenSearch service.
- Click Service settings. Scroll to the Advanced configuration section and click Configure.
- Click Add configuration options, then select an
ml_commons_*option from the list. - Set the value and click Save configuration.
Use the avn service update command
to configure ML Commons settings:
avn service update SERVICE_NAME \
-c opensearch.ml_commons_only_run_on_ml_node=false \
-c opensearch.ml_commons_model_access_control_enabled=true
Replace SERVICE_NAME with your Aiven for OpenSearch service name.
Call the ServiceUpdate endpoint to configure ML Commons settings:
curl --request PUT \
--url "https://api.aiven.io/v1/project/PROJECT_NAME/service/SERVICE_NAME" \
--header "Authorization: Bearer API_TOKEN" \
--header "Content-Type: application/json" \
--data '{
"user_config": {
"opensearch": {
"ml_commons_only_run_on_ml_node": false,
"ml_commons_model_access_control_enabled": true
}
}
}'
Replace PROJECT_NAME, SERVICE_NAME, and API_TOKEN with your values.
The full list of ml_commons_* options, including their types and defaults, is in
Advanced parameters for Aiven for OpenSearch.
Run ML tasks on the right nodes
By default, ml_commons_only_run_on_ml_node is true, so ML tasks only run on nodes
with the ml role.
- If your service uses a cluster plan with dedicated ML nodes, leave this setting at its default so ML tasks run there instead of on data nodes.
- If your service doesn't have dedicated ML nodes, set
ml_commons_only_run_on_ml_nodetofalseso ML tasks can run on data nodes.
Deploy a pretrained model
To deploy an OpenSearch-provided pretrained model, use the ML Commons REST API through the standard OpenSearch API. Registering, deploying, and running predictions with a pretrained model follows the same steps as registering a pretrained model in OpenSearch.
Connect to an externally hosted model
To connect to a model hosted outside your Aiven for OpenSearch service, such as a third-party LLM API:
- Add the endpoint of your model provider to
ml_commons_trusted_connector_endpoints_regex. - Create a connector and register a remote model using the OpenSearch connector blueprint for your model provider.
Restrict access to ML models and connectors
By default, all service users have the ml_full_access role and can register, deploy,
use, and delete any model or connector. To restrict which users can manage or use ML
models and connectors:
- Enable OpenSearch Security management for your service.
- Set
ml_commons_model_access_control_enabled,ml_commons_connector_access_control_enabled, or both totrue. - Map the
ml_full_accessandml_readonly_accessroles to specific backend roles using OpenSearch Security.
Redeploy models after node changes
- If some of a service's
ml-role nodes (ordatanodes, on plans without dedicated ML nodes) become unavailable, for example during a node replacement or plan change, affected models move to aPARTIALLY_DEPLOYEDstate. They redeploy automatically ifml_commons_model_auto_redeploy_enableistrue(the default). - If all of those nodes become unavailable at once, for example during a power cycle,
models move to
DEPLOY_FAILED. Call_deployfor each model to make it available again.
ml_commons_model_auto_deploy_enable (true by default) only applies to externally
hosted models that haven't been deployed yet: it deploys them automatically on their
first prediction request. Pretrained built-in models always need an explicit _deploy
call.
Related pages