Migrate schemas to Karapace
Migrate schemas from Confluent Schema Registry or another Karapace instance to Aiven for Apache Kafka® and keep their original schema IDs and version numbers.
To keep the IDs and versions, you set the target Schema Registry to IMPORT mode
during the migration.
Each message serialized with Schema Registry contains a schema ID. Consumers use this ID to get the schema they need to read the message. When you keep the same IDs, existing consumers can keep reading existing messages after you switch registries.
You can migrate schemas in one of the following ways:
| Method | Use it for | Soft-deleted versions | Reserves the source's highest ID |
|---|---|---|---|
| Script, recommended | All subjects in a registry or an export | Reproduced | Yes |
| Schema Registry API | A few schema versions, manually with curl | Not imported | No |
Reserving the highest ID that the source issued keeps the target from issuing an ID that the source already used.
Prerequisites
- An Aiven for Apache Kafka service with Schema Registry enabled and Karapace 6.2.4 or later.
- The target Schema Registry URL and credentials from the Schema Registry tab of
Connect information on the Overview page of your
service in the Aiven Console. The default
avnadminuser can run the migration. - The source registry URL and credentials that can read schemas and configuration.
- Network access to the source and target registries.
curl.- Python 3, if you use the script. The script needs only the Python standard library.
Depending on your setup, you also need the following:
- If you use a user other than
avnadminwith Schema Registry authorization, add ACL entries with theschema_registry_writepermission forConfig:and for the subjects that you migrate, for exampleSubject:*. - If you use
role-based authorization with OAuth 2.0/OIDC,
define roles in your identity provider that allow
GET,POST,PUT, andDELETErequests on the target, and make sure the access token includes them.
Prepare for the migration
- Run the full migration on a non-production target first, such as a test Aiven service. Then repeat it for production.
- Stop schema changes on the source, including registrations, deletions, and compatibility changes. Stop producers that register schemas automatically, or configure them not to register schemas. Keep schema changes paused until the migration is complete and your clients use the target registry.
- Don't register schemas on the target subjects during the migration.
An export is a snapshot of the source. If anything changes on the source after you export, create another export. Create a separate export for the production run too.
The script can't export versions that were hard deleted on the source.
How import mode works
A subject is a named collection of schema versions. A registry or subject runs in a mode that controls how it handles schema registrations. A migration uses the following modes:
READWRITE: The default mode. The registry assigns IDs and version numbers, reuses the existing ID for an identical schema, and runs compatibility checks.IMPORT: You provide the original IDs and version numbers, and the registry uses them. The registry rejects registrations that don't include an ID, and it doesn't run compatibility checks.
To keep the IDs, set the target registry or subjects to IMPORT mode during the
migration, and return them to READWRITE mode when you finish. The script handles
these mode changes for you. With the API, you change the modes yourself.
You can set IMPORT mode with one of the following scopes:
- Global scope: Sets
IMPORTmode for the whole registry. Use it to migrate a whole registry into a new, empty target. - Subject scope: Sets
IMPORTmode for individual subjects. Use it to import into a registry that already has other subjects.
In both scopes, the registry or the subjects that you import can't have live schemas on
the target. A live schema is a schema version that hasn't been soft-deleted.
Soft-deleted versions don't block IMPORT mode, but they keep their ID and version
number, so they can still conflict with the schemas you import.
To set IMPORT mode on a registry or subject that has live schemas, see
Import into a non-empty registry.
Migrate with the script
The
sr_migrate.py
script exports schemas from the source, imports them into your Aiven service with their
original IDs and versions, and verifies the result.
Script workflow
The import command runs the following nine steps. It shows a
plan and asks you to confirm each one. The step names match the names that the script
prints.
- Export from source or Load export: With
--source, exports the schemas from the source, including soft-delete status, compatibility levels, and the highest schema ID that the source registry issued. With--file, loads an existing export. - Check target: Checks for conflicting subjects, versions, and IDs. It also reports
whether the target has live schemas that block
IMPORTmode. If the target has live schemas, the script asks whether to continue with--force. This step doesn't change the target. - Enter
IMPORTmode: Sets the registry or subjects toIMPORTmode. - Register versions: Registers each version with its original ID and version number, referenced schemas first.
- Reproduce soft deletes: Soft-deletes the versions that are soft-deleted on the source.
- Reserve source max ID: If the source issued an ID higher than any exported ID,
the script reserves that ID in the
--reserve-subjectsubject. For example, this happens when the source hard-deleted the version with the highest ID. Reserving the ID prevents the target from reusing it. The script imports a placeholder schema with that ID and soft-deletes it. - Apply compatibility levels: Applies source compatibility levels that differ from the target.
- Verify: Checks that every imported version has its source schema ID.
- Leave
IMPORTmode: With global scope, sets the registry mode back toREADWRITE. With subject scope, removes the subject mode overrides, so the subjects follow the registry mode. Then checks that nothing in scope is still inIMPORTmode.
With global scope, the script also applies the source's global compatibility level. With subject scope, the script doesn't change the target's global compatibility level, and subjects without their own setting use it. After the import, compare the target's global level with the source's. For more information, see Verify the migration.
Download the script and set authentication
-
Review the
sr_migrate.pysource on GitHub. -
Download the script:
curl --fail --location \https://raw.githubusercontent.com/Aiven-Open/karapace/main/bin/sr_migrate.py \--output sr_migrate.py -
Set
SRC_AUTHfor the source registry andDST_AUTHfor the target registry. Each variable holds the complete value of the HTTPAuthorizationheader.For HTTP Basic authentication:
export SRC_AUTH="Basic $(printf '%s' 'SOURCE_USER:SOURCE_PASSWORD' | base64 | tr -d '\n')"export DST_AUTH="Basic $(printf '%s' 'SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD' | base64 | tr -d '\n')"Replace the following:
SOURCE_USERandSOURCE_PASSWORD: credentials for the source registry.SCHEMA_REGISTRY_USERandSCHEMA_REGISTRY_PASSWORD: credentials for the target registry, from the Schema Registry tab of Connect information in the Aiven Console.
For OAuth 2.0/OIDC authentication:
export SRC_AUTH="Bearer SOURCE_ACCESS_TOKEN"export DST_AUTH="Bearer TARGET_ACCESS_TOKEN"Replace
SOURCE_ACCESS_TOKENandTARGET_ACCESS_TOKENwith the access tokens for the source and target registries.The registry reads authorization claims from the same token to authorize the requests that set the
IMPORTmode and register schemas. Define the roles from the prerequisites in your identity provider, and make sure the token includes them.
The registries can use different authentication methods. Leave a variable unset if its
registry doesn't require authentication. If you import from an export file with
--file, you don't need SRC_AUTH.
To import directly from the source registry, go to Import the schemas. The next two sections are optional.
Optional: Review an export before importing
Review the schemas before you change the target:
-
Export the schemas from the source to a file:
python3 sr_migrate.py export \--source SOURCE_REGISTRY_URL \--file export.jsonReplace
SOURCE_REGISTRY_URLwith the URL of the source registry. -
Confirm that
export.jsoncontains the subjects and versions you expect, including referenced schemas. -
Compare the schema IDs that your consumers read with the schema IDs in the export. An ID that isn't in the export can belong to a hard-deleted version. The script also prints a note when the source issued an ID that no exported version uses.
-
Continue with Import the schemas, and import from the file.
Optional: Export from a topic dump
If you can't reach the source registry API, create the export file from a dump of the
source _schemas topic.
Don't replicate the source _schemas topic directly to the target. Create an export
file and import it instead.
-
Dump the topic as tab-separated key-value pairs. Include all records from the start of the topic, including records with null values, which mark deleted schemas. For example, with
kcat:kcat -C -b SOURCE_KAFKA_SERVER -t _schemas -o beginning -e -q -Z \-f '%k\t%s\n' > schemas.logReplace
SOURCE_KAFKA_SERVERwith the bootstrap server of the Kafka cluster that hosts the source registry. Add the connection options that your source cluster needs, such as TLS or SASL settings. -
Create the export file from the dump:
python3 sr_migrate.py export-topic \--dump schemas.log \--file export.jsonIf the topic has no global compatibility record, the script uses
BACKWARD. To use a different level, add--global-config COMPATIBILITY_LEVEL. ReplaceCOMPATIBILITY_LEVELwith a level such asFULL. -
Continue with Import the schemas, and import from the file.
Import the schemas
-
Run the import. The following example exports from the source and imports into a new, empty target:
python3 sr_migrate.py import \--source SOURCE_REGISTRY_URL \--target SCHEMA_REGISTRY_URL \--scope global \--reserve-subject _migration_id_reservationReplace the following:
SOURCE_REGISTRY_URL: URL of the source registry.SCHEMA_REGISTRY_URL: Schema Registry URL of your Aiven service, which is the target.
With
--source, the command also writes the export toexport.jsonin the current directory. To use a different path, add--filewith that path. If you already created an export file, replace--source SOURCE_REGISTRY_URLwith--file export.json.The command uses the following options:
--scope:globalfor a new, empty target. Usesubjectif the target already has other subjects. The default issubject. Both scopes import every subject in the export. For more information, see How import mode works.--reserve-subject: the name of a dedicated subject that reserves the highest schema ID that the source registry issued. This prevents the target from reusing IDs that the source registry already issued. Use a name that isn't already in use. The script imports a placeholder schema into this subject and soft-deletes it. In the example,_migration_id_reservationstays on the target as a subject with one soft-deleted version.
The script also supports these options:
--force: allowsIMPORTmode on a target that has live schemas. If you don't use this option and the target has live schemas, the script shows a summary and asks whether to continue. If you entery, the script applies--force. For more information, see Import into a non-empty registry.--yes: answers yes to every confirmation prompt except the--forceprompt. The script never applies--forcewithout your confirmation.
-
Review the plan and confirm each step. If the Check target step reports conflicting IDs or versions, including soft-deleted ones, stop. Use a new target, or fix the conflicting subjects on the target before you run the import again.
-
When the script finishes, it prints a summary. Confirm that the Verify and Leave
IMPORTmode steps both reportdone:Verify done 226 versions carry their source idLeave IMPORT mode done nothing in scope is in IMPORTOther possible results are
skipped,FAILED,stopped, andnot run. If the migration stops before it finishes, see Troubleshoot migration errors.
Migrate with the API
Use this procedure to migrate a small number of schema versions with curl. You
register each version manually.
You can set IMPORT mode with one of the following scopes:
- Subject scope: Sets the mode for one subject at a time. This procedure uses this scope.
- Global scope: Sets the mode once for the whole registry. Use it only if the target registry is empty. See the tip in Import each subject.
This method has the following limits:
- It imports only live versions, so it doesn't recreate versions that are soft-deleted on the source.
- It doesn't reserve the highest schema ID that the source registry issued. Without the
reservation, the target can issue an ID that the source already used. If the source
issued IDs higher than the highest ID you import, use the script with
--reserve-subject.
The curl examples use HTTP Basic authentication with the -u option. For OAuth
2.0/OIDC, replace the -u option and its credentials in each request with
-H "Authorization: Bearer ACCESS_TOKEN". Replace ACCESS_TOKEN with the source
access token for requests to the source registry, and with the target access token for
requests to the target registry.
In the commands in this section, replace the following:
SOURCE_REGISTRY_URL: URL of the source registry.SOURCE_USERandSOURCE_PASSWORD: credentials for the source registry.SCHEMA_REGISTRY_URL: Schema Registry URL of your Aiven service, which is the target.SCHEMA_REGISTRY_USERandSCHEMA_REGISTRY_PASSWORD: credentials for the target registry, from the Schema Registry tab of Connect information in the Aiven Console.SUBJECT_NAME: name of the subject. URL-encode it when you use it in a request path.VERSION: version number of a schema.SCHEMA_ID: schema ID from the source.COMPATIBILITY_LEVEL: a compatibility level, such asFULL.
Retrieve the source schemas
-
List the subjects, and choose the ones to migrate:
curl -u SOURCE_USER:SOURCE_PASSWORD \SOURCE_REGISTRY_URL/subjects -
For each subject, list its versions:
curl -u SOURCE_USER:SOURCE_PASSWORD \SOURCE_REGISTRY_URL/subjects/SUBJECT_NAME/versions -
For each version, retrieve the schema and save the response:
curl -u SOURCE_USER:SOURCE_PASSWORD \SOURCE_REGISTRY_URL/subjects/SUBJECT_NAME/versions/VERSIONThe response includes the
id,version, andschemafields. It also includes theschemaTypeandreferencesfields when the schema has them. You need these fields for the import. Retrieve the versions that your schemas reference as well, even if they're in other subjects. -
For each subject, record any compatibility setting:
curl -u SOURCE_USER:SOURCE_PASSWORD \SOURCE_REGISTRY_URL/config/SUBJECT_NAMEIf the subject has no setting of its own, it uses the source's global compatibility level. To get that level, send a
GETrequest to/configon the source.
Check the target
Use an empty target subject when possible. Schema IDs are global, so look for ID conflicts even if the target subject is empty.
-
For each source ID that you saved from the source responses, look up the ID on the target:
curl -u SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD \SCHEMA_REGISTRY_URL/schemas/ids/SCHEMA_IDA
404 Not Foundstatus code means the ID isn't used on the target. If the ID exists and its schema type, definition, and references match the source, you can import it again without changes. If they differ, you can't import this schema with its original ID. -
If the target subject exists, list its versions, including soft-deleted ones, and compare them with the source versions:
curl -u SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD \"SCHEMA_REGISTRY_URL/subjects/SUBJECT_NAME/versions?deleted=true"Soft-deleted versions keep their ID and version number. A soft-deleted version conflicts only if its ID or content differs from the version you import. If they're identical, importing the version again restores it on the target.
-
If the target holds the same ID or version with different content, use a new target.
Import each subject
Import the subjects that other schemas reference first. The registry rejects a schema if the versions that it references aren't on the target yet. For each subject, do the following:
-
Set
IMPORTmode for the target subject:curl -u SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD \-X PUT \-H "Content-Type: application/vnd.schemaregistry.v1+json" \-d '{"mode":"IMPORT"}' \SCHEMA_REGISTRY_URL/mode/SUBJECT_NAMEIf the request fails with error code
40901and a message that starts withCannot import, the subject has live schemas. See Import into a non-empty registry.TipIf the target registry is empty, you can avoid setting the mode on each subject by setting
IMPORTmode for the whole registry:curl -u SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD \-X PUT \-H "Content-Type: application/vnd.schemaregistry.v1+json" \-d '{"mode":"IMPORT"}' \SCHEMA_REGISTRY_URL/modeAfter you set the registry mode, skip this step for every subject. When you finish importing all subjects, return the registry to
READWRITEmode, as described in the last step of this procedure. -
Import each version of the subject, starting with the versions that other schemas reference. For each version:
-
Create the
schema-import.jsonfile with the source values for the ID, version, schema definition, and any schema type and references. For example:{"schemaType": "AVRO","schema": "{\"type\":\"string\"}","id": 1001,"version": 5}Replace the example
schema,id, andversionvalues with the values from the source response. Theschemavalue is an escaped JSON string. Copy it from the source response without changing it. If the source response includesreferences, copy that field to the file. -
Register the version:
curl -u SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD \-X POST \-H "Content-Type: application/vnd.schemaregistry.v1+json" \--data @schema-import.json \SCHEMA_REGISTRY_URL/subjects/SUBJECT_NAME/versionsInclude both
idandversionto keep the source values. The registry accepts values from 1 to 2,147,483,647. Confirm that the ID in the response matches the source ID.
-
-
After you import all versions, set the target subject's compatibility level to the level that you recorded from the source:
curl -u SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD \-X PUT \-H "Content-Type: application/vnd.schemaregistry.v1+json" \-d '{"compatibility":"COMPATIBILITY_LEVEL"}' \SCHEMA_REGISTRY_URL/config/SUBJECT_NAMEIf the source subject used the source's global level, choose one of the following:
- Set that level on the target subject, using the preceding command. The effective level stays the same.
- Leave the target subject without a setting, so it inherits the target's global level. First confirm that the target's global level matches the source's.
-
Return the target to
READWRITEmode.For subject scope, delete each subject's mode setting, so that the subject follows the registry mode:
curl -u SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD \-X DELETE \SCHEMA_REGISTRY_URL/mode/SUBJECT_NAMEThe script does the same. To pin the subject to
READWRITEmode regardless of later changes to the registry mode, set an explicit mode instead:curl -u SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD \-X PUT \-H "Content-Type: application/vnd.schemaregistry.v1+json" \-d '{"mode":"READWRITE"}' \SCHEMA_REGISTRY_URL/mode/SUBJECT_NAMEFor global scope, after you import all subjects, set the registry mode to
READWRITE:curl -u SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD \-X PUT \-H "Content-Type: application/vnd.schemaregistry.v1+json" \-d '{"mode":"READWRITE"}' \SCHEMA_REGISTRY_URL/mode
Verify the migration
In the commands in this section, replace SCHEMA_REGISTRY_URL, SCHEMA_REGISTRY_USER,
SCHEMA_REGISTRY_PASSWORD, SUBJECT_NAME, and VERSION with the values that you used
in the migration.
Verify a script migration
The script's Verify step checks that every imported version has its source schema ID.
To confirm that the script reproduced the soft-delete status, compare the versions,
including soft-deleted ones, with the source. Add ?deleted=true to the versions
request:
curl -u SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD \
"SCHEMA_REGISTRY_URL/subjects/SUBJECT_NAME/versions?deleted=true"
Verify an API migration
-
Confirm that the target registry is in
READWRITEmode:curl -u SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD \SCHEMA_REGISTRY_URL/modeFor a subject migration, verify each subject:
curl -u SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD \SCHEMA_REGISTRY_URL/mode/SUBJECT_NAMEIf the subject has no mode setting of its own, the response shows the registry mode.
-
Get each target version and compare its ID, version, schema type, definition, and references with the saved source response:
curl -u SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD \SCHEMA_REGISTRY_URL/subjects/SUBJECT_NAME/versions/VERSION
Run final checks
After either method, do the following:
- Confirm that the target has every subject that you migrated, for example with
GET /subjects, and that each subject's compatibility setting matches the source, for example withGET /config/SUBJECT_NAME. - Unless you used the script with
--scope global, compare the target's global compatibility level with the source's. Send aGETrequest to/configon each registry. Subjects without their own setting use the global level. - Configure a test consumer to use the Schema Registry of your Aiven service. Confirm that it can deserialize existing messages, including messages that use schemas with references.
- If you used
--reserve-subject, confirm that the Reserve source max ID step reported the ID that it holds, or that it skipped because the source issued no higher ID.
Switch clients to the target
After you verify the migration, switch your clients to the target registry:
- Update your producers and consumers to use the Schema Registry URL and credentials of your Aiven service.
- Restart any producers that you stopped.
- Resume schema changes against the target registry.
Keep the source registry running until your clients work with the target. You can switch clients back to it only until they register new schemas on the target, because those schemas aren't on the source.
Schema migration doesn't copy Kafka topics or messages. If you migrate Kafka data separately, plan the cutover for both migrations together.
Handle special migration scenarios
Import into a non-empty registry
By default, IMPORT mode requires a registry or subject with no live schemas.
Import into an empty registry or subject when you can. Use force=true only after you
confirm that the IDs and versions you import don't conflict with schemas that are
already on the target.
To set IMPORT mode on a subject that already has live schemas, add force=true:
curl -u SCHEMA_REGISTRY_USER:SCHEMA_REGISTRY_PASSWORD \
-X PUT \
-H "Content-Type: application/vnd.schemaregistry.v1+json" \
-d '{"mode":"IMPORT"}' \
"SCHEMA_REGISTRY_URL/mode/SUBJECT_NAME?force=true"
Replace the following:
SCHEMA_REGISTRY_USERandSCHEMA_REGISTRY_PASSWORD: credentials for the target registry, from the Schema Registry tab of Connect information in the Aiven Console.SCHEMA_REGISTRY_URL: Schema Registry URL of your Aiven service, which is the target.SUBJECT_NAME: name of the subject. URL-encode it when you use it in a request path.
The force=true parameter also works with PUT /mode, which sets the mode for the whole
registry. In the script, use the --force option. The parameter and the option skip only
the requirement for no live schemas. Schema Registry still rejects the request if an ID
or version is already bound to different content.
Troubleshoot migration errors
If an import stops, the script marks the remaining steps not run and prints a hint.
Fix the cause and run the same import command again. The script repeats the versions
that are already imported without changing them. If you use the API, list the versions
that the registry already imported, register the remaining ones, and set the subject or
registry back to READWRITE mode.
If the import stops partway, the target can stay in IMPORT mode. Keep ordinary
registrations paused until the import is complete. The script summary shows which
subjects are still in IMPORT mode. With the API, get the registry mode with
GET SCHEMA_REGISTRY_URL/mode and each subject's mode with
GET SCHEMA_REGISTRY_URL/mode/SUBJECT_NAME.
Common migration errors
The registry returns an error_code field in the response body, in addition to the HTTP
status code. The following table lists common problems:
| Error | What to do |
|---|---|
HTTP 401 Unauthorized or 403 Forbidden status code | Authentication failed, or the user lacks permission. If you don't use avnadmin, confirm that the user has the schema_registry_write ACL entries for Config: and the subjects that you migrate. For more information, see Schema Registry authorization. If you use OIDC, verify the roles. |
Error code 42205 with not allowed | The target doesn't allow mode changes because mode_mutability is false. See Settings that affect the migration. Contact the Aiven support team and include the service name and the full error response. |
Error code 42205 during registration | The subject isn't in IMPORT mode, for example because the mode changed during the run. Set IMPORT mode again and resume. |
Error code 40901 with a message that starts with Cannot import | The target has live schemas in the scope you chose. Use --force only after you confirm that the IDs and versions don't conflict. See Import into a non-empty registry. |
Error code 40901 with any other message | The target already has this ID or version with different content. --force doesn't help. Resolve the conflict that the Check target step reports, or use a new target. |
Error code 42207 with already registered | The target doesn't allow the same schema content under more than one ID because allow_duplicate_schema_ids is false. See Settings that affect the migration. Contact the Aiven support team and include the service name and the full error response. |
Error code 42202 or 42207 with any other message | The ID or version is outside the range 1 to 2,147,483,647. You can't import this schema with its original ID. |
HTTP 422 Unprocessable Entity status code | The registry rejected the schema. If it has references, confirm that the referenced schemas are in the export, and import them first. |
Settings that affect the migration
Two Karapace settings change how the target handles a migration. If one of them blocks your migration, contact the Aiven support team.
mode_mutability: Controls whether the registry accepts mode changes. The default istrue. If it'sfalse, the registry rejects every request that sets or deletes a mode, including settingIMPORTmode, with error code42205. A subject that's already inIMPORTmode can't return toREADWRITEmode.allow_duplicate_schema_ids: Controls whether the registry accepts the same schema content under more than one ID. The default istrue. If it'sfalse, the registry rejects a schema that's already registered under a different ID with error code42207. Importing a schema again under the ID it already has still works.
For more information, see Operator flags in the Karapace documentation.
Related pages