Topic 8.2
Schema Registry: Subjects, IDs, and Serializers
In one line
A schema registry stores versioned schemas under subjects (usually <topic>-value), assigns each schema a global ID, and enforces compatibility on registration. Serializers put the schema ID in each record (a magic byte plus 4-byte ID), and deserializers fetch and cache the schema to decode.
Think of it like this
A library of official form templates, each with a number. Instead of stapling the whole template to every filled form, you write the template number at the top; anyone can look up the template to read the form.
Key ideas
- 01
Registries: Confluent Schema Registry (the common standard), Apicurio Registry, AWS Glue Schema Registry, and others with compatible APIs.
- 02
Subject naming strategies:
TopicNameStrategy(orders-value, one schema per topic),RecordNameStrategy(by record type, allowing several event types in one topic),TopicRecordNameStrategy(both). - 03
Wire format (Confluent): byte 0 = magic byte 0, bytes 1–4 = schema ID, then the Avro/Protobuf/JSON payload. Consumers cache schemas by ID, so the registry is only hit for new IDs.
- 04
Registration: in production, register schemas from CI (not auto-registration by producers at runtime,
auto.register.schemas=false), so incompatible changes fail the build instead of production. - 05
Availability: the registry is a dependency for producers (first use of a schema) and consumers (first sight of an ID). Run it highly available; clients cache aggressively.
Code & diagrams
# Register a schema (normally done by CI with the Maven/Gradle plugin)
curl -s -X POST -H "Content-Type: application/vnd.schemaregistry.v1+json" \
--data '{"schemaType":"AVRO","schema":"{\"type\":\"record\",\"name\":\"UserCreated\",\"fields\":[{\"name\":\"id\",\"type\":\"long\"},{\"name\":\"name\",\"type\":\"string\"}]}"}' \
http://registry:8081/subjects/users-value/versions
{"id":17}
curl -s http://registry:8081/config/users-value
{"compatibilityLevel":"BACKWARD"}
# Wire format of a record value: 00 | 00 00 00 11 | <avro bytes>
# magic schema id 17value.serializer=io.confluent.kafka.serializers.KafkaAvroSerializer
schema.registry.url=http://registry:8081
auto.register.schemas=false
use.latest.version=true
# consumer
value.deserializer=io.confluent.kafka.serializers.KafkaAvroDeserializer
specific.avro.reader=trueWhen it breaks
Producers auto-register schemas at runtime
What you see
A developer's local change registers a new version in production's registry, or an incompatible change fails at runtime instead of in CI.
Fix & prevent
Register schemas from CI with compatibility checks; set auto.register.schemas=false in production.
Explain it without notes
How does a consumer know which schema to use for each record?
Practice
Run Schema Registry in your lab, register a v1 schema, and produce and consume Avro records with kcat or the Avro console tools.
Trade-offs
- ↔
A registry adds a component to operate but turns breaking changes into CI failures.
Done when you can
I can explain subjects, schema IDs, the wire format and CI-based registration.