Command Palette

Search for a command to run...

Hectal
PHASE 8Intermediate ~7 min· topic 2 of 4

Topic 8.2

Schema Registry: Subjects, IDs, and Serializers

In one line

A schema registry stores versioned schemas under subjects (usually <topic>-value), assigns each schema a global ID, and enforces compatibility on registration. Serializers put the schema ID in each record (a magic byte plus 4-byte ID), and deserializers fetch and cache the schema to decode.

0/4 · 0%

Think of it like this

A library of official form templates, each with a number. Instead of stapling the whole template to every filled form, you write the template number at the top; anyone can look up the template to read the form.

Key ideas

  1. 01

    Registries: Confluent Schema Registry (the common standard), Apicurio Registry, AWS Glue Schema Registry, and others with compatible APIs.

  2. 02

    Subject naming strategies: TopicNameStrategy (orders-value, one schema per topic), RecordNameStrategy (by record type, allowing several event types in one topic), TopicRecordNameStrategy (both).

  3. 03

    Wire format (Confluent): byte 0 = magic byte 0, bytes 1–4 = schema ID, then the Avro/Protobuf/JSON payload. Consumers cache schemas by ID, so the registry is only hit for new IDs.

  4. 04

    Registration: in production, register schemas from CI (not auto-registration by producers at runtime, auto.register.schemas=false), so incompatible changes fail the build instead of production.

  5. 05

    Availability: the registry is a dependency for producers (first use of a schema) and consumers (first sight of an ID). Run it highly available; clients cache aggressively.

Code & diagrams

registry.shbash
# Register a schema (normally done by CI with the Maven/Gradle plugin)
curl -s -X POST -H "Content-Type: application/vnd.schemaregistry.v1+json" \
  --data '{"schemaType":"AVRO","schema":"{\"type\":\"record\",\"name\":\"UserCreated\",\"fields\":[{\"name\":\"id\",\"type\":\"long\"},{\"name\":\"name\",\"type\":\"string\"}]}"}' \
  http://registry:8081/subjects/users-value/versions
{"id":17}

curl -s http://registry:8081/config/users-value
{"compatibilityLevel":"BACKWARD"}

# Wire format of a record value:  00 | 00 00 00 11 | <avro bytes>
#                                magic   schema id 17
avro-producer.propertiesproperties
value.serializer=io.confluent.kafka.serializers.KafkaAvroSerializer
schema.registry.url=http://registry:8081
auto.register.schemas=false
use.latest.version=true
# consumer
value.deserializer=io.confluent.kafka.serializers.KafkaAvroDeserializer
specific.avro.reader=true

When it breaks

Producers auto-register schemas at runtime

What you see

A developer's local change registers a new version in production's registry, or an incompatible change fails at runtime instead of in CI.

Fix & prevent

Register schemas from CI with compatibility checks; set auto.register.schemas=false in production.

Explain it without notes

01

How does a consumer know which schema to use for each record?

Practice

01

Run Schema Registry in your lab, register a v1 schema, and produce and consume Avro records with kcat or the Avro console tools.

Trade-offs

  • ↔

    A registry adds a component to operate but turns breaking changes into CI failures.

Done when you can

  • I can explain subjects, schema IDs, the wire format and CI-based registration.