Topic 8.3
Serialization: Size, Speed, Compatibility, Security
In one line
How you encode values decides memory, CPU, cross-language readability, safety of rolling deploys, and security. Prefer explicit, schema-aware formats (JSON for readability, MessagePack or Protobuf for size), version your keys, compress large values, and never deserialize untrusted native object formats.
Think of it like this
Shipping goods abroad. You can send items in a language only your own warehouse understands (Java serialization), in a universal format everyone can read but that's bulky (JSON), or vacuum-packed with a standard label (Protobuf with compression). What matters is who has to unpack it, and whether they'll still understand the label next year.
Key ideas
- 01
Options: plain strings and numbers (use native types where possible: counters as integers, hashes for flat objects), JSON (readable, universal, larger, slower), MessagePack (binary JSON-like, smaller, schemaless), Protobuf/Avro (schema-based, compact, strong evolution rules), and language-native formats (Java serialization, Python pickle), which you should avoid.
- 02
Security: deserializing Java-serialized or pickled data can execute code if an attacker can write values into Redis (or if Redis is compromised). Jackson with default typing has similar gadget risks. Use formats that produce plain data, and restrict who can write to Redis.
- 03
Schema evolution: rolling deploys mean old and new code read each other's values. Add fields as optional, ignore unknown fields, never repurpose a field, and change the key version when a change is incompatible.
- 04
Size and speed: for a 2 KB product JSON, MessagePack might be ~1.4 KB and Protobuf ~0.9 KB; compression (LZ4, zstd, Snappy) often gives 3–5× on JSON. Compress only above a threshold (~1 KB), and mark compressed values with a prefix byte so readers know.
- 05
Redis-native modelling is often the best "serialization": a hash with fields, a sorted set, or JSON type lets you update parts atomically and avoids read-modify-write of blobs.
Code & diagrams
A versioned, compressed codec with a one-byte header, readable by any language that implements it.
import json, zlib
V1_JSON, V1_JSON_Z = b"\x01", b"\x02"
THRESHOLD = 1024
def encode(obj: dict) -> bytes:
raw = json.dumps(obj, separators=(",", ":")).encode()
if len(raw) >= THRESHOLD:
return V1_JSON_Z + zlib.compress(raw, 6)
return V1_JSON + raw
def decode(data: bytes) -> dict:
tag, body = data[:1], data[1:]
if tag == V1_JSON:
return json.loads(body)
if tag == V1_JSON_Z:
return json.loads(zlib.decompress(body))
raise ValueError(f"unknown codec tag {tag!r}")Product document (same data) Bytes Encode+decode (relative)
JSON (pretty) 2,480 1.0x
JSON (compact) 2,050 0.9x
MessagePack 1,430 0.5x
Protobuf 910 0.3x
Compact JSON + zstd 640 1.2x
Java serialization 3,900 1.4x (and unsafe, Java-only)
(Illustrative; measure with your own data.)Interview problem
The problem
Choose a serialization strategy for a shared cache
A product cache is written by a Java service and read by Java, Node and Python services. It holds 30M entries averaging 3 KB of JSON. Memory costs are high and a recent deploy broke readers for 20 minutes. Propose a serialization strategy.
The interviewer follows up
Does compression hurt latency?
Explain it without notes
Why is Java native serialization a bad fit for Redis caches?
Practice
Measure one of your cached objects encoded as JSON, MessagePack and compressed JSON with MEMORY USAGE.
Trade-offs
- ↔
Readable formats help debugging; compact formats save memory and bandwidth. Schema formats add tooling but make evolution safer.
Done when you can
I pick serialization formats based on readers, size, evolution and security.
I version key namespaces and add a codec header for safe migrations.