Topic 4.6
Vector Search and Semantic Caching for AI Features
In one line
Redis can store embedding vectors and find the nearest ones with HNSW or exact (FLAT) indexes, which powers RAG retrieval, recommendations and semantic caching of LLM responses. Redis 8 also adds vector sets (VADD, VSIM), a native data type for similarity search.
Think of it like this
A librarian who understands meaning, not spelling. Ask for "books about fixing cars" and they bring you "Auto Repair Basics" even though no word matches. Embeddings turn text into points in space where similar meanings sit close together, and vector search finds the closest points.
Key ideas
- 01
An embedding model turns text (or images) into a vector of floats, for example 768 or 1,536 dimensions. Similar content has small cosine distance. Storing 1M vectors of 1,536 float32 values takes ~6 GB for the raw vectors alone, plus index overhead, so dimensions, precision and quantisation matter for cost.
- 02
Query Engine vectors: declare a field
VECTOR HNSW 6 TYPE FLOAT32 DIM 1536 DISTANCE_METRIC COSINE(orFLATfor exact brute force on small sets), store vectors as binary blobs in hashes or JSON arrays, and query with KNN:FT.SEARCH idx "*=>[KNN 5 @embedding $vec AS dist]" PARAMS 2 vec <blob> SORTBY dist DIALECT 2. Combine with filters (@tenant:{acme}) for hybrid queries. - 03
HNSW (hierarchical navigable small world graphs) gives approximate nearest neighbours in roughly logarithmic time with tunable accuracy (
M,EF_CONSTRUCTION,EF_RUNTIME); FLAT is exact but O(N) per query. Approximate is almost always fine for retrieval. - 04
Vector sets (Redis 8, introduced as beta):
VADD key VALUES 3 0.1 0.2 0.3 itemadds an element;VSIM key ELE item COUNT 5 WITHSCORESorVSIM key VALUES ...finds similar ones, with optional attribute filters. They default to 8-bit quantisation to save memory. They're simpler than a full index when you just need similarity over one set. - 05
Semantic caching: embed each incoming LLM prompt; if a previously answered prompt is within a distance threshold (for example cosine distance < 0.1), return the stored answer instead of calling the model. It cuts cost and latency for repetitive questions, but a threshold that's too loose returns wrong answers for subtly different questions. Scope caches per tenant and per context, add TTLs, and never share cached answers containing one user's private data.
- 06
Libraries like RedisVL (Python) wrap these patterns (vector indexes, semantic cache, semantic routing). On Valkey, vector search comes from the valkey-search module; managed services differ in support, so check before designing around it.
Code & diagrams
127.0.0.1:6379> FT.CREATE idx:docs ON HASH PREFIX 1 doc: SCHEMA tenant TAG content TEXT embedding VECTOR HNSW 6 TYPE FLOAT32 DIM 1536 DISTANCE_METRIC COSINE
OK
# HSET doc:1 tenant acme content "How to reset a password" embedding <6144-byte float32 blob>
127.0.0.1:6379> FT.SEARCH idx:docs "(@tenant:{acme})=>[KNN 3 @embedding $vec AS dist]" PARAMS 2 vec <blob> SORTBY dist RETURN 2 content dist DIALECT 2
1) (integer) 3
2) "doc:1"
3) 1) "content" 2) "How to reset a password" 3) "dist" 4) "0.0712"
...
# Vector sets (Redis 8)
127.0.0.1:6379> VADD movies VALUES 3 0.9 0.1 0.2 "inception"
(integer) 1
127.0.0.1:6379> VADD movies VALUES 3 0.85 0.15 0.25 "interstellar"
(integer) 1
127.0.0.1:6379> VSIM movies ELE "inception" COUNT 2 WITHSCORES
1) "inception"
2) "1"
3) "interstellar"
4) "0.9968"A minimal semantic cache with redis-py. `embed()` and `call_llm()` stand in for your model client.
import numpy as np, redis, hashlib
from redis.commands.search.query import Query
r = redis.Redis()
THRESHOLD = 0.10 # cosine distance: lower = more similar
def cached_answer(tenant: str, prompt: str) -> str:
vec = np.asarray(embed(prompt), dtype=np.float32).tobytes()
q = (Query(f"(@tenant:{{{tenant}}})=>[KNN 1 @embedding $vec AS dist]")
.sort_by("dist").return_fields("answer", "dist").dialect(2))
res = r.ft("idx:semcache").search(q, query_params={"vec": vec})
if res.docs and float(res.docs[0].dist) < THRESHOLD:
return res.docs[0].answer # cache hit
answer = call_llm(prompt)
key = "semcache:" + hashlib.sha1((tenant + prompt).encode()).hexdigest()
r.hset(key, mapping={"tenant": tenant, "prompt": prompt,
"answer": answer, "embedding": vec})
r.expire(key, 86400) # answers go stale
return answerInterview problem
The problem
Semantic cache for a support chatbot
A support chatbot answers 2M questions a day and LLM costs are high. Many questions are paraphrases of each other. Design a semantic cache with Redis, including the threshold, multi-tenancy, freshness, and how you'd measure whether it's working.
You're given
- 2M questions/day
- ~40% are near-duplicates
- Multi-tenant (per customer knowledge base)
- Answers change when docs change
The interviewer follows up
Why use Redis instead of a dedicated vector database?
When it breaks
A semantic cache shared across tenants
What you see
A question from customer B gets customer A's answer, possibly exposing A's internal information: a data leak.
Fix & prevent
Filter every lookup by tenant (TAG field) and include tenant in keys; test isolation explicitly.
Embedding model upgraded without re-indexing
What you see
New query vectors are compared with old-model vectors; distances become meaningless and hit quality collapses.
Fix & prevent
Store the model name/version with each index; build a new index for the new model and switch over.
Explain it without notes
Explain HNSW vs FLAT vector indexes.
Practice
Estimate memory for 5M vectors at 768 dims in float32 vs 8-bit quantised.
Trade-offs
- ↔
Looser similarity thresholds raise hit rate and savings, and also the rate of wrong answers.
- ↔
In-memory vector search is fast but RAM-priced; disk-based vector stores are cheaper for very large corpora.
Done when you can
I can create a vector index and run filtered KNN queries.
I can design a tenant-safe semantic cache with thresholds, TTLs and quality checks.
I know what vector sets are and when they're enough.