Topic 11.5
Elasticsearch: Inverted Indexes, Relevance and Reindexing
In one line
Elasticsearch (and OpenSearch) stores JSON documents in shards of Lucene segments with inverted indexes: analyzers turn text into terms, BM25 scores relevance, and filters narrow results cheaply. It's near-real-time (documents become searchable after a refresh, 1 s by default) and should be fed from the database as a derived index that can be rebuilt with an alias switch.
Think of it like this
The index at the back of a book. Each word points to the pages where it appears, and pages mentioning the word many times, in a short chapter, rank higher.
Key ideas
- 01
Mapping defines field types:
text(analyzed for full-text) vskeyword(exact, for filters, sorting, aggregations). Use multi-fields (nameas text plusname.rawas keyword). Dynamic mapping in production causes mapping explosions; define mappings explicitly. - 02
Analyzer = char filters + tokenizer + token filters (lowercase, stemming, synonyms, edge n-grams for autocomplete). The same analyzer runs at query time, so "Running Shoes" matches "running shoe".
- 03
Queries:
boolwithmust(scored),filter(not scored, cached, fast),should,must_not. Relevance: BM25 (term frequency with saturation, inverse document frequency, field length). Fuzzy matching (fuzziness: AUTO) tolerates typos. - 04
Internals: an index has primary shards (fixed at creation) and replicas; each shard is a Lucene index of immutable segments. Refresh makes new segments searchable; merges combine segments; the translog provides durability between flushes.
- 05
Reindexing: mappings of existing fields can't change in place. Create
products_v2with the new mapping, backfill from the source of truth (or_reindex), catch up from the change stream, then atomically move theproductsalias.
Code & diagrams
PUT products_v2
{
"settings": {
"number_of_shards": 3, "number_of_replicas": 1,
"analysis": {
"analyzer": {
"autocomplete": { "tokenizer": "standard", "filter": ["lowercase", "edge_3_15"] }
},
"filter": { "edge_3_15": { "type": "edge_ngram", "min_gram": 3, "max_gram": 15 } }
}
},
"mappings": {
"dynamic": "strict",
"properties": {
"name": { "type": "text", "fields": { "raw": { "type": "keyword" },
"ac": { "type": "text", "analyzer": "autocomplete", "search_analyzer": "standard" } } },
"brand": { "type": "keyword" },
"price": { "type": "scaled_float", "scaling_factor": 100 },
"in_stock": { "type": "boolean" },
"updated_at": { "type": "date" }
}
}
}
GET products/_search
{
"query": { "bool": {
"must": [ { "match": { "name": { "query": "runing shoes", "fuzziness": "AUTO" } } } ],
"filter": [ { "term": { "in_stock": true } }, { "range": { "price": { "lte": 5000 } } } ]
} }
}
POST _aliases
{ "actions": [ { "remove": { "index": "products_v1", "alias": "products" } },
{ "add": { "index": "products_v2", "alias": "products" } } ] }Interview problem
The problem
Product search for 50M products
Design product search with typo tolerance, autocomplete, facets (brand, price range), in-stock filtering and freshness within seconds of catalogue changes, with PostgreSQL as the source of truth.
When it breaks
Elasticsearch used as the primary store
What you see
A mapping mistake or cluster loss means data is gone; there's no way to rebuild documents that exist only in the index.
Fix & prevent
Keep the source of truth in a database; treat indexes as rebuildable projections with a tested reindex path.
Oversharding
What you see
Thousands of tiny shards consume heap and slow cluster state updates; the cluster turns yellow or red under load.
Fix & prevent
Aim for shards of ~10–50 GB, use ILM rollover for time-based indexes, and shrink or merge small indexes.
Explain it without notes
What's the difference between text and keyword fields?
Practice
Why put price and stock conditions in filter rather than must?
Trade-offs
- ↔
Rich relevance and fast text search as an eventually consistent projection, with the operational cost of another cluster and a sync pipeline.
Done when you can
I can design mappings and analyzers, write bool queries, and plan CDC-fed indexing and zero-downtime reindexing.