Command Palette

Search for a command to run...

Hectal
PHASE 11Intermediate ~9 min· topic 5 of 7

Topic 11.5

Elasticsearch: Inverted Indexes, Relevance and Reindexing

In one line

Elasticsearch (and OpenSearch) stores JSON documents in shards of Lucene segments with inverted indexes: analyzers turn text into terms, BM25 scores relevance, and filters narrow results cheaply. It's near-real-time (documents become searchable after a refresh, 1 s by default) and should be fed from the database as a derived index that can be rebuilt with an alias switch.

0/7 · 0%

Think of it like this

The index at the back of a book. Each word points to the pages where it appears, and pages mentioning the word many times, in a short chapter, rank higher.

Key ideas

  1. 01

    Mapping defines field types: text (analyzed for full-text) vs keyword (exact, for filters, sorting, aggregations). Use multi-fields (name as text plus name.raw as keyword). Dynamic mapping in production causes mapping explosions; define mappings explicitly.

  2. 02

    Analyzer = char filters + tokenizer + token filters (lowercase, stemming, synonyms, edge n-grams for autocomplete). The same analyzer runs at query time, so "Running Shoes" matches "running shoe".

  3. 03

    Queries: bool with must (scored), filter (not scored, cached, fast), should, must_not. Relevance: BM25 (term frequency with saturation, inverse document frequency, field length). Fuzzy matching (fuzziness: AUTO) tolerates typos.

  4. 04

    Internals: an index has primary shards (fixed at creation) and replicas; each shard is a Lucene index of immutable segments. Refresh makes new segments searchable; merges combine segments; the translog provides durability between flushes.

  5. 05

    Reindexing: mappings of existing fields can't change in place. Create products_v2 with the new mapping, backfill from the source of truth (or _reindex), catch up from the change stream, then atomically move the products alias.

Code & diagrams

products-index.jsonjson
PUT products_v2
{
  "settings": {
    "number_of_shards": 3, "number_of_replicas": 1,
    "analysis": {
      "analyzer": {
        "autocomplete": { "tokenizer": "standard", "filter": ["lowercase", "edge_3_15"] }
      },
      "filter": { "edge_3_15": { "type": "edge_ngram", "min_gram": 3, "max_gram": 15 } }
    }
  },
  "mappings": {
    "dynamic": "strict",
    "properties": {
      "name":     { "type": "text", "fields": { "raw": { "type": "keyword" },
                     "ac": { "type": "text", "analyzer": "autocomplete", "search_analyzer": "standard" } } },
      "brand":    { "type": "keyword" },
      "price":    { "type": "scaled_float", "scaling_factor": 100 },
      "in_stock": { "type": "boolean" },
      "updated_at": { "type": "date" }
    }
  }
}

GET products/_search
{
  "query": { "bool": {
    "must":   [ { "match": { "name": { "query": "runing shoes", "fuzziness": "AUTO" } } } ],
    "filter": [ { "term": { "in_stock": true } }, { "range": { "price": { "lte": 5000 } } } ]
  } }
}

POST _aliases
{ "actions": [ { "remove": { "index": "products_v1", "alias": "products" } },
               { "add":    { "index": "products_v2", "alias": "products" } } ] }

Interview problem

The problem

Product search for 50M products

Design product search with typo tolerance, autocomplete, facets (brand, price range), in-stock filtering and freshness within seconds of catalogue changes, with PostgreSQL as the source of truth.

When it breaks

Elasticsearch used as the primary store

What you see

A mapping mistake or cluster loss means data is gone; there's no way to rebuild documents that exist only in the index.

Fix & prevent

Keep the source of truth in a database; treat indexes as rebuildable projections with a tested reindex path.

Oversharding

What you see

Thousands of tiny shards consume heap and slow cluster state updates; the cluster turns yellow or red under load.

Fix & prevent

Aim for shards of ~10–50 GB, use ILM rollover for time-based indexes, and shrink or merge small indexes.

Explain it without notes

01

What's the difference between text and keyword fields?

Practice

01

Why put price and stock conditions in filter rather than must?

Trade-offs

  • ↔

    Rich relevance and fast text search as an eventually consistent projection, with the operational cost of another cluster and a sync pipeline.

Done when you can

  • I can design mappings and analyzers, write bool queries, and plan CDC-fed indexing and zero-downtime reindexing.