How hybrid search works
Hybrid search is the next step after pure semantic search: the query goes to two indexes in parallel. The lexical branch looks up the query words in an inverted index and ranks with BM25 — a formula that weighs how often a word appears in a product record, how rare it is across the catalogue and how long the text is. The vector branch turns the query into an embedding and finds the nearest product vectors in a vector database using approximate nearest neighbor search. The two lists are then merged, and the top of the final list is often refined by reranking.
Which method handles which query
Site search receives both kinds of queries, and a hybrid covers the weak spots of each method:
| Query | Lexical (BM25) | Vector |
|---|---|---|
| “RTX 4070 Ti” | finds it precisely | may mix in the 4060 and 4080 |
| SKU “AB-10234” | finds it precisely | almost useless |
| “summer wedding guest dress” | looks for the words “wedding” and “summer” | finds light, dressy dresses |
| “gift for a dad who loves fishing” | almost useless | finds fishing gear |
| “winter footwear” | only items with the word “winter” | also finds “insulated boots” |
How the result lists are merged
- Reciprocal Rank Fusion. RRF(d) = Σ 1/(k + r(d)), where r(d) is the product’s position in each list. The method was described by Gordon Cormack, Charles Clarke and Stefan Büttcher at SIGIR 2009; the authors fixed k = 60 in a pilot experiment. In Elasticsearch it is the rrf retriever, where k (rank_constant) also defaults to 60.
- Weighted sum of scores. The scores of both branches are brought to one scale and added with a weight. In Weaviate the weight is alpha: 0 is pure BM25, 1 is pure vector search; since version 1.24 the default is Relative Score Fusion, with Ranked Fusion by position as the alternative. OpenSearch offers a hybrid query and a score-normalisation processor for the same purpose.
- Reranking. A few dozen top candidates from both branches are reordered by a model that sees the query and the product together and takes business signals into account.
RRF example: product A is first in BM25 and absent from the vector list, so it scores 1/61 ≈ 0.0164. Product B is fifth in both lists and scores 2/65 ≈ 0.0308. B ranks higher: being present in both lists weighs more than first place in one.
In practice: how to roll it out
- Start with RRF — it has no weights to tune.
- Build a labelled set of 200–300 real queries in three classes — SKUs and model numbers, brands, descriptive queries — and measure quality for each class separately.
- If a query looks like a SKU (digits, hyphens), boost the lexical branch.
- Apply stock, category and price filters in both branches before merging; otherwise one of them may have almost nothing left after filtering.
- Compare the result with your current search in an A/B test on real traffic.