Reciprocal Rank Fusion (RRF) for hybrid search: the formula and a full worked example
A hybrid RAG pipeline runs two searches against the same query — a BM25 keyword search and a dense vector search — and gets back two ranked lists on completely incompatible scales: BM25 scores are unbounded, cosine similarity is bounded between -1 and 1. Reciprocal Rank Fusion (RRF) sidesteps that mismatch by throwing the scores away entirely and fusing on rank position alone. Here's the exact formula, where its k=60 default actually comes from, and a full numeric example you can check by hand.
The formula, and where k = 60 comes from
RRF was introduced by Cormack, Clarke and Büttcher in a two-page SIGIR 2009 paper. For a document d and a set of ranked lists R (one per retrieval method), the score is:
RRFscore(d) = Σ 1 / (k + rank(d)), summed over every list d appears in
A document missing from a list simply contributes nothing from that list. The paper fixed k = 60 "during a pilot investigation and not altered during subsequent validation," explaining that "the constant k mitigates the impact of high rankings by outlier systems" — without it, a single method ranking a document #1 could dominate the fused score no matter how the other methods ranked it. Their own pilot table shows exactly how little the choice matters: mean average precision moved from 0.2139 at k=30 to 0.2144 at k=50 to 0.2145 at k=60 to 0.2142 at k=100 — a near-flat curve, with 60 simply the best of a wide, forgiving plateau. Elasticsearch's RRF retriever ships the identical default (rank_constant: 60) today, so this isn't just an academic footnote — it's the number you'll meet in production.
Worked example: fusing a BM25 list and a vector list
Say the query is "database backup schedule" and each method returns its top 5 matches from a shared document pool. The two lists only partly overlap:
| Rank | BM25 keyword search | Vector similarity search |
|---|---|---|
| 1 | Doc A | Doc B |
| 2 | Doc B | Doc D |
| 3 | Doc C | Doc F |
| 4 | Doc D | Doc A |
| 5 | Doc E | Doc G |
Apply 1 / (60 + rank) to every appearance and sum per document. Doc A: rank 1 in BM25 gives 1/61 = 0.01639, plus rank 4 in vector gives 1/64 = 0.01563, for a total of 0.03202. Doc B: rank 2 in BM25 gives 1/62 = 0.01613, plus rank 1 in vector gives 1/61 = 0.01639, for a total of 0.03252. Doing the same for every document:
| Doc | BM25 term | Vector term | RRF score |
|---|---|---|---|
| B | 1/62 = 0.01613 | 1/61 = 0.01639 | 0.03252 |
| A | 1/61 = 0.01639 | 1/64 = 0.01563 | 0.03202 |
| D | 1/64 = 0.01563 | 1/62 = 0.01613 | 0.03175 |
| C | 1/63 = 0.01587 | — (not retrieved) | 0.01587 |
| F | — (not retrieved) | 1/63 = 0.01587 | 0.01587 |
| E | 1/65 = 0.01538 | — (not retrieved) | 0.01538 |
| G | — (not retrieved) | 1/65 = 0.01538 | 0.01538 |
The fused ranking is B, A, D, then C and F tied, then E and G tied. Notice what happened to Doc A: it was the single best keyword match, ranked #1 by BM25, yet it finishes behind Doc B, which was never better than #2 on either list. Showing up solidly on both lists beat being the top result on just one — that's the entire point of fusing two independent signals instead of trusting either one alone.
The alternative: weighted score fusion with alpha
RRF isn't the only way to combine two ranked lists. The other common approach normalizes each method's raw scores into a shared 0-to-1 range (typically min-max over the candidate set) and combines them as a weighted sum: score = (1 - alpha) × bm25_norm + alpha × vector_norm. Weaviate's hybrid search uses exactly this convention, with alpha defaulting to 0.5 — equal weight to both searches, adjustable toward pure keyword (alpha = 0) or pure vector (alpha = 1) per query.
Reuse the same seven documents, now with raw scores attached — BM25's unbounded term-frequency score and the vector search's cosine similarity:
| Doc | BM25 raw score | Vector raw score |
|---|---|---|
| A | 18.2 | 0.81 |
| B | 15.4 | 0.91 |
| C | 9.7 | — |
| D | 7.1 | 0.88 |
| E | 5.3 | — |
| F | — | 0.85 |
| G | — | 0.78 |
Min-max normalizing each column (BM25 ranges 5.3-18.2, a span of 12.9; vector ranges 0.78-0.91, a span of 0.13) and averaging at alpha = 0.5 gives: B = 0.5×0.783 + 0.5×1.000 = 0.891; A = 0.5×1.000 + 0.5×0.231 = 0.615; D = 0.5×0.140 + 0.5×0.769 = 0.454; F = 0.5×0 + 0.5×0.538 = 0.269; C = 0.5×0.341 + 0.5×0 = 0.171; E and G both = 0.
The ranking is still B, A, D first, matching RRF — but now F clearly outranks C, whereas RRF scored them identically. The reason: RRF only ever sees "rank 3 in one list," and a rank-3 finish is a rank-3 finish regardless of how strong the underlying match was. Weighted fusion keeps the magnitude — F's 0.85 cosine similarity sits much closer to the field's top score than C's 9.7 term-frequency score does to its field's top — so it breaks the tie in F's favor.
RRF vs weighted fusion: the trade-off
| Reciprocal Rank Fusion | Weighted score fusion | |
|---|---|---|
| Needs score normalization | No — works on rank position only | Yes — min-max or similar, recomputed whenever the candidate set changes |
| Tunable per query | Barely — only k, rarely changed from 60 | Yes — alpha shifts weight toward keyword or vector search |
| Sensitive to an outlier raw score | No, by construction | Yes — one unusually high or low score shifts every other document's normalized value in that batch |
| Uses how strong a match was, not just where it ranked | No | Yes |
Both defaults above come from production systems, not just the original paper, so they're worth knowing by heart: Elasticsearch's RRF retriever defaults rank_constant to 60, and Weaviate's hybrid search defaults alpha to 0.5. For more on how BM25's own scoring works, why vocabulary mismatch breaks a single retrieval method, and other hybrid-retrieval mechanics, practice against the RAG & embeddings quiz.
Source: Cormack, Clarke & Büttcher, "Reciprocal Rank Fusion outperforms Condorcet and Individual Rank Learning Methods," SIGIR 2009 (formula, k=60 derivation, pilot MAP-vs-k table); Elastic, Elasticsearch RRF retriever reference documentation (rank_constant default); Weaviate, hybrid search documentation (alpha default).