passdrill

Reciprocal Rank Fusion (RRF) for hybrid search: the formula and a full worked example

A hybrid RAG pipeline runs two searches against the same query — a BM25 keyword search and a dense vector search — and gets back two ranked lists on completely incompatible scales: BM25 scores are unbounded, cosine similarity is bounded between -1 and 1. Reciprocal Rank Fusion (RRF) sidesteps that mismatch by throwing the scores away entirely and fusing on rank position alone. Here's the exact formula, where its k=60 default actually comes from, and a full numeric example you can check by hand.

The formula, and where k = 60 comes from

RRF was introduced by Cormack, Clarke and Büttcher in a two-page SIGIR 2009 paper. For a document d and a set of ranked lists R (one per retrieval method), the score is:

RRFscore(d) = Σ 1 / (k + rank(d)), summed over every list d appears in

A document missing from a list simply contributes nothing from that list. The paper fixed k = 60 "during a pilot investigation and not altered during subsequent validation," explaining that "the constant k mitigates the impact of high rankings by outlier systems" — without it, a single method ranking a document #1 could dominate the fused score no matter how the other methods ranked it. Their own pilot table shows exactly how little the choice matters: mean average precision moved from 0.2139 at k=30 to 0.2144 at k=50 to 0.2145 at k=60 to 0.2142 at k=100 — a near-flat curve, with 60 simply the best of a wide, forgiving plateau. Elasticsearch's RRF retriever ships the identical default (rank_constant: 60) today, so this isn't just an academic footnote — it's the number you'll meet in production.

Worked example: fusing a BM25 list and a vector list

Say the query is "database backup schedule" and each method returns its top 5 matches from a shared document pool. The two lists only partly overlap:

RankBM25 keyword searchVector similarity search
1Doc ADoc B
2Doc BDoc D
3Doc CDoc F
4Doc DDoc A
5Doc EDoc G

Apply 1 / (60 + rank) to every appearance and sum per document. Doc A: rank 1 in BM25 gives 1/61 = 0.01639, plus rank 4 in vector gives 1/64 = 0.01563, for a total of 0.03202. Doc B: rank 2 in BM25 gives 1/62 = 0.01613, plus rank 1 in vector gives 1/61 = 0.01639, for a total of 0.03252. Doing the same for every document:

DocBM25 termVector termRRF score
B1/62 = 0.016131/61 = 0.016390.03252
A1/61 = 0.016391/64 = 0.015630.03202
D1/64 = 0.015631/62 = 0.016130.03175
C1/63 = 0.01587— (not retrieved)0.01587
F— (not retrieved)1/63 = 0.015870.01587
E1/65 = 0.01538— (not retrieved)0.01538
G— (not retrieved)1/65 = 0.015380.01538

The fused ranking is B, A, D, then C and F tied, then E and G tied. Notice what happened to Doc A: it was the single best keyword match, ranked #1 by BM25, yet it finishes behind Doc B, which was never better than #2 on either list. Showing up solidly on both lists beat being the top result on just one — that's the entire point of fusing two independent signals instead of trusting either one alone.

The alternative: weighted score fusion with alpha

RRF isn't the only way to combine two ranked lists. The other common approach normalizes each method's raw scores into a shared 0-to-1 range (typically min-max over the candidate set) and combines them as a weighted sum: score = (1 - alpha) × bm25_norm + alpha × vector_norm. Weaviate's hybrid search uses exactly this convention, with alpha defaulting to 0.5 — equal weight to both searches, adjustable toward pure keyword (alpha = 0) or pure vector (alpha = 1) per query.

Reuse the same seven documents, now with raw scores attached — BM25's unbounded term-frequency score and the vector search's cosine similarity:

DocBM25 raw scoreVector raw score
A18.20.81
B15.40.91
C9.7—
D7.10.88
E5.3—
F—0.85
G—0.78

Min-max normalizing each column (BM25 ranges 5.3-18.2, a span of 12.9; vector ranges 0.78-0.91, a span of 0.13) and averaging at alpha = 0.5 gives: B = 0.5×0.783 + 0.5×1.000 = 0.891; A = 0.5×1.000 + 0.5×0.231 = 0.615; D = 0.5×0.140 + 0.5×0.769 = 0.454; F = 0.5×0 + 0.5×0.538 = 0.269; C = 0.5×0.341 + 0.5×0 = 0.171; E and G both = 0.

The ranking is still B, A, D first, matching RRF — but now F clearly outranks C, whereas RRF scored them identically. The reason: RRF only ever sees "rank 3 in one list," and a rank-3 finish is a rank-3 finish regardless of how strong the underlying match was. Weighted fusion keeps the magnitude — F's 0.85 cosine similarity sits much closer to the field's top score than C's 9.7 term-frequency score does to its field's top — so it breaks the tie in F's favor.

RRF vs weighted fusion: the trade-off

Reciprocal Rank FusionWeighted score fusion
Needs score normalizationNo — works on rank position onlyYes — min-max or similar, recomputed whenever the candidate set changes
Tunable per queryBarely — only k, rarely changed from 60Yes — alpha shifts weight toward keyword or vector search
Sensitive to an outlier raw scoreNo, by constructionYes — one unusually high or low score shifts every other document's normalized value in that batch
Uses how strong a match was, not just where it rankedNoYes

Both defaults above come from production systems, not just the original paper, so they're worth knowing by heart: Elasticsearch's RRF retriever defaults rank_constant to 60, and Weaviate's hybrid search defaults alpha to 0.5. For more on how BM25's own scoring works, why vocabulary mismatch breaks a single retrieval method, and other hybrid-retrieval mechanics, practice against the RAG & embeddings quiz.

Source: Cormack, Clarke & Büttcher, "Reciprocal Rank Fusion outperforms Condorcet and Individual Rank Learning Methods," SIGIR 2009 (formula, k=60 derivation, pilot MAP-vs-k table); Elastic, Elasticsearch RRF retriever reference documentation (rank_constant default); Weaviate, hybrid search documentation (alpha default).

Drill RAG & Embeddings practice questions →