CustomLabs
Retrieval

How do you get both broad recall and precise ranking out of retrieval?

Intent

A cheap, high-recall first pass — vector search, optionally fused with keyword search — pulls a wide candidate set of 50 to 100 documents likely to contain the right answer somewhere. A precision reranker, usually a cross-encoder that scores the query and each candidate jointly, then re-sorts that set and applies a relevance floor allowed to return nothing rather than force a weak match to the top.

Also known as two-stage retrieval, cross-encoder rerank, candidate-then-rank

Problem

What problem does this solve?

A single-stage vector search has to be both broad enough to find the right document among thousands and precise enough to rank it above near-misses, using one similarity score computed independently for each document against the query. Bi-encoder similarity is good at "about the same topic" and much weaker at the specific detail that actually decides correctness — an identifier, a negation, a specific date — so a topically similar but wrong document can outrank the correct one.

Retrieval that always returns its top-k, with no floor, treats "nothing in the corpus actually answers this" identically to "here is the answer" — both come back as a ranked list, and a downstream model reading that list has no signal that it should say "I don't know" instead of confidently answering from the best-available-but-wrong match.

Mechanics

How does it work?

Stage one casts a wide net cheaply: vector search, optionally fused with BM25 keyword search for the identifier and exact-phrase cases vectors miss, pulls a candidate set of 50-100 documents, favoring recall over precision — false positives here are fine, because the second stage will filter them.

Stage two scores the query jointly against each candidate with a cross-encoder, which is more expensive per comparison than bi-encoder similarity but far more accurate, because it can actually attend to both texts together instead of comparing two independently-computed vectors. This is affordable specifically because it only runs on the narrowed candidate set, not the whole corpus.

A relevance floor on the reranked scores decides whether anything qualifies as an answer at all — if the top reranked result still falls under the floor, the system returns "no confident match" rather than the best of a bad set. That floor is a real design decision, usually validated against a labelled eval set, and returning nothing is treated as correct behavior, not a failure to log.

Shape

What are the moving parts?

  1. Retrieve a wide candidate set

    Pull 50-100 candidates with a cheap, high-recall method, favoring recall over precision at this stage.

  2. Fuse in keyword retrieval where useful

    Add BM25 keyword search alongside vector search specifically for identifier and exact-phrase queries.

  3. Score jointly with a cross-encoder reranker

    Re-rank the narrowed candidate set with a model that attends to the query and each candidate together.

  4. Apply a relevance floor

    Decide, against a labelled eval set, the reranked score below which nothing qualifies as an answer.

  5. Return "no confident match" below the floor

    Treat returning nothing as correct behavior when nothing clears the floor, not as a failure to paper over.

  6. Feed only floor-cleared results downstream

    Generation only ever sees the reranked, floor-passing results, never the raw first-stage candidate set.

Fit

When does it fit, and when doesn't it?

Use it when

  • Retrieval quality complaints trace back to a topically-similar-but-wrong document outranking the correct one
  • The corpus is large enough that single-stage similarity search has real precision problems, not just recall problems
  • Some queries genuinely have no answer in the corpus, and a confident wrong answer is worse than admitting that
  • You can afford the extra latency and compute of a second scoring pass on a narrowed candidate set

Don't use it when

  • The corpus is small enough that single-stage retrieval already has near-perfect precision — a reranker adds latency for no measurable gain
  • Latency budget genuinely can't absorb a second model call in the retrieval path (hard real-time constraints)
  • You have no labelled eval set to calibrate the relevance floor against — an uncalibrated floor either blocks good answers or lets bad ones through, and you will not know which
Trade-offs

What does it actually cost?

Gain

Meaningfully better precision than single-stage similarity search, especially on near-miss and identifier-heavy queries

Cost

A second model call (the cross-encoder) on every request, adding latency and compute cost

Gain

A relevance floor lets the system say 'no answer' instead of guessing

Cost

Calibrating that floor takes a real labelled eval set — get it wrong and you either suppress good answers or let bad ones through

Gain

Keyword fusion recovers exact-match and identifier queries pure vector search misses

Cost

Rank fusion across two different scoring methods is its own tuning problem, not a free combination

Gain

Reranking only the narrowed candidate set keeps the expensive model affordable

Cost

Recall is capped by whatever the first-stage retrieval already missed — the reranker cannot recover a document the first stage never surfaced

Signals

How do you know it's working?

  • Precision@k on the final reranked, floor-applied result set against a labelled eval set
  • Rate of 'no confident match' responses, watched for both under- and over-triggering
  • Added latency from the rerank stage (p50/p95)
  • Recall@k for the first-stage candidate set specifically, since it caps everything downstream
Prevents

What failure modes does this prevent?

Not sure this is the right pattern?

A Ship Audit checks which of these patterns your system actually needs, prioritized against what's most likely to break first.

Book a Ship Audit

navigate select esc close