Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Designing a Retrieval Benchmark

A university library holds 50 books. A student submits one information need: which books give a beginning master’s student solid computer science foundations? Fifteen of the 50 books are useful for that need. Two retrieval systems answer the same request, and their first five results look like this.

RankSystem A resultUseful?System B resultUseful?
1Computer Organization and DesignyesDatabase Systems: The Complete Bookyes
2Operating System ConceptsyesPattern Recognition and Machine Learningyes
3The Handmaid’s TalenoIntroduction to Algorithmsyes
4Narrative of the Life of Frederick DouglassnoOn the Origin of Speciesno
5Towards a New ArchitecturenoIntroduction to the Theory of Computationyes

System A does not stop at rank 5. It returns 25 books in total, half the library, and finds 12 of the 15 useful ones. System B returns only 8 books, of which 6 are useful. So System A surfaces more of what exists, and System B wastes less of the reader’s attention.

Which system is better? Two students with the same assignment can reasonably disagree. One wants a short list to read end to end and prefers System B. Another wants to be certain that nothing important was missed and prefers System A. Neither is wrong, so before we can measure anything, we have to say what counts as a good result.

What Counts as a Better Result

Relevance is not a property of a book. It is a judgment about whether a book serves one person’s need at one moment. “Introduction to Algorithms” is essential to our student and useless to someone looking for a novel for a train journey. Every measurement in this chapter therefore rests on a human decision taken before any system ran.

That subjectivity is unavoidable, but it is containable. Once the need is written down and the judgments are recorded, the counting becomes mechanical: anyone who applies the same procedure to the same lists obtains the same numbers. Evaluation does not remove the subjective step, it isolates it.

What counts as success then varies enormously with the need. At one extreme only the first few results matter. Someone checking when BM25 was published reads two or three entries and stops, so anything below rank 5 might as well not exist. A short and precise list wins here, in the style of System B, and whatever relevant material it leaves behind costs the reader nothing. At the other extreme, missing anything is the real failure. A patent lawyer searching for prior art will work through a long and noisy list rather than risk overlooking one document, so broad coverage wins, in the style of System A, and a very short list would be useless however precise it was.

Many information needs are in between these extremes, including our student’s. A foundations reading list has to be broad enough to cover the main areas of the field, yet short enough to work through in a semester. Neither system serves that well: System A buries the useful books among thirteen novels, and System B leaves out nine of the fifteen that would have helped. Web search leans towards the first extreme, biomedical literature search towards the middle, and the lesson is the same in every domain: “better” has no meaning until the intended use is named.

Why Anecdotes Are Not Evidence

Product advertising is full of comparative claims with no shared basis. A vendor states that a search engine is forty percent more accurate, without naming the collection it was tested on, the queries that were used, who judged the results, or what “accurate” counts. Such a claim cannot be checked, and it is usually easy to construct: pick the queries where your system happens to win, and the number follows.

Retrieval systems invite exactly this failure, because single-example demonstrations are trivially available. For almost any pair of systems, someone can find one query where the first beats the second and another query where the reverse holds. The two result lists at the start of this section are a demonstration, not evidence. They describe one need in a 50-book collection, so they support no general statement about either system.

A scientific comparison replaces the anecdote with a fixed experimental setup. The collection is frozen, the information needs are written down in advance, the relevance judgments are made independently of the systems being compared, and the measurement procedure is stated before any results are seen. Others can then rerun the same experiment and obtain the same numbers. This is what a benchmark provides: not a guarantee that the winner is better for every purpose, but a claim that is reproducible and open to challenge.

Comparable benchmarks let retrieval research accumulate: each generation preserves the fixed collection, information needs, and judgments that make its results comparable with the last. The history below shows how Cranfield, TREC, multilingual campaigns, MS MARCO, and BEIR adapted that shared-evaluation idea to new collections and new failure modes.

The Four Components of a Benchmark

A benchmark is built from four parts: a document collection, a set of information needs, relevance judgments, and a measurement procedure. Each part answers one question, and a weakness in any of them undermines everything measured on top.

The document collection defines the universe the system may search. It can hold news articles, scientific papers, legal filings, product descriptions, or, as in our running example, library records. Two properties matter most. The collection must be stable, because results are only comparable over time if the searched material does not change underneath them. It must also be representative, reflecting the document lengths, content types, and topic mix that the target users really encounter. Well-known examples include the TREC collections, which range from newspaper archives to biomedical abstracts, MS MARCO for web-style passage retrieval, and domain corpora such as PubMed Central and arXiv for scientific literature.

The information needs state what the users actually want. It is worth separating the need from the query string: the need is “which books give a beginning master’s student solid computer science foundations”, while the query is whatever short text the user types into the search box. Assessors judge against the need, not the keywords, which is why the need must be written out clearly enough that two people would interpret it the same way. Needs are best derived from real user behaviour such as query logs. When no log exists, large language models can draft candidate needs from samples of the collection, which expands coverage cheaply, though the drafts still require review to stay realistic.

The relevance judgments record, for each need, which documents satisfy it. They are the ground truth, so their quality bounds the quality of every conclusion. Human assessors work from written guidelines, receive training on example cases, and are checked for agreement, because vague criteria produce inconsistent labels that no metric can repair. Large language models can pre-screen documents and propose judgments for clear-cut cases, leaving human effort for the borderline ones. Judgments must then be frozen: revising them later silently invalidates every result measured earlier.

The measurement procedure turns a result list into a number. It fixes which measure is computed, at which cutoff, and how values are combined across needs. The next three sections develop these measures, starting with results treated as an unordered set, then as a ranking, then as a ranking with graded relevance. What matters here is that the procedure is chosen before the experiment, not after the results are known.

The following table collects the practical rules for each component.

ComponentWhat it providesDesign ruleCommon failure
Document collectionThe universe that may be searchedFreeze it for the benchmark’s lifetime, make it representative of the target domain, preprocess consistently, and give every document a stable identifierBuilt around one feature, so it flatters the methods that exploit that feature
Information needsWhat the target users wantDerive needs from real user behaviour, state each one unambiguously, and provide 25 to 50 for a first experiment or 100 and more for a serious comparisonWording so vague that assessors disagree and the ground truth becomes noisy
Relevance judgmentsThe ground truth for scoring resultsWrite assessor guidelines, train assessors, check their agreement, and freeze the labels once publishedJudgments revised mid-project, which invalidates all earlier measurements
Measurement procedureThe rule that turns results into numbersState the measure, the cutoff, and the averaging method before running any experimentMeasure chosen after seeing the results, which turns evaluation into advocacy

Building all four parts is expensive, which is why established benchmarks are reused heavily and why a small, carefully judged collection is often more valuable than a large, carelessly labeled one.

Judging at Scale

One practical obstacle remains. Judging every document for every need is impossible at realistic scale: fifty books can be assessed exhaustively, fifty million cannot. Pooling is the standard response. The organizers collect the documents that participating systems actually retrieved, judge that pool, and treat everything unjudged as not relevant. Figure 1 shows the situation this assumes. Relevant documents that no participant retrieved never get judged, so the measured values differ somewhat from what exhaustive judging would produce. For a comparison that matters less than it first appears: every participant is scored against the same judged subset, so the ordering between them survives even though the individual numbers are approximations.

Pooled assessment when participating systems retrieve most relevant documents.

Figure 1:Pooled assessment when participating systems retrieve most relevant documents.

This assumption is sound while the pool covers most of the relevant material, and it weakens as the judged fraction shrinks. Pooling is one of several places where a benchmark can mislead rather than inform, and When Evaluation Misleads returns to all of them once we have the measures needed to describe what gets distorted.

With a benchmark in place, the remaining question is mechanical: how do we turn a result list into a number that reflects how well the need was served? The next section answers it for results treated as an unordered set, using the same library collection and the same two systems.