Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Precision and Recall

Boolean retrieval returns a set of documents with no order attached: a document either matches the query or it does not. This section evaluates that kind of result. Ranking comes in Evaluating Ranked Results; until then, a result list is treated purely as a set.

Recall System A and System B from Designing a Retrieval Benchmark, both answering the same need: which books give a beginning master’s student solid computer science foundations? System A returned 25 books, System B returned 8. To turn “returned” and “useful” into numbers, every book in the 50-book collection falls into exactly one of four groups, depending on whether the system retrieved it and whether it is actually relevant.

Counting Outcomes

RelevantNot relevant
RetrievedTrue Positive (TP)False Positive (FP)
Not retrievedFalse Negative (FN)True Negative (TN)

A true positive is a relevant book the system retrieved: a hit. A false positive is a book the system retrieved that turns out not to be relevant: noise in the result. A false negative is a relevant book the system failed to retrieve: a miss. A true negative is a book the system correctly left out. Figure 1 shows this partition as overlapping sets for both systems: the 50-book collection, the 15 relevant books, the retrieved books, and their intersection.

Set-theoretic view of retrieval performance for Systems A and B over the 50-book collection.

Figure 1:Set-theoretic view of retrieval performance for Systems A and B over the 50-book collection.

Counting each book in the collection against the two systems gives:

TPFPFNTN
System A1213322
System B62933

Every row sums to 50, the size of the collection, and every count came directly from checking the 15 known relevant books against each system’s result list.

Precision and Recall

The raw counts alone are hard to compare. System A has 12 true positives and System B has 6, but A also retrieved three times as many books and operates in a collection where 15 are relevant. A system retrieving the entire library would produce 15 true positives and look best by that count alone, which shows that an absolute count cannot separate a good system from one that simply returns everything. Ratios solve this: dividing by the appropriate total produces a value between 0 and 1 that is independent of how many books the system returned or how many relevant books the collection contains.

This is the trade-off from Designing a Retrieval Benchmark stated in numbers rather than in words. System B’s list is three-quarters useful material; System A’s list finds four-fifths of everything useful but buries it among thirteen novels. Whether 48% precision at 80% recall is better than 75% precision at 40% recall still depends on what the user prefers (quick fact check or exhaustive library scan).

A third quantity, fallout, measures the opposite kind of mistake: how much of the irrelevant material got pulled in.

f=FPFP+TNf = \frac{FP}{FP + TN}

Fallout is the fraction of non-relevant documents that the system retrieved. For our two systems, fA=13/35=37.1%f_A = 13/35 = 37.1\% and fB=2/35=5.7%f_B = 2/35 = 5.7\%. Fallout measures the reading cost imposed on someone who inspects every result: the patent lawyer from Designing a Retrieval Benchmark needs high recall, because missing a document is expensive, but also wants low fallout, because every irrelevant document in the list costs time and attention. System A delivers recall of 80% but pulls in more than a third of the irrelevant collection; the ideal for an exhaustive search would push recall toward 1 while keeping fallout near 0.

Balancing the Two

Comparing two numbers side by side, as above, does not say which system wins. The FβF_\beta-measure collapses precision and recall into a single value, letting a scenario’s priority set the trade-off explicitly.

At β=0\beta = 0 the formula reduces to precision alone; as β→∞\beta \to \infty it reduces to recall alone. F1F_1, the case β=1\beta = 1, is the most common default and appears throughout machine learning wherever a single number is needed to compare classifiers.

The winner changes between β=0.5\beta = 0.5 and β=1\beta = 1. Neither system is better in general: a fact-checker who wants a short, clean list would set β\beta below 1 and prefer System B, while a patent lawyer would set β\beta above 1 and prefer System A. F1F_1 is a convenient default, not a neutral one, since it silently commits to weighing a false positive and a false negative equally.

Averaging Across Needs

One need produces one precision and one recall value. A real benchmark evaluates many needs, and summarising all of them into a single score requires an averaging method. The natural first idea is to compute precision and recall for each need separately and then take the arithmetic mean.

Example: Consider two additional needs alongside the CS-foundations query, answered by the same two systems:

NeedRelevantSystem A: retrieved / foundSystem B: retrieved / found
CS foundations1525 / 128 / 6
Stage plays56 / 44 / 3
Russian novels13 / 01 / 1

The straightforward summary is the arithmetic mean of the per-need values:

The problem is visible: macro-averaging gives each need the same weight 1/N1/N in the sum, so a need with one relevant document contributes as much to the average as a need with fifteen. A single zero from the Russian-novels miss lowers the three-need average by a full 1/31/3, which is a large negative impact relative to the system’s overall performance across 21 relevant documents. This is intentional when every need matters equally, but it can misrepresent a system’s aggregate capability.

Micro-averaging addresses this by pooling all outcome counts before computing the ratio. Instead of averaging per-need ratios, it sums the numerators and denominators across needs:

pmicro=∑i=1NTPi∑i=1N(TPi+FPi)rmicro=∑i=1NTPi∑i=1N(TPi+FNi)p_{\text{micro}} = \frac{\sum_{i=1}^{N} TP_i}{\sum_{i=1}^{N} (TP_i + FP_i)} \qquad r_{\text{micro}} = \frac{\sum_{i=1}^{N} TP_i}{\sum_{i=1}^{N} (TP_i + FN_i)}

Every document counts equally. Needs with more relevant documents contribute more to the result.

Same systems, same needs, opposite conclusions on recall: macro says B is better, micro says A is better. The difference reflects a genuine choice about what matters.

Architecture and the Metrics

Retrieval Systems and Flows introduced retriever-only, retriever-with-filter, and retriever-ranker architectures. Precision and recall explain why each stage exists and how each is evaluated. The retriever is optimized for recall: it casts a wide net over the collection, accepting low precision because a false positive can still be removed later, while a false negative is lost for good. The ranker or filter is optimized for precision: it reorders or removes candidates so that the top of the list is dense with relevant material. Even a user who cares only about a precise top-ten result depends on the retriever’s recall, because the ranker can only surface documents that the retriever found in the first place.

Precision, recall, and their averages describe a result set with no notion of order. The next section brings rank back into the picture: whether a system found the right documents matters less if they never appear near the top of the list.