Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Summary

Metric Lookup Table

MetricFormulaWhat it rewardsWhen to use
PrecisionTPTP+FP\frac{TP}{TP+FP}Clean results (few false positives)Any binary decision; especially when FP is costly
RecallTPTP+FN\frac{TP}{TP+FN}Completeness (few misses)Exhaustive search; patent, legal, medical
FβF_\beta(β2+1)⋅p⋅rβ2⋅p+r\frac{(\beta^2+1) \cdot p \cdot r}{\beta^2 \cdot p + r}Balances precision and recall with tuneable weightSingle-number summary; β=1\beta=1 default
P@kP@kPrecision in top kkClean top-of-listWeb search, short result pages
MRRMean of 1/rankfirst relevant1/\text{rank}_{\text{first relevant}}First relevant result highQA, known-item search
R-Precisionp@Rp@RBalance at a natural cutoffSelf-adjusting single-point summary
APMean of pkp_k at relevant positions, divided by $\text{Rel}$
MAPMean of per-query APSystem-level ranking qualityBenchmark comparison (TREC standard)
nDCG@knDCG@kDCGk/IDCGkDCG_k / IDCG_kHigh-grade documents placed earlyWeb search, graded relevance, user experience
AccuracyTP+TNP+N\frac{TP+TN}{P+N}Overall correct decisionsBalanced classification; equals micro-precision and micro-recall
Specificity (TNR)TNTN+FP\frac{TN}{TN+FP}Few false alarmsMedical tests, biometric security
AUCArea under ROC curveScore separation across all thresholdsComparing classifiers without committing to a threshold

Key Takeaways

  1. Evaluation requires a benchmark: a collection, information needs, relevance judgments, and a performance goal. Without any component, a number is not interpretable.

  2. Precision and recall are in tension. Improving one typically costs the other, and the FβF_\beta-measure makes the trade-off explicit.

  3. Ranked evaluation adds position: AP and nDCG reward systems that place relevant (or high-grade) material where users will see it.

  4. Graded relevance distinguishes “somewhat useful” from “essential”. nDCG captures what binary metrics cannot: the quality of the relevant documents, not only their presence.

  5. Classification evaluation reuses precision and recall but adds accuracy, specificity, and the confusion matrix. Accuracy equals micro-precision equals micro-recall in single-label settings.

  6. Prevalence distorts accuracy: when one class dominates, a trivial baseline can score higher than a real classifier. Always check class balance before trusting accuracy.

  7. A continuous score becomes a binary decision only after choosing a threshold. The ROC curve shows all possible trade-offs; AUC summarizes overall score separation.

  8. No benchmark score is a claim about generalization. Incomplete judgments, data contamination, prevalence mismatch, and overfitting to static test sets all limit what a number can mean.

Key Formulas

p=TPTP+FPr=TPTP+FNp = \frac{TP}{TP+FP} \qquad r = \frac{TP}{TP+FN}

Precision and recall: the fundamental pair from which all other metrics derive.

AP=1∣Rel∣∑k:dk∈RelpkAP = \frac{1}{|\text{Rel}|} \sum_{k: d_k \in \text{Rel}} p_k

Average precision: area under the precision-recall curve, the standard per-query measure.

nDCGk=∑i=1krelilog⁡2(i+1)IDCGknDCG_k = \frac{\sum_{i=1}^k \frac{rel_i}{\log_2(i+1)}}{\text{IDCG}_k}

Normalized discounted cumulative gain: the standard for graded relevance evaluation.

AUC=P(s(x+)>s(x−))AUC = P(s(x_+) > s(x_-))

Area under the ROC curve: the probability that a random positive scores higher than a random negative.

Self-Check Questions

  1. (Understand) A system achieves 95% accuracy on a dataset where 95% of items belong to class A. Is this system useful? What additional metric would reveal the problem?

  2. (Analyze) System X has MAP = 0.60 and System Y has MAP = 0.55, both measured on 50 queries. Under what circumstances might Y actually be the better system?

  3. (Evaluate) A face-unlock system reports AUC = 0.99 and a spam filter reports AUC = 0.99. Are they equally good? What additional information do you need to judge each for its intended use?

  4. (Apply) Given a ranked list of 10 documents where relevant documents appear at positions 1, 4, and 7 (out of 8 relevant in the collection), compute P@5P@5, AP, and explain why AP is less than the average of the three precision values.

  5. (Analyze) Explain why macro-averaged recall across 4 information needs can disagree with micro-averaged recall, and give a concrete scenario where the disagreement is large.

Further Reading