Summary
Metric Lookup Table¶
| Metric | Formula | What it rewards | When to use |
|---|---|---|---|
| Precision | Clean results (few false positives) | Any binary decision; especially when FP is costly | |
| Recall | Completeness (few misses) | Exhaustive search; patent, legal, medical | |
| Balances precision and recall with tuneable weight | Single-number summary; default | ||
| Precision in top | Clean top-of-list | Web search, short result pages | |
| MRR | Mean of | First relevant result high | QA, known-item search |
| R-Precision | Balance at a natural cutoff | Self-adjusting single-point summary | |
| AP | Mean of at relevant positions, divided by $ | \text{Rel} | $ |
| MAP | Mean of per-query AP | System-level ranking quality | Benchmark comparison (TREC standard) |
| High-grade documents placed early | Web search, graded relevance, user experience | ||
| Accuracy | Overall correct decisions | Balanced classification; equals micro-precision and micro-recall | |
| Specificity (TNR) | Few false alarms | Medical tests, biometric security | |
| AUC | Area under ROC curve | Score separation across all thresholds | Comparing classifiers without committing to a threshold |
Key Takeaways¶
Evaluation requires a benchmark: a collection, information needs, relevance judgments, and a performance goal. Without any component, a number is not interpretable.
Precision and recall are in tension. Improving one typically costs the other, and the -measure makes the trade-off explicit.
Ranked evaluation adds position: AP and nDCG reward systems that place relevant (or high-grade) material where users will see it.
Graded relevance distinguishes “somewhat useful” from “essential”. nDCG captures what binary metrics cannot: the quality of the relevant documents, not only their presence.
Classification evaluation reuses precision and recall but adds accuracy, specificity, and the confusion matrix. Accuracy equals micro-precision equals micro-recall in single-label settings.
Prevalence distorts accuracy: when one class dominates, a trivial baseline can score higher than a real classifier. Always check class balance before trusting accuracy.
A continuous score becomes a binary decision only after choosing a threshold. The ROC curve shows all possible trade-offs; AUC summarizes overall score separation.
No benchmark score is a claim about generalization. Incomplete judgments, data contamination, prevalence mismatch, and overfitting to static test sets all limit what a number can mean.
Key Formulas¶
Precision and recall: the fundamental pair from which all other metrics derive.
Average precision: area under the precision-recall curve, the standard per-query measure.
Normalized discounted cumulative gain: the standard for graded relevance evaluation.
Area under the ROC curve: the probability that a random positive scores higher than a random negative.
Self-Check Questions¶
(Understand) A system achieves 95% accuracy on a dataset where 95% of items belong to class A. Is this system useful? What additional metric would reveal the problem?
(Analyze) System X has MAP = 0.60 and System Y has MAP = 0.55, both measured on 50 queries. Under what circumstances might Y actually be the better system?
(Evaluate) A face-unlock system reports AUC = 0.99 and a spam filter reports AUC = 0.99. Are they equally good? What additional information do you need to judge each for its intended use?
(Apply) Given a ranked list of 10 documents where relevant documents appear at positions 1, 4, and 7 (out of 8 relevant in the collection), compute , AP, and explain why AP is less than the average of the three precision values.
(Analyze) Explain why macro-averaged recall across 4 information needs can disagree with micro-averaged recall, and give a concrete scenario where the disagreement is large.
Further Reading¶
Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to information retrieval, Chapter 8: Evaluation in information retrieval. Cambridge University Press. Read online. A concise foundation for precision, recall, MAP, and benchmark design.
Järvelin, K., & Kekäläinen, J. (2002). Cumulated gain-based evaluation of IR techniques. ACM Transactions on Information Systems, 20(4), 422–446. ACM record. The paper that introduced DCG and nDCG for graded relevance evaluation.
Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8), 861–874. Read the paper. A practical guide to ROC curves, AUC, and threshold selection for binary classifiers.
Armstrong, T. G., Moffat, A., Webber, W., & Zobel, J. (2009). Improvements that don’t add up: Ad-hoc retrieval results since 1998. In Proceedings of CIKM 2009 (pp. 601–610). ACM record. An influential warning that static benchmarks can conceal a lack of cumulative progress.
TREC (Text REtrieval Conference). NIST’s continuing evaluation campaigns provide the pooled collections, shared topics, and comparative methodology behind much of this chapter’s retrieval evaluation practice.