Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

2 - Performance Evaluation

A student searches a university library for books that provide the computer science foundations needed for a master’s degree. One search system returns many useful books mixed with unrelated titles. Another returns fewer useful books, but places them near the top. Which system is better? The answer depends on what the student needs: broad coverage, a clean result set, or the best books in the first few positions.

Performance evaluation turns claims such as “better results” into measurements that can be checked and reproduced. A fair comparison requires more than a metric. It needs a representative collection, realistic information needs, consistent relevance judgments, and a clear definition of success. Without this shared benchmark, a higher score may reflect an easier test rather than a better retrieval method.

Evaluation also applies when a text system assigns labels rather than retrieving documents. A spam filter decides whether an email is unwanted, while an AI assistant may route a query to a knowledge base, web search, calculator, or clarification step. Retrieval evaluates decisions about query-document pairs; classification evaluates labels assigned to text items. Both expose the same fundamental trade-off: finding more desired items often means accepting more incorrect decisions.