Summary
Model Comparison¶
| Property | Boolean | Extended Boolean | Vector Space | BIR | BM25 |
|---|---|---|---|---|---|
| Query form | Boolean expression | Boolean expression | Free text | Term set | Free text |
| Ranked output | No | Yes | Yes | Yes | Yes |
| Partial matches | No | Yes | Yes | Yes | Yes |
| Term frequency | Ignored | Normalized TF-IDF | Linear TF-IDF | Ignored | Saturating |
| Document length | Ignored | Indirect normalization | Cosine or none | Ignored | Explicit parameter |
| Foundation | Set theory | Soft-logic heuristic | Geometric heuristic | Relevance probability | BIR plus TF and length models |
| Typical role | Exact filters | Specialized soft constraints | Baseline and vector similarity | Feedback model and foundation | Production lexical ranking |
| Main weakness | No ranking | Operator choices | Length and repetition effects | Binary evidence and feedback burden | Lexical matching and tuning |
The progression is cumulative rather than a sequence of complete replacements. Boolean predicates remain useful for mandatory filters, vector operations remain central to dense retrieval, BIR explains relevance feedback and probabilistic term weights, and BM25 provides the strongest general-purpose lexical ranking baseline.
Key Takeaways¶
Text retrieval represents documents and queries using the same vocabulary, then measures similarity to estimate relevance. This direct word-level matching makes text the least affected modality by the semantic gap.
The feature extraction pipeline (Extract → Split → Tokenize → Summarize) transforms raw text into searchable representations. The bag-of-words model records term frequencies; the set-of-words model records only term presence.
IDF weighting captures term discrimination power: rare terms that appear in few documents are more informative for distinguishing relevant from non-relevant results.
Retrieval models evolved from Boolean (no ranking, exact match) through the Extended Boolean Model (partial matches, simple scoring) to the Vector Space Model (tf-idf, cosine similarity) and Probabilistic Retrieval (BIR model, relevance feedback).
BM25 combines the best ideas: saturating term frequency, document-length normalization, and probabilistically grounded IDF. It remains the default baseline in modern search engines (Lucene, Elasticsearch, OpenSearch).
Key Formulas¶
Self-Check Questions¶
(Understand) Explain in your own words why the Boolean model cannot rank documents, and what the Extended Boolean Model changes to enable ranking.
(Analyze) Given a collection of 10,000 documents where the term “retrieval” appears in 50 documents and the term “the” appears in 9,800 documents, calculate the IDF for both terms. Why does this difference matter for relevance scoring?
(Analyze) A document is 3× longer than the average document in the collection. How does BM25 (with ) adjust the effective term frequency compared to an average-length document? What happens when ?
(Evaluate) An e-commerce search system currently uses Boolean retrieval with faceted filtering. A developer proposes switching to BM25. What are the trade-offs? In which scenarios might Boolean retrieval still be preferable?
Further Reading¶
Spärck Jones, K. (1972). A Statistical Interpretation of Term Specificity and Its Application in Retrieval. Journal of Documentation, 28(1), 11-21. PDF. Introduced inverse document frequency; foundational for every term-weighting scheme discussed in this chapter, including TF-IDF and BM25.
Salton, G., Wong, A., & Yang, C. S. (1975). A Vector Space Model for Automatic Indexing. Communications of the ACM, 18(11), 613-620. PDF. The original paper introducing the Vector Space Model; still readable and surprisingly concise.
Robertson, S. E., & Spärck Jones, K. (1976). Relevance Weighting of Search Terms. Journal of the American Society for Information Science, 27(3), 129-146. Publisher page. Formalizes the Binary Independence Model and its / relevance-feedback estimates covered in this chapter.
Salton, G., Fox, E. A., & Wu, H. (1983). Extended Boolean Information Retrieval. Communications of the ACM, 26(11), 1022-1036. Free access. Introduces the p-norm model, positioning graded Boolean retrieval between the strict Boolean and vector-space models.
Porter, M. F. (1980). An Algorithm for Suffix Stripping. Program, 14(3), 130-137. Author’s page. The classic stemming algorithm still used as a baseline in English text processing.
Robertson, S. E., & Walker, S. (1994). Some Simple Effective Approximations to the 2-Poisson Model for Probabilistic Weighted Retrieval. In Proceedings of the 17th ACM SIGIR Conference, 232-241. PDF. Derives BM25’s term-frequency saturation and length normalization from the 2-Poisson model discussed in the Probabilistic Ranking and BM25 section.
Manning, C. D., Raghavan, P., & Schütze, H. (2008). Introduction to Information Retrieval. Cambridge University Press. Free online edition. The standard graduate textbook; chapters 1 to 6 cover Boolean, VSM, and probabilistic models with rigorous treatment.
Robertson, S. E., & Zaragoza, H. (2009). The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends in Information Retrieval, 3(4), 333-389. PDF. Definitive survey of BM25’s theoretical foundations and practical extensions.