Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Summary

Model Comparison

PropertyBooleanExtended BooleanVector SpaceBIRBM25
Query formBoolean expressionBoolean expressionFree textTerm setFree text
Ranked outputNoYesYesYesYes
Partial matchesNoYesYesYesYes
Term frequencyIgnoredNormalized TF-IDFLinear TF-IDFIgnoredSaturating
Document lengthIgnoredIndirect normalizationCosine or noneIgnoredExplicit parameter bb
FoundationSet theorySoft-logic heuristicGeometric heuristicRelevance probabilityBIR plus TF and length models
Typical roleExact filtersSpecialized soft constraintsBaseline and vector similarityFeedback model and foundationProduction lexical ranking
Main weaknessNo rankingOperator choicesLength and repetition effectsBinary evidence and feedback burdenLexical matching and tuning

The progression is cumulative rather than a sequence of complete replacements. Boolean predicates remain useful for mandatory filters, vector operations remain central to dense retrieval, BIR explains relevance feedback and probabilistic term weights, and BM25 provides the strongest general-purpose lexical ranking baseline.

Key Takeaways

  1. Text retrieval represents documents and queries using the same vocabulary, then measures similarity to estimate relevance. This direct word-level matching makes text the least affected modality by the semantic gap.

  2. The feature extraction pipeline (Extract → Split → Tokenize → Summarize) transforms raw text into searchable representations. The bag-of-words model records term frequencies; the set-of-words model records only term presence.

  3. IDF weighting captures term discrimination power: rare terms that appear in few documents are more informative for distinguishing relevant from non-relevant results.

  4. Retrieval models evolved from Boolean (no ranking, exact match) through the Extended Boolean Model (partial matches, simple scoring) to the Vector Space Model (tf-idf, cosine similarity) and Probabilistic Retrieval (BIR model, relevance feedback).

  5. BM25 combines the best ideas: saturating term frequency, document-length normalization, and probabilistically grounded IDF. It remains the default baseline in modern search engines (Lucene, Elasticsearch, OpenSearch).

Key Formulas

Self-Check Questions

  1. (Understand) Explain in your own words why the Boolean model cannot rank documents, and what the Extended Boolean Model changes to enable ranking.

  2. (Analyze) Given a collection of 10,000 documents where the term “retrieval” appears in 50 documents and the term “the” appears in 9,800 documents, calculate the IDF for both terms. Why does this difference matter for relevance scoring?

  3. (Analyze) A document is 3× longer than the average document in the collection. How does BM25 (with b=0.75b = 0.75) adjust the effective term frequency compared to an average-length document? What happens when b=0b = 0?

  4. (Evaluate) An e-commerce search system currently uses Boolean retrieval with faceted filtering. A developer proposes switching to BM25. What are the trade-offs? In which scenarios might Boolean retrieval still be preferable?

Further Reading