Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Intent Routing and Classification

Recall the second query from the chapter opener: “Bücher von Goethe”. By this point in the chapter, the pipeline has enough machinery to see it clearly. The tokenizer splits it into three tokens. The rule-based language detector recognizes German characters and function words. POS tagging identifies “Bücher” as a plural noun, “von” as a preposition, and “Goethe” as a proper noun. Named entity recognition promotes “Goethe” to a person entity. What has not been decided yet is what to do with all this structure. Should the query go to a book catalogue, a web index, an author database, or a general-purpose LLM? This section turns the extracted features into that routing decision.

The end-to-end query pipeline

A query enters the retrieval system as a raw string. Before any matching happens, it flows through the classical pipeline this chapter has been building:

raw query
    -> tokenize                       (section 1)
    -> normalize (case, Unicode)      (section 1)
    -> segment sentences              (section 1, if multi-sentence)
    -> detect language                (section 1 rules + this section's classifier)
    -> stop-word filter               (section 2, language-specific)
    -> stem or lemmatize              (section 2, language-specific)
    -> split compounds                (section 3, if applicable)
    -> POS-tag                        (section 4)
    -> extract named entities         (section 4)
    -> spell-correct                  (section 4)
    -> classify intent                (this section)
    -> route to backend               (this section)

The order matters. Language detection has to come before any language-specific step; POS tagging has to come before selective stop-word filtering; NER has to come before intent classification because the entity types are input features to the classifier. What comes out at the end is a structured representation of the query that a downstream system can act on: a language, a set of entities, an intent label, and a corrected token list.

Two of these classification steps use the same underlying machinery. Language detection classifies a query into one of the languages the system knows; intent classification classifies a query into one of the backends the system supports. Both are text classification problems and both can be solved with the same Naive Bayes model. This section develops that model once, then applies it to both problems.

Naive Bayes for text classification

Naive Bayes chooses the most probable class given the observed features by applying Bayes’ theorem and assuming that features are conditionally independent given the class. The derivation lives in the ML foundations appendix; here we state only the decision rule.

For a feature vector x=(x1,…,xM)\mathbf{x} = (x_1, \ldots, x_M) and classes C1,…,CKC_1, \ldots, C_K:

Three things need to be learned from a labelled training corpus:

The product in the decision rule is a direct consequence of the independence assumption: taking the features to be conditionally independent given the class turns the joint likelihood P(x∣Ck)P(\mathbf{x} \mid C_k) into the product ∏jP(xj∣Ck)\prod_j P(x_j \mid C_k) of per-feature terms. This is what makes the estimation possible at all. In classical text classification the feature space has tens or hundreds of thousands of dimensions, so any particular full vector x\mathbf{x} almost never recurs in the training data, and P(x∣Ck)P(\mathbf{x} \mid C_k) could never be estimated directly. Each individual dimension xjx_j, by contrast, is observed many times across the corpus, so the per-feature factors P(xj∣Ck)P(x_j \mid C_k) can be counted reliably. Independence trades a joint probability we cannot measure for a product of marginals we can.

The assumption is not literally true, since tokens in real text are not independent, but Naive Bayes still performs well on text because the errors it introduces tend to affect all classes similarly and rarely change which class wins the argmax. Its speed and small memory footprint are the reasons it stays in production stacks two decades after neural classifiers became viable.

Language detection as classification

Language detection uses character-based n-grams as features. For nn from 1 to 5, count how many times each n-gram appears in the input text. This produces a bag-of-n-grams representation exactly like the bag-of-words representation used for topic classification, except that the tokens are short character strings instead of whole words.

Character n-grams work well as features because each language has characteristic short sequences that rarely appear in other languages:

Some trigrams like ent appear in multiple languages but with different frequencies, which is enough for a Naive Bayes classifier to disambiguate.

Setting up the classifier:

The decision rule then takes the standard log-form:

C^=arg⁡max⁡k(log⁡P(Ck)+∑jcjlog⁡P(tj∣Ck))\hat{C} = \arg\max_{k} \left( \log P(C_k) + \sum_{j} c_j \log P(t_j \mid C_k) \right)

where cjc_j is the observed count of n-gram jj in the input.

For very short queries with only one or two trigrams that discriminate between languages, confidence drops. lingua reports its confidence as a normalized posterior, and the retrieval system can fall back to a default language (English, or the user’s browser locale) when the top confidence is too close to the second-best.

from lingua import Language, LanguageDetectorBuilder

detector = LanguageDetectorBuilder.from_all_languages().build()
detector.detect_language_of("Bücher von Goethe")
# Language.GERMAN

# Restrict to expected languages and get confidence scores
euro = [Language.ENGLISH, Language.GERMAN, Language.FRENCH, Language.ITALIAN]
d2 = LanguageDetectorBuilder.from_languages(*euro).build()
d2.compute_language_confidence_values("Bücher von Goethe")
# GERMAN: 0.96, ENGLISH: 0.02, FRENCH: 0.01, ITALIAN: 0.01

langdetect is a lighter alternative: rule- and n-gram-based, 55 languages, ISO-code output.

from langdetect import detect
detect("Bücher von Goethe")   # 'de'
detect("Livres de Goethe")    # 'fr'
detect("Books by Goethe")     # 'en'

Intent routing as classification

The same machinery routes the query to a backend. Once the language is known, an intent classifier decides which of the search system’s supported intents the query expresses:

The training data is query logs annotated with the backend the user actually clicked into. Features are richer than for language detection because more of the pipeline’s output can be used:

With these features, “Bücher von Goethe” scores highly on book_search because it contains a book-domain plural noun (“Bücher”) and a PERSON entity (“Goethe”) in an author-attribution construction. “What to do in Basel?” scores highly on map_search and web_search because it contains a WH-word, a LOCATION entity, and no product or person entity. “Who is Albert Einstein?” scores highly on people_search because of the WH-word combined with a PERSON entity.

The classifier itself is exactly the same Naive Bayes model as for language detection, just with different features and different classes. In practice, a modern production stack fine-tunes a small neural classifier on top of the same features because it captures conjunctions (“has-PERSON AND WH-word AND is-question”) that Naive Bayes cannot express under the independence assumption. In a later chapter, we study the neural variants using LLM to make the routing decision (agentic RAG).

End-to-end walkthrough

Putting all five sections of the chapter together on the running query:

Input:              "Bücher von Goethe"

Tokenize:           ['Bücher', 'von', 'Goethe']
Normalize (case):   ['bücher', 'von', 'goethe']
Detect language:    German (confidence 0.96)
Stop words:         drop 'von'  ->  ['bücher', 'goethe']
Stem (Snowball DE): ['buch', 'goeth']
POS tag:            [(Bücher, NOUN), (Goethe, PROPN)]
NER:                [(Goethe, PERSON)]
Spell-check:        no correction needed
Intent classify:    book_search  (features: language=DE, PERSON present,
                                  plural NOUN of book-domain lemma)
Route:              library catalog with filters
                        language = "de"
                        author   = "Goethe"
                        content  = *

The library catalog answers a structured query with all Goethe titles it holds. No general-purpose keyword match against “Bücher von Goethe” was ever needed.

A second walkthrough on the other chapter-opener query. This one combined a WH-question form, tokens that would not appear in the answer, and a domain-vocabulary gap.

Input:                       "Who won the F1 race on the weekend?"

Tokenize:                    ['Who', 'won', 'the', 'F1', 'race', 'on',
                              'the', 'weekend', '?']
Normalize (case):            ['who', 'won', 'the', 'f1', 'race', 'on',
                              'the', 'weekend']
Detect language:             English
POS tag:                     [(who, WH-PRON), (won, VERB), (the, DET),
                              (F1, PROPN), (race, NOUN), (on, ADP),
                              (the, DET), (weekend, NOUN)]
NER:                         [(F1, EVENT), (the weekend, DATE)]
Selective stop-word filter:  drop 'the', 'on'
                             keep 'who' as a classification feature
                                  (WH-word signals question form)
Stem/lemma (English):        ['win', 'F1', 'race', 'weekend']
Synonym expansion (domain):  F1     -> {Formula 1, Grand Prix}
                             race   -> {Grand Prix, GP}
Temporal normalization:      the weekend -> [Sat, Sun of the past weekend]
                                          -> ['2026-08-08', '2026-08-09']
Intent classify:             news_search
                                 features: WH-word + EVENT entity +
                                           DATE entity + motorsport terms
                                 features against 'question' subclass:
                                           WH-word + won + who
Route:                       news index with
                                 query = "Formula 1" OR "Grand Prix" OR "GP"
                                 date  = [2026-08-08, 2026-08-09]
                                 sort  = date desc

The classical pipeline turned three tokens of question form and one date reference into a scoped keyword query the news index can answer: a Boolean OR over motorsport aliases, a date filter for the past weekend, and a sort by date. Each of the three gaps named in the chapter opener has a specific step above that addresses it. The missing “who” is absorbed into the classification features rather than searched as a keyword. The missing “weekend” becomes a temporal range that a race-results page dated in that range will match. The missing “race” is replaced by the domain-synonym expansion.

That is as far as this chapter takes the F1 example. What the news backend returns is a ranked list of race-result pages, not the two-word answer “Verstappen won” that the user actually requested. Two later chapters carry the example the rest of the way:

The classical pipeline of this chapter is still there in both cases. It prepares the query and prunes the search space so the more expensive downstream steps only see the pages that matter. The router itself does not do the retrieval; it picks a backend and passes the structured query to it. Everything upstream in this chapter was building the structure that the router now dispatches on.

Where this sits in modern retrieval

Classical intent routing runs in every large search stack in 2025. Google Search, Amazon Search, and enterprise search products all have an intent-classification stage before the retriever. The classifier is rarely still Naive Bayes: production systems have moved to gradient-boosted trees or small transformers because the feature interactions matter and the classifier is on the query-latency critical path only in a few-millisecond budget. Naive Bayes remains the pedagogical baseline and a common first pass in cost-sensitive deployments (mobile, embedded).

Large language models can perform the whole pipeline in one prompt: extract entities, detect the language, classify the intent, and generate the structured backend query. This is a live area of production adoption, especially for open-ended assistant systems. The trade-offs against classical pipelines are latency (an LLM call is 100-1000 milliseconds versus sub-millisecond for the classical stack), cost (per-query LLM cost is measurable, per-query classical cost is not), and confidence quantification (Naive Bayes gives calibrated posteriors, LLMs do not). Production stacks in 2025 usually run the classical pipeline first, escalate to LLMs for the queries where the classical stack’s confidence is low, and reserve LLMs for high-value verticals like agentic search and complex reformulation.

Two forward references pick up where this section leaves off:

The next page summarizes the chapter and points to the material that follows it in the book.