Module 02 · Lesson 2 of 4
How Language Models Retrieve and Choose Sources
Retrieval is the step where a language model fetches candidate material before it writes a word. Modern answer systems are built as retrieval plus generation: a retrieval layer finds passages relevant to the question, then the language model composes an answer drawing on what was fetched. The generation gets the attention; the retrieval decides the outcome.
Passages, not pages
Classic ranking scores a URL against a query. Retrieval for generation is finer-grained: it hunts for passages that address the question, wherever they sit. Two consequences follow. A weak page can win on the strength of one excellent passage. And a strong page whose answer is buried under three paragraphs of throat-clearing can lose to a plainer page that answers in the second sentence.
The exact ranking mechanics inside any given system are not public. What you can observe, repeatedly, is what sources get used: pages that answer the actual question directly, in structured, self-contained chunks, with named entities and supporting evidence. Work to the observable behaviour and label your inference as inference.
What a fetchable passage looks like
| Trait | Fetchable version | Unfetchable version |
|---|---|---|
| Opening | "Query fan-out is the practice of issuing related searches across subtopics." | "In this section we will explore an interesting concept..." |
| Entity naming | Says "Google AI Overviews" every time it means it | "this innovative feature", "the platform" |
| Structure | Question heading, direct answer, then detail | Narrative that reveals the answer in paragraph four |
| Independence | Passage makes sense quoted alone | Depends on the previous page to mean anything |
You will practice writing these in Module 4. For now the point is recognition: start seeing pages as collections of passages in competition with other passages.
Conversational prompts change the fetch
A prompt like "I run a bakery in Leeds, how do I get mentioned in AI answers?" carries context (a bakery, a city, a goal) that a bare keyword never would. Retrieval for that prompt hunts for material matching the situation, not just the topic words. Content that speaks to specific situations ("for local businesses", "for beginners", "for agencies") gives retrieval more handles to grab. That is also why query universes in Module 3 include prompts, not just keywords.
- Pick one question you know cold, ideally in your niche. For example: "how often should you check AI citations?"
- Search it in two AI assistants and one AI Overview if it appears.
- Open the sources each answer credits. Find the exact passage the answer leaned on: usually one to three sentences.
- Note the shared traits: how the passage opens, how it names entities, how it is structured.
Ten minutes of this teaches more about retrieval than any diagram. You are reading the winners.
- Choose one target question for your own site, phrased the way a customer would ask it aloud.
- Find three passages anywhere on the web that answer it directly. Copy them into your log with their URLs.
- Write one line per passage: which trait from the table does it demonstrate best?
- Mark whether your own site currently contains a passage that could compete. Honest answer only; you will fix it in Module 4.
Why does passage quality matter more than page length?
Do we know exactly how each system ranks retrieval candidates?
Why do conversational prompts reward situation-specific content?
Sources used in this lesson
Google Search Central: AI features and your website
Google Search Central: helpful, reliable, people-first content