AISEOCourse.academy Module 2

Module 02 · Lesson 2 of 4

How Language Models Retrieve and Choose Sources

About 12 minutesPrerequisite: Lesson 2.1
After this lesson you canexplain in plain terms how a generated answer gets its material, why retrieval works at passage level, and what writing traits make a passage easy to fetch and use.

Retrieval is the step where a language model fetches candidate material before it writes a word. Modern answer systems are built as retrieval plus generation: a retrieval layer finds passages relevant to the question, then the language model composes an answer drawing on what was fetched. The generation gets the attention; the retrieval decides the outcome.

DefinitionRetrieval-augmented generation (RAG) is the pattern of fetching relevant documents or passages first and generating the answer from them. The answer can only lean on what retrieval found, so being findable and fetchable is the whole game.

Passages, not pages

Classic ranking scores a URL against a query. Retrieval for generation is finer-grained: it hunts for passages that address the question, wherever they sit. Two consequences follow. A weak page can win on the strength of one excellent passage. And a strong page whose answer is buried under three paragraphs of throat-clearing can lose to a plainer page that answers in the second sentence.

The exact ranking mechanics inside any given system are not public. What you can observe, repeatedly, is what sources get used: pages that answer the actual question directly, in structured, self-contained chunks, with named entities and supporting evidence. Work to the observable behaviour and label your inference as inference.

What a fetchable passage looks like

TraitFetchable versionUnfetchable version
Opening"Query fan-out is the practice of issuing related searches across subtopics.""In this section we will explore an interesting concept..."
Entity namingSays "Google AI Overviews" every time it means it"this innovative feature", "the platform"
StructureQuestion heading, direct answer, then detailNarrative that reveals the answer in paragraph four
IndependencePassage makes sense quoted aloneDepends on the previous page to mean anything

You will practice writing these in Module 4. For now the point is recognition: start seeing pages as collections of passages in competition with other passages.

Conversational prompts change the fetch

A prompt like "I run a bakery in Leeds, how do I get mentioned in AI answers?" carries context (a bakery, a city, a goal) that a bare keyword never would. Retrieval for that prompt hunts for material matching the situation, not just the topic words. Content that speaks to specific situations ("for local businesses", "for beginners", "for agencies") gives retrieval more handles to grab. That is also why query universes in Module 3 include prompts, not just keywords.

Worked example: collect the passages that win
  1. Pick one question you know cold, ideally in your niche. For example: "how often should you check AI citations?"
  2. Search it in two AI assistants and one AI Overview if it appears.
  3. Open the sources each answer credits. Find the exact passage the answer leaned on: usually one to three sentences.
  4. Note the shared traits: how the passage opens, how it names entities, how it is structured.

Ten minutes of this teaches more about retrieval than any diagram. You are reading the winners.

Workbench 2.2
  1. Choose one target question for your own site, phrased the way a customer would ask it aloud.
  2. Find three passages anywhere on the web that answer it directly. Copy them into your log with their URLs.
  3. Write one line per passage: which trait from the table does it demonstrate best?
  4. Mark whether your own site currently contains a passage that could compete. Honest answer only; you will fix it in Module 4.
Self-check
Why does passage quality matter more than page length?
Retrieval fetches passages, not pages. One excellent self-contained passage can carry a modest page, while a long page with buried answers gives retrieval nothing clean to grab.
Do we know exactly how each system ranks retrieval candidates?
No, and any course claiming otherwise is guessing. We work from published platform statements and observed behaviour, and we label inference as inference.
Why do conversational prompts reward situation-specific content?
Prompts carry context (who, where, why). Content that addresses named situations gives retrieval more matching handles than generic topic text.
Key principleRetrieval selects passages, not websites: write the passage you would want fetched.

Sources used in this lesson
Google Search Central: AI features and your website
Google Search Central: helpful, reliable, people-first content