Module 01 · Lesson 1 of 6
How Search Engines Work: Crawl, Index, Rank
A search engine works in three stages: it crawls the web, it indexes what it finds, and it ranks indexed pages against queries. Everything in this course, including the AI parts, depends on those three stages running correctly for your site.
Crawling: discovery
Crawling is how a search engine discovers URLs. Automated crawlers follow links, read sitemaps, and fetch pages. If nothing links to a page and it is not in an XML sitemap, the crawler has no way to find it. The crawler also reads robots.txt first: that file can block crawling, and a blocked page stays invisible to the engine no matter how good it is.
Indexing: understanding
Once fetched, the page is analysed. The engine works out what the page is about, which version of duplicate URLs is canonical, and whether the page should be indexed at all. A page carrying a noindex directive, or canonicalising to another URL, will not enter the index. Indexing is the gate: unindexed pages cannot rank, and they cannot be retrieved as sources for AI answers either.
Ranking: choosing
For each query, the engine scores indexed candidates on relevance to the words and intent, authority and trust, and usability signals such as page experience. Ranking is comparative. You are not being graded in isolation; you are being compared with every other indexed candidate for that query, in that moment.
Where AI answers fit
AI Overviews and AI Mode generate answers, but they do not replace this pipeline. Google states that its AI features draw on the same core search systems, that no special schema or technical work is required to be eligible, and that eligibility does not guarantee inclusion. In practice that means the AI layer is a new set of placements sitting on top of the same competition. A site that cannot be crawled and indexed is invisible to all of it.
| Stage | What happens | What goes wrong |
|---|---|---|
| Crawl | Crawler finds URLs via links and sitemaps | No internal links, robots.txt block, dead URLs |
| Index | Page is analysed and stored | noindex, wrong canonical, duplicate content, soft 404 |
| Rank | Indexed pages are scored per query | Wrong intent, weak relevance, thin content, no authority |
The fastest check needs one search. Take a real URL from your own site and search for site:example.com/your-page. If the URL appears, the page is indexed. If it does not, the page is either not crawled or not indexed, and the fix belongs to stage one or two, not to content. Search Console gives you the same answer with more detail: paste the URL into URL Inspection and read the indexing report. Both checks are free and take under a minute.
Pick one page you care about, ideally a page that should bring traffic. Then:
- Run the
site:check for its exact URL. Record indexed or not. - Fetch
yourdomain.com/robots.txtand check nothing blocks the page's path. - If you have Search Console, run URL Inspection and note the stated reason if indexing fails.
- Write one sentence in your log: which stage (crawl, index, rank) is the current bottleneck for this page, and how you know.
Why can a well-written page fail to appear in search?
Does an AI Overview replace the index?
What single check tells you whether a page is indexed?
site: search for the exact URL, or Search Console's URL Inspection report.Sources used in this lesson
Google Search Central: crawling and indexing overview
Google Search Central: AI features and your website