Module 01 · Lesson 3 of 6
Crawlability and Indexability
Crawlability is whether a search engine is allowed to fetch a page. Indexability is whether a fetched page is allowed into the index. They are separate gates with separate controls, and mixing them up wastes hours of debugging.
The four controls
| Control | Gate | What it does | Common mistake |
|---|---|---|---|
| robots.txt | Crawl | Blocks fetching of paths | Blocking CSS and JS, or blocking pages you want indexed |
| XML sitemap | Crawl | Tells the crawler which URLs exist | Listing redirects, 404s or noindex pages |
| noindex directive | Index | Removes or keeps a page out of the index | Hiding it in robots.txt so the directive is never seen |
| rel=canonical | Index | Names the preferred version of duplicates | Pointing every page at the homepage |
That last mistake in the noindex row deserves its own sentence, because it traps real sites constantly: if robots.txt blocks a page, the crawler never reads it, so a noindex directive on that page does nothing. To keep a page out of the index, the crawler must be allowed to fetch it and see the directive.
Soft 404s and duplicates
Two quieter problems deserve attention. A soft 404 is a page that returns a 200 status while behaving like an error: empty results, expired listings, thin placeholders. Engines detect them and treat them as junk. Duplicate content is similar: when multiple URLs carry the same or near-same content, the engine picks one canonical version, and it may not pick yours. Parameter URLs, print versions and www/apex splits are the usual suspects. Fix by consolidating: one canonical URL, 301 redirects for the rest.
What a clean access layer looks like
- Every page worth ranking is reachable through internal links and listed in the XML sitemap.
- robots.txt blocks only what should never be fetched, and never blocks assets the page needs to render.
- noindex is used deliberately for pages you want seen but not ranked, such as internal search results.
- One canonical version of every page, with duplicates redirected or canonicalised.
- Fetch
yourdomain.com/robots.txt. Read every rule and ask which of your important paths it touches. - Open a key page, view source, and search for
noindexand forrel="canonical". Record what you find. - Check the sitemap: does it list the page exactly once, in its canonical form, with a 200 status?
- Search
site:yourdomain.comand spot-check that the pages you care about appear.
Four steps, no tools, and it catches the majority of access failures on small and mid-sized sites.
Audit five URLs on your site: your homepage, two money pages, one blog post, one page you suspect is struggling.
- For each, record: robots.txt status, canonical target, noindex present or not, sitemap listing yes or no,
site:result. - Mark every red flag you find and note which control fixes it.
- Fix the single worst one now, if you can. Access fixes are usually minutes, not days.
A page is blocked in robots.txt and also carries noindex. Will it stay out of the index?
Why can a 200 page still be treated as junk?
What belongs in an XML sitemap?
Sources used in this lesson
Google Search Central: crawling and indexing overview
Google Search Central: consolidating duplicate URLs