AISEOCourse.academy Module 1

Module 01 · Lesson 3 of 6

Crawlability and Indexability

About 13 minutesPrerequisite: Lesson 1.1
After this lesson you canaudit any URL for crawl and index blocks, explain what robots.txt, sitemaps, noindex and canonicals each control, and fix the four most common access problems.

Crawlability is whether a search engine is allowed to fetch a page. Indexability is whether a fetched page is allowed into the index. They are separate gates with separate controls, and mixing them up wastes hours of debugging.

The four controls

ControlGateWhat it doesCommon mistake
robots.txtCrawlBlocks fetching of pathsBlocking CSS and JS, or blocking pages you want indexed
XML sitemapCrawlTells the crawler which URLs existListing redirects, 404s or noindex pages
noindex directiveIndexRemoves or keeps a page out of the indexHiding it in robots.txt so the directive is never seen
rel=canonicalIndexNames the preferred version of duplicatesPointing every page at the homepage

That last mistake in the noindex row deserves its own sentence, because it traps real sites constantly: if robots.txt blocks a page, the crawler never reads it, so a noindex directive on that page does nothing. To keep a page out of the index, the crawler must be allowed to fetch it and see the directive.

DefinitionIndexability is a page's real-world eligibility to appear in search results. A page can be perfectly crawlable and still unindexable because of a directive, a canonical elsewhere, or duplicate content that collapses it into another URL.

Soft 404s and duplicates

Two quieter problems deserve attention. A soft 404 is a page that returns a 200 status while behaving like an error: empty results, expired listings, thin placeholders. Engines detect them and treat them as junk. Duplicate content is similar: when multiple URLs carry the same or near-same content, the engine picks one canonical version, and it may not pick yours. Parameter URLs, print versions and www/apex splits are the usual suspects. Fix by consolidating: one canonical URL, 301 redirects for the rest.

What a clean access layer looks like

Worked example: the two-minute access audit
  1. Fetch yourdomain.com/robots.txt. Read every rule and ask which of your important paths it touches.
  2. Open a key page, view source, and search for noindex and for rel="canonical". Record what you find.
  3. Check the sitemap: does it list the page exactly once, in its canonical form, with a 200 status?
  4. Search site:yourdomain.com and spot-check that the pages you care about appear.

Four steps, no tools, and it catches the majority of access failures on small and mid-sized sites.

Workbench 1.3

Audit five URLs on your site: your homepage, two money pages, one blog post, one page you suspect is struggling.

  1. For each, record: robots.txt status, canonical target, noindex present or not, sitemap listing yes or no, site: result.
  2. Mark every red flag you find and note which control fixes it.
  3. Fix the single worst one now, if you can. Access fixes are usually minutes, not days.
Self-check
A page is blocked in robots.txt and also carries noindex. Will it stay out of the index?
Probably not reliably. The block stops the crawler from fetching the page, so it never sees the noindex directive. To de-index deliberately, allow crawling and let the directive be read.
Why can a 200 page still be treated as junk?
Soft 404s: pages that technically load but deliver nothing, such as empty filters or expired listings. Engines classify them as errors regardless of the status code.
What belongs in an XML sitemap?
Canonical, indexable, 200-status URLs you want discovered and ranked. Redirects, 404s and noindex pages pollute it.
Key principleEvery crawlable URL should be worth indexing: fix the gates before you touch the content.

Sources used in this lesson
Google Search Central: crawling and indexing overview
Google Search Central: consolidating duplicate URLs