Before a page can rank, a search engine has to reach it. robots.txt, crawl budget, and the structural issues that stop bots discovering your content: get this wrong and the rest of your SEO never gets a chance.
User-agent: *Disallow: / ← left over from stagingDisallow: /admin/# Sitemap: (missing)
Each step depends entirely on the one before it. A brilliant page is worth nothing in search if Googlebot can’t reach it.
Disallow rule: often a Disallow: / left over from staging: stops crawlers reaching pages you want indexed. The most common self-inflicted crawl wound.Disallow: /Overlaps with Links, Redirects and Sitemaps: discovery depends on all three.
Crawling happens first: a blocked page never reaches the indexing check.
Blocking stops crawling, not indexing. To remove a page, allow crawling and add noindex.
Look in the Pages and Crawl Stats reports for these four signs.
robots.txt allows fetching; it doesn’t help a crawler find a page with no links in.
It bites on large sites, faceted URLs, and sites wasting crawls on redirects and 404s.
Web crawling is the automated process by which search engines discover and download web pages, and crawlability is how easily a site lets them do so. Google’s main crawler, Googlebot, finds URLs through links and sitemaps, fetches them and, where needed, renders them before anything can be indexed. Access is governed by the Robots Exclusion Protocol, published as an IETF standard in 2022.
Free to start. Find blocked pages, crawl traps and discovery gaps across your site.