Home/SEO Guide/Crawlability
SEO guide · category 01 · foundation

Crawlability for SEOCan search engines reach your pages?

Before a page can rank, a search engine has to reach it. robots.txt, crawl budget, and the structural issues that stop bots discovering your content: get this wrong and the rest of your SEO never gets a chance.

Last reviewed · 1 Oct 2026Reference article · 19 sources →Kalenux engineering
robots.txt · yourdomain.com1 critical rule
User-agent: *Disallow: / ← left over from stagingDisallow: /admin/# Sitemap: (missing)
Pages blocked1,284
Sitemap declaredno
CSS / JS blockedno
illustrative
why crawlability is the foundation

Crawl, index, rank: in that order.

Each step depends entirely on the one before it. A brilliant page is worth nothing in search if Googlebot can’t reach it.

1THIS CATEGORY
CrawlA bot fetches the page. Blocked or undiscoverable? It stops here: nothing downstream happens.robots.txt · links · server
2INDEXABILITY
IndexThe page is analysed and stored. A noindex or bad canonical removes it here.noindex · canonicalIndexability →
3EVERYTHING ELSE
RankOnly indexed pages compete. Content, links and performance finally pay off.content · links · speed
most crawlability problems trace back to

Three causes, one of them is usually yours.

⊘Blocked by robots.txt
An over-broad Disallow rule: often a Disallow: / left over from staging: stops crawlers reaching pages you want indexed. The most common self-inflicted crawl wound.Disallow: /
◌No path to discover
Pages with no internal links (orphans) sit outside the crawl. Bots find pages by following links; if nothing links to a page, it may never be found.0 inbound linksRead the explainer →
∞Crawl traps
Redirect chains, infinite calendars and endless parameter URLs waste a bot’s limited time on junk instead of your real content./cal?m=2031-04…Read the explainer →
read the crawlability explainers

Four deep dives, the ones people get wrong most.

the crawlability checks we run

13 checks, grounded in the real crawler.

Overlaps with Links, Redirects and Sitemaps: discovery depends on all three.

robots.txt7 checks
✓robots.txt missing✓robots.txt blocks the sitemap✓robots.txt blocks indexable pages✓robots.txt blocks CSS or JS needed to render✓No sitemap declared in robots.txt✓noindex page blocked by robots.txt✓Sitemap URL blocked by Disallow
Structure4 checks
✓Broken or orphaned rel=next / prev targets✓Redirect chains and loops✓Orphan pages with no path for discovery✓Pages more than four clicks from home
URL consistency2 checks
✓www and non-www both resolving✓Trailing-slash mismatch
WITH A SERVER LOG CONNECTEDReal Googlebot request patterns: wasted crawl activity, hits that returned 4XX/5XX or a redirect, and spoofed bots claiming to be Googlebot from outside Google’s published IP ranges.
crawlability in practice

The three questions that come up most.

01robots.txt can’t remove a page from GoogleBlocking stops crawling, not indexing, and it stops Google seeing a noindex you add later. To remove a page: allow crawling, add noindex.
robots.txt → crawlingnoindex → indexing
02When crawl budget is worth worrying aboutRarely for a few thousand URLs. It bites on large sites, faceted or parameter URLs, and sites wasting crawls on redirects and 404s.
Small site → fineLarge / faceted → matters
03Checking whether a page is being crawledURL Inspection for one page; Crawl Stats for site-wide patterns. “Discovered, currently not indexed” usually points to discovery or budget, not content.
URL InspectionCrawl Stats
crawlability questions, answered

Five questions, answered at a glance.

Q1TWO DIFFERENT GATES

Crawlability vs indexability?

1 · CRAWLCan it be fetched?robots.txt · links · server
→
2 · INDEXIs it allowed in?noindex · canonical

Crawling happens first: a blocked page never reaches the indexing check.

Q2NO

Does robots.txt remove a page from Google?

Disallow: /old-page/✕ still indexed
yourdomain.com/old-page/No information is available for this page.
Allow + noindex✓ removed

Blocking stops crawling, not indexing. To remove a page, allow crawling and add noindex.

Q3SEARCH CONSOLE

How do I spot a crawl problem?

Discovered – not indexedCrawled – not indexedURLs vs indexed gapOld last-crawl date

Look in the Pages and Crawl Stats reports for these four signs.

Q4YES

Can orphan pages hurt crawlability?

HomeBlogPostPostOrphan · 0 links in

robots.txt allows fetching; it doesn’t help a crawler find a page with no links in.

Q5RARELY

Does crawl budget affect small sites?

A few thousand URLsrarely matters
Large or faceted sitematters

It bites on large sites, faceted URLs, and sites wasting crawls on redirects and 404s.

Reference article · 19 sources

How crawling really works, sourced line by line.

Web crawling is the automated process by which search engines discover and download web pages, and crawlability is how easily a site lets them do so. Google’s main crawler, Googlebot, finds URLs through links and sitemaps, fetches them and, where needed, renders them before anything can be indexed. Access is governed by the Robots Exclusion Protocol, published as an IETF standard in 2022.

KALENUX REFERENCEWeb crawling and crawlability
  1. 1Overview
  2. 2History
  3. 3Google’s crawlers
  4. 4URL discovery
  5. 5robots.txt
  6. 6Crawl budget
  7. 7Controlling crawl rate
  8. 8Rendering JavaScript
  9. 9Common crawl problems
  10. 10Common misconceptions
19 references · cite-readyRead the article →
next in the guide

Make sure search engines can reach every page.

Free to start. Find blocked pages, crawl traps and discovery gaps across your site.

Start my free audit →