Home/SEO Guide/Reference/Web crawling
Reference article · 19 sources · last reviewed 1 Oct 2026

Web crawling and crawlability

PRACTICAL GUIDEPrefer diagrams and quick fixes? The Web crawling guide covers this visually.Open the guide →

Web crawling is the automated process by which search engines discover and download web pages, and crawlability is how easily a site lets them do so. Google’s main crawler, Googlebot, finds URLs through links and sitemaps, fetches them and, where needed, renders them before anything can be indexed [1][3]. Access is governed by the Robots Exclusion Protocol, published as an IETF standard in 2022 [5].

IN SHORT
  • Crawling is the first of three stages in Google Search: crawl, index, serve [1].
  • robots.txt controls crawling, not indexing: a blocked URL can still appear in results [7].
  • Most sites don’t need to manage crawl budget; Google’s guidance targets sites with about a million pages or more [9].
  • Google only reliably follows <a href> links to resolvable URLs [15].

1Overview

Google describes Search as three stages: crawling, where it downloads text, images and video from pages it has found; indexing, where it analyses that content; and serving, where it returns results for queries [1]. Not every page that is crawled gets indexed [1].

A crawler starts from URLs it already knows, fetches each one, extracts the links it contains and adds new URLs to a queue. Google uses algorithms to decide which sites to crawl, how often and how many pages to fetch from each [1].

KEY FACTGoogle doesn’t accept payment to crawl a site more often [1].

2History

The Robots Exclusion Protocol was proposed in 1994 and became the de facto way for site owners to tell crawlers which paths to avoid, but it remained an informal convention for 25 years [6].

In July 2019 Google proposed formalising the protocol as an internet standard and open-sourced the parser it uses to read robots.txt files [6]. The same year Googlebot moved to an “evergreen” version of Chromium, so it renders pages with a current browser engine [14]. The standard was published as RFC 9309 in September 2022 [5].

YearEvent
1994Robots Exclusion Protocol proposed [6]
2019Google proposes REP as a standard and open-sources its robots.txt parser [6]; Googlebot becomes evergreen [14]
2022RFC 9309 published by the IETF [5]
2023–24Search Console’s crawl rate limiter tool retired [12]

3Google’s crawlers

Google runs several crawlers. Common crawlers such as Googlebot obey robots.txt; special-case crawlers serve specific products; and user-triggered fetchers act on a user’s request and may ignore robots.txt [2].

Googlebot has two main types, Googlebot Smartphone and Googlebot Desktop. Because Google uses mobile-first indexing, most pages are crawled primarily with the smartphone agent [3].

User-agent strings can be faked. Google recommends verifying a crawler with a reverse DNS lookup to a googlebot.com or google.com hostname, followed by a forward lookup, or by matching published IP ranges [4].

KEY FACTGooglebot fetches only the first 15 MB of an HTML or supported text file; anything after that is not considered for indexing [3].

4URL discovery

Most new URLs are found by following links from pages Google already knows. Others come from sitemaps submitted by site owners [1][16].

Google can only reliably follow links written as <a> elements with an href attribute that resolves to a URL. Links that exist only as JavaScript event handlers may not be discovered [15].

Some other search engines, including Bing, also accept instant URL notifications through the IndexNow protocol [19]. Google does not list IndexNow among its supported methods; it relies on links and sitemaps [1][16].

5robots.txt

A robots.txt file sits at the root of a host and tells crawlers which paths they may fetch. Google says it is mainly for managing crawler traffic, and is not a mechanism for keeping pages out of Google [7].

Fields supported by Google [8]. Google ignores crawl-delay [8].
FieldPurpose
user-agentWhich crawler the following rules apply to [8]
disallowA path prefix the crawler should not fetch [8]
allowA path that may be fetched despite a broader disallow [8]
sitemapAbsolute URL of a sitemap; not tied to a user-agent [8]

When several rules match a URL, Google applies the most specific, the longest matching path, and uses allow when an allow and disallow rule are equally specific [8]. Paths are case-sensitive [8].

How failures are handled

robots.txt responseGoogle’s behaviour
2XXRules are applied [8]
3XXUp to five redirects followed, then treated as 404 [8]
4XX (except 429)Treated as if no robots.txt exists: everything may be crawled [8]
5XX or 429Crawling of the site is paused; after a prolonged outage Google falls back to a cached copy or assumes no restrictions [8]
KEY FACTGoogle caches robots.txt for up to 24 hours, and ignores content beyond 500 KiB [8].

6Crawl budget

Google defines crawl budget as the set of URLs it can and wants to crawl on a site, shaped by two things: the crawl capacity limit, based on how much load the server can handle, and crawl demand, based on how popular and how fresh the content is [9].

Its guidance is written for large or fast-changing sites, roughly a million or more unique pages changing weekly, or ten thousand or more changing daily, and for sites with many pages reported as discovered but not indexed. Most sites don’t need to think about it [9].

  • Consolidate duplicate URLs and block low-value infinite spaces such as faceted filters [9][17].
  • Return 404 or 410 for permanently removed pages [9].
  • Keep sitemaps current and avoid long redirect chains [9].
  • Watch response times: a faster, more reliable server raises the crawl capacity limit [9].

Search Console’s Crawl Stats report shows Googlebot’s requests, response codes, file types and average response time for a site [10].

7Controlling crawl rate

To slow Googlebot down urgently, Google recommends temporarily returning 500, 503 or 429 status codes; prolonged errors, though, can lead to URLs being dropped [11][18].

Search Console previously offered a crawl rate limiter setting. Google announced its removal in November 2023, saying its crawling now adjusts to server load automatically [12].

8Rendering JavaScript

Google processes JavaScript pages in three phases: crawling, rendering and indexing. Pages may wait in a queue before rendering, which runs the page in a headless, evergreen Chromium [13][14].

Content and links that only appear after JavaScript runs can be indexed, but they depend on rendering succeeding. Server-side or pre-rendered HTML makes discovery faster and more reliable [13].

9Common crawl problems

ProblemEffectFix
Faceted filters creating endless URL combinationsCrawl time spent on near-duplicatesLimit crawlable combinations [17]
CSS or JavaScript blocked in robots.txtGoogle can’t render the page properlyAllow page resources [13]
Persistent 5XX errorsCrawl rate falls; URLs may be droppedFix server stability [18]
JavaScript-only linksLinked pages never discoveredUse <a href> [15]
robots.txt returning 5XXWhole site may stop being crawledServe robots.txt reliably [8]

10Common misconceptions

BeliefWhat the sources say
Disallowing a page removes it from Googlerobots.txt blocks crawling; a linked URL can still be indexed without content. Use noindex instead [7]
Every site has a crawl budget problemGoogle’s guidance targets very large or fast-changing sites [9]
Crawl-delay slows GooglebotGoogle ignores crawl-delay [8]
A user-agent saying “Googlebot” is GooglebotUser-agents can be spoofed; verify with DNS [4]

See also

References

  1. [1]“In-depth guide to how Google Search works”. Google Search Central. developers.google.com/search/docs/fundamentals/how-search-works
  2. [2]“Overview of Google crawlers and fetchers”. Google Search Central. developers.google.com/search/docs/crawling-indexing/overview-google-crawlers
  3. [3]“Googlebot”. Google Search Central. developers.google.com/search/docs/crawling-indexing/googlebot
  4. [4]“Verifying Googlebot and other Google crawlers”. Google Search Central. developers.google.com/search/docs/crawling-indexing/verifying-googlebot
  5. [5]“RFC 9309: Robots Exclusion Protocol”. IETF, September 2022. www.rfc-editor.org/rfc/rfc9309
  6. [6]“Formalizing the Robots Exclusion Protocol specification”. Google Search Central Blog, July 2019. developers.google.com/search/blog/2019/07/rep-id
  7. [7]“Introduction to robots.txt”. Google Search Central. developers.google.com/search/docs/crawling-indexing/robots/intro
  8. [8]“How Google interprets the robots.txt specification”. Google Search Central. developers.google.com/search/docs/crawling-indexing/robots/robots_txt
  9. [9]“Crawl budget management for large sites”. Google Search Central. developers.google.com/search/docs/crawling-indexing/large-site-managing-crawl-budget
  10. [10]“Crawl Stats report”. Search Console Help. support.google.com/webmasters/answer/9679690
  11. [11]“Reduce the Googlebot crawl rate”. Google Search Central. developers.google.com/search/docs/crawling-indexing/reduce-crawl-rate
  12. [12]“Saying goodbye to the crawl rate limiter tool in Search Console”. Google Search Central Blog, November 2023. developers.google.com/search/blog/2023/11/sc-crawl-limiter-byebye
  13. [13]“Understand the JavaScript SEO basics”. Google Search Central. developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics
  14. [14]“The new evergreen Googlebot”. Google Search Central Blog, May 2019. developers.google.com/search/blog/2019/05/the-new-evergreen-googlebot
  15. [15]“Link best practices for Google”. Google Search Central. developers.google.com/search/docs/crawling-indexing/links-crawlable
  16. [16]“Learn about sitemaps”. Google Search Central. developers.google.com/search/docs/crawling-indexing/sitemaps/overview
  17. [17]“Managing crawling of faceted navigation URLs”. Google Search Central. developers.google.com/search/docs/crawling-indexing/crawling-managing-faceted-navigation
  18. [18]“How HTTP status codes and network errors affect Google Search”. Google Search Central. developers.google.com/search/docs/crawling-indexing/http-network-errors
  19. [19]“IndexNow documentation”. IndexNow.org. www.indexnow.org/documentation
CITE THIS ARTICLEKalenux. (2026). Web crawling and crawlability. In Kalenux SEO Guide. Last reviewed 1 Oct 2026. https://kalenux.com.tr/reference/web-crawling