Web crawling is the automated process by which search engines discover and download web pages, and crawlability is how easily a site lets them do so. Google’s main crawler, Googlebot, finds URLs through links and sitemaps, fetches them and, where needed, renders them before anything can be indexed [1][3]. Access is governed by the Robots Exclusion Protocol, published as an IETF standard in 2022 [5].
- Crawling is the first of three stages in Google Search: crawl, index, serve [1].
- robots.txt controls crawling, not indexing: a blocked URL can still appear in results [7].
- Most sites don’t need to manage crawl budget; Google’s guidance targets sites with about a million pages or more [9].
- Google only reliably follows <a href> links to resolvable URLs [15].
1Overview
Google describes Search as three stages: crawling, where it downloads text, images and video from pages it has found; indexing, where it analyses that content; and serving, where it returns results for queries [1]. Not every page that is crawled gets indexed [1].
A crawler starts from URLs it already knows, fetches each one, extracts the links it contains and adds new URLs to a queue. Google uses algorithms to decide which sites to crawl, how often and how many pages to fetch from each [1].
2History
The Robots Exclusion Protocol was proposed in 1994 and became the de facto way for site owners to tell crawlers which paths to avoid, but it remained an informal convention for 25 years [6].
In July 2019 Google proposed formalising the protocol as an internet standard and open-sourced the parser it uses to read robots.txt files [6]. The same year Googlebot moved to an “evergreen” version of Chromium, so it renders pages with a current browser engine [14]. The standard was published as RFC 9309 in September 2022 [5].
3Google’s crawlers
Google runs several crawlers. Common crawlers such as Googlebot obey robots.txt; special-case crawlers serve specific products; and user-triggered fetchers act on a user’s request and may ignore robots.txt [2].
Googlebot has two main types, Googlebot Smartphone and Googlebot Desktop. Because Google uses mobile-first indexing, most pages are crawled primarily with the smartphone agent [3].
User-agent strings can be faked. Google recommends verifying a crawler with a reverse DNS lookup to a googlebot.com or google.com hostname, followed by a forward lookup, or by matching published IP ranges [4].
4URL discovery
Most new URLs are found by following links from pages Google already knows. Others come from sitemaps submitted by site owners [1][16].
Google can only reliably follow links written as <a> elements with an href attribute that resolves to a URL. Links that exist only as JavaScript event handlers may not be discovered [15].
Some other search engines, including Bing, also accept instant URL notifications through the IndexNow protocol [19]. Google does not list IndexNow among its supported methods; it relies on links and sitemaps [1][16].
5robots.txt
A robots.txt file sits at the root of a host and tells crawlers which paths they may fetch. Google says it is mainly for managing crawler traffic, and is not a mechanism for keeping pages out of Google [7].
| Field | Purpose |
|---|---|
| user-agent | Which crawler the following rules apply to [8] |
| disallow | A path prefix the crawler should not fetch [8] |
| allow | A path that may be fetched despite a broader disallow [8] |
| sitemap | Absolute URL of a sitemap; not tied to a user-agent [8] |
When several rules match a URL, Google applies the most specific, the longest matching path, and uses allow when an allow and disallow rule are equally specific [8]. Paths are case-sensitive [8].
How failures are handled
| robots.txt response | Google’s behaviour |
|---|---|
| 2XX | Rules are applied [8] |
| 3XX | Up to five redirects followed, then treated as 404 [8] |
| 4XX (except 429) | Treated as if no robots.txt exists: everything may be crawled [8] |
| 5XX or 429 | Crawling of the site is paused; after a prolonged outage Google falls back to a cached copy or assumes no restrictions [8] |
6Crawl budget
Google defines crawl budget as the set of URLs it can and wants to crawl on a site, shaped by two things: the crawl capacity limit, based on how much load the server can handle, and crawl demand, based on how popular and how fresh the content is [9].
Its guidance is written for large or fast-changing sites, roughly a million or more unique pages changing weekly, or ten thousand or more changing daily, and for sites with many pages reported as discovered but not indexed. Most sites don’t need to think about it [9].
- Consolidate duplicate URLs and block low-value infinite spaces such as faceted filters [9][17].
- Return 404 or 410 for permanently removed pages [9].
- Keep sitemaps current and avoid long redirect chains [9].
- Watch response times: a faster, more reliable server raises the crawl capacity limit [9].
Search Console’s Crawl Stats report shows Googlebot’s requests, response codes, file types and average response time for a site [10].
7Controlling crawl rate
To slow Googlebot down urgently, Google recommends temporarily returning 500, 503 or 429 status codes; prolonged errors, though, can lead to URLs being dropped [11][18].
Search Console previously offered a crawl rate limiter setting. Google announced its removal in November 2023, saying its crawling now adjusts to server load automatically [12].
8Rendering JavaScript
Google processes JavaScript pages in three phases: crawling, rendering and indexing. Pages may wait in a queue before rendering, which runs the page in a headless, evergreen Chromium [13][14].
Content and links that only appear after JavaScript runs can be indexed, but they depend on rendering succeeding. Server-side or pre-rendered HTML makes discovery faster and more reliable [13].
9Common crawl problems
| Problem | Effect | Fix |
|---|---|---|
| Faceted filters creating endless URL combinations | Crawl time spent on near-duplicates | Limit crawlable combinations [17] |
| CSS or JavaScript blocked in robots.txt | Google can’t render the page properly | Allow page resources [13] |
| Persistent 5XX errors | Crawl rate falls; URLs may be dropped | Fix server stability [18] |
| JavaScript-only links | Linked pages never discovered | Use <a href> [15] |
| robots.txt returning 5XX | Whole site may stop being crawled | Serve robots.txt reliably [8] |
10Common misconceptions
| Belief | What the sources say |
|---|---|
| Disallowing a page removes it from Google | robots.txt blocks crawling; a linked URL can still be indexed without content. Use noindex instead [7] |
| Every site has a crawl budget problem | Google’s guidance targets very large or fast-changing sites [9] |
| Crawl-delay slows Googlebot | Google ignores crawl-delay [8] |
| A user-agent saying “Googlebot” is Googlebot | User-agents can be spoofed; verify with DNS [4] |
See also
- Crawlability guide: diagrams and quick fixes
- Search indexing: what happens after a page is crawled
- XML sitemaps: the second discovery path
- Core Web Vitals: how server speed affects users and crawl rate
References
- [1]“In-depth guide to how Google Search works”. Google Search Central. developers.google.com/search/docs/fundamentals/how-search-works
- [2]“Overview of Google crawlers and fetchers”. Google Search Central. developers.google.com/search/docs/crawling-indexing/overview-google-crawlers
- [3]“Googlebot”. Google Search Central. developers.google.com/search/docs/crawling-indexing/googlebot
- [4]“Verifying Googlebot and other Google crawlers”. Google Search Central. developers.google.com/search/docs/crawling-indexing/verifying-googlebot
- [5]“RFC 9309: Robots Exclusion Protocol”. IETF, September 2022. www.rfc-editor.org/rfc/rfc9309
- [6]“Formalizing the Robots Exclusion Protocol specification”. Google Search Central Blog, July 2019. developers.google.com/search/blog/2019/07/rep-id
- [7]“Introduction to robots.txt”. Google Search Central. developers.google.com/search/docs/crawling-indexing/robots/intro
- [8]“How Google interprets the robots.txt specification”. Google Search Central. developers.google.com/search/docs/crawling-indexing/robots/robots_txt
- [9]“Crawl budget management for large sites”. Google Search Central. developers.google.com/search/docs/crawling-indexing/large-site-managing-crawl-budget
- [10]“Crawl Stats report”. Search Console Help. support.google.com/webmasters/answer/9679690
- [11]“Reduce the Googlebot crawl rate”. Google Search Central. developers.google.com/search/docs/crawling-indexing/reduce-crawl-rate
- [12]“Saying goodbye to the crawl rate limiter tool in Search Console”. Google Search Central Blog, November 2023. developers.google.com/search/blog/2023/11/sc-crawl-limiter-byebye
- [13]“Understand the JavaScript SEO basics”. Google Search Central. developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics
- [14]“The new evergreen Googlebot”. Google Search Central Blog, May 2019. developers.google.com/search/blog/2019/05/the-new-evergreen-googlebot
- [15]“Link best practices for Google”. Google Search Central. developers.google.com/search/docs/crawling-indexing/links-crawlable
- [16]“Learn about sitemaps”. Google Search Central. developers.google.com/search/docs/crawling-indexing/sitemaps/overview
- [17]“Managing crawling of faceted navigation URLs”. Google Search Central. developers.google.com/search/docs/crawling-indexing/crawling-managing-faceted-navigation
- [18]“How HTTP status codes and network errors affect Google Search”. Google Search Central. developers.google.com/search/docs/crawling-indexing/http-network-errors
- [19]“IndexNow documentation”. IndexNow.org. www.indexnow.org/documentation