Indexing is the stage in which a search engine analyses a crawled page and decides whether to store it in its index, the database it draws results from. A page is indexable when nothing prevents that: it returns a successful status code, carries no noindex directive, and is chosen as the canonical URL among any duplicates [1][2][5].
- Only pages returning 2XX are considered for indexing [9].
- noindex removes a page from results, but only if Google can crawl the page to see it [2].
- A canonical tag is a strong hint, not a command: Google may choose another URL [5].
- Search Console’s Page indexing report lists every excluded URL and the reason [10].
1Overview
After crawling a page, Google processes its text, tags and attributes, and media to understand what it is about. During this stage it also decides whether the page is a duplicate of another and which version is canonical. The result may be stored in the Google index, a large database spread across many computers [1].
Indexing is not guaranteed. Google says not every page it processes will be indexed, and that whether a page is indexed depends on its content and metadata [1].
To be eligible at all, a page must meet Google’s technical requirements: Googlebot isn’t blocked, the page works and returns HTTP 200, and it has indexable content [17].
2Status codes
The HTTP status code a URL returns sets an upper limit on whether it can be indexed.
| Response | Effect on indexing |
|---|---|
| 2XX success | Content may be considered for indexing [9] |
| 3XX redirect | Google follows the redirect; the target is considered instead [9][16] |
| 4XX client error | Content is not indexed; already-indexed URLs are gradually removed [9] |
| 5XX server error / 429 | Crawling slows; persistent errors lead to removal [9] |
| 200 with error content | Treated as a soft 404 and generally not indexed [9] |
3The noindex directive
A noindex rule tells search engines not to include a page in their results. It can be set in an HTML robots meta tag or in an X-Robots-Tag HTTP response header; the header form also works for non-HTML files such as PDFs and images [2][3].
<meta name="robots" content="noindex"> applies to all crawlers that support it, while <meta name="googlebot" content="noindex"> targets Google only [3].
When Googlebot next crawls the page and sees the rule, Google drops it from results, regardless of whether other sites link to it [2].
noindex and robots.txt
robots.txt controls crawling, not indexing. A URL disallowed in robots.txt can still be indexed if linked from elsewhere, usually shown without a description [4]. In July 2019 Google announced it would stop supporting noindex rules written inside robots.txt, which had never been part of the protocol [13].
4Canonicalization
When several URLs show the same or very similar content, Google groups them and selects one as canonical: the version it crawls most regularly and shows in results. The others are treated as duplicates [5].
Duplicates arise for many reasons: http and https versions, www and non-www hosts, tracking parameters, sorting and filtering options, separate mobile URLs, and printer-friendly pages [5].
Signals Google uses
| Method | Strength | Notes |
|---|---|---|
| Redirect (301/308) | Strong | Use when the duplicate should no longer be reached [6][16] |
| rel="canonical" link | Strong | In HTML or an HTTP Link header [6] |
| Sitemap inclusion | Weak | List only canonical URLs [6] |
Google introduced support for the rel="canonical" link element in February 2009 [8]. It was later standardised by the IETF as RFC 6596 in 2012 [7].
5Canonical best practice
- Use absolute URLs, and point to a page that returns 200, not a redirect or an error [6].
- Don’t use robots.txt for canonicalisation, and don’t use the URL removal tool to choose canonicals [6].
- Don’t specify different canonicals for the same page through different methods [6].
- In hreflang clusters, each language version should self-canonicalise, not point to another language [15].
- Keep internal links, sitemaps and canonicals pointing at the same URL [6].
6Mobile-first indexing
Google predominantly uses the mobile version of a page’s content, crawled with the smartphone agent, for indexing and ranking [14].
Content, structured data, metadata and directives such as noindex should therefore be the same on mobile and desktop. Material present only on desktop may not be indexed [14].
7Diagnosing indexing
Search Console’s Page indexing report shows how many of a site’s URLs are indexed and groups excluded URLs by reason, such as “Excluded by ‘noindex’ tag”, “Alternate page with proper canonical tag” or “Crawled: currently not indexed” [10].
The URL Inspection tool shows Google’s indexed version of a single URL, including the user-declared canonical and the Google-selected canonical, and can test the live page [11].
For urgent cases, the Removals tool hides a URL from results for about six months while a permanent fix such as noindex or deletion takes effect [12].
8Common misconceptions
| Belief | What the sources say |
|---|---|
| robots.txt keeps a page out of Google | It blocks crawling only; use noindex [4][2] |
| A canonical tag is a directive | It is a strong hint Google can override [5] |
| Removals deletes pages from the index | It hides them temporarily [12] |
| Every crawled page gets indexed | Indexing is not guaranteed [1] |
See also
- Indexability guide: diagrams and quick fixes
- Web crawling: the stage before indexing
- HTTP redirects: the strongest canonical signal
- Indexability Checker: check one URL instantly
References
- [1]“In-depth guide to how Google Search works”. Google Search Central. developers.google.com/search/docs/fundamentals/how-search-works
- [2]“Block Search indexing with noindex”. Google Search Central. developers.google.com/search/docs/crawling-indexing/block-indexing
- [3]“Robots meta tag, data-nosnippet and X-Robots-Tag specifications”. Google Search Central. developers.google.com/search/docs/crawling-indexing/robots-meta-tag
- [4]“Introduction to robots.txt”. Google Search Central. developers.google.com/search/docs/crawling-indexing/robots/intro
- [5]“What is URL canonicalization”. Google Search Central. developers.google.com/search/docs/crawling-indexing/canonicalization
- [6]“How to specify a canonical URL with rel="canonical" and other methods”. Google Search Central. developers.google.com/search/docs/crawling-indexing/consolidate-duplicate-urls
- [7]“Canonical Link Relation (RFC 6596)”. IETF, April 2012. www.rfc-editor.org/rfc/rfc6596
- [8]“Specify your canonical”. Google Search Central Blog, February 2009. developers.google.com/search/blog/2009/02/specify-your-canonical
- [9]“How HTTP status codes and network errors affect Google Search”. Google Search Central. developers.google.com/search/docs/crawling-indexing/http-network-errors
- [10]“Page indexing report”. Search Console Help. support.google.com/webmasters/answer/7440203
- [11]“URL Inspection tool”. Search Console Help. support.google.com/webmasters/answer/9012289
- [12]“Removals and SafeSearch reports tool”. Search Console Help. support.google.com/webmasters/answer/9689846
- [13]“A note on unsupported rules in robots.txt”. Google Search Central Blog, July 2019. developers.google.com/search/blog/2019/07/a-note-on-unsupported-rules-in-robotstxt
- [14]“Mobile-first indexing best practices”. Google Search Central. developers.google.com/search/docs/crawling-indexing/mobile/mobile-sites-mobile-first-indexing
- [15]“Tell Google about localized versions of your page”. Google Search Central. developers.google.com/search/docs/specialty/international/localized-versions
- [16]“Redirects and Google Search”. Google Search Central. developers.google.com/search/docs/crawling-indexing/301-redirects
- [17]“Technical requirements (Search Essentials)”. Google Search Central. developers.google.com/search/docs/essentials/technical