noindex is the right way to keep a page out of search, and one of the most dangerous tags in SEO when it ends up where it shouldn't. Here's noindex explained: when to use it, when it silently kills rankings, and how to catch a stray one fast.
noindex tells search engines: crawl this page if you like, but don't put it in search results. It's a precise tool for excluding pages from the index without hiding them from users.
You can apply noindex two ways. The common one is a meta robots tag in the page <head>; the other is an HTTP response header, which works for non-HTML files like PDFs where you can't add a meta tag:
Either way, search engines that crawl the page will drop it from their index. The critical catch: the page must be crawlable for noindex to work. If it's blocked in robots.txt, Google never fetches the page, never sees the tag, and the URL can stay indexed. Use robots.txt for crawling and noindex for indexing; they are not interchangeable, and combining them backfires.
The line between a smart noindex and a costly one.
Thank-you and order-confirmation pages, internal search-results pages, thin tag/filter/pagination pages that add no unique value, staging or duplicate utility pages, and private account or admin areas. These are pages users may need but that would clutter or dilute your presence in search. noindex keeps them accessible while out of the index.
If two URLs are duplicates and you want their ranking signals combined onto one, use a canonical tag, not noindex. noindex removes a page entirely; canonical consolidates it. Mixing them, a noindex plus a canonical to another URL: sends conflicting signals.
This sounds obvious, but it's the single most common SEO disaster, because noindex usually arrives by accident via a template or global setting, not a deliberate page-level choice. The next section covers exactly how that happens.
| Meta robots tag | X-Robots-Tag header | |
|---|---|---|
| Where it lives | Inside the page's <head> | An HTTP response header, set by the server |
| Works on | HTML pages only | Any file type: HTML, PDF, images, XML feeds |
| How you'd typically set it | CMS template or page-level meta field | Server/CDN config, applies to a whole path in one rule |
| Easiest to audit visually | Yes, visible in page source | No, requires a header check (curl -I) |
| Needs the page crawlable to work | Yes | Yes |
How a single tag deindexes a whole site, and the check that stops it.
Most sites are built with shared templates. If a noindex tag lands on a template, a staging-environment "noindex everything" rule that ships to production, a CMS visibility toggle left on, a plugin default, or a developer's leftover, it doesn't deindex one page. It deindexes every page using that template at once, a whole blog, a whole product category, sometimes the entire site. The cruel part is the delay. Nothing breaks visibly; the pages load fine. Then over the following days, as Google re-crawls and honours the tag, rankings and traffic collapse, and by the time someone notices the traffic drop, the cause is buried days in the past.
The defence is monitoring for change. An audit that flags pages which became non-indexable since the last crawl turns a silent, delayed catastrophe into an alert the same day the tag ships. Run a crawl after every deploy that touches templates, and the stray noindex never gets the chance to do real damage.
Both are checks the audit runs directly, because both are common enough to be worth naming.
A page marked noindex should not also carry a canonical pointing somewhere else: noindex says remove this page, a canonical elsewhere says this page is a duplicate whose signals belong on another URL. The two directives disagree about what should happen to the page, and Google resolves the conflict on its own terms. The audit flags a canonical target that is itself noindex specifically because that combination leaves an entire cluster with no indexable home at all.
Applying noindex to page 2 onward of an archive is sometimes suggested as a way to keep thin listing pages out of search. It works, and it also removes those pages as a discovery path, since a crawler reaches page 3's items by following page 2's link. If the underlying catalogue depends on pagination to be fully reachable, this quietly caps how much of it Google can ever see. See pagination and SEO for the fuller picture, including what replaced rel=next/prev as the relevant signal.
These two look interchangeable and are not. noindex keeps a page out of the index while still allowing crawling, which is exactly what you want when the goal is removal from search. A robots.txt Disallow does the opposite: it blocks crawling but does not remove a page from the index, and worse, it stops Google ever seeing a noindex tag you add later.
The practical rule is to allow crawling and add noindex when you want a page out of search, and to reserve robots.txt for managing crawl load. Applying both to the same URL is the combination that leaves pages stubbornly indexed with no description attached.
noindex,follow keeps the page out of search but still lets link equity flow through its outbound links. That is useful for pagination and utility pages you want excluded from results while still passing authority onward to the pages they link to.
noindex,nofollow excludes the page and also stops search engines following its links, cutting off that flow entirely. The audit distinguishes between the two variants precisely because choosing the wrong one can silently sever internal link flow through a whole section of a site.
A noindex directive applies the next time Google crawls the page and processes the tag, which can be anywhere from hours to weeks depending on how often that page is crawled. To speed up removal of an important page, request indexing or use the removal tool in Search Console.
That same delay is what makes an accidental noindex so dangerous. The mistake ships on one day and the damage lands several days later, by which time the deploy that caused it is well behind you and no longer the obvious suspect.
Use the meta robots tag for HTML pages, it lives in the head and is the standard, easiest-to-audit method. Use the X-Robots-Tag HTTP header for anything that has no HTML head to put a tag in: PDFs, images, XML feeds, or when you need to apply noindex to a whole path at the server or CDN level in one rule instead of editing every template. Both are read the same way by crawlers; the header is simply the only option when there is no markup to place a tag in.
No. robots.txt only controls crawling with Disallow and Allow rules, it has no noindex directive. Google briefly supported an unofficial noindex line in robots.txt and stopped honouring it in September 2019. noindex only works as a meta tag in the page head or as an X-Robots-Tag HTTP header, and either one requires the page to be crawlable so the directive can be read.
The two directives contradict each other. noindex says remove this page from the index; a canonical to a different URL says this page is a duplicate, credit the other one instead. Google resolves the conflict on its own terms rather than yours, and the outcome is unpredictable. Pick one job for the page: noindex it if you want it gone, or canonicalise it if you want its signals merged elsewhere, never both.
It keeps thin archive pages out of the index, but it also removes them as a discovery path, since paginated pages are how crawlers reach items listed only on page 3 or beyond. If a category depends on pagination to expose its full catalogue, noindexing page 2 onward can quietly cut off crawl access to everything past page 1. Reserve it for archives that genuinely add nothing beyond page 1, and prefer fixing thin pagination with better page sizes or richer categories first.
Search Console's URL Inspection tool reports the indexing status and the exact reason a URL is excluded, including a noindex detected in the meta tag or in the X-Robots-Tag header. A quick manual check is curl -I on the URL to see the X-Robots-Tag header directly, plus viewing the page source for a meta robots tag, since the two are set independently and a page can carry one without the other.
No. Removing the tag only takes effect once Google re-crawls the page and reads the change, which depends on how often that URL is crawled and can take anywhere from hours to weeks. Requesting indexing through Search Console's URL Inspection tool prompts a faster recrawl for a single important URL, but there is no way to force instant re-inclusion at scale.
Free to start. Get alerted when any page becomes non-indexable since your last crawl.
Start my free audit