Home/SEO Guide/Indexability
SEO guide · category 02 · foundation

Indexability & canonicalsIs Google allowed to index it?

A page can only rank if Google is allowed to index it, and if the right URL is marked as the master copy. Indexability bugs are the most damaging in SEO because they make good content invisible.

Last reviewed · 1 Oct 2026Reference article · 17 sources →Kalenux engineering
/pricing/ · indexable?✕ not indexable
1No noindexno meta noindex, no X-Robots-Tag✓
2Crawlablenot blocked in robots.txt✓
3Self-canonicalpoints to itself✕ → /pricing?ref=nav
4Returns 200not 3XX / 4XX / 5XX✓
all four must hold at once · illustrative
why indexability comes before everything else

A gate, not a scale.

These failures are silent: the page loads for you, ranks nowhere, and there’s no error: just a quiet absence. Miss any one of four and it won’t appear.

⊘1 / 4
No noindexNo noindex meta tag and no X-Robots-Tag: noindex header.
↓2 / 4
CrawlableNot blocked in robots.txt, so Google can fetch and read the page.
⟲3 / 4
Self-canonicalA canonical that points to itself, or none sending Google elsewhere.
2004 / 4
Returns 200A successful HTTP 200, not a redirect, a 4XX or a 5XX.
WHY IT HURTSOne line in a shared template can remove a whole section, and because traffic lags crawling by days, the drop is noticed well after the change.
how the two failure modes differ

Same result, different route, and different fix.

noindex and a broken canonical both end with a page missing from search, but the fix isn’t interchangeable.

noindexCanonical to another URL
What it tells GoogleDon’t put this page in the indexThis page is a copy: credit the other URL
Still crawled?Yes: crawling and indexing are separateYes
Where signals end upNowhere: the page is excludedConsolidated onto the canonical target
Right tool whenYou want the page out of search entirelyTwo URLs are duplicates and one should rank
Combining the two is a conflict: noindex says remove, canonical says merge.
read the indexability explainers

The two checks that deindex more pages than anything else.

Downstream of crawlability, a page must be reachable first, and upstream of sitemaps, where listing a non-indexable URL asks Google to recrawl something it will reject.
the indexability checks we run

18 checks, on every page.

noindex5 checks
✓noindex pages with a robots directive✓X-Robots-Tag: noindex in the header✓noindex-follow / noindex-nofollow variants✓Newly non-indexable since last crawl✓noindex blocked by robots.txt
Canonicals9 checks
✓Missing canonical on an indexable page✓Multiple canonical tags✓Canonical changed since last crawl✓Canonical to a redirect, 4XX or 5XX✓Canonical to a non-indexable target✓Canonical chains (A → B → C)✓Off-domain canonical✓Protocol-crossing canonical✓Canonical with no internal links in
Conflicts4 checks
✓Non-canonical URL getting organic traffic✓Sitemap lists a URL that canonicalises elsewhere✓noindex page listed in the sitemap✓Trailing-slash mismatch
indexability in practice

Four situations, each with its rule of thumb.

01Why pages suddenly drop outA stray noindex on a shared template, a canonical to the wrong URL, a robots.txt block, or new 4XX/5XX after a migration. All ship silently.Flag newly non-indexable pages the day they ship
02When “not indexed” is correctThank-you pages, internal search results and thin filter pages should stay out. Excluded ≠ broken.Only flag pages you want ranking
03noindex vs robots.txtTo remove a page, use noindex and allow crawling. A robots.txt block stops Google seeing the noindex.robots.txt = crawling · noindex = indexing
04Canonical and sitemap conflictsA sitemap says “index this”; a canonical elsewhere contradicts it. Over time Google trusts the sitemap less.Sitemap, canonical and links all agree
indexability questions, answered

Three questions, answered at a glance.

Q1NO

Does a canonical tag do the same job as noindex?

noindexRemove
canonicalMerge

No. A canonical tells search engines a page is a duplicate and to consolidate ranking signals onto another URL, while the page can still be crawled and even shown for some queries.

Q2IT CONTRADICTS

Why does a noindex page in the sitemap matter?

sitemap.xml → “index /old/”/old/ → noindex↳ contradiction · wasted crawl

A sitemap is a claim that a URL is worth indexing. Listing a noindex page in it contradicts the page’s own directive and wastes part of the crawl budget your sitemap is meant to direct efficiently..

Q3YES

Can robots.txt stop a noindex tag from working?

Disallow: /x/→noindex never read

Yes. A crawler has to fetch a page to read its noindex meta tag or X-Robots-Tag header.

Reference article · 17 sources

noindex, canonicals and status codes: documented in full.

Indexing is the stage in which a search engine analyses a crawled page and decides whether to store it in its index, the database it draws results from. A page is indexable when nothing prevents that: it returns a successful status code, carries no noindex directive, and is chosen as the canonical URL among any duplicates.

KALENUX REFERENCESearch indexing, noindex and canonical URLs
  1. 1Overview
  2. 2Status codes
  3. 3The noindex directive
  4. 4Canonicalization
  5. 5Canonical best practice
  6. 6Mobile-first indexing
  7. 7Diagnosing indexing
  8. 8Common misconceptions
17 references · cite-readyRead the article →
next in the guide

Make sure your best pages can actually be indexed.

Free to start. Catch stray noindex tags and broken canonicals before they cost you rankings.

Start my free audit →