Crawlability · Explainer

Pagination and SEO

Paginated archives are how most sites expose their deepest content, and how most sites accidentally hide it. Page 2 onward is where products, articles and listings go missing from search. Here is what actually governs whether those pages get crawled.

Why pagination decides what gets found

On a large site, most pages are reachable only through a paginated list, so the health of that list sets the ceiling on what search engines can see.

A category with 40 products across four pages puts 30 of them behind ?page=2 and beyond. If those pages are crawlable and their links are real, the whole catalogue is reachable. If the pagination is broken, blocked, or rendered only after a click, the crawler sees ten products and stops. Nothing about the missing thirty looks broken in a browser, which is why this failure survives so long: the site works perfectly for humans and truncates silently for search engines.

What happened to rel=next and rel=prev

The advice most guides still repeat was retired years ago. Knowing what replaced it matters more than the markup itself.

Google announced in 2019 that it no longer uses rel=next and rel=prev as an indexing signal, and said it had not done so for some time before the announcement. The tags are not harmful and other search engines may still read them, but adding them will not make page 2 rank, and removing them will not cause a drop.

What replaced them is plainer: search engines treat paginated pages as ordinary pages discovered through ordinary links. That shifts the whole question. Instead of asking whether your pagination markup is correct, ask whether each paginated URL is a real, crawlable, linked page that returns 200 and is not excluded by other means. That is a harder standard to meet by accident, and it is where real pagination problems live.

Kalenux still checks rel=next and rel=prev where you use them, because a tag pointing at a broken or redirected URL is evidence that the pagination itself is broken, whatever search engines do with the tag.

The patterns that hide deep pages

Four failures account for most pagination that does not get crawled.

Infinite scroll with no links

When more items load on scroll and there is no anchor pointing at page 2, there is nothing to follow. Googlebot does not scroll or trigger the JavaScript event that loads the next batch, so the content exists in the browser but is reachable only by an interaction a crawler never performs.

The fix is to back the scroll with real paginated URLs: keep the scroll for people, and expose /category?page=2 as a genuine link that loads a genuine page independent of the scroll behaviour. This pattern, sometimes called paginated infinite scroll, gives both audiences a path: people scroll, crawlers follow links.

Canonicalising every page to page 1

Pointing the canonical on page 2, 3 and 4 back at page 1 looks tidy and removes those pages from consideration. Items that appear only on page 3 lose their route into the index.

Each paginated page should be self-canonical. They are not duplicates of each other: they hold different items, which is exactly why they need to exist separately. See what a canonical tag does for the underlying rule.

noindex on page 2 onward

Applying noindex past page 1 is sometimes recommended to keep thin archive pages out of the index. It works, and it also stops those pages being a reliable discovery path for what they list.

If archive pages genuinely add nothing, the better answer is usually fewer, richer categories rather than a blocked trail. noindex explained covers when the directive earns its place.

Pagination behind parameters that are blocked

A Disallow rule written to stop faceted-navigation crawling often catches the pagination parameter too, because both live in the query string.

Check that the rule blocking filter parameters does not also block page. This is a common way a site blocks its own catalogue while intending to save crawl budget: see robots.txt for SEO.

What good pagination looks like

The requirements are unglamorous and they are the whole job.

Every paginated URL returns 200 and renders its items in the HTML. Each page is self-canonical. Each is reachable by a real <a href> from the page before it, so a crawler can walk the sequence without running scripts. Page numbers stay stable enough that a URL crawled today holds broadly the same items tomorrow. And the deepest page is reachable in a sensible number of clicks from a category root.

That last point is where large catalogues fail quietly. A 200-page archive walked one page at a time puts the final items 200 clicks from the entry point, far beyond the depth at which crawling stays thorough. Sub-categories, filters exposed as crawlable pages, or a denser page size all shorten the walk. The goal is not prettier pagination, it is fewer steps between the category root and the item.

Where a "view all" page fits

An alternative to paginating at all, not something layered on top of it.

A single page listing every item removes the pagination problem entirely: there is nothing to discover past page 1 because there is no page 2. When the total item count is modest, this can be the simplest, most robust option, one URL, fully crawlable, nothing to canonicalise or self-reference.

It stops being the right choice once the list gets long enough that the page becomes slow to load, since a page that loads poorly can cost you more in Core Web Vitals and user experience than pagination ever would have. There is no fixed item count where that trade tips, judge it against your own load performance. If you offer both a paginated set and a view-all page for the same content, treat the view-all page as canonical and have the paginated pages either self-canonicalise or canonicalise to it consistently, whichever you decide is the one you want ranking, and apply that decision everywhere rather than mixing approaches across the site.

Pagination questions, answered

Should I still add rel=next and rel=prev to my pagination?

It is no longer an indexing signal, Google confirmed in 2019 that it had stopped using the tags and had done so for some time before saying it publicly. Adding them will not help page 2 rank and removing them will not cause a drop. They are not harmful, and Kalenux still checks them where present, because a rel=next pointing at a broken or redirected URL is a useful signal that the underlying pagination is broken, independent of what search engines do with the tag itself.

How does Google crawl infinite scroll pages?

Googlebot does not scroll or run the interactions a human would to trigger more items loading. If infinite scroll has no equivalent paginated URLs behind it, such as /category?page=2, the items that load only on scroll are invisible to a crawler. The fix is to pair infinite scroll with real, crawlable paginated URLs: keep the scrolling experience for visitors and expose the same content through links a crawler can follow.

Is a "view all" page better for SEO than paginated pages?

A single view-all page that lists every item can work well when the total count is manageable, since it removes pagination entirely and gives crawlers and users one comprehensive page to work with. It stops being a good idea once the list is long enough to hurt page load time, since a slow page can cost you more than pagination ever did. There is no fixed threshold, judge it against your own load performance rather than a rule of thumb.

Should paginated pages (page 2, 3, 4) be noindexed?

Not by default. noindex removes the page from search, but paginated pages are also how a crawler reaches items listed only on page 3 or beyond, so noindexing them can quietly cut off discovery for whatever they contain. It is a reasonable choice for archives that add no unique value beyond page 1, but the better first move is usually improving what those pages contain rather than hiding them.

Should each paginated page have its own canonical or point to page 1?

Each paginated page should carry a self-referencing canonical pointing at itself, not at page 1. Paginated pages are not duplicates of each other, they list different items, and canonicalising them all to page 1 tells Google to treat pages 2 and beyond as copies, which removes them and anything only reachable from them.

Find out whether your page 2 is reachable

Free to start. Checks paginated URLs for broken and redirected targets, and flags pagination that crawlers cannot follow.

Start my free audit