Content · Explainer

Thin content

Thin content is a page that adds too little value to earn a place in search. It is rarely about word count alone. Here is what actually signals it, why a pile of thin pages drags down the pages you care about, and whether to improve, merge, or remove.

What is thin content?

Thin content is a page that offers little unique value to the person who lands on it. Word count is one hint, but a 200-word answer that fully resolves a query is not thin, and a 2,000-word page that says nothing is.

Google's concern is value per page, not length. A page is thin when a searcher arrives, finds nothing they could not get faster elsewhere, and leaves. The classic cases are auto-generated pages, doorway pages built only to rank, near-empty category or tag pages, scraped or spun text, and pages that exist to hold an ad rather than to answer anything. What they share is that a real person gets little out of them.

Because "value" is hard to measure directly, an audit looks at the signals that tend to travel with thin pages, and flags a page when several stack up at once rather than relying on any single one. That is what the "thin content composite" finding is: not a word-count trip-wire, but a cluster of weak signals appearing together.

The signals that stack up

One of these alone is often fine. Several together is the pattern that flags a page.

Low word count

Very little body text, or an "empty" page that returns 200 but has almost no words, so there is not enough substance to answer anything.

Too few headings

A long page with no structure: no headings breaking it into sections, which usually means it is not actually covering the subject in depth.

Shallow paragraph depth

Content that skims the surface, thin paragraphs that raise a point and move on without developing it, so the page never resolves the query it targets.

Templatized skeleton

The page follows a skeleton shared by many others with only a few words swapped in, so it is near-duplicate filler rather than a distinct answer.

Why thin pages hurt the whole site

The damage is not limited to the thin page itself.

A single thin page just does not rank, which on its own is harmless. The problem is scale. When a site has hundreds or thousands of thin, templated pages, Google's assessment of overall quality drops, and that can weigh on pages that are genuinely good. This is the core lesson behind the principle that one useful page beats a thousand near-empty ones: mass-produced thin content does not add a long tail of traffic, it adds a drag on the whole domain.

It also wastes crawl budget. Search engines spend a finite amount of effort crawling any site, and every thin page they fetch is effort not spent on the pages you want indexed and refreshed. Pruning thin content often improves how quickly your important pages get crawled and updated, an effect covered in the crawl budget guide.

How do you check a site for thin content?

Thin content is a judgement about value, so no tool can hand you a final answer. What a crawl can do is narrow thousands of URLs down to a shortlist worth a human look, and give you the evidence to decide each one quickly.

Step 1: pull the signals, not the verdict

Crawl the site and export, for every indexable URL, the main-content word count, the number of headings, the count of internal links pointing at the page, and a similarity score against the rest of the site. Word count alone is the trap everyone falls into. Sorting ascending by word count gives you a list topped by pages that are perfectly fine: a contact page, a pricing table, a well-made tool page whose value is the tool rather than the prose. Those are short, not thin.

The signal that actually discriminates is several weak indicators appearing on the same URL. A page with 180 words, no subheadings, two internal links and 80 percent template similarity to forty other URLs is thin in a way a 180-word contact page is not. Filter for the intersection rather than any single column.

Step 2: group by pattern before you judge pages

Thin pages are usually produced, not written, so they arrive in families. Sort the shortlist by URL pattern and you will typically find the whole problem lives in two or three directories: /tag/, /author/, a filter parameter, a location-page generator that spun up one URL per town. Judging the pattern is far faster than judging four hundred individual pages, and the fix is different in kind: you change the generator or noindex the pattern once, rather than editing pages one at a time.

This is also where you catch the pages that should never have been crawlable. Internal search result pages are the classic case. They are generated on demand, infinite in number, and thin by construction, and they belong behind a robots rule rather than in a content review.

Step 3: apply the three-way test per page

For each page or pattern that survives, ask one question in three parts. Does a real search demand exist for this topic? Can this page be made the best available answer to it? And is anyone with the time and knowledge going to actually do that work?

Three yeses means improve. A yes on demand but a no on capacity means merge into a stronger page that already covers the ground. A no on demand means remove, regardless of how much effort went into creating the page originally. The sunk cost is the most common reason thin directories survive audits: nobody wants to delete work, so the pages sit there indefinitely, and the site carries the drag.

What the numbers do and do not tell you

Treat every threshold as a prompt to look, never as a verdict. A crawler flagging "low word count" is telling you where to point your attention, not that the page is bad. The reverse error matters more: pages that clear every numeric threshold can still be thin, because a page can be long, well-structured, heavily linked and still say nothing a reader could not get faster elsewhere. No crawler catches that. A person reading the page catches it in about fifteen seconds.

The practical discipline is to use the crawl to build the shortlist and to spend your human attention on the shortlist rather than on the whole site. A few hundred pages reviewed properly beats tens of thousands sorted by a metric that was never measuring value in the first place.

Is there a thin content penalty?

Almost never, and the distinction matters because it changes what you should do about it. Thin content overwhelmingly costs you rankings algorithmically rather than through any punishment applied to your site.

Algorithmic suppression is not a penalty

A penalty in the strict sense is a manual action: a human reviewer at Google looks at your site, decides it violates the spam policies, and applies a sanction. You are told when this happens. It appears in the Manual Actions report in Search Console, with a description and a route to request reconsideration. If that report is empty, you do not have a penalty, whatever your traffic is doing.

What thin content usually causes is different. Google's ranking systems assess helpfulness continuously, and pages judged unhelpful simply rank below pages judged more helpful. Nothing is applied to you; you are just outranked. The effect looks identical from the traffic graph, which is why the two get confused, but the remedy is not the same.

Why the difference changes your response

If you believe you have been penalised, the instinct is to find the offending pages, remove them, and file a reconsideration request. With a manual action that is exactly right. With algorithmic assessment there is nothing to file, no one reviewing, and no switch to be flipped back.

Recovery is also shaped differently. A lifted manual action can restore rankings quickly. Algorithmic recovery waits for re-evaluation, which happens as pages are re-crawled and, for the broader helpfulness signals, when Google next runs a core update. That can mean months rather than weeks, and it is the single most common reason people conclude their fix did not work when it simply has not been reassessed yet.

When thin content does draw a manual action

Manual actions for thin content exist, but they target deliberate manipulation rather than weak pages. The named categories are auto-generated content with no added value, thin affiliate pages that reproduce a merchant feed without original contribution, scraped content republished from elsewhere, and doorway pages built to funnel traffic from many near-identical entry points.

A site with a directory of shallow tag archives is not in that category. A site running a programmatic generator that spun up ten thousand keyword-shaped URLs may well be. If you are unsure which describes you, the Manual Actions report answers it definitively in about five seconds.

What to do either way

The work is the same in both cases, which is the reassuring part: decide for each thin page whether to improve it, merge it into a stronger page, or remove it. Doing that resolves a manual action and improves your standing in the algorithmic assessment at once.

Expect a lag before you can judge the outcome. Google has to re-crawl the changed pages, notice the removals, and re-evaluate the site as a whole. Judging a content cleanup two weeks after shipping it tells you almost nothing.

Thin content in practice

How do you fix thin content pages?

Every thin page gets one of three decisions. Improve it if the topic deserves a page and you can genuinely make it the best answer: add depth, examples, structure. Merge it if it overlaps with a stronger page, combine them and 301-redirect the thin URL, so the surviving page inherits its signals. Remove it if the page has no reason to exist, return a 410 or redirect it to the most relevant page.

The wrong move is to leave a directory of thin pages sitting there "just in case". Doing nothing keeps the drag on site quality and the crawl-budget waste. Pick one of the three for each page and act on it.

Is thin content the same as short content?

Do not read a thin-content flag as "add more words". Padding a page to hit a word count makes it worse, not better: now it is thin and bloated. A concise page that completely answers its query is not thin, and no audit should push you to inflate it.

The fix for genuine thinness is added value, not added length: a real example, a comparison, a step someone can follow, a piece of information the reader could not easily find elsewhere. If a page cannot be given that, it is a candidate for merge or removal, not expansion.

What causes thin content pages?

It is rarely written deliberately. It accumulates: a programmatic page generator that spun up a URL per keyword, a CMS that publishes a page for every tag and filter combination, an old strategy of "more pages equals more traffic" that left a directory of near-empty stubs behind.

That is why the fix usually starts at the source, turning off the generator or noindexing the auto-created pattern, rather than editing pages one by one. An audit that groups the thin pages together shows you the pattern, so you can address the cause and then clean up what it already produced. Pages that are also duplicates are covered in the duplicate content guide.

Find the thin pages dragging your site down

Free to start. The audit stacks the weak-content signals and flags the pages where they cluster.

Start my free audit