Most sites of any size now have two kinds of page. There are the articles someone wrote, and there are the pages a template produced from a dataset: one per city, per product, per comparison, per question. A content agent pointed at the site adds to the second kind faster than anyone can read it. The two kinds fail differently, and an audit built for the first will pass the second while it is quietly the larger risk.
This post quotes the search policy that governs the second kind, as Google publishes it and with the date it took effect. It then sets out the audit that policy implies: classify every URL into one of the two classes, sample the template rather than the page, measure four things per template, and decide per template whether to fix it, thin the set or retire the class. No site is named, and nothing here promises a ranking outcome.
- 01Google's policy is about scale and purpose, not about how the pages were made.Its own wording: many pages generated primarily to manipulate rankings and not help users, focused on unoriginal content of little value, 'no matter how it's created'. Announced March 5, 2024. Human, automated or mixed production is explicitly covered.
- 02Classify first, because the two classes need different audits.An editorial page is judged on its own merits. A template-generated page is judged as one instance of a class, and the class is what can be thin, stale or pointless. Four signals separate them, and an agent can apply all four.
- 03Sample templates, and measure four things per template.Unique-content ratio, boilerplate share, whether the data behind the template is real and current, and whether the page answers the query it targets. Each has a definition and a way to compute it below.
- 04Per template, the decision is fix, thin or retire.Fix the template when the data is good and the rendering is thin. Thin the set when only some instances have enough data to be worth a page. Retire the class when no instance answers a query the site should be answering.
01 — The anchorThe policy, as published
Google's spam policies page, as read on September 22, 2026, defines scaled content abuse in one sentence, quoted in full below, and follows it with a list of examples that begins with using generative AI or similar tools to generate many pages without adding value for users, and includes scraping feeds or search results into pages, stitching content from other pages, spreading scaled content across multiple sites to hide its scale, and pages that make little sense to a reader but contain search keywords. The page tells a site hosting such content to exclude it from Search.
Scaled content abuse is when many pages are generated for the primary purpose of manipulating search rankings and not helping users. This abusive practice is typically focused on creating large amounts of unoriginal content that provides little to no value to users, no matter how it's created.Google Search Central, spam policies page, as read September 22, 2026
The policy was announced in Google's March 5, 2024 post alongside the March 2024 core update, as one of three new spam policies. That post says the policy builds on the earlier one about automatically generated content so that action can be taken whether content is produced by automation, by people, or by a combination. Google's February 2023 guidance on AI content makes the same distinction from the other side: automation is not spam by itself, and has long produced useful pages such as sports scores, weather and transcripts; using it to generate content primarily to manipulate rankings is.
Two things follow for an audit. The test is about purpose and value at scale, so it applies to a class of pages, not to a page. And the examples are all things a template does: renders a feed, combines sources, produces many near-identical pages. A per-article review answers none of that.
02 — Step oneClassify every URL first
Before measuring anything, split the site's URLs into two classes. This is a job an agent does well because the signals are mechanical and the answer for most URLs is obvious. Four signals, checked in order; a URL that shows two or more of the template signals goes in the template class.
| Signal | Editorial | Template-generated |
|---|---|---|
| URL shape | A unique slug per page with no repeating pattern | A path with a variable segment: a city, a product, a category, a comparison pair |
| Page structure | Heading order and section count vary between pages | The same heading order, section count and word count band across hundreds of URLs |
| Text overlap | Low shared text with any other page on the site | High shared text with sibling pages once the variable terms are masked |
| Where the words come from | An author, a byline, a publication history | A data source rendered through a layout, with or without a generated paragraph |
The output of this step is a list of templates, each with the count of URLs it produces. That count is the first number that matters: a template producing twelve pages and one producing forty thousand are different risks under a policy about scale, even if each page looks identical. The editorial class then goes to the audit in our post-core-update content audit template; the rest of this post is for the template class.
03 — Step twoSample templates, not pages
A template-generated page is one draw from a distribution the template defines. Reading a hundred of them tells you about the hundred; reading twenty from each template, chosen to span the data, tells you about the class. Pick the sample to include the instances with the most underlying data, the least, and a few from the middle, because a template that looks fine on its richest instance is usually thin on its poorest.
For each sampled instance, render the page as a user sees it, strip the site chrome, and keep the article body. That body is what the four measures below run on. Record the template, the instance, the variable values and the four measurements as one row, so the results can be aggregated per template at the end.
An editorial audit scores each page on its own: is it accurate, is it well written, does it have a source. A template page can pass all three and still be one of forty thousand near-copies that exist because the data had forty thousand rows. The policy test is about the class, and no per-page score can see the class. Sites that audit template pages one at a time conclude the pages are fine, because individually they are.
04 — Step threeThe four measures
Four numbers per template, each with a definition and a way to compute it. The first two are mechanical. The third needs someone who knows the data. The fourth needs someone who knows the query.
- Unique-content ratioMask the variable terms, then compute the share of each page's body text that does not appear on any sibling page from the same template. Average across the sample. A template whose pages are mostly shared text once the city or product name is masked is producing one page many times.
- Unique words / total words
- Boilerplate shareThe share of the rendered page, chrome included, that is identical across every instance: navigation, disclaimers, generic intros, calls to action. High boilerplate with low unique content is the shape the policy's examples describe.
- Shared characters / page characters
- Data reality and currencyFor each variable field, is the value a real observation with a source and a date, or a placeholder, a default or a guess? What share of instances have a value updated within the field's shelf life? A page rendered from stale or empty fields answers nothing, however well it is written.
- Fields sourced and current / fields
- Query answeredTake the query the template targets, substitute the instance's terms, and ask whether the page answers it better than the site's own category or search page would. Score yes, partly or no per instance. A template that scores no on most of its sample has no reason to exist as separate pages.
- Yes / partly / no per instance
The third measure is where agent-generated pages most often fail, and it is the one a text-only audit cannot see. A page can have a high unique-content ratio because a model wrote a fresh paragraph for every instance, and every paragraph can be describing data that is empty, invented or two years old. Our post on expired offers and what to do with the page works through the currency question for one common template.
05 — Step fourFix, thin or retire
Aggregate the four measures per template and route each template to one of three outcomes. The decision is made once per template, applied to every instance it produces, and recorded with the numbers that drove it.
Retirement is the outcome teams resist, because the pages were cheap to make and each one seems harmless. The policy is written about the aggregate, and the audit's first output, the count per template, is what makes the aggregate visible. If you want the classification and the four measures run across a site before a content agent adds to it, that is part of what we do under agentic SEO.
Per-page audit
Each page's accuracy, writing and sourcing, one at a time. Passes a template class whose every instance is individually fine and collectively unoriginal.
Per-template audit
The count of pages per template, the unique-content ratio, the boilerplate share, the state of the data, and whether the query is answered. Produces one decision per template.
06 — ConclusionThe policy is about a class of pages, so the audit has to be too
Classify the URLs, count the pages per template, measure four things on a sample, and decide once per template
The policy Google publishes is short and clear, and it applies to scale and purpose rather than to how the pages were made. An audit that matches it works at the level of the template: how many pages it makes, how much of each is shared, whether the data underneath is real and current, and whether anyone is asking the question the page answers. Run that before a content agent adds the next ten thousand instances, and the decision to fix, thin or retire is yours rather than a search engine's.