Every organisation with a content archive is being offered the same thing this year: an agent that reads all of it and proposes what to keep, what to rewrite and what to delete. The offer is real. An agent can read two thousand posts in an afternoon and produce a plan a person could not. The mistake is treating the plan as a decision. It is a proposal, and the data has to decide.
This post sets out how to make that work in practice: three dispositions with rules behind them, the measurable signals each rule reads and where each signal comes from, a reconciliation step that turns an agent's plan into a short list of disagreements for a person, and the mechanics of merging and removing pages taken from Google's own documentation. Every threshold in it is illustrative. Your archive sets its own.
- 01Three dispositions, not one: keep or rewrite, merge, retire.Each has a different trigger. A page that still receives impressions but no longer answers its query is rewritten; a page a newer page has replaced is merged; a page with no traffic, no links and no job is retired.
- 02Four signals decide, and none of them is editorial taste.Organic entrances over a stated window, referring domains and internal links, whether the page still answers its query, and duplication against newer pages. Each comes from a named source you already have.
- 03An agent's plan needs a second pass from a different data source, and only the disagreements go to a person.Re-derive the classification independently, diff the two, and review the rows that differ. On most archives that is a small fraction of the whole, which is the only way a human review of thousands of URLs is real.
- 04Never retire before the redirect audit, and never retire a page that holds links to a 404.Google's documentation says a permanent redirect is a canonical signal and a 4xx removes the URL from the index. The order matters: links first, redirects second, removals last.
01 — The choiceThe three dispositions and the rule behind each
Keep, or rewrite
Keep when the page receives organic entrances and its facts hold. Rewrite when it still receives impressions for its query but the content is dated, wrong or thin, which usually shows as impressions holding while clicks fall. The URL stays. The words change. Nothing in the redirect layer moves.
Merge
Two pages targeting one query split the links and the signals between them. Move anything unique from the older page into the survivor, then permanently redirect the older URL to it. The older page's referring domains now point at the page that ranks.
Retire
The page receives no entrances over the window, no referring domain points at it, no internal link depends on it and no newer page needs its content. Remove it from the sitemap and return a 404 or 410. If any link exists, it is not a retirement; it is a merge or a redirect.
The rules are strict on purpose. An agent asked to "clean up the archive" will propose deletions, because deletion is the visible outcome of cleaning. The rules above make deletion the hardest disposition to earn: a page has to fail every signal, not just look old. Counting the links tends to turn proposed retirements into merges, which is the opposite of the first draft an agent produces.
02 — The evidenceThe four signals, where each comes from, and what it says
Each signal is a number or a fact you can pull from a system you already run. The fourth column shows the shape of a threshold, labelled illustrative because the right value depends on the archive's size, subject and seasonality. Set yours before the agent reads a single page, and write it down, so the plan is reproducible.
| Signal | Where it comes from | What it tells you | Threshold shape |
|---|---|---|---|
| Organic entrances over a stated window | Your analytics, landing-page report, filtered to organic; Search Console clicks by page as the second source | Whether anyone arrives from search at all. Use a window long enough to cover seasonality for your subject. | Illustrative: fewer than five entrances in twelve months |
| Referring domains and internal links pointing at the URL | A backlink index for external domains; your own crawl or site index for internal links | What the URL is worth to other pages even if nobody reads it. A page with links is never retired to a 404. | Illustrative: any referring domain, or more than three internal links |
| Whether the page still answers the query it targets | Search Console queries by page; a person, or an agent with the current facts, reading the page against them | Whether the content is wrong, dated or thin for the query it still receives. This is the rewrite trigger. | Illustrative: impressions rising while clicks fall, or a fact that is no longer true |
| Duplication against newer pages | A similarity pass across titles, headings and body text; Search Console showing two of your URLs on one query | Whether a newer page already does this page's job. This is the merge trigger. | Illustrative: a newer page targeting the same query with more entrances |
Two of the four signals have a clock on them. Analytics and Search Console keep a limited window of history, and a page that had entrances two years ago may show none simply because the data expired. Our retention table lists the windows; the practical rule is to export the landing-page and query history for the whole archive before the audit starts, and to date the export in the plan.
03 — The controlThe reconciliation step, and why an agent's plan needs one
An agent that reads the pages and the analytics export will produce a confident table: URL, disposition, reason. It will be mostly right and wrong in ways that are hard to spot by reading the table, because the errors look like the correct rows. A person cannot review two thousand rows with any care. A person can review sixty.
The control is to produce the table twice, independently, and review only the rows where the two disagree. The second pass must use a different data source or a different method: if the first classification read the analytics export, the second reads Search Console and the backlink index; if the first was an agent reading page text, the second is a rule applied to the numbers with no page text at all. Then diff the two. Rows where both passes say retire are retired after a spot check. Rows where they differ go to a person with both reasons attached. The same principle, an independent signal the agent cannot influence, is the one we set out for agents whose metrics look too good.
Two passes of the same agent over the same export agree with each other for the same reasons they are wrong. The reconciliation only works when the second classification could not have inherited the first one's mistake, which means a different source, a different method, or both.
04 — The mechanicsHow to merge and remove, from Google's documentation
Once a disposition is approved, the mechanics are not a matter of opinion. Google documents what its systems do with a redirect, a status code and a noindex tag, and the table below pairs each disposition with the mechanism and the sentence from Google that justifies it. The pages are the redirects guide, the HTTP status code reference, the page-removal guide and the canonicalization guide, all read on September 25, 2026.
| Disposition | Mechanism | What Google's documentation says |
|---|---|---|
| Merge into a newer page | Move any unique content across, then a permanent server-side redirect (301 or 308) from the old URL to the survivor. | Google's redirects page says a permanent redirect is used by the indexing pipeline as a signal that the target should be canonical, and that a server-side redirect has the highest chance of being interpreted correctly. |
| Retire a page that still holds links | Redirect to the nearest page that serves the same reader. If none exists, keep the URL live with a noindex tag rather than a 404. | Google's removal page lists noindex as a way to keep a page out of Search results while people and other crawlers can still reach it. The redirect-over-404 preference for linked pages is ours, not Google's. |
| Retire a page with no links and no traffic | Return 404 or 410 and remove it from the sitemap. | Google's HTTP status page says URLs already indexed that return a 4xx status are removed from the index, and that crawling frequency gradually decreases. |
| Any disposition | Do not use robots.txt to retire or consolidate, and do not use the Removals tool as a permanent measure. | The canonicalization page says not to use robots.txt or the URL removal tool for consolidation; the removal page says Removals tool requests last about six months. |
The one place we go beyond Google's text is the linked-page row. Google says a 4xx removes a URL from the index; it does not say what to do with the links that pointed at it. Our answer is that a link is a reason for the page's existence even when the reader is gone, so a linked page is redirected to the nearest page that serves the same person, and only a page nobody links to gets a 404. The consolidation guide's own reasoning supports the instinct: it lists consolidating signals such as links into one preferred URL as the first reason to specify a canonical.
05 — The orderThe order of operations: links, redirects, then removals
The dispositions above are safe only in a particular sequence. Export the history first, because the signals expire and the capture is free today and impossible later; our post on what to capture before a migration has the list. Build the source index before any agent rewrites a page, as set out in our source-index post, so a rewrite does not lose the fact that made the page worth keeping. Run the redirect audit before any retirement, so you know which URLs already carry a redirect chain that a 404 would break. Then merges, which create redirects. Then, last, retirements.
Reversed, the same steps do damage. A retirement before the redirect audit can 404 a URL that three older URLs redirect into. A rewrite before the source index can replace a sourced claim with a plausible one. The sequence is the safeguard, and it fits inside the content-audit template we published for post-update reviews.
"The agent proposes a disposition for every URL from the analytics export. A rules pass re-derives it from Search Console and the link index. We review only the disagreements. No URL is retired before the redirect audit, and no linked URL is retired at all." Our SEO practice runs the audit in that order.
06 — ConclusionThe plan is the agent's; the decision is the data's; the disagreements are yours
Set the thresholds before the agent reads a page, classify twice from different sources, review only the rows that differ, and retire nothing until the links and redirects are known
An archive audit done this way is faster than a manual one and safer than an automated one. The agent does the reading. The rules do the deciding. A person does the sixty rows that matter.