SEOFramework16 min readPublished August 2, 2026

A six-signal hub depth diagnostic · the AI-citation mechanism argued, not asserted · 8 sources cited

Your Category Pages Are Thin: A Hub Depth Diagnostic

Most category and hub pages are a heading over a grid — and in classic search, that often works. Whether it works for AI answers is a different question, and here is the honest part: no study we located measures citations by hub-versus-leaf depth. This piece separates what the research proves from what we argue, then gives you a six-signal diagnostic and a repair pattern that stand on their own.

DA
Digital Applied Team
Senior strategists · Published Aug 2, 2026
PublishedAugust 2, 2026
Read time16 min
Sources8 cited
Avg unique words · #1 UK category pages
310
n=300 · Digitaloft, Aug 2025 snapshot
Quality-score citation odds ratio
4.2×
GEO-16 · 1,702 citations · 3 engines · English B2B SaaS
Comparison-page citations per retrieval
1.87
DeltaV · 240-URL sample
+45% vs avg
Studies segmenting by hub-vs-leaf depth
0
among the sources reviewed here

Thin hub pages — category, collection, and topic-index URLs that offer a heading, a grid of links, and little else — are the most common structural weakness we find in site-architecture audits. They sit at the exact URLs a site nominates as canonical for its most valuable topics, yet many of them answer no question a buyer actually has.

Here is what makes this topic harder than most SEO writing admits: the evidence is genuinely mixed. A 300-page study of UK e-commerce category pages found that pages ranking #1 in classic search average just 310 words of unique content — thin hubs demonstrably win rankings. Meanwhile, two large AI-citation studies and one quality-signal framework suggest that generic, low-signal pages are cited by AI answer engines at lower rates. Neither body of research directly tests the question this post cares about, and we will say so plainly rather than pretend otherwise.

This guide covers what the research actually measures, why hubs go thin at scale, a six-signal diagnostic you can run on your own architecture this week, and a repair pattern — rebuilding hubs as decision documents — that improves the page for buyers regardless of how the AI-citation question eventually resolves.

Key takeaways
  1. 01
    No study we located measures citations by hub-vs-leaf depth.Every AI-citation study we reviewed segments by page format (product, article, listicle, comparison, homepage, category) or funnel stage — never by position in a site hierarchy. DeltaV does report a category-page class — 5.2% of its citations at the lowest per-retrieval rate of its ten page types — but that is a template label, not a depth axis. The hub-citation mechanism in this post is argued from adjacent evidence, and we label it that way.
  2. 02
    Thin hubs already win classic search.Across 300 UK e-commerce category pages ranking #1 for commercial keywords, unique content averaged 310 words; 44% of those pages had under 200 words and 10% had none beyond the H1. Thinness is not universally punished — the argument has to be about AI answers, not rankings.
  3. 03
    Quality signals, not word count, predict citation.In the GEO-16 framework study (1,702 citations, three engines, English B2B SaaS pages), overall page quality carried a citation odds ratio of 4.2. Metadata and freshness, semantic HTML, and structured data showed the strongest associations of the 16 on-page pillars it audits — word count was not one of them.
  4. 04
    The diagnostic stands on its own.Six signals — hub-vs-leaf word gap, zero-content check, entity-coverage gap, decision-content presence, faceted URL sprawl, and one deliberately weak metric included as a cautionary example — form a repeatable self-audit no located source combines.
  5. 05
    The repair is a decision document, not word stuffing.A rebuilt hub answers four questions: what the options are, how to choose, what each costs, and who each suits. That template serves buyers in any channel — which is why it is the load-bearing half of this post.

01Evidence StatusWhat is measured, and what we argue.

Before any diagnostic, the claim inventory. A popular line in AEO and GEO content goes something like: AI engines cite your category page because it is the canonical URL for the topic, find nothing quotable there, and your brand loses the citation. It is a tidy mechanism. It is also, as far as we can determine, unproven — none of the citation studies we reviewed for this piece segments AI citations by site-architecture depth, hub versus leaf.

What the located research does support is narrower. AI answer engines cite certain page formats — product pages, comparison pages, articles, and category pages as a template class — at rates that vary by funnel stage and industry. Page quality signals showed the strongest associations with citation among the on-page pillars the one framework study we located audited. And faceted navigation is, per Google’s own crawling team, the most common source of overcrawl issues site owners report to them. Those three findings are adjacent to the hub question without answering it directly.

Honesty check
The direct study — AI citations segmented by hub versus leaf position in a site hierarchy — does not exist among the sources we reviewed. Two current, well-documented citation datasets (768,000 citations and 25,337 citations) both had the opportunity to report a depth axis and segmented by page format instead. The DeltaV dataset does carry a “category page” label — 5.2% of its 25,337 citations, at 1.06 citations per retrieval, the lowest rate of the ten page types it lists — and that is the closest adjacent number in this post. But it is a template classification, not a hierarchy position: a category page four clicks deep and one linked straight off the homepage land in the same bucket. Everything in section 06 is therefore our reasoned inference, and we recommend treating any competing content that asserts a measured hub-versus-leaf citation preference with suspicion until someone publishes the depth data.

Why write the post anyway? Because the practical problem is real whether or not the citation mechanism is ever measured. A hub page that answers nothing is a weak asset in every channel it appears — classic SERPs where a competitor’s richer page wins the long-click, AI answers where quotable substance appears to matter, and direct traffic where a buyer bounces off a wall of tiles. The diagnostic and the repair earn their place on buyer value alone.

02Root CausesWhy hubs go thin at scale.

Hub thinness is rarely a content decision. It is an architectural byproduct. Most platforms generate category pages from a template: H1, optional intro field nobody fills in, product or article grid, pagination. Every new category inherits the template, and the template has no opinion about substance. Multiply by faceted navigation — size, color, price, brand filters — and each filter combination generally creates its own URL showing largely the same inventory, which is how large sites accumulate thousands of technically unique but sparse category-adjacent pages.

This is not a fringe concern. In a December 2024 Search Central post on faceted navigation, Google’s Gary Illyes wrote that faceted navigation is “by far the most common source of overcrawl issues site owners report to us” — and Google’s recommended fix is a binary choice: block the filter URLs outright, or optimize them for crawling with consistent parameter order, canonical tags, and 404s on empty result sets. Google is watching this URL class closely; that is documented. What it means for AI citation is not, and we return to that distinction in section 06.

One vocabulary note worth getting right: Google’s current spam policies do not use the phrase “thin content” at all — that is industry shorthand, not Google’s term. The closest defined violations are doorway abuse, described as pages “created to rank for specific, similar search queries” that “lead users to intermediate pages that are not as useful as the final destination,” and thin affiliation, where product descriptions are “copied directly from the original merchant without any original content or added value.” A zero-content category page usually violates neither policy. The problem with thin hubs is almost never a penalty — it is opportunity cost.

Any honest version of this argument has to start with data that cuts against it. In a study by Digitaloft, 300 UK e-commerce category pages ranking #1 for commercial-intent keywords (collected via Semrush from a 524-keyword list, each page manually verified) averaged just 310 words of unique content — excluding product-grid text, buttons, faceted navigation, and the H1. The distribution is even more striking than the average.

Unique content on #1-ranking UK e-commerce category pages

Source: Digitaloft category-page content study, n=300, data collected Aug 2025
All pages analyzed300 UK category pages ranking #1 · Aug 2025 collection
300
Under 400 words unique content66% of the 300 pages
66%
Under 200 words unique content44% of the 300 pages
44%
Zero unique content beyond the H110% of the 300 pages — 29 URLs, mostly household brands
10%

Two caveats keep this honest. First, the author’s own framing: “The aim wasn’t to prove that word count directly affects rankings, but to see what top-ranking pages are actually doing in practice.” It is a correlational snapshot of one country and vertical mix, not a causal claim. Second, the word-count benchmarks in this space disagree: a separate practitioner claim — which we could not verify against a disclosed methodology — puts top-ranking category pages at 500 to 700 words, roughly double the Digitaloft figure. We present the sourced number, note the disagreement, and deliberately settle on neither as a target.

The interpretive point matters more than either number: many hubs are thin by design and still perform in classic search — most of the 29 zero-content pages in the sample, 18 of them, were household-name brands whose authority, in the study author’s reading, likely explains the ranking. So the case for rebuilding a hub cannot rest on “thin pages lose rankings.” For plenty of strong brands, they demonstrably do not. The case has to rest on how thin hubs fail differently — for buyers who land and find nothing, and plausibly for AI answers that need something quotable. That distinction shapes everything that follows.

04Citation ResearchWhat the citation studies actually measure.

Two independent datasets dominate the AI-citation evidence base at the time of writing, and they must not be blended. The first, reported by Search Engine Journal from research by XFunnel, tracked 768,000 citations over 12 weeks across ChatGPT, Google AI Overviews, and Perplexity. Within that dataset, product content captured 46 to 70 percent of citations depending on funnel stage — rising past 70 percent at bottom-of-funnel queries — while blog content took only 3 to 6 percent and PR materials under 2 percent. B2B and B2C mixes differed sharply: product pages earned 56 percent of B2B citations in the sample versus 35 percent for B2C.

The second, a DeltaV Digital study published in July 2026, analyzed 25,337 citations from 21,075 AI-engine responses across five engines, narrowed to the 30 most-cited URLs per brand across 8 brands — a 240-URL sample. In that sample, articles took 23.7 percent of citations, listicles 19.6 percent, and product pages 16.3 percent — but comparison pages, at only a 4.1 percent share, posted the highest citation rate per retrieval at 1.87, about 45 percent above the portfolio average. Small footprint, high intensity: the page format built around helping someone choose gets cited hardest when it is retrieved at all.

The table below puts the two side by side — not to reconcile them, but to show why they cannot be reconciled, and what each one does and does not tell you about hubs.

Citation share by page type across two independent AI-citation studies — XFunnel via Search Engine Journal (768,000 citations, 2025) and DeltaV Digital (25,337 citations, 2026) — with caveats explaining why the two datasets are not directly comparable.
Page typeXFunnel via SEJ · 768K citations, 2025DeltaV Digital · 240-URL sample, 2026Comparability caveat
Product pages (leaf-type in our framing)46–70% of citations by funnel stage; 70%+ at bottom-funnel16.3% share · 1.22 per retrievalDifferent denominators, engine mixes, and sampling — never merge these percentages
ArticlesNews/research articles 5–16% each; blog content 3–6%23.7% share · 1.43 per retrievalSegment definitions differ — SEJ splits news, research, and blog; DeltaV uses one article class
ListiclesNot reported as a separate segment19.6% share · 1.45 per retrievalDeltaV-only segment; 61% of B2B-services citations in its sample were listicles
Comparison pagesNot reported as a separate segment4.1% share · highest rate at 1.87 per retrieval (+45% vs portfolio average)Small share, high intensity — the strongest per-page signal in the DeltaV sample
HomepagesNot reported as a separate segment10.8% portfolio-wide share · 1.11 per retrieval; 55% of citations for the single multi-location medical aesthetics brand in the sampleThe 55% is one brand, not a class — no single page type dominated across DeltaV’s 8 verticals
Category pages (a format label)Not reported as a separate segment5.2% share · 1.06 per retrieval — the lowest rate of the ten page types DeltaV listsA template classification, not a position: a category page four clicks deep and one off the homepage share this bucket
Hub vs leaf position in a hierarchyNot measuredNot measuredThe gap this post is built around — neither study segments citations by how deep a URL sits in the architecture

DeltaV’s self-disclosed limitations deserve passing on: the top-30-per-brand cut captures the head of a power-law distribution but not the long tail, the prompt sets skew commercially relevant, page classifications are automated and imperfect at the margins, and 8 brands in 8 industries is a directional sample, not a census. Directional is still useful — both studies agree that generic, undifferentiated content under-cites relative to specific, decision-useful formats. One DeltaV row is the closest adjacent number this post has: pages its classifier labelled category pages took 5.2 percent of the 25,337 citations at 1.06 citations per retrieval, the lowest rate of the ten page types it reports. We use that as suggestive, not decisive — it is a template label applied by an automated classifier DeltaV itself calls imperfect at the margins, and it says nothing about where in your architecture that content lives.

"The practical takeaway for marketers: copy the fingerprint of your industry, not a generic AI content best practices checklist."— Brandon Kidd, DeltaV Digital AI citation study, July 2026

05Quality SignalsQuality predicts citation — within limits.

The closest thing to a mechanism-level account comes from the GEO-16 framework paper on arXiv, which audited 1,702 citations across 1,100 unique URLs generated by 70 product-intent prompts, spanning three engines: Brave Summary, Google AI Overviews, and Perplexity. Its headline finding: among the 16 on-page pillars it audits, those covering metadata and freshness, semantic HTML, and structured data showed the strongest associations with citation. A logistic regression put the citation odds ratio for overall quality score at 4.2, with a 95 percent confidence interval of 3.1 to 5.7. The paper describes itself as observational, and every pillar in it is an on-page quality signal — page length, URL type, and hierarchy position are not among the variables it models.

Strongest predictor
Quality-score citation odds ratio
4.2×

Logistic regression across 1,702 audited citations found overall page quality a strong citation predictor (95% CI 3.1–5.7), with metadata, semantic HTML, and structured data the strongest associations among the 16 on-page pillars audited.

GEO-16 · arXiv
High-quality pages
Cross-engine citation rate
78%

Pages scoring at least 0.70 on the framework's 0–1 quality scale with 12+ of 16 pillar hits reached a 78% citation rate across the three engines studied.

G ≥ 0.70 · 12/16 pillars
Engines differ
Perplexity's mean cited quality
0.300

Mean quality score of cited pages: Brave 0.727, Google AI Overviews 0.687, Perplexity just 0.300 — Perplexity's citation bar was measurably lower in this sample.

Per-engine means

Two more findings and one large hedge. Pages cited by more than one engine — 134 of the 1,100 audited URLs — scored 71 percent higher on quality than single-engine citations, suggesting that quality is what generalizes across engines while individual engines tolerate individual weaknesses. The hedge: GEO-16’s sample is narrow. English-language B2B SaaS pages only, 70 prompts, three engines — no ChatGPT, no Gemini — and it predates the 2026 model generations. Treat its odds ratio and thresholds as directional, not universal. What it contributes to the hub question is the direction of the gradient: pages weak on metadata, structure, and entity coverage under-cited even when topically relevant. A template-generated hub page is, almost by definition, weak on exactly those signals.

06Our ReasoningThe hub-citation mechanism, argued step by step.

What follows is our inference, not a measured finding — we want that unmistakable. Built from the adjacent evidence above, the argument runs in four steps:

  • Step 1 — hubs are the nominated topic URLs. Your internal linking, breadcrumbs, and canonical structure all tell crawlers the category page is the authoritative URL for the category-level topic. That is what site architecture is for, and Google’s crawling guidance confirms this URL class gets heavy crawler attention — heavy enough that faceted navigation, the URL family category pages spawn, is what Google calls by far the most common source of overcrawl issues site owners report to it.
  • Step 2 — retrieval and citation are separate hurdles. The DeltaV data distinguishes citation share from citations per retrieval — comparison pages were retrieved rarely but cited at the highest rate (1.87 per retrieval in its 240-URL sample) when they were. Being fetched is not being quoted.
  • Step 3 — quotable substance predicts the quote. GEO-16 found quality signals carried a 4.2× citation odds ratio within its sample, and both citation studies found generic content under-cited relative to specific, decision-useful formats. The nearest directly relevant number: the category-page class in DeltaV’s sample posted the lowest citations-per-retrieval of its ten page types, 1.06. A grid of product tiles offers an answer engine nothing to extract.
  • Step 4 — the inference. If a hub is the URL your architecture nominates for a category query, and it contains nothing quotable, the plausible outcomes are that the engine cites a competitor’s richer page, cites one of your leaves out of context, or skips the citation entirely. We consider that mechanism likely. We cannot show you a study that measures it — and any competing article that claims one exists is, as of our review, ahead of the evidence.

The way to convert this from argument to observation on your own site is measurement: track which of your URLs AI engines already cite and watch how that changes after a hub rebuild. Our brand citation audit checklist covers exactly that instrumentation. One honest n=1 beats a fabricated n=768,000.

07The FrameworkThe Hub Depth Diagnostic: six signals.

This table is ours. No located source combines these signals into one self-audit: existing coverage measures word count alone, or citation-quality signals alone, never a working diagnostic aimed at a site owner deciding which hubs to rebuild. The verdict thresholds are our recommended starting heuristics — explicitly not sourced numeric cutoffs, because as section 03 showed, the published word-count benchmarks disagree with each other by roughly 2×. One signal, text-to-HTML ratio, is included deliberately as a cautionary example of a metric to drop.

The Hub Depth Diagnostic: six signals for auditing category and hub pages — what each measures, how to check it, its known limitation, and our recommended verdict threshold, labeled as heuristics rather than sourced cutoffs.
SignalHow to check itKnown limitationThin verdict (our heuristic)
Hub-vs-leaf unique word gapCrawl the section; compare unique on-page words on the hub against the average of the leaves beneath it, excluding template and grid textWord count is not value — #1-ranking UK hubs averaged 310 words (n=300), while an unverified competing benchmark says 500–700; the benchmarks disagree ~2×Hub carries under a third of its average leaf’s unique words and answers no buyer question on its own
Zero-unique-content checkManual: does the page render anything beyond H1, grid, filters, and pagination?10% of the 300 Digitaloft pages had zero content beyond the H1 and still ranked #1 — most of them, 18 of the 29, household-name brandsZero unique content without dominant brand authority to carry it
Entity-coverage gapEntity extraction (InLinks, Semrush, or manual): list the entities a buyer needs for this decision; compare hub coverage against its leaves collectivelyA practitioner heuristic — documented in vendor explainers, with no engine-published threshold behind itHub mentions under half the entity set its own leaves cover between them
Decision content presentManual read: does the hub say what the options are, how to choose, what each costs, and who each suits?Qualitative; our own template, informed by comparison pages’ high per-retrieval citation rate in the DeltaV sample, not externally validated as a thresholdNone of the four questions answered anywhere on the page
Faceted / filter URL sprawlCrawl or Search Console: count crawlable filter-combination URLs the hub spawns; check canonical and robots handling against Google’s faceted-navigation guidanceA crawl-efficiency signal, not a content-quality measure — but it multiplies however thin the hub already isFilter URLs crawlable, uncanonicalized, and returning near-duplicate inventory
Text-to-HTML ratio (weak — do not use)Legacy audit tools still report it; included here only to show why it failsGoogle’s John Mueller has said the metric “makes absolutely no sense at all for SEO” — a public Reddit comment, not official documentation, but blunt enoughNo verdict — drop it from your audit rather than act on it

The entity-coverage signal deserves one expansion, because it is the best depth measure that is not word count. The heuristic — covered in entity-based SEO explainers and implemented by commercial tools — checks entities per thousand words and gap-analyzes a page against the entity set its topic implies. A running-shoes hub whose leaves collectively cover cushioning classes, drop heights, surface types, and brand lines, while the hub itself names none of them, has a measurable depth gap no word counter sees.

Running this across a large site by hand is tedious, which is why we fold it into automated audits — the deep-research audit prompt method we published alongside this piece surfaces thin hubs as a standard finding category. And once the diagnostic produces a list, resist fixing it top-to-bottom: our refresh prioritization matrix is the tool for deciding which hubs earn the rebuild first — commercial value times depth gap, not alphabetical order.

08The RepairRebuild hubs as decision documents.

The repair is not “add 500 words.” Word-count targets are exactly the trap section 03 warned about. The repair is a structural template we call the decision document: a hub that answers, above or alongside its grid, the four questions a buyer at category level actually has — what the options are, how to choose between them, what each costs, and who each one suits. That template is informed by the one strong per-page signal in the citation research: comparison content’s 1.87 citations per retrieval in the DeltaV sample — and it improves the page for human buyers regardless of what any engine does with it. Not every hub needs the same treatment, so match the repair to the diagnostic verdict:

Zero-content hub
High commercial value, nothing on the page

Full decision-document rebuild: options overview, selection criteria, honest cost ranges, fit-by-buyer-type. Keep the grid; add the substance above and beside it. This is the highest-leverage case the diagnostic finds.

Rebuild as decision document
Faceted sprawl
One thin hub, thousands of thinner children

Architecture first, content second. Apply Google's binary — block filter URLs or optimize them with consistent ordering and canonicals — before writing a word. A rebuilt hub buried under crawl waste stays buried.

Fix architecture first
Ranking thin hub
Already #1 with 100 words

Do not stuff it — the Digitaloft data says thinness is not what classic rankings punish. Add a decision layer only where it answers a real buyer question, and measure before-and-after rather than assuming improvement.

Add decision layer, measure
Strong leaves, weak hub
Depth exists — one level too deep

Surface it. Pull the comparison logic your leaf pages already contain up to the hub, link each claim to its leaf, and close the entity-coverage gap the diagnostic measured. Mostly an editing job, not a writing job.

Promote leaf depth upward

Three build notes. First, the cost question is the one most brands skip and the one with the clearest bottom-funnel case — we made the full argument in our piece on cost transparency pages, and a hub that includes honest price ranges inherits it. Second, a rebuilt hub only pays off inside a coherent structure: the pillar-and-cluster build method is the constructive counterpart to this post’s diagnostic — this piece finds the weak hubs, that one builds the architecture they should sit in. Third, do not neglect the usability layer: Nielsen Norman Group maintains a long-running body of paywalled e-commerce UX research on category and listing pages — two report volumes with 62 and 139 design guidelines respectively — which we cite here as evidence that category-page usability is a mature research field, not as a source of specific guidelines we have not read.

Looking forward, our expectation — clearly labeled as projection — is that someone will publish the missing study within the next research cycle: citation datasets already carry URL paths, and segmenting by architecture depth is one classifier away. If that study lands and shows hub URLs under-cited relative to their retrieval rates, the diagnostic above becomes more urgent. If it shows the opposite, you will have rebuilt your most valuable category pages into genuinely useful buyer documents — which is the rare SEO bet that pays out even when the thesis is wrong. This is the standard we hold agentic SEO engagements to: recommendations that survive the evidence changing.

09ConclusionDiagnose honestly, rebuild for buyers.

The bottom line

The mechanism is argued. The diagnostic is yours to run either way.

The honest summary of the evidence: thin hubs demonstrably survive classic search — 310 average words across 300 #1-ranking UK category pages says so. Generic, low-signal pages under-cite in AI answers — two independent citation datasets and a quality-signal framework point the same direction. And the specific claim that AI engines fetch hub URLs and fail to cite them for lack of quotable substance is our inference, stated as such, because no study we located segments citations by hierarchy depth and settles it.

That honesty is not a weakness of the method — it is the method. Run the six-signal diagnostic, rank the failures by commercial value, and rebuild the winners as decision documents that answer what the options are, how to choose, what each costs, and who each suits. Every one of those steps improves an asset your buyers already land on, whatever the citation research eventually proves.

The pages your architecture nominates as most important should be the pages that say the most. Most sites have it exactly backwards — and now you have a repeatable way to measure by how much.

Fix the pages your architecture says matter most

Your most valuable URLs should answer real questions.

Our team runs hub-depth diagnostics, rebuilds category pages as decision documents, and instruments AI-citation tracking so you can see what changed — delivered in days, not quarters.

Free consultationExpert guidanceTailored solutions
What we work on

Site architecture engagements

  • Hub depth diagnostics across full site architectures
  • Decision-document rebuilds for category pages
  • Faceted-navigation and crawl-budget cleanup
  • Entity-coverage gap analysis, hub vs leaves
  • AI citation tracking before and after rebuilds
FAQ · Hub depth diagnostic

The questions we get every week.

A thin hub page is a category, collection, or topic-index URL that offers little beyond a heading and a grid of links to deeper pages — no explanation of the options, no selection guidance, no unique substance. Worth knowing: “thin content” is industry shorthand, not Google’s term. Google’s current spam policies define related but distinct violations — doorway abuse (pages created to rank for similar queries that route users to intermediate pages less useful than the final destination) and thin affiliation (merchant descriptions republished without added value). A typical zero-content category page violates neither policy, which is why the practical problem with thin hubs is opportunity cost — a weak asset at your most important URLs — rather than penalty risk.