AI DevelopmentAnalysis6 min readPublished October 1, 2026

One paper, 800 small models, and a number that changes sign

AI-Written Web Text Helps Small Models, Then Starts to Hurt

Researchers pretrained 800 models and found AI-written web text helps small, data-starved runs, then hurts as human data grows. What the study shows and omits.

DA
Digital Applied Team
Research and practical guidance
CoverageOctober 1, 2026

A paper submitted to arXiv on September 30, 2026 asks a question every lab now faces: what happens when the web you crawl for training data is partly written by the models you trained last year? The authors, from the University of Maryland and the AI detection company Pangram Labs, label almost a third of quality-filtered August 2026 web tokens as AI-generated, then pretrain 800 small models on mixtures of human and AI text. Their finding is a sign change. AI text helps a model that lacks human data, stops helping at the compute-optimal budget, and raises loss beyond it.

Editorial note: Prepared October 3 as an October 1, 2026 dispatch from the paper’s arXiv abstract and HTML text, version 1. One paper, not yet peer reviewed. It makes no claim about how search engines treat AI text, and neither does this post.

Key takeaways
  1. 01
    31.1% of filtered web tokensThe share of tokens passing FineWeb’s quality filter that Pangram labels AI-generated, for August 2026. Two years earlier it was 10.1%.
  2. 02
    Help below 10 tokens per parameterAI text lowers loss on human text only for models trained on fewer than about 10 human tokens per parameter. At the standard 20, the benefit is gone.
  3. 03
    1.6 times the computeBy the paper’s law, training on unfiltered web at today’s AI share costs 1.6 times the compute of training on its human subset, and the gap grows.
  4. 04
    Repeating human text beats adding AI textAfter eight added epochs, loss was 6.3 to 6.4% lower when human data was repeated than when the same volume of AI text was added.

01 — The measurementHow much of the web is AI-written now

The paper, How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text, starts by measuring. The authors take Common Crawl snapshots, apply the FineWeb quality filter that many open pretraining pipelines use, and run two detectors over what survives: a small open classifier to label everything, then Pangram’s own detector as a second check on the confident cases. The result is the chart below. The two later bars are the authors’ forecasts, not measurements.

Share of quality-filtered web tokens labelled AI-generated

arXiv:2609.40295 v1, figure 1 and section 5.1. Measured shares are Pangram labels on FineWeb-filtered Common Crawl; the 2027 and 2028 figures are the paper’s forecasts.
June 2024
10.1%
June 2026
27.5%
August 2026
31.1%
End 2027, forecast
42%
End 2028, forecast
51%

Two things about the measurement matter for everything after it. The share is of tokens that pass a quality filter, and the paper reports that AI text survives filtering more often than human text does, so the filtered share is higher than the raw share. And the label comes from a detector sold by one of the authors’ employers. The paper is open about both; a reader should carry them forward.

02 — The experimentWhat 800 models showed

The experiment holds a human corpus fixed and adds AI text to it, up to 64 AI tokens for every human token, across models from 19.9 million to 973 million parameters. Loss is measured on held-out human text from three sources and, separately, on held-out AI text. The design is what distinguishes the paper from earlier model-collapse work: the AI text is not generated by the model under test or curated to help, it is whatever the web contained, from many models, written for people.

Models
19.9M to 973M parameters
800

Varying human tokens per parameter and the ratio of AI to human tokens.

Pretrained from scratch
Corpus
Wild AI, released with the paper
83Btokens

42.2B human and 35.3B AI tokens, labelled by source, topic and format.

Public
Threshold
Below this, AI text helps
~10tokens / param

At the compute-optimal 20, the benefit is gone; above it, AI tokens raise loss.

Human-text loss

The pattern the authors describe is consistent across sizes. For a model starved of human data, adding AI text lowers loss on human text at first, then the benefit saturates and reverses as more is added. For a model already trained at or past the compute-optimal budget of about 20 human tokens per parameter, AI text raises loss almost immediately, and does so more sharply at larger sizes, while the same number of fresh human tokens keeps lowering it. The paper notes that today’s production models train far past that budget, citing one open model at 1,100 tokens per parameter, which is the regime where the harm applies.

03 — The modelA scaling law where a token can hurt

Existing scaling laws cannot express this, and the paper says why. The Chinchilla law treats an AI token as a human token. Laws built for repeated data let a token’s value fall towards zero but never below it. The authors fit a law with separate benefit and harm terms, so the value of an AI token can change sign, and which reduces to Chinchilla when no AI text is present. Fit on the smaller models, it predicts the effect of AI text on models up to 3.6 times larger with 41% lower error than the best existing law, on the authors’ own evaluation.

The law yields the paper’s most quotable figure. Training on unfiltered web text at August 2026’s AI share requires 1.6 times the compute of training on the human subset alone, at the compute-optimal budget, and the gap widens with more human data. At the forecast shares of 42% and 51%, the multipliers become 2.1 and 3.0. Those last two numbers rest on a forecast and a fitted law, and should be quoted as such.

04 — RecommendationsWhat the authors recommend

The abstract ends with four recommendations for anyone building a pretraining corpus. They are short enough to table.

Source: arXiv:2609.40295 v1, abstract and section 5, September 30, 2026. Paraphrased; the supporting figures are the paper’s.
RecommendationWhenThe paper’s support
Filter AI text outWhen the target is human-written text.The 1.6 times compute gap at today’s share, growing with the human budget.
Repeat human text before adding AI textWhen the human corpus runs out before the compute does.Loss 2.3% lower after one added epoch of repeated human data than after the same volume of AI data, widening to 6.3 to 6.4% after eight.
Report validation loss on human and AI text separatelyAlways.At a 22.3% AI share in the evaluation crawl, a mixed validation set hid the harm in 95.5% of the runs where harm occurred.
Keep AI text when the target is AI textWhen the model will mostly read or score machine output.AI tokens keep lowering loss on held-out AI text; the two are treated as separate domains.

The third row is the one with consequences outside research labs. If a mixed validation set hides the harm, then a team that fine-tunes on scraped web text and checks only an aggregate loss will not see the problem it has introduced. Our decision guide on synthetic training data makes the same point from the curated side: generated data is a tool with a target, not a free supply.

05 — FencesWhat the study cannot say

It measures loss, not capability. Every result is a held-out cross-entropy on text; the paper does not report whether the models answer questions worse or follow instructions worse. Lower loss on human text is the standard proxy, and the authors use it as one, but a reader should not translate the 1.6 figure into a benchmark score.

Its largest model is under one billion parameters. The scaling law is validated on models 3.6 times larger than those it was fit on, which is still far below production scale. Whether the harm term keeps its shape at tens of billions of parameters is a prediction the paper makes, not a measurement it reports.

Its labels come from a detector. A token is “AI-generated” when Pangram says so, and the company that sells Pangram co-authored the paper. The authors disclose this and release the corpus with its labels so others can relabel it. Until someone does, the share figures are one detector’s view.

And it says nothing about search. The paper is about what AI text does to a model that trains on it. It does not test, and does not claim, that search engines rank AI-written pages differently; the question of what publishers should do with AI text is covered in our note on Google’s updated guidance on reviewing AI content, which is a separate matter with a separate primary. For the data-supply side, our piece on publishers and Common Crawl is the context for who decides what gets crawled.

For teams that fine-tune

The practical lesson transfers down from pretraining. If your fine-tuning set includes scraped web pages, forum threads or documentation written since 2024, assume a material share is model-written, hold out human-written and machine-written examples separately, and look at both losses. The paper’s evidence says the aggregate number can look fine while the half you care about gets worse. Our AI transformation work builds that split into the evaluation from the start.

Next step

Split your validation set before you trust the loss

Whatever you train, keep a human-written holdout and a machine-written holdout and report both. That one change is the paper’s most transferable finding, and it costs nothing but the labelling.

AI model evaluation

Know what your training data is made of

Digital Applied audits fine-tuning corpora for provenance, builds split holdouts, and designs the evaluation that shows harm an aggregate score hides.

Provenance labellingSplit holdoutsLoss by domain
Start with the holdout

One afternoon

  • →Label a sample by source
  • →Separate human and machine text
  • →Report both validation losses
  • →Decide what to filter
Questions and answers

Practical questions

It says the effect depends on how much human data the model already has. AI text helps models with fewer than about 10 human tokens per parameter and raises loss for models at or past the compute-optimal 20, which is where production models sit.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

When AI Research Sources Disagree, What Can You Publish?

Resolve conflicting AI research claims by checking dates, definitions and primary evidence. Publish a supported answer or show exactly what remains uncertain.

September 12, 2026 · 5 minRead
AI Development

AI Research Claims: What Has Actually Been Verified?

Assess AI research claims with a practical evidence matrix. Separate formal proofs, measured results and demos, and record what each check establishes.

September 9, 2026 · 6 minRead
AI Development

An AI Agent Read an Old Web Page: How to Spot the Gap

Check whether an AI research answer uses the right web revision. Separate retrieval time, publication date and effective date when primary pages disagree.

September 6, 2026 · 4 minRead
AI Development

AI Research Sources: Original, Syndicated or Repeated

Trace AI research claims to their original evidence. Distinguish copies, new analysis and independent observations without discarding useful follow-ups.

September 6, 2026 · 6 minRead
AI Development

AI Video Generation 2026: Omni vs Sora vs Veo 3 Compared

Gemini Omni, OpenAI Sora 2, and Google Veo 3.1 compared for video — quality, per-second cost spread of 17x, and the September 24 Sora API sunset clock.

May 22, 2026 · 15 minRead
AI Development

Google Intelligent Eyewear: Gemini AI Glasses Fall 2026

Google announces Gemini-powered smart glasses with Samsung, Gentle Monster, and Warby Parker at I/O. Audio glasses ship fall 2026; display tier TBD.

May 20, 2026 · 18 minRead
Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source