AI DevelopmentMethodology6 min readPublished October 1, 2026

Years of tested advice, filed so a model can find the sentence and its source

Turn Your Blog Archive Into a Knowledge Base AI Can Use

Your old posts hold years of tested advice. How to turn a blog archive into a sourced knowledge base that AI tools can use for briefs, rewrites and links.

DA
Digital Applied Team
Research and practical guidance
CoverageOctober 1, 2026

Most companies that have blogged for a few years own something they never use: a record of what they advised, when, and on what evidence. AI writing tools do not see it. They see the brief, a search result or two, and whatever the model remembers. The fix is not to feed the model every post. It is to extract what the posts claim into records that carry a source and a date, then put those records where a tool can retrieve them. This is the method.

Editorial note: A general method, written October 3 as an October 1, 2026 dispatch. It uses no corpus, count or result from any client engagement. The record schema is illustrative.

Key takeaways
  1. 01
    Inventory from the outsideThe CMS list is what you meant to publish. The sitemap plus a crawl is what readers and crawlers can reach. Start from the second.
  2. 02
    Records, not postsA knowledge base is a set of claims with a URL and a date each, not a folder of articles. Extraction is the work.
  3. 03
    Staleness is a fieldEvery record carries the date the claim was true and a confidence tag. Retrieval without those two fields repeats old advice with new confidence.
  4. 04
    Small enough to read wholeAnthropic’s own retrieval guidance says a base under about 200,000 tokens can go into the prompt directly. Many archives, once extracted, fit.

01 — Step oneInventory from the sitemap and a crawl

The CMS export is the wrong starting list. It contains drafts, posts that redirect, pages that were unpublished but still indexed, and nothing about the URLs a migration left behind. Pull the XML sitemap, then run a crawl from the home page, and keep the union. Every URL that appears in one list and not the other is a finding before the work has begun: an orphan the crawl could not reach, or a live page the sitemap forgot.

For each URL, record the publish and modified dates from the page itself rather than from the database. The schema.org Article vocabulary defines datePublished as the date of first publication and dateModified as the date the work was most recently changed, and most sites emit both in structured data. They are the two fields that make staleness computable later. If a site’s dates are missing or wrong, that is the first repair, because nothing downstream can be trusted without them.

02 — Step twoExtract claims, advice and examples

Read each post for three kinds of sentence and ignore the rest. A claim is a statement about the world that could be checked: a figure, a rule a platform enforces, a date. Advice is a recommendation the post makes, with its condition if it states one. An example is a worked case, a template or a before-and-after that the advice points at. Introductions, transitions and conclusions are not extracted. They are the part of a post that the model can write again; the three kinds above are the part it cannot.

This is a reading job that a model can do at scale and a person must sample. Give the model one post and the schema below, ask for records only, and have an editor check one post in ten against the original. The two failure modes to look for are invented precision, where a vague sentence becomes a number, and dropped conditions, where “if the site is under a thousand pages” falls off the front of the advice. Our note on preserving facts through AI rewrites describes the same two failures from the other direction.

03 — The artefactOne record schema

Keep the schema small enough that a model fills it reliably and a spreadsheet can hold it. These are the fields we would start with for any archive; a specialist site adds its own.

An illustrative record schema. Field names are generic; the example values are invented to show the shape.
FieldWhat goes in itExample
typeclaim, advice or exampleadvice
textThe statement in one sentence, in the post’s own terms, with its condition kept.Add a visible last-reviewed date to any page that cites a platform rule.
source_urlThe post it came from, and the section anchor where one exists./blog/example-post#review-dates
evidenceWhat the post cited for it: a primary, a vendor figure, an observation, or nothing.vendor documentation, linked
as_ofThe date the claim was true, from the post’s published or modified date.2025-03-14
confidencehigh, medium or low, set by the evidence field and the editor’s sample.medium
stalenessstable, dated or expired, judged on whether the subject changes yearly, quarterly or never.dated
topicsTwo or three terms from a taxonomy you control, not free tags.content-review, platform-policy

Two fields do most of the work. The evidence field is what lets a later reader tell a measured figure from a remembered one. The as_of field is what lets a tool say “this was true in March 2025” rather than asserting it as current. A base without them is a quote bank; with them it is a record.

04 — Step threeTag, group and deduplicate

An archive repeats itself. The same advice appears in a 2023 guide, a 2024 checklist and a 2025 refresh, worded three ways and sometimes contradicting itself in a detail. Group records by topic, then within a topic sort by as_of and read the duplicates together. Keep the newest version that has evidence; mark the older ones as superseded rather than deleting them, because the change itself is information. A claim that moved from “always” to “usually” between two posts is worth a sentence in the next one.

The taxonomy should be small, with a few dozen terms at most, and written before tagging starts. Free tags drift into synonyms within a week. If the archive already has a category and tag structure, start from it and prune; our keep, rewrite or retire method is the companion pass for the posts themselves, and the two share a spreadsheet comfortably. If the immediate job is a bulk rewrite rather than a standing base, our note on indexing the facts before an agent rewrites an archive is the narrower, faster version of the same extraction.

A size check worth doing early

Anthropic’s engineering note on contextual retrieval makes a practical point before it gets to retrieval: if a knowledge base is smaller than about 200,000 tokens, roughly 500 pages, it can simply be included in the prompt. An archive of a few hundred posts often extracts to fewer tokens than that, because the records are a fraction of the prose. Measure before building a retrieval layer. Many teams do not need one.

05 — PayoffFour things the base is for

The base earns its cost in four jobs that every content team already does by memory and search.

01
Briefs
Before a new post is written

Pull every record on the topic. The brief now says what the company has already claimed, with dates, so the new post extends the position rather than restating or contradicting it.

Retrieve by topic
02
Refreshes
When a post is due for review

Filter records by staleness and as_of. The list of expired claims is the edit list, and the evidence field says which ones need a new source rather than a new sentence.

Filter by date
03
Internal links
While drafting

A record that supports a sentence in the draft is a link candidate to its source_url. The anchor text is the record’s text; the link lands on the section, not the home page.

Match by claim
04
Voice and position
When a model drafts

The advice records, read together, are the company’s stated positions. Hand them to the model with the brand voice guide and the draft stops inventing opinions the company never held.

Ground the draft

The fourth use pairs with a brand voice guide: the guide says how the company sounds, the base says what it thinks. A model given only the first writes fluent pages with no spine. Given both, it writes pages the editor recognises.

06 — Serving itHow AI tools read it

Once the records exist, there are three ways to put them in front of a tool, and the right one depends on size and on who is asking.

Under ~200K tokens and used by your own team
Put the whole base in the prompt or a project file. Anthropic’s note says this is the simplest option at that size, and prompt caching makes repeated use cheap.
Whole base
Larger, or many topics queried separately
Index the records in a retrieval store. OpenAI’s Retrieval API, for example, is built on vector stores that files are uploaded into; Anthropic’s note describes adding chunk context and a keyword index to cut failed retrievals.
Retrieval
Outside agents should find your positions
Publish a curated summary as an llms.txt file: an H1, a short blockquote, then sections of links with one-line notes, per the proposal. Point each link at the post, and keep the file small enough to fit in a context window.
llms.txt

The third option deserves a caveat the proposal itself makes. The llms.txt format is for inference, when an agent needs information about a topic while helping a user. It is not a search-ranking input, and the proposal does not claim to be one. Our review of llms.txt in practice has the adoption evidence. Publish one because it is a cheap, readable index of what the company stands behind, not because it will move a ranking.

Where a team wants this built as a repeatable system rather than a one-off spreadsheet, our content engine work sets up the extraction, the editor sample and the refresh filter so the base stays current as the archive grows.

Next step

Extract ten posts by hand before you automate anything

Take the ten most-linked posts, fill the schema for each, and read the records together. That afternoon tells you whether the taxonomy holds, how often the evidence field is empty, and whether the archive is worth the rest of the work. Usually it is.

AI content operations

Make your archive the source your AI tools draw from

Digital Applied builds the extraction, review and refresh workflow that turns an archive into a sourced knowledge base your writers and models share.

Sourced recordsEditor samplingRefresh filters
Start with ten posts

The pilot

  • →Sitemap plus crawl inventory
  • →One schema, filled by hand
  • →Confidence and staleness tags
  • →A brief written from the records
Questions and answers

Practical questions

Because the tool then retrieves prose, not claims, and cannot tell a 2022 figure from a current one. Extracting records with a source and an as-of date is what makes the archive usable rather than merely available.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Give an Agent the Facts Before It Rewrites Your Archive

Before an AI agent rewrites hundreds of pages, index every claim, figure and source already in them. The schema, three rules, and what to do when a claim fails.

September 21, 2026 · 7 minRead
AI Development

What an OpenRouter Usage Chart Can and Cannot Show You

OpenRouter documents what its rankings count and what they exclude. A source-pinned inventory of what a usage chart measures, and what it cannot tell you.

August 26, 2026 · 18 minRead
AI Development

Can a Buyer Reproduce a Vendor Benchmark Row?

We collected 42 vendor-published benchmark rows from 9 vendors and checked what each page discloses: version, subset, harness, effort. Full table in the post.

August 16, 2026 · 20 minRead
AI Development

RAG Chunking Strategies: A 2026 Retrieval Playbook

A practical playbook for chunking documents in RAG pipelines, comparing fixed, semantic, recursive, and late-chunking with retrieval-quality benchmarks.

May 27, 2026 · 12 minRead
AI Development

AI Browser Landscape 2026: Atlas vs Comet vs Arc vs Dia

AI browser landscape 2026 — Atlas, Comet, Arc, Dia, Brave Leo, and Opera Neon. Feature matrix, market share estimates, and how agencies should prepare.

April 16, 2026 · 16 minRead
AI Development

AI Agent Marketplaces 2026: Discovery and Distribution

AI agent marketplace landscape — Claude Skills, GPT Store, MCP Hubs, Hugging Face Spaces, Replit Agent Market. Distribution strategy for agency builds.

April 16, 2026 · 16 minRead
Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source