Tagged "ai-crawlers"
Cross-cutting reads on this topic
From September 15, 2026 new ad-supported Cloudflare domains block AI agents on ad pages and refuse AI training by default. A 20-bot census of who is affected.
#AI Crawlers#Cloudflare+2 more
2026-09-15
Read Article
User-agent strings are trivial to spoof and many user-triggered fetchers ignore robots.txt. How to verify AI crawlers and enforce access at the edge.
#ai-crawlers#robots-txt+5 more
2026-08-13
Read Article
Digiday reports Perplexity blocked Time's markdown ads, copy served to AI crawlers but not to humans, as deceptive. Why this is a cloaking question.
#ai-search#perplexity+5 more
2026-08-13
Read Article
Cloudflare's AEO Visibility Dashboard shipped 6 August in early access with no published price, and it probes only two assistants: Claude and GPT.
#cloudflare#aeo+5 more
2026-08-08
Read Article
Adoption of llms.txt splits hard by cohort, and the best traffic studies found no real AI-bot requests. Who reads the file, and a fifteen-minute setup.
#llms-txt#ai-search-visibility+4 more
2026-08-07
Read Article
Blocking AI crawlers is now an economic call. Cloudflare's Pay Per Crawl beta lets you charge per request; crawl-to-referral ratios show what each bot returns.
#ai-crawlers#pay-per-crawl+5 more
2026-08-06
Read Article
During a pre-launch audit of an independent writer's publication, every archive post returned a 302 to crawlers. How to detect and fix this AI-crawler wall.
#AI crawlers#Common Crawl+6 more
2026-07-14
Read Article
DCN sent Common Crawl a cease-and-desist on June 3, 2026. Why the WARC archive makes removal hard and why blocking CCBot alone won't protect your content.
#Common Crawl#AI training data+6 more
2026-06-12
Read Article
Eight major AI crawlers, eight different calls. A bot-by-bot matrix to block training crawlers like GPTBot and ClaudeBot while keeping AI-search citations.
#ai-crawlers#robots-txt+6 more
2026-06-04
Read Article
Log files capture what crawlers actually do — not a simulation. A 2026 guide to crawl-budget analysis and the new training-vs-retrieval AI crawler split.
#log-file-analysis#crawl-budget+6 more
2026-05-30
Read Article
Why agentic crawlers prefer Markdown, the llms.txt + AGENTS.md pattern, and file-tree spec for agent-ready sites. Reference implementation + audit.
#llms-txt#agents-md+6 more
2026-04-27
Read Article
How GPTBot, ClaudeBot, Google-Extended, and PerplexityBot crawl real sites — frequency, depth, paths preferred, and refresh cadence. Server-log evidence.
#ai-crawlers#gptbot+8 more
2026-04-26
Read Article