AI DevelopmentPricing Tracker6 min readPublished October 2, 2026

Four ways to rent search over your own documents

Managed RAG Services Compared: Cloudflare, Google, AWS

Cloudflare AI Search is now generally available. How it compares with Google Agent Search, Bedrock Knowledge Bases and OpenAI file search on price and limits.

DA
Digital Applied Team
Research and practical guidance
CoverageOctober 2, 2026

Cloudflare made AI Search generally available on October 1, 2026, with usage billing from November 1. It is a managed retrieval service: upload documents or point it at a site, and it handles the chunking, embedding, indexing and search an AI assistant needs to answer from your own content. This page compares it with Google Agent Search, Amazon Bedrock Managed Knowledge Base and OpenAI file search, priced from each vendor’s own pages.

Key takeaways
  1. 01
    Cloudflare is cheapest per query$0.75 per 1,000 semantic queries against $1 to $4 at Google and AWS and $2.50 at OpenAI.
  2. 02
    Billing starts November 1Cloudflare AI Search is GA, but usage is not billed until November 1, 2026.
  3. 03
    Storage is measured differentlyRaw data at Google and AWS, chunks plus embeddings at OpenAI. The same corpus bills differently.
  4. 04
    Connectors split the fieldAWS connects to SharePoint, Confluence and Drive. OpenAI takes uploads only.

01 — The releaseWhat Cloudflare made generally available

AI Search started as an open beta called AutoRAG in April 2025. Cloudflare’s GA announcement reduces the bill to three meters: content ingested, data stored and queries run. The limits and pricing page sets them at $0.75 per million ingestion tokens, $2 per GB-month of storage, and $0.75 per 1,000 semantic, vector or hybrid queries, or $0.10 for plain full-text queries. Embedding and reranking with Cloudflare’s own Workers AI models are included rather than billed separately.

Each month is free up to 5 million ingestion tokens, 10 GB of storage, 1,000 semantic queries and 1,000 full-text queries. The GA release also raised the size limit for text files and PDFs read with OCR to 10 MiB, from 4 MiB. Hybrid search, which combines keyword and vector matching, is on by default for new instances.

The date that matters

Usage before November 1, 2026 is not billed. A team that wants a free month of real traffic has October; a team planning a budget should price from November, when the free allowances above become the only free usage.

02 — ContextWhat is being compared

Two of the four need a naming note, which matters when searching for their docs. Google renamed Vertex AI Search to Agent Search on April 22, 2026, and says the functionality is unchanged. AWS now sells two kinds of knowledge base. The table uses the newer one, which is priced per unit like Cloudflare’s.

A
Cloudflare AI Search
GA October 1, 2026

Ingestion, storage and query meters. Sources are uploads, R2 storage or your own site.

Anchor
B
Google Agent Search
Formerly Vertex AI Search

Per-query and per-GiB pricing, or a subscription. Google’s RAG Engine is a separate, assemble-it-yourself option with no list price of its own.

Google Cloud
C
Bedrock Managed Knowledge Base
GA June 17, 2026

Per-GB and per-call pricing. The older customer-managed knowledge base bills the vector store you run instead.

AWS
D
OpenAI file search
Vector stores, Responses API

Per-GB-day storage and per-call search, plus model tokens. Upload only, no connectors.

OpenAI

Microsoft’s Azure AI Search is left out of the tables. It bills by provisioned search units or compute hours rather than per query, so its prices do not line up with the four above.

03 — The dataStorage and query prices

Query prices are per 1,000. Storage is per month except at OpenAI, which bills per GB per day; at 30 days that is about $3 per GB-month. The storage figures do not measure the same thing: Google and AWS count the raw data you load, Cloudflare counts what sits in its index, and OpenAI counts the parsed chunks plus their embeddings.

Sources: Cloudflare, Google Cloud, AWS and OpenAI pricing pages, read October 3, 2026. US dollars; queries per 1,000.
ServiceStorageQueries per 1,000Free each month
Cloudflare AI Search$2.00 per GB-month$0.75 semantic or hybrid; $0.10 full-text5M ingestion tokens, 10 GB, 1,000 + 1,000 queries
Google Agent Search (Standard)About $5 per GiB-month$1.5010,000 queries, 10 GiB
Google Agent Search (Enterprise)About $5 per GiB-month$4.00, generative answers included10,000 queries, 10 GiB
Amazon Bedrock Managed Knowledge Base$5.00 per GB-month$1.00; agentic $4.00 plus $1.00 per underlying callNone stated
OpenAI file search$0.10 per GB-day$2.50 per 1,000 tool calls, plus model tokens1 GB of storage

Google’s 10,000 free queries are described as a monthly trial allowance per account and exclude its advanced generative answers; the same page’s annual example deducts them only once, so treat the Google totals below as a best case. AWS publishes no free tier for the managed knowledge base. Google also offers a subscription model with a minimum commitment of 1,000 queries a minute and 50 GB of storage, aimed at steady high volume.

04 — Worked exampleOne workload, four bills

Take 10 GB of documents and 100,000 retrieval queries a month. Apply each list price and free allowance, and leave out ingestion, generation and model tokens. This is our arithmetic, not a vendor quote. Cloudflare: storage fits the free 10 GB, and 99,000 billable queries cost $74.25. Google Standard: storage is free, and 90,000 billable queries cost $135. AWS: $50 of storage plus $100 of queries is $150. OpenAI: 9 GB beyond the free one for 30 days is $27, and 100,000 calls are $250, a total of $277 before model tokens.

Monthly cost for 10 GB and 100,000 queries, US dollars, lower is cheaper

Digital Applied arithmetic from list prices read October 3, 2026. Excludes ingestion, generation and model tokens; OpenAI also bills model tokens on top. Storage bases differ by vendor.
Cloudflare AI Searchsemantic queries
$74.25
Google Agent SearchStandard
$135
Bedrock Managed KBstandard retrieval
$150
OpenAI file searchplus model tokens
$277
Google Agent SearchEnterprise
$360

The ranking flips with the shape of the workload. A large corpus queried rarely favours the cheapest storage; a small corpus queried constantly favours the cheapest query. AWS’s own pricing page gives a larger example, 50 GB and 100,000 standard retrievals at $350 a month, or $850 with agentic retrieval, in which a managed model plans the retrieval and each underlying search is billed as well.

05 — The dataIngestion, files and sources

Sources: vendor pricing pages and documentation, read October 3, 2026.
ServiceMax fileIndexing costsData sources
Cloudflare AI Search10 MiB text and OCR PDFs; 4 MiB others$0.75 per 1M tokens (+$0.50 for images); Workers AI embedding and reranking includedUpload, R2 buckets, your own site on the same Cloudflare account
Google Agent Search200 MBIncluded in storage; Layout Parser $10 per 1,000 pages; Ranking API $1 per 1,000Websites, Cloud Storage, BigQuery and other Google databases, Workspace; third-party connectors no longer supported
Bedrock Managed Knowledge Base30 MB of extracted textManaged parsing, embeddings and reranker at $0S3, SharePoint, Confluence, Google Drive, OneDrive, web crawler, custom
OpenAI file search512 MB; 5M tokens per fileNot priced separatelyFile upload only

Connectors are where the products differ most. Agent Search has dropped outside data sources, according to Google’s docs, and those connectors now live in its Gemini Enterprise product. AWS’s managed knowledge base reads from SharePoint, Confluence, Google Drive and OneDrive and can filter results by each document’s permissions, which its documentation says the customer-managed version cannot. Cloudflare crawls only websites on the same Cloudflare account. For a team turning its own content into a source, our method for turning a blog archive into a knowledge base covers the preparation that comes before any of these services.

06 — The dataScale, rate and data limits

Sources: vendor quota, limits and data-control documentation, read October 3, 2026.
ServiceScaleRateData controls
Cloudflare AI Search1M files per instance (500K with hybrid); 5,000 instances on paid plansNot publishedNo index residency statement; EU-jurisdiction R2 buckets accepted as a source
Google Agent Search10M documents per project per location300 searches a minute per project and locationData at rest in us or eu multi-regions, or global
Bedrock Managed Knowledge Base10 TB per knowledge base; 200 data sources each100 retrievals a second per account; 600 a minute per knowledge baseEight regions including GovCloud; document-level permission filtering
OpenAI file searchNot published on the pages read100 to 1,000 a minute by usage tierData residency supported; vector stores not eligible for zero retention

OpenAI’s data controls page lists vector stores as not eligible for zero data retention: stored files stay until they are deleted. Its retrieval guide documents an expiry setting that deletes a store and stops its charges. Whichever service you choose, deletion is part of the design, as our guide to where deleted agent data survives explains.

07 — Practical implicationsWhich one to use

Content on your own site or in R2, high query volume
Cloudflare AI Search, priced from November
Cloudflare
Documents in SharePoint, Confluence or Drive
Bedrock Managed Knowledge Base with permission filtering
AWS
Data already in BigQuery or Cloud Storage
Google Agent Search, Standard edition first
Google
A small corpus behind an OpenAI assistant
File search, with an expiry on every store
OpenAI

Retrieval quality is not on this page, because none of the vendors publishes a comparable measure and we did not run one. Test with 50 real questions from your users before committing, and check that each answer cites the right document. For agents that also need the open web, the sibling comparison of web search APIs for AI agents prices that side. Our AI transformation work covers the evaluation, the build and the cost controls.

08 — MethodMethod and as-of date

Methodology

A comparison of published prices and documented limits. Nothing on this page was benchmarked by Digital Applied.

What was collected
Storage, query, ingestion and add-on prices, free allowances, file size limits, scale and rate quotas, data sources and data controls for four managed retrieval services, with Google’s Standard and Enterprise editions as separate price rows.
Sources
Each vendor’s own pricing pages, quota pages, product documentation and announcements: Cloudflare’s GA post and AI Search docs, Google Cloud’s Agent Search pricing, quota and release notes, AWS’s Bedrock pricing, What’s New post and user guide, and OpenAI’s pricing and retrieval guides.
As-of date
October 3, 2026. Cloudflare’s pages were updated October 1 and Google’s release notes and AWS’s pricing page September 30. Google’s pricing and quota pages, AWS’s documentation and OpenAI’s pages show no content date, and OpenAI’s prices could not be confirmed against an archived copy from before October 2.
Units
US dollars. Queries per 1,000. Google prices storage per GiB-hour; the GiB-month figure is the page’s own rounded example. OpenAI’s per-day storage is converted at 30 days.
Exclusions
Azure AI Search, which bills by capacity rather than per query; Google RAG Engine, which has no list price of its own; the customer-managed Bedrock knowledge base, whose cost is the vector store you provision; enterprise and committed-use discounts.
Limitations
No retrieval quality, latency or indexing speed was measured. Cloudflare publishes no query rate limit or index residency statement. OpenAI’s file and store counts were not on the pages read.
Refresh
Re-read every pricing page monthly, and after November 1 when Cloudflare billing begins. Correct any figure in place with a dated note.
Next step

Price your corpus and query volume before choosing

Measure two numbers first: how many gigabytes you will index and how many queries a month you expect. Run them through the four price rows, rule out any service that cannot reach where your documents live, and test the survivors on real questions.

Agentic AI implementation

Ground your assistant in your own documents

Digital Applied prepares the content, chooses the retrieval service on your real workload and builds the assistant that cites it.

Content preparationCost modellingRetrieval testing
Before you choose

Know four numbers

  • →Gigabytes to index
  • →Queries per month
  • →Where documents live
  • →Who may see each one
Questions and answers

Practical questions

Usage-based billing begins November 1, 2026. The service became generally available on October 1, and usage before November 1 is not billed.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Web Search APIs for AI Agents Compared: Price and Limits

Cloudflare added a Web Search API to AI Gateway on October 2. How it compares with search tools from OpenAI, Anthropic, Google and others on price and limits.

October 2, 2026 · 6 minRead
AI Development

Cohere Embed 5: Pro and Fast Embedding Models Compared

Cohere Embed 5 comes in Pro and Fast versions with 128K context. How the two compare on price, speed and use cases, and when to switch embedding models.

September 30, 2026 · 5 minRead
AI Development

AI Model API Price Changes, Q3 2026: Five Labs, One Table

API price changes from Anthropic, OpenAI, Google, DeepSeek and xAI, July to September 2026: launches, cuts, a rise and promotions per million tokens, dated.

October 3, 2026 · 4 minRead
AI Development

Open Decision Models Compared: Clef, Decider 2B and Jev

Cloudflare's Clef and Amazon's Decider 2B put decision models like closed Jev into open weights. Licence, size, context, latency and benchmarks in one table.

October 1, 2026 · 6 minRead
AI Development

Cloudflare Blocks AI Agents on Ad Pages: Which Bots Are Hit

From September 15, 2026 new ad-supported Cloudflare domains block AI agents on ad pages and refuse AI training by default. A 20-bot census of who is affected.

September 15, 2026 · 10 minRead
AI Development

A Proxy Stripped One Header and Claude Code Paid Twice

Claude Code v2.1.239 fixed what its changelog calls silently doubled billed API calls: behind a proxy stripping Content-Type, it re-ran turns non-streaming.

August 21, 2026 · 18 minRead
Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source