Free AI tier data terms are where vendor privacy policies diverge most — and the divergence is almost never on the pricing page. This census reads the published terms of eleven hosted model APIs and records, in each vendor’s own words where a page states them, what happens to prompts sent on a free tier, a trial key, a free quota, or a promo-priced window.
The stakes are concrete. Two vendors state outright that their free tier runs on different data terms than their paid tier. Two more have no free tier to compare at all, which is itself a finding. One vendor’s policy reads as training-by-default across every tier. And on the day this census was compiled, a promo-priced model at the center of the week’s news cycle raised the question all over again: does a discount change what happens to your data?
What follows is the method, the complete eleven-row table, the two clean free-versus-paid asymmetries quoted verbatim, a five-way typology of vendor postures, and the two places where a gateway’s summary of a vendor disagrees with the vendor’s own documentation — arguably the most useful finding in the set.
- 01Free tiers are where the terms actually diverge.Google and Mistral both state, in their own published words, that the free tier runs on a different training default than the paid tier. Google’s free Gemini API tier uses submitted content to improve products, with disclosed human review; the paid tier does not.
- 02Every fixed retention day-count is 30 days.Four of the eleven vendors publish a fixed day-count for API prompt retention — OpenAI, Anthropic (covered models), Cohere, and Groq — and all four figures are 30 days, two of them phrased as ‘up to’. The other seven publish a limited-period phrase, a user-configurable window, or no day-count this census could locate.
- 03Two vendors have no free tier to compare.OpenAI and Anthropic run trial credits under the same commercial terms as pay-as-you-go. ‘There is no separate free-tier policy’ is the accurate cell for both — a different finding from a vendor that publishes one policy covering a real free product.
- 04The gateway and the vendor disagree twice.OpenRouter’s provider table says Mistral prompts are retained 30 days while Mistral’s own docs describe a configurable Never-to-one-year window; it marks Together as zero-retention while Together’s own page frames ZDR as an active opt-in. This census records both sides rather than picking one.
- 05Coding agents are excluded by design.This is the hosted-model-API companion to our 17-agent coding-agent data terms census — the same documentation-only form applied to a different population. Nothing from that table is restated here, and no coding agent appears as a row.
01 — The QuestionThe cheapest tier is the least read tier.
Free tiers, trial keys, and promo windows are how most teams first touch a hosted model API — and they are precisely the surfaces where data terms are least likely to be read. A paid enterprise contract gets a legal review. A free API key gets pasted into a prototype the same afternoon. If a vendor’s terms treat those two surfaces differently, the difference lands exactly where nobody is looking.
The pattern this census found is not a scandal; it is an asymmetry of disclosure. Some vendors state the free-versus-paid split plainly and in detail. Some publish one policy and say it covers everything. Some publish a policy that never mentions the free surface at all. And for one vendor, no vendor-primary page could be located at all. The census’s job is to separate those four situations — because “the vendor discloses a training default we dislike” and “the vendor discloses nothing” are different procurement facts that call for different responses.
Hosted model APIs
OpenAI, Google (Gemini API), Anthropic, Mistral, Alibaba Cloud Model Studio, Z.ai, DeepSeek, xAI, Cohere, Groq, and Together AI. OpenRouter is examined separately as the aggregator layer, not counted as a row.
Every fixed day-count
Four of eleven vendors publish a fixed retention day-count for API traffic, and every one of the four figures is 30 days — two of them hedged as ‘up to 30’.
Pages with no absolute date
Four of the cited primary pages carry no absolute last-updated date — Groq’s data page, OpenRouter’s logging docs, a Mistral help article, and Alibaba’s free-quota page. The citations below say ‘undated’ rather than implying an as-of date that cannot be supported.
02 — MethodologyScope, sourcing, and a four-way vocabulary.
Scope. This census covers hosted model APIs only — the raw inference endpoints a developer calls with an API key. Coding agents and harnesses are excluded entirely. That population — seventeen agents, six fixed questions — is our coding-agent data terms census from August 17, and this post is its hosted-model-API companion: the same documentation-only form applied to a different population, not an extension of it. If you came here for what a coding agent does with your code, that is the table you want.
Sourcing rule. Where a commonly repeated figure could not be confirmed on a fetched vendor-owned page, the cell reads not established — never an inference from the paid tier, a sibling product, or a forum report. Where no vendor-primary page was located for a row at all, the row says so rather than borrowing terms from a sibling surface.
Vocabulary. Each cell uses a four-way standard: disclosed (the vendor’s own page states the behaviour, in words quoted here), not disclosed (the page was located and read; it does not address the question), not established (a claim circulates but was not confirmed on a fetched vendor-owned page), or page not located (no vendor-primary page was found for that specific question). Four cited pages carry no absolute last-updated date and are labelled undated. And this is deliberately a census — rows and citations — not a checklist; the questions-to-ask form of this same concern is our AI procurement questions guide, and the single-vendor deep-dive form is our look at Fable 5's 30-day retention and enterprise ZDR.
03 — The DatasetThe complete table: eleven vendors, row by row.
The table below is the asset — all eleven rows, no summary substitution. Cells read not disclosed, not established, or page not located per the vocabulary above — never an inference. Data as of August 26, 2026, from vendor-owned pages only.
| Vendor · primary source | Free or promo surface | Trains on your inputs? | Retention (vendor-stated) | Opt-out / ZDR |
|---|---|---|---|---|
| Rows 1–2 · Explicit free-vs-paid split, in the vendor’s own words | ||||
| Google — Gemini APIai.google.dev/gemini-api/termsUpdated Apr 28, 2026 | Free / unpaid tier of the Gemini API and AI Studio — a named free product with its own terms section | Free tier: yes — content used to “provide, improve, and develop Google products and services”; human review disclosed. Paid tier: prompts and responses not used to improve products | Paid tier: logged “for a limited period of time” for abuse prevention — no day-count. Free tier: no day-count published | Regional, not toggled: EEA / Switzerland / UK users get the paid-tier data terms even on the free service |
| Mistral — La Plateformehelp.mistral.ai · article 455207Undated page | Free “Experiment” tier vs paid pay-as-you-go / Scale | Free tier: opted in by default — users “may opt out of training” manually via the admin Privacy menu. Pay-as-you-go: opted out by default | Chat retention user-configurable: Never / 30 / 60 / 90 / 180 days / 1 year — no single fixed default (see the gateway tension in section 06) | Manual toggle: "Anonymous improvement data" (Studio/API) or the training toggle (Vibe) |
| Rows 3–4 · No separate free tier — trial credits under the same terms | ||||
| OpenAI — API Platformopenai.com/policies/api-data-usage-policiesUpdated Jan 8, 2026 | No free-tier API product — new-account trial credits run under the same policy as pay-as-you-go | “We do not train our models on your data by default” — stated once for the API Platform and business products as one list, with no free-vs-paid carve-out | API inputs and outputs retained “up to 30 days” to provide the services and identify abuse, then removed unless legally required | ZDR by request — approval-shaped, not a free-tier toggle |
| Anthropic — Claude APIanthropic.com/legal/commercial-termsEffective Jun 17, 2025 | No free-tier API product — console trial credits run under the Commercial Terms | “Anthropic may not train models on Customer Content from Services” — uniform, no tier carve-out | Covered models: retained 30 days to support safety work (per the privacy center). Standard-tier day-count: not established | Human review: “By default, no Anthropic personnel can read your retained conversations” — controlled access path only |
| Rows 5–10 · One policy, no tier carve-out (xAI: no page located) | ||||
| Z.ai — GLM APIdocs.z.ai/legal-agreement/terms-of-useUpdated Apr 14, 2026 | API including the promo-priced GLM-5.3-Flash window — promo pricing appears nowhere in the data terms (see section 07) | API surface: “We will not use End User Content for developing or improving Services, unless you explicitly agree to such use.” Consumer surface differs — training-permissive | Not disclosed — no day-count found | Training use is opt-in by construction on the API surface |
| DeepSeek — API + appcdn.deepseek.com · privacy policyUpdated Feb 10, 2026 | One policy across app and API — no tier distinction stated for training purposes | Reads as yes by default: data processed “to train and improve our technology, such as our machine learning models and algorithms” | No fixed day-count — retained “for as long as necessary” per stated purposes | Policy states a “right to opt-out” of training use. Human review: not disclosed |
| xAI — Grok APINo primary page located | Page not located — no distinct free-tier API terms page was found; the “free tier” language xAI publishes concerns the consumer Grok app, not the API | Page not located — no vendor-primary page stating a training posture for the Grok API was located this pass | Page not located | Page not located |
| Cohere — Trial keyscohere.com/data-usage-policyDated Dec 5, 2025 | Trial API keys — explicitly governed by the same Terms of Use and Privacy Policy as production | Opt-out mechanism disclosed (“in your dashboard settings at any time”); the trial-key default direction itself is not disclosed | “We automatically delete logged prompts and generations after 30 days,” absent legal or contractual need | Enterprise Data Commitments scoped to “commercial, paying customers” only — trial users textually excluded |
| Groq — GroqCloudconsole.groq.com/docs/your-dataUndated page | Free developer usage — the no-training wording is account-scoped, with no tier carve-out on the page | No — Groq “is not permitted to use Inputs or Outputs for training or fine-tuning” absent explicit customer permission | “Up to 30 days,” and only when troubleshooting reliability or investigating suspected abuse — conditional, not an always-on log | “All customers may enable Zero Data Retention (ZDR) in Data Controls settings” — not enterprise-gated |
| Together AItogether.ai/privacyUpdated Dec 17, 2025 | Free credits widely offered — no separate data-terms carve-out found; one policy for all | No — “We do not use any data collected from you to train our models without your explicit opt-in and consent” | No fixed day-count on this page — retained “only for as long as is necessary” | ZDR is an active choice, not a default: “By choosing ‘No’, you are enabling Zero Data Retention” (see section 06) |
| Row 11 · A quota, not a tier | ||||
| Alibaba Cloud Model Studio (Qwen)alibabacloud.com · new-free-quotaUndated page | Time-boxed free quota, not a standing tier: “The free quota is valid for 90 days,” and eligibility is limited to models in the China (Beijing) region and models in the Singapore region | Not established — a no-training claim circulates in secondary sources but was not confirmed on a fetched vendor terms page this pass | Not established | Not established. A circulating claim that the developer free tier ended in April 2026 is not reflected on the vendor’s quota page, which describes the quota as live |
One scope note on row 11: Alibaba Cloud Model Studio’s raw API is a different commercial surface from OpenRouter’s hosting of Qwen models, and the hosted Qwen3.8-Flash is a different artifact again from the open-weight Qwen3.8-Flash-Next — a split we unpack in our companion piece on the Qwen open-weights-versus-hosted split. This row is about Model Studio’s own terms and nothing else.
04 — The Clean SplitsGoogle and Mistral, verbatim.
Only two of the eleven vendors publish an explicit free-versus-paid data-terms split under their own name — and both deserve credit for the disclosure, because the disclosure is exactly what lets a buyer decide.
Google is the marquee example. The Gemini API Additional Terms of Service (updated April 28, 2026) state that for the free tier, “Google uses the content you submit to the Services and any generated responses to provide, improve, and develop Google products and services,” and that “human reviewers may read, annotate, and process your API input and output.” For paid services, the same document flips the default: “Google doesn’t use your prompts (including associated system instructions, cached content, and files such as images, videos, or documents) or responses to improve our products,” with logging “for a limited period of time, solely for detecting and preventing violations of the Prohibited Use Policy.” One document, two regimes, both in plain language. There is also a regional override: users in the EEA, Switzerland, and the UK get the paid-tier data-protection terms even on the free service.
Mistral is the second clean case, with a wrinkle. Its help center states that free Experiment-tier users may opt out of training but must do so manually, while pay-as-you-go customers are opted out by default — the same shape as Google’s split, expressed as toggle defaults. The wrinkle: a second Mistral page on privacy and data controls says flatly that “data sent through the API isn’t used for model training,” which sits in tension with the help article’s free-tier opt-in default. Per this census’s method, we record both statements rather than silently picking one; the two pages may describe different surfaces, but the vendor’s own docs do not say which.
05 — The PatternsFive vendor postures, not two.
“Does the free tier train on your data?” sounds like a yes/no question. Across the ten vendors with a located primary page it resolves into five distinct postures — and knowing which posture a vendor occupies tells you what to check next. A sixth posture below belongs to the aggregator layer, not to any vendor row. The patterns overlap at the edges: Anthropic, for instance, sits in both the no-carve-out group and the no-free-tier group, because its trial credits run under one uniform commercial document. xAI is placed in no posture at all: with no vendor-primary page located, there is no disclosed posture to place it in.
Explicit asymmetry
The vendor names the free tier and states a different training default than the paid tier, in its own words. The cleanest situation for a buyer: the trade is disclosed, so it can be priced.
One policy, training off
A single policy with no tier carve-out and a no-training default — though Cohere leaves the trial-key default direction undisclosed, and Together makes training an explicit opt-in for everyone.
One policy, training on
A single policy whose stated purposes include training its models, with no tier distinction and a stated opt-out right. OpenRouter’s provider table independently marks DeepSeek ‘may train’ with an unknown retention period.
No free tier exists
Trial credits run under the same commercial terms as paid usage, so ‘does the free tier differ’ resolves to ‘there is no separate free tier’. That is the accurate cell — different from a vendor choosing one policy for a real free product.
A quota, not a tier
A time-boxed, region-gated free quota — valid for 90 days, and limited to eligible models and eligible regions — rather than a standing free product. Training and retention terms for it: not established this pass.
The aggregator setting
The routing layer turns the free-versus-paid question into a literal account setting, with separate training-permission toggles for paid and free models. The census’s core question is a documented design decision at the gateway, not a curiosity.
The trend worth interpreting: disclosure quality does not track vendor size. The two clearest free-tier disclosures come from the largest company in this table and a mid-sized European lab; the vaguest cells belong to vendors of every scale. What predicts clarity is whether the vendor decided to productize the free tier at all — vendors that treat free usage as a named product write terms for it, and vendors that treat it as marketing spillover mostly do not.
06 — The Cross-CheckWhen the gateway disagrees with the vendor.
OpenRouter publishes a live per-provider table recording each upstream provider’s retention and train-on-prompts posture on its privacy and logging page (undated). Reading that table against each vendor’s own documentation is the census’s best cross-check: where the two agree, you get independent corroboration; where they disagree, you get a genuine finding. It disagrees twice, and leans stronger than the vendor once.
“On your account settings page, you can set whether you would like to allow routing to providers that may train on your data (according to their own policies). There are separate settings for paid and free models.”— OpenRouter, provider privacy and logging documentation (undated)
| Provider | OpenRouter’s table says | The vendor’s own page says | How to read the pair |
|---|---|---|---|
| MistralTension | “Retained for 30 days” / does not train | A user-configurable chat retention window: Never / 30 / 60 / 90 / 180 days / 1 year | A fixed figure versus a configurable menu. Possibly the gateway describes API routing specifically — but neither source says so. Recorded, not reconciled |
| Together AITension | “Zero retention” by default | ZDR is an active opt-in: “By choosing ‘No’, you are enabling Zero Data Retention” | The gateway’s framing is stronger than the vendor’s own. If ZDR matters to you, take the weaker of the two claims until the vendor states the default |
| Z.aiStronger claim | “Zero retention” / does not train | API training is opt-in; no retention day-count or ZDR claim found on the terms page | Consistent on training, but zero-retention is an additional claim not found verbatim on Z.ai’s own page this pass |
| Anthropic | “Retained for 30 days” / does not train | Covered models retained 30 days for safety work; no training on Customer Content | Corroboration — two independent sources land on the same 30-day figure |
| DeepSeek | “Prompts are retained for unknown period” / may train | Processing purposes include training; no fixed retention day-count published | Corroboration — the gateway’s blunt cell matches the vendor’s own policy text |
| Google AI Studio | “Retained for 55 days” / does not train | Free tier trains; paid tier logs “for a limited period” with no day-count | Not a contradiction: 55 days is OpenRouter’s own stated retention for its AI Studio routing, a different claim from Google’s free/paid split — the two should not be merged into one sentence |
The disagreement between a gateway’s summary of a vendor and the vendor’s own docs is not a nuisance to be edited away — it is arguably the most useful thing in this census. It tells you that the retention number you see in an aggregator UI is a characterization, made in good faith, of documents that are sometimes vaguer or more configurable than a single cell can hold. OpenRouter is also explicit that its free/paid training toggle “has no bearing on OpenRouter’s own policies and what we do with your prompts” — the setting governs routing eligibility to upstream providers, not the gateway’s own logging, which is a separate question we examine in our companion piece on what OpenRouter’s charts actually measure.
07 — Reveal-Day FootnoteDoes a promo price change the terms? Z.ai, one row.
This census shipped on the day the stealth listing ox-alpha was revealed as Z.ai’s GLM-5.3-Flash, which makes row 5 topical: the model is running a promotional pricing window through September 9, 2026, and promo windows are exactly the kind of surface this census exists to check. The answer from the terms page is clean. Z.ai’s Terms of Use (updated April 14, 2026) draw one line — consumer surface versus API surface — and the promo window appears nowhere in it. For API and business users: “We will not use End User Content for developing or improving Services, unless you explicitly agree to such use.” For consumer users, the terms are training-permissive: Z.ai reserves the right to process User Content to improve and develop its services. A promo price is a cost change, not a documented data-terms change.
The stealth window itself is a separate story with its own data-terms angle — the listing’s no-training promise during the anonymous period, which we covered when the listing first appeared. Here, post-reveal, the relevant fact is the now-attributed vendor’s standing API terms — a different question with a documented answer. And what stealth listings in general do and do not disclose is territory we mapped in our identity-blind evaluation guide; this census deliberately does not re-cover it.
08 — Using the CensusWhat to do with each finding.
A census is only useful if it changes behaviour. The honest decision rules that follow from the table:
Treat the free tier as a different product
Where the vendor states a free/paid split (Google, Mistral), never prototype on the free tier with data you would not paste into a public form. Upgrading to paid — or opting out where a toggle exists — is a data-governance action, not just a billing one.
Price the silence
A cell reading not disclosed or not established is not an accusation — it is a work item. Ask the vendor directly, in writing, and pin the answer to a dated document. Silence that survives a direct question is itself an answer.
Set both toggles, trust neither summary
If you route through OpenRouter, set the training-permission toggles for paid and free models explicitly, and remember the setting governs upstream routing, not the gateway’s own logging. Where the gateway’s cell is stronger than the vendor’s own page, assume the weaker claim.
Re-read on a calendar, not on a scandal
Four of the pages cited here are undated. Terms drift. Put the vendor pages your stack depends on into a quarterly review, and re-verify before any new data category touches an API.
Looking forward: expect the free-tier terms gap to narrow, because the aggregator layer is already forcing the issue. When a routing gateway operationalizes the question of whether a provider may train on your prompts as separate account settings for paid and free models, vendors whose pages are silent start to look worse than vendors who disclose an unflattering default — and the commercial pressure runs toward disclosure. Until then, the census form is the right defence: rows, citations, and a refusal to infer. If your team is standing up AI workloads and needs the data-terms diligence done against your own risk profile — including which tiers your prototypes are actually allowed to touch — our AI transformation engagements start with exactly this kind of vendor-terms review.
09 — ConclusionRead the tier, not the brand.
The free tier is a different product. Read it like one.
Eleven vendors, five distinct postures across the ten with a located primary page. Two state a free-versus-paid split outright. Two have no free tier to compare. One trains by default across every surface. One offers a quota instead of a tier. Four publish a fixed retention day-count, and every one of the four is 30 days. The rest is limited-period phrasing, configurable windows, and silence — each labelled here as exactly what it is.
The method matters as much as the cells. Undated pages are called undated; figures that could not be confirmed on a fetched primary read not established rather than shipping on a search snippet. Where the gateway’s summary and the vendor’s documentation disagree, both sides are on the record. That discipline is what makes a census citable — and it is the same discipline to demand from any vendor summary you are handed.
The practical takeaway fits in a sentence: before a prompt with anything real in it touches a free tier, a trial key, or a promo window, read the page this table links for that row — because the cheapest tier is the one nobody reads, and the vendors know it.