An AI procurement checklist for 2026 has to answer a different question than the one most buying guides cover: not “which vendor should we pick?” but “what must be in the contract before we sign with the vendor we picked?” The two stages fail differently. A weak selection wastes an evaluation cycle; a weak contract locks in training-data rights, silent model swaps, and exit terms you will live with for years.
We have already published the selection-stage toolkit — the vendor-selection scorecards, RFP template and evaluation rubric that get you from a long list to a shortlist. This post starts where that one stops: the moment a preferred vendor is chosen and a draft order form lands in your inbox. Everything below is signing-stage — clause by clause, with what a good answer looks like and what should make you pause.
The checklist covers seven clause families: training-data use, retention and zero data retention, model-change and deprecation policy, usage-limit mechanics, exit and export, sub-processor disclosure, and benchmark claims. Wherever possible, the “good answer” is anchored to what OpenAI, Anthropic and Google actually publish — because the strongest negotiating position is knowing what the market already puts in writing.
- 01Selection scorecards and signing checklists differ.A selection rubric compares vendors; a signing checklist governs the one you chose. The seven clause families here decide cost and risk after the scorecard is filed away.
- 02Every major retention promise is tier-qualified.OpenAI, Anthropic and Google all qualify their no-training and retention promises by product tier and carve out exceptions. An unqualified blanket promise is a red flag, not a reassurance.
- 03Deprecation floors vary by an order of magnitude.OpenAI commits to at least 6 months for GA models, 3 for specialized variants, as little as 2 weeks for previews; Anthropic commits to at least 60 days. Get your floor in writing, in months.
- 04Usage limits are becoming rolling windows, not seats.Continuously replenishing windows are now a standard AI-vendor limit shape, materially different from per-seat SaaS quotas. Ask what mechanic applies and what happens at the ceiling.
- 05Exit, sub-processors, benchmarks: demand it in writing.None of the three major vendors publishes a numeric data-export SLA in its public terms at the time of writing. If a commitment matters to you, it belongs in the agreement, not the sales call.
01 — Selection vs SigningTwo checklists — and the expensive one comes second.
Most AI-buying content operates at the selection stage: scorecards, demos, reference calls, an eight-axis rubric. That work matters, and our Stage 4 vendor-selection templates cover it in template form. But selection artifacts have a short shelf life — once the vendor is chosen, the scorecard is history. The contract is what your team lives inside for the next one to three years.
The signing stage has its own failure modes, and they are quieter. Nobody writes a post-mortem titled “we agreed to a retention clause we never read.” The damage shows up eighteen months later as a model retired on two weeks’ notice, a usage ceiling nobody understood, a sub-processor added silently, or a fine-tuned model that cannot leave the vendor’s infrastructure. General SaaS procurement discipline — the kind we covered in our vendor-management and procurement guide — still applies. What follows is the AI-specific layer on top.
One framing rule before the clauses: for every question below, the useful output is not a verbal reassurance. It is either a citation to the vendor’s own published policy — pinned by version and date in the agreement — or a negotiated term in the order form. Anything that lives only in a sales thread does not exist.
02 — Training DataWho trains on your data — and on which tier.
The first question is the one every buyer asks and most vendors answer too loosely: do you train on our data? The honest answer at every major lab is tier-dependent — the same company can have three different answers across its consumer, paid API and enterprise products. That makes “which SKU does this answer apply to?” the real question.
Business + API: no training by default
OpenAI’s enterprise-privacy page (updated January 2026) states “We do not train our models on your data by default” for its business tiers and the API Platform. Data from consumer ChatGPT and other individual services is used for training unless the user opts out — the tier boundary is the whole story.
Commercial content: not trained on
Anthropic’s Commercial Terms of Service state it does not train on customer content from its commercial products, and its data-retention docs add that retained data is never used for training without express permission — two separate commitments worth citing separately in a contract.
Paid vs unpaid decides it
The Gemini API terms split on whether a billing account is attached to the Cloud project. On Paid Services, Google does not use prompts or responses to improve its products or train base Gemini models. On the unpaid tier, content is used to improve Google products, including model training — the cleanest binary data-for-access trade among the majors.
Google’s free-versus-paid split is the canonical example of a broader 2026 pattern: vendors increasingly price the training-data right explicitly, trading cheaper or free access for the right to learn from your prompts. Procurement teams should treat any discounted or free tier as presumptively carrying a data trade until the terms prove otherwise — and confirm which lane the specific order form actually puts you in, because the tier boundary rarely appears on the invoice.
One structural note for open-weight vendors: Meta-style Acceptable Use Policies are license conditions, a different contractual mechanism from a DPA training clause. An AUP restricts what you may do with the model; a data-processing agreement restricts what the vendor may do with your data. They live in different documents, and a procurement review needs to read both — clearing one tells you nothing about the other.
03 — Retention & ZDRRetention windows, ZDR, and the carve-outs.
“Do you train on our data?” and “do you keep our data?” are different questions with different clauses. On the second, the two primary-documented positions look similar at the surface: OpenAI states it may retain API inputs and outputs for up to 30 days to provide the service and identify abuse, after which they are removed unless legal requirements apply. Anthropic’s standard commercial terms auto-delete API inputs and outputs within 30 days, with the same style of exceptions. Both offer Zero Data Retention (ZDR) — and neither offers it as a self-serve toggle. OpenAI describes ZDR as available on request for eligible endpoints with a qualifying use case; Anthropic enables it per organization through sales, and enabling it for one org does not extend it to sibling orgs under the same account.
“Retained data is never used for model training without your express permission.”— Anthropic, API and data-retention documentation
Auto-delete window
OpenAI retains API inputs and outputs up to 30 days for service delivery and abuse monitoring, then removes them; Anthropic auto-deletes within 30 days on standard commercial terms. Both are vendor-stated defaults — cite the page and date in your agreement.
Trust-and-safety ceiling
Anthropic’s docs carve out flagged content: material flagged by trust-and-safety systems may be retained up to two years regardless of the ZDR arrangement. Every vendor has some equivalent — ask for the exception list and its retention ceiling in writing.
Not a toggle — and not universal
Anthropic enables ZDR per organization via sales; new orgs need separate enablement. Its Covered Models — Claude Fable 5 and Claude Mythos 5 at the time of writing — require mandatory 30-day retention and are not available under ZDR at all, opt-in per workspace.
A defensible retention answer specifies (a) which SKU it applies to, (b) whether ZDR requires a separate signed arrangement, (c) which specific models or features are carved out, and (d) the legal-and-safety exception with its ceiling. A vendor who answers “we don’t store your data” with none of those qualifiers is either oversimplifying or hasn’t documented its own pipeline — both are problems you want surfaced before signature, not after an incident.
04 — DeprecationModel change: get the notice floor in months.
AI vendors retire models on schedules that would be unthinkable in traditional enterprise software, and the notice you are entitled to varies enormously by vendor and model class. “Unless safety or compliance concerns require a faster timeline,” OpenAI’s deprecations page reads, “we provide the following minimum notice periods before model retirement” — and then commits to at least 6 months for generally available models, at least 3 months for specialized variants of GA models, and as little as 2 weeks for models with “preview” in the name.
Anthropic’s model-deprecations page commits to at least 60 days’ notice before retirement for publicly released models — a materially shorter floor than OpenAI’s GA tier, and one it uses: on June 5, 2026 Anthropic notified developers that Claude Opus 4.1 would retire on August 5, 2026, a 61-day window. Meanwhile Microsoft’s documentation for Azure OpenAI reportedly guarantees GA model versions for a minimum of 12 months, with existing customers getting a further 6 months on the older version after that — a secondary-sourced figure, but directionally important: the same model family can carry very different lifecycle commitments depending on which surface you buy it through. Anthropic’s docs make the same point from the other side, noting that partner-operated platforms set their own retirement schedules.
Minimum notice before model retirement · stated floors by surface
Sources: OpenAI + Anthropic deprecation pages (vendor-stated); Azure figure via Microsoft Learn coverage, secondary“We don’t recommend using preview models for business-critical production workloads unless you can migrate on short notice.”— OpenAI, model deprecations documentation
The policy is not theoretical — OpenAI’s live deprecations log shows it operating. A June 11, 2026 announcement scheduled the GPT-5-era models (gpt-5, gpt-5-mini, gpt-5-nano, gpt-5-pro, o3, o3-pro) for shutdown on December 11, 2026 — exactly the 6-month GA window. A May 8, 2026 announcement scheduled gpt-5.2-chat-latest and gpt-5.3-chat-latest for August 10, 2026 — roughly the 3-month specialized-variant tier. Two details in the same policy are easy to miss and worth contract language: deprecation is not shutdown (a model becomes “deprecated” when announced, with a separate future shutdown date), and some customers can negotiate dedicated capacity for continued access after shutdown — an exit-adjacent lever to raise at signing, not after the retirement email arrives.
Three asks for the contract. First, a minimum-notice commitment denominated in months, per model class — OpenAI’s three-tier structure (GA / specialized / preview) is a reasonable template to request from any vendor. Second, clarity on what happens to fine-tuned models: OpenAI’s stated pattern is that fine-tuned model inference continues until the base model is deprecated — ask your vendor to confirm in writing how its fine-tuning product behaves. Third, a named migration path: deprecation notices that recommend a replacement model are operationally very different from notices that just end service. Industry observers describe support windows compressing as release cadence accelerates, so assume the clause will be exercised during your term — for the operational side of surviving a retirement, see our model-deprecation calendar and API-sunset survival guide.
05 — Usage LimitsRolling windows are the new seat counts.
Traditional SaaS usage math is simple: seats times price, quota resets on the billing date. AI vendors are quietly replacing that shape with something procurement teams have less practice pricing: the rolling window — a cap that replenishes continuously rather than resetting at a calendar boundary. Usage trackers report that Anthropic’s Claude subscription plans, for example, layer session-based windows with longer rolling caps that run concurrently and replenish continuously rather than resetting weekly on a fixed day. The specific numbers change frequently — which is itself the point: the mechanic, not the current figure, is what belongs in your contract review.
The question set for any AI usage clause: does “included” or “unlimited” usage mean a fixed monthly quota, a rolling window, or a per-seat allocation that multiplies with headcount — and what exactly happens when the team hits the ceiling? Each mechanic prices differently, each fails differently under load, and vendors rarely volunteer which one applies.
Fixed monthly quota
The classic shape: N units per billing period, reset on the invoice date. Predictable to budget, brutal at month-end if the team front-loads usage. Confirm whether unused quota rolls over and how mid-term seat additions prorate.
Rolling window
A cap measured over a continuously sliding window — it never “resets,” it replenishes. Now a standard AI-vendor shape for subscription and discounted tiers. Harder to budget, impossible to game with month-end timing, and the window length can change under you if it lives in a docs page rather than the contract.
Per-seat allocation
Usage entitlements attached to named users, multiplying with headcount. The familiar SaaS math — but AI seats vary wildly in consumption, so blended per-seat pricing quietly subsidizes heavy users. Confirm seat-sharing rules and true-up terms before the renewal conversation does it for you.
Ceiling behavior
Whatever the mechanic, the ceiling has one of three behaviors: hard stop, throttle, or automatic overage billing. Each is a different operational risk — a hard stop mid-launch, silent degradation, or a surprise invoice. The behavior should be named in the order form, not discovered in production.
Rate limits and seat economics also interact with pricing surfaces: the same vendor can expose standard API rates, discounted batch processing, and subscription seats with entirely different limit mechanics on each. We mapped the seat-versus-API math in detail in our AI coding-tool seat-economics guide — the procurement takeaway is that “what does it cost?” is unanswerable until “which surface, under which limit mechanic?” is settled in writing.
06 — Exit & ExportExit rights: if it isn’t written, it doesn’t exist.
Standard SaaS exit hygiene applies to AI contracts — a defined data-export format, a defined export window after termination, and a deletion-on-termination commitment with a stated timeline. What is striking about the AI majors specifically: at the time of writing, we could not locate a numeric data-export SLA (“exports delivered within N days of a termination request”) in the public enterprise terms of OpenAI, Anthropic or Google. Such commitments may exist inside negotiated enterprise agreements — which is precisely the point. If an export window matters to you, it will only exist because you asked for it in writing.
AI contracts also add an exit asset that traditional SaaS never had: fine-tuned model artifacts. In our analysis, this is where exit clauses most often fall short — the training data you submitted is usually exportable, but the fine-tuned model built from it is generally not portable to another vendor’s infrastructure, because the underlying base model is proprietary. OpenAI’s enterprise-privacy page states that fine-tuned models are yours alone to use and that data submitted for fine-tuning is retained until the customer deletes the files — ownership and deletion answered, portability not. An exit clause that inventories what actually leaves with you — prompt and conversation history, uploaded files, evaluation data, fine-tuning datasets — and what does not, is the difference between a migration plan and a renegotiation under duress.
Exit terms are also where the build-versus-buy calculus resurfaces: the harder it is to leave a vendor, the stronger the case for owning more of the stack yourself. If the exit conversation with a vendor goes badly, that is useful signal for the architecture decision — our build-vs-buy decision framework treats vendor lock-in as a first-class input rather than a footnote.
07 — Sub-ProcessorsWho else touches your data.
Your AI vendor is not one company — it is a chain of them. Every major provider relies on sub-processors for infrastructure, analytics and support tooling, and the mature disclosure pattern is a public, dated, versioned list. OpenAI publishes one at openai.com/policies/sub-processor-list; a third-party monitoring service that tracks such pages daily reported it held 24 entries with its most recent change on July 10, 2026 — treat the count as the aggregator’s figure rather than OpenAI’s own stated number, but the existence of the public, versioned list is OpenAI’s own page. The list spans cloud and inference infrastructure, data tooling and support systems, plus OpenAI’s regional affiliates.
The procurement question is not “do you have sub-processors?” — everyone does. It is: where is the list, how is it versioned, and how do we hear about changes? A published list with an opt-in change-notification mechanism (OpenAI offers a subscription form for new sub-processor announcements) means you learn about a new data-handling party before it processes your data. A static PDF supplied on request, with no update mechanism, means you learn about it in an audit — or in an incident report. Your DPA should reference the list by URL and date, require advance notice of additions, and preserve a right to object.
08 — Benchmark ClaimsBenchmark claims that survive a fact-check.
AI contracts are increasingly sold on benchmark numbers, and 2026 buyers have learned to distrust them. The proliferation of benchmarks — coding, agentic, reasoning, task-specific — has given vendors more configurations to selectively report from, not fewer. Sophisticated buyers now discount any headline percentage that arrives without three qualifiers: the exact benchmark version and date, whether the comparison models were run by the vendor or independently, and whether the tested configuration matches production use — tool access, context length and retry policy included.
This week supplied its own live illustration. Qwen3.8-Max shipped with benchmark figures circulating widely in press coverage while at least one report noted that no official vendor-published benchmark table accompanied the release at the time of writing — meaning the numbers in circulation could not be pinned to a primary source. We covered the release itself in our Qwen3.8-Max analysis; the procurement lesson stands on its own: before a benchmark figure enters your decision memo, ask which of the vendor’s claims live on the vendor’s own published pages versus only in press coverage.
The good answer, then: the vendor can point to the result on its own published page, states the benchmark version and date, and — ideally — documents methodology well enough to reproduce: dataset, prompt template, retry policy. The red-flag answer is third-party press coverage with no primary table behind it. In contract terms, benchmark claims that materially influenced the purchase belong in the agreement as representations, not in a slide deck — vendors confident in their numbers rarely object, and the objection itself is information.
09 — The Question SetThe printable question set.
Everything above, compressed into the seven questions to put in front of a vendor before signature — with what a good answer looks like and the pattern that should stop the process. Each “good answer” is modeled on something at least one major vendor already publishes, which is your negotiating leverage: none of this is an exotic ask.
| Clause family | The question | A good answer looks like | Red flag |
|---|---|---|---|
| Training data | Do you train on our prompts, outputs or files — on the exact SKU we are buying? | A tier-named, written commitment citing the specific product — the pattern OpenAI, Anthropic and Google’s paid tier all publish today. | An unqualified “we never use your data” with no SKU named and no document cited. |
| Retention & ZDR | What is retained, for how long — and which models or features are carved out? | A stated default window (30 days at OpenAI and Anthropic), a written ZDR mechanism, and named exceptions with ceilings. | “ZDR” promised verbally, with no carve-out list and no answer for models shipped after signature. |
| Deprecation | What minimum notice do we get, per model class, before a model we depend on retires? | Floors in months by class (OpenAI: 6 GA / 3 specialized; Anthropic: 60 days), plus the fine-tune inheritance rule. | “We’ll give reasonable notice” — undated, unclassed, unwritten. |
| Usage limits | Is “included” usage a monthly quota, a rolling window, or per-seat — and what happens at the ceiling? | The mechanic named in the order form, with ceiling behavior specified: hard stop, throttle, or overage at stated rates. | “Unlimited” plus a fair-use clause that names no mechanic and reserves the right to change limits. |
| Exit & export | What exports on termination, in what format, within what window — and what happens to fine-tuned artifacts? | Enumerated data types, a named format, a window in days, a deletion timeline, and explicit fine-tune ownership terms. | Export “on request” with no SLA; silence on fine-tuned models and uploaded datasets. |
| Sub-processors | Where is your sub-processor list, and how do we hear about changes? | A public, dated, versioned list referenced in the DPA, with an opt-in change-notification mechanism and a right to object. | A static PDF on request, no update mechanism, no advance notice of additions. |
| Benchmark claims | Which of your performance claims are on your own published pages, with version and date? | Vendor-published results with benchmark version, date and methodology — willing to stand as representations in the agreement. | Only third-party press citations, no primary table, and resistance to putting numbers in the contract. |
If your organization runs a formal third-party risk program, the seven families map cleanly onto the NIST AI Risk Management Framework’s four functions — Govern, Map, Measure, Manage — released in January 2023 and now widely described as a common backbone for AI governance despite being voluntary. Training data, retention and sub-processors sit under Govern and Map; usage limits and benchmark verification under Measure; deprecation and exit under Manage. That mapping gives the checklist a recognized home in an existing TPRM process rather than a bespoke sidecar — useful, because reviewers increasingly expect vendors to demonstrate RMF-aligned practice, and because vendor risk does not end at signature: as we argued in our vendor-risk reassessment after the Pacing letter, the ground under an AI contract can shift mid-term.
Looking forward, we expect the leverage to keep moving toward buyers who ask early. Vendor-published policy pages have become more specific every quarter — notice floors, retention windows and sub-processor mechanisms that were unwritten a year before are now citable documents — and buyers who pin those pages into contracts by version and date are effectively getting enterprise terms without enterprise negotiation. Teams that want a second set of eyes on an AI contract, or an evaluation pipeline that feeds it, can engage our AI transformation practice — contract-stage diligence is where most of the avoidable cost hides.
10 — ConclusionSign what is written, not what is said.
The contract is the product: buy what is written down.
The selection stage picks a vendor; the signing stage decides what you actually bought. The seven clause families here — training data, retention, deprecation, usage limits, exit, sub-processors, benchmarks — are where AI contracts differ most from the SaaS agreements your procurement process was built for, and every one of them has a documented good answer somewhere in the market already.
That is the practical leverage: you are almost never asking a vendor to invent a commitment, only to put in writing what the market leaders already publish. Tier-scoped no-training defaults, 30-day retention windows, notice floors measured in months, versioned sub-processor lists — these exist as citable pages today. A vendor unwilling to match in contract what its competitors publish on the open web has told you something valuable before you signed.
And the clauses will be exercised. Models retire mid-term, limits change mechanics, sub-processor lists grow, and the model your ZDR arrangement covered is not the model your team will be using at renewal. The checklist is not paranoia — it is the assumption, borne out by the vendors’ own logs, that the contract you sign will meet every one of these clauses during its life. Sign accordingly.