A successful AI demo is a reason to investigate, not enough evidence to buy. Ask the supplier to run tasks you choose, show any human intervention and deliver the result in the form your team would actually use. Agree on acceptance before the trial starts, while neither side knows which examples will look best.
The buying question is specific: can this agent complete the work you intend to delegate under conditions you can support? A polished presentation can answer a different question—whether a carefully prepared example is possible.
- 01Choose the tasks together.Include routine work and meaningful exceptions from the intended scope.
- 02Record intervention.A result rescued by a specialist is different from unattended completion.
- 03Inspect the outcome.Buy against usable results and explicit limits, not a narrated success screen.
01 — Ask for an acceptance packetAsk for an acceptance packet
Use the following packet to make a pilot reviewable. These are proposed evidence requirements; they are not universal purchasing standards or measured predictors of commercial success.
| Evidence | What the buyer can check | Question it answers |
|---|---|---|
| Task brief and source inputs | Input versions and agreed completion rule | Did the trial address our actual work? |
| Trial record | Every attempted case, including failures | Are we seeing the run history or selected highlights? |
| Intervention log | Who corrected, restarted or completed the work | How much support did the result require? |
| Final artifact or state | Open the output or inspect the destination | Did the requested result exist and remain usable? |
| Exception behavior | Missing input, unavailable tool or ambiguous target | Does the agent stop or recover appropriately? |
| Operating assumptions | Access, tools, review work and expected workload | Can we reproduce the conditions after purchase? |
02 — Use repeated trials without inventing a magic pass rateUse repeated trials without inventing a magic pass rate
Anthropic’s evaluation guide distinguishes a task from a trial and recommends multiple trials because outputs vary. It also separates a narrated success from the final state of the environment. Those distinctions make a useful foundation for a buyer-controlled pilot.
Choose cases that represent the work you expect to assign. Keep a record of how they were selected. A pilot made entirely of easy, clean inputs cannot answer how the system handles incomplete attachments or conflicting instructions.
Decide which failures block adoption and which are acceptable with review. A wrong recipient may be disqualifying even when most routine cases pass. Report the counts and conditions rather than turning a small, selected sample into a general reliability percentage.
03 — Include the work hidden around the demoInclude the work hidden around the demo
Record setup, retries, manual corrections and final review separately. If an engineer quietly repairs an export, the resulting file may be excellent, but the operating model includes that engineer’s work. The buyer should know whether the proposed service includes such support.
Ask what happens when the agent encounters a new exception. Is there a named escalation owner? Is the partial result preserved? Can the buyer understand the state without replaying a long chat? The tool-error decision table can help define these cases.
Keep cost analysis grounded in the pilot’s observed inputs. Do not call the saved time or return on investment measured unless you collected a comparable baseline and accounted for review. This article supplies no invented savings estimate.
04 — Review the result through the buyer’s workflowReview the result through the buyer’s workflow
Anthropic’s harness report describes cases where code changes and limited checks did not establish working end-to-end behavior. It is a vendor engineering report, not evidence about the supplier you are evaluating.
For a content agent, open the exported artifact. For a coding agent, run the agreed user action. For a research agent, inspect consequential citations. Choose the acceptance check for the deliverable rather than asking for the same dashboard screenshot in every pilot.
The file acceptance reference makes this concrete for documents. A downloadable deck that cannot be edited may be a successful export and an unsuccessful delivery.
05 — Make a bounded purchasing decisionMake a bounded purchasing decision
Write the decision in one of three forms: ready for the agreed scope, continue a limited pilot to resolve named gaps, or not suitable for this task. List the evidence supporting that decision and the conditions that would invalidate it, such as removing required human review.
A narrow successful pilot can justify a narrow purchase. It does not establish fitness for every department or a higher-risk use. Expand only when the added work has its own acceptance evidence.
For selecting trial cases, the guide to testing on your own work provides a related framework. Keep contractual and procurement review separate from the empirical claim that the agent can complete the job.
06 — DecisionWhat to do next
Buy the supported scope, with its operating conditions.
Request a small, reviewable pilot and make the acceptance rule explicit. The best result is a buying decision whose evidence remains understandable after the demonstration ends.
For implementation support, explore our AI transformation services.