A cheaper model can cost more per accepted task if people spend the difference inspecting, correcting and rechecking its work. It can also be the better choice. You need a record of the review work to tell which is happening; the inference bill alone cannot settle the decision.
Keep the acceptance standard fixed and compare the complete route to a usable result. The ledger below is a proposed measurement method, not evidence that inexpensive models are generally worse or that a particular model saves money.
- 01Count accepted work.Keep the spending on rejected attempts in the comparison.
- 02Separate effort from waiting.Active reviewer minutes and time spent in a queue answer different business questions.
- 03Avoid double-counting.A correction followed by a recheck should have distinct recorded intervals.
01 — Record where the human time goesRecord where the human time goes
Use one record per assigned task, with any model reruns attached to that same record. The categories below are our operational definitions. Apply them consistently across candidates rather than adjusting the categories to fit a preferred result.
| Review activity | Include | Keep separate |
|---|---|---|
| Initial inspection | Reading or running checks to judge the first result | Time the draft waits unopened. |
| Evidence retrieval | Opening sources or artifacts needed for review | Automatic source fetch time without active attention. |
| Correction | Human editing or explanation needed to fix the output | A model-only rerun already counted in inference spend. |
| Recheck | Inspecting the repaired result against the same standard | The earlier correction interval. |
| Escalation | Specialist review and the handover needed to resolve an issue | Unrelated meetings or general team overhead. |
| Rejection | Review effort on work that is abandoned or redone | Do not silently drop it from the task total. |
02 — Use research to choose the measurement, not the winnerUse research to choose the measurement, not the winner
Anthropic’s agent-evaluation guide distinguishes code-based, model-based and human grading, with different costs and limitations. That supports naming the kind of review performed rather than treating every check as an interchangeable quality score.
METR’s July 2025 developer study measured work in a specific setting using early-2025 tools. The authors explicitly limit generalization beyond those developers, repositories and tools. We use it as a reason to measure actual work, not as a current estimate of model productivity.
Neither source supports a rule that a cheaper model creates more review. Prompt design, source access, task difficulty and the required output can change the result. The title describes a possibility to investigate, not a finding about a model tier.
03 — Keep the comparison fair enough to interpretKeep the comparison fair enough to interpret
Choose representative tasks and write the acceptance rule before the outputs arrive. Keep source access and output requirements comparable. Record whether the reviewer knows which model produced the result; expectations can influence how much scrutiny they apply.
Separate active time from elapsed time. A reviewer might spend a short interval checking a draft that waited until the next morning. The labor cost follows active work; the delivery delay follows elapsed time. Both matter, but adding them together would count waiting as paid effort without justification.
Keep difficult tasks and rejected outputs. If one model’s weak drafts are removed before computing the average, the comparison rewards the model for work that did not make it through review. Report task mix and acceptance counts alongside any average.
04 — Calculate the break-even point from your own inputsCalculate the break-even point from your own inputs
For a defined task set, add inference spending, tool spending and active review cost, then divide by the number of accepted deliverables. Active review cost is review minutes multiplied by the chosen hourly labor rate and divided by sixty. State whether that rate includes overhead. If nothing is accepted, cost per accepted result is undefined, not zero.
A model’s inference saving can be compared with its extra review cost only when the deliverable count and quality standard are comparable. The break-even additional review minutes equal the inference saving multiplied by sixty and divided by the hourly review rate. This is arithmetic, not a forecast; insert your measured inputs and disclose them.
Show a range when labor rates or review timing are uncertain. Do not turn a small trial into a precise annual saving. The API price index supplies a different input; it cannot tell you how much checking your team needs.
05 — Route work according to the review bottleneckRoute work according to the review bottleneck
If a lower-cost model produces usable results with similar review effort, there is no reason to penalize it for being inexpensive. If it creates frequent specialist escalations, the scarce resource may be specialist attention rather than compute spending.
Use separate results for distinct task types. A model can be suitable for formatting verified material and unsuitable for source-heavy synthesis. Change the route for the affected work instead of declaring a universal winner.
Use the model-switch testing guide for the broader evaluation setup and the reviewer-independence guide when adding a second automated reviewer. More review steps should earn their cost by answering a defined question.
06 — DecisionWhat to do next
Budget for accepted results and the people who verify them.
Measure inspection, correction and rechecking before calling a lower model bill a saving. Keep the quality bar fixed and route tasks according to the total work they require.
For implementation support, explore our AI transformation services.