On September 25, 2026, OpenAI fixed a bug that had degraded how GPT-6 Sol and Luna read images, and recommended that customers re-run affected evaluations and retry affected workflows. For an agency that had already delivered image-based work to a client, the re-run is the easy part. In the worked example below, the model costs about $22 to run again. The people who check the results cost between $580 and $6,580.
That gap is the whole problem. Most agency contracts say nothing about a model provider’s defect, so the cost falls wherever the pricing model happens to put it. This post prices one illustrative job, shows who pays under three common pricing models, and suggests a clause that settles the question before the next bug.
- 01Re-running the model is cheap.About $22 for 4,000 items on Batch at OpenAI’s published GPT-6 Sol prices, in our illustrative job.
- 02Re-checking the output is not.$400 for a 5% sample, and $6,000 more for a full review if the sample finds errors.
- 03Your pricing model decides who pays.Fixed fees and per-item prices leave it with the agency; hourly billing sends it to a client who did nothing wrong.
- 04A vendor-defect clause settles it in advance.Agree what a re-run costs, who reviews, and for how long results are covered before the next bug arrives.
01 — The scenarioThe job: 4,000 invoices read by a model with a fault
The job is invented to make the arithmetic concrete; every volume, time and rate below is illustrative, and only the model prices are real. An agency extracts supplier, date, amount and tax fields from 4,000 scanned invoices for a client, using GPT-6 Sol in the days after its September 22 launch. It delivers the data to the client’s finance system. On September 25, OpenAI’s changelog reports the image fix and advises re-running affected work.
Nobody yet knows whether the delivered data is wrong. OpenAI did not say how badly image reading was affected, as we noted in our post on the fix. The responsible course is to re-run, compare and check. The question is what that costs and whose budget it comes from.
02 — The billWhat redoing the work costs
OpenAI’s GPT-6 Sol model page lists $2 per million input tokens and $10 per million output tokens, with Batch at half those rates. We assume 3,000 input tokens per invoice for the image and instructions and 500 output tokens for the extracted fields. That is $0.011 an invoice at standard rates, or $44 for the job, and $22 on Batch, which suits a re-run with no deadline.
| Cost line | Cost | When | Basis |
|---|---|---|---|
| Model re-run on Batch | $22 | Always | 4,000 items × (3,000 in + 500 out tokens) at $1 / $5 per million |
| Sample review | $400 | Always | 200 items (5%) × 2 minutes at $60 an hour |
| Client handling | $180 | Always | 3 hours of account time at $60 an hour |
| Full review | $6,000 | Only if the sample finds errors | 4,000 items × 1.5 minutes at $60 an hour |
If the sample is clean, the total is $602, of which the model is about 4%. If the sample finds errors and every invoice is reviewed, the total is $6,602, and the model is about a third of 1%. The same pattern appears whenever AI output needs checking, which we worked through in when checking AI output costs more than generating it. A vendor bug simply makes you pay the checking cost twice.
03 — Who paysWho absorbs the cost under each pricing model
Fixed fee per project
The job is already paid for. Every re-run and review hour comes out of the agency’s margin unless the contract says otherwise.
Per item processed
The client paid per invoice for a correct result. Reprocessing the same invoices earns nothing new.
Time and materials
Review hours are billable, but the client is paying for a defect neither party caused, and many will refuse.
None of the three is fair by default. Under a fixed fee, a clean sample costs the agency a tolerable $602; a full review wipes out $6,602 of margin on a job that may have been priced at a few times that. Under hourly billing, the agency is covered on paper, but asking a client to pay for a model provider’s fault is a hard conversation. The changelog entry does not mention credits or refunds, so neither party should count on recovering the cost from the vendor.
The pricing models themselves are covered in our AI agency pricing guide. What none of them includes by default is a rule for this case.
04 — The clauseA vendor-defect clause, agreed before it is needed
(1) When a model provider publicly acknowledges a defect affecting delivered work, the agency re-runs it at no charge for model usage. (2) A sample review, sized in advance, is included; a full review is split or billed at an agreed discounted rate. (3) The cover applies to work delivered within a stated window, such as 30 days before the vendor’s notice. (4) The agency records which model and date produced each deliverable, so affected work can be identified.
The fourth point is the one that makes the others workable. GPT-6 Sol has no dated snapshot to pin, so the model name alone does not tell you which version produced a file. A dated log of runs does. Without it, an agency cannot say which deliverables fall inside the window, and the default becomes reviewing everything.
The clause also helps with pricing. An agency that knows its exposure can price a small reserve into fixed-fee AI work, rather than discovering the risk the week a bug is announced. Our statement-of-work framework shows where terms like these sit in a contract.
05 — ConclusionThe model is cheap to rerun; the review is what someone must fund
Add a vendor-defect clause to your next AI statement of work and start logging model and run date for every deliverable
Model bugs will keep happening, and vendors will keep advising customers to re-run their work. The tokens will be cheap each time; the review will not. Deciding in advance who pays for that review turns an awkward client call into a line in a contract. If you want help building AI delivery processes with this kind of traceability, our AI transformation team sets them up with clients.