On September 25, 2026, OpenAI’s API changelog recorded a fix for a bug in image encoding that had degraded how GPT-6 Sol and GPT-6 Luna understood images. Both models had launched three days earlier. OpenAI says the fix improves results on visual tasks in the API and in Codex, computer use included, and it recommends that anyone with image inputs re-run their evaluations and retry the workflows the bug touched.
That advice is short, and it matters most to the teams that moved fastest. If you tested GPT-6 on screenshots, scanned documents, charts or a computer-use agent in its first days, some of what you concluded may have come from a broken input path rather than from the model.
- 01OpenAI fixed an image-encoding bug in GPT-6 Sol and Luna on September 25.The changelog entry covers image understanding in the API and in Codex, including computer use.
- 02Image results from the launch window are the ones to distrust.Anything scored, compared or shipped on GPT-6 image inputs between the September 22 launch and the fix.
- 03Re-running is cheap; re-deciding is the real cost.A model rejected for a vision task in its first three days deserves a second test before the decision sticks.
- 04OpenAI has not said how large the degradation was.The entry gives no scores, no start date for the bug and no word on whether launch benchmarks were affected.
01 — The fixWhat OpenAI fixed
- Models affectedNamed in the changelog entry.
- GPT-6 Sol, GPT-6 Luna
- Where the fix appliesVisual tasks, including computer use.
- API and Codex
- Changelog dateLaunch entry for both models: September 22.
- September 25, 2026
- OpenAI’s recommendationFor use cases with image inputs.
- Re-run evals, retry workflows
The OpenAI API changelog files the entry as a fix under both model names. It places the fault in image encoding, the step that turns a picture into input the model can read, and OpenAI shipped the correction without a new model identifier. Both models take text and images as input and return text; neither generates images, so the bug affected reading images, not making them.
The GPT-6 Sol model page lists a single identifier, gpt-6-sol, with no dated snapshot to pin, and the Luna page is the same. So the fix reached every caller without a code change. That is convenient now, and it also means there is no old snapshot to reproduce the broken behaviour if you want to measure how much it cost you.
02 — The exposureWhich results to distrust
OpenAI has not said when the bug began. The models went live in the API on September 22 and the fix is dated September 25, so the work to check is any image task run on either model before the fix. Four kinds of output deserve a second look:
Evaluation scores
Any internal benchmark with screenshots, charts, forms or photos. A low score may describe the bug, not the model.
Model comparisons
A head-to-head that put GPT-6 behind another model on vision may have been decided by the encoding fault.
Agent runs
Agents that read the screen to click or type. Failed or looping runs from launch week are worth retrying.
Shipped extractions
Fields pulled from invoices, receipts or scans and passed on to a client or a database without human review.
The second box is the one teams forget. A comparison run in launch week becomes a routing rule, and a routing rule is rarely retested. If your router sends vision work away from GPT-6 because of a test from that window, the rule may rest on a fault that no longer exists. Our guide to telling a real regression from a broken evaluation covers the same trap from the other direction.
03 — The re-runRe-running it without re-running everything
The token cost of a re-run is small at these prices. OpenAI lists GPT-6 Sol at $2 per million input tokens and $10 per million output tokens, and Luna at $0.10 and $0.50, with Batch processing at half the standard rate. The expensive part is the human review that follows a changed result, so narrow the re-run before you start:
- Filter your logs for requests to
gpt-6-solorgpt-6-lunathat carried an image, dated before the fix. - Re-run the evaluation set, not production traffic. A fixed test set gives a before-and-after comparison; production requests do not repeat.
- Use Batch for the bulk. An evaluation re-run has no deadline, which is what the half-price tier is for.
- Record which runs predate the fix so a future comparison does not mix the two populations.
If you do not have a fixed image test set, this is a reason to build one. Our walkthrough on building an evaluation harness to qualify new models shows the minimum version.
04 — The gapsWhat the changelog leaves out
The entry gives no measure of how much image understanding was degraded, no date when the fault started, and no statement on whether any published launch scores were produced with it. Our launch coverage reported OpenAI’s computer-use results, which depend on reading screenshots. Treat any vision-heavy launch figure as unconfirmed until OpenAI says whether it was affected.
None of this is unusual for a changelog, which exists to record changes rather than explain them. It does mean the only reliable measure of the fix is the one you run on your own images. When OpenAI recommends “rerunning your evaluations,” it is telling you, in effect, that it cannot give you the answer for your workload.
05 — ConclusionLaunch-week image results are provisional
Pull the image requests made before September 25, re-run your fixed test set on Batch, and retest any routing rule set in launch week
A fix like this is good news with a short to-do list. The model you evaluated in launch week may not be the model you have now, and decisions made in that window should be checked once before they harden. If you want help setting up evaluation sets that survive vendor fixes like this one, our AI transformation team builds them as part of model rollouts. For how GPT-6 compares on cost with its closest rival, see GPT-6 Sol against Claude Opus 5.5.