AI DevelopmentPlaybook5 min readPublished September 27, 2026

GPT-6 Sol · GPT-6 Luna · API and Codex · computer use · three days of image results to recheck

GPT-6 Sol and Luna Image Bug Fixed: Re-run Your Image Evals

OpenAI fixed an image-encoding bug in GPT-6 Sol and Luna on September 25. Which results to re-run, which decisions to revisit, and what OpenAI did not say.

DA
Digital Applied Team
Research and practical guidance
Fix datedSeptember 25, 2026
Models launchedSeptember 22, 2026

On September 25, 2026, OpenAI’s API changelog recorded a fix for a bug in image encoding that had degraded how GPT-6 Sol and GPT-6 Luna understood images. Both models had launched three days earlier. OpenAI says the fix improves results on visual tasks in the API and in Codex, computer use included, and it recommends that anyone with image inputs re-run their evaluations and retry the workflows the bug touched.

That advice is short, and it matters most to the teams that moved fastest. If you tested GPT-6 on screenshots, scanned documents, charts or a computer-use agent in its first days, some of what you concluded may have come from a broken input path rather than from the model.

Key takeaways
  1. 01
    OpenAI fixed an image-encoding bug in GPT-6 Sol and Luna on September 25.The changelog entry covers image understanding in the API and in Codex, including computer use.
  2. 02
    Image results from the launch window are the ones to distrust.Anything scored, compared or shipped on GPT-6 image inputs between the September 22 launch and the fix.
  3. 03
    Re-running is cheap; re-deciding is the real cost.A model rejected for a vision task in its first three days deserves a second test before the decision sticks.
  4. 04
    OpenAI has not said how large the degradation was.The entry gives no scores, no start date for the bug and no word on whether launch benchmarks were affected.

01 — The fixWhat OpenAI fixed

Models affectedNamed in the changelog entry.
GPT-6 Sol, GPT-6 Luna
Where the fix appliesVisual tasks, including computer use.
API and Codex
Changelog dateLaunch entry for both models: September 22.
September 25, 2026
OpenAI’s recommendationFor use cases with image inputs.
Re-run evals, retry workflows

The OpenAI API changelog files the entry as a fix under both model names. It places the fault in image encoding, the step that turns a picture into input the model can read, and OpenAI shipped the correction without a new model identifier. Both models take text and images as input and return text; neither generates images, so the bug affected reading images, not making them.

The GPT-6 Sol model page lists a single identifier, gpt-6-sol, with no dated snapshot to pin, and the Luna page is the same. So the fix reached every caller without a code change. That is convenient now, and it also means there is no old snapshot to reproduce the broken behaviour if you want to measure how much it cost you.

02 — The exposureWhich results to distrust

OpenAI has not said when the bug began. The models went live in the API on September 22 and the fix is dated September 25, so the work to check is any image task run on either model before the fix. Four kinds of output deserve a second look:

1
Evaluation scores
Scored before Sep 25

Any internal benchmark with screenshots, charts, forms or photos. A low score may describe the bug, not the model.

Re-run first
2
Model comparisons
GPT-6 vs a rival

A head-to-head that put GPT-6 behind another model on vision may have been decided by the encoding fault.

Revisit the decision
3
Agent runs
Computer use, Codex

Agents that read the screen to click or type. Failed or looping runs from launch week are worth retrying.

Retry failures
4
Shipped extractions
Data already delivered

Fields pulled from invoices, receipts or scans and passed on to a client or a database without human review.

Spot-check

The second box is the one teams forget. A comparison run in launch week becomes a routing rule, and a routing rule is rarely retested. If your router sends vision work away from GPT-6 because of a test from that window, the rule may rest on a fault that no longer exists. Our guide to telling a real regression from a broken evaluation covers the same trap from the other direction.

03 — The re-runRe-running it without re-running everything

The token cost of a re-run is small at these prices. OpenAI lists GPT-6 Sol at $2 per million input tokens and $10 per million output tokens, and Luna at $0.10 and $0.50, with Batch processing at half the standard rate. The expensive part is the human review that follows a changed result, so narrow the re-run before you start:

  • Filter your logs for requests to gpt-6-sol or gpt-6-luna that carried an image, dated before the fix.
  • Re-run the evaluation set, not production traffic. A fixed test set gives a before-and-after comparison; production requests do not repeat.
  • Use Batch for the bulk. An evaluation re-run has no deadline, which is what the half-price tier is for.
  • Record which runs predate the fix so a future comparison does not mix the two populations.

If you do not have a fixed image test set, this is a reason to build one. Our walkthrough on building an evaluation harness to qualify new models shows the minimum version.

04 — The gapsWhat the changelog leaves out

Unanswered, as of the fix

The entry gives no measure of how much image understanding was degraded, no date when the fault started, and no statement on whether any published launch scores were produced with it. Our launch coverage reported OpenAI’s computer-use results, which depend on reading screenshots. Treat any vision-heavy launch figure as unconfirmed until OpenAI says whether it was affected.

None of this is unusual for a changelog, which exists to record changes rather than explain them. It does mean the only reliable measure of the fix is the one you run on your own images. When OpenAI recommends “rerunning your evaluations,” it is telling you, in effect, that it cannot give you the answer for your workload.

05 — ConclusionLaunch-week image results are provisional

You tested GPT-6 on images before September 25
Re-run the same test set on Batch and compare. Keep both result sets, labelled by date.
Re-run
You routed vision work away from GPT-6 after a launch-week test
Retest before the rule becomes permanent. The rejection may describe the bug.
Revisit
You delivered image extractions made before the fix
Sample and review a slice by hand; widen the review only if the sample finds errors.
Spot-check
You send GPT-6 text only
Nothing to do. The fix concerns image inputs.
No action
What to do this week

Pull the image requests made before September 25, re-run your fixed test set on Batch, and retest any routing rule set in launch week

A fix like this is good news with a short to-do list. The model you evaluated in launch week may not be the model you have now, and decisions made in that window should be checked once before they harden. If you want help setting up evaluation sets that survive vendor fixes like this one, our AI transformation team builds them as part of model rollouts. For how GPT-6 compares on cost with its closest rival, see GPT-6 Sol against Claude Opus 5.5.

Digital Applied

Know what your models actually do on your data.

We build evaluation sets and routing rules that are retested when a vendor changes something, so a launch-week result never becomes a permanent decision by accident.

Evaluation setsModel routingVendor change reviews
Your next project

Model decisions you can defend

  • →Fixed test sets per workload
  • →Dated before-and-after results
  • →Routing rules with review dates
Questions and answers

The questions we get about the GPT-6 image fix

A bug in image encoding that degraded image understanding. OpenAI's API changelog dates the fix to September 25, 2026 and says it improves visual tasks in the API and in Codex, including computer use.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

GPT-6 Sol and Luna: API Prices, Benchmarks and Trade-offs

GPT-6 Sol costs $2/$10 and Luna $0.10/$0.50 per million tokens, half GPT-5.6's price. What OpenAI's own charts show about scores and effort.

September 22, 2026 · 8 minRead
AI Development

Testing a Vision Model on Screenshots of Your Own App

Evaluate a vision model on real interface screenshots, including missing text and ambiguous controls. Score reading, target location and unsupported claims.

September 9, 2026 · 6 minRead
AI Development

OpenAI DevDay Is September 29: What to Have Ready Now

OpenAI's DevDay is September 29 in San Francisco. What already shipped on the API, what retires on September 28, and the checklist to finish before the event.

September 25, 2026 · 6 minRead
AI Development

GPT-6 Sol vs Claude Opus 5.5: Cost per Task and Benchmarks

Artificial Analysis ran GPT-6 Sol and Claude Opus 5.5 on the same tests. Sol is cheaper up to a point; Opus 5.5 at its default outscores Sol at max.

September 22, 2026 · 6 minRead
AI Development

A Cheaper AI Model Can Leave You With More Review Work

Compare AI models using the review work needed for an accepted result. Track inspection, corrections and rechecks before treating a lower bill as savings.

September 6, 2026 · 4 minRead
AI Development

Give AI Reviewers Different Checks Before Trusting Them

Two AI reviewers can repeat one mistake. Design reviews around separate evidence checks, clear rubrics and independent calculations instead of votes.

September 5, 2026 · 4 minRead