GPT-5.6 Sol Ultrafast is OpenAI’s newest service tier, previewed on August 13, 2026: the company’s most capable GPT-5.6 model running on Cerebras hardware at what OpenAI describes as “up to 14× faster than Standard processing,” peaking at 750 output tokens per second. It launches first in the API, behind a waitlist, for a select group of customers.
The announcement matters because it attacks a trade-off every team building real-time AI products knows well: until now, getting conversational-speed responses usually meant stepping down to a smaller or more specialized model. Ultrafast claims to remove that compromise — frontier-model output at speeds that keep pace with a live conversation, an incident bridge, or a checkout flow.
This guide covers what OpenAI actually announced, why Cerebras hardware is the enabling story, where every headline speed claim comes from (and who measured it), what 750 tokens per second changes for real-time agentic products, and what a limited preview does — and does not — let you act on today. Every number below is traced to the vendor post that stated it.
- 01Ultrafast is a speed tier, not a new model.It runs the existing GPT-5.6 Sol — previewed in June 2026, broadly available since July — on Cerebras hardware. OpenAI and Cerebras both position it as a latency-only change with the same claimed intelligence.
- 02The headline numbers are vendor ceilings.Up to 14x standard speed and up to 750 output tokens per second are OpenAI’s own figures, with no disclosed baseline workload, prompt set, or reasoning-effort setting behind the multiple.
- 03Every competitive multiple is vendor-reported.The 11x-vs-Fable-5, 5x-vs-Opus-4.8-Fast, roughly-7x HLE, and 5.6x GDP-Val figures all come from Cerebras — self-run benchmarks or Cerebras’ characterization of Artificial Analysis data, not independent verification.
- 04There is no price, no model ID, and no GA date.All three are confirmed absent from the primary announcement and press coverage. Access is a waitlist for a select group of customers, expanding as capacity grows.
- 05Ultrafast is not Ultra mode.OpenAI now ships two similarly named features: Ultra mode is a parallelism feature from the GPT-5.6 GA release that fans work out to as many as 64 subagents; Ultrafast is this Cerebras-powered speed tier. They are unrelated.
01 — The AnnouncementOpenAI previews a Cerebras-powered speed tier.
OpenAI’s announcement post frames Ultrafast as a new way to run GPT-5.6 Sol, the company’s most capable model in the family — not a new model. The claim: “up to 14× faster than Standard processing,” powered by Cerebras and generating up to 750 output tokens per second. The wording is worth reading precisely. “Up to” is a ceiling, not an average, and neither the post nor any coverage discloses the workload, prompt length, or reasoning-effort setting the multiple was derived from.
Access is deliberately narrow. Ultrafast launches first in the OpenAI API as a limited preview for what the company calls a “select group of customers,” with a public signup form and a promise to “expand access as capacity grows.” Notably, Cerebras runs its own separate signup funnel for Ultrafast updates — a small but telling detail that this is a coordinated joint announcement between the model vendor and its chip partner, each courting the same waitlist.
Vendor-stated ceiling
OpenAI’s stated maximum output speed for GPT-5.6 Sol on the Ultrafast tier. A peak figure — no sustained-throughput or per-workload numbers were published alongside it.
Up to, vs Standard
OpenAI’s own framing against its Standard processing tier. The baseline workload behind the multiple is undisclosed, so treat it as a marketing ceiling until you can measure your own prompts.
Price, model ID, GA date
None of the three exist anywhere in the primary post or press coverage. Ultrafast currently reads as a service-tier selection on the existing model, not a separately priced product.
02 — The HardwareWhy wafer-scale silicon makes Sol fast.
The speed story is a hardware story. Cerebras attributes Ultrafast’s throughput to its Wafer-Scale Engine architecture, which keeps model weights on-chip — 44 GB of SRAM per wafer-scale chip, per Cerebras — instead of shuttling them to off-chip memory between tokens. That memory-bandwidth round trip is the bottleneck that typically caps GPU-based frontier-model inference speed, and removing it is where Cerebras claims the order-of-magnitude gain comes from.
This is not the first OpenAI-Cerebras speed play. Earlier in 2026, GPT-5.3-Codex-Spark shipped on Cerebras hardware as a real-time coding model — but that was a smaller, specialized model built for speed. Ultrafast’s differentiating claim is that the flagship itself now runs at those speeds: Cerebras describes the tier as “delivering up to 750 output tokens per second and without any quality compromise.” Rohan Varma, who works on product at OpenAI, put the pitch this way: “With GPT-5.6 Sol Ultrafast, Cerebras enables AI that keeps up with how you think, code, and collaborate. We’re excited to see how workflows and applications are transformed by Ultrafast inference.”
The “no quality compromise” claim is the load-bearing one — it is what separates a speed tier from a distillation — and it is also the one nobody outside the two vendors can check yet. No third-party quality evaluations of the Ultrafast preview existed at the time of writing, because almost nobody outside the select preview group can run one.
03 — ProvenanceEvery speed claim, traced to its source.
Most coverage of this launch repeats the same numbers without saying who measured them. That distinction is the difference between a benchmark and a press release, so the table below traces each circulating claim to the party that produced it and states whether anyone outside the vendor pairing can reproduce it today. The short version: the 14x and 750 figures are OpenAI’s, and every competitive comparison — including the ones invoking Claude models — comes from Cerebras, the party with the strongest incentive to look fastest.
| Claim | Who measured it | Disclosed method | Independently verifiable today? |
|---|---|---|---|
| OpenAI-stated — the announcement’s own numbers | |||
| Up to 14x faster than Standard processing | OpenAI | None — no baseline workload, prompt set, or reasoning-effort setting disclosed; “up to” marks it as a ceiling | No. The preview is waitlist-gated, so outside teams cannot yet run the comparison. |
| Up to 750 output tokens per second | OpenAI, with Cerebras | Peak figure only — no sustained-throughput or per-workload distribution published | No, for the same access reason. |
| Cerebras-run — self-reported competitive benchmarks | |||
| 11x faster than Claude Fable 5 (output speed) | Cerebras, citing Artificial Analysis data | Cerebras’ characterization of a third-party aggregator’s output-speed numbers; the underlying dataset was not independently pulled for this analysis | Partially — Artificial Analysis publishes data, but the Ultrafast side of the comparison is preview-only. |
| 5x faster than Claude Opus 4.8 on Fast mode | Cerebras, citing Artificial Analysis data | Same characterization, same caveat | Partially, with the same access limitation. |
| HLE, 2,500 questions: 11h 11m vs 78h 27m for Fable 5 — roughly 7x (our arithmetic: 671 vs 4,707 minutes) | Cerebras | Self-run: Sol Ultrafast with Codex on xhigh reasoning on July 10; Claude Fable 5 with Claude Code on xhigh reasoning on July 13-15 — different dates, “comparable accuracy” asserted, not shown | No. Self-reported, non-simultaneous runs by the party with the incentive to win. |
| 5.6x end-to-end on GDP-Val, “no quality degradation” | Cerebras | Self-run July 31, 2026: Sol vs Sol Ultrafast on medium reasoning within Codex | No — though as a same-model comparison it avoids the cross-vendor apples-to-oranges problem. |
The chain of custody is the point. A number that reads as “OpenAI says Ultrafast beats Claude” is actually OpenAI’s chip partner characterizing a third-party aggregator’s data — three hops from an independent measurement. That does not make the claims false; Cerebras hardware has posted striking throughput numbers before. It makes them unverified, which is a different planning input. For how Sol and Claude Fable 5 already compared on the dimensions you can verify — price and access — see our Sol vs Fable 5 comparison.
04 — Real-Time AgentsWhat 750 tokens per second actually changes.
Here is the practical math, and it is ours, not OpenAI’s. At 750 output tokens per second, a typical ~500-token agent response lands in roughly two-thirds of a second. Dividing OpenAI’s own numbers — 750 divided by the claimed 14x — implies a Standard-tier baseline in the ballpark of 54 tokens per second, meaning that same response takes nine-plus seconds today. One of those latencies holds a phone conversation or a checkout flow; the other does not. That derived baseline is our inference from the vendor’s two figures, not a number OpenAI states.
OpenAI names five workflow categories it built the tier for, and they share one property: a human or a system is actively waiting on the answer. The company’s own dogfooding anecdotes are qualitative but pointed — incident responders reading logs and traces, synthesizing conversations, and preparing fixes “in a fraction of the time,” and researchers turning overnight experiment-review loops into same-day iteration.
Incident response & reliability
OpenAI’s lead example and its own internal dogfooding case: reading logs and traces, synthesizing incident conversations, and preparing fixes while the outage clock is running.
Financial research & security
Workloads where the value of an answer decays in minutes — market-moving research, fraud and security triage — and a nine-second response is a materially worse product than a one-second one.
Customer support & voice
Voice agents are the harshest latency judges: sub-second response is the difference between a conversation and an answering machine. Frontier quality at voice speed is the specific promise here.
Commerce
Recommendation, checkout assistance, and post-purchase flows where every second of added latency measurably erodes conversion — a domain where teams previously defaulted to small models.
Live research & experimentation
OpenAI’s internal anecdote: experiment-review cycles that used to run overnight now iterate the same day. Qualitative, vendor-reported — but the compounding effect on research velocity is the argument.
For agencies and product teams, the support-and-voice category is the one to watch. Real-time frontier-quality agents would change what a CRM and support automation build can promise: today those systems route hard conversations to humans partly because a capable-enough model is too slow to hold the line. If the “same intelligence” claim survives contact with independent testing, that routing calculus changes.
“Whereas formerly I might have to wait a couple minutes for a task to finish, it now finishes for me before I even have the opportunity to context-switch. It makes me way more productive.”— Jeffrey Wang, OpenAI researcher, quoted on Cerebras’ launch post
The circulating speed multiples · every one is a vendor claim
Source: OpenAI + Cerebras announcement posts, Aug 13, 2026 — all vendor-stated claims, none independently verified05 — Preview RealityWhat a preview does — and doesn’t — give you.
A press cycle this loud makes it easy to forget what was actually shipped: a waitlist. The table below turns the preview’s caveats into a checklist you can act on, because the honest answer to “should I build against Ultrafast today” depends on which of these rows your plan touches.
| Dimension | What the vendors state | Status at preview |
|---|---|---|
| Access | Limited preview, select customers | Waitlist via OpenAI’s public signup form; the company says access expands as capacity grows. Cerebras runs a parallel signup for updates. |
| Pricing | Nothing disclosed | Confirmed absent across the primary post and press coverage. GPT-5.6 Sol’s standard API list remains $5 input / $30 output per 1M tokens — Ultrafast’s premium over that, if any, is unknown. |
| API model ID | None published | No separate model string appears anywhere. Ultrafast reads as a service-tier selection on the existing model — our inference from absence, not a vendor statement. |
| GA timeline | None given | No date, no quarter, no commitment beyond capacity-based expansion. |
| Quality guarantee | “Same intelligence,” “no quality compromise” | Vendor-stated only. No third-party quality evaluations of the preview existed at the time of writing. |
| Surfaces | API-first | Launches first in the OpenAI API. OpenAI named no ChatGPT or Codex availability for customers at preview, though Cerebras ran its own benchmarks with Ultrafast inside Codex. |
Read as a whole, the table says: this is an announcement you plan around, not on. There is nothing to price, nothing to call, and no committed date. What a preview does give you is lead time — a concrete signal of where OpenAI is pushing, early enough to prepare the benchmark set and the latency budget you’ll evaluate it with when access opens.
06 — DisambiguationUltrafast is not Ultra mode.
OpenAI now ships two similarly named GPT-5.6 features, and the press cycle is already blurring them. Ultra mode arrived with the GPT-5.6 GA release — a parallelism feature that fans work out to as many as 64 subagents, covered in our Ultra mode deep dive. Ultrafast is this announcement: a Cerebras-powered hardware speed tier for a single model. One multiplies the number of workers; the other makes one worker faster. They solve different problems, and neither implies the other.
The timeline context also clarifies what kind of launch this is. GPT-5.6 Sol, Terra, and Luna were first previewed in June 2026 and reached broad availability across ChatGPT, Codex, and the API in July, per 9to5Mac’s reporting timeline. Ultrafast is a speed tier bolted onto an already-GA flagship — an infrastructure announcement wearing a launch-day headline, which is exactly why the missing price and model ID matter more than they would for a true model launch.
07 — Team PlaybookWhat teams should do this week.
The decision tree depends on whether latency is currently forcing you into a smaller model — because that is the exact trade-off Ultrafast claims to dissolve.
Voice, support & checkout builds
If latency currently forces you down-model, this preview is aimed at you. Join the waitlist and prepare a benchmark set from your own prompts — but commit no architecture until a price and model ID exist.
Budget owners
Treat Ultrafast as unpriced, not free. Speed tiers across the industry typically carry a premium over standard rates, and Sol’s standard list is already $5/$30 per 1M. Scenario-model it rather than penciling in a number.
Standard-tier workloads
Nothing changes today. Standard processing, current pricing, and existing model IDs are untouched by this announcement. No migration exists to plan, because there is nothing yet to migrate to.
Teams that need speed now
Fast options already exist across vendors — smaller models, fast modes, other accelerated hardware. Measure TTFT and throughput on your own workloads instead of waiting on a waitlist or trusting a vendor multiple.
Two resources pair naturally with that last row: our latency and throughput benchmark data for measuring what today’s tiers actually deliver, and our inference FinOps playbook for budgeting a faster-but-likely-pricier tier before the price exists.
The forward-looking read: speed is becoming its own competitive axis, separate from capability. The same day Ultrafast previewed, Google shipped its own speed-and-price play with Gemini 3.7 Flash — a different bet on the same customer anxiety about latency and cost. Anthropic already sells a fast mode for Claude, which TechCrunch’s coverage argued “doesn’t deliver the kind of speed that OpenAI is offering here” — an editorial judgment, not a measured comparison, but a sign of how the narrative is forming. Expect every frontier vendor to ship a named speed tier within a few quarters, and expect the marketing multiples to keep arriving faster than the independent verification. Teams that maintain their own latency benchmarks will be the only ones able to tell the difference — if your stack decisions are riding on it, our AI transformation engagements start with exactly that kind of comparative eval.
08 — ConclusionA speed race with vendor-graded scorecards.
Fast is plausible; verified is absent — treat Ultrafast as lead time, not a product.
The Ultrafast preview is a genuinely significant signal: OpenAI putting its most capable GPT-5.6 model on Cerebras silicon and claiming frontier quality at up to 750 tokens per second is a direct attack on the last structural reason to run smaller models in real-time products. If the “no quality compromise” claim holds up under independent testing, the latency tax that shaped a generation of agent architectures starts to disappear.
But every load-bearing number in this launch is vendor-graded. The 14x is a ceiling with no disclosed baseline; the competitive multiples are Cerebras characterizing third-party data or running its own non-simultaneous tests; and the tier itself has no price, no model ID, and no date. That is not a criticism of the engineering — it is the normal shape of a preview announcement, and the correct response is the normal one: excitement on a delay.
The practical move is to use the lead time. Build the benchmark set from your own prompts, set the latency budgets your products actually need, join the waitlist if real-time frontier quality would change your roadmap — and grade Ultrafast against your own numbers the day access arrives, not against anyone’s launch-day multiple.