AI DevelopmentNew Release10 min readPublished August 13, 2026

Cerebras-powered speed tier · up to 14x vendor-claimed · limited preview

GPT-5.6 Sol Ultrafast: OpenAI Previews 14x Inference

OpenAI previewed Ultrafast on August 13, 2026 — a Cerebras-powered service tier that runs GPT-5.6 Sol at a vendor-claimed up-to-14x standard speed, peaking at 750 output tokens per second. The fine print matters: it’s a waitlist-gated limited preview with no price, no GA date, no separate model ID — and every competitive multiple traces back to the vendors themselves.

DA
Digital Applied Team
Senior strategists · Published Aug 13, 2026
PublishedAugust 13, 2026
Read time10 min
Sources5 vendor + press
Peak output speed
750tok/s
vendor-stated ceiling
up to
Claimed multiple
14x
vs Standard processing
up to
Ultrafast pricing
None
disclosed at preview
GA timeline
TBD
waitlist-gated access

GPT-5.6 Sol Ultrafast is OpenAI’s newest service tier, previewed on August 13, 2026: the company’s most capable GPT-5.6 model running on Cerebras hardware at what OpenAI describes as “up to 14× faster than Standard processing,” peaking at 750 output tokens per second. It launches first in the API, behind a waitlist, for a select group of customers.

The announcement matters because it attacks a trade-off every team building real-time AI products knows well: until now, getting conversational-speed responses usually meant stepping down to a smaller or more specialized model. Ultrafast claims to remove that compromise — frontier-model output at speeds that keep pace with a live conversation, an incident bridge, or a checkout flow.

This guide covers what OpenAI actually announced, why Cerebras hardware is the enabling story, where every headline speed claim comes from (and who measured it), what 750 tokens per second changes for real-time agentic products, and what a limited preview does — and does not — let you act on today. Every number below is traced to the vendor post that stated it.

Key takeaways
  1. 01
    Ultrafast is a speed tier, not a new model.It runs the existing GPT-5.6 Sol — previewed in June 2026, broadly available since July — on Cerebras hardware. OpenAI and Cerebras both position it as a latency-only change with the same claimed intelligence.
  2. 02
    The headline numbers are vendor ceilings.Up to 14x standard speed and up to 750 output tokens per second are OpenAI’s own figures, with no disclosed baseline workload, prompt set, or reasoning-effort setting behind the multiple.
  3. 03
    Every competitive multiple is vendor-reported.The 11x-vs-Fable-5, 5x-vs-Opus-4.8-Fast, roughly-7x HLE, and 5.6x GDP-Val figures all come from Cerebras — self-run benchmarks or Cerebras’ characterization of Artificial Analysis data, not independent verification.
  4. 04
    There is no price, no model ID, and no GA date.All three are confirmed absent from the primary announcement and press coverage. Access is a waitlist for a select group of customers, expanding as capacity grows.
  5. 05
    Ultrafast is not Ultra mode.OpenAI now ships two similarly named features: Ultra mode is a parallelism feature from the GPT-5.6 GA release that fans work out to as many as 64 subagents; Ultrafast is this Cerebras-powered speed tier. They are unrelated.

01The AnnouncementOpenAI previews a Cerebras-powered speed tier.

OpenAI’s announcement post frames Ultrafast as a new way to run GPT-5.6 Sol, the company’s most capable model in the family — not a new model. The claim: “up to 14× faster than Standard processing,” powered by Cerebras and generating up to 750 output tokens per second. The wording is worth reading precisely. “Up to” is a ceiling, not an average, and neither the post nor any coverage discloses the workload, prompt length, or reasoning-effort setting the multiple was derived from.

Access is deliberately narrow. Ultrafast launches first in the OpenAI API as a limited preview for what the company calls a “select group of customers,” with a public signup form and a promise to “expand access as capacity grows.” Notably, Cerebras runs its own separate signup funnel for Ultrafast updates — a small but telling detail that this is a coordinated joint announcement between the model vendor and its chip partner, each courting the same waitlist.

Peak output speed
Vendor-stated ceiling
750tok/s

OpenAI’s stated maximum output speed for GPT-5.6 Sol on the Ultrafast tier. A peak figure — no sustained-throughput or per-workload numbers were published alongside it.

OpenAI + Cerebras, Aug 13
Speed multiple
Up to, vs Standard
14x

OpenAI’s own framing against its Standard processing tier. The baseline workload behind the multiple is undisclosed, so treat it as a marketing ceiling until you can measure your own prompts.

No baseline disclosed
Disclosures
Price, model ID, GA date
0

None of the three exist anywhere in the primary post or press coverage. Ultrafast currently reads as a service-tier selection on the existing model, not a separately priced product.

Confirmed absent
OpenAI’s own framing
“Until now, getting real-time speed typically meant choosing a smaller or more specialized model. Ultrafast points to progress in a new direction: more useful work per second.” — OpenAI, Previewing Ultrafast, August 13, 2026. That sentence is the entire pitch: frontier capability without the historical latency tax.

02The HardwareWhy wafer-scale silicon makes Sol fast.

The speed story is a hardware story. Cerebras attributes Ultrafast’s throughput to its Wafer-Scale Engine architecture, which keeps model weights on-chip — 44 GB of SRAM per wafer-scale chip, per Cerebras — instead of shuttling them to off-chip memory between tokens. That memory-bandwidth round trip is the bottleneck that typically caps GPU-based frontier-model inference speed, and removing it is where Cerebras claims the order-of-magnitude gain comes from.

This is not the first OpenAI-Cerebras speed play. Earlier in 2026, GPT-5.3-Codex-Spark shipped on Cerebras hardware as a real-time coding model — but that was a smaller, specialized model built for speed. Ultrafast’s differentiating claim is that the flagship itself now runs at those speeds: Cerebras describes the tier as “delivering up to 750 output tokens per second and without any quality compromise.” Rohan Varma, who works on product at OpenAI, put the pitch this way: “With GPT-5.6 Sol Ultrafast, Cerebras enables AI that keeps up with how you think, code, and collaborate. We’re excited to see how workflows and applications are transformed by Ultrafast inference.”

The “no quality compromise” claim is the load-bearing one — it is what separates a speed tier from a distillation — and it is also the one nobody outside the two vendors can check yet. No third-party quality evaluations of the Ultrafast preview existed at the time of writing, because almost nobody outside the select preview group can run one.

03ProvenanceEvery speed claim, traced to its source.

Most coverage of this launch repeats the same numbers without saying who measured them. That distinction is the difference between a benchmark and a press release, so the table below traces each circulating claim to the party that produced it and states whether anyone outside the vendor pairing can reproduce it today. The short version: the 14x and 750 figures are OpenAI’s, and every competitive comparison — including the ones invoking Claude models — comes from Cerebras, the party with the strongest incentive to look fastest.

Provenance audit of every speed claim in the GPT-5.6 Sol Ultrafast announcement, grouped by which vendor produced the claim, with the disclosed method and whether the claim is independently verifiable at the time of the preview.
ClaimWho measured itDisclosed methodIndependently verifiable today?
OpenAI-stated — the announcement’s own numbers
Up to 14x faster than Standard processingOpenAINone — no baseline workload, prompt set, or reasoning-effort setting disclosed; “up to” marks it as a ceilingNo. The preview is waitlist-gated, so outside teams cannot yet run the comparison.
Up to 750 output tokens per secondOpenAI, with CerebrasPeak figure only — no sustained-throughput or per-workload distribution publishedNo, for the same access reason.
Cerebras-run — self-reported competitive benchmarks
11x faster than Claude Fable 5 (output speed)Cerebras, citing Artificial Analysis dataCerebras’ characterization of a third-party aggregator’s output-speed numbers; the underlying dataset was not independently pulled for this analysisPartially — Artificial Analysis publishes data, but the Ultrafast side of the comparison is preview-only.
5x faster than Claude Opus 4.8 on Fast modeCerebras, citing Artificial Analysis dataSame characterization, same caveatPartially, with the same access limitation.
HLE, 2,500 questions: 11h 11m vs 78h 27m for Fable 5 — roughly 7x (our arithmetic: 671 vs 4,707 minutes)CerebrasSelf-run: Sol Ultrafast with Codex on xhigh reasoning on July 10; Claude Fable 5 with Claude Code on xhigh reasoning on July 13-15 — different dates, “comparable accuracy” asserted, not shownNo. Self-reported, non-simultaneous runs by the party with the incentive to win.
5.6x end-to-end on GDP-Val, “no quality degradation”CerebrasSelf-run July 31, 2026: Sol vs Sol Ultrafast on medium reasoning within CodexNo — though as a same-model comparison it avoids the cross-vendor apples-to-oranges problem.

The chain of custody is the point. A number that reads as “OpenAI says Ultrafast beats Claude” is actually OpenAI’s chip partner characterizing a third-party aggregator’s data — three hops from an independent measurement. That does not make the claims false; Cerebras hardware has posted striking throughput numbers before. It makes them unverified, which is a different planning input. For how Sol and Claude Fable 5 already compared on the dimensions you can verify — price and access — see our Sol vs Fable 5 comparison.

Reading vendor benchmarks
The HLE comparison ran the two models on different dates with different harnesses — Sol Ultrafast via Codex on July 10, Claude Fable 5 via Claude Code on July 13-15 — and asserts “comparable accuracy” without publishing the scores. Speed measured under non-identical conditions is a demonstration, not a benchmark. Ask any vendor quoting a competitive multiple the same three questions: who ran it, on what workload, and can I reproduce it?

04Real-Time AgentsWhat 750 tokens per second actually changes.

Here is the practical math, and it is ours, not OpenAI’s. At 750 output tokens per second, a typical ~500-token agent response lands in roughly two-thirds of a second. Dividing OpenAI’s own numbers — 750 divided by the claimed 14x — implies a Standard-tier baseline in the ballpark of 54 tokens per second, meaning that same response takes nine-plus seconds today. One of those latencies holds a phone conversation or a checkout flow; the other does not. That derived baseline is our inference from the vendor’s two figures, not a number OpenAI states.

OpenAI names five workflow categories it built the tier for, and they share one property: a human or a system is actively waiting on the answer. The company’s own dogfooding anecdotes are qualitative but pointed — incident responders reading logs and traces, synthesizing conversations, and preparing fixes “in a fraction of the time,” and researchers turning overnight experiment-review loops into same-day iteration.

Ops
Incident response & reliability
logs → synthesis → fix prep

OpenAI’s lead example and its own internal dogfooding case: reading logs and traces, synthesizing incident conversations, and preparing fixes while the outage clock is running.

OpenAI-named workflow
Finance
Financial research & security
time-sensitive analysis

Workloads where the value of an answer decays in minutes — market-moving research, fraud and security triage — and a nine-second response is a materially worse product than a one-second one.

OpenAI-named workflow
Support
Customer support & voice
conversational latency budgets

Voice agents are the harshest latency judges: sub-second response is the difference between a conversation and an answering machine. Frontier quality at voice speed is the specific promise here.

OpenAI-named workflow
Commerce
Commerce
checkout & product flows

Recommendation, checkout assistance, and post-purchase flows where every second of added latency measurably erodes conversion — a domain where teams previously defaulted to small models.

OpenAI-named workflow
Research
Live research & experimentation
overnight loops → same-day

OpenAI’s internal anecdote: experiment-review cycles that used to run overnight now iterate the same day. Qualitative, vendor-reported — but the compounding effect on research velocity is the argument.

OpenAI-named workflow

For agencies and product teams, the support-and-voice category is the one to watch. Real-time frontier-quality agents would change what a CRM and support automation build can promise: today those systems route hard conversations to humans partly because a capable-enough model is too slow to hold the line. If the “same intelligence” claim survives contact with independent testing, that routing calculus changes.

“Whereas formerly I might have to wait a couple minutes for a task to finish, it now finishes for me before I even have the opportunity to context-switch. It makes me way more productive.”— Jeffrey Wang, OpenAI researcher, quoted on Cerebras’ launch post

The circulating speed multiples · every one is a vendor claim

Source: OpenAI + Cerebras announcement posts, Aug 13, 2026 — all vendor-stated claims, none independently verified
Up to 14x vs Sol StandardOpenAI · no baseline workload disclosed
14x
11x vs Claude Fable 5Cerebras, citing Artificial Analysis output-speed data
11x
~7x on HLE completion timeCerebras-run · 11h 11m vs 78h 27m — our arithmetic
7x
5.6x end-to-end on GDP-ValCerebras-run, Jul 31 · Sol vs Sol Ultrafast, medium reasoning
5.6x
5x vs Claude Opus 4.8 Fast modeCerebras, citing Artificial Analysis output-speed data
5x

05Preview RealityWhat a preview does — and doesn’t — give you.

A press cycle this loud makes it easy to forget what was actually shipped: a waitlist. The table below turns the preview’s caveats into a checklist you can act on, because the honest answer to “should I build against Ultrafast today” depends on which of these rows your plan touches.

Dimension
Access
What the vendors state
Limited preview, select customers
Status at preview
Waitlist via OpenAI’s public signup form; the company says access expands as capacity grows. Cerebras runs a parallel signup for updates.
Dimension
Pricing
What the vendors state
Nothing disclosed
Status at preview
Confirmed absent across the primary post and press coverage. GPT-5.6 Sol’s standard API list remains $5 input / $30 output per 1M tokens — Ultrafast’s premium over that, if any, is unknown.
Dimension
API model ID
What the vendors state
None published
Status at preview
No separate model string appears anywhere. Ultrafast reads as a service-tier selection on the existing model — our inference from absence, not a vendor statement.
Dimension
GA timeline
What the vendors state
None given
Status at preview
No date, no quarter, no commitment beyond capacity-based expansion.
Dimension
Quality guarantee
What the vendors state
“Same intelligence,” “no quality compromise”
Status at preview
Vendor-stated only. No third-party quality evaluations of the preview existed at the time of writing.
Dimension
Surfaces
What the vendors state
API-first
Status at preview
Launches first in the OpenAI API. OpenAI named no ChatGPT or Codex availability for customers at preview, though Cerebras ran its own benchmarks with Ultrafast inside Codex.

Read as a whole, the table says: this is an announcement you plan around, not on. There is nothing to price, nothing to call, and no committed date. What a preview does give you is lead time — a concrete signal of where OpenAI is pushing, early enough to prepare the benchmark set and the latency budget you’ll evaluate it with when access opens.

06DisambiguationUltrafast is not Ultra mode.

OpenAI now ships two similarly named GPT-5.6 features, and the press cycle is already blurring them. Ultra mode arrived with the GPT-5.6 GA release — a parallelism feature that fans work out to as many as 64 subagents, covered in our Ultra mode deep dive. Ultrafast is this announcement: a Cerebras-powered hardware speed tier for a single model. One multiplies the number of workers; the other makes one worker faster. They solve different problems, and neither implies the other.

The timeline context also clarifies what kind of launch this is. GPT-5.6 Sol, Terra, and Luna were first previewed in June 2026 and reached broad availability across ChatGPT, Codex, and the API in July, per 9to5Mac’s reporting timeline. Ultrafast is a speed tier bolted onto an already-GA flagship — an infrastructure announcement wearing a launch-day headline, which is exactly why the missing price and model ID matter more than they would for a true model launch.

Naming collision
When evaluating vendor claims this quarter, check which feature a number belongs to: Ultra mode claims are about parallel subagents and task decomposition; Ultrafast claims are about single-stream output speed on Cerebras hardware. A “GPT-5.6 is faster now” headline could mean either — and the evidence standards for each are different.

07Team PlaybookWhat teams should do this week.

The decision tree depends on whether latency is currently forcing you into a smaller model — because that is the exact trade-off Ultrafast claims to dissolve.

Real-time agents
Voice, support & checkout builds

If latency currently forces you down-model, this preview is aimed at you. Join the waitlist and prepare a benchmark set from your own prompts — but commit no architecture until a price and model ID exist.

Join the waitlist
Cost planning
Budget owners

Treat Ultrafast as unpriced, not free. Speed tiers across the industry typically carry a premium over standard rates, and Sol’s standard list is already $5/$30 per 1M. Scenario-model it rather than penciling in a number.

Model it as unpriced
Current Sol users
Standard-tier workloads

Nothing changes today. Standard processing, current pricing, and existing model IDs are untouched by this announcement. No migration exists to plan, because there is nothing yet to migrate to.

No action needed
Speed shoppers
Teams that need speed now

Fast options already exist across vendors — smaller models, fast modes, other accelerated hardware. Measure TTFT and throughput on your own workloads instead of waiting on a waitlist or trusting a vendor multiple.

Benchmark today’s options

Two resources pair naturally with that last row: our latency and throughput benchmark data for measuring what today’s tiers actually deliver, and our inference FinOps playbook for budgeting a faster-but-likely-pricier tier before the price exists.

The forward-looking read: speed is becoming its own competitive axis, separate from capability. The same day Ultrafast previewed, Google shipped its own speed-and-price play with Gemini 3.7 Flash — a different bet on the same customer anxiety about latency and cost. Anthropic already sells a fast mode for Claude, which TechCrunch’s coverage argued “doesn’t deliver the kind of speed that OpenAI is offering here” — an editorial judgment, not a measured comparison, but a sign of how the narrative is forming. Expect every frontier vendor to ship a named speed tier within a few quarters, and expect the marketing multiples to keep arriving faster than the independent verification. Teams that maintain their own latency benchmarks will be the only ones able to tell the difference — if your stack decisions are riding on it, our AI transformation engagements start with exactly that kind of comparative eval.

08ConclusionA speed race with vendor-graded scorecards.

The shape of inference, August 2026

Fast is plausible; verified is absent — treat Ultrafast as lead time, not a product.

The Ultrafast preview is a genuinely significant signal: OpenAI putting its most capable GPT-5.6 model on Cerebras silicon and claiming frontier quality at up to 750 tokens per second is a direct attack on the last structural reason to run smaller models in real-time products. If the “no quality compromise” claim holds up under independent testing, the latency tax that shaped a generation of agent architectures starts to disappear.

But every load-bearing number in this launch is vendor-graded. The 14x is a ceiling with no disclosed baseline; the competitive multiples are Cerebras characterizing third-party data or running its own non-simultaneous tests; and the tier itself has no price, no model ID, and no date. That is not a criticism of the engineering — it is the normal shape of a preview announcement, and the correct response is the normal one: excitement on a delay.

The practical move is to use the lead time. Build the benchmark set from your own prompts, set the latency budgets your products actually need, join the waitlist if real-time frontier quality would change your roadmap — and grade Ultrafast against your own numbers the day access arrives, not against anyone’s launch-day multiple.

Build real-time AI on verified numbers

Speed tiers only pay off when the latency budget is yours.

Our team helps businesses design, benchmark, and ship real-time AI agents — voice, support, and commerce — with latency budgets and model routing grounded in measurements you own, not vendor multiples.

Free consultationExpert guidanceTailored solutions
What we work on

Real-time AI engagements

  • Latency benchmarking across frontier tiers on your prompts
  • Voice & support agents with strict latency budgets
  • Model routing — speed tiers vs smaller models per workload
  • Inference cost modeling for unpriced preview tiers
  • Vendor-claim verification before stack commitments
FAQ · GPT-5.6 Sol Ultrafast

The questions we get every week.

Ultrafast is a new OpenAI service tier, previewed on August 13, 2026, that runs GPT-5.6 Sol — the most capable model in the GPT-5.6 family — on Cerebras hardware at what OpenAI describes as up to 14x the speed of Standard processing, peaking at 750 output tokens per second. It is not a new model: Sol itself was previewed in June 2026 and became broadly available in July. Ultrafast launches first in the OpenAI API as a limited preview for a select group of customers, with access managed through a waitlist that OpenAI says will expand as capacity grows.
Related dispatches

Continue exploring frontier releases.