AI DevelopmentNew Release8 min readPublished September 3, 2026

A 1.05-million-token context window, a Critical cyber rating and a benchmark table that needs reading, not repeating

GPT-6 Astra: Price, Access and What the Benchmarks Show

OpenAI’s launch numbers put GPT-6 Astra at the front of computer use, scientific terminal work, long-context retrieval and several cyber evaluations. They also show losses on general and coding indices. The buying decision turns on those limits, a long-context price multiplier and safety controls that can stop a task.

DA
Digital Applied Team
Senior strategists · Published Sep 3, 2026
PublishedSep 3, 2026
Read time8 min
SourcesOpenAI + SRE-Bench
Standard API price, input / output per M
$10 / $50
$1 cached input; $12.50 cache write (OpenAI)
Context / maximum output
1.05M / 128K
922K maximum input tokens (OpenAI)
ARC-AGI-3
99.9%
vendor-run with an OpenAI Responses API harness
Preparedness rating
Critical
cybersecurity; High for biological and chemical capability

OpenAI released GPT-6 Astra on September 3, 2026 at $10 per million input tokens, $1 per million cached input tokens and $50 per million output tokens. The API model is gpt-6-astra, with a 1,050,000-token context window, 922,000 maximum input tokens and 128,000 maximum output. ChatGPT access is rolling out over several days to Plus, Pro, Business and Enterprise rather than appearing everywhere at once.

The headline score is 99.9% on ARC-AGI-3. The more useful story is wider: 72.6% on OSWorld 2.0, 64.6% on Terminal-Bench Science and 88% pass@1 on the public SRE-Bench, all reported in OpenAI’s launch material. Astra also trails Claude Fable 5.1 on the Artificial Analysis Intelligence Index shown by OpenAI and trails several models on Humanity’s Last Exam with tools. This is a specialist frontier model with broad capability, not an automatic replacement for every cheaper route.

Key takeaways
  1. 01
    $10 in, $1 cached and $50 out is only the base lane.OpenAI also lists a $12.50 cache-write rate. Above 272,000 input tokens, input and cache prices double and output rises 50% for the full request, not just the excess tokens.
  2. 02
    The clearest gains are in tool-heavy work.OpenAI reports substantial gains over GPT-5.6 Sol on computer use, scientific terminal tasks, reverse engineering and recent-vulnerability cyber work. The company also changed harnesses or conditions on several evaluations.
  3. 03
    Astra does not lead every row OpenAI published.It trails on Humanity's Last Exam with tools, FrontierCode Extended and both Artificial Analysis indices in the launch table. Maximum-at-any-effort scores are not same-cost comparisons.
  4. 04
    Critical cyber capability brings operational friction.OpenAI says monitoring may slow, pause or stop legitimate work. Better observed alignment came with lower chain-of-thought monitorability under adversarial testing, so the safety story has two directions.

01The releaseWhat launched, and who gets it.

GPT-6 Astra is available through the OpenAI API and, according to OpenAI, Amazon Bedrock. OpenAI names GPT-6 Astra Pro for ChatGPT Pro, Business and Enterprise, but its launch sources do not publish a separate Astra Pro API model ID, price or specification sheet. Plus receives Astra, not the Pro variant. Enterprise access is off by default and requires an administrator to enable it.

The rollout begins with a limited set of organisations and expands over several days. Zero Data Retention is supported for eligible API customers; “eligible” matters because it is not a universal promise. OpenAI does not disclose Astra’s parameter count, architecture, training compute or training-data size. Its published April 30, 2026 knowledge cutoff is not the same thing as a training-data cutoff.

The name appeared before the product. We covered Astra’s August mathematics disclosure and OpenAI’s pre-launch cyber-risk assessment. Those articles record what was known then; today is the first broad product launch.

02The rate cardThe million-token window has a 272K price line.

Astra and Claude Fable 5.1 share the same headline input and output rates: $10 and $50 per million tokens. They are not identically priced. OpenAI lists Astra cache reads at $1 per million; Anthropic lists Fable 5.1 cache reads at $0.25. GPT-5.6 Sol’s current promotional rates remain much cheaper at $4 input, $0.40 cached input and $20 output in OpenAI’s published rate card.

Model
Input / cached / output
Context
What changes the bill
GPT-6 Astra
$10 / $1 / $50
1.05M
A $12.50 cache-write rate and a full-request premium above 272K input.
Claude Fable 5.1
$10 / $0.25 / $50
1M
Lower cache reads; two cache-write TTL tiers and separate long-context rules.
GPT-5.6 Sol
$4 / $0.40 / $20
1.05M
Promotional through at least Nov 21; Astra may still win on cost per solved task.

Published API rates per million tokens, retrieved September 3, 2026. Sources: OpenAI Astra model page and Anthropic Fable 5.1 overview. Context is maximum context, not base-rate entitlement.

The long-input premium

Once an Astra request crosses 272,000 input tokens, OpenAI charges twice the applicable input and cache rates and 1.5 times the output rate for the entire request. At Standard rates, that means $20 per million uncached input tokens and $75 per million output tokens. Batch and Flex cost half the applicable rate; Fast costs twice the applicable rate and is unavailable with EU data residency.

The API accepts text and image input and produces text. It supports Responses, Chat Completions and Batch, but tool calling requires the Responses API. Realtime, Assistants, fine-tuning, embeddings and native image, video or audio generation are not supported. The model exposes low, medium, high, xhigh and max reasoning; there is no none setting.

03The evidenceThe strongest case is work across tools.

OpenAI’s clearest launch-day gains sit in environments where a model has to inspect state, act, recover and finish. Astra scores 59.3% on Agents’ Last Exam against 53.6% for Sol, and 72.6% on the OSWorld 2.0 offline partial score against Sol’s 65.7%. OpenAI also reports roughly 40 minutes per OSWorld task for Astra versus 75 for Sol, though runtime depends on the harness as well as the model.

EvaluationAstraGPT-5.6 SolReading
Agents’ Last Exam59.3%53.6%Computer-use lead in the table shown
OSWorld 2.0 partial72.6%65.7%Lead plus lower reported task time
Terminal-Bench Science 0.164.6%22.4%Large gap on scientific terminal work
SRE-Bench pass@188.0%55.9%Public reverse-engineering benchmark
MRCR v2, 512K–1M96.3%73.8%Eight-needle long-context retrieval
ARC-AGI-399.9%7.8%Extraordinary, harness-specific result

Scores are OpenAI-reported launch results, generally the maximum at any effort, not a same-cost comparison. SRE-Bench is a public benchmark; OpenAI ran the model. Source: GPT-6 Astra launch, September 3, 2026.

The cyber and professional-work rows point in the same direction. OpenAI reports 95.9% geometric overlap on BenchCAD, 41.4% on AutomationBench and 39% on a recent-vulnerability V8 evaluation where Sol scored 5.5%. Astra also reached 88% pass@1 and 99.2% pass@4 on SRE-Bench. OpenAI says the V8 work found two previously unknown vulnerabilities and that both are being disclosed to maintainers; the operational details do not belong in a product comparison.

04The caveatA maximum score is not a same-cost test.

Most launch scores came from OpenAI’s research environment or API, not an independent replication of production ChatGPT. OpenAI takes each model’s best score at any reasoning effort, so the table says which configuration reached the highest point, not which one reached it for the least money or time. Research and ChatGPT also use different system prompts, tools and inference settings.

Astra loses rows in OpenAI’s own appendix. It scores 57.2% on Humanity’s Last Exam with tools, behind Sol at 65% and Fable 5.1 at 63.8%. Its 64.5% on FrontierCode Extended trails Fable 5 at 64.9%, while its 67 on the Artificial Analysis Coding Agent Index trails Fable 5 at 68.1. On the broader Artificial Analysis Intelligence Index, Astra’s 61.2 trails Fable 5.1 at 65.7.

Methodology

We treat OpenAI’s launch table as vendor-reported evidence and use public benchmark documentation only where it is available.

Score selection
Values use the highest effort shown unless stated otherwise. That favours capability ceilings over equal-budget comparisons.
Exact values
OpenAI’s prose says 57.9% on Terminal-Bench 4.0 while its table says 57.7%; we use the table. FrontierMath is 97.6% in the table and 98% when rounded in the headline.
Harness changes
ARC-AGI-3 uses two Responses API settings intended to better match real use. The 1.9× Mind2Web speed claim combines Astra with an updated Codex harness.
Comparison limits
Some Claude results were reproduced by OpenAI, some use different settings, and ScreenSpot-Pro and ExploitGym substitute the less-restricted Mythos configuration for Fable.
Cyber conditions
ExploitBench and ExploitGym remove production safeguards to measure raw capability. Astra and Sol also ran ExploitGym without its six-hour limit.

ARC-AGI-3 is the sharpest example. OpenAI explains that Astra’s 99.9% used a Responses API harness with two settings meant to reflect real-world use; they were not designed specifically for the benchmark. That makes the score relevant and still not equivalent to a plain default run. Extraordinary numbers deserve more context, not less.

05The APIThe useful changes happen mid-task.

Astra introduces asynchronous tool calling in the Responses API. A model can issue a tool call, continue reasoning about independent work and later consume the result when the application returns it under the original call_id. For agents with slow searches, code execution or external approvals, that changes the amount of idle time inside a run.

Mid-turn steering lets an application add instructions over a WebSocket without discarding completed work. A configuration_update can also change reasoning effort during a conversation while preserving the prompt prefix, subject to compatibility rules. Those are orchestration features, not reasons to skip workload evals: their value appears only when the surrounding agent loop uses them correctly.

Migration checks

Remove temperature, top_p and top_logprobs; Chat Completions also removes logprobs. When moving from GPT-5.5 or earlier, replace prompt_cache_retention with prompt_cache_options.ttl: "30m". Use Responses, not Chat Completions, for tool calling.

The published knowledge cutoff is April 30, 2026. For work that depends on current state, that still means retrieval or tools. The larger context window reduces how often a team has to compact, but it does not remove the need for context design—and above 272K input, the decision also changes the rate lane.

06The safeguardsBetter alignment, lower monitorability.

OpenAI rates Astra Critical for cybersecurity capability, High for biological and chemical capability, and below High for AI self-improvement. Critical is a capability assessment, not a claim that the production endpoint will perform unrestricted cyber work. At launch, the shipping system refuses advanced exploit generation and places tool use under universal monitoring.

The alignment evidence is encouraging but internal. In a simulation of more than 54,000 Codex tasks, OpenAI reports roughly half as many higher-severity misalignment flags as Sol. In an impossible cyber task without production safeguards, Astra showed 0% unauthorised- target behaviour against 48% for Sol. Neither number is a production incident rate.

The counterweight sits in the GPT-6 Astra system card: under adversarial tests, chain-of-thought monitorability was lower than Sol’s. OpenAI spent at least 200,000 A100-equivalent GPU-hours on one measured portion of red teaming, excluding attacker and helper inference, and built controls that can slow, pause or stop a task. For API users, a triggered task stops. That interruption rate belongs in the eval alongside accuracy and cost.

The distinction matters because an agent can appear aligned while becoming harder to inspect. Our earlier analysis of agents manipulating their own evidence explains why behaviour and observability need separate measures.

07The decisionRoute the task, not the announcement.

Astra has a credible case when a task is long, tool-heavy and expensive to fail: computer use, reverse engineering, scientific terminal work, complex professional workflows and retrieval across hundreds of thousands of tokens. It has a weaker case when a cheaper model already clears the acceptance bar, when cache reads dominate the bill or when a safety stop would be costly.

Hard tool workflows
Run Astra against your longest computer-use, terminal, reverse-engineering or cross-document tasks. Measure solved tasks and elapsed time, not only benchmark rank.
Evaluate Astra
Large cached prefixes
Compare cache-hit share and total output. Astra's $1 cache-read rate is four times Fable 5.1's, and requests over 272K input move the full request into a higher rate lane.
Price the trace
Routine production
Keep Sol or another lower-cost route when it already meets the quality bar. Astra's own launch table does not establish a universal lead on general or coding indices.
Stay or route
Cyber-sensitive work
Include refusal and safety-stop rates, authorisation boundaries and human-review time. Raw cyber benchmark capability is not equivalent to production availability.
Test controls

A useful bake-off records completed-task success, latency, uncached and cached input, output tokens, tool-call cost, retries, safety-stop rate and human-review time for every route. Run at least two reasoning efforts. The model with the higher token rate can be cheaper per solved task; the model with the higher benchmark can be more expensive on your distribution.

This is the same routing discipline behind our GPT-5.6 Sol, Terra and Luna analysis. A model family is useful when the application can change routes without rewriting the workflow. Our AI transformation practice builds that instrumentation and evaluation layer around real tasks.

08ConclusionAstra is a ceiling to test, not a default to assume.

GPT-6 Astra

The launch makes a strong case for hard, tool-heavy work—and a weak case for blanket migration.

The release is consequential. A 1.05-million-token context window, 128,000 maximum output, asynchronous tools and large gains on computer use and scientific terminal work expand what one model run can finish. The recent-vulnerability and SRE-Bench results also explain why OpenAI classifies the model at the highest cyber capability tier.

The rate card and benchmark notes narrow the claim. The base price is $10/$50, cache writes cost more, cache reads cost four times Fable 5.1’s, and a request above 272K input moves wholesale into a higher rate lane. OpenAI’s own table includes losses, harness changes and maximum-at-any-effort reporting.

So the correct first move is not “switch.” It is to send Astra the work your current route fails, capture the whole trace, and ask whether higher completion offsets price, controls and review time. Astra may be the best model for that task. The launch does not prove it is the best model for every task around it.

Evaluate the workload

Find the tasks where Astra earns the premium.

We turn vendor rate cards and launch benchmarks into a workload eval: completed-task cost, latency, cache behaviour, safety stops and the routing rules that survive the next release.

Free consultationExpert guidanceTailored solutions
What we measure

Frontier-model evaluation

  • Completed-task cost across reasoning levels
  • Cache hits and long-context rate changes
  • Tool latency, retries and failure recovery
  • Safety-stop and human-review overhead
  • Routing rules for production workloads
FAQ · GPT-6 Astra

The questions we get about GPT-6 Astra.

OpenAI lists Standard API rates of $10 per million input tokens, $1 per million cached input tokens, $12.50 per million cache-write tokens and $50 per million output tokens. Above 272,000 input tokens, input and cache rates double and output rises 50% for the full request. Batch and Flex are half the applicable price; Fast is twice the applicable price.
Related dispatches

Continue exploring OpenAI and frontier models.