AI DevelopmentDecision Matrix16 min readPublished October 4, 2026

A maintained guide to model selection

Which Frontier AI Model Should You Use?

Choose a model against the work it must finish, the API you can access, and the bill you can explain. Current provider records, explicit cost assumptions, and the dated comparisons behind the discussion.

DA
Digital Applied Team
Model selection reference
PublishedOct 4, 2026
Sources checkedOct 4, 2026
Read time16 min
Selected records
16
direct APIs and an access-limited announcement
Workload classes
5
each needs its own acceptance test
Historical chapters
42
original dates and URLs preserved

Start with the task you can judge. A coding change needs working tests and a reviewable diff; an invoice extraction needs correct fields; a voice assistant needs to handle interruptions. The useful frontier model is the one that meets that specific standard through an API you can use, at a total cost you have measured. A launch ranking cannot make that decision for you.

This guide joins current provider evidence to a practical selection method. The tables cover selected models from OpenAI, Anthropic, Google, SpaceXAI, and DeepSeek, including lower-cost candidates and voice endpoints that solve different parts of a workflow. They are a bounded reference, not a claim that these are the only capable models or a universal performance ranking.

For an initial coding evaluation, compare GPT-6.1 Sol and Claude Sonnet 5.5 with Claude Opus 5.5 on the same work. For narrowly specified extraction, include GPT-6 Luna and Gemini 3.5 Flash-Lite. Treat these as candidate sets chosen from documented capabilities and prices. We have not run a cross-provider benchmark for this guide, and none has earned a production recommendation merely by appearing here.

Key takeaways
  1. 01
    Buy the endpoint, not the model nameAn API rate, a coding subscription, a cloud reseller, and a restricted preview are different products. Record the identifier and surface before comparing cost.
  2. 02
    Test a starting route and an escalation routeUse representative examples with explicit pass conditions. A lower token rate is useful only after the route can finish the work to your standard.
  3. 03
    Count the whole attemptInput, cache reads and writes, generated reasoning, tool calls, retries, and human review can all change the economics. The worked examples below isolate token charges.
  4. 04
    Keep history separate from current evidenceThe dated chapters preserve earlier comparisons. Current choices use the fresh source dates, access labels, and limitations in this page.

01 — October viewWhat changed in the decision

The current decision starts with availability. Google has announced Gemini 4 Argon, but its initial rollout is through the Fairwind Program for trusted cyber defenders. A forthcoming public API price does not make it a route an ordinary application can deploy today. Keep an announced model in the watch list until the required access, identifier, and service terms exist for your account.

The other easily missed distinction is the provider surface. Our frontier model API price index remains the canonical pricing reference and retains its dated observations. This page independently checks the selected cells below against the vendors on October 4. An older reseller observation in the index is not a fresh direct-provider quote; its source date and buying surface still matter.

Monthly changelog

October 4, 2026 — first edition.

Added the selected provider table, workload criteria, and dated chapter directory. GPT-6.1 Sol's current API page supplies the Sol row; Anthropic's current lineup supplies Fable 5.1, Opus 5.5, Sonnet 5.5, and Haiku 4.5. Argon remains marked restricted. See the October release tracker for the month's dated events.

02 — Buying surfaceThe current API table

Prices below are US dollars per million tokens on the direct provider's stated tier. Read input and output as separate meters, not as a combined price. Cache is the price for reading eligible cached input; it does not establish a hit rate or include writing and storing a cache. Ordinary text rows cannot be compared directly with audio-token rows.

The table keeps an access-limited announcement and the live voice endpoints visible because excluding them would hide useful boundaries. It does not count an announced rate as a purchasable offer. For the ordinary APIs, documented availability means that the provider describes the route; we did not authenticate an account, buy credits, or submit inference requests to establish individual access.

Selected direct-provider records, as of October 4, 2026. USD per 1M tokens. Each model link identifies its specification source; each rate link identifies pricing. Promotional, clock-dependent, voice, and announced rates retain their scope.
Model / API IDInput / outputCache readRate scopeAccess
GPT-6.1 SolOpenAIgpt-6.1-sol$2 / $10$0.10Standard; ≤272K inputDocumented API; account access not tested
GPT-6 AstraOpenAIgpt-6-astra$10 / $50$1Standard; ≤272K inputDocumented API; account access not tested
GPT-6 LunaOpenAIgpt-6-luna$0.10 / $0.50$0.01Standard; ≤272K inputDocumented API; account access not tested
Claude Fable 5.1Anthropicclaude-fable-5-1$10 / $50$0.25Standard Messages APIActive; model-specific safeguards and access terms
Claude Opus 5.5Anthropicclaude-opus-5-5$4 / $20$0.20Standard Messages APIActive (latest)
Claude Sonnet 5.5Anthropicclaude-sonnet-5-5$2 / $10$0.20Standard Messages APIActive (latest)
Claude Haiku 4.5Anthropicclaude-haiku-4-5-20251001$1 / $5$0.10Standard Messages APIActive (latest)
Gemini 3.8 FlashGooglegemini-3.8-flash$0.75 / $3.75$0.075Paid Standard promotion through 2026-12-31Stable API
Gemini 3.5 Flash-LiteGooglegemini-3.5-flash-lite$0.30 / $2.50$0.03Paid StandardStable API
Gemini 3.1 Pro PreviewGooglegemini-3.1-pro-preview$2 / $12$0.20Paid Standard; ≤200K inputPreview API
Gemini 4 ArgonGoogleNo public API ID located$2 / $10 announced$0.10 derivedIntroductory; public API start/end dates unstatedRestricted Fairwind rollout; broader API access forthcoming
Grok 4.7SpaceXAIgrok-4.7$2 / $6$0.50Global Standard; below 200K inputPublic API documented; account access not tested
DeepSeek V4.1 FlashDeepSeekdeepseek-flash$0.15 / $0.60 off-peak; $0.30 / $1.20 peak$0.003 / $0.006Direct API; clock-dependentListed API; older Flash IDs route here
DeepSeek V4 ProDeepSeekdeepseek-v4-pro$0.66 / $1.98 off-peak; $1.32 / $3.96 peak$0.022 / $0.044Direct API; version V4-Pro-0813Listed API; no account request made
Gemini 3.8 LiveGooglegemini-3.8-live$0.75 / $4.50 text; $3 / $12 audioUnsupportedPaid Standard; separate text/audio token metersStable Live API
Gemini 3.8 Live Extended ThinkingGooglegemini-3.8-live-extended-thinking$0.75 / $4.50 text; $3 / $12 audioUnsupportedPaid Standard; separate text/audio token metersStable Live API; async tool calls only

Long prompts change several rows. The OpenAI model pages apply twice the input and cache rates and 1.5 times the output rate to the full request above 272K input tokens. Gemini 3.1 Pro moves to $4 input, $18 output, and $0.40 cached input above 200K. Grok's price table marks its long-context threshold at 200K and lists $4/$12 with $1 cached input; accompanying prose says above 200K, so budget the higher rate at the boundary until your endpoint confirms it.

Promotions and clocks need a separate budget. Google's price page lists Gemini 3.8 Flash at $1.50/$7.50, with $0.15 cached input, from January 1, 2027. DeepSeek's peak windows are 01:00–04:00 and 06:00–10:00 UTC on weekdays except Chinese public holidays; other periods use off-peak rates. Its current pricing page maps older Flash identifiers to V4.1 Flash. Pinning a familiar string does not always pin the underlying model.

Argon's announcement is not a live price test. Google says $2/$10 initially, then $4/$20 after an unstated introductory period. The displayed $0.10 cache figure is our arithmetic: $2 multiplied by the remaining 5% after the announced 95% discount. Neither a public API start date nor the introductory end date was supplied in the captured announcement.

These qualifications come from the linked OpenAI model specification, Gemini price schedule, xAI price schedule, DeepSeek price schedule, and Argon announcement. Residency premiums, tools, batch, priority processing, and negotiated contracts are outside these base rows.

03 — Practical limitsContext, effort, and the exit date

A context window is the space available to a request and its conversation; it is not a promise that every fact inside a long document will be used correctly. An output cap is a ceiling, not a target. Google publishes separate input and output limits, which we label explicitly. Do not add an output allowance to another vendor's context figure and assume the sum is supported.

Knowledge cutoffs also need their labels. Anthropic distinguishes reliable knowledge from the broader training-data cutoff; the Haiku row preserves both. A blank in a model card is not permission to substitute the release date. “Not located” means the captured source set did not establish the field, while “not published” is reserved here for fields absent from Argon's announcement. Neither means unlimited support.

Specifications as of October 4, 2026. Limits are provider-stated tokens, not tested capacity. Effort means the direct API default unless a cell says otherwise. Anthropic retirement floors cover its own operated platforms; Google rows follow its deprecation page.
ModelContext / max outputAPI effortKnowledge cutoffLifecycle evidence
GPT-6.1 Sol1,050,000 / 128,000medium2026-04-30Not located in captured sources
GPT-6 Astra1,050,000 / 128,000Not located; set explicitly2026-04-30Not located in captured sources
GPT-6 Luna1,050,000 / 128,000medium2026-05-18Not located in captured sources
Claude Fable 5.11,000,000 / 128,000highReliable: Jun 2026Not before 2027-09-01
Claude Opus 5.51,000,000 / 128,000mediumReliable: Jun 2026Not before 2027-09-22
Claude Sonnet 5.51,000,000 / 128,000highReliable: Jun 2026Not before 2027-09-28
Claude Haiku 4.5200,000 / 64,000Effort parameter unsupportedReliable: Feb 2025; training: Jul 2025Not before 2026-10-15
Gemini 3.8 Flash1,048,576 input / 65,536 outputmediumNot located in captured sourcesNo shutdown date announced
Gemini 3.5 Flash-Lite1,048,576 input / 65,536 outputminimalNot located in captured sourcesNo shutdown date announced
Gemini 3.1 Pro Preview1,048,576 input / 65,536 outputhighNot located in captured sourcesNo shutdown date announced
Gemini 4 ArgonContext not published / 1,000,000 output announcedNot publishedNot located in captured sourcesNot published
Grok 4.7500,000 / no separate text-output cap statedhighMay 2026Not located in captured sources
DeepSeek V4.1 Flash1M / 384K maximumhigh; thinking onNot located in captured sourcesNot located in captured sources
DeepSeek V4 Pro1M / 384K maximumhigh; thinking onNot located in captured sourcesNot located in captured sources
Gemini 3.8 Live131,072 input / 65,536 outputthinking_level unsupportedNot located in captured sourcesNo shutdown date announced
Gemini 3.8 Live Extended Thinking131,072 input / 65,536 outputlow/medium/high; default not locatedNot located in captured sourcesNo shutdown date announced

Effort and lifecycle cells also use Anthropic's lineup, Anthropic's lifecycle policy, Google's thinking table, Google's deprecation schedule, and DeepSeek's effort guide. OpenAI's captured model pages do not establish a model-specific retirement floor. Grok's guide says there is no separate text-output limit; the context budget still constrains a request.

For procurement, an announced earliest retirement date is a planning floor, not a shutdown appointment. Haiku's October floor deserves attention, but it does not establish that the API will disappear on that date. Keep the alias and retirement ledger beside the context and output census, and recheck the provider before committing a migration deadline.

04 — Task firstRoute by the work you can verify

The routing suggestions below are editorial evaluation designs. They use the documented modalities, limits, and prices above to choose candidates; they do not assert measured superiority. Before testing, define what counts as a correct answer, which failures are unacceptable, and how long a user can wait. Keep the prompt, tools, documents, and review rules consistent across candidates.

Agentic coding. Start with a bounded task: one bug, a reproducible failure, and an agreed test. Compare Sol or Sonnet with Opus, then include Astra or Fable when difficult cases remain unresolved. Judge the final diff, tests, scope control, and review effort. A model that passes a test by deleting the failing assertion has failed the task; a longer reasoning trace does not repair that failure.

Long-document work. Shortlist only routes whose documented limits fit the input and expected output. For document-heavy multimodal work, include Gemini 3.8 Flash or Pro Preview where the preview lifecycle is acceptable; compare a Claude or OpenAI route on the same extracted text when that is the actual application input. Plant questions with known answers and require page-level evidence. Test contradictory passages and missing facts, not just an easy summary at the beginning of a file.

High-volume extraction. Start the comparison with Luna and Flash-Lite, alongside an already approved route if you have one. Require exact field names, valid types, faithful values, and an explicit missing-value policy. Escalate only the records that fail validation or need interpretation. A valid JSON object with the wrong invoice total is still wrong; schema validity alone is an incomplete acceptance test.

Customer-facing chat. Evaluate answer correctness, retrieval citations, latency, tone, and the decision to hand off. Include ambiguous requests and situations where the knowledge base lacks an answer. Choose a richer model when it demonstrably resolves those cases, not because a general benchmark ranks it higher. Keep account changes and other consequential actions behind application checks regardless of model choice.

Voice. Compare an audio-native route with a cascade of transcription, text reasoning, and speech generation if both architectures meet your needs. The two Gemini Live rows have the same published token rates but different reasoning and tool behavior. The Extended Thinking model supports asynchronous tool calls; it needs a client that understands background work. Test interruptions, silence, noisy input, and a tool result arriving after a spoken response.

Coding
Sol or Sonnet versus Opus; add Astra or Fable for unresolved hard cases. Accept only a correct, scoped change.
Diff + tests + review
Documents
Choose compatible input support and limits, then test citation accuracy and contradictory evidence.
Evidence fidelity
Extraction
Luna or Flash-Lite as cost candidates; route failed or ambiguous records to a separately tested fallback.
Field-level accuracy
Chat
Use your knowledge base, response deadline, and handoff rules as the test. Keep business actions separately validated.
Useful, grounded answers
Voice
Test Live and Extended Thinking against an equivalent cascade. Include interruptions and delayed tools.
Conversation completion

Keep an untouched evaluation set for the final decision. If you tune prompts after every failure and score on those same examples, you are measuring how well you adapted to that set. Report the number and type of tasks, the model and surface, effort, token use, failures, and reviewer time. That is enough to make a local decision without pretending that your result is a market-wide leaderboard.

05 — Cost arithmeticWhat an illustrative task costs

Consider an illustrative request with 8,000 input tokens and 2,000 billed output tokens, including any billed reasoning. The uncached estimate is input tokens times the input rate, plus output tokens times the output rate, divided by one million. These are chosen accounting assumptions, not observed tasks. Different tokenizers and reasoning behavior mean that identical text will not necessarily produce identical token counts.

The warm-cache column instead assumes 2,000 uncached input tokens, 6,000 eligible cache-read tokens, and the same 2,000 billed output tokens. It represents a hit on an already established cache. It excludes cache creation, storage, expiry, misses, tools, retries, tax, and human review. Include those items before treating it as an operating budget; the table deliberately isolates the token arithmetic.

Illustrative token charges using October 4, 2026 rates above. Both scenarios stay below the listed long-input thresholds. Gemini 3.8 Flash uses its current promotional rate. No measured quality, latency, or completion-rate claim is implied.
ModelUncached requestAlready-warm cache hit
GPT-6 Luna$0.00180$0.00126
Gemini 3.5 Flash-Lite$0.00740$0.00578
Gemini 3.8 Flash$0.01350$0.00945
GPT-6.1 Sol$0.03600$0.02460
Claude Opus 5.5$0.07200$0.04920
Claude Fable 5.1$0.18000$0.12150
GPT-6 Astra$0.18000$0.12600

Published output-token rates

USD per 1M · selected standard rows · Oct 4, 2026
GPT-6 Luna
$0.50
Gemini 3.5 Flash-Lite
$2.50
Grok 4.7Below the long-context threshold
$6.00
GPT-6.1 SolAt or below 272K input
$10.00
Claude Opus 5.5
$20.00
Claude Fable 5.1
$50.00
GPT-6 AstraAt or below 272K input
$50.00

The chart plots the provider rates recorded above, scaled to the highest displayed rate. It is a price chart, not a performance chart. Sources: OpenAI, Anthropic, Google, and SpaceXAI.

Under the uncached assumptions, moving from Sol to Opus adds $0.036 per attempt; moving from Opus to Fable adds $0.108. Those are useful price gaps when deciding what to test. They do not tell you whether the second attempt will succeed, or whether a more capable first attempt would have avoided the retry. A fallback that repeats the entire prompt creates another bill.

For a real batch, divide all attempt charges plus review cost by the number of accepted outcomes. Keep rejected work in the numerator. Otherwise a route that cheaply produces unusable answers looks artificially efficient. If no output is accepted, there is no meaningful cost per accepted task to report. Voice needs its own accounting because audio and text use different meters and a conversation contains more than one turn.

06 — Control knobsSet effort before you compare

Effort labels are provider controls, not a shared unit of thought. “High” on one model need not spend the same tokens, take the same time, or achieve the same result as “high” elsewhere. Anthropic's effort guide describes a behavioral signal rather than a strict token budget. Its current API default is medium for Opus 5.5 and high for Sonnet 5.5 and Fable 5.1. Comparing omitted settings is already comparing different choices.

OpenAI's model pages list medium for Sol and Luna. We did not locate an Astra default in the captured model page, so the table does not guess one. The API's documented values also need not match a Codex or ChatGPT model picker. A subscription grants usage under that product's rules; it is not evidence of the direct API's identifier, price, or available effort values.

Google's thinking table gives Flash 3.8 a medium default, Flash-Lite 3.5 minimal, and Pro Preview high. The ordinary Live model does not accept the thinking-level control. Its Extended Thinking sibling does, but that is a different endpoint with different client behavior. Use the effort-ladder reference for surface-specific historical context, then verify the current surface before copying a setting.

A useful evaluation varies one factor at a time: keep the task and tools fixed, compare supported effort levels on one model, then compare the best acceptable configuration with another model. Log billed output and completion time rather than assuming that lowering the label always saves the same percentage. Extra reasoning may help on some tasks and merely lengthen others.

Fallback boundary

A fallback needs its own acceptance test.

Check that the destination accepts your message structure, tools, files, and reasoning state. Do not assume provider-specific thinking blocks transfer. Validate a neutral task summary and the necessary evidence before switching. A policy refusal needs an allowed resolution or human handoff; a more permissive route is not an automatic retry policy.

07 — Dated chaptersThe comparisons behind the current view

These 42 chapters retain their original publication dates and URLs. They span model generations, deployment choices, assistants, and related commercial comparisons; the directory is the selected historical series, not a claim that every row is a like-for-like model benchmark. The short descriptions identify the subject without certifying old prices, availability, benchmark populations, or conclusions.

Use a chapter to understand the question people were asking at the time. Use the current vendor links above for a purchase or implementation decision. The August tracker, September tracker, and October tracker preserve the surrounding release chronology.

2025 and early 2026

Historical chapter directory, checked October 4, 2026. Dates are original publication dates; descriptions identify topics, not verified current findings.
PublishedChapterWhat it covered
2025-08-01ChatGPT vs Claude vs Gemini vs Grok AI ComparisonConsumer assistants across ChatGPT, Claude, Gemini, and Grok.
2025-10-03Claude Sonnet 4.5 vs GPT-5 Pro: Complete 2025 ComparisonClaude Sonnet 4.5 and GPT-5 Pro.
2025-10-07Zhipu GLM 4.6 vs Claude Sonnet: Coding Model GuideGLM 4.6 and Claude Sonnet.
2026-01-06DeepSeek R1 vs Qwen 3 vs Mistral Large: LLM ComparisonDeepSeek R1, Qwen 3, and Mistral Large.
2026-01-11Claude Code vs Aider vs Gemini CLI: AI CLI ComparisonTerminal coding tools: Claude Code, Aider, and Gemini CLI.
2026-02-07OpenClaw vs ChatGPT vs Siri: AI Assistants ComparedAssistant workflows across OpenClaw, ChatGPT, and Siri.
2026-03-05GPT-5.4 vs Opus 4.6 vs Gemini 3.1 Pro: Best AI Model?GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro.
2026-03-27GLM-5.1 and Claude Opus: the dated coding comparisonGLM 5.1 coding claims beside Claude Opus.
2026-03-30Qwen 3.5-Omni vs Gemini 3.1 vs GPT-5.4 ComparisonQwen 3.5 Omni, Gemini 3.1, and GPT-5.4.

April–June 2026

Historical chapter directory, checked October 4, 2026. Dates are original publication dates; descriptions identify topics, not verified current findings.
PublishedChapterWhat it covered
2026-04-02Gemma 4 vs Llama 4 vs Mistral Small 4: Full ComparisonOpen-weight families: Gemma 4, Llama 4, and Mistral Small 4.
2026-04-02Qwen 3.6 Plus vs Claude Opus 4.6 vs GPT-5.4 ComparedQwen 3.6 Plus, Claude Opus 4.6, and GPT-5.4.
2026-04-12Anthropic's Cost Problem: Opus Spend vs Xiaomi VolumeAnthropic spending and Xiaomi usage as different measures.
2026-04-16Claude Opus 4.7 vs GPT-5.4: Agentic Coding ComparedOpus 4.7 and GPT-5.4 for coding agents.
2026-04-23GPT-5.5 vs Claude Opus 4.7: Benchmarks & PricingGPT-5.5 and Claude Opus 4.7.
2026-04-24MoE Architecture: GPT, Claude, DeepSeek, Qwen ComparedMixture-of-experts architecture across model families.
2026-05-19Gemini 3.5 Flash vs GPT-5.5 vs Opus 4.7: Agentic CodingGemini 3.5 Flash, GPT-5.5, and Opus 4.7.
2026-05-21Claude Ad-Free Pledge: Anthropic vs OpenAI Ads in 2026Anthropic and OpenAI advertising approaches.
2026-05-28Claude Opus 4.8 vs Gemini 3.5 Flash: AI Agent RoutingOpus 4.8 and Gemini 3.5 Flash routing.
2026-05-28Claude Opus 4.8 vs GPT-5.5: Benchmarks & Cost ComparedOpus 4.8 and GPT-5.5.
2026-06-03MiniMax M3 vs Opus 4.8 vs GPT-5.5: Coding ShowdownMiniMax M3, Opus 4.8, and GPT-5.5.
2026-06-09Claude Fable 5 vs GPT-5.5: Benchmarks & Cost ComparedClaude Fable 5 and GPT-5.5.
2026-06-16GLM-5.2 Benchmarks: Open Weights vs Claude Opus 4.8GLM 5.2 and Claude Opus benchmark and access claims.

July 2026

Historical chapter directory, checked October 4, 2026. Dates are original publication dates; descriptions identify topics, not verified current findings.
PublishedChapterWhat it covered
2026-07-01Sonnet 5 vs Opus 4.8 vs Fable 5: Which to Use WhenSonnet 5, Opus 4.8, and Fable 5 workload selection.
2026-07-02Claude Code Routines vs n8n and Zapier: Real CostsClaude Code Routines, n8n, and Zapier automation costs.
2026-07-02GPT-5.6 Sol vs Fable 5: Price, Access and BenchmarksGPT-5.6 Sol and Claude Fable 5 price and access.
2026-07-03GLM-5.2 API Access Compared: Z.ai vs OpenRouter vs HostsGLM 5.2 provider access and pricing.
2026-07-08Grok 4.5 vs Opus 4.8 vs GPT-5.5: Which Model Wins?Grok 4.5, Opus 4.8, and GPT-5.5.
2026-07-09Muse Spark 1.1 vs Grok 4.5: Which Agent Model Wins?Muse Spark 1.1 and Grok 4.5.
2026-07-17Kimi K3 vs Claude Fable 5: Frontier Comparison 2026Kimi K3 and Claude Fable 5.
2026-07-17Kimi K3 vs GPT-5.6 Sol: Agentic AI Comparison 2026Kimi K3 and GPT-5.6 Sol.
2026-07-22Gemini 3.6 Flash Benchmarks: The Per-Task Price Shake-UpGemini 3.6 Flash, GPT-5.6, Sonnet 5, and Kimi K3.
2026-07-24FLUX 3 vs Seedance, Omni, and Kling: The Video FieldFLUX 3, Seedance 2.5, and Gemini Omni video models.

August 2026

Historical chapter directory, checked October 4, 2026. Dates are original publication dates; descriptions identify topics, not verified current findings.
PublishedChapterWhat it covered
2026-08-12Grok 4.6 vs Sol vs Opus 5 vs Fable 5: Read the TiersGrok 4.6, GPT-5.6 Sol, Opus 5, and Fable 5 effort.
2026-08-12Grok Build vs Claude Code: What Each Can Actually SeeGrok Build and Claude Code verification workflows.
2026-08-14Gemini 3.7 Flash vs Sonnet 5 vs GPT-5.6 Terra: Real WinsGemini 3.7 Flash, Sonnet 5, and GPT-5.6 Terra.
2026-08-26The Weights and the API Are Not the Same Qwen ModelQwen3.8-Flash-Next open weights and hosted API identity.

September 2026

Historical chapter directory, checked October 4, 2026. Dates are original publication dates; descriptions identify topics, not verified current findings.
PublishedChapterWhat it covered
2026-09-08GPT-6 Astra vs Claude Fable 5.1: Which Model Fits BestGPT-6 Astra and Claude Fable 5.1.
2026-09-22Claude Opus 5.5 vs GPT-6 Astra: Benchmarks, Price and FitClaude Opus 5.5 and GPT-6 Astra.
2026-09-22GPT-6 Sol vs Claude Opus 5.5: Cost per Task and BenchmarksGPT-6 Sol and Claude Opus 5.5.
2026-09-22Opus 5.5 vs Grok 4.7 vs Muse Spark 1.3: Real Cost per TaskOpus 5.5, Grok 4.7, and Muse Spark 1.3 task costs.
2026-09-30OpenAI vs Claude Marketplace: Using AI Spend CommitmentsOpenAI and Claude marketplace spending commitments.
2026-09-30Personal AI Agents Compared: Dots, Muse, Cue and GrokPersonal agents including Dots, Muse, Cue, and Grok Team Bots.

08 — Collection methodWhat the reference can establish

This is a provider-document census for a selected routing set. It establishes what the cited sources stated when retrieved, not whether a model will satisfy a task or whether your organization will receive access. The practical value is in keeping the rate, endpoint, limit, date, and missing field together. A table loses that value when a number is copied without the condition beside it.

Methodology

Provider-published specifications and prices, checked against complete captured pages; routing criteria and worked costs are explicitly separate.

As-of
October 4, 2026. All current model and pricing sources were retrieved that day. Publication is October 4; historical chapter dates come from their existing metadata.
Population
Sixteen selected records across five providers: the current OpenAI and Claude candidate families, selected Gemini text and voice endpoints, Grok 4.7, DeepSeek's two listed routes, and restricted Argon. This is not every frontier model.
Collection
Read provider model pages, pricing tables, effort documentation, and lifecycle pages. Preserve URL, retrieval time, full response, and SHA-256. Compare identifiers and rate conditions before transcribing; distinguish vendor list schedules from reseller observations.
Units
USD per million tokens. Prices use the indicated direct API tier. Google input limits are labelled separately; other context figures follow the provider. Knowledge cutoffs retain the provider's definition.
Missing fields
Not located means this source set did not establish the value. Not published refers to absent Argon announcement details. Unsupported describes an explicit capability limit. No shutdown announced is not a perpetual availability guarantee.
Arithmetic
Task scenarios use fixed assumed token counts. Cache-hit examples exclude cache creation and storage. The rate chart scales each stated output price against $50. Argon's cache rate is derived from its announced discount.
Exclusions
No inference runs, benchmark results, provider account checks, negotiated rates, or subscription entitlements were measured. Open-weight hosting, image/video generation, and other providers fall outside this selected set. Historical titles are inventory evidence only.
Limitations
Provider pages can change or disagree with account-specific terms. Tokenizers, hidden reasoning, prompts, and harnesses differ. Equal token assumptions do not establish equal work, quality, latency, or cost per accepted outcome.
Refresh
The monthly refresh batch maintains this URL: review within a week of a frontier launch and at least monthly, update changed rows and modified time, and add a dated changelog entry. This is an editorial maintenance commitment.

09 — Your next testTurn the table into a decision

Write a short selection record before sending a production workload: the task, acceptable errors, required modality, data terms, exact endpoint, effort, expected token mix, latency requirement, and fallback behavior. Pick a current candidate and a plausible challenger. The table narrows that choice; your evaluation settles it.

Keep the decision reversible. Save the prompt and tool configuration, record the returned model version where the provider exposes it, and retain a small regression set. Re-run the relevant cases when an alias moves, a model changes, or your workload expands. Rechecking only the price page will miss a change that makes an earlier success criterion fail.

When a route passes, set a budget from actual usage rather than the illustrative rows. When it fails, name the failure: missing evidence, invalid action, wrong field, unacceptable latency, or excessive review. That tells you whether to change the model, the prompt, the retrieval, or the surrounding application. Buying more reasoning cannot fix every input or workflow defect.

For related implementation scope, see our AI transformation service.

Make it reproducible

Choose an accepted outcome, then price the route

A sound model choice has a concrete task, a documented endpoint, a testable standard, and an explainable bill. Keep those four together and the next launch becomes a candidate to evaluate, rather than a reason to rebuild the whole workflow.

Applying the decision

Build a model evaluation around your workload.

Translate business requirements into acceptance tests, model choices, and an operating budget.

Workload definitionEvaluation designCost measurement
Decision record

What to specify

  • →The task and acceptable outcome
  • →The exact API and effort setting
  • →The complete attempt cost
  • →The fallback and review boundary
Practical questions

Choosing a model without guessing.

Compare GPT-6.1 Sol or Claude Sonnet 5.5 with Opus 5.5 on representative changes and your own tests. Add Astra or Fable for unresolved difficult cases. These are evaluation candidates, not measured winners in this guide.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

More model reference work.

AI Development

Switch Off Claude Fable 5.1 Mid-Task and It Forgets Why

Claude Fable 5.1 reasoning cannot be read by Opus 5 or Sonnet 5. Any router, retry or refusal fallback that moves a task down loses it silently. What to change.

September 1, 2026 · 12 minRead
AI Development

AI API Pricing, August 2026: Cuts, Promos, and Traps

GPT-5.6 Luna fell 80% and Terra 20% on July 30. Sonnet 5's scheduled $3/$15 increase was cancelled. OpenRouter list prices mix standard with batch rates.

August 5, 2026 · 18 minRead
AI Development

Claude Opus 4.8 vs Gemini 3.5 Flash: AI Agent Routing

Gemini 3.5 Flash beats Claude Opus 4.8 on MCP-Atlas and Finance Agent at a third of the price — but a 61% hallucination rate complicates the routing call.

May 28, 2026 · 15 minRead
AI Development

Why an AI’s Reasoning Can’t Follow You to Another Model

Anthropic, OpenAI and Google now bind a model’s reasoning to the model that produced it. What each locks, what breaks on a switch, and how a router copes.

September 2, 2026 · 7 minRead
AI Development

Deleting AI Agent Memory: Where Stored Copies Survive

Deleting AI agent memory takes more than clearing a chat. Map stored copies, retrieval indexes and backups, then verify what your system can still recover.

September 4, 2026 · 7 minRead
AI Development

Preview, Beta, GA: What Vendors Said vs What Coverage Said

Thirty-six AI vendor announcements from 17-22 August 2026, each scored on the vendor's own status word against the word its coverage used, where located.

August 22, 2026 · 27 minRead
Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source