Start with Fugu Max when a task already has a clear acceptance test. Pay for Ultra v2 when a harder task repeatedly fails that test and the additional reasoning reduces the total cost of reaching an accepted result. A higher token price can be worthwhile, but the model name cannot establish that trade-off for you.
Sakana announced both variants on September 11, 2026. Prices and documentation below were checked September 12. This is a buying framework with illustrative calculations; we did not run a head-to-head benchmark.
- 01Compare accepted work.Include retries, review and elapsed time alongside token charges.
- 02Watch the context boundary.Ultra v2 has higher rates above 272K context; Max lists fixed rates.
- 03Treat rankings as a hypothesis.Sakana’s launch results identify tasks to test, not a guaranteed winner for your work.
01 — Practical decisionWhat the new variants actually change
Sakana’s announcement presents Max and Ultra v2 as two objectives within its orchestration architecture: cost efficiency and higher capability on difficult tasks. The company says both are immediately available through its OpenAI-compatible API. These are hosted system choices; using open models inside a service does not make the complete service an open-weight download.
The distinction matters if you buy complete agent work rather than isolated answers. An orchestration system can select and coordinate underlying models, but you still need to supply the right tools, data and acceptance criteria. Our earlier Fugu introduction covers that architecture. This comparison asks which new price tier deserves a place in a workflow.
Do not confuse the Fugu Max model with a subscription carrying the word Max. A model rate and a monthly account allowance answer different questions. Compare the actual API variant on the invoice and keep subscription entitlement outside the token arithmetic below.
02 — Practical decisionCompare the rate card before comparing results
The current Fugu pricing page lists the following US-dollar rates per million tokens. Cached input is a separate billing category, not an automatic discount for all repeated text. Record the cache usage the service actually reports.
Ultra v2’s long-context condition is expressed as context greater than 272K. The page does not fully explain every boundary-accounting case. Before pricing a request close to that threshold, confirm what the service counts and how it applies the rate. The table faithfully records the published categories; it is not a reconstructed billing specification.
The same page lists Max web_search and web_fetch at $0.007 per call. Add these when used. Do not apply the ordinary configurable Fugu pool’s blended-price explanation to these fixed-rate variants, or assume that all downstream tool spending is included.
| Billing category | Fugu Max | Ultra v2 |
|---|---|---|
| Uncached input / 1M | $2 | $5 |
| Output / 1M | $6 | $30 |
| Cached input / 1M | $0.25 | $0.50 |
| Input / 1M above 272K context | $2 | $10 |
| Output / 1M above 272K context | $6 | $45 |
| Cached input / 1M above 272K context | $0.25 | $1.00 |
03 — Practical decisionPut a realistic task budget around those rates
For an illustrative request with 100,000 uncached input tokens and 10,000 output tokens, Max costs $0.26: 0.1 × $2 plus 0.01 × $6. Ultra v2 costs $0.80 at its ordinary rates: 0.1 × $5 plus 0.01 × $30. That is about 3.08 times the token cost for the same billed token counts. It is not a measured difference in the cost of completing a task.
For an illustrative 300,000-input, 10,000-output request, Max costs $0.66. Ultra v2 costs $3.45 if its published long-context rates apply to the full request. This assumption is explicit because threshold billing needs confirmation. Neither example includes cache reads, tools, failed attempts or human review.
Equal token counts are useful for explaining a rate card. They are a poor forecast when the systems produce different reasoning lengths, call different tools or require different numbers of retries. Use actual usage records for a pilot and retain failed attempts in the total. The review-cost guide explains why a cheap draft can still be expensive to accept.
04 — Practical decisionUse the vendor benchmarks to choose test cases
Sakana reports stronger benchmark performance for Ultra v2 and positions Max around cost efficiency. Those are vendor-reported results. This article does not independently validate the benchmark runs, their task distribution or their relevance to your repository. We therefore do not turn the announcement into a universal ranking.
Choose examples that resemble the decision you face: a multi-file repair with regression checks, a research answer with conflicting evidence, or a structured document requiring several connected deductions. Also include routine cases. A test set consisting only of exceptionally hard tasks cannot tell you whether to upgrade everyday traffic.
Keep the brief, tools and acceptance rules comparable. Preserve each system’s actual settings rather than pretending differently named effort controls are equivalent. The benchmark interpretation guide provides the distinction between a published score and evidence from your own task distribution.
05 — Practical decisionEscalate a task for a reason you can record
A useful escalation rule names the failure that more reasoning is expected to address. For example: the cheaper run has the correct files and permissions but cannot produce a patch that passes the agreed regression case. Record that condition before asking Ultra v2 to continue or retry.
Escalation is less promising when the missing ingredient is a credential, a source document or a decision only the user can make. More reasoning cannot establish an unavailable fact. Fix the information gap first. Otherwise the expensive run may simply produce a more elaborate version of the same unsupported answer.
Treat the fallback as another attempt in the same task budget. Count the original run, the escalation, repeated tool work and the final review. If you pass a prior draft to Ultra v2, label that as an assisted escalation test. It is different from independently giving both variants the same starting brief.
06 — Practical decisionAdopt the smallest change supported by the pilot
Before starting, write down the task categories, what counts as acceptance and who will review the work. Afterward, report accepted and rejected cases, total charge, waiting time and review effort separately. Averages can hide a category that becomes less reliable even when the overall bill improves.
If Ultra v2 helps only on complex structured reasoning, route that category to it and keep routine work on Max. If results are mixed, keep the selection reversible and collect more representative cases. A small pilot supports a local routing decision, not a broad claim that one system is best.
An AI transformation project should turn the model choice into a measurable acceptance decision. Keep the dated rates with the pilot record so a later price change can be evaluated without confusing it with a change in quality.
Evidence and scope
- As-of date
- September 12, 2026: sources retrieved and reviewed. September 12 is the editorial allocation. Verified event dates are stated separately.
- Sources and method
- Two Sakana primary pages checked September 12, 2026. Published rate categories transcribed and illustrative arithmetic recalculated.
- Limits
- No model runs or independent benchmark reproduction. Boundary accounting and actual account terms must be confirmed for a purchasing decision.
07 — Next stepBuy additional reasoning where it changes acceptance
Buy additional reasoning where it changes acceptance
Max provides the lower published token rates. Ultra v2 earns its additional cost only when the completed-work evidence supports it. Keep long-context billing, vendor benchmarks and your own acceptance results separate, then route the tasks for which the difference matters.