DevelopmentFramework16 min readPublished August 2, 2026

Converged enough for most work · not all of it · the brief carries the rest

The Brief Is the Product: Specification Quality for AI

Frontier models now sit within a few points of each other on most general benchmarks — and roughly 20x apart on a few specific ones. That means the brief you write does more work than the model you pick for most tasks, with named exceptions this essay spells out rather than hides. Here is the anatomy of a specification that reliably raises output quality, and the honest boundary where model choice still dominates.

DA
Digital Applied Team
Senior strategists · Published Aug 2, 2026
PublishedAug 2, 2026
Read time16 min
Sources14
AA Index spread, top 3
3pts
three labs — a pre-Opus 5 snapshot
Jul 17, 2026
ARC-AGI-3 spread
~20x
same Jul 24 launch table
30.2% vs 1.5%
PR-outcome study
265
validated developer-AI interactions
Spec Kit stars
~125K
125,000+ stars at the time of writing

Specification quality — the precision of the brief you hand an AI model — has quietly become the biggest lever most teams control over AI output quality. The top three models on the Artificial Analysis Intelligence Index came from three different labs and spanned just 3 points in mid-July 2026 — a snapshot taken before Claude Opus 5 shipped on July 24 and reshuffled the order at the top. When the models cluster that tightly, the variable that separates a usable result from a mediocre one is increasingly the specification, not the model picker.

But that claim only partially holds, and this essay is honest about the boundary. The same July 2026 launch table that shows four current models within 6.5 points of each other on one benchmark shows a roughly 20x spread between two of them on another. The capability gap between frontier models did not vanish — it moved, from general-purpose work into specific task families where model choice still dominates everything you write in the prompt.

This guide covers both halves: the evidence that models have converged on most general work, the named exceptions where they emphatically have not, and — the practical core — the anatomy of a brief that reliably raises output quality: hard constraints, acceptance criteria, negative space, a worked example, and a self-check rubric. It is the anchor essay for a set of concrete method posts we publish alongside it.

Key takeaways
  1. 01
    Frontier models converged on most general benchmarks.The Artificial Analysis Intelligence Index's top three models — from three different labs — spanned 3 points in mid-July 2026, and six labs each fielded a model above 50. On BrowseComp, four current models sit within 6.5 points.
  2. 02
    The gap moved; it didn't vanish.On Anthropic's own July 24, 2026 launch table, ARC-AGI-3 shows a roughly 20x spread (30.2% vs 1.5% of tasks solved) and a held-out legal benchmark a roughly 5x spread. Abstract reasoning and niche domains still reward model choice.
  3. 03
    Brief quality is measurable leverage.An empirical arXiv study of 265 validated developer-AI interactions from real pull requests found specificity and context most strongly correlated with actionable code — and verification cues were the strongest predictor of whether code was actually adopted.
  4. 04
    A working brief has five parts.Hard constraints, measurable acceptance criteria, negative space (what not to do), a worked example, and a self-check rubric. Each has a sourced failure mode when missing — and real caveats, like rubric self-preference bias.
  5. 05
    Spec-driven development went industry-wide.GitHub's open-source Spec Kit had passed 125,000 stars with support for 30+ AI coding agents at the time of writing, and multiple vendors have independently shipped spec-driven modes. The industry is converging on the same conclusion: write the spec first.

01ConvergenceThe convergence that actually happened.

Start with the strongest independent evidence. Artificial Analysis reported on July 17, 2026 that four frontier launches had landed in eight days, and that six labs — Anthropic, OpenAI, Moonshot AI, SpaceXAI, Z.ai, and Meta — now each field a model scoring above 50 on its Intelligence Index. The top three models (Claude Fable 5 at 60, GPT-5.6 Sol at 59, Kimi K3 at 57) came from three different labs and spanned just 3 points. In the same eight-day window, the leader's margin over second place compressed from 4 points to 1. That snapshot predates Claude Opus 5's July 24 launch, so the exact top-three line-up has since moved — the clustering is the durable finding, not the ordering.

Vendor tables tell the same story on general-purpose agentic work. Anthropic's Opus 5 launch table (July 24, 2026) — notable because it includes a competitor's model in the same comparison — shows four current-generation models within 6.5 points on BrowseComp: Opus 5 solving 90.8% of its tasks, GPT-5.6 Sol 90.4%, Fable 5 87.4%, and Opus 4.8 84.3%. And on GPQA Diamond, the graduate-level science QA benchmark, aggregated leaderboards show most frontier models bunched roughly between 92% and 95% of questions answered correctly, against a human-PhD-expert baseline of 69.7% — the benchmark itself is close to saturated at the frontier, so exact rankings shuffle by aggregator and snapshot.

Index spread
Top 3 models, 3 labs
3pts

Claude Fable 5 (60), GPT-5.6 Sol (59), and Kimi K3 (57) topped the Artificial Analysis Intelligence Index in mid-July 2026 — three labs separated by three points on the composite.

AA Index · Jul 17, 2026
Labs above 50
A crowded frontier
6

Anthropic, OpenAI, Moonshot AI, SpaceXAI, Z.ai, and Meta each fielded a model above 50 on the same index — the leader's margin over #2 narrowed from 4 points to 1 in eight days.

Four launches in 8 days
BrowseComp spread
Four models, one table
6.5pts

On Anthropic's own July 24 launch table, four current models solve between 84.3% and 90.8% of BrowseComp tasks — single-digit territory on general agentic web research.

Cross-lab vendor table

The pattern also holds within model families. Thinking Machines' own table for Inkling-Small (July 30, 2026) reports its 12B-active-parameter model resolving 80.2% of SWE-bench Verified issues versus 77.6% for the larger 41B-active flagship Inkling — within 3 points, on roughly a third of the active parameters. When a model a third the size lands within the noise band of its own flagship on a major coding benchmark, "pick the biggest model" stops being a strategy for work of that shape. What differentiates the output is what you asked for, and how precisely.

02The ConcessionWhere model choice still decides the outcome.

Here is the section most "the model doesn't matter anymore" content leaves out — and the reason this essay argues partial convergence, not blanket convergence. The very same Anthropic launch table that shows a 6.5-point BrowseComp spread shows a roughly 20x spread on ARC-AGI-3, the abstract visual-reasoning benchmark: Opus 5 solved 30.2% of its tasks, GPT-5.6 Sol 7.8%, and Opus 4.8 just 1.5%. On a held-out Legal Agent Benchmark in the same table, Fable 5 completed 13.3% of tasks against GPT-5.6 Sol's 2.5% — a roughly 5x gap on a specialised domain, between two models that sit within single digits of each other on BrowseComp.

The divergence even shows up within a single vendor's lineup. Thinking Machines' own table has Inkling-Small at near-parity with its larger sibling on general coding — then notably worse on the specialised τ³-Banking agentic task, completing 15.5% of tasks versus Inkling's 23.7%. And no lab sweeps its own comparisons: OpenAI's GPT-5.6 GA tables (July 9, 2026) show Claude models leading several of OpenAI's own metrics — Claude Mythos 5 resolving 80.3% of SWE-Bench Pro tasks and Fable 5 80% against Sol's 64.6%, and Fable 5 ahead on GDPval Elo (1,759.6 vs 1,747.8) — while Sol leads others, including Terminal-Bench 2.1 at 88.8 and OSWorld 2.0 at 62.6. The honest read of the cross-vendor evidence is "no universal leader," which is a different claim from "model choice doesn't matter."

We assembled the spread table below from three primary vendor tables — Anthropic's July 24 Opus 5 launch table, OpenAI's July 9 GPT-5.6 GA table, and Thinking Machines' July 30 Inkling-Small table. To our knowledge no other coverage has put these spreads side by side; most picks one benchmark and declares convergence or divergence accordingly. The point of the table is that both are true, on different rows.

The convergence spread: benchmarks where current frontier models sit within a few points of each other versus benchmarks where the spread is 5x to 20x, compiled from Anthropic, OpenAI, and Thinking Machines launch tables published in July 2026.
BenchmarkCompared modelsSpreadVerdict
Near parity — brief quality likely dominates
AA Intelligence Index (composite)Top 3 models, 3 labs · Jul 173 pts (60 vs 57)Near parity
BrowseComp (agentic web research)4 models, Anthropic table · Jul 246.5 pts (90.8% vs 84.3%)Near parity
GPQA Diamond (graduate science QA)Frontier cluster, aggregators~3 pts — most bunched ~92–95%; rankings vary by snapshotNear saturation
SWE-bench Verified (agentic coding)Inkling-Small vs Inkling (same vendor) · Jul 302.6 pts (80.2% vs 77.6%)Near parity at ~⅓ the size
Wide gap — model choice still dominates
ARC-AGI-3 (abstract visual reasoning)3 models, Anthropic table · Jul 24~20x (30.2% vs 1.5%)Model choice dominates
Legal Agent Benchmark (held-out)Fable 5 vs GPT-5.6 Sol, same table · Jul 24~5x (13.3% vs 2.5%)Model choice dominates
τ³-Banking (specialised agentic)Inkling-Small vs Inkling (same vendor) · Jul 308.2 pts (15.5% vs 23.7%)Larger sibling wins
SWE-Bench Pro (OpenAI's own GA table)Claude models vs Sol · Jul 915.7 pts (80.3% vs 64.6%)No universal leader

The pattern is not new, either. On the earlier ARC-AGI-2 benchmark — figures that are aggregator-sourced and reflect the pre-Opus 5 model generation, so treat the decimals as approximate — Google's Gemini 3 Deep Think (around 85% of tasks, February 2026) and OpenAI's GPT-5.4 Pro (around 83%) sat well ahead of Anthropic's then-flagship Opus 4.6 (roughly 69%). That ordering did not hold: on July 24, 2026, ARC Prize independently administered ARC-AGI-2's 120-task Public Eval and recorded Opus 5 at 90.4% of those tasks solved at Max reasoning effort, ahead of every pre-Opus-5 figure above. Read together, the two snapshots make a structural point, not a directional one: on abstract reasoning the spreads stay wide and the leader changes hands, so the model you pick still decides the outcome. The practical rule that falls out of all this: for general coding, research, writing, and agentic web work, the models have converged enough that your brief is the bigger lever. For abstract reasoning, niche professional domains, and specialised agentic tasks, run your own eval before believing any convergence story — including this one.

03Methodology CautionVendor tables vs independent reruns.

Before building anything on benchmark numbers — including the table above — one live example of why single tables deserve suspicion. Thinking Machines' own launch table for Inkling-Small reported it edging out its larger sibling on Terminal-Bench 2.1, completing 64.7% of tasks versus Inkling's 63.8%. Artificial Analysis's independent rerun of the same comparison scored both models at 55% — a tie, at a meaningfully lower level. Same two models, same benchmark, opposite conclusions about which is better, depending on who ran the harness.

That is not an accusation of bad faith; harness configuration, scaffolding, and effort settings legitimately move agentic benchmark scores. But it is a reason to treat every vendor table — and every aggregator snapshot — as one measurement, not ground truth. It is also, conveniently, an argument for this essay's thesis: if the same model pair can score 64.7 or 55 depending on how the evaluation was specified, then specification quality is visibly a first-order variable in the very benchmarks people use to argue about models.

Read every table this way
Vendor benchmark tables and independent reruns can disagree even on the same comparison — Inkling-Small vs Inkling went from a vendor-reported win (64.7% vs 63.8% of Terminal-Bench 2.1 tasks) to an independent tie at 55%. Every number in this post carries its source and date for exactly that reason. Treat our spread table as three vendors' measurements compiled honestly — not as settled physics.

04The EvidenceWhy the brief does the work.

The claim that specification quality drives output quality is not just craft folklore — it has an unusually good empirical source. An arXiv study of real open-source pull requests (a June 2026 preprint) manually validated 265 developer-AI interactions and traced which prompt characteristics predicted which outcomes. Specificity and context correlated most strongly with actionable code generation. But the strongest predictor of whether generated code was actually adopted was something else: verification — giving the model a way to check its own work. In other words, the part of the brief most teams skip (how will we know this is right?) is the part most associated with output that ships.

The vendors' own guidance points the same direction. Anthropic's prompting documentation says a few well-crafted input/output examples improve accuracy and consistency, and recommends including 3–5 of them for best results — wrapped in explicit tags and diverse enough to cover the edge cases you care about. And OpenAI's updated GPT-5.6 Sol guidance, as reported in July 2026, claims that in OpenAI's own internal coding-agent testing — no independent methodology has been published, so treat it as OpenAI's claim — leaner, more specific system prompts improved evaluation scores by roughly 10–15% while cutting total tokens by 41–66% and cost by 33–67% relative to the heavier prompts they replaced. The same guidance makes a subtler point: conflicting instructions are worse than missing ones, because contradictory rules force the model to burn reasoning tokens trying to reconcile both.

Strongest correlate
Specificity + context
265 validated real-PR interactions

In the PR-outcome study, prompt specificity and contextual grounding correlated most strongly with actionable code generation — vague asks produced output that needed reworking before it was usable.

arXiv 2606.19644
Strongest predictor
Verification cues
Predicts adoption, not just generation

Giving the model a way to check its own work was the strongest predictor of whether generated code was actually adopted. Evaluability is what separates output that ships from output that stalls.

The most-skipped brief section
Vendor-reported
Leaner, specific prompts
OpenAI internal coding-agent tests — unaudited

OpenAI reports leaner, more specific system prompts improved its internal coding-agent eval scores ~10–15% while cutting tokens 41–66% and cost 33–67%. Its guidance also warns that conflicting rules are worse than missing ones.

Treat as OpenAI's own claim

One more evidence class deserves a careful framing: worked examples. Published few-shot-prompting studies report gains that are real but wildly task-dependent — one 2025 clinical-text benchmark found few-shot prompting beat zero-shot in 95.8% of tested model configurations, with two-thirds of those gains exceeding 20 percentage points, while a sensor-data classification study found a more modest 16.6% relative gain from optimized example selection versus random selection. There is no honest single number for "how much a good brief improves output" — the defensible claim is that the direction is consistent and the magnitude depends heavily on task, model, and baseline. Anyone selling you one universal percentage is compressing away the most important variable.

05The FrameworkAnatomy of a brief that raises output quality.

Pulling the research together, a working brief has five components. Each row below names what the component specifies, the sourced failure mode when it is missing, and the one question to ask before you hit send. This is the checklist we run on our own briefs before handing work to a model — and it is deliberately model-agnostic, because per the convergence evidence above, for most work the brief travels better than the model pick does.

Anatomy of a brief: five components — hard constraints, acceptance criteria, negative space, worked example, and self-check rubric — with what each specifies, the failure mode when missing, and a self-check question.
ComponentWhat it specifiesWhen it's missingAsk before you send
Hard constraintsStack, scope boundaries, formats, non-negotiables — the walls of the problemThe model guesses at unstated requirements — GitHub's spec-kit team calls this out as the core failure of vague promptsCould two reasonable engineers read this and build different things?
Acceptance criteriaMeasurable pass/fail conditions — Anthropic's own worked example of a good one: less than 0.1% of outputs across 10,000 trials flagged for toxicity by a content filter (the bad version it contrasts this with is simply "safe outputs")Output can't be evaluated — and verification cues were the strongest predictor of adoption in the 265-PR studyHow will the model — or I — know this is done?
Negative spaceKnown failure modes, out-of-scope moves, what not to touchThe model repeats the failure you didn't name — but this one is not a free win: badly specified negative constraints can suppress valid edge cases, and one forecasting study found two structured prompt patterns decreased accuracyAm I naming real failure modes, or piling on rules that contradict each other?
Worked example3–5 input/output pairs showing the target shape, diverse enough to cover edge casesFormat drift and unhandled edge cases — Anthropic's docs say a few well-crafted examples improve accuracy and consistency, and recommend 3–5 for best resultsDoes my example set cover the ugliest edge case I actually care about?
Self-check rubricCriteria the model iterates its own draft against before returning — OpenAI's cookbook recommends building the rubric first, then checking the output against itYou get the first draft, not the checked draft — but see the self-preference-bias caveat below: a self-graded rubric is not independent verificationAre the criteria specific enough to fail a bad draft, not just generic virtues?

Two companions to the five components, both from OpenAI's official prompting guide. First, the escape hatch: a clause that lets the model proceed under uncertainty rather than stalling — the guide's own example wording is "even if it might not be fully correct." Hard constraints without an escape hatch produce brittle briefs. Second, the rubric pattern itself — the guide tells developers to have the model spend time thinking of a rubric until it is confident, then iterate its own output against that rubric before returning a result. The rubric stays internal; the user just gets better output. If you maintain a shared prompt library, these components are also the backbone of a scoring system — our 100-point prompt-library audit operationalises the same anatomy as an evaluation checklist.

06The CaveatA rubric is a self-check, not a verdict.

The rubric pattern is vendor-endorsed and genuinely useful — and it has a sourced failure mode that most prompting advice omits. A 2026 study on rubric-based LLM evaluation found a systematic self-preference bias: models rate their own outputs more favorably than competing outputs even when applying the same standardized rubric to both. A rubric sharpens what the model is checking for; it does not make the model an objective judge of its own work. Related rubric-evaluation research adds one practical refinement: detailed, instance-specific criteria let LLM judges distinguish good from bad responses more accurately than generic axis-level rubrics. What the evidence does not support is stacking up negative criteria as the fix — the same self-preference-bias paper reports that negative rubrics, along with subjective topics like communication, are among the categories most susceptible to the bias.

The operational consequence: keep the rubric in the brief, and add independent verification on top for anything that matters. In our own workflow that means a different model — or a human — grades high-stakes output against the same rubric, which is the premise of cross-model review and consensus verification. One more boundary worth naming: briefs do not automatically survive model upgrades. Simon Willison's write-up of OpenAI's GPT-5.5 prompting guide highlights OpenAI's own advice to treat a new release as a new model family to tune for, not a drop-in replacement — the vendor itself concedes that a brief tuned for one generation does not transfer untouched to the next.

Self-preference bias
A model grading its own work against its own rubric is not independent verification. The self-preference-bias study found models score their own outputs more favorably than competitors' even under an identical standardized rubric. Use the rubric to raise the floor of the first draft — then have a different model, or a human, apply the same rubric before anything high-stakes ships.

07Industry SignalSpec-driven development went industry-wide.

If the brief-is-the-product thesis were just our take, it would be worth less. The strongest external signal is that the tooling industry has independently converged on it. GitHub's announcement of its open-source Spec Kit toolkit (September 2025, by principal product manager Den Delimarsky) states the core problem plainly: a vague prompt forces the model to guess at potentially thousands of unstated requirements. At the time of writing, Spec Kit had passed 125,000 GitHub stars and 11,200-plus forks, with support for more than 30 AI coding agents including Copilot, Claude Code, Cursor, Gemini CLI, and Codex CLI.

"Specifications don't serve code — code serves specifications."— Den Delimarsky, Principal Product Manager, GitHub

It is not one vendor's bet. Multiple tools — GitHub Spec Kit, AWS Kiro, Claude Code, Cursor, OpenSpec, BMAD, Tessl, Google Antigravity — have each, independently, shipped their own version of a spec-driven mode (that enumeration is our synthesis of trade coverage; no single source lists them all together). AWS's own case study describes a life-sciences target-identification agent built end-to-end by three developers in three weeks using Kiro's Requirements → Design → Tasks → Code workflow, with each generated component traceable back to a specific requirement — a vendor case study with no independent audit of the timeline, but a concrete picture of the workflow. You will find claims online that spec-driven development cuts rework by some specific percentage; we could not trace any such figure to a named source with a methodology, so we are not repeating one. The observable fact is the adoption curve, not a rework statistic.

The anti-pattern this whole movement is reacting to has a name too — prompt-and-pray coding, where underspecified requests compound into unmaintainable output. We catalogued that failure family in our vibe-coding anti-patterns guide; spec-driven development is, in effect, the industry writing the antidote into its tooling.

Spec Kit stars
At the time of writing
~125K

GitHub's open-source spec-driven toolkit had passed 125,000 stars with more than 11,200 forks — snapshot figures that keep climbing — signalling how mainstream write-the-spec-first has become.

11,200+ forks
Supported agents
One spec, many models
30+

Spec Kit supports 30+ AI coding agents — Copilot, Claude Code, Cursor, Gemini CLI, Codex CLI and others. The spec is deliberately the portable artifact; the model behind it is swappable.

Model-agnostic by design
Kiro case study
Spec to production
3wks

AWS's published case study: three developers, three weeks, a life-sciences agent built through a Requirements → Design → Tasks → Code spec workflow. Vendor-reported, not independently audited.

Traceable to requirements

08In PracticeThe method in practice: four briefs we publish today.

This essay is the argument; the method posts publishing alongside it are the practice. Each one is a concrete, reusable brief built on the five-component anatomy above — constraints, acceptance criteria, negative space, worked example, self-check — applied to a different deliverable. If you want to see what "the brief is the product" looks like as working method rather than thesis, start with these.

Theming
Design-system extraction

Extract an existing site's real tokens — color, type, spacing, radius — into a specification an AI app builder can apply, so generated UI inherits the brand instead of the tool's defaults. The brief is the design system.

The theming brief
Research
Deep-research audit prompt

A structured deep-research prompt that turns a model loose on a website audit with explicit scope, evidence standards, and output shape — the acceptance-criteria component doing the heaviest lifting.

The research brief
Voice
Brand-voice guide extraction

Distill a brand's actual published writing into a voice guide an AI content workflow can follow — worked examples and negative space (what this brand never says) as the load-bearing components.

The voice brief
UI review
Screenshot-driven UI loop

Use vision models to critique rendered UI against the spec — a visual rubric variant where the screenshot comparison is the self-check the model runs before you accept the change.

The visual rubric

The four briefs, in reading order: extracting a design system for AI app-builder theming, the deep-research website-audit prompt, extracting a brand-voice guide for AI content, and screenshot-driven UI development with vision models. One adjacent discipline is deliberately out of scope here: managing what an agent holds in its context window at runtime — retrieval budgeting, compaction, memory — is a distinct problem from authoring the brief, and it has its own playbook in context engineering for agent reliability. This post is about what a human writes before the run starts; that one is about what the system feeds the model while it runs.

Where this lands for client work: we treat the brief as a deliverable in its own right — versioned, reviewed, and reused — because it is the artifact that survives model swaps. When a client's AI transformation engagement produces a working specification library, the next model generation is an upgrade, not a rebuild — with the Willison caveat from Section 06 applied: re-tune, don't assume. The same logic drives our content engine work, where the voice guide and the acceptance criteria are the product and the model is an implementation detail. Looking forward, we expect the spec layer to keep absorbing value as models continue to commoditize general-purpose capability: if the last two years moved the differentiator from model access to model choice, the period ahead looks likely to move it from model choice to specification quality — with the stubborn exceptions in Section 02 as the standing reminder to keep running evals.

09ConclusionWrite the spec like it's the product.

The shape of the argument

Converged enough for most work — with named exceptions.

The evidence supports a bounded claim, and the bound is the point. On composite indices, agentic web research, and much of coding, current frontier models sit within a few points of each other — close enough that the brief you write is the bigger lever. On abstract reasoning, held-out professional domains, and specialised agentic tasks, the same launch tables show 5x and roughly 20x spreads. A post that conceded nothing here would be less credible, not more.

The practical program follows directly. Write briefs with all five components: hard constraints, measurable acceptance criteria, honestly-chosen negative space, a worked example, and a self-check rubric. Add independent verification on top of the rubric, because self-graded work carries self-preference bias. Re-tune briefs when model generations change. And for any task in the wide-gap families, run your own eval before trusting anyone's table — vendor or independent, including ours.

The industry's tooling is already voting this way — spec-driven modes across at least eight tools, a toolkit past 125,000 stars built on the premise that the specification, not the code, is the source of truth. The teams that win the next cycle will not be the ones with secret model access; everyone has the same models. They will be the ones whose specifications are precise enough that any frontier model can execute them — and checkable enough that the output actually ships.

Make the brief your asset

The models converged. Your advantage is now the specification.

We build specification libraries, evaluation rubrics, and AI-assisted delivery workflows — the briefs, criteria, and self-checks that make model output shippable, portable across vendors, and durable across upgrades.

Free consultationExpert guidanceTailored solutions
What we work on

Specification-first engagements

  • Brief and spec libraries for AI-assisted delivery
  • Acceptance criteria + evaluation rubric design
  • Cross-model verification workflows
  • Model-swap portability audits for prompt libraries
  • Task-family evals for wide-gap domains
FAQ · The brief method

The questions we get every week.

Partially — and the partial matters. On the Artificial Analysis Intelligence Index (July 17, 2026, a snapshot taken before Claude Opus 5 shipped and reshuffled the top three), the top three models came from three different labs and spanned just 3 points, with six labs fielding a model above 50. On BrowseComp, four current models sit within 6.5 points on Anthropic's own July 24 launch table. But the same table shows a roughly 20x spread on ARC-AGI-3 (30.2% vs 1.5% of tasks solved) and a roughly 5x spread on a held-out legal benchmark. The honest summary: converged enough that brief quality dominates for most general work — coding, research, writing, agentic web tasks — with named exceptions in abstract reasoning and specialised domains where model choice still decides the outcome.
Related dispatches

Continue exploring AI-assisted development.