SWE-2 is worth a bounded coding trial if your team is considering Devin, but the buying decision needs more than a favorable benchmark. Check the exact product surface, the offer expiry and the work required to accept its changes. A cheaper model run can still leave an expensive review queue.
Cognition announced SWE-2 on September 10, 2026. This guide belongs to the September 11 editorial allocation; access and pricing were checked on September 12. The trial below is a proposed evaluation, not a report of hands-on SWE-2 results.
- 01The surface matters.Desktop and CLI availability does not establish an identical rollout or billing arrangement on every Devin surface.
- 02The offer has an end date.Treat the current SWE-2 promotion as a trial opportunity, not a permanent operating cost.
- 03Buy accepted work.Keep subscription, extra usage and reviewer time in the same decision record.
01 — Practical decisionSeparate the model from the product you are buying
The Cognition launch announcement says SWE-2 was available that day in Devin Desktop and CLI, with Web and Fusion rolling out. That wording establishes a release date and a rollout boundary. It does not establish that every account already has identical access.
SWE-2 is the model; the surrounding agent determines what it can read, run and change. A result from a particular harness combines the model with its tools, instructions, environment and budget. When comparing a desktop workflow with a cloud agent, keep those differences visible.
Before a trial, write down the surface, selected model, effort setting, repository revision and permissions. If the product chooses or changes a setting automatically, record that limitation. Our Devin overview provides platform background; this decision is about whether the new model improves a specific workflow enough to justify adopting it.
02 — Practical decisionPut current prices beside the offer expiry
The current pricing page lists Free at $0, Pro at $20 per month and Max at $200 per month. It advertises free SWE-2 use in Desktop and CLI through October 10, 2026. The scope and expiry are part of the price: do not extend that offer to cloud work or assume it continues afterward.
Teams needs a written clarification. The pricing banner describes an $80 monthly team plan plus $40 per full developer seat. The self-serve billing guide describes an $80 minimum and $40 full seats. They also differ on the member limit. This guide therefore does not calculate a Teams total.
For an individual trial, record the amount actually allocated to it and any extra usage. For a team purchase, resolve the billing formula, included access and member limit before accepting a quote. Historical Core-plan ACU examples should not silently become the current self-serve budget.
| Buying question | Evidence to retain | Decision it supports |
|---|---|---|
| Can this account use SWE-2? | Selected model and product surface | Start the trial only on the intended surface. |
| When does the promotion end? | Offer wording and October 10 expiry | Separate promotional and ongoing costs. |
| What happens beyond quota? | Current extra-usage terms and account limits | Set the trial spending boundary. |
| What does Teams actually cost? | Written seat and minimum-charge clarification | Leave the total unresolved until confirmed. |
03 — Practical decisionRead the unfavorable benchmark result too
Cognition reports SWE-2 at 50.0% on FrontierCode 1.1 Main, beside Fable 5.1 at 50.9%, with a claimed 64% lower cost. Its same launch table reports SWE-2 at 27.3% on Terminal-Bench 4, versus 55.8% for Fable 5.1. These are vendor-run results, not our measurements.
The useful lesson is task dependence. A close score on one suite does not imply equal performance on another. A benchmark cost comparison also does not tell you the subscription charge, review effort or failure cost of your application. Keep the benchmark name and version attached whenever you repeat a score.
Choose your trial task class before choosing the most flattering chart. A routine bug fix, a cross-module refactor and an unfamiliar environment can stress different abilities. The Terminal-Bench 4 interpretation guide explains why a changed benchmark should be read as a different test rather than a simple model regression.
04 — Practical decisionBuild a trial that can change your decision
Select a small set of representative repository tasks that you are allowed to use. For each, preserve the starting revision, requirement, acceptance checks and environment setup. Include an ordinary task and a difficult case that matters to your team. Do not quietly remove failures because the task seemed unfair after the run.
Give both workflows the same task boundary and comparable access. If you compare different harnesses, call it a workflow comparison. If the goal is to isolate the model, hold the harness and available tools fixed where the product allows it. Record unavoidable differences instead of pretending the test is perfectly controlled.
Review the patch against the requirement, not merely the agent's final message. Run the relevant checks and inspect changed behavior. Keep rejected changes, repairs and abandoned attempts in the record. A task completed only after a human rewrites the solution belongs in a different outcome category from a change accepted as delivered.
Illustrative working record
Task: correct a reproducible export bug on a fixed repository revision. Acceptance: the failing example succeeds, the existing export contract remains intact and the relevant regression checks pass.
Record: model, surface, effort, elapsed time, extra usage, reviewer minutes, repair minutes and accepted or rejected outcome. Leave every result blank until the trial is run.
05 — Practical decisionCount review effort in the same cost calculation
An illustrative calculation makes the accounting boundary clear. Suppose you allocate a $20 subscription charge and $10 of extra usage to a trial. Suppose review and repair take two hours, valued at an assumed $60 per hour. The trial cost is $20 + $10 + $120 = $150. If three tasks are accepted, the illustrative cost is $50 per accepted task.
Those usage, time and outcome numbers are invented inputs for the example, not SWE-2 measurements. If the subscription also supports unrelated work, choose and disclose an allocation rule. If no task is accepted, report the spend and failures; dividing by zero does not produce a useful task price.
This prevents a common buying mistake: celebrating cheap generation while excluding the labor that makes the result usable. Our human review cost guide covers that tradeoff in more detail. Keep risk separate too: a low average cost does not make a severe escaped defect acceptable.
06 — Practical decisionDecide what would justify expanding the trial
Before starting, define the evidence that would support a larger rollout. That might be acceptable patches on the selected task class, manageable review effort and predictable spending after the promotional period. Use thresholds that reflect your own work; this article supplies no universal pass rate.
At the end, distinguish three decisions. Adopt for the tested task class when the evidence supports it. Extend the trial when a material uncertainty can be resolved cheaply. Stop when the required access, quality or economics do not fit. A mixed result can justify narrow routing without standardizing the entire team on one agent.
The Astra skills and prompts guide is relevant when instruction changes accompany a model switch. Change one meaningful variable at a time where possible. Otherwise you may attribute an improvement to SWE-2 that actually came from a clearer task or a repaired test environment.
Evidence and scope
- Dates
- Editorial allocation: September 11, 2026. Sources retrieved and article reviewed September 12, 2026. Event dates are stated separately.
- Sources
- Cognition’s September 10 launch, Devin pricing and self-serve billing documentation. Benchmark figures are attributed to the vendor; account eligibility was not tested.
- Method
- Compared release wording, current commercial terms and benchmark scope. Trial fields and acceptance criteria are an original proposed buying method. Illustrative arithmetic was recalculated.
- Limits
- No SWE-2 run or benchmark replication was performed. Teams billing wording conflicts across primary pages; obtain written terms. No standalone model API rate is asserted.
07 — Next stepChoose a task class before choosing a plan
Choose a task class before choosing a plan
Use the promotion to answer a bounded question about your own coding work. Preserve the unfavorable outcomes, include review cost and confirm the price that applies when the offer ends. Expand only as far as the evidence supports.