Topic

#swe-bench-pro

10 articles tagged swe-bench-pro. Browse the full set below, or see all topics.

Tagged "swe-bench-pro"

Cross-cutting reads on this topic

10 articles
Anthropic shipped Claude Sonnet 5 on June 30, 2026. Its most agentic Sonnet yet lands near Opus 4.8 on coding and knowledge work at roughly half the price.
#Claude Sonnet 5#Anthropic+6 more
2026-06-30
Read Article
Claude Fable 5 tops SWE-bench Verified at 95%, but 99 of 100 results are self-reported and the scaffold gap can exceed 28 points. How to read the numbers.
#swe-bench#ai-coding-benchmarks+6 more
2026-06-16
Read Article
MiniMax M3 lands at 5-17x lower cost, but Opus 4.8 leads SWE-bench Pro and GPT-5.5 wins Terminal-Bench. A full three-way agentic coding routing matrix.
#minimax-m3#claude-opus-4-8+6 more
2026-06-03
Read Article
MiniMax M3 fuses frontier coding, a 1M-token context window, and native multimodality. Inside its Sparse Attention design, vendor benchmarks, and pricing.
#minimax-m3#open-weight-models+6 more
2026-05-31
Read Article
StepFun's Apache-2.0 Step 3.7 Flash pairs a 196B MoE backbone with a 1.8B vision encoder, activating ~11B params per token. The cost case for agentic teams.
#stepfun-step-3-7-flash#mixture-of-experts+6 more
2026-05-30
Read Article
Opus 4.8 tops the Artificial Analysis index, but GPT-5.5 still leads Terminal-Bench. An evidence-graded roundup of the first 48 hours of independent evals.
#claude-opus-4-8#ai-benchmarks+5 more
2026-05-30
Read Article
Alibaba's Qwen 3.7 Max ships with 1M context, $2.50/$7.50 pricing, and benchmarks topping Opus 4.6 on Terminal-Bench, SWE-Bench Pro, and MCP-Atlas.
#qwen-3-7-max#alibaba-qwen+7 more
2026-05-25
Read Article
Agentic coding head-to-head: Gemini 3.5 Flash vs GPT-5.5 vs Opus 4.7. MCP Atlas, SWE-Bench Pro, Terminal-Bench, plus Antigravity 2.0 launch context.
#gemini-3-5-flash#gpt-5-5+8 more
2026-05-19
Read Article
What SWE-Bench Pro, Terminal-Bench, CursorBench, and MCP Atlas actually measure — why vendor self-evals deserve skepticism. Developer guide.
#swe-bench-verified#swe-bench-pro+8 more
2026-05-17
Read Article
Alibaba's Qwen3.6-Max-Preview tops six coding benchmarks including SWE-bench Pro and Terminal-Bench 2.0. Closed-weights pivot and agency playbook inside.
#qwen-3-6-max#alibaba-qwen+8 more
2026-04-20
Read Article