Tagged "benchmarks"
Cross-cutting reads on this topic
xAI's launch table sets Grok 4.6's default effort against GPT-5.6 Sol's and Fable 5's documented ceilings. Why the rung changes what a score means.
#grok-4-6#gpt-5-6-sol+5 more
2026-08-12
Read Article
Four levers decide what a vendor benchmark table means: which benchmark version, which task subset, what shape the metric takes, and who got a column.
#benchmarks#vendor-claims+5 more
2026-08-12
Read Article
Mem0, Letta and Zep bet differently on what remembering means. A licence, architecture and benchmark comparison, including why the published scores disagree.
#agent-memory#mem0+5 more
2026-08-04
Read Article
Claude Opus 4.7 beats GPT-5.4 on SWE-bench Pro, tool use, and computer use. Full agentic coding benchmark comparison with migration guidance.
#Claude#GPT-5+4 more
2026-04-16
Read Article
Gemini 3.1 Pro scores 77.1% on ARC-AGI-2 and 2887 Elo on LiveCodeBench at $2/$12M tokens. Full benchmarks, pricing, and competitive comparison guide.
#Gemini 3.1 Pro#Google+6 more
2026-02-19
Read Article
Gemini 3.1 Pro vs Claude Opus 4.6 vs GPT-5.3-Codex for agentic coding. SWE-Bench, Terminal-Bench, LiveCodeBench, and pricing comparison with recommendations.
#Gemini 3.1 Pro#Claude Opus 4.6+6 more
2026-02-19
Read Article
Claude Sonnet 4.6 scores 72.5% on OSWorld and 79.6% on SWE-bench Verified at $3/$15M tokens. Complete benchmarks, coding, computer use, and pricing guide.
#Claude Sonnet 4.6#Anthropic+6 more
2026-02-17
Read Article
ByteDance Seed 2.0 Pro scores 98.3 on AIME25, 87.8 on LiveCodeBench, and 3020 Codeforces. Full benchmarks, agentic capabilities, and Volcano Engine API.
#Seed 2.0#ByteDance+6 more
2026-02-16
Read Article
Qwen 3.5-397B scores 83.6 on LiveCodeBench v6 and 91.3 on AIME26 with 17B active MoE params. Benchmarks vs GPT-5.2, Claude, and pricing details.
#Qwen 3.5#Alibaba+6 more
2026-02-16
Read Article
DeepSeek V4 brings 1 trillion parameters, 1M token context, and Engram O(1) memory. Architecture details, leaked benchmarks, and what it means for developers.
#DeepSeek#DeepSeek V4+6 more
2026-02-14
Read Article
Claude Opus 4.5, GPT-5.2 Codex, and Gemini 3 Pro compared. SWE-bench scores, pricing, context windows, and coding tests to choose your AI partner.
#Claude Opus 4.5#GPT-5.2 Codex+4 more
2026-01-02
Read Article