Topic

#model-evaluation

12 articles tagged model-evaluation. Browse the full set below, or see all topics.

Tagged "model-evaluation"

Cross-cutting reads on this topic

12 articles
A free, anonymous 1M-context reasoning model appeared on OpenRouter on August 20. Nobody has claimed it, and its no-training promise is a one-off exception.
#ox-alpha#openrouter+4 more
2026-08-20
Read Article
We collected 42 vendor-published benchmark rows from 9 vendors and checked what each page discloses: version, subset, harness, effort. Full table in the post.
#ai-benchmarks#vendor-claims+5 more
2026-08-16
Read Article
Four levers decide what a vendor benchmark table means: which benchmark version, which task subset, what shape the metric takes, and who got a column.
#benchmarks#vendor-claims+5 more
2026-08-12
Read Article
Anthropic says its Fable 5 retune cut biology-related fallbacks about 85% across product surfaces. A rare published guardrail false-positive figure.
#Anthropic#AI Safety+4 more
2026-08-09
Read Article
Kimi K3 leads open-weights models on intelligence, but its hallucination rate climbed to 51%. Why you must run your own evals before adopting it.
#Kimi K3#AI Benchmarks+4 more
2026-07-27
Read Article
ARC Prize independently administered Claude Opus 5's 30.16% ARC-AGI-3 result. Almost every other launch-day benchmark number is vendor-run on a vendor harness.
#arc-prize#ai-benchmarks+5 more
2026-07-26
Read Article
Between July 17 and 23, seven notable models shipped from five vendors — most efficiency or positioning plays, not capability jumps. A triage framework.
#ai-model-releases#model-wave-2026+5 more
2026-07-24
Read Article
Four AIs hit a perfect IMO 2026 score in one VC's own test, not official IMO grading. What benchmark saturation means for how you should judge models.
#imo 2026#ai benchmarks+4 more
2026-07-23
Read Article
Claude Fable 5 tops SWE-bench Verified at 95%, but 99 of 100 results are self-reported and the scaffold gap can exceed 28 points. How to read the numbers.
#swe-bench#ai-coding-benchmarks+6 more
2026-06-16
Read Article
Epoch AI found errors in 42% of FrontierMath problems and shipped v2 on June 12, 2026. Scores jumped, rankings held — here is what that means for model choice.
#frontiermath#ai-benchmarks+4 more
2026-06-14
Read Article
Opus 4.8 tops the Artificial Analysis index, but GPT-5.5 still leads Terminal-Bench. An evidence-graded roundup of the first 48 hours of independent evals.
#claude-opus-4-8#ai-benchmarks+5 more
2026-05-30
Read Article
How to read AI model leaderboards without being fooled by benchmark contamination, eval gaming, and cherry-picked MMLU, GPQA, and SWE-bench scores.
#llm-benchmarks#ai-evaluation+6 more
2026-05-27
Read Article