Topic

#ai-evaluation

16 articles tagged ai-evaluation. Browse the full set below, or see all topics.

Tagged "ai-evaluation"

Cross-cutting reads on this topic

16 articles
Update GPT-6 Astra skills, AGENTS.md and task prompts with a practical audit, clear completion criteria and a rollout plan for multiple projects.
#GPT-6 Astra#Codex+4 more
2026-09-11
Read Article
Test whether an AI agent uncovers missing client requirements before building. Compare interview quality, code inspection, evidence and acceptance criteria.
#Requirements Discovery#Coding Agents+3 more
2026-09-09
Read Article
Assess AI research claims with a practical evidence matrix. Separate formal proofs, measured results and demos, and record what each check establishes.
#AI Research#AI Evaluation+3 more
2026-09-09
Read Article
Evaluate a vision model on real interface screenshots, including missing text and ambiguous controls. Score reading, target location and unsupported claims.
#Vision Models#AI Evaluation+3 more
2026-09-09
Read Article
DeepSeek V4.1 Flash adds vision and lower prices. Compare 19 official benchmarks, agent frameworks and API costs; V4 Pro now continues past September 14.
#DeepSeek#Model Releases+3 more
2026-09-09
Read Article
Compare AI models using the review work needed for an accepted result. Track inspection, corrections and rechecks before treating a lower bill as savings.
#AI model selection#Human review+3 more
2026-09-06
Read Article
Two AI reviewers can repeat one mistake. Design reviews around separate evidence checks, clear rubrics and independent calculations instead of votes.
#AI agents#AI evaluation+3 more
2026-09-05
Read Article
OSWorld measures desktop GUI control, Terminal-Bench measures CLI work. Same-named benchmarks differ by version and harness, so check before you compare.
#osworld#terminal-bench+5 more
2026-08-05
Read Article
How to read AI model leaderboards without being fooled by benchmark contamination, eval gaming, and cherry-picked MMLU, GPQA, and SWE-bench scores.
#llm-benchmarks#ai-evaluation+6 more
2026-05-27
Read Article
Stage 5 of the agentic AI pipeline — prototype brief, eval harness, success criteria, and the templates that get you from vendor to live demo.
#agentic-ai-pipeline#prototype-templates+7 more
2026-05-07
Read Article
200 agentic AI terms — agents, MCP, memory, planning, evaluation, governance. Examples, source links, cross-references. The reference glossary.
#agentic-ai#glossary+8 more
2026-04-30
Read Article
Eighty AI eval metrics — quality (BLEU, ROUGE, BERTScore), agentic (success rate, tool accuracy), safety, fairness, RAG (faithfulness). Defined.
#ai-evaluation#eval-metrics+8 more
2026-04-30
Read Article
Defining ASR for production agents — completion vs partial vs hallucinated, cost-adjusted scoring, and benchmark suites. Formal methodology + reference dataset.
#agent-evaluation#agent-success-rate+6 more
2026-04-27
Read Article
We measured low/medium/high reasoning effort across 5 frontier models on math, code, and analysis. Quality lift, latency tax, and cost-per-correct-answer data.
#reasoning-effort#ai-benchmarks+8 more
2026-04-23
Read Article
Cross-model hallucination rates on factual recall, citation accuracy, and code reference. 5,000 prompts tested across 5 frontier models with confidence bands.
#ai-hallucination#ai-benchmarks+8 more
2026-04-23
Read Article
Why $/token is the wrong unit and $/successful-task is the right one. Formulas, worked examples across 6 task families, and a downloadable scoring template.
#ai-evaluation#cost-per-task+8 more
2026-04-23
Read Article