Category

AI Development Articles

Deep dives into AI-assisted and agentic development. Coding agents, frontier model releases, SDKs, prompting patterns, and the engineering workflows behind building production software with AI.

Latest Articles

The newest AI Development guides and analysis

Showing 1-24 of 829 articles
Artificial Analysis ran GPT-6 Sol and Claude Opus 5.5 on the same tests. Sol is cheaper up to a point; Opus 5.5 at its default outscores Sol at max.
#AI Models#OpenAI+2 more
2026-09-22
Read Article
GPT-6 Sol costs $2/$10 and Luna $0.10/$0.50 per million tokens, half GPT-5.6's price. What OpenAI's own charts show about scores and effort.
#AI Models#OpenAI+2 more
2026-09-22
Read Article
Opus 5.5 outscores Fable 5.1 on every coding benchmark Anthropic published, at 40% of the token price. When Fable 5.1 still earns its place.
#AI Models#Claude+3 more
2026-09-22
Read Article
Grok 4.7 lists at under a third of Opus 5.5's output price, yet costs twice as much per task in Artificial Analysis tests. Where Muse Spark 1.3 fits.
#AI Models#Claude+4 more
2026-09-22
Read Article
Claude Opus 5.5 lists at 40% of GPT-6 Astra's price with no long-context surcharge. Where each leads, and where the vendors' benchmarks disagree.
#AI Models#Claude+3 more
2026-09-22
Read Article
Claude Opus 5.5 costs $4/$20 per million tokens, defaults to medium effort and breaks four Opus 5 integrations. What the 40% saving assumes.
#AI Models#Claude+2 more
2026-09-22
Read Article
A real capability regression and a broken measurement pipeline look identical on one dashboard. Two tests separate them, five common breaks, a first-hour plan.
#Agentic AI#Evaluation+2 more
2026-09-21
Read Article
20% of US businesses use AI; nearly nine in ten survey respondents say theirs does. Both are real. Four evidence classes explain the gap, with eight numbers.
#AI Adoption#Statistics+2 more
2026-09-21
Read Article
Before an AI agent rewrites hundreds of pages, index every claim, figure and source already in them. The schema, three rules, and what to do when a claim fails.
#Agentic AI#Content Operations+2 more
2026-09-21
Read Article
Six signals that a model is shipping, five checkable and one not, from stealth tests to alias flips, with the API field for each and three launches scored.
#AI Models#Model Catalogs+2 more
2026-09-21
Read Article
xAI released Grok 4.7 on September 21, 2026 at Grok 4.6's $2/$6 list price. The vendor-run benchmark table, the route price you are billed, three questions.
#AI Models#Grok+2 more
2026-09-21
Read Article
16 models in the open-model conversation, checked against the licence file, OSI's list and what else was released: weights, data, recipes, evaluation harness.
#AI Models#Open Source+2 more
2026-09-21
Read Article
Fan a batch out across subagents and items vanish with zero errors. The fault is the hand-written work list, not the agents. Three corruption points, one fix.
#Agentic AI#Multi-Agent Systems+2 more
2026-09-21
Read Article
A cost formula for AI jobs at volume: generation, verification and rework. Worked at three failure rates and three checking designs on published 2026 rates.
#Agentic AI#AI Costs+2 more
2026-09-20
Read Article
A census of 12 external evaluation arrangements at Anthropic, OpenAI and Google DeepMind: who evaluates, who pays, what access they get, what gets published.
#Agentic AI#AI Governance+2 more
2026-09-20
Read Article
npm added stage-only tokens on September 18, 2026 and targets January 2027 to end direct publishing by bypass-2FA tokens. Who moves to what, and what stays.
#Agentic AI#Supply Chain+2 more
2026-09-20
Read Article
Claude Code 2.1.278 stopped charging for its auto-mode classifier where checks run server-side. What Claude Code, OpenRouter and Copilot bill for routing.
#Agentic AI#AI Costs+2 more
2026-09-20
Read Article
A September 17 advisory shows four coding agents resolving a pinned commit to a same-named branch. How git decides, which versions fix it and what a pin means.
#Agentic AI#AI Security+2 more
2026-09-20
Read Article
A matrix from vendor documentation: which filenames nine coding agents read, what wins when several exist, nested files, size caps and what is not documented.
#Agentic AI#Coding Agents+2 more
2026-09-20
Read Article
A census of 15 industry packages from OpenAI, Anthropic and xAI, decomposed into index, instructions, tools, safeguards, gate and price. None ships a new model.
#Enterprise AI#OpenAI+2 more
2026-09-19
Read Article
Proof-of-Control, a 127-requirement draft standard open for comment, grades agent evidence by who you must trust. The tiers, the six domains, what to do now.
#Agentic AI#AI Governance+2 more
2026-09-19
Read Article
A census of 20 AI agents by the identity each acts under: 11 get their own account or token, 9 reuse your login. What it means for revocation and blast radius.
#Agentic AI#AI Security+2 more
2026-09-19
Read Article
IBM Research ran an agent five times per task: 77.4% of runs passed but only 53.0% of tasks passed every time. What the gap is and how to measure yours.
#Agentic AI#AI Evaluation+2 more
2026-09-19
Read Article
Anthropic's Life Sciences Verification Program gives vetted labs two grant types and moves enforcement from blocking to monitoring. Who qualifies, what it asks.
#Anthropic#Life Sciences+2 more
2026-09-19
Read Article
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.