Category

AI Development Articles

Page 17 of 34. Deep dives into AI-assisted and agentic development. Coding agents, frontier model releases, SDKs, prompting patterns, and the engineering workflows behind building production software with AI.

Page 17 of 34

The newest AI Development guides and analysis

Showing 385-408 of 804 articles
MiniMax M3 fuses frontier coding, a 1M-token context window, and native multimodality. Inside its Sparse Attention design, vendor benchmarks, and pricing.
#minimax-m3#open-weight-models+6 more
2026-05-31
Read Article
NVIDIA's May 31 keynote pushed every compute tier into the agentic era: RTX Spark runs 120B models locally, Vera Rubin delivers 10x agent throughput.
#nvidia-computex-2026#rtx-spark+6 more
2026-05-31
Read Article
StepFun's Apache-2.0 Step 3.7 Flash pairs a 196B MoE backbone with a 1.8B vision encoder, activating ~11B params per token. The cost case for agentic teams.
#stepfun-step-3-7-flash#mixture-of-experts+6 more
2026-05-30
Read Article
Opus 4.8 tops the Artificial Analysis index, but GPT-5.5 still leads Terminal-Bench. An evidence-graded roundup of the first 48 hours of independent evals.
#claude-opus-4-8#ai-benchmarks+5 more
2026-05-30
Read Article
Claude Opus 4.8 lands May 28 with stronger coding benchmarks, a major honesty gain, new effort controls, and dynamic workflows in Claude Code.
#claude-opus-4-8#anthropic+6 more
2026-05-28
Read Article
We compare Claude Opus 4.8 and GPT-5.5 on coding, agents, reasoning, and real cost — including where GPT-5.5 still wins and which model fits which job.
#claude-opus-4-8#gpt-5-5+6 more
2026-05-28
Read Article
Gemini 3.5 Flash beats Claude Opus 4.8 on MCP-Atlas and Finance Agent at a third of the price — but a 61% hallucination rate complicates the routing call.
#claude-opus-4-8#gemini-3-5-flash+6 more
2026-05-28
Read Article
How to read AI model leaderboards without being fooled by benchmark contamination, eval gaming, and cherry-picked MMLU, GPQA, and SWE-bench scores.
#llm-benchmarks#ai-evaluation+6 more
2026-05-27
Read Article
When self-hosting open-weight models beats API calls: a cost-crossover model, GPU sizing tables, and a deployment matrix for vLLM, SGLang, and Ollama.
#self-hosting-llm#open-weight-models+6 more
2026-05-27
Read Article
A practical playbook for chunking documents in RAG pipelines, comparing fixed, semantic, recursive, and late-chunking with retrieval-quality benchmarks.
#rag#chunking+6 more
2026-05-27
Read Article
What to log, trace, and alert on when running AI agents in production: an observability-stack comparison covering spans, token cost, eval gates, replay.
#ai-observability#agent-tracing+6 more
2026-05-27
Read Article
A reference architecture for layering input, output, and tool-call guardrails on production LLM systems: prompt-injection, PII, and jailbreak defense.
#llm-guardrails#ai-safety+6 more
2026-05-26
Read Article
A playbook for engineering production-agent context windows: retrieval budgeting, compaction, memory tiering, and tool-result pruning with keep-or-drop rules.
#context-engineering#ai-agents+6 more
2026-05-26
Read Article
A decision guide for when to generate synthetic training and eval data versus collecting real data: distillation, bootstrapping, and model-collapse risk.
#synthetic-data#llm-training+6 more
2026-05-26
Read Article
A technical reference for hybrid retrieval: BM25 keyword scoring, dense vector search, reciprocal rank fusion, and cross-encoder reranking for better RAG.
#hybrid-search#bm25+6 more
2026-05-26
Read Article
Alibaba's Qwen 3.7 Max ships with 1M context, $2.50/$7.50 pricing, and benchmarks topping Opus 4.6 on Terminal-Bench, SWE-Bench Pro, and MCP-Atlas.
#qwen-3-7-max#alibaba-qwen+7 more
2026-05-25
Read Article
50+ AI coding adoption stats for 2026 — sourced from Stack Overflow, JetBrains, GitHub Octoverse, DORA, McKinsey, GitClear, Veracode, and DX. Each cited.
#ai-coding-adoption#developer-statistics-2026+6 more
2026-05-25
Read Article
Seven production patterns for Anthropic's self-hosted sandbox from Code with Claude London — Docker isolation, MCP tunnels, HITL gates, evals, rollback.
#anthropic-sandbox#self-hosted-agents+6 more
2026-05-25
Read Article
Five days after Gemini 3.5 Flash GA — independent benchmarks from Artificial Analysis, llm-stats, WaveSpeed, and Aider. Google-claimed vs verified.
#gemini-3-5-flash#independent-benchmarks+6 more
2026-05-25
Read Article
12-dimension decision matrix for AI agent deployment — when to choose SaaS (Agentforce 360, Copilot Studio), self-hosted (Anthropic, LangGraph), or hybrid.
#ai-agent-deployment#decision-matrix+6 more
2026-05-25
Read Article
Issue #1 of the MCP server tracker — 56 production-ready servers cataloged across 10 categories with auth model, transport, maintainer, and status.
#mcp-server-tracker#model-context-protocol+6 more
2026-05-25
Read Article
Eight biggest AI industry stories from May 18-25, 2026 — Google I/O, Composer 2.5, Code with Claude London, Copilot Studio GA, SpaceX S-1, Qwen 3.7 Max.
#weekly-recap#ai-industry-may-2026+6 more
2026-05-25
Read Article
Every AI model released in May 2026 — Gemini 3.5 Flash, Composer 2.5, Grok Build, Gemini Omni, Antigravity 2.0. Pricing, specs, and availability tracker.
#ai-model-releases-may-2026#gemini-3-5-flash+8 more
2026-05-24
Read Article
Forward-looking forecast for agentic coding in H2 2026 — Microsoft Build June 2-3, GPT-5.6 by June 30, Cursor's Colossus 2 model, IDE consolidation wave.
#agentic-coding-h2-2026#microsoft-build-2026+8 more
2026-05-24
Read Article
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.