Topic

#model-routing

27 articles tagged model-routing. Browse the full set below, or see all topics.

Tagged "model-routing"

Cross-cutting reads on this topic

27 articles
GPT-5.6 Luna fell 80% and Terra 20% on July 30, Sonnet 5's promo ends August 31, and OpenRouter list prices quietly mix standard with batch rates.
#ai-api-pricing#gpt-5-6+5 more
2026-08-05
Read Article
Qwen3.7 Flash listed at $0.03 per million input tokens on its base tier, 1M context and video input — but no published benchmarks. Where it fits in subagents.
#Qwen#Model Pricing+5 more
2026-07-28
Read Article
Microsoft's MAI-Cyber-1-Flash claims 96% on CyberGym at half the cost. The routing thesis behind it matters more than the unverified benchmark.
#Microsoft AI#MAI-Cyber-1-Flash+4 more
2026-07-27
Read Article
Runway's Media Router sends one API call to its own models or to rivals, with hard USD price caps per second or per image and dry-run routing validation.
#runway#generative-media+5 more
2026-07-26
Read Article
Claude Opus 5 launched July 24 at $5/$25 per million tokens — Anthropic's new state of the art on coding and knowledge work at half Fable 5's price.
#claude-opus-5#anthropic+6 more
2026-07-25
Read Article
AI cost control has split into four stackable levers: routing, planner/worker splits, token efficiency, and model choice. A 2026 playbook for cutting AI spend.
#ai cost optimization#model routing+5 more
2026-07-23
Read Article
Cursor Router auto-classifies each coding request and picks the right model. Cursor claims up to 60% savings in A/B tests versus running everything on Opus.
#cursor#model routing+4 more
2026-07-23
Read Article
Seedream 5.0 Pro joins Vercel AI Gateway alongside GPT-5.6 and Muse Spark 1.1. How to route image models for on-brand ad creative at scale in 2026.
#vercel ai gateway#seedream+5 more
2026-07-13
Read Article
Tencent's Hy3 shipped Apache 2.0 at a sub-300GB FP8 footprint. A routing guide: send agentic tool-use to Hy3 and repo-scale coding to a frontier model.
#hunyuan-hy3#open-weights+5 more
2026-07-07
Read Article
Fable 5 meters from July 7. Spend the remaining included window on durable plan artifacts that cheaper models can execute for months, not on routine drafts.
#Claude Fable 5#AI Strategy+5 more
2026-07-02
Read Article
Claude Fable 5 was included through July 7 — now extended to July 12 — then moves to metered usage credits. How to sequence the window before the meter starts.
#Claude Fable 5#Anthropic+5 more
2026-07-02
Read Article
Run Claude Fable 5 as the planner and route execution to cheaper models in Hermes and OpenClaw. The exact config keys, cost math, and marketplace hardening.
#Claude Fable 5#Hermes Agent+5 more
2026-07-02
Read Article
Ten one-shot audit and upgrade jobs for Claude Fable 5, each drawn from Anthropic's own prompting guide, with the cheaper model that should execute the plan.
#Claude Fable 5#Prompt Engineering+6 more
2026-07-02
Read Article
Claude Fable 5 costs $10/$50 per million tokens, double Opus 4.8. How individuals and small firms recoup that metered spend — with the honest counterexamples.
#Claude Fable 5#AI Costs+5 more
2026-07-02
Read Article
Claude Fable 5 pricing resolved July 18: Max and Team Premium keep 50% included access; Pro and Team Standard get $10/$50 usage credits from July 20 on.
#claude-fable-5#usage-credits+6 more
2026-07-01
Read Article
Why AI services don't earn SaaS margins: inference COGS, the 50-60% gross-margin reality, model-routing levers, and hybrid pricing for agency work.
#ai-unit-economics#ai-pricing+5 more
2026-06-30
Read Article
Uber burned a year's AI budget in four months; Microsoft cut Claude Code. A 4-gate playbook that right-sizes model spend and cuts AI bills 60 to 80 percent.
#ai-cost-optimization#model-routing+5 more
2026-06-23
Read Article
The Fable 5 export shutdown showed single-vendor AI can halt your business overnight. A four-step second-source playbook with open-weight failover backups.
#AI vendor resilience#open-weight models+5 more
2026-06-21
Read Article
NVIDIA shipped Nemotron 3 Ultra, a 550B open MoE reasoning model with weights, data and recipes under a permissive license. It runs fast but trails Kimi K2.6.
#nvidia#nemotron+6 more
2026-06-05
Read Article
MiniMax M3 lands at 5-17x lower cost, but Opus 4.8 leads SWE-bench Pro and GPT-5.5 wins Terminal-Bench. A full three-way agentic coding routing matrix.
#minimax-m3#claude-opus-4-8+6 more
2026-06-03
Read Article
The LLM gateway is now critical AI infrastructure. Compare LiteLLM, Portkey, Cloudflare, Vercel, and OpenRouter on caching, routing, and build-vs-buy economics.
#llm-gateway#litellm+6 more
2026-06-03
Read Article
Gemini 3.5 Flash beats Claude Opus 4.8 on MCP-Atlas and Finance Agent at a third of the price — but a 61% hallucination rate complicates the routing call.
#claude-opus-4-8#gemini-3-5-flash+6 more
2026-05-28
Read Article
A FinOps playbook for cutting AI inference spend without quality loss: model routing, prompt and KV caching, batching, quantization, and unit-cost tracking.
#inference-cost#ai-finops+6 more
2026-05-26
Read Article
Hands-on deep dive into Zed's AI coding — Parallel Agents, channels, threads, performance-first editor, model routing, and the workflows enabled.
#zed-editor#deep-dive+7 more
2026-05-10
Read Article
Hands-on deep dive into Continue.dev — the open-source AI coding assistant, model routing, context providers, slash commands, IDE integrations.
#continue-dev#deep-dive+7 more
2026-05-10
Read Article
Hands-on deep dive into Cursor 3 — Agents and Cloud Agents, Composer evolution, MCP integration, model routing, and what's new vs Cursor 2.
#cursor-3#deep-dive+7 more
2026-05-10
Read Article
OpenRouter Fusion sends queries to multiple AI models, analyzes outputs, and fuses optimal results. Deep Research agents preferred Fusion to their own outputs.
#openrouter-fusion#multi-model-ai+5 more
2026-04-01
Read Article