Topic

#ai-cost-optimization

10 articles tagged ai-cost-optimization. Browse the full set below, or see all topics.

Tagged "ai-cost-optimization"

Cross-cutting reads on this topic

10 articles
Sonnet 5's $2/$10 intro price looks cheap until effort climbs. A decision matrix for Sonnet 5 vs Opus 4.8 vs Fable 5, with cost-per-task data by effort level.
#claude-sonnet-5#opus-4-8+6 more
2026-07-01
Read Article
Uber burned a year's AI budget in four months; Microsoft cut Claude Code. A 4-gate playbook that right-sizes model spend and cuts AI bills 60 to 80 percent.
#ai-cost-optimization#model-routing+5 more
2026-06-23
Read Article
Route each request to the cheapest model that can handle it. Model routing cuts real LLM bills 40-85% with no visible quality loss. The 2026 engineering guide.
#llm-routing#ai-cost-optimization+5 more
2026-06-14
Read Article
The LLM gateway is now critical AI infrastructure. Compare LiteLLM, Portkey, Cloudflare, Vercel, and OpenRouter on caching, routing, and build-vs-buy economics.
#llm-gateway#litellm+6 more
2026-06-03
Read Article
Tokens, prefill, decode, KV cache, prefix cache, batch tier, reserved capacity, context-window discount — every LLM billing term, defined.
#token-economics#llm-cost+8 more
2026-04-30
Read Article
GPU spend, ops headcount, latency, and break-even volume for hosting Llama, Qwen, DeepSeek, and Mistral yourself vs API. With per-token cost curves at 4 scales.
#self-hosting-llm#ai-tco+8 more
2026-04-24
Read Article
Paged attention, prefix caching, MQA/GQA, MLA, and quant-aware caching — when each technique pays off and the inference-cost numbers behind it.
#kv-cache#llm-inference+8 more
2026-04-24
Read Article
Cross-model quality regression, throughput lift, and VRAM savings at GPTQ-4, AWQ-4, INT8, and FP8 — benchmark data across 6 open-weight models.
#quantization#gptq+8 more
2026-04-24
Read Article
When 1M context pays off — and when it bankrupts you. Token-spend math, prompt-cache strategy, and break-even tables for agentic Claude Opus 4.7 workloads.
#claude-opus-4-7#anthropic+8 more
2026-04-23
Read Article
Business guide to small language models for on-device deployment. Gemma 4 E2B/E4B, Microsoft Phi-4, and Qwen 3.5 compared for cost savings and privacy.
#small-language-models#gemma-4+5 more
2026-04-03
Read Article