Topic

#llm-pricing

26 articles tagged llm-pricing. Browse the full set below, or see all topics.

Tagged "llm-pricing"

Cross-cutting reads on this topic

26 articles
Inception's Mercury 2.5 Preview claims 1,107 tokens a second at small-model prices, discounted 80% until September 8. What a diffusion LLM changes for agents.
#inception-labs#mercury-2-5+4 more
2026-09-01
Read Article
Identical dollar-per-million pricing does not mean identical cost: tokenizers split the same text into different counts. A vocabulary reference table.
#tokenizers#llm-pricing+5 more
2026-08-24
Read Article
DeepSeek shipped a separate experimental V4 Flash vision model on August 21 at text prices, on a peak and off-peak clock that doubles the rate.
#deepseek#vision models+5 more
2026-08-21
Read Article
OpenAI cut GPT-5.6 Sol to $4 in and $20 out per million tokens on August 21. The same Vercel changelog window added a $0.10/GB-month storage line.
#vercel#gpt-5.6 sol+5 more
2026-08-21
Read Article
Price-change notice is a contract term, not a courtesy, and the periods range from a fortnight to none at all. What to check before committing model spend.
#ai-procurement#llm-pricing+5 more
2026-08-15
Read Article
Crossing a context threshold can reprice an entire request, not just the tokens above the line. Where the cliffs sit and how to keep prompts underneath.
#long-context#llm-pricing+5 more
2026-08-15
Read Article
Time-of-day LLM pricing is arriving, and a cheaper window is not the same as a cheaper bill. How the announced schedules and peak hours actually work.
#llm-pricing#off-peak-pricing+5 more
2026-08-15
Read Article
Gemini 3.6 and 3.7 Flash both bill $0.75/$3.75 until December 31, with $1.50/$7.50 posted for January. Budget at the standard rate, not the promo.
#ai-cost-management#gemini-3-7-flash+5 more
2026-08-14
Read Article
Google, OpenAI and Anthropic all cut batch pricing 50 percent, while Grok 4.6 gets no batch discount at all. What the cheap lane actually costs you.
#batch-api#llm-pricing+5 more
2026-08-14
Read Article
Mistral's Aug 11 launch pins inference to the EU or US for 10% more, adds a 99.5% SLA tier at 75% more, and hosts Z.ai's GLM-5.2 alongside its own.
#mistral-ai#eu-data-residency+5 more
2026-08-13
Read Article
Gemini 3.7 Flash ships at $0.75/$3.75 per M through Dec 31, 2026 — but Google cut 3.6 Flash to that same intro rate, so half-price needs an asterisk.
#gemini-3-7-flash#google-deepmind+5 more
2026-08-13
Read Article
Qwen3.8-Max ships as a 2.4T-parameter MoE with 95B active and a 1M context. What the vendor benchmark table shows, and what was only promised at launch.
#qwen#alibaba+5 more
2026-08-03
Read Article
One week repriced the cheap end of the model market. A surface-labelled table of what Luna, Qwen3.7 Flash and DeepSeek V4 Flash now cost, and what that buys.
#llm-pricing#cost-optimization+5 more
2026-08-01
Read Article
Meta's Muse Spark 1.1 and SpaceXAI's Grok 4.5 launched a day apart — two cheap, agentic value models. We compare price, context, tool use and coding.
#muse-spark#grok-4-5+6 more
2026-07-09
Read Article
Meta enters the paid model-API market with Muse Spark 1.1 — a $1.25/$4.25, 1M-context agent model that tops tool-use benchmarks but trails on pure coding.
#meta-muse-spark#meta-superintelligence-labs+6 more
2026-07-09
Read Article
Grok 4.5, Opus 4.8, and GPT-5.5 compared on benchmarks, cost, context, and reliability — an honest, source-checked read on which model wins which job.
#grok-4-5#claude-opus-4-8+5 more
2026-07-08
Read Article
SpaceXAI shipped Grok 4.5 on July 8: co-trained with Cursor, $2/$6 per Mtok, 4.2x fewer tokens than Opus 4.8. Where it fits an agency stack.
#grok-4-5#spacexai+6 more
2026-07-08
Read Article
Dev shops quote $8K-$500K for the same 'simple agent' with no math. Our itemized 2026 index: stated-assumption build and monthly run costs you can recompute.
#ai-agent-cost#ai-agent-development+4 more
2026-07-03
Read Article
Anthropic shipped Claude Sonnet 5 on June 30, 2026. Its most agentic Sonnet yet lands near Opus 4.8 on coding and knowledge work at roughly half the price.
#Claude Sonnet 5#Anthropic+6 more
2026-06-30
Read Article
Uber burned a year's AI budget in four months; Microsoft cut Claude Code. A 4-gate playbook that right-sizes model spend and cuts AI bills 60 to 80 percent.
#ai-cost-optimization#model-routing+5 more
2026-06-23
Read Article
We compare Claude Opus 4.8 and GPT-5.5 on coding, agents, reasoning, and real cost — including where GPT-5.5 still wins and which model fits which job.
#claude-opus-4-8#gpt-5-5+6 more
2026-05-28
Read Article
Forecast monthly LLM token spend by client tier, workflow, model mix. Downloadable forecasting model + reconciliation cadence for agency CFOs.
#token-budget#ai-cost-management+6 more
2026-04-27
Read Article
Side-by-side input, output, cached, and batch pricing for 30 frontier and open-weight models across 12 providers. Updated April 2026 with 200+ price points.
#ai-model-pricing#llm-pricing+8 more
2026-04-23
Read Article
Xiaomi's MiMo-V2-Pro has 1T+ parameters with 42B active, ranking #8 worldwide. $1/$3 per million tokens. Complete model guide and deployment.
#mimo-v2-pro#xiaomi-ai+5 more
2026-03-18
Read Article
GPT-5.4 Nano at $0.20 per million input tokens is OpenAI's cheapest model. API-only, designed for classification, extraction, and subagents.
#gpt-5-4-nano#openai-api+5 more
2026-03-17
Read Article
Qwen 3.5 medium series: Flash, 35B-A3B, 122B-A10B, and 27B. Benchmarks vs GPT-5 mini and Claude Sonnet 4.5, pricing from $0.10/M tokens.
#qwen-3-5#alibaba-ai+6 more
2026-02-25
Read Article