OpenRouter's July 2026 intake reads like a full quarter of releases compressed into three weeks. Between July 1 and July 22, ten new model families reached the catalog — roughly one every 2.2 days — spanning OpenAI's renamed GPT-5.6 trio, xAI's Grok 4.5, Moonshot's Kimi K3, Meta's first paid-API model, Thinking Machines' debut, two coding specialists from Kuaishou, Meituan's LongCat 2.0, Google's triple Gemini launch, and poolside's open-weight Laguna line.
The spread is the story. The cheapest paid model in the wave, poolside's Laguna XS 2.1, lists at $0.06 per million input tokens; the most expensive, GPT-5.6 Sol, lists at $5.00 input and $30.00 output. That is roughly an 83x gap on input and a 250x gap on output — and a free Laguna S 2.1 variant sits below both at $0 with a 256K context window. A month ago, our June OpenRouter roundup tracked a 25x spread across five models in ten days. July widened the range and sustained the cadence for an entire month.
This tracker catalogs all fifteen July listings with live-verified pricing, computes the listing-lag column nobody else publishes — the gap between when a vendor shipped a model and when OpenRouter's created field says it arrived — and closes with a workload-by-workload routing posture. Every benchmark figure this month is vendor-reported, and the post labels each one as such.
- 01Ten model families in 22 days.Grok 4.5 (Jul 8), the GPT-5.6 trio (Jul 9), KAT-Coder V2.5 (Jul 10–11), Kimi K3 and Muse Spark 1.1 (Jul 16), Inkling (Jul 17), LongCat 2.0 (listed Jul 20), and a four-model July 21 — Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Laguna S 2.1 in paid and free variants.
- 02Pricing spans ~83x on input, ~250x on output.Laguna XS 2.1 at $0.06/$0.12 per million tokens anchors the floor; GPT-5.6 Sol at $5/$30 anchors the ceiling. A 500K-input, 100K-output task costs about $0.042 on Laguna XS versus about $5.50 on Sol — roughly 131x — before quality enters the decision.
- 03Listing dates lie; vendor dates are authoritative.LongCat 2.0 launched June 30 but OpenRouter listed it July 20 — a 20-day lag. KAT-Coder V2.5 was listed a day before its stated vendor date. Anyone using OpenRouter's created field as a launch calendar will misdate releases by up to three weeks.
- 04Every benchmark in the wave is vendor-reported.No independent audit existed for GPT-5.6, Grok 4.5, or Gemini 3.6 Flash benchmark tables at publication. poolside is the partial exception — its harness and trajectories are publicly disclosed — while Muse Spark 1.1's Terminal-Bench setup is actively disputed.
- 05Three models earn production traffic now.On the month's evidence, Gemini 3.6 Flash, Kimi K3, and Grok 4.5 are our picks for frontier production routing, with the specialists — Gemini 3.5 Flash-Lite, KAT-Coder Air, Laguna XS — covering the cost tiers below. At this cadence, revisit that call monthly.
01 — The WaveTen families in a single 22-day window.
Walk the calendar. poolside opened the month quietly, listing Laguna XS 2.1 on July 2 at the lowest paid price of the wave. xAI shipped Grok 4.5 on July 8. OpenAI took GPT-5.6 to general availability on July 9 — just under two weeks after its June 26 preview — and all six OpenRouter listings (Sol, Terra, Luna, and their pro-mode twins) went live the same day. Kuaishou's Kwaipilot released KAT-Coder V2.5 in Pro and Air tiers around July 10–11. Then the back half: Kimi K3 and Muse Spark 1.1 listed July 16, Thinking Machines' Inkling July 17, LongCat 2.0's delayed listing July 20, and a four-model July 21 that combined Google's triple announcement with poolside's Laguna S 2.1.
Count the distinct families that hit the catalog between July 1 and July 22 and you get ten — one roughly every 2.2 days. June's wave, by comparison, packed five major models into a single ten-day burst; July held nearly the same density for an entire month. The practical consequence is the same one we drew in June, amplified: any model-selection decision made at the start of July was stale by the end of it.
GPT-5.6 Sol · Terra · Luna
OpenAI's new naming convention: the number is the generation, the name is a durable capability tier — Sol the flagship, Terra balanced, Luna high-volume. Each tier has a pro reasoning mode reached via an API setting rather than a separate model slug.
Grok 4.5
Billed in xAI's launch post as “SpaceXAI's smartest model” and trained alongside Cursor on tens of thousands of NVIDIA GB300 GPUs (vendor-stated). Launched without EU availability, with xAI pointing to mid-July for EU access — verify before routing EU traffic.
Kimi K3
Moonshot's new flagship line: 2.8T total parameters, native multimodal, and open weights promised for July 27, 2026. API-available now — the weights commitment is an announced date, not yet a downloadable artifact.
02 — The MatrixEvery July listing, price and lag included.
The table below is the canonical reference for the July window: all fifteen listings with pricing per million tokens, context window, the date OpenRouter's created field reports, and the vendor's own launch date — plus the lag between the two. Prices were read from the live OpenRouter models API at publication and cross-checked against each vendor's primary announcement. No other roundup we have seen publishes the lag column, and it is the column that prevents the most mistakes.
| Model | OpenRouter ID | $/Mtok in · out | Context | Listed | Vendor launch · lag |
|---|---|---|---|---|---|
| Frontier tier · Jul 8 – 16 | |||||
| GPT-5.6 Sol | openai/gpt-5.6-sol | $5.00 · $30.00 | 1.05M | Jul 9 | Jul 9 (GA) · same-day |
| GPT-5.6 Terra | openai/gpt-5.6-terra | $2.50 · $15.00 | 1.05M | Jul 9 | Jul 9 (GA) · same-day |
| GPT-5.6 Luna | openai/gpt-5.6-luna | $1.00 · $6.00 | 1.05M | Jul 9 | Jul 9 (GA) · same-day |
| Grok 4.5 | x-ai/grok-4.5 | $2.00 · $6.00 | 500K | Jul 8 | Jul 8 · same-day |
| Kimi K3 | moonshotai/kimi-k3 | $3.00 · $15.00 | 1M | Jul 16 | Jul 16 · same-day |
| New entrants · Jul 9 – 18 | |||||
| Muse Spark 1.1 | meta/muse-spark-1.1 | $1.25 · $4.25 | 1M | Jul 16 | Jul 9 (preview) · 7 days |
| Inkling | thinkingmachines/inkling | $1.00 · $4.05 | 1M | Jul 17 | Jul 17–18 · same-day |
| KAT-Coder-Pro V2.5 | kwaipilot/kat-coder-pro-v2.5 | $0.74 · $2.96 | 256K | Jul 10 | Jul 11 · listed 1 day early |
| KAT-Coder-Air V2.5 | kwaipilot/kat-coder-air-v2.5 | $0.15 · $0.60 | 256K | Jul 10 | Jul 11 · listed 1 day early |
| Google's triple launch · Jul 21 | |||||
| Gemini 3.6 Flash | google/gemini-3.6-flash | $1.50 · $7.50 | 1M | Jul 21 | Jul 21 · same-day |
| Gemini 3.5 Flash-Lite | google/gemini-3.5-flash-lite | $0.30 · $2.50 | 1M | Jul 21 | Jul 21 · same-day |
| Open-weight value lane · Jul 2 – 21 | |||||
| LongCat 2.0 | meituan/longcat-2.0 | $0.30 · $1.20 | ~1M | Jul 20 | Jun 30 · 20 days (~3 wks) |
| Laguna S 2.1 | poolside/laguna-s-2.1 | $0.10 · $0.20 | 1M | Jul 21 | Jul 21 · same-day |
| Laguna S 2.1 (free) | poolside/laguna-s-2.1:free | $0.00 · $0.00 | 256K | Jul 21 | Jul 21 · same-day |
| Laguna XS 2.1 | poolside/laguna-xs-2.1 | $0.06 · $0.12 | 256K | Jul 2 | Jul 2 · same-day |
Two footnotes keep the table honest. GPT-5.6's pro reasoning modes (Sol Pro, Terra Pro, Luna Pro) list at the same base rates as their standard tiers — pro mode bills through reasoning tokens rather than a separate price — so the trio's three rows cover all six listings. And the KAT-Coder pair shows OpenRouter's created date one day before the stated vendor date — almost certainly a timezone or rounding artifact, not a leak. Treat July 10–11 as the release window.
03 — Frontier TierThe frontier tier gets three new anchors.
July's top shelf belongs to three launches. OpenAI's GPT-5.6 went GA on July 9 with its new two-axis naming — generation number plus a durable tier name, so Sol is always the flagship, Terra the balanced tier, Luna the high-volume tier — and its announcement promised "full availability over the next 24 hours" across ChatGPT, Codex, and the API. On OpenAI's own preview benchmark tables (vendor-stated, no independent audit at GA), Sol posts roughly 88.8% on Terminal-Bench 2.1 with an ultra mode near 91.9% — while the same tables show Anthropic's Claude line still leading several agentic-coding rows, including SWE-Bench Pro, by a wide margin.
Grok 4.5 landed a day earlier at half of Sol's input price. xAI's vendor-reported numbers put it at 83.3% on Terminal-Bench 2.1 — just behind Claude Fable 5's 84.3% on the same table — and 64.7% on SWE-Bench Pro against Fable 5's 80.4%, while claiming roughly 4.2x fewer output tokens than Opus 4.8's max mode on that benchmark. The efficiency claim matters as much as the score: at $2/$6 with a $0.50 cached-input rate, Grok 4.5 is priced as a workhorse, not a trophy. Its 500K context is the smallest of the frontier three, and EU availability had not arrived at launch.
Kimi K3 is the wave's open-weight wildcard. Moonshot shipped it July 16 as a 2.8T-parameter flagship with native multimodal input and a 1M-token window at $3/$15 — premium pricing by open-lane standards — with weights promised for July 27. If that date holds, K3 becomes the largest open-weight release of the summer; until the files land, it is a closed API model with an announced commitment.
$30/M out · 1.05M context
The price ceiling of the wave. Terra at $2.50/$15 and Luna at $1/$6 step the same architecture down the cost curve, and pro reasoning mode bills via reasoning tokens at the same base rate. Vendor benchmarks only at publication — no independent audit existed at GA.
$6/M out · 500K context
Vendor tables concede several rows to Claude while claiming large output-token efficiency gains. Not available in the EU at its July 8 launch, with xAI pointing to mid-July — confirm current EU status before routing European traffic through it.
$3/$15 per Mtok · 1M context
Open weights promised for July 27, 2026 — an announced date, not a shipped artifact. Premium-priced for the open lane, and the model to watch if the release lands as stated. Multimodal input is native rather than bolted on.
04 — New EntrantsTwo debuts and a coding specialist pair.
The mid-tier is where July got structurally interesting. Meta Superintelligence Labs launched Muse Spark 1.1 on July 9 through Meta's first-ever paid Model API — $20 in free credits at launch, closed weights, no Hugging Face release — and OpenRouter listed it July 16 under a new meta/ vendor prefix, distinct from the older meta-llama/ namespace. At $1.25/$4.25 with a 1M context it is priced to compete, but it launched US-only with a single Meta-hosted provider. Thinking Machines Lab — Mira Murati's outfit — followed on July 17 with Inkling, an open-weight 975B-total / 41B-active multimodal MoE at $1/$4.05, the lab's first major OpenRouter listing.
The quiet value story is Kwaipilot's KAT-Coder V2.5, released around July 10–11 in two tiers positioned to take a whole coding issue autonomously. Pro lists at $0.74/$2.96 and Air at $0.15/$0.60 — both with 256K context. Air's pricing is the notable one: it undercuts every general-purpose model in the wave except poolside's Laguna line while targeting agentic coding specifically, which makes it the budget candidate for issue-to-PR automation pipelines.
05 — Price FloorThe open-weight lane keeps lowering the floor.
Two releases define July's bottom rungs. Meituan's LongCat 2.0 — a sparse MoE with 1.6T total and roughly 48B active parameters, a ~1M-token window, and a repo-level coding and long-horizon agent focus — reached OpenRouter on July 20 at $0.30/$1.20, three weeks after its actual June 30 launch. And poolside shipped Laguna S 2.1 on July 21: an open-weight 118B-total / 8B-active MoE at $0.10/$0.20, trained start-to-launch in under nine weeks on 4,096 H200 GPUs (vendor-stated), with weights in four quantizations on Hugging Face on day one under the OpenMDW-1.1 license — plus a free OpenRouter variant at $0/$0 with a 256K window.
The West's most capable open-weight model.— poolside, Laguna S 2.1 launch framing, July 21, 2026
That line is poolside's own framing — a vendor self-description, not an independent ranking. The disclosed numbers behind it are more modest and more credible than the slogan: 70.2% on Terminal-Bench 2.1 against Fable 5's 88.0% and GPT-5.6 Sol's 88.8%, and 78.5% on SWE-Bench Multilingual, roughly level with much larger models. poolside is the one lab in this wave that publishes its harness and full trajectories, which makes its self-reported figures the most auditable of the month even though they remain vendor-run. The full technical read is in our Laguna S 2.1 launch analysis.
Input price per 1M tokens · the July 2026 wave
Source: live OpenRouter models API, read at publication (July 2026)Run the arithmetic on a representative agentic task — 500K input tokens plus 100K output. On Laguna XS 2.1 that costs about $0.042 (500K × $0.06/M + 100K × $0.12/M). On GPT-5.6 Sol the same token volume costs about $5.50 (500K × $5/M + 100K × $30/M) — roughly 131x more. The June wave's equivalent spread was about 25x. The gap between the cheapest capable model and the most expensive frontier model widened by a multiple in one month, which makes routing discipline — not model loyalty — the thing that controls an AI budget.
06 — GoogleGoogle's mid-cycle reset, in one announcement.
Google compressed its entire July move into a single July 21 announcement: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a security-focused 3.5 Flash Cyber variant, shipped together. 3.6 Flash replaces 3.5 Flash as the Gemini app and API default at $1.50/$7.50 — with the output rate cut from a prior $9 — and is currently Google's strongest publicly available model, because Gemini 3.5 Pro remains unreleased and in partner testing. Google frames the release as an efficiency gain (roughly 17% fewer output tokens per task, per Artificial Analysis measurement it cites) rather than a frontier jump. Flash-Lite lists at $0.30/$2.50, streams around 350 tokens per second, and is explicitly positioned for subagents in multi-agent systems.
We keep this section short deliberately, because the per-model depth already exists: the full Gemini 3.6 Flash launch analysis, how it stacks up against GPT-5.6 and Kimi K3 on benchmarks, and Flash-Lite's subagent economics. The roundup-level takeaway is positioning: Google now covers the $0.30 and $1.50 rungs of the ladder with same-day OpenRouter listings, and it told the market plainly where the next jump comes from.
Most ambitious pre-training run yet.— Google, on Gemini 4 pre-training, July 21, 2026 announcement
That is Google confirming Gemini 4 pre-training is officially underway — in training only, with no release date — alongside a launch it explicitly framed as an efficiency play. Read together, the message is that the 3.x line is now Google's cost-performance ladder while the capability bet rides on 4. For teams routing production traffic, that makes 3.6 Flash a stable default to adopt rather than a stopgap to wait out.
07 — How To RouteWhat to route where, workload by workload.
A pricing table is only useful once it becomes a routing decision. The matrix below pairs each major workload class with the July models we would shortlist for it, grounded in the live pricing above and with vendor benchmark claims treated as claims. It follows the same discipline as our AI cost-optimization playbook: match the task to the cheapest model that clears your quality bar, and verify on your own prompts before committing a default.
Hard, high-stakes engineering work
GPT-5.6 Sol, Kimi K3, and Grok 4.5 are July's candidates, with incumbent Claude still leading several vendor-published agentic rows. Sol costs 2.5x Grok's input rate — benchmark both on your own repos before paying the premium.
High-volume workhorse traffic
Gemini 3.6 Flash at $1.50/$7.50 with a 1M window and an output-price cut is the strongest default-tier launch of the month, and Google's own efficiency framing matches the workhorse role. GPT-5.6 Luna at $1/$6 is the closest cross-vendor rival.
Multi-agent orchestration layers
Gemini 3.5 Flash-Lite was built and marketed for exactly this — $0.30 input, ~350 tok/s, 1M context for cheap parallel subagents under a stronger orchestrator. KAT-Coder Air at $0.15 covers the coding-specific fan-out case.
Classification, extraction, drafts
Laguna XS 2.1 at $0.06/$0.12 is the cheapest paid rung in the wave, and the free Laguna S 2.1 variant lets you eval at zero token cost before committing. LongCat 2.0 at $0.30/$1.20 adds a ~1M-context option for long-document batch runs.
Multi-model routing is also no longer exotic plumbing — OpenRouter's Fusion ensemble endpoint makes cross-model synthesis a config option rather than an engineering project. For teams that want the eval harness, routing policy, and spend governance built rather than described, that is the core of our AI transformation engagements.
08 — The ComparisonHow July compares to June.
Per our June roundup, that wave was a ten-day burst — five major models from May 27 to June 4, spanning a roughly 25x input-price range. July sustained nearly the same per-day cadence across a full month: ten families in 22 days, an ~83x input spread, and a four-model single day on July 21. The structural reading we made in June — that every price tier now carries a credible long-context option — hardened in July into something stronger: every price tier now has competition within it, from three models clustered at or below $0.30 input to three more contesting the $1–$2 workhorse band.
The absences are as telling as the arrivals. DeepSeek, Mistral, and the Meta-Llama open-weight line shipped nothing new on OpenRouter in this window — Meta's July move was the closed, API-only Muse Spark under a brand-new vendor prefix, a strategic U-turn from the lab that defined open weights. DeepSeek V4's GA status remained secondary-sourced-only as of July 21, with no primary confirmation. Meanwhile the catalog itself churns: a live pull at publication returned 343 models, a point-in-time figure that fluctuates weekly as free and paid variant pairs and deprecations come and go. For the pre-wave baseline, see our April 2026 OpenRouter rankings.
Projecting forward: if the July cadence holds even approximately, the second half of 2026 delivers a new production-relevant model roughly every week, and the shelf life of any "best model" claim compresses toward days. Kimi K3's promised July 27 weights release is the next scheduled event on that calendar. The durable posture is the one both roundups now point to — maintain a routing policy with named fallbacks per workload, and re-run the comparison monthly rather than re-litigating it per launch.
09 — ConclusionThe month the ladder got crowded.
Ten launches later, routing discipline matters more, not less.
July 2026 put ten new model families on OpenRouter in 22 days and stretched the paid price ladder from $0.06 to $5.00 per million input tokens — roughly 83x, with a free rung below it. Reading the month as one dataset surfaces what per-launch coverage misses: same-day listings are now the norm for major vendors, the open-weight lane keeps compressing the floor, and the listing lag on stragglers like LongCat 2.0 can misdate a launch by three weeks for anyone treating the catalog as a calendar.
The honest caveats travel with the numbers. Every benchmark cited this month is vendor-reported; poolside's disclosed harness is the closest thing to auditability, and Muse Spark 1.1's coding scores are actively contested. Prices and free variants change weekly, so every figure here is a point-in-time read from the live API, not a permanent fact. Verify before you route.
Our working call stands: Gemini 3.6 Flash for the production default, Kimi K3 and Grok 4.5 contesting the frontier band, and the specialist rungs — Flash-Lite, KAT-Coder Air, the Laguna line — covering fan-out and batch work. At one release every 2.2 days, that call has a shelf life measured in weeks, which is exactly why the policy, not the pick, is the asset worth maintaining.