Chinese models accounted for a majority of token usage in the OpenRouter figures reported by CNBC on September 26. The finding is significant for teams choosing models, but it describes traffic through a gateway rather than the entire AI market. Keep the platform, period and unit attached whenever you cite it.
Editorial note: Prepared October 1 from information published by September 26, 2026. Later product developments are outside this article’s scope.
- 01Token share is a volume measureIt does not directly measure revenue, customer count or the quality of completed work.
- 02The periods differDo not combine weekly OpenRouter figures with monthly Vercel figures into one market estimate.
- 03Popularity earns an evaluationA widely used model still needs to pass your task, data-handling and cost checks.
01 — The evidenceThe reported figures, with their limits
CNBC’s September 26 report gives Chinese models 57–67% of OpenRouter tokens in the week of September 14, compared with 6–13% in February. The figures cover companies in the United States, Europe and OpenRouter’s “Global South” grouping. CNBC separately reports Vercel’s share reaching 55% in August from 11% in January; its geographic breakdown is unspecified. These are platform-supplied figures reported by CNBC, not measurements collected by Digital Applied.
| Platform | Earlier share | Later share | Reporting period |
|---|---|---|---|
| OpenRouter | 6–13% | 57–67% | February → week of September 14 |
| Vercel | 11% | 55% | January → August |
The published range should remain a range. Without the underlying group weights, a midpoint would be an invented aggregate. The two platforms also cannot simply be averaged: the larger traffic population would need a different weight, and the observation windows would need to match.
02 — Practical implicationsWhat a token share can and cannot tell you
A token is a unit of text processed by a model. A workload that sends long documents or runs repeated agent steps can consume more tokens than many short conversations. A change in workload mix can therefore move the volume share without a matching change in customer count. That is a reason to label the measure precisely, not a reason to dismiss it.
How much text is processed?
Useful for understanding the traffic mix inside a defined platform and time window.
How much money is paid?
Requires price and billing data, including the relevant input, output and cache treatment.
How much useful work finishes?
Requires a task definition, acceptance test and a record of failed or repeated work.
A hypothetical example shows the distinction. Suppose two models each receive half the tokens, but one costs twice as much per token under an otherwise identical workload. Their volume shares are equal while their spending shares are not. Real billing is more complicated because input, output and cached tokens can have different prices, which strengthens the need to keep the units separate.
Our earlier Chinese-model landscape article addresses the provider context. This report concerns gateway usage; neither should be cited as a census of every enterprise deployment, direct API customer or self-hosted installation.
03 — Practical implicationsWhy there is no reconstructed historical ranking
We have not independently reproduced the September usage shares. A live rankings page observed later cannot establish what it showed during the reported week. To recreate that measurement, a researcher would need a dated export or archived snapshot, the grouping rules, the reporting window and a way to reconcile every included row with the stated denominator.
The table is a compact summary of reported figures. It is not an original dataset, a benchmark or a measured estimate of global AI market share.
An honest follow-up measurement would choose its own observation period and state it explicitly. It would record whether the unit is input tokens, output tokens or both, how cached tokens are handled, which requests are included, and how model families are assigned. A label such as “Chinese model” also needs a definition: model developer, serving provider and location of processing are separate properties.
Until those details are available, do not turn platform rankings into a national-market estimate. The useful statement is narrower and still worth knowing: a particular usage report shows a substantial shift within the systems it covers.
04 — Practical implicationsUse adoption as a reason to test a candidate
For a buyer, adoption can help identify a model worth evaluating. It cannot answer whether that model follows your output contract, handles your documents well or meets the rules for the data you plan to send. Build a small task set from work the team actually performs and define the pass criteria before seeing the outputs.
| Decision | Evidence to request |
|---|---|
| Quality | Accepted outputs on representative tasks, including difficult cases. |
| Economics | Total cost including retries, review and fallback calls. |
| Data handling | The specific serving provider, processing terms and permitted data. |
| Reliability | Observed failures, latency and fallback behavior for the workload. |
Keep the task and acceptance criteria fixed when comparing providers. If one configuration gets extra retries or a human correction step, record that as part of its cost. Our model-routing guide explains how to connect those checks to routing decisions.
A practical next step is to identify one lower-risk workload where a different model could compete on accepted-result cost. Route only that workload after evaluation, with a fallback for failures. The permission boundary should remain the same when the model changes.
05 — Practical implicationsKeep the metric attached to the claim
When presenting this finding internally, put the platform and period in the same sentence as the percentage. Keep the source nearby and distinguish the reported observation from your recommendation. “This candidate deserves a trial” is a defensible purchasing response; “this proves it is the best model” requires evidence this report does not supply.
Teams building a multi-model workflow can use our AI transformation service to define the tests, data boundary and fallback behavior before changing production routing.
Evaluate the model and preserve the denominator
The reported shift deserves attention without being enlarged into a whole-market claim. Use it to choose candidates for a controlled evaluation, and make the final decision on accepted work, total cost and the data terms of the actual provider.