AI DevelopmentMethodology5 min readPublished September 26, 2026

Gateway traffic is useful evidence when its denominator stays attached

Chinese AI Models Lead Reported OpenRouter Token Usage

Chinese models reached 57–67% of OpenRouter tokens in reported regional groups. What CNBC’s September figures measure, their limits and what buyers can learn.

DA
Digital Applied Team
Research and practical guidance
CoverageThrough September 26, 2026

Chinese models accounted for a majority of token usage in the OpenRouter figures reported by CNBC on September 26. The finding is significant for teams choosing models, but it describes traffic through a gateway rather than the entire AI market. Keep the platform, period and unit attached whenever you cite it.

Editorial note: Prepared October 1 from information published by September 26, 2026. Later product developments are outside this article’s scope.

Key takeaways
  1. 01
    Token share is a volume measureIt does not directly measure revenue, customer count or the quality of completed work.
  2. 02
    The periods differDo not combine weekly OpenRouter figures with monthly Vercel figures into one market estimate.
  3. 03
    Popularity earns an evaluationA widely used model still needs to pass your task, data-handling and cost checks.

01 — The evidenceThe reported figures, with their limits

CNBC’s September 26 report gives Chinese models 57–67% of OpenRouter tokens in the week of September 14, compared with 6–13% in February. The figures cover companies in the United States, Europe and OpenRouter’s “Global South” grouping. CNBC separately reports Vercel’s share reaching 55% in August from 11% in January; its geographic breakdown is unspecified. These are platform-supplied figures reported by CNBC, not measurements collected by Digital Applied.

Source: CNBC, September 26, 2026, reporting data supplied by the platforms. Different populations and periods; no combined estimate.
PlatformEarlier shareLater shareReporting period
OpenRouter6–13%57–67%February → week of September 14
Vercel11%55%January → August

The published range should remain a range. Without the underlying group weights, a midpoint would be an invented aggregate. The two platforms also cannot simply be averaged: the larger traffic population would need a different weight, and the observation windows would need to match.

02 — Practical implicationsWhat a token share can and cannot tell you

A token is a unit of text processed by a model. A workload that sends long documents or runs repeated agent steps can consume more tokens than many short conversations. A change in workload mix can therefore move the volume share without a matching change in customer count. That is a reason to label the measure precisely, not a reason to dismiss it.

Volume
How much text is processed?
Token share

Useful for understanding the traffic mix inside a defined platform and time window.

Reported unit
Spending
How much money is paid?
Revenue share

Requires price and billing data, including the relevant input, output and cache treatment.

Different measure
Outcomes
How much useful work finishes?
Accepted tasks

Requires a task definition, acceptance test and a record of failed or repeated work.

Buyer measure

A hypothetical example shows the distinction. Suppose two models each receive half the tokens, but one costs twice as much per token under an otherwise identical workload. Their volume shares are equal while their spending shares are not. Real billing is more complicated because input, output and cached tokens can have different prices, which strengthens the need to keep the units separate.

Our earlier Chinese-model landscape article addresses the provider context. This report concerns gateway usage; neither should be cited as a census of every enterprise deployment, direct API customer or self-hosted installation.

03 — Practical implicationsWhy there is no reconstructed historical ranking

We have not independently reproduced the September usage shares. A live rankings page observed later cannot establish what it showed during the reported week. To recreate that measurement, a researcher would need a dated export or archived snapshot, the grouping rules, the reporting window and a way to reconcile every included row with the stated denominator.

Evidence boundary

The table is a compact summary of reported figures. It is not an original dataset, a benchmark or a measured estimate of global AI market share.

An honest follow-up measurement would choose its own observation period and state it explicitly. It would record whether the unit is input tokens, output tokens or both, how cached tokens are handled, which requests are included, and how model families are assigned. A label such as “Chinese model” also needs a definition: model developer, serving provider and location of processing are separate properties.

Until those details are available, do not turn platform rankings into a national-market estimate. The useful statement is narrower and still worth knowing: a particular usage report shows a substantial shift within the systems it covers.

04 — Practical implicationsUse adoption as a reason to test a candidate

For a buyer, adoption can help identify a model worth evaluating. It cannot answer whether that model follows your output contract, handles your documents well or meets the rules for the data you plan to send. Build a small task set from work the team actually performs and define the pass criteria before seeing the outputs.

Digital Applied purchasing checklist; adoption statistics alone do not establish these properties.
DecisionEvidence to request
QualityAccepted outputs on representative tasks, including difficult cases.
EconomicsTotal cost including retries, review and fallback calls.
Data handlingThe specific serving provider, processing terms and permitted data.
ReliabilityObserved failures, latency and fallback behavior for the workload.

Keep the task and acceptance criteria fixed when comparing providers. If one configuration gets extra retries or a human correction step, record that as part of its cost. Our model-routing guide explains how to connect those checks to routing decisions.

A practical next step is to identify one lower-risk workload where a different model could compete on accepted-result cost. Route only that workload after evaluation, with a fallback for failures. The permission boundary should remain the same when the model changes.

05 — Practical implicationsKeep the metric attached to the claim

When presenting this finding internally, put the platform and period in the same sentence as the percentage. Keep the source nearby and distinguish the reported observation from your recommendation. “This candidate deserves a trial” is a defensible purchasing response; “this proves it is the best model” requires evidence this report does not supply.

Teams building a multi-model workflow can use our AI transformation service to define the tests, data boundary and fallback behavior before changing production routing.

Next step

Evaluate the model and preserve the denominator

The reported shift deserves attention without being enlarged into a whole-market claim. Use it to choose candidates for a controlled evaluation, and make the final decision on accepted work, total cost and the data terms of the actual provider.

Agentic AI implementation

Build a workflow you can evaluate and control

Digital Applied helps teams connect AI capabilities to useful work, clear acceptance checks and responsible operating limits.

Task evaluationsCost visibilityControlled access
Start with one task

Define the pilot

  • →Approved source material
  • →A named reviewer
  • →A clear acceptance check
  • →Spending and permission limits
Questions and answers

Practical questions

No. A token-share statistic describes usage volume within its stated population. It does not include every channel through which AI is used.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Qwen3.7 Flash: A Cheap Multimodal Tier for Subagents

Qwen3.7 Flash listed at $0.03 per million input tokens on its base tier, 1M context and video input — but no published benchmarks. Where it fits in subagents.

July 28, 2026 · 10 minRead
AI Development

GLM 4.6 API Deployment Guide: Local & Cloud Setup

Deploy GLM 4.6 with current Z.ai, OpenRouter, vLLM, and SGLang guidance covering endpoints, pricing, MIT licensing, model IDs, and production caveats.

October 14, 2025 · 9 minRead
AI Development

Gemini 4 Argon: Price, Benchmarks and When You Can Use It

Gemini 4 Argon launches at $2/$10 per million tokens, rising to $4/$20 later. Who can use it now, what Google's benchmarks show and how to prepare.

September 30, 2026 · 6 minRead
AI Development

GPT-6.1 Sol: Pricing, Benchmarks and the Upgrade Case

GPT-6.1 Sol approaches Astra on key agent tasks at $2/$10 per million tokens. Compare pricing, cache costs, benchmark limits and the case for switching.

September 29, 2026 · 5 minRead
AI Development

Model Aliases and Retirements in 2026: The Full Ledger

Twenty-five floating aliases across five vendor surfaces and 92 dated 2026 model-ID retirements, with notice periods computed from each vendor's own dates.

August 22, 2026 · 20 minRead
AI Development

Do Not Single-Source Your AI: A Second-Source Playbook

The Fable 5 export shutdown showed single-vendor AI can halt your business overnight. A four-step second-source playbook with open-weight failover backups.

June 21, 2026 · 11 minRead
Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source