Alibaba’s Token Plan is the most aggressively priced agent subscription any major AI vendor has shipped: $6 a month buys 2,500 credits per rolling 7-day window, 700 per rolling 5-hour window, and the right to run one or two concurrent coding agents against Qwen3.8-Max-Preview — plus GLM-5.2 and DeepSeek-V4-Pro, two models Alibaba does not even make.
The plan went up on Alibaba’s English-language pricing surfaces alongside the Qwen3.8-Max-Preview launch we covered in our launch analysis. That post handles the model. This one handles the money — because the packaging is genuinely novel, and in places genuinely opaque. Alibaba is not copying Anthropic’s or OpenAI’s subscription shape; it is tiering on how many agents you can run at once, metering in credits that map to no published token quantity, and bundling competitors’ models like an aggregator rather than a lab.
Below: the complete Individual and Team price sheet, how the rolling 7-day-plus-5-hour credit windows differ structurally from Anthropic’s and OpenAI’s caps, a discrepancy between Alibaba’s own two pricing pages that no other coverage we found has flagged, and a four-vendor comparison against Claude, ChatGPT/Codex, and Kimi. Every Token Plan figure here was pulled directly from Alibaba’s live pricing pages on July 21, 2026.
- 01Entry pricing starts at $6/mo — with tiers to $200/seat.Individual runs Lite $6, Standard $18, Pro $68 (list $8/$25/$80); Team seats run $20, $75, and $200 with 25,000 to 250,000 credits per month. Annual Individual plans are $65, $195, and $770.
- 02Agent concurrency, not model access, is the tier axis.Lite allows 1–2 concurrent agents, Standard 3–4, Pro 6–8. All tiers get the same model roster. Neither Anthropic nor OpenAI prices consumer tiers by how many agents run at once.
- 03Credits run on rolling 7-day + 5-hour windows.A genuinely different shape from Anthropic’s fixed weekly ceilings — though Alibaba’s own second pricing page undercuts the framing by describing the same tiers as flat monthly totals.
- 04No token math is published anywhere.Nothing on either pricing page says what one credit buys in tokens, per model. You cannot convert credits to tokens, or tokens to dollars, from Alibaba’s public materials as of July 21, 2026.
- 05It is an aggregator play, not just a Qwen plan.One subscription bundles Zhipu AI’s GLM-5.2 and DeepSeek-V4-Pro alongside the Qwen roster, exposed through OpenAI- and Anthropic-protocol endpoints that drop into Claude Code, Cursor, Codex, and five other named tools.
01 — The Price SheetEvery tier, both pages, one table.
Alibaba publishes the Token Plan across two properties: qwencloud.com’s pricing page carries the Individual tiers with their credit-window mechanics, and the alibabacloud.com campaign page carries the Team seats, annual pricing, and the bundled-model roster. No press coverage we found assembles both — most secondary write-ups mention only the Individual monthly prices. Here is the complete sheet, with the discount math recomputed from the stated list prices.
| Tier | Monthly (vs list) | Annual | Credits | Concurrent agents |
|---|---|---|---|---|
| Individual (qwencloud.com framing) | ||||
| Lite | $6 (list $8, −25%) | $65/yr (list $96) | 2,500 / rolling 7-day + 700 / rolling 5-hour | 1–2 |
| Standard | $18 (list $25, −28%) | $195/yr (list $300) | 10,000 / rolling 7-day + 3,000 / rolling 5-hour | 3–4 |
| Pro | $68 (list $80, −15%) | $770/yr (list $960) | 40,000 / rolling 7-day + 12,000 / rolling 5-hour | 6–8 |
| Team — per seat (alibabacloud.com campaign page) | ||||
| Team Standard | $20 (list $30, −33%) | — | 25,000 / month | Not stated |
| Team Pro | $75 (list $100, −25%) | — | 100,000 / month | Not stated |
| Team Max | $200 (no discount listed) | — | 250,000 / month | Not stated |
Three details in the fine print are worth surfacing. First, the Standard and Pro Individual tiers are labeled on the page as “4x the credits of Lite” and “16x the credits of Lite” — and the arithmetic checks out exactly (10,000 and 40,000 against Lite’s 2,500). Second, the Team ladder has a quiet mid-tier incentive: Team Pro works out to $0.75 per 1,000 credits versus $0.80 on both Team Standard and Team Max, matching the page’s own “6% savings” label — the top seat buys headroom, not a better rate. Third, annual Individual billing discounts against list by 32%, 35%, and 20% respectively — the mid tier, again, is where the sharpest deal sits.
02 — Credit WindowsTwo rolling windows, stacked.
The qwencloud.com framing meters every Individual tier against two simultaneous rolling windows: a 7-day cap and a 5-hour cap. Burn through the 5-hour allowance and you wait for it to roll; burn through the 7-day allowance and no amount of short-window patience helps. Both windows are described as rolling continuously rather than resetting on a fixed calendar day.
That is a structurally different shape from what Anthropic runs. Claude’s paid tiers meter a rolling 5-hour session window stacked with a weekly ceiling that behaves as a fixed allowance — the shape whose practical consequences we mapped in our Claude usage-credits pricing guide. A fixed weekly ceiling concentrates pain at the end of the week: once you hit it, you are done until reset. A genuinely rolling 7-day window instead releases allowance continuously as usage ages out of the trailing week — smoother for steady daily users, less forgiving for burst-then-rest patterns that a calendar reset would wipe clean overnight.
Individual-tier credits per rolling 7-day window
Source: qwencloud.com Token Plan pricing page, retrieved July 21, 2026One caveat belongs next to that analysis, and it comes from Alibaba itself: the company’s other pricing page describes these exact tiers as flat monthly totals with no window mechanics at all. We unpack that inconsistency in section 04 — but treat the rolling behavior as Alibaba’s stated design, not yet as independently verified metering behavior.
03 — The Tier AxisSelling concurrency, not capability.
The genuinely unusual commercial choice is what separates the tiers. It is not model access — every Individual tier gets the same roster, including Qwen3.8-Max-Preview. It is not context length or feature gates. It is how many agents you may run at the same time: 1–2 on Lite, 3–4 on Standard, 6–8 on Pro.
1–2 agents
One coding agent, maybe a second in parallel. The shape of a single developer running Claude Code or Qwen Code against the Token Plan endpoint on one project at a time.
3–4 agents
Parallel-agent territory: a main task plus background workers, or multi-repo work. 4x Lite’s credits at 3x the price — the per-credit rate improves as you move up.
6–8 agents
Fleet mode: subagent swarms, long autonomous runs, several concurrent sessions. 16x Lite’s credits. This is the tier aimed at how agentic workflows actually run in mid-2026.
Read as a bet, the concurrency axis says something specific about where Alibaba thinks agentic usage is going. Anthropic and OpenAI price consumer tiers by volume multiplier — 5x or 20x of a base allowance — which assumes the scarce resource is total tokens. Alibaba’s bands assume the scarce resource is parallelism: the moment a developer stops babysitting one agent and starts orchestrating several, the constraint that bites first is how many can run at once, not how much any single one consumes. That maps cleanly onto how multi-agent workflows have actually evolved this year — and it gives Alibaba a natural upsell trigger that fires precisely when a user’s workflow matures.
What the pages do not define is what counts as one “concurrent agent” — a session, an API key, a stream of parallel requests through the compatible endpoint? For a Claude Code or Cursor user routing through the Token Plan base URL, the honest answer as of July 21 is that you find out by hitting the limit. The bands are also ranges (1–2, 3–4, 6–8) rather than hard integers, which itself goes undefined.
04 — The Opacity ProblemCredits with no token math.
Here is the number you will not find anywhere in this article, because it does not exist in public: how many tokens one credit buys. Neither Alibaba page publishes a credit-to-token conversion for any bundled model, a per-model credit-consumption rate, or a dollar-equivalent for a credit. The subscription is priced in a currency whose exchange rate is unpublished. For the broader industry pattern of subscriptions colliding with usage meters — and why vendors keep landing on abstractions like credits — see our subscriptions-versus-usage-credits explainer; this post stays on what Alibaba’s specific implementation does and does not disclose.
The opacity is compounded by the discrepancy we could not find flagged anywhere else: Alibaba’s own two pricing pages describe the same product with two different mental models.
To be fair to Alibaba: no major vendor publishes full subscription-side token math. Anthropic prices API tokens to the cent but has never published what its Pro or Max weekly allowances equal in tokens. The difference is that Alibaba invented a new named unit and then declined to define it on either of the pages selling it — and its cheapest competitors-included bundle makes the undefined unit do even more work, since one credit presumably buys different amounts of Qwen3.8, GLM-5.2, and DeepSeek-V4-Pro inference. Presumably. The pages do not say.
05 — The BundleAn aggregator wearing a lab’s badge.
The Token Plan’s model roster is the second structural surprise. Alongside Alibaba’s first-party lineup — Qwen3.8-Max-Preview, Qwen3.7-Max, Qwen3.7-Plus, Qwen3.6-Flash, and Text-Embedding-V4 — the campaign page bundles Zhipu AI’s GLM-5.2 and DeepSeek-V4-Pro. Those are competitors’ flagship-class models, included in the same subscription rather than sold separately. Neither Anthropic nor OpenAI resells a rival’s model inside its own consumer plan; Alibaba is behaving less like a lab defending its models and more like a cloud aggregating whatever Chinese frontier inference its customers might want, under one bill.
Qwen + Zhipu AI + DeepSeek
GLM-5.2 and DeepSeek-V4-Pro ship inside the Qwen-branded plan. Audio (Qwen-Audio-3.0-TTS-Plus, Fun-ASR), image (Wan2.7-Image), and video (HappyHorse1.1) models round out the roster.
Drop-in compatibility row
qwencloud.com lists Qwen Code, Cline, Claude Code, Cursor, OpenCode, Codex, Kilo CLI, and OpenClaw as supported tools — pointing rivals’ own clients at Alibaba’s endpoint.
Base URL + API key
Per the FAQ: subscribe, grab a key from the API Keys page, paste the base URL into any tool speaking the OpenAI or Anthropic protocol. No SDK, no plugin, no migration.
The compatibility row deserves a second look, because it is a strategy in eight names. Qwen Code is Alibaba’s own CLI; the other seven are ecosystem tools built around OpenAI’s and Anthropic’s APIs — including Claude Code and Codex themselves. Alibaba is not asking developers to adopt new tooling. It is asking them to change two config lines and keep their existing workflow, with the switch priced at $6 to find out. The protocol-compatibility layer turns every popular agent harness into a distribution channel for Alibaba’s bundle — and makes the marginal cost of trying (or leaving) close to zero, which cuts both ways.
06 — Head to HeadAgainst Claude, Codex, and Kimi.
The table below lines up all four vendors’ consumer-tier ladders against the same structural axes: price, what actually caps usage, the window shape, and whether third-party models ride along. Qwen figures are primary-fetched from Alibaba’s pages on July 21, 2026. Claude Max and ChatGPT Pro dollar figures (marked *) are as listed at the time of writing, corroborated across pricing coverage rather than re-verified character-for-character on partially script-rendered pricing pages — treat them as reference points, not gospel.
| Tier | Price / month | What caps usage | Window shape | 3rd-party models |
|---|---|---|---|---|
| Qwen Token Plan (Alibaba) — Individual | ||||
| Lite | $6/mo | Credits + 1–2 concurrent agents | Rolling 7-day + rolling 5-hour | Yes — GLM-5.2, DeepSeek-V4-Pro |
| Standard | $18/mo | Credits + 3–4 concurrent agents | Rolling 7-day + rolling 5-hour | Yes — same bundle |
| Pro | $68/mo | Credits + 6–8 concurrent agents | Rolling 7-day + rolling 5-hour | Yes — same bundle |
| Claude (Anthropic) | ||||
| Pro | $20/mo | Usage allowance (base) | 5-hour session + weekly ceiling | No |
| Max 5x | $100/mo* | 5x Pro usage multiplier | 5-hour session + weekly ceiling | No |
| Max 20x | $200/mo* | 20x Pro usage multiplier | 5-hour session + weekly ceiling | No |
| ChatGPT / Codex (OpenAI) | ||||
| Plus | $20/mo* | Usage allowance (base); Codex bundled | Rate limits per tier | No |
| Pro 5x / 20x | $100–$200/mo* | 5x / 20x Plus limits; Codex bundled | Rate limits per tier | No |
| Kimi (Moonshot AI) | ||||
| Moderato | $15/mo (list $19) | Model-access tier + 1x agent credits | Membership allowance | No |
| Allegro | $79/mo (list $99) | Model-access tier + 5x agent credits | Membership allowance | No |
| Vivace | $159/mo (list $199) | Model-access tier + 10x agent credits | Membership allowance | No |
Three readings. On raw entry price, nothing touches $6 — Kimi’s Moderato at $15 is the nearest, and Claude and ChatGPT both start at $20. On seats, Qwen’s Team Standard at $20 undercuts Claude’s Team seats ($25 standard, $125 premium, as listed at the time of writing) while attaching a stated 25,000-credit monthly pool — though with an undefined unit on one side of the comparison, the per-dollar-of-actual-inference winner is unknowable, which is rather the point. And on packaging philosophy, the column that separates the vendors most cleanly is the last one: only Alibaba bundles someone else’s frontier model.
The timing sharpens the contrast. The same weekend this plan appeared, Anthropic raised Claude Code’s weekly limits by 50% and added Fable 5 access to Max and Team Premium — moves we verified in our Claude Code limits coverage. Anthropic’s response to agent-era demand is more of the same allowance, on the same weekly-ceiling shape. Alibaba’s is a different shape entirely. Meanwhile Moonshot’s Kimi — whose K3 launch two days before Qwen3.8’s debut made it the same week’s other Chinese frontier story — tiers its memberships by model access plus agent-credit multipliers, closer to the Western pattern than to Alibaba’s concurrency bands.
07 — The Model Behind ItA preview is doing the selling.
Everything above is packaging; the headline product inside it is Qwen3.8-Max-Preview, which debuted July 19 at the World AI Conference in Shanghai — vendor-stated at 2.4 trillion parameters, with the active-parameter count undisclosed, no model card, no official benchmark table, and no license published as of July 21. One naming trap is worth stating plainly: Qwen3.8 is not Qwen3-8B, the unrelated 8-billion-parameter dense model from 2025 — the similarity of the names is an accident of versioning. Our launch analysis covers the model in full; what matters for the pricing story is that every capability claim currently attached to it is Alibaba’s self-assessment or early-user anecdote, because there is nothing independent to check yet. Context-window figures circulating for the model conflict across unofficial sources and remain unconfirmed, so we omit them.
"During Preview, Qwen3.8 is getting better by the day. Latest version is live now, with broad gains and a big step up on web frontend."— Qwen’s official account, July 20, 2026, as embedded by TechnoSports (X itself is not directly fetchable)
We tried the preview ourselves on consecutive days — July 20 and July 21 — and day-over-day improvement was not perceptible in our brief testing. That is not a gotcha: daily-update claims are genuinely hard for any single user to verify, and that is the honest point about preview-phase models. A model that is “still evolving daily,” in Qwen’s own words, is also a model whose behavior under your workload next week is not guaranteed to match this week’s — worth weighing before routing production agents through it, however cheap the subscription.
The open question hanging over the whole bundle is the open-weight promise. Qwen says Qwen3.8 is going open-weight, with no date confirmed — which would reverse the closed-model posture Alibaba has held since the Qwen3.6-Max-Preview pivot, through a Qwen3.7-Max that also stayed API-only despite speculation. Until weights actually land, the Token Plan is the only way to run the flagship — which is precisely what makes the $6 entry price a rational acquisition spend for Alibaba.
One more pricing signal hides on the Chinese-facing platform: platform.qianwenai.com carries a banner offering night-time calls from 20% of standard pricing — an off-peak discount that appears nowhere on the English-language pages. It echoes, from the opposite direction, the peak-hour surcharge DeepSeek has reportedly announced for its own API — announced rather than live, as of this writing. Two Chinese vendors moving toward time-of-day pricing in the same week suggests where constrained inference capacity is actually binding — and time-of-day metering may be the next pricing surface Western vendors are forced to consider.
08 — Decision GuideWhat to do with a $6 question.
The practical question is not whether the Token Plan is a good product — nobody outside Alibaba can answer that yet, given the credit opacity and the preview-phase model. The question is whether it is worth $6 to generate your own answer. By workload:
Curious, price-sensitive
Lite with the $2 new-user coupon is an effectively-$4 first month. Point Qwen Code or an existing CLI at the endpoint, run a week of real tasks, and measure how far 2,500 credits actually go on your work — the only credit-to-token data that exists is what you generate yourself.
Parallel orchestration workflows
The 6–8 agent Pro band at $68 is the only consumer tier from any vendor priced explicitly for concurrency. If weekly ceilings on Western plans are your bottleneck, this is the cheapest structured experiment available.
Happy with current stack
The drop-in endpoint means switching cost is two config lines — in either direction. Run a shadow evaluation on non-critical tasks if curious, but a preview-phase model with undefined credits is not a reason to migrate a working production setup.
Compliance-bound workloads
No model card, no license, no published benchmarks, a model updating daily, and a metering unit with no public definition. Every one of those is individually disqualifying for governed production use today. Revisit if weights and documentation land.
For teams treating this as a live routing question — which models and plans to put behind which workloads, and how to benchmark an opaque credit system against a known one — this is exactly the comparative-evaluation work our AI transformation engagements start with: instrument your real tasks, measure spend per outcome across vendors, and decide on evidence rather than on a pricing page.
09 — ConclusionPricing is the product now.
Alibaba changed the axis, not just the price.
The cheap headline number is the least interesting thing about the Token Plan. What is actually new: tiers built on agent concurrency instead of usage multipliers, dual rolling windows instead of weekly ceilings, competitors’ models bundled into one subscription, and protocol compatibility that turns rivals’ own tools into distribution. Each choice reads like a bet that the agent era rewards a different pricing geometry than the chat era did.
The honest caveats are equally structural. Credits have no published token math; Alibaba’s own two pricing pages describe the metering differently; and the flagship model inside the bundle is a preview with no model card, no benchmarks, and daily-moving behavior. A $6 subscription whose unit of account is undefined is cheap in dollars and unpriceable in inference — both facts are true simultaneously, and both belong in any evaluation.
Looking forward, the aggregator pattern is the piece we expect to travel. If bundling GLM-5.2 and DeepSeek-V4-Pro under one Qwen bill wins developers, pressure grows on every cloud with distribution to sell rivals’ inference the same way — and the competitive unit stops being the model and becomes the bundle, the endpoint, and the meter. Watch whether the promised open weights land, whether anyone reverse-engineers the credit math, and whether a Western vendor answers with a concurrency-priced tier of its own.