BusinessDecision Matrix13 min readPublished July 26, 2026

One combined number · 3 undefined caveats · the CFO’s own counter-metric

OpenAI Says 10 Million Agent Users. What Does That Mean?

Bloomberg reported on July 21 that OpenAI’s agent products — Codex and ChatGPT Work combined — reached 10 million users, roughly double the count from two weeks earlier. The growth is real. But the figure carries three caveats OpenAI has not resolved, and four days before the milestone, OpenAI’s own CFO published a framework arguing that usage counts are the wrong way to measure AI value.

DA
Digital Applied Team
Senior strategists · Published Jul 26, 2026
PublishedJuly 26, 2026
Read time13 min
SourcesBloomberg · TNW · OpenAI · Axios
Combined agent users
10M
Codex + ChatGPT Work · Jul 21
~2x in two weeks
Codex weekly users
~8M
mid-July · OpenAI-posted
+3M since late May
ChatGPT weekly users
900M+
reported scale context
Independent audits
0
company estimate, unaudited

OpenAI’s 10 million agent users milestone, reported by Bloomberg on July 21, 2026, is the largest agent-adoption figure any vendor has published that we could find. It is also a combined tally across two different products, with no product split, no stated definition of “active,” and no independent audit. Reading it correctly — neither dismissing it nor repeating it uncritically — is a skill every AI buyer now needs.

The stakes are practical, not academic. Adoption milestones like this one flow directly into board decks, procurement arguments, and “everyone is already doing this” pressure on budget owners. Meanwhile the same week produced the counter-current: OpenAI CFO Sarah Friar published a framework arguing that seats and usage are the wrong measures of AI value, and Axios reported that executives are increasingly routing everyday work to the cheapest model that can do the job. Adoption headlines and value measurement have split.

This guide covers what was actually reported and by whom, the three caveats that define the figure’s limits, the verified Codex trajectory underneath it, a claim-by-claim verification table for the whole news cycle, Friar’s "Useful Intelligence per Dollar" framework and its self-serving tension, and a buyer’s scorecard for interrogating any vendor adoption number — this one included.

Key takeaways
  1. 01
    10 million is a combined, two-product figure.Bloomberg reported OpenAI’s agent products — Codex and ChatGPT Work together — reached 10 million users on July 21, roughly double the count from two weeks earlier. OpenAI does not break out the two products.
  2. 02
    Three caveats define the number’s limits.No product split; no stated definition of whether ‘active’ means daily, weekly, or point-in-time; and it is a company estimate rather than an audited metric. None of the three is a scandal — all three are things a buyer should notice.
  3. 03
    The verified trajectory is Codex-specific.OpenAI-executive-posted thresholds put Codex at 5 million weekly users at the end of May, 6 million on July 12, and roughly 8 million by mid-July. The 10 million headline is the combined tally, not a Codex-only figure.
  4. 04
    OpenAI’s own CFO proposed a competing yardstick.Four days before the milestone, Sarah Friar’s ‘A scorecard for the AI age’ argued for measuring cost per successful task instead of seats or tokens — while Axios noted OpenAI’s incentive: its models are among the most expensive.
  5. 05
    Buyers are already voting with routing tables.Executives told Axios they route everyday tasks to the cheapest model that can do the job, reserving frontier models for the hardest work. The practical response to vendor metrics is instrumentation, not belief.

01The ReportWhat Bloomberg actually reported.

The primary claim comes from Bloomberg’s Shirin Ghaffary on July 21, 2026: OpenAI’s agent products — the Codex coding agent and the ChatGPT Work agent platform combined — reached 10 million users, roughly double the count from two weeks earlier. The Next Web and a Yahoo Finance syndication of the Bloomberg piece corroborated the figure the same day. ChatGPT Work itself launched on July 9 for Pro, Enterprise, and Edu plans and reached Plus and Business within days — we covered ChatGPT Work’s own launch separately, so this post treats it only as one half of the combined tally.

The scale context matters in both directions. Against ChatGPT’s reported base of over 900 million weekly users, 10 million agent users is a small fraction — this is an early market, not a saturated one. Against the agentic-tools category itself, it is the largest adoption claim any vendor has published that we could find, and it landed amid a visible competitive push: the same syndicated reporting names Anthropic’s Cowork as the direct competitive backdrop for OpenAI’s agent expansion.

Two supporting figures from the same reporting cycle are worth holding onto. More than 1 million Codex users now apply it to work outside software development — a figure OpenAI itself stated in its July 9 launch post. And agentic-tool usage reportedly rose about 2.5 times in the week after GPT-5.6 shipped on July 9 — that one is reported via TNW rather than published by OpenAI. One is vendor-stated and one is reported, but both are at least specific and dated.

Why this number travels
Adoption milestones are the most portable artifact in enterprise AI sales. A single figure like 10 million users moves from a press report into analyst notes, then into vendor decks, then into your CFO’s inbox — usually shedding its caveats at every hop. The figure OpenAI disclosed is legitimate corporate practice; the version that reaches a budget meeting three hops later often is not.

02The CaveatsThree things the number does not say.

These three caveats are the spine of this story, not footnotes to it. Each one is checkable against the primary reporting, and each one changes what conclusions the number can support.

1. It is a combined figure with no product split

OpenAI does not break out how many of the 10 million are Codex users and how many are ChatGPT Work users. That matters because the two products have entirely different buyer profiles — one is an established developer tool, the other a twelve-day-old knowledge-work product at the time of the report. A combined number can’t tell you whether the newer product is landing.

2. “Active” is undefined

OpenAI has not stated whether the figure counts daily, weekly, or point-in-time users. The Next Web’s own reading is that it refers to weekly active users of the two products combined — not enterprise seats, not paying customers, and not the number of agents actually run. A weekly-active figure is a legitimate metric; an unlabeled one forces every reader to guess.

3. It is a company estimate, not an audited metric

No third party has verified the count, and TNW notes that OpenAI’s usage claims have drawn scepticism before, including over the self-reported figures behind its internal adoption research. TNW’s editorial framing goes further: the milestone reveals little about how many of the 10 million sit on paid plans or how often they return once the novelty of a new model fades. That framing is the outlet’s, not an OpenAI disclosure — but the underlying absence of paid-share and retention data is a fact.

"For now the number reads as a marker of momentum, self-reported and taken mid-race, rather than proof of a market OpenAI has already won."— The Next Web, July 21, 2026

None of this is an accusation. Disclosing a combined, company-estimated figure is normal corporate practice — every platform vendor does it, and the underlying growth here is real by every independent signal available. The point is narrower: a number’s usefulness to a buyer depends on its definition, and this one ships without three of the definitions that matter most.

03Verified TrajectoryThe Codex curve underneath the combined tally.

Strip the combined figure away and a cleaner, better-sourced series remains. OpenAI executives have posted Codex weekly-user thresholds publicly as they were crossed, and The Next Web collated them: 5 million weekly users at the end of May, 6 million on July 12, and roughly 8 million by mid-July. That is about 3 million added weekly users in roughly six weeks — with the steepest stretch coming after GPT-5.6 shipped on July 9. OpenAI’s own July 9 baseline, stated in its launch post, matches the series: more than 5 million weekly Codex users, of whom more than 1 million use it for non-software work.

OpenAI agent adoption · verified Codex series vs the combined headline

Sources: OpenAI executive posts collated by The Next Web; Bloomberg, July 21, 2026
Codex · end of May5M weekly users · OpenAI-executive-posted
5M
Codex · July 126M weekly users
6M
Codex · mid-July~8M weekly users · steepest stretch after GPT-5.6 shipped
~8M
Combined · July 21Codex + ChatGPT Work together · NOT a Codex-only figure
10M

One arithmetic point is worth stating plainly, because a post about metric literacy should not paper over it: the two published series do not reconcile. Doubling backwards from the July 21 combined figure implies roughly 5 million combined in early July, which sits below the Codex-only readings posted in the days that followed. Both series come from the sources as published, and OpenAI has not reconciled them.

One trap to avoid explicitly: this trajectory is sometimes misread as treating 10 million as a Codex-only figure. That reading contradicts the corroborated series above and the combined framing in the original report — the 10 million is Codex plus ChatGPT Work, full stop. If you see the milestone cited as a single-product number, the citation is wrong.

Can you derive the split yourself? Not cleanly. Subtracting the mid-July Codex reading from the July 21 combined figure looks tempting, but the dates don’t align, Codex kept growing between the two readings, and OpenAI has published no split. Treat any derived product split as a guess, and treat anyone presenting one as a fact with appropriate suspicion.

04Claim-by-ClaimThe metric stack: every claim, one verification standard.

No single piece of coverage lines up the whole news cycle’s claims against a consistent standard — most outlets take the 10 million figure at face value, and none connect it to the value-metric post OpenAI published four days earlier. The table below applies one standard to every claim: how it’s defined, whether anyone independent has audited it, and what it structurally cannot tell you. It pairs naturally with our guide to what counts as agent-washing, which applies the same discipline to capability claims.

Verification table for every claim in the OpenAI 10 million agent users news cycle, showing how each figure is defined, whether it is independently audited, and what it does not tell a buyer.
ClaimHow it’s definedIndependently audited?What it doesn’t tell you
Adoption claims
10M combined agent users (Jul 21)Undefined activity window; TNW reads it as weekly active across both products combinedNo — company estimateProduct split, paid share, retention
Codex 5M → 6M → ~8M weeklyWeekly active users, OpenAI-executive-posted thresholds with datesNo — vendor-postedPaid vs. free mix, depth of use per user
>1M Codex users on non-software workWeekly, per OpenAI’s July 9 launch postNo — vendor-statedWhich tasks, at what success rate
~2.5x agentic-tool usage jump after GPT-5.6“Usage,” baseline unspecified; the week after July 9No — reported via TNWWhether the launch-week spike retained
Value & efficiency claims (from OpenAI’s Jul 17 scorecard post)
GPT-5.6 Sol: “54% fewer output tokens than another leading model”Vendor benchmark on the Artificial Analysis Coding Agent Index; the rival is unnamed by OpenAINo — self-reportedWhich model, which harness, which task mix
DeepSWE v1.1: Sol 72.7% vs. Claude Fable 5’s 69.9%, at 36.2% lower estimated API costVendor self-graded chart inside the post proposing OpenAI’s own value frameworkNo — no independent corroboration foundReplication, cost-estimate assumptions, variance

The pattern in the third column is uniform: six claims, zero independent audits. Again — that is not unusual for a fast-moving product cycle, and vendor-posted milestones with dates are still far better than undated ones. But a buyer who treats any row of this table as settled fact is doing the vendor’s marketing for free.

05The ScorecardFour days earlier, OpenAI’s CFO said stop counting users.

Here is the twist that makes this cycle worth a full post. On July 17 — four days before the 10 million milestone hit the press — OpenAI CFO Sarah Friar published “A scorecard for the AI age,” proposing a metric she calls Useful Intelligence per Dollar. Her argument: seats, tokens, and usage counts don’t measure whether AI is producing value; what matters is the cost of each successful task — the full cost of completing the work divided by the number of tasks that met the quality bar. The cheapest per-token model, she argues, isn’t always cheapest per outcome.

"The basic economic question facing CFOs and other business leaders is whether the value of the work AI completes grows faster than the cost of producing it."— Sarah Friar, CFO, OpenAI · ‘A scorecard for the AI age,’ July 17, 2026
Question 1
Work that matters
Is AI completing work that matters?

Not logins, not prompts — completed tasks with business consequence. Friar’s worked example is a finance team’s forecast-review workflow: finding the latest forecast, moving data into Excel or Sheets, identifying changes, reconciling tabs, rebuilding slides.

Measure completions, not seats
Question 2
Cost per successful task
What does each successful task cost?

Full cost of completing the work divided by the number of tasks that met the quality bar. Failed attempts and human rework count in the numerator — which is exactly why cheap-per-token models can lose on this measure.

The core formula
Question 3
Dependability
Can people depend on the result?

Reliability as a first-class economic variable. A workflow that succeeds four times in five still needs a human check on all five runs — dependability is what lets review time actually fall.

Reliability = economics
Question 4
Value at scale
Does each AI dollar produce more value as usage grows?

The compounding test: as adoption widens, does marginal value per dollar rise or fall? This is the question adoption counts alone can never answer — including OpenAI’s own 10 million figure.

The compounding test

Friar points to GPT-5.6’s three tiers as the built-in mechanism for this optimization — in her framing, a customer might use Luna for a fast, high-volume workflow, Terra for work requiring greater depth, or Sol when stronger reasoning delivers the best result with fewer attempts. The list prices make the spread concrete: Sol runs 5x Luna’s rate on both input and output.

Frontier tier
GPT-5.6 Sol · per Mtok in / out
$5/ $30

The strongest-reasoning tier, and the priciest in OpenAI’s own family. Friar’s pitch: fewer attempts per successful task can beat a lower sticker price.

Highest list price
Middle tier
GPT-5.6 Terra · per Mtok in / out
$2.50/ $15

OpenAI pitches Terra as delivering competitive performance to GPT-5.5 while being 2x cheaper — a vendor claim, but a checkable one on your own workloads.

OpenAI’s value pitch
Volume tier
GPT-5.6 Luna · per Mtok in / out
$1/ $6

The high-volume tier — one-fifth of Sol’s rate on both sides. The existence of this tier is itself evidence that per-token price competition is alive and well.

1/5 of Sol’s rate
Vendor self-grading — read with care
The same post carries two efficiency claims that deserve explicit labels. OpenAI says GPT-5.6 Sol at max reasoning set a new state of the art on the Artificial Analysis Coding Agent Index while using 54% fewer output tokens than another leading model — the rival is unnamed by OpenAI, and we won’t guess. And a DeepSWE v1.1 chart shows Sol at 72.7% against Claude Fable 5’s 69.9% at a 36.2% lower estimated API cost — a vendor self-reported comparison, published inside the very post arguing for OpenAI’s value framework, with no independent corroboration we could find. Neither claim is implausible; neither is verified.

Taken on its merits, the framework is genuinely useful — it is close to the approach we recommend in measuring agent ROI beyond task completion, and cost-per-successful-outcome is a better unit of account than anything in the adoption headline. The problem isn’t the framework. The problem is who’s holding it, and what they published four days later.

06Reality CheckAxios, and the tension OpenAI can’t message away.

Axios got an exclusive pre-publication look at Friar’s post and framed it bluntly: as executives question whether their massive AI bills are paying off, OpenAI’s CFO is pitching a new scorecard for measuring AI’s success. Axios reads the framework as OpenAI’s response to a growing enterprise AI cost reckoning, as companies increasingly route work to cheaper models and demand clearer returns on their investments. And it names the conflict directly: “OpenAI has an incentive to steer customers toward measuring outcomes instead of sticker prices, since its models are among the most expensive.”

The behavioral evidence runs the other way from the messaging. Rather than defaulting to the most capable frontier model, executives told Axios they’re increasingly routing everyday tasks to the cheapest model that can do the job, using frontier models only for the most intensive work — with others trying open-weight models or AI routers that auto-balance cost against performance. Axios also reported that Sam Altman describes cost as the second biggest challenge customers raise with him, behind AI deployment within their organizations — reported speech, not a direct quote, but a striking admission either way. For added color, Axios has separately reported on an executive who oversaw a half-billion-dollar accidental Claude bill over the course of a single month — the kind of story that explains why cost discipline arrived so abruptly.

Put the two publications side by side and the tension is the whole story. On July 17, OpenAI argues that value per dollar — not usage — is the measure that matters, and discloses no value-per-dollar data. On July 21, OpenAI’s milestone reaches the press as a pure usage number — no paid share, no retention, no cost-per-task, no product split. The company proposing the better yardstick is, so far, only publishing numbers measured with the old one.

07Buyer’s ScorecardWhat to actually ask when a vendor cites adoption.

Friar’s four questions are a good start — and they conveniently stop where a vendor’s disclosure comfort ends. A complete buyer’s scorecard adds the questions her framework skips. The next time an adoption milestone shows up in a sales deck (any vendor’s, not just OpenAI’s), these four demands separate a meaningful metric from a momentum headline.

Product split
Which product is this number for?

Combined tallies hide whether the product you’re actually buying is the one growing. OpenAI’s 10M spans a mature developer tool and a twelve-day-old work product with no split — a buyer evaluating ChatGPT Work learns almost nothing from it.

Demand per-product actives
Activity window
Daily, weekly, or point-in-time?

An undefined ‘active’ can inflate a figure severalfold versus daily actives. OpenAI hasn’t defined its window; TNW had to infer ‘weekly.’ Ask for the definition in writing, and ask for the same series three months back.

Demand the definition + trend
Paid share & retention
Who pays, and who returns?

Launch-week curiosity and durable adoption look identical in a users count. The two numbers that distinguish them — paid conversion and return rate after the novelty fades — are exactly the ones the milestone omits.

Ask paid ratio + 90-day retention
Verification
Who audited this?

Company estimates are fine as directional signals and worthless as procurement evidence. If a vendor’s number matters to your business case, ask whether any third party has reviewed the methodology — and weight it accordingly if not.

Ask for third-party review

Then turn the lens inward, because the strongest response to vendor metrics is instrumentation you own. Friar’s four questions convert directly into steps you can run this week without buying anything: pick one workflow that matters (question 1) and define what a successful completion looks like before the AI touches it; compute fully-loaded cost per successful task (question 2) — model spend plus human review time, divided by outputs that met your quality bar, with failed attempts in the numerator; track dependability (question 3) as the share of runs needing human correction, week over week; and re-run the numbers as usage grows (question 4) to see whether marginal value per dollar is rising or falling. A concrete example metric: cost per successfully closed support ticket, fully loaded including review time. If you want the spreadsheet version, our walkthrough on building an enterprise agent business case operationalizes the same math.

08ImplicationsWhat the split between headlines and measurement means for you.

The trend underneath this news cycle is a genuine phase change in how enterprise AI gets bought. Through 2025, adoption itself was the KPI — being able to say “our teams use AI” carried internal value regardless of measured output. The July 2026 picture is different: the vendor publishing the largest agent-adoption figure we could find now feels compelled to argue, through its own CFO, that adoption figures aren’t the point. That reversal only happens when buyers have already stopped paying for usage stories — and the Axios reporting on cheapest-adequate-model routing shows exactly that behavior hardening into procurement policy.

Looking forward, expect the two currents to keep diverging. Vendors will keep publishing bigger combined adoption numbers — they are cheap to produce and still move perception — while the buying side keeps building routing tables, cost dashboards, and outcome metrics that make sticker-price loyalty irrational. Our projection: within a few quarters, value-per-task disclosures (however imperfect) will start appearing in vendor marketing because undefined user counts will have stopped converting sophisticated buyers. The vendors that publish audited task-success economics first will win the enterprise argument, for the simple reason that they’ll be the only ones a CFO can put in a model.

For your own stack, the playbook follows from the scorecard above: treat adoption milestones as market-timing signals, not procurement evidence; instrument cost per successful task on your top three workflows before the next renewal conversation; and route by measured outcome, not by brand. If you want senior help standing that up — model routing, cost baselines, and outcome-level measurement across your AI spend — our AI transformation engagements start with exactly this instrumentation, and the same discipline applies whether the vendor on the invoice is OpenAI, Anthropic, or anyone else.

09ConclusionReal growth, unfinished number.

The shape of the agent market, July 2026

Believe the momentum. Interrogate the metric.

OpenAI’s 10 million agent users is a real milestone sitting on top of a real curve — the executive-posted Codex series from 5 million to roughly 8 million weekly users in about six weeks is the best-evidenced adoption run we could find in the category. Nothing in this post argues otherwise, and OpenAI disclosing a combined estimate is normal corporate practice, not a deception.

But the number ships without the three definitions that would make it decision-grade: no product split, no stated activity window, no independent verification. And the company published, four days before the milestone, its own CFO’s argument that usage counts are the wrong yardstick — a framework the milestone itself conspicuously fails to meet. When a vendor’s measurement philosophy and its press numbers point in opposite directions, the gap is the story.

The buyer’s job hasn’t changed, only sharpened: demand the definitions, instrument cost per successful task on your own workflows, and let measured outcomes — not milestones — decide where your AI budget goes. Ten million users tells you the agent market is arriving. It tells you nothing about whether it’s working. That second question is the only one your CFO will ultimately pay for — and it’s the one you can answer yourself, starting this week.

Measure value, not vendor milestones

Adoption headlines are the vendor’s metric. Build your own.

Our team helps businesses instrument AI spend at the outcome level — cost per successful task, model routing by measured value, and vendor-claim verification — delivered in days, not quarters.

Free consultationExpert guidanceTailored solutions
What we work on

AI measurement engagements

  • Cost-per-successful-task baselines on your workflows
  • Multi-vendor routing by measured outcome, not brand
  • Vendor-claim verification for procurement decisions
  • AI spend dashboards your CFO will actually accept
  • Renewal-ready ROI evidence packs
FAQ · OpenAI’s 10M agent users

The questions we get every week.

Bloomberg reported on July 21, 2026 that OpenAI’s agent products — the Codex coding agent and the ChatGPT Work platform combined — reached 10 million users, roughly double the count from two weeks earlier. The Next Web and a Yahoo Finance syndication of the Bloomberg piece corroborated the figure the same day. Three caveats apply: OpenAI does not break out how the total splits between the two products; it has not stated whether ‘active’ means daily, weekly, or point-in-time users (TNW reads it as weekly actives); and the figure is a company estimate rather than an audited metric. The milestone landed twelve days after ChatGPT Work launched on July 9, 2026, and four days after OpenAI’s CFO published a post arguing that usage counts are the wrong way to measure AI value.