OpenAI Signals published its first country-by-country ChatGPT usage release on August 6, 2026, and buried inside it is the most useful public evidence yet about how AI adoption at work actually differs from AI adoption everywhere else: at work, people are more than twice as likely to ask ChatGPT to complete a task or create something than they are outside work, where asking for information remains the largest category.
That single sentence is worth more to anyone sizing an internal AI rollout than most of the adoption headlines published this year, because it describes a behaviour rather than a subscriber count. It is also easy to over-read. The dataset behind it covers ChatGPT Free, Go, Plus and Pro accounts — the ones OpenAI describes as generally managed by individuals rather than organisations. It is not a study of enterprise deployments, and OpenAI says plainly that it excludes enterprise and Codex usage.
This piece does three things: states what the release actually found and on what denominator, cross-reads it against an independently collected dataset from a different lab, and then converts the honest residue into a sequencing decision — where an internal rollout should start if task-completion demand is the signal you are following.
- 01At work, people ask the model to do, not to explain.OpenAI states that people are more than twice as likely to use ChatGPT to complete a task or create something at work than outside work, where seeking information and clarification remains the largest category.
- 02The denominator is consumer accounts, not companies.The Signals dataset covers messages in Free, Go, Plus and Pro accounts. Enterprise and Codex usage are excluded, and OpenAI says the dataset therefore likely underrepresents business and technical use cases.
- 03No exact ratio was published for the headline claim.OpenAI states the work-versus-outside-work gap qualitatively — more than twice as likely — without a percentage on the page. Anyone quoting a precise figure is inventing one.
- 04A second lab points the same direction on its own data.Anthropic’s Economic Index, collected independently from Claude surfaces between April 10 and June 10, 2026, reports that work sessions skew more automated than personal ones — a different taxonomy reaching a compatible conclusion.
- 05Use it to sequence, not to justify.This is directional evidence about where task-completion demand concentrates. It cannot carry a productivity, ROI or headcount claim — that still requires your own baseline on your own workflows.
01 — The FindingOne sentence worth more than a subscriber count.
The release, titled From asking to doing: How the world is putting ChatGPT to work, is the first country-by-country usage publication from OpenAI Signals, the public data hub run by OpenAI’s Economic Research team. Its headline claim is stated directly: at work, people are more than twice as likely to use ChatGPT to complete a task or create something — from writing and coding to analysis — than they are outside work. Outside work, use is more exploratory, and asking remains the largest category.
Most adoption numbers published by AI vendors describe reach: how many people signed in, how many seats were sold, how many weekly actives were counted in some undisclosed window. This one describes intent. It says something about what a person wanted from the model in the moment they typed, and it says that the answer differs systematically depending on whether the moment was a work moment. That is a more durable finding than a count, because it survives pricing changes, bundling and promotional spikes.
One discipline before going further: OpenAI states the ratio qualitatively. There is no percentage attached to “more than twice as likely” on the page, and no precise figure was published alongside it. If you carry this claim into a deck, carry it in OpenAI’s own words. A specific-sounding number here would be fabricated, and fabricated precision is exactly how a directional signal gets turned into a business case it cannot support.
“AI is no longer just helping people find answers. It is helping more people, in more places, get things done.”— OpenAI, From asking to doing: How the world is putting ChatGPT to work, August 6, 2026
The second thing the release does is establish a cadence. Signals is positioned as a recurring publication rather than a one-off report, with a companion data page documenting scope and definitions and a citable dataset version. For practitioners, a recurring vendor-published usage series with stated definitions is more useful than an occasional press release, provided the definitions stay stable and the scope stays legible. Both are worth watching across future releases.
02 — ScopeThe whole reading hinges on who is counted.
OpenAI is explicit about the boundary, and the sentence is load-bearing: the Signals dataset reflects messages sent within ChatGPT Free, Go, Plus and Pro accounts — the group of accounts generally managed by individuals rather than organisations. The companion data page adds that the sample is drawn from messages between July 2024 and June 2026 and that it excludes enterprise and Codex usage, and therefore likely underrepresents business and technical use cases.
Read that against the headline and the tension becomes obvious. The finding is about work. The dataset excludes the products companies actually buy for work. “At work” here means work-shaped usage occurring on a personal plan — someone drafting a client email on their own Plus subscription, an analyst pasting a spreadsheet into a Free session. That is a real and interesting population. It is not the population a CIO is measuring when they ask how AI is going inside the company.
There is a second-order consequence worth naming. Because Codex is excluded, the release under-counts precisely the coding work it partly credits itself with — the headline lists coding among the tasks people have the model do. The most agentic, most task-completion-heavy surface OpenAI ships is outside the frame. If anything, that cuts in favour of the direction of the finding while cutting against any attempt to treat the magnitude as complete.
| Dimension | What the release measures | What it is easy to assume | What the assumption would need |
|---|---|---|---|
| Who is in the sample | |||
| Account types | Messages sent within ChatGPT Free, Go, Plus and Pro accounts — accounts OpenAI describes as generally managed by individuals | Company-provisioned ChatGPT seats across an organisation | Workspace-level telemetry, which OpenAI reports separately from this release |
| Excluded surfaces | Enterprise and Codex usage are excluded; OpenAI states the dataset likely underrepresents business and technical use cases | A complete picture of AI at work, agentic coding included | Codex and workspace usage reported on the same denominator and in the same period |
| Sample window | A sample of messages spanning July 2024 to June 2026, with country comparisons drawn on quarterly slices | A live read on how people are using AI this month | Per-period counts with the sampling frame and eligibility rules published alongside them |
| How the numbers are produced | |||
| Work context | Inferred from the content of a message by an automated classifier — not from an account type, employer or seat | A verified work account, or usage logged against an employer | Account-level or seat-level attribution, which a consumer dataset does not have |
| Classification | Automated, privacy-preserving classification with noise-addition safeguards, as described by OpenAI in its own methodology documentation | An audited measurement pipeline comparable to a survey instrument | An independent audit, or a released classifier others can reproduce and stress-test |
| The headline ratio | Stated qualitatively — more than twice as likely — with no exact percentage on the page | A precise figure you can quote in a slide footnote | The underlying ratio published in the release itself or its companion dataset |
| What the numbers can carry | |||
| Direction | Work-shaped usage on consumer plans skews towards task completion rather than information seeking | Task completion is delivering measurable productivity gains | Outcome data — cycle time, rework rate, cost per completed task |
| Rollout value | A hint about where task-completion demand already concentrates before any rollout starts | A business case for a specific deployment and budget | Your own baseline, measured on your own workflows, before and after |
| The billion-person framing | The closing paragraph’s “more than 1 billion people” line is a broader, all-surface framing than the consumer-only dataset the statistics are drawn from | A billion people counted inside this dataset, doing what the statistics describe | A stated denominator, period and surface list for the billion figure specifically |
03 — TaxonomyAsking, doing and expressing, defined.
The three-way split doing the work here is not new to this release, and keeping the two publications separate matters. The Signals data page defines the categories precisely: asking is when a user is seeking information or clarification from ChatGPT; doing is when a user wants ChatGPT to produce an output or perform a task; expressing is when a user expresses views or feelings without seeking any information or action.
That taxonomy originates in an earlier and entirely separate OpenAI publication — a working paper circulated through the National Bureau of Economic Research, co-authored by OpenAI’s Economic Research team with Harvard economist David Deming, built on roughly 1.5 million conversations with data running through about mid-2025. That prior study reported an overall split of about 49% asking, 40% doing and 11% expressing, and estimated that roughly 30% of consumer usage was work-related against roughly 70% non-work.
Those numbers belong to that study, not to the August 2026 release. Different sample, different window, different question. The continuity between them is methodological rather than statistical: Deming appears among the named authors on both, which is why the category definitions carry across. Fusing the 49/40/11 split into the country-level release would produce a claim neither publication makes.
Overall message split, all contexts · earlier study, not the August 2026 release
Source: OpenAI's earlier NBER-circulated study of ~1.5 million conversations, data through approximately mid-2025 — a separate publication from the August 2026 country-level releasePut the two publications side by side and the interesting thing is not a contradiction but a decomposition. In the earlier study, across all contexts, asking was the largest category. In the later release, once you separate work contexts from everything else, task-completion use is more than twice as likely inside work as outside it. Both can be true at once, and the earlier study's estimate that roughly 30% of consumer usage is work-related is what makes the distinction matter: an all-contexts average is dominated by the non-work majority.
The practical reading: an average taken across all consumer usage will systematically understate how task-oriented AI use is inside a working context, because it is diluted by a much larger volume of personal, exploratory sessions. Anyone benchmarking their team against a headline “most people just ask it questions” statistic is benchmarking against the wrong denominator.
04 — The Other SignalsThree findings that travelled less far.
The doing-versus-asking line took the coverage, but three other findings in the same release matter more to anyone planning content, product or market entry. Each carries the same consumer-account denominator, and one of them carries an additional caveat worth stating up front.
Multimedia use, globally
Generation, analysis and retrieval of images and other media is the fastest-growing use case globally, at 7.8% of classified consumer messages. In Brazil and Colombia it exceeds one in ten messages.
Over-35 share, France and Czechia
The over-35 share of messages rose by more than ten percentage points year-over-year in France and Czechia, with over-35 usage rising in nearly every country measured. This analysis covers only users who self-reported an age.
Countries in the rank comparison
Latin America, Africa and Oceania are described as catching up to early adopters. Peru, Uruguay and Costa Rica climbed most in per-capita usage rank between Q1 and Q2 2026.
The age finding needs its caveat repeated rather than footnoted: it covers only the users who volunteered an age on the platform. That is a self-selected subsample, and self-selection on a demographic question is rarely random. The direction — older cohorts adopting faster than the stereotype suggests — is plausible and consistent with what agencies see in their own client data, but the specific percentage-point movements describe the people who answered, not the user base.
The multimedia figure is the one with the clearest commercial consequence. If 7.8% of consumer messages globally, and more than one in ten in two large Latin American markets, now involve generating or interrogating media rather than text, then the assumption that AI usage is a text phenomenon is already stale. For teams planning creative operations, that shifts the question from whether to build a media pipeline to which parts of one are worth owning.
The geographic movement is the least surprising and the most structurally important. Per-capita rank movement across 144 countries is a coarse measure, but the countries named — Peru, Uruguay, Costa Rica — sit outside the markets where AI product decisions are usually made. Growth concentrating there while product roadmaps assume North American and Western European usage patterns is the kind of mismatch that shows up later as a localisation problem.
05 — Second LensA different lab, a different taxonomy, the same direction.
A vendor measuring its own product and publishing a flattering reading of the result is not evidence in the strong sense. It becomes more interesting when a competitor, measuring a different product with a different method, lands in a compatible place. Anthropic’s Economic Index report, published on June 26, 2026, does exactly that — and it is worth being precise about how little the two studies have in common methodologically.
Anthropic’s report draws on continuous hourly sampling of conversations across Claude surfaces between April 10 and June 10, 2026, paired with a survey of approximately 9,700 Claude users linked to their usage data, and classified by privacy-preserving classifiers that do not expose transcripts to human reviewers. Its relevant finding: work sessions skew more automated than personal ones, and sessions on Claude Code — an agentic coding surface — are on average more automated than those on chat.
Two studies, two companies, two products, two classification schemes, and a compatible conclusion: the work context is where task-completion usage concentrates. That is closer to a real signal than either release is on its own. It is emphatically not the same measurement — Anthropic’s automation and augmentation taxonomy (directive, feedback loop, task iteration, learning, validation) is not OpenAI’s asking, doing and expressing, and the two should never be presented as measuring identically defined categories.
Country-level consumer usage
Automated classification of message intent into asking, doing and expressing. Work context inferred from message content. Enterprise and Codex usage excluded by design.
Cross-surface automation share
Continuous hourly sampling paired with a linked survey of approximately 9,700 users. Automation and augmentation taxonomy, classified without human transcript review.
Direction, not magnitude
Agreement across independent datasets raises confidence that work context predicts task-completion usage. It does not license averaging the two, comparing their percentages, or treating either magnitude as validated.
One more contrast from Anthropic’s data is useful as a sanity check on how strongly context shapes usage: personal conversations rise from roughly 35% on weekdays to roughly 50% at weekends on its chat surfaces. Usage composition is not a fixed property of a user base; it is a property of the moment. That is a caution against reading any single-period snapshot — including this one — as a stable description of how people use AI.
06 — MethodologyWho classified these messages, and audited by whom?
Every percentage in this release is the output of a classifier deciding what a message was for. OpenAI describes that pipeline in its own methodology documentation: classification performed automatically by language models rather than human reviewers, with noise-addition safeguards intended to provide message-level privacy guarantees. That design is defensible — it is arguably the only way to study usage at this scale without exposing conversations to people — and it is also entirely self-described.
No independent audit of the classification methodology was locatable at the time of writing. That does not make the numbers wrong. It does mean the correct citation form is “OpenAI states” rather than “research shows”, and it means the classifier’s definition of a work context is a design choice by the company whose product is being measured, not an externally validated construct.
The commercial incentive is worth naming plainly rather than insinuating. OpenAI benefits from a reading in which its consumer product is already doing serious work. That does not imply the data is manipulated; vendor-published usage research is often the only usage research that exists, and it is frequently useful. It does mean the burden of scepticism sits with the reader, and that the specific claims to distrust are the expansive framings rather than the narrow, scoped ones.
We have applied the same discipline before. When OpenAI disclosed a 10 million weekly-user figure for its agent products earlier this quarter, the useful analysis was not the number but the questions it left unanswered — which products were combined, what activity period counted, who verified it. That post, on what the 10 million agent-user number hides, is a separate metric on a separate denominator from anything in this release, and the two should never be added together or used to corroborate each other.
07 — What It ArguesThe claims this data can carry.
The most common failure with a release like this is not misquotation, it is over-extension: a scoped behavioural finding gets promoted into a productivity claim somewhere between the source and the board pack. The following separation is the one worth internalising before the finding leaves your hands.
Work context predicts task-completion usage
Across consumer accounts, work-shaped sessions skew towards having the model produce an output rather than explain something. Independently echoed by Anthropic's separately collected index across Claude surfaces.
Demand is already there before you deploy
People are doing task-completion work on personal plans, unmanaged and unmeasured. That is a live signal about latent internal demand — and about shadow usage your policy may not cover yet.
A productivity or ROI claim
Nothing in this release measures output quality, cycle time, rework or cost. Message intent is not an outcome. No study here connects a doing message to a completed piece of work of known quality.
An enterprise adoption benchmark
Enterprise and Codex usage are excluded by construction. Using this to benchmark your organisation's deployment maturity compares a consumer message sample against a corporate rollout.
A precise ratio behind the headline claim
OpenAI published the work-versus-outside-work gap qualitatively — more than twice as likely, with no percentage. Any specific figure attached to it in secondary coverage was not in the release, and should be traced to a primary before it travels further.
The billion-person headline
The closing framing is broader and all-surface, while the statistics are consumer-only. If you cite the billion figure, say which claim it belongs to and do not present it as validated by the same dataset.
There is a forward-looking read worth putting on the record. If the asking-to-doing shift is real and continues, the constraint on enterprise AI value moves from model capability to workflow plumbing — permissions, data access, review gates, the ability to hand a model a task with enough context to finish it. A population that already prefers delegation over explanation will hit the limits of an under-integrated deployment quickly, and will route around it with personal accounts, which is precisely the behaviour this dataset is made of.
The second forward read concerns measurement. As agentic surfaces grow, message-count taxonomies get progressively less informative: a single message that dispatches an hour of autonomous work counts the same as a one-line question. Expect future releases in this series to either add a unit of work alongside the message count or to under-describe the phenomenon they are tracking. That is worth watching as the honest test of whether the series stays useful.
08 — SequencingTurning a directional signal into a starting point.
The usable output of this release is a sequencing hint, not a business case. If task-completion demand concentrates in work contexts, then the first place to put an agentic workflow is wherever that demand is already visible without you having created it — the teams whose people are quietly pasting work into personal accounts. That is cheap to find and it is a far better starting signal than a capability matrix.
The order below is the one we run with clients, and it is deliberately front-loaded with measurement you own rather than statistics somebody else published. It pairs with the fuller treatment in our guide to hybrid human-agent adoption in professional services, which covers the staffing and review-gate side of the same decision.
Find the existing demand
Ask which tasks people already hand to a model on their own account. Unmanaged usage is the highest-signal indicator you have of where a supported workflow will be adopted, and it costs nothing to collect.
Baseline before you deploy
Measure the three tasks at the top of that list as they run today. Without a pre-deployment baseline, no later number can be attributed to the rollout, and the vendor's usage statistics cannot fill that gap.
Ship one completion workflow
Pick a task where the desired outcome is an artefact, not an answer, and instrument the review gate. Doing-shaped work fails in review, not in generation, so the gate is the part worth engineering first.
Then decide on platform
Vendor pricing for agent capacity is still moving quarter to quarter. Committing to a seat or credit model before you know your own task volume is how organisations end up paying for capacity they never route work through.
Step four deserves emphasis because the pricing environment is genuinely unsettled. Vendors are actively re-shaping how agent capacity is sold — credits, seats, consumption, or some hybrid — and the same week this release landed, our analysis of HubSpot’s Q2 2026 AI pricing and credit changes showed how quickly the commercial terms around agent adoption can move. Buying capacity against a usage statistic from someone else’s product is the wrong order of operations.
There is also a role-design consequence. Task-completion usage does not respect job descriptions — OpenAI’s own earlier research on task crossover between professional roles found a substantial share of job-specific requests belonging to a different profession entirely. If your rollout is scoped by department, expect the demand to cross the boundary within weeks, and plan the review gates accordingly rather than treating the crossover as misuse.
None of this requires believing OpenAI’s framing. It requires only accepting the narrow, scoped version of the finding: in work contexts, people want the model to finish something. Design for finishing — context, permissions, review — and the rollout meets the demand that already exists. If you want help running that sequence, our AI and digital transformation practice is built around exactly this order of operations, and our content engine work covers the writing and analysis workflows where task-completion demand tends to surface first.
09 — ConclusionA good signal, worth reading narrowly.
The finding is useful precisely because it is narrow — keep it that way.
OpenAI’s country-level Signals release contains one genuinely valuable observation: people are more than twice as likely to ask ChatGPT to complete a task at work than outside work, where asking for information still dominates. Anthropic’s independently collected index points the same way from a different product with a different taxonomy. Two labs, two methods, one compatible direction is about as good as public evidence on this question currently gets.
It is also a consumer dataset, self-classified, without an independent audit, published by a company with an interest in an expansive reading, and stated without the precise ratio that would let anyone check the magnitude. Every one of those limits is disclosed in the release itself, which is to OpenAI’s credit and is also the reason the honest use of this data is directional rather than evidentiary.
So use it for what it is. Let it tell you where to look first — the teams already delegating tasks on personal accounts — and then measure your own workflows before and after, because that is the only number that will ever belong to you. The organisations that get value from this shift will not be the ones that quoted the statistic fastest; they will be the ones that built the review gates and data access that let a task actually finish.