Task crossover is OpenAI’s name for something most agency operators have already watched happen on their own delivery floor: a marketer asking an AI assistant to model a payment schedule, a designer drafting release notes, a customer-experience lead writing a SQL query. On July 27, 2026, OpenAI put a number on it — 43.5% of occupation-specific ChatGPT requests from U.S. business users involve tasks associated with a different profession.
The report, “Work at the Frontier: How AI is Expanding What People Do at Work,” is drawn from an analysis of more than 800,000 messages and is the first instalment of what OpenAI says will be a recurring labour-market series. Axios got it first; the trade press ran the 43.5% figure within hours. What almost nobody carried is the denominator underneath it, which changes the size of the claim by more than half.
This guide does three things. It restates exactly what OpenAI measured and on which basis, so the number can be quoted without overstating it. It applies an independent economist’s named critique of vendor-authored labour research to this specific report. And it translates the finding into the only question that matters for an agency: whether role definitions, scoping and capacity planning are still measuring the right unit of work.
- 01OpenAI coined “task crossover” and put 43.5% on it.Across 800,000+ messages from U.S. business ChatGPT users, OpenAI reports that 43.5% of occupation-specific requests involve tasks associated with a different occupation, mapped via O*NET activity codes.
- 02The denominator does most of the work.43.5% is calculated on occupation-specific messages only — which are 38.5% of all work messages. On an all-messages basis the crossover rate is 16.8%. Both numbers are correct; only one of them is the headline.
- 03Confirmed per-occupation shares: CX 77%, design 75%, HR 69%, legal 56%, marketing 53%.All five are on the occupation-specific basis and appear verbatim in OpenAI’s report. They describe the share of job-specific requests, not the share of a person’s working time.
- 04Engineering runs the pattern in reverse.OpenAI states that 18.5% of engineering messages involve tasks from other fields, while engineering tasks account for 7.4% of messages from workers in other occupations. Engineering supplies work outward more than it borrows inward.
- 05It is vendor-authored research on its own product.The report is internally produced and not peer-reviewed, and OpenAI states it cannot tell whether crossover work is new or pre-existing, nor whether it matches specialist quality. Treat it as a well-instrumented signal, not settled economics.
01 — The FindingWhat OpenAI actually published.
The method is straightforward and worth understanding before the headline number means anything. An LLM-based classifier reads each message transcript, summarises the underlying task, and maps it to an activity code in O*NET, the U.S. Department of Labor’s occupational database of work activities. A message counts as crossover when its mapped activity belongs to a different occupation than the one inferred for the user.
Before that comparison runs, OpenAI filters out what it calls generic messages — writing, summarising, scheduling and similar activities that are shared so broadly across occupations that they cannot signal crossover in either direction. What remains is the occupation-specific slice, and it is inside that slice that the 43.5% figure lives.
Trade coverage of the launch reports that OpenAI also opened an External Research Exchange alongside the report, inviting independent labour economists to propose studies using its platform data — an acknowledgement, implicitly, that internally-run analysis of your own product has limits. The full report is on openai.com, and Axios carried the first interview with OpenAI chief economist Ronnie Chatterji.
Messages analysed
Work-related ChatGPT messages from U.S. business users. Each transcript is summarised by a classifier and mapped to an O*NET work-activity code, then compared against the occupation inferred for the sender.
Occupation-specific crossover
The share of occupation-specific messages whose mapped task belongs to a different occupation. Axios rounds this to roughly 44%; 43.5% is the figure in OpenAI's own report text and the one to quote.
All work messages
The same crossover measure calculated across every work-related message, generic ones included. This is the figure that answers 'how much of AI use at work crosses role lines' without a filter applied first.
02 — Denominator DisciplineTwo numbers, both true, wildly different sizes.
Generic messages make up 61.5% of all work-related messages in the study. The remaining 38.5% are classified occupation-specific, and the 43.5% crossover rate is calculated inside that narrower slice. The same phenomenon measured across every work message comes out at 16.8%.
Those two figures reconcile, and checking that they do is the fastest way to confirm you have understood the structure. Dividing the all-messages rate by the occupation-specific rate recovers the occupation-specific share of the corpus: 16.8 ÷ 43.5 = 38.6%, against the 38.5% occupation-specific share reported for the corpus — a rounding step apart. Run it the other way and 43.5% × 38.5% = 16.7%, again within rounding of the published 16.8%. The arithmetic is internally consistent; the interpretation is where posts go wrong.
The failure mode to avoid is time. None of these percentages describe how a person spends their day. They describe the composition of messages sent to one AI assistant, by users whose occupation was inferred rather than declared, filtered through a classifier’s judgement about which O*NET activity a request most resembles. “Designers spend 75% of their time on other people’s work” is not what the report says, and it is not what the data can support.
03 — By OccupationWhere crossover concentrates.
Five per-occupation shares appear verbatim in OpenAI’s report text. All five are on the occupation-specific basis — the narrower denominator — so they should be read as “of the job-specific requests this group sends, this share maps to another occupation.”
Outside-occupation share · occupation-specific messages
Source: OpenAI, Work at the Frontier (July 27, 2026) — occupation-specific (non-generic) messages onlyRedraw the same groups on the all-messages basis and the picture compresses hard. Only three occupations have a published all-messages figure, and the study-wide rate sits below all of them.
Outside-occupation share · all work messages
Source: OpenAI, Work at the Frontier (July 27, 2026) — design, marketing and the study-wide rate are on the all-messages basis; OpenAI's report text names no denominator for the engineering figureOne honesty note on engineering. Secondary coverage of this report circulated a precise occupation-specific figure for engineers, ranking them lowest of the groups charted. We could not confirm that figure against either OpenAI’s own report page or the Axios article text, and the outlet carrying it contradicts itself within the same piece. So we are not printing it. What OpenAI does state directly is the 18.5% / 7.4% pair above and below — which is a different and more interesting measurement anyway.
04 — Borrow vs SupplyCrossover has two directions, not one.
Every write-up we read treated crossover as a single ranked list. It is not. OpenAI publishes an inbound measure (how often this group’s requests belong to someone else’s occupation) and, for some groups, an outbound measure (how often this group’s tasks show up in everyone else’s requests). Those two axes describe genuinely different roles in an organisation, and the extremes sit at opposite corners.
Design borrows heavily and is rarely borrowed from: 35.2% of all designer messages involve another occupation’s tasks, while design-associated tasks make up just 1.7% of other workers’ messages. Engineering is the mirror image — 18.5% of engineering messages involve outside tasks, but engineering tasks appear in 7.4% of messages from workers in other occupations. Marketing is the only group that scores high on both: 24.3% inbound on the all-messages basis, and an 8.9% outward share that is the highest of any occupation measured.
The table below is our own synthesis. No single source in the coverage set combines the inbound and outbound metrics into one comparison; each outlet reported either the ranked list or the directionality paragraphs. Cells are left blank where OpenAI did not publish a figure rather than filled with an inference.
| Occupation | Inbound · occupation-specific | Inbound · all messages | Outbound · supplied to others | Derived · occupation-specific density |
|---|---|---|---|---|
| Inbound published on the occupation-specific basis only | ||||
| Customer experience | 77% | Not published | Not published | Not derivable — needs both denominators |
| Human resources | 69% | Not published | Not published | Not derivable — needs both denominators |
| Legal | 56% | Not published | Not published | Not derivable — needs both denominators |
| Both directions published | ||||
| Design | 75% | 35.2% | 1.7% of other occupations’ messages — the lowest of the three outward shares OpenAI published | 35.2 ÷ 75 = 46.9% |
| Marketing | 53% | 24.3% | 8.9% of other occupations’ messages — the highest outward share reported | 24.3 ÷ 53 = 45.8% |
| Engineering | Not published in comparable form — see note above | 18.5%, stated by OpenAI as a share of engineering messages | 7.4% of other occupations’ messages | Not derivable — needs both denominators |
| Study-wide baseline | ||||
| All occupations | 43.5% | 16.8% | Not applicable — outward share is defined per occupation | 16.8 ÷ 43.5 = 38.6%, against the 38.5% occupation-specific share reported for the corpus |
The derived column is the one we find most useful, and it is simple arithmetic on OpenAI’s own figures: divide a group’s all-messages crossover share by its occupation-specific share and you recover how much of that group’s AI volume was classified occupation-specific in the first place. Designers land at 46.9% and marketers at 45.8%, both meaningfully above the 38.6% study-wide figure. In plain terms, these two groups send a higher proportion of job-specific requests than the average worker — which is exactly why their headline percentages look so dramatic.
Two task types travel almost everywhere. OpenAI reports that financial calculation and technology troubleshooting each rank among the top three outside-occupation tasks for workers in all seven other occupation groups studied. “Creating marketing materials” appears as an outside-occupation task across five other groups, and is especially prominent among designers. Those three activities — money maths, fixing the tooling, and making the collateral — are the connective tissue of the whole finding.
05 — Company SizeSmall teams cross more lines.
OpenAI segmented the result by workspace size and found a modest but directionally clear gradient. Among typical (median) users, the outside-occupation task share falls from 18.9% at workspaces of two to five seats to 16.3% at workspaces of 100 or more. Among the heaviest users, the report notes the pattern is not monotonic — so this is a tendency, not a law.
The interpretation OpenAI offers is intuitive: “AI may be especially useful as a generalist tool where specialist resources are scarce.” In a five-person business there is no finance function to route the question to, so the person nearest the problem handles it. In a large organisation there is, and they do.
For a boutique agency this is the most operationally honest part of the report, because it describes the actual failure mode rather than the aspiration. Small teams do not cross role lines because AI unlocked latent range; they cross them because nobody else is going to. Whether the output survives contact with a specialist reviewer is a question this dataset cannot answer.
Small-workspace crossover
Outside-occupation task share among typical (median) users at workspaces of two to five seats. The highest band OpenAI reports on the workspace-size gradient.
Large-workspace crossover
The same measure at workspaces of 100 or more seats — a 2.6-point gap versus the smallest band. Among the heaviest users, OpenAI notes the pattern does not hold monotonically.
06 — Read the IncentiveThe company measuring this sells the thing being measured.
This is not a hostile framing; it is a methodological fact that OpenAI itself partly concedes. The report is internally produced and has not been peer-reviewed. It measures usage of OpenAI’s own product, by OpenAI’s own classifier, against an occupation label OpenAI inferred. And the conclusion it reaches — that AI is expanding what people can do at work — is the most commercially useful conclusion available from that data.
The sharpest available critique of exactly this pattern predates the report by more than four months. Writing for the Peterson Institute for International Economics in March 2026, economist Jed Kolko argued that the evidence on AI’s labour-market effects is inconclusive and that claims about harm to particular groups are premature. He also named a specific failure mode — narrator’s bias — for the way researchers, journalists and content producers who are themselves heavy LLM users can unconsciously colour the interpretation and tone of AI-and-labour research. A vendor studying its own product’s usage is the purest instance of the problem he described.
To OpenAI’s credit, the stated limitations are unusually candid. Per the company’s own framing reported by Axios, the analysis does not determine whether AI is creating new cross-occupation work or simply helping workers perform responsibilities they already had. It does not measure whether crossover work produces higher productivity. It does not assess output quality against a specialist baseline. And it says nothing about hiring decisions. Every one of those gaps sits directly underneath the operational conclusions people are already drawing from the headline.
The vendor-scrutiny discipline generalises. We applied the same reading to OpenAI’s own agent-adoption numbers, where the headline figure combined two products, defined no activity period, and was unaudited. The pattern is consistent: vendor telemetry is genuinely valuable evidence about usage, and genuinely weak evidence about outcomes.
"The boundaries between jobs are likely already becoming more flexible due to AI."— Ronnie Chatterji, Chief Economist at OpenAI, speaking to Axios, July 27, 2026
07 — TriangulationThree other lenses on the same shift.
A single vendor’s message-classification study is one lens. It gets considerably more persuasive when a competitor, using a completely different method, lands in a compatible place — and considerably more sobering when the government’s own projections decline to move. The table below puts all four next to each other, including what each one explicitly cannot tell you.
| Study | Method | Sample | Headline finding | What it does not establish |
|---|---|---|---|---|
| AI-vendor research · self-collected data | ||||
| OpenAI — Work at the Frontier (Jul 27, 2026) | LLM classifier summarises each message and maps it to an O*NET activity code, compared against an inferred occupation | 800,000+ work messages from U.S. business ChatGPT users | 43.5% of occupation-specific messages involve another occupation’s tasks; 16.8% on an all-messages basis | Whether the crossover work is new or pre-existing; productivity; output quality; hiring effects. Not peer-reviewed |
| Anthropic — Economic Index, June 2026 report (Jun 26, 2026) | Linked self-report worker survey — a different instrument entirely from message classification | 9,700 surveyed workers | More than one-third rated it likely or very likely their own responsibilities would change significantly within 12 months; about 10% feared losing their own role | Whether expectations match outcomes; what specifically changes; anything measured rather than self-reported |
| Anthropic — Labor Market Impacts (Mar 5, 2026) | Occupational AI-exposure measure tested against labour-market outcomes | Not restated in our source set — see the original study | No systematic increase in unemployment for highly exposed workers since late 2022, alongside a roughly 14% drop in the job-finding rate for workers aged 22–25 entering the most exposed occupations | Causation. Anthropic describes the young-worker signal as just barely statistically significant |
| Government statistical baseline | ||||
| BLS — AI impacts in employment projections (published Mar 11, 2025; 2023–33 window) | Official occupational employment projections with AI-exposed occupations flagged | All U.S. occupations | Professional-services growth stays positive: personal financial advisors +17.1%, business and financial operations +6.9%, lawyers +5.2%, paralegals +1.2%, against +4.0% for all occupations | Anything about task-level reallocation inside a role. BLS itself frames these trajectories as remaining uncertain |
Read together, the four lenses converge on something narrower than the headlines suggest. Task and role boundaries are loosening — OpenAI sees it in message composition, Anthropic’s surveyed workers expect it of their own jobs. But nothing in the set demonstrates that headcount follows. BLS’s projections for AI-flagged professional-services occupations are broadly positive over the 2023–33 window; the only negative projections in its AI-flagged list are narrow transaction-processing roles, with claims adjusters at −4.4%, credit analysts at −3.9% and auto-damage insurance appraisers at −9.2%. Those are the roles where the task is the job.
The one genuinely uncomfortable data point is Anthropic’s hiring-pipeline signal for 22-to-25-year-olds, and it deserves its hedge — Anthropic itself calls it just barely statistically significant. Read alongside the same survey’s finding that respondents worried far more about junior colleagues than themselves, it points at the entry-level rung rather than the profession. That is the part of an agency’s staffing model most exposed here: not senior specialists, but the apprentice work that used to train them. We wrote about the skill side of this in our look at what separates an expert AI user from a novice.
"AI may change the work people do before it changes the number of people who do it."— Axios, bottom-line framing on the OpenAI report, July 27, 2026
08 — Agency OperationsThe org chart is measuring the wrong unit of work.
OpenAI, Axios and the trade press are all writing this story for a labour-economics audience. The operational reading is different and nobody has written it, so here it is. If a marketer’s AI volume is 24.3% outside-occupation tasks and a designer’s is 35.2% — both on the conservative all-messages basis — then scoping, resourcing and pricing built around rigid role silos are describing something that no longer matches how the work is actually being produced.
That does not mean restructure. OpenAI’s report explicitly cannot tell you whether the crossover work is new, whether it is any good, or whether it changes what you should hire. Any recommendation to collapse specialist roles on the strength of this data is claiming more than the data claims. What it does justify is instrumentation: find out where crossover is already happening in your own delivery, and whether the output is being reviewed by anyone qualified to catch it when it is wrong.
Four decisions actually change on the back of this. The rest is noise.
Scope by task cluster, not by title
If financial calculation and technology troubleshooting are the two tasks travelling into every other occupation, they belong in a role definition as named responsibilities with a named reviewer — not as invisible work that a specialist absorbs and nobody prices.
Add a review gate before you add range
Nothing in this dataset says crossover output matches specialist quality. Before celebrating a marketer who can now draft a contract clause, decide who reads it. The cheapest version of this is a named reviewer per crossover task type.
Measure delivery in tasks, not in seats
Utilisation models that assume a designer produces design hours will systematically mis-forecast when a third of that person's AI volume is other-occupation work. Track task mix for one quarter before changing any staffing ratio.
Protect the apprentice rung deliberately
The one hiring signal in the wider evidence base points at entry-level entry, not at senior displacement — and it is a weak signal. Treat junior scope as something to design on purpose rather than something crossover quietly erodes.
There is a service-design version of this argument too. If the unit of delivery is a task cluster rather than a job title, then the way you package and price work should follow — which is the logic behind how we structure productised content engine engagements and the broader operating-model work in our AI transformation practice. A hybrid model, where agents carry the repeatable task volume and senior humans hold judgement and review, is the pattern we set out in our hybrid adoption model for AI agents in professional-services delivery.
09 — What To Do NowA measured response, in four moves.
The right posture is neither dismissal nor restructuring. It is instrumentation — running the same measurement on your own delivery data that OpenAI ran on its message corpus, then deciding from evidence you own rather than evidence a vendor published about its own product.
Audit your own crossover
Tag delivery work by task cluster rather than by who did it. You are looking for the same two axes OpenAI measured: which roles are absorbing outside tasks, and which roles' tasks are being absorbed by everyone else.
Name a reviewer per crossover type
For every task type crossing a role line — financial modelling, technical troubleshooting, contract language — name the person who signs it off. This costs nothing and is the only defensible answer to the quality question the report leaves open.
Re-price the unit
Where a task cluster is genuinely being delivered by whoever is nearest with AI assistance, price it as a unit of output rather than as hours of a job title. Where it is not, leave the model alone.
Watch the cost curve
Crossover work is not free. We would expect a non-specialist reaching outside their domain to burn more assistant tokens and more review time to reach an acceptable answer, though OpenAI publishes no data on this. Track it before assuming range is a margin win.
The fourth move is the one most teams skip, and it is measurable today. Different roles doing adjacent work consume very different amounts of assistant capacity to reach the same standard, which is a direct input to whether crossover improves delivery economics or quietly degrades them — we looked at how token spend differs across roles doing adjacent work in more detail. Our forward read: over the next two to three quarters, the agencies that come out ahead will not be the ones that flattened their org charts fastest. They will be the ones that instrumented task mix early, kept a specialist review gate on every crossover cluster, and could therefore tell the difference between genuine range and confident output nobody checked.
10 — ConclusionA real signal, honestly sized.
The finding is credible. The denominator, the source and the silence about outcomes all matter.
OpenAI’s task-crossover report is the best-instrumented look yet at how AI use redistributes work across job boundaries, and the core observation matches what agency operators already see. But 43.5% is a share of occupation-specific requests, not of anyone’s time or workload, and the same phenomenon measured across all work messages is 16.8%. Quote the one that matches your claim, and say which one it is.
The directional structure is more useful than the ranking anyway. Design borrows heavily and supplies almost nothing back; engineering supplies broadly and borrows little; marketing does both, with the highest outward share of any group measured. That two-axis picture tells you where review capacity needs to sit far better than a league table of percentages does.
And the caveat is not decoration. This is a vendor grading its own product, unreviewed, with an explicit acknowledgement that it cannot say whether crossover work is new, whether it is any good, or whether it changes hiring. The right response is to run the measurement on your own delivery data, put a named reviewer behind every task that crosses a line, and let the org chart follow the evidence rather than the press release.