BusinessMethodology6 min readPublished September 23, 2026

20 studies · RCTs, field and natural experiments, administrative data · 12 find novices gain most, 3 go the other way

Who Gains Most From AI at Work? What 20 Studies Found

Twenty studies that measured how AI productivity gains differ by worker and task, in one table. Novices usually gain most, with exceptions the rows show.

DA
Digital Applied Team
Research and practical guidance
PublishedSeptember 23, 2026
Sources readSeptember 25, 2026

Every article about AI and work argues about who benefits, and almost none of them cite the studies that measured it. This page is the table those articles are missing: twenty studies that measured how the productivity effect of AI differs across workers or across tasks, each with its design, sample, average effect, and the exact figures for who gained more or less.

The short answer is that gains are uneven and usually largest for the least experienced. Twelve of the twenty rows find larger gains for lower-skilled, less experienced or lower-performing workers. Three go the other way or show experts slowed. The direction flips by task, and no row links a productivity gain to that worker's pay.

The scope is deliberately narrow. Usage logs and self-report surveys are not causal estimates, so they sit in a separate table. Surveys on whether workers admit to using AI sit in a third. One widely cited study is excluded because it was withdrawn, and it is named so that nobody adds it back.

Key takeaways
  1. 01
    Twelve of twenty studies find larger gains for less experienced or lower-performing workers.The biggest gaps: support agents in the lowest skill quintile gained 36% against 15% on average; below-median consultants 43% against 17%; translation gains were four times larger for slower workers.
  2. 02
    The exceptions are real and setting-specific.Kenyan entrepreneurs who started weak did nearly 10% worse with an AI assistant. Experienced open-source developers took 19% longer in METR's 2025 trial. Less experienced Italian developers produced more when ChatGPT was banned.
  3. 03
    Task type matters as much as the person.Consultants outside AI's frontier were 19 points less likely to be right. Agentic tool-use tasks gained far less than analytical ones. Maintenance coding rose more than new functionality.
  4. 04
    It has not yet shown up in pay in the one study that measured it.Danish administrative data covering about 25,000 workers rule out earnings effects over 2% two years after ChatGPT, with the null also holding for intensive users and early-career jobs. Between 32% and 57% of workers say they hide their AI use.

01 — The tableThe census: 20 studies that split the effect by worker or task

Rows are ordered roughly by how often they are cited. Peer-reviewed journal versions are marked; the rest are working papers or research-organisation reports. "Who gained" quotes the study's own split. Every figure was read at a primary source, which for several paywalled journals means the author's working-paper text or the journal abstract; the methodology section lists which.

Sources: the primary listed in each row, read September 25, 2026. Effects are on each study's own outcome and are not on a common scale.
StudyDesign and sampleAverage effectWho gained more or less
Generative AI at Work. Brynjolfsson, Li, Raymond. Quarterly Journal of Economics, 2025 (peer-reviewed).Quasi-experiment, staggered rollout of an AI assistant. 5,172 customer-support agents, one firm.+15% issues resolved per hourLowest skill quintile +36%. Most skilled: no productivity gain, small quality declines. Under one month's tenure +0.7 resolutions an hour; over a year, no effect.
Navigating the Jagged Technological Frontier. Dell'Acqua and colleagues (HBS/BCG). Working paper 24-013, September 2023.Pre-registered field experiment. 758 BCG consultants, GPT-4.Inside the frontier: +12.2% tasks, 25.1% faster, over 40% higher qualityBottom-half performers +43%, top-half +17% against own baseline. Outside the frontier, AI users were 19 points less likely to be correct.
Experimental evidence on the productivity effects of generative AI. Noy, Zhang. Science, July 2023 (peer-reviewed).Pre-registered online RCT. 453 college-educated professionals, writing tasks, ChatGPT.Time −40%, quality +18%Inequality between workers decreased. Short incentivised tasks, not on-the-job work.
The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. Peng and colleagues. arXiv, February 2023.Controlled experiment. 95 freelance developers, one JavaScript task.55.8% faster (95% CI 21% to 89%)Less experienced, older, and heavier-hours developers benefited most. Vendor-affiliated authors, small sample.
The Effects of Generative AI on High-Skilled Work. Cui and colleagues. Management Science, online February 2026 (peer-reviewed).Three company-run RCTs (Microsoft, Accenture, a Fortune 100 firm). 4,867 developers, Copilot.+26.08% completed tasks (standard error 10.3%)Less experienced developers adopted more and gained more. The tenure splits in the 2025 draft are noted by the authors as noisy.
Generative AI and labour productivity: a quasi experiment on coding. Gambacorta and colleagues (BIS). Journal of Financial Stability, 2026 (peer-reviewed).Matched treatment and control. 1,219 Ant Group programmers, 335 treated, 12 weeks.+55% lines of code; tasks completed +13% to +22%Gains significant only among junior or entry-level staff. Authors attribute the smaller senior effect to lower engagement, not lower usefulness.
Generative AI and the Nature of Work. Hoffmann and colleagues. HBS working paper 25-021, late 2024.Regression discontinuity on Copilot free-access eligibility. 187,489 open-source developers.Coding share of activity +12.4%; project management −24.9%Effects greater for lower-ability developers. Measures how time is allocated, not output.
The Impact of LLMs on Open-source Innovation. Yeverechyahu, Mayya, Oestreicher-Singer. arXiv, v4 May 2026.Natural experiment: Copilot supported Python, not R.Contributions +28% to +40%By task type: maintenance contributions rose more than new functionality. Ecosystem level, not per worker.
The Heterogeneous Productivity Effects of Generative AI. Kreitmeir, Raschky (Monash). arXiv, June 2024.Difference-in-differences around Italy's 2023 ChatGPT ban. 36,000+ GitHub users.No systematic effect on experienced developers' outputOpposite direction: when ChatGPT was removed, less experienced users produced more. Very short window.
The Uneven Impact of Generative AI on Entrepreneurial Performance. Otis and colleagues. Management Science, online July 2026 (peer-reviewed).Five-month field RCT. 640 Kenyan entrepreneurs, GPT-4 assistant.No significant average effect on revenue or profitOpposite direction: low performers nearly 10% worse, high performers may have gained over 15%. The gap came from which advice they chose to act on.
Lawyering in the Age of Artificial Intelligence. Choi, Monahan, Schwarcz. Minnesota Law Review, November 2024.RCT. Law students, four legal tasks, GPT-4. Sample size not confirmed at primary.Quality only slightly improved; large, consistent time savingsQuality gains, where any, went to the lowest-skilled. Time saved was roughly the same regardless of baseline speed.
AI-Powered Lawyering. Schwarcz and colleagues. Journal of Law and Empirical Analysis, April 2026 (peer-reviewed).Three-arm RCT (RAG tool, o1-preview, no AI). Upper-level law students.+50% to +130% productivity in five of six tasks; quality also improvedBy task type: largest on complex drafting and analysis. No split by student skill in the abstract.
Generative AI enhances individual creativity but reduces collective diversity. Doshi, Hauser. Science Advances, July 2024 (peer-reviewed).Online experiment. 293 writers, 600 evaluators, short stories.Novelty +8.1% with five AI ideasLess creative writers: novelty +10.7%, usefulness +11.5%, writing quality up to +26.6%, which equalises them with the most creative, who saw little effect.
Scaling Laws for Economic Productivity: LLM-Assisted Translation. Merali (Yale). arXiv, September 2024.Pre-registered online RCT. 300 professional translators, 1,800 tasks, 13 models.Per 10x model compute: speed +12.3%, earnings per minute +16.1%Low-skill (slower at baseline) time −21.1% vs high-skill −4.9%, about four times larger for lower-skilled workers.
Scaling Laws for Economic Productivity: Consulting, Data Analyst and Management Tasks. Merali. arXiv, December 2025.Pre-registered RCT. 500+ consultants, analysts and managers, 13 models.Each year of model progress: −8% task time. Pooled: quality +18%By task type: earnings per minute +$1.58 on analytical tasks vs +$0.34 on agentic tool-use tasks. No worker-skill split reported.
The Cybernetic Teammate. Dell'Acqua and colleagues. NBER working paper 33641, April 2025.Pre-registered 2x2 field experiment. 776 P&G professionals.Individual with AI +0.37 SD; team with AI +0.39 SD; time −16.4%Employees outside the core job, working alone with AI, matched teams that included a core-job expert.
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR. July 2025.RCT at the task level. 16 developers, 246 tasks.+19% longer with AI (CI +2% to +39%)Experts on their own mature codebases were slowed. They forecast a 24% speedup and afterwards believed 20%.
METR late-2025 follow-up and design change. METR blog, February 24, 2026.Same design. 57 developers, 143 repositories, 800+ tasks.Original cohort −18% time (CI −38% to +9%); new recruits −4% (CI −15% to +9%)30 to 50% of developers withheld tasks they did not want to do without AI. METR calls the estimate a lower bound. Time logging unreliable with concurrent agents.
Does Generative AI Narrow Education-Based Productivity Gaps? Cruces and colleagues. NBER working paper 34851, 2026.Randomised online experiment. 1,174 adults, business problem-solving task.AI raises performance for all participantsEducation gap 0.548 SD without AI, 0.139 SD with it. Part of the gap re-emerges once AI is removed.
Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI (first circulated as Large Language Models, Small Labor Market Effects). Humlum, Vestergaard. NBER working paper 33777, revised March 2026.Danish administrative data plus surveys. About 25,000 workers, 11 exposed occupations.Precise null on earnings and hours; rules out effects over 2% two years after ChatGPTThe null holds for the subgroups tested, including intensive users, early adopters, workers reporting large gains and early-career jobs. Measures pay, not task output.
Excluded: the withdrawn study

The MIT materials-science paper by Toner-Rodgers, which claimed top researchers nearly doubled output while the bottom third gained little, was a widely cited "experts gain more" result of 2024. On May 16, 2025 MIT said it had no confidence in the provenance, reliability or validity of the data, and arXiv marks it withdrawn. It is not in the table and should not be cited as evidence.

02 — The findingThe pattern the rows support

Where a study splits workers by skill, experience or prior performance, the larger gain usually goes to the weaker group. The chart shows the three rows with the cleanest paired figures.

Reported gain, lower-skilled vs higher-skilled group

Rows 1, 2 and 14 of the census. Each pair is on that study's own outcome; bars are not comparable across studies.
Support agents, lowest skill quintileBrynjolfsson, Li, Raymond, QJE 2025
+36%
Support agents, averageSame study
+15%
Consultants, below medianDell'Acqua et al., 2023
+43%
Consultants, above medianSame study
+17%
Translators, slower at baseline (time saved)Merali, 2024, per 10x compute
−21.1%
Translators, faster at baseline (time saved)Same study
−4.9%

The mechanism most authors propose is diffusion of best practice. The support-agent paper describes the assistant as spreading what the best agents already did. The creativity study finds the tool "effectively equalizes" weaker and stronger writers. The P&G experiment finds employees outside the core job, working alone with AI, matching teams that included an expert.

Part of the senior gap is adoption rather than usefulness. The Ant Group study attributes the smaller effect on senior programmers to lower engagement, with acceptance rates the same across experience levels. The three-firm Copilot trials found short-tenure developers more likely to keep using the tool.

Speed gains are more uniform than quality gains. In the first legal study, time saved was roughly the same regardless of baseline speed while quality gains went to the weakest. The most skilled support agents got small speed gains and small quality declines. And task type matters as much as the person: consultants outside the frontier got worse, agentic tasks gained less than analytical ones, and maintenance coding rose more than new features.

03 — The limitsWhat the pattern is not

  • Not "AI makes experts worse". Most rows show smaller gains for experts, not losses. Losses appear in four settings: small quality declines for top support agents, the outside-the-frontier consulting task, the Kenyan trial, and METR's experienced developers.
  • Not settled for agentic tools. Most rows used 2022 to 2024 chat or autocomplete tools. The one agentic-era RCT is flagged as biased by its own authors, and the one study that compared agentic with analytical tasks found agentic tool-use tasks gained least.
  • Not evidence of higher pay for those who gain. No row links a worker's productivity gain to that worker's earnings. The Danish administrative study finds no pay effect at all.
  • Not a universal levelling law. The effect narrows gaps in assisted output. The education-gap study found the gap partly re-emerged once AI was removed, and two rows run the other way.
  • Not one comparable number. Outcomes range from resolutions per hour to lines of code, graded quality and revenue. Averages from −19% to +55% are not on a common scale, and we do not average across studies.
  • Not representative of all work. Eight of the twenty rows are coding. Two have vendor-affiliated authors.

04 — Kept apartUsage and self-report sources, kept apart from the census

These sources are cited as often as the studies above and measure something different: what people say they saved, or what a vendor's logs show. They are here so they are not confused with causal estimates.

Not causal evidence. Read September 25, 2026 at each source.
SourceFigure relevant to who gainsCaveat
St. Louis Fed, Bick, Blandin, Deming. February 2025, updated November 2025. Nationally representative US survey.Users saved 5.4% of work hours; across all workers 1.4% (1.6% in the 2025 update). 33.5% of daily users saved four hours or more.Self-reported counterfactual time.
METR AI usage survey, May 2026. 349 technical workers.Median self-reported 1.4 to 2x value uplift and 3x speed.METR's own 2025 RCT found people overestimated time savings by about 40 percentage points.
Anthropic Economic Index, March 2026. Claude usage data.Users with six or more months' tenure have a 10% higher success rate.Correlational, and points the opposite way: experience with the tool helps.
Anthropic, Estimating AI productivity gains, November 2025.Claude estimates AI cuts task time by 80%, from 56% to 90% by task.The model's own estimates, excluding human validation time.
Microsoft 2026 Work Trend Index, May 2026. 20,000 AI-using knowledge workers. Vendor-commissioned.16% of AI users are Frontier Professionals; they share tips and agents more (61% vs 36%). Only 13% of all AI users say they are rewarded for reinventing work with AI even if results aren't met; 45% say it feels safer to focus on current goals.Self-classified. Frontier Professionals are twice as likely to say they are rewarded.

The gap between the two tables is the point of our post on why AI adoption statistics disagree. A self-reported 3x speed-up and a measured 19% slowdown can describe the same developers, as METR's two exercises show.

05 — DisclosureWho hides AI use, and why

The census measures gains the researchers could see. A large share of workers say they keep their AI use from their employer, which means employer-visible gains may understate who is benefiting, or attribute them to the wrong person. The figures are not comparable with each other, because each survey asked a different question.

Question wording differs: "reluctant to admit", "uncomfortable telling a manager", "hide and present as own", "keep secret".
SurveyFinding
Microsoft and LinkedIn 2024 Work Trend Index, May 2024. 31,000 people, 31 countries. Vendor-commissioned.52% of people who use AI at work are reluctant to admit using it for their most important tasks. 53% worry it makes them look replaceable.
Slack Workforce Lab, Fall 2024 Workforce Index, November 2024. 17,372 desk workers, 15 countries. Vendor-commissioned.48% would be uncomfortable admitting to their manager that they used AI. Reasons: it feels like cheating (47%), fear of looking less competent (46%) or lazy (46%).
University of Melbourne with KPMG, Trust, attitudes and use of AI, April 2025. 48,000+ people, 47 countries. Academic-led, KPMG-funded.57% of employees hide their use of AI and present AI-generated work as their own. 66% rely on output without checking accuracy; 56% have made mistakes because of AI.
Ivanti 2025 Technology at Work, May 2025. 6,000+ office workers, 1,200 IT staff. Vendor survey.32% of generative AI users keep their use secret. Reasons: the secret advantage (36%), fear of job cuts (30%), imposter syndrome (27%).

What an employer should do with a hire who is quietly twice as productive is the subject of our companion post on two hires and one time tracker. The broader statistics on agent productivity and return are in our agent productivity data points.

Methodology

A census of rigorous studies that measure differential effects, compiled from the primary listed in each row.

What was collected
Randomised trials, field experiments, quasi-experiments and administrative-data studies public by September 23, 2026 that report how AI productivity effects differ by worker skill, experience, prior performance or education, or by task type.
Sources
Journal abstracts and author working papers. Where a publisher page blocked automated reading, the figure was verified at the nearest primary: the arXiv, NBER, BIS or institution-hosted PDF, or the journal abstract via OpenAlex or Crossref metadata.
As-of date
Sources were read on September 25, 2026.
Exclusions
Usage logs and self-report surveys (kept in a separate table). Exposure studies that estimate which jobs could be affected rather than measured productivity. The withdrawn MIT materials-science paper.
Known limitations
Effects are on each study's own outcome and are never averaged. Eight rows are coding studies. The tenure split in the three-firm Copilot study comes from the authors' 2025 draft. The law-student sample size in the 2024 study was not confirmed at a primary.
Refresh
Refreshed in place when a new study with a worker or task split is published, or when a listed working paper reaches a journal.

07 — ConclusionThe gains are uneven, usually favour the less experienced, and have not reached pay

How to use this table

Cite the row, not the average, and match the study's task and tool to the work you are planning for

A manager setting expectations for an AI rollout should expect the least experienced staff to move most on routine, in-frontier work, experts to move little or to be slowed on work they already do well, and the whole effect to depend on task type. A writer citing "AI helps novices most" should cite a row and its caveat. Teams that want this evidence turned into a rollout plan can see how we approach it in our AI transformation practice.

Digital Applied

Plan an AI rollout on measured effects, not vendor averages.

We design pilots, control groups and output measures so your organisation learns who gains from AI on your own work before it sets targets.

Pilot designOutput measurementRollout planning
Your next project

An AI rollout with a control group

  • →A measured baseline per role
  • →Task types matched to the evidence
  • →Targets set after the pilot, not before
Questions and answers

The questions we get about who gains from AI

In most measured studies, beginners. Twelve of the twenty in this census find larger gains for less experienced or lower-performing workers, such as 36% for the lowest skill quintile of support agents against 15% on average. Three find the opposite or show experts slowed, so the answer depends on the task.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading