Two people join a digital agency in the same week, on the same pay, in the same role. Each gets the same first test: produce 100 content pieces. One opens a chat assistant and works through them one at a time. It takes ten working days, and the time tracker logs ten full, active days. The other spends two days building a pipeline that drafts, checks and formats in batches, then reviews the output. That hire finishes in five days, and the tracker logs two busy days and three quiet ones.
This is a thought experiment. No agency, client or person in it is real. It is here because the question it raises is real for every business that pays people by time or by visible activity, and because the answer is not obvious once you look at what the tools record and what the research says.
The post takes four questions in order. What does the tracker actually see? What does the evidence say about who gains from AI and why the gap is uneven? What can the agency do about the faster hire, and what does each option teach the next hire? And what would a fair measurement look like?
- 01Activity-based trackers score keyboard and mouse input, so the slower method looks more diligent.Hubstaff's own documentation defines activity as active seconds divided by 600 per ten-minute block, and says a 25% score and a 75% score can both be productive. A pipeline that runs while its builder reviews output scores low.
- 02The studies say gains from AI are large, uneven, and often largest for the least experienced.Customer-support agents gained 15% on average and the lowest skill quintile 36%. Consultants below the median improved 43% against 17% above it. In one field trial, low performers did worse. Experienced developers in METR's 2025 trial took 19% longer.
- 03Every reward option teaches the builder something, and 'more work at the same pay' teaches them to hide the method.Frederick Taylor documented in 1911 that workers whose piece rate was cut after working faster learned to conceal how fast work could be done. The incentive has not changed.
- 04Fair measurement counts accepted output and its quality per unit of time, not activity.The tracker is not wrong about what it measures. It measures the wrong thing for work that a pipeline does. Change the unit before you change the pay.
01 — The scenarioThe scenario, and what the dashboard makes of it
The table below is the thought experiment written out. The numbers are the scenario's, not measurements from any real team. The point is the last row: two hires with the same output, one of them twice as fast, and the tracker's summary favours the slower one.
| What the tracker records | Hire A, chat window | Hire B, pipeline |
|---|---|---|
| Calendar days to deliver 100 pieces | 10 | 5 |
| Tracked days logged | 10 | 5 |
| Typical activity score | High all day: typing and clicking in a chat window | High for two days of building, then low while the pipeline runs |
| What screenshots show | Drafts in progress, every interval | A terminal, a spreadsheet, a review queue |
| Pieces per tracked day | 10 | 20 |
| How the dashboard reads it | Diligent | Idle for half the week |
02 — The instrumentWhat the tracker sees
Time-tracking products fall into three groups, and only one of them has the problem in the scenario. Timesheet-only tools such as Toggl Track and Harvest record hours and refuse, in their own published statements, to capture screenshots or monitor activity. Newer "work intelligence" products such as ActivTrak and Insightful sell measurement of AI adoption itself. The third group, which includes Hubstaff, Time Doctor and Clockify, scores activity from input devices and takes periodic screenshots. Everhour takes optional screenshots but says it records no keystrokes or other activity.
Hubstaff's support documentation is the clearest statement of how that score works. Activity is the share of seconds in each ten-minute block with any keyboard or mouse input: active seconds divided by 600. The same page carries the caveat that matters for our scenario. It says that "people with 75% scores and those with 25% scores can often times both be working productively", that meeting-heavy and research-heavy roles score lower, and that Hubstaff discourages "the use of quotas or scoring based on activity rates". The vendor is telling its customers not to do the thing the dashboard invites.
Freelance marketplaces build pay on the same signal. Upwork's Hourly Payment Protection, per its help centre, requires that screenshot activity relates to the contract, that memos describe the work, and that the freelancer "maintained adequate and fair activity levels". A contractor whose pipeline does the work while they review it can fail that test with the invoice fully earned. Fixed-price contracts with milestones in escrow avoid it, which is one reason the output-based option in section 04 is not exotic.
One more figure shows how little of AI's value the trackers can see. Hubstaff's own 2026 report on its platform data, covering about 140,000 workers at 17,000 organisations, says the share of its users on AI tools rose from 65% to 73% between 2024 and 2025 while time spent in AI apps slipped from around 4% to 3% of tracked time. That is vendor platform data, not a study, but the direction is the scenario's: more people using AI, less visible time doing it.
Across the nine vendor pages reviewed for this post, none offers accepted output per hour, or work completed by an agent on the user's behalf, as a first-class metric. The instruments measure inputs. A method that removes inputs looks, to them, like absence.
03 — The evidenceWhat the studies found about who gains
The scenario assumes the pipeline builder is simply better. The research says the gap between two people using AI is usually large, and that which person gains more depends on the task and on what each chooses to do with the tool. Four headline studies, each with its sample and date. The full census of twenty is in our companion post on who gains most from AI at work.
- Customer support agents, one firmBrynjolfsson, Li and Raymond, Quarterly Journal of Economics 2025. 5,172 agents, staggered rollout of an AI assistant.
- +15% averageLowest skill quintile +36%Most skilled: no productivity gain
- Management consultants, GPT-4Dell'Acqua and colleagues, HBS working paper 24-013, September 2023. 758 consultants, pre-registered field experiment.
- +12.2% tasks, 25.1% fasterBelow-median performers +43%, above-median +17%Outside AI's frontier: 19 points less likely to be correct
- Small-business owners, GPT-4 assistantOtis and colleagues, Management Science 2026. Five-month field trial with Kenyan entrepreneurs.
- No average effectHigh performers may have gained over 15%Low performers did nearly 10% worse
- Experienced open-source developersMETR, July 2025. 16 developers, 246 tasks, randomised at the task level. Developers had forecast a 24% speedup.
- +19% longer with AIConfidence interval +2% to +39%Afterwards they still believed they had been 20% faster
Two things in that list bear on the scenario. First, the direction flips by task. In the customer-support and consulting studies, the least experienced gained most; the support-agent authors attribute this to AI spreading the habits of the best workers. In the Kenyan trial the weaker performers lost ground because of the advice they chose to act on. Second, people misjudge their own gain. METR's developers were slower and believed they were faster. A tracker that scores activity will not correct that belief; it will confirm it.
The support-agent paper's published abstract, at the Quarterly Journal of Economics, adds a detail worth holding onto: the most skilled agents saw small gains in speed and small declines in quality. The tool did not help everyone, and where it helped least it also cost a little. METR's 2025 write-up makes the same point about experts on codebases they know well.
None of these studies measured a pipeline builder against a chat user. They measured people with and without a tool, at a single sitting or over months. The scenario's gap, one person automating a batch while another works by hand, is a gap in method rather than in access, and the published trials mostly predate the agentic tools that make it possible. Our earlier piece on the productivity paradox in developer work covers why measured gains lag the tools.
04 — The decisionThe agency's options, and what each one teaches
The agency now knows one hire finished in half the time. Whatever it does next is a message to that hire and to the next one. The table lists the usual responses with the incentive each creates. The last column is the part managers skip.
| Option | What the builder gets | What the agency gets | Incentive it creates |
|---|---|---|---|
| More work at the same pay | Five more days of tasks | Double the output for the same salary | Hide the method next time. This is the piece-rate cut Taylor described in 1911. |
| One-off bonus | A payment tied to this test | Output now, no ongoing commitment | Build once, then wait to see whether the second bonus arrives before improving again. |
| Raise or promotion | Higher base pay for the same hours | Keeps the builder, and pays more whether or not the method spreads | Keep building. Nothing yet rewards teaching others. |
| Time back | Five days of the week returned | Same output as before, no cash cost | Finish fast and stop. Output stays flat at the old baseline. |
| Output-based contract | Paid per accepted piece, at a rate that does not fall | Pays for results, not hours | Maximise accepted pieces. Quality gates become the whole contract. |
| Share of saved hours | A stated fraction of the five saved days, as cash or time | The rest of the saving, and a reason for the builder to keep finding more | Keep improving, and share the method if sharing is what earns the fraction. |
The first option is the default, and it is the one with a century-old record. Frederick Winslow Taylor's 1911 book on scientific management describes workers who had their price per piece cut after working faster, and who responded by making sure the employer never learned how fast the work could really be done. He called it systematic soldiering. The full text is on Project Gutenberg.
The greater part of the systematic soldiering, however, is done by the men with the deliberate object of keeping their employers ignorant of how fast work can be done.Frederick Winslow Taylor, The Principles of Scientific Management, 1911
The output-based option has its own evidence. Edward Lazear's study of Safelite Glass, published in the American Economic Review in 2000, found that moving from hourly pay to piece rates raised output per worker by 44%, attracted more able workers, and widened the spread of output across individuals. The scenario's hires already differ by a factor of two. A per-piece contract would make that visible in pay, which is the point, and would also make quality gates the entire contract, which is the risk.
The share-of-saved-hours option is the old gainsharing idea: the firm sets a standard for how long a job should take, and when it takes less, a stated fraction of the saving goes to the people who made it faster. Its advantage over a raise is that it pays for the next improvement as well as this one. Its advantage over a bonus is that the fraction is known in advance, so the builder is not guessing whether honesty will be rewarded twice. Our agency pricing guide covers the client-facing side of the same shift.
05 — The fixWhat a fair measurement looks like
The tracker in the scenario is not lying. It measures input, and the pipeline removed input. A fair measurement for work that machines can do in batches has to count what left the building, not what happened at the keyboard. Four properties, in order of importance.
Count accepted output
A piece counts when it clears review, not when it is drafted. This is the only number both hires can be compared on, and it is the number the client pays for.
Score quality per unit
Track revision rate and rejection rate per piece. A pipeline that ships faster and gets sent back more often has not saved time; it has moved it to the reviewer.
Divide by elapsed time, not active time
Calendar days from brief to acceptance. Active-seconds scores penalise anyone whose method runs without them at the keyboard.
Record the method
Ask how the work was done and write it down. The method is the asset the agency wants to keep. It cannot be measured, but it can be named and credited.
Our content-team metrics guide sets out the acceptance and revision measures in detail. The change that matters is the unit. Once output per elapsed day is the score, the activity dashboard becomes what Hubstaff itself says it should be: a trend to glance at, not a number to manage by.
06 — The other seatIf you are the faster hire
The scenario ends with a choice for the builder as well. Show the method, or keep it. The surveys in our companion post say a large share of workers already keep AI use to themselves, for reasons that range from feeling like cheating to liking the edge. Taylor's workers had a sharper reason: showing speed got the rate cut.
Three practical moves before deciding. Keep your own record of accepted output per day, because the tracker will not. Ask, before the review, how output will be measured and what the agency does when someone finishes early; the answer tells you which row of the options table you are in. And put the method in writing under your own name, whether or not you share it yet. Credit is the one reward that does not depend on the agency choosing well. The question of whether the agency should ask you to teach everyone, and what it owes you if it does, is a separate post.
07 — ConclusionThe tracker rewards the slower method because it measures the wrong thing
Change the unit to accepted output per elapsed day, decide in advance what a faster hire earns, and write that rule down before the next test
The scenario is invented; the mechanism is not. Activity scores, hourly protection and screenshot review all reward input, and the most valuable thing an employee can now do is remove input. An agency that decides its reward rule before it needs one keeps the builder and the method. One that decides after, under the default of more work at the same pay, keeps neither. Teams building the measurement layer for AI-assisted work can see how we approach it in our AI transformation practice.