Rising AI usage tells you that more work is passing through an AI system. It does not tell you whether the team is completing more useful work, spending less effort or moving its bottleneck into review. Before expanding access, connect usage to a defined set of accepted tasks and keep unfinished work visible.
This guide proposes a decision record for a team considering wider AI access. It does not estimate a universal productivity effect or present a new company dataset. The question is practical: what evidence would justify the next expansion, and what would tell you to improve the workflow first?
- 01Count outcomes with acceptance rules.A generated artifact or finished model response is not automatically accepted work.
- 02Keep comparable tasks together.Changing task difficulty, staffing or review standards can distort a before-and-after comparison.
- 03Decide where to expand.Useful evidence can support a specific task class without supporting a company-wide rollout.
01 — Practical guideDefine the expansion decision before choosing metrics
Write the proposed change in concrete terms: which people, which task class, which access level and which period. “Use more AI” is too broad to evaluate. Giving another design team access to a coding agent is a different decision from allowing an existing agent to execute external actions.
Define what would count as accepted work for that task. A reviewed change merged into the codebase, a checked report delivered to its intended reader and an approved design artifact are distinct outcomes. Do not combine them into a raw output count unless the grouping is justified.
The SPACE productivity framework explains why activity cannot stand in for productivity and why more than one dimension is necessary. Token consumption is an activity signal; it needs context from completed work and the effort required to accept it.
02 — Practical guideLink usage to the work item people recognize
Use a stable task identifier across agent runs, review and acceptance. One task may involve several prompts, retries and agents. Keep those attempts together so a workflow that needs repeated repair is not credited with several completed tasks.
Record the outcome at the end of the observation period, including work still pending. If the agent produced many drafts and reviewers accepted only some, the unfinished queue is part of the result. Carry it forward explicitly rather than letting it vanish at the reporting boundary.
Keep tool usage in its native units and preserve model/version context when it matters. A tokenizer or model change can alter token counts without an equivalent change in the work. Compare the accepted task and its total effort before interpreting consumption trends.
| Record | Why it matters | Minimum evidence |
|---|---|---|
| Task and cohort | Defines the unit and comparable group. | Task ID, type, difficulty criteria and period. |
| Attempts and usage | Avoids crediting retries as extra work. | Linked runs, model configuration and usage records. |
| Human effort | Shows where work moved. | Preparation, review, correction and acceptance effort. |
| Elapsed time | Shows queues as well as execution. | Start, review-ready and accepted timestamps. |
| Quality and rework | Prevents speed from hiding downstream repair. | Acceptance check, rejection reason and reopened work. |
| Unfinished work | Prevents incomplete tasks disappearing. | Pending status, owner and reason at period end. |
03 — Practical guideCompare task groups without hiding selection effects
Where feasible, compare similar tasks under a predeclared assignment rule. Keep task type, experience, review standards and observation windows as consistent as possible. Record when participants choose whether to use AI, because that choice can be related to difficulty or expected benefit.
METR’s February 2026 experiment update describes selection effects as more developers became unwilling to work without AI. It also identifies problems measuring task time when developers use multiple agents concurrently. The researchers caution that their newer data does not reliably establish the current effect size.
The lesson for a team is methodological, not a borrowed productivity percentage. Record who and what is missing from the comparison. An AI-assisted cohort consisting only of enthusiastic users on easy tasks cannot support the same claim as a broader, controlled assignment.
When a controlled comparison is impractical, use an observational record and say so. A before-and-after trend can support a narrower operational decision, but staffing, task mix or process changes may explain part of the difference.
04 — Practical guideCheck whether review became the constraint
An agent can increase draft production while accepted output remains flat. That does not necessarily mean the agent is useless; it may mean the process now produces work faster than reviewers can assess it. Look at queue age, review effort and the reasons for rejection before buying more generation capacity.
Also inspect rework after acceptance. If a change passes initial review but creates later repair, the first acceptance event alone overstates its contribution. Choose a follow-up window appropriate to the task and retain reopened work rather than quietly reclassifying it as a new task.
Our agent ROI measurement guide covers broader value and cost models. This decision record stops earlier: is there enough evidence that the proposed group can turn additional AI access into accepted work without an unacceptable review burden?
05 — Practical guideChoose expand, repair or keep learning
Expand the specific task class when acceptance and quality hold, the total effort is acceptable and there is capacity to review the extra work. Keep the conclusion scoped to the people and tasks examined. A result for routine UI changes does not automatically justify the same access for financial reporting or customer-facing actions.
Repair the workflow first when failures cluster around missing context, unclear requirements or review congestion. Better input or a clearer acceptance rule may be more useful than a higher model tier. If the evidence is incomplete, extend the observation rather than converting uncertainty into a confident rollout claim.
In a hypothetical team, more drafts with a growing unreviewed queue would support examining review capacity. Faster accepted work with stable quality could support a bounded expansion. These are decision examples, not measured team outcomes.
The tool-result context reference provides one specific intervention to test when agents repeatedly reread irrelevant records. The managed Agents API guide covers an infrastructure change that may alter engineering effort without changing the accepted task itself.
06 — Practical guideWrite a short decision record with a recheck date
Include the proposed expansion, task cohort, acceptance rule, observation window, comparable baseline, result and limitations. Record the owner of the decision and the condition that would cause you to reconsider it. A decision to keep learning should identify the missing evidence.
Download the blank expansion decision record to begin. It is a template, not an ROI calculator or populated dataset. Use it to connect the usage report to artifacts and review evidence a colleague can inspect.
Keep the record proportional to the choice. A small access expansion may need a brief review; a change that permits external actions needs evidence about those actions too. Our AI transformation service can help frame the pilot around a task the business already knows how to accept.
Evidence and scope
- As-of date
- September 12, 2026. September 10 is the editorial allocation; current documentation was reviewed later.
- Method
- Primary documentation and research were reviewed for the cited distinctions. Tables, worksheets and pilot checks are Digital Applied proposed methods, not observed deployment results.
- Limitations
- No production API workflow, vendor benchmark or participant study was executed for this article. Documentation can change; verify the selected configuration before implementation.
07 — Next stepPut the decision into practice
Define the expansion decision before choosing metrics
Choose the next access decision, then collect the accepted work and effort that would justify it. Consumption can tell you where to investigate; the decision should rest on what the team can deliver and sustain.