01The same technology produces +14% and −19% in different settings

Empirical studies measuring AI's productivity effects have accumulated rapidly since 2023. The results point in different directions. Brynjolfsson, Li, and Raymond studied 5,179 customer support agents and found that AI tools raised the number of issues resolved per hour by 14% on average. Among novice agents, the improvement was 34%. In contrast, METR's 2025 experiment tracked 16 experienced developers working on 246 real issues in large open-source repositories. Those who used AI took 19% longer than those who did not.

Both findings are valid. The question is not whether AI works but under what conditions it works.

Figure 1 The conditions that determine AI's productivity effect
CombinedeffectConditionsalignConditions reversedTaskstructureUser skilllevelStructured +noviceLargest gains(+14–34%)Unstructured +expertNo gain orslowdownCombined effectConditions alignConditions reversedTask structureUser skill levelStructured + noviceLargest gains (+14–34%)Unstructured + expertNo gain or slowdown
AI's productivity effect depends on the combination of task structure and user skill level. Uniform deployment produces inconsistent results.

02Effect size depends on task structure and user skill level

Lining up the firm-level experiments reveals two axes that determine effect size: how structured the task is, and how skilled the user is.

In Brynjolfsson et al.'s customer support study, the gains concentrated among novice agents. Top performers saw almost no improvement. The AI functioned as a mechanism for disseminating best practices, pulling less experienced agents toward the level of their more experienced peers.

A joint study by Harvard Business School and BCG gave 758 consultants access to GPT-4. On tasks within AI's capability frontier, completion speed rose by 25.1% and output quality improved by 40%. But on tasks outside that frontier, the AI-assisted group's accuracy fell by 19%. Dell'Acqua, McFowland, Mollick, and their co-authors called this boundary the "jagged technological frontier." The line between what AI handles well and what it does not is irregular and hard for users to predict.

Software development shows the same pattern. Peng, Kalliamvakou, and colleagues had 95 developers implement an HTTP server—a well-defined task. Those using GitHub Copilot finished 55.8% faster. Cui, Demirer, and co-authors ran experiments across three companies with 4,867 developers and found a 26% increase in completed tasks. These were relatively structured coding assignments. By contrast, METR's experiment targeted real issues in large repositories—unstructured tasks requiring deep contextual understanding.

StudySubjectsTask typeProductivity effect
Brynjolfsson et al. (2025)5,179 support agentsStructured, scripted responses+14% average, +34% for novices
Dell'Acqua et al. (2025)758 consultantsKnowledge work (inside frontier)+25.1% speed, +40% quality
Cui, Demirer et al. (2024)4,867 developersStructured coding tasks+26% completed tasks
METR (2025)16 developersReal issues in large OSS repos−19% (slowdown)

03Task-level efficiency does not automatically translate to organizational productivity

Firm-level experiments measure task-level efficiency. Faster individual tasks do not necessarily mean higher productivity for an entire organization or an economy. Three mechanisms intervene.

First, task allocation shifts. When AI accelerates structured tasks, workers spend more time on unstructured ones. If unstructured tasks have lower productivity, aggregate efficiency may stagnate or decline. Second, coordination costs increase in the short term. Introducing new tools requires changes in training, review processes, and meeting structures. Third, there is a quality-speed tradeoff. In the BCG experiment, the diversity of consultants' output fell by 41% when they used AI. Faster output is less valuable if it is also more uniform.

Figure 2 The gap between task-level efficiency and organizational productivity
Task levelOrganization/ macro…Individual taskspeedupMeasured in experimentsTask reallocationShift to unstructuredworkQuality/diversitydeclineUniform outputCoordination costsTraining, review changesInstitutionalrigidityReimbursementstructures etc.Macro statisticsDelayed or unclearimpactTask levelOrganization / macrolevelIndividual taskspeedupMeasured inexperimentsTask reallocationShift to unstructuredworkQuality/diversitydeclineUniform outputCoordinationcostsTraining, reviewchangesInstitutionalrigidityReimbursementstructures etc.Macro statisticsDelayed or unclearimpact
Task-level efficiency gains are mediated by coordination costs and institutional rigidity before reaching macro productivity statistics.

04Macro estimates range from 0.5% to 7% of GDP—a tenfold gap

When task-level evidence is translated into macroeconomic estimates, projections diverge sharply.

Daron Acemoglu, in a paper published in Economic Policy in 2025, estimated that AI would raise total factor productivity by no more than 0.53% over ten years. His assumptions are threefold: only about 5% of tasks can be profitably automated by AI; task-level cost savings are drawn conservatively from existing experiments; and results from "easy-to-learn" tasks are not extrapolated to "hard-to-learn" tasks.

Goldman Sachs estimated that AI could raise global GDP by 7%. The assumptions differ substantially. They project that 7% of US employment will be replaced by AI and 63% will be complemented by it. They assume annual labor productivity growth rises by 1.5 percentage points, cumulated over a decade. McKinsey Global Institute projected $2.6 to $4.4 trillion in annual economic value, with labor productivity gains of 0.1 to 0.6% per year.

Assumption 01

Share of automatable tasks

Acemoglu: ~5% / Goldman Sachs: 7% replaced + 63% complemented

Acemoglu limits the count to tasks where AI substitution is profitable. Goldman Sachs includes both substitution and complementarity, broadening the scope of affected tasks considerably.

Assumption 02

Cost savings per task

Acemoglu: conservative extrapolation / Goldman Sachs: broad productivity gains assumed

Acemoglu restricts experimental findings to easy-to-learn tasks and avoids extrapolating to hard-to-learn ones. Goldman Sachs incorporates expected technological progress into savings rates.

Assumption 03

Adoption speed

Acemoglu: limited in 10 years / Goldman Sachs: broad adoption in 10 years

Adoption rates differ dramatically across countries, industries, and firm sizes. This assumption alone accounts for a large share of the gap between estimates.

Assumption 04

Treatment of new tasks

Acemoglu: excluded / Investment banks: implicitly included

Acemoglu covers only substitution and complementarity of existing tasks. He excludes the creation of entirely new tasks or industries by AI. This conservatism lowers his estimate.

05The gap between estimates reflects different questions, not different errors

Acemoglu's 0.53% and Goldman Sachs' 7% are not different answers to the same question. They answer different questions.

Acemoglu asks: what happens if current AI technology is applied to existing tasks within the current economic structure? Goldman Sachs asks: what happens if AI technology improves over a decade, adoption becomes widespread, and economic structures adapt? The former approximates a lower bound. The latter is a conditional upper-bound scenario.

Which turns out to be closer to reality depends on three uncertainties: the pace of technological progress, the pace of adoption, and the pace of organizational adaptation. Viewed through the J-curve hypothesis covered in the first article of this series, the speed of complementary investment is the key variable. Acemoglu's estimate corresponds to a world where complementary investment is slow. Goldman Sachs' corresponds to a world where it is fast.

06Healthcare and drug development have high shares of structured tasks and stand to gain early

Applying the findings from firm-level experiments to specific industries, the sectors most likely to see early productivity gains are those with a high proportion of structured tasks and large existing inefficiencies.

Healthcare fits both criteria. Radiological image interpretation is a highly structured task; FDA-cleared AI medical devices exceeded 1,000 by the end of 2024. Clinical documentation, billing, and test result summarization are also highly structured.

In drug development, target identification and compound screening are the stages where AI's effects are most pronounced. As of 2026, more than 200 AI-derived compounds are in clinical stages. AI-discovered candidates have reported Phase I success rates of 80–90%, compared to 40–65% for conventionally discovered candidates. These figures come from early-stage data, and whether later-stage success rates will hold is not yet established.

At the same time, some tasks in these fields resist AI acceleration. Complex diagnostic reasoning, treatment planning that requires nuanced patient context, and patient communication are unstructured tasks—the kind where METR's experiment suggests AI may impose net costs rather than savings.

Figure 3 An evidence-based deployment sequence
Identifyhigh-effect…AlongsidedeploymentScale onevidenceSort tasksDraw thestructured/unstructured…Deploy tonovices firstTarget the groupwith most room to…Build qualitymonitoringTrack diversityand accuracyRun smallexperimentsScale based onyour own dataIdentify high-effectareasAlongside deploymentScale on evidenceSort tasksDraw the structured/unstructured boundaryDeploy to novices firstTarget the group with most room to growBuild quality monitoringTrack diversity and accuracyRun small experimentsScale based on your own data
Investment decisions should be based on organization-level experiments, not macro estimates. Estimates are scenarios, not forecasts.

07Extracting productivity gains requires designing the conditions on your side

What the empirical studies consistently show is that AI's effect is determined by conditions on the adopter's side. Technological capability alone does not raise productivity.

First, tasks must be sorted. The BCG experiment's distinction between tasks inside and outside the frontier must be applied to an organization's own operations. Which tasks are structured enough for AI, and which are too unstructured? Deploying AI uniformly without this sorting produces inconsistent results and, in some cases, slowdowns.

Second, there is a rational case for prioritizing deployment among less experienced workers. The largest effects in Brynjolfsson et al.'s study appeared among novices. If the goal is raising organization-wide productivity, targeting the group with the most room for improvement offers the highest return per unit of investment.

Third, quality monitoring must be built into the workflow. The BCG experiment showed a 41% decline in output diversity, and accuracy dropped on tasks outside the frontier. AI output adopted without verification degrades the quality of organizational decisions.

Fourth, macro estimates should not drive firm-level investment decisions. Whether global GDP rises by 0.53% or 7% tells an individual organization nothing about its own situation. The actionable path is to design small-scale experiments, measure task-level effects, and scale based on data. Estimates are scenarios, not forecasts.

Key Points ── 3 to take away
  1. AI's productivity effect varies sharply with task structure and user skill level. Structured tasks and less experienced users see the largest gains (+14 to +34%), while unstructured tasks and experienced users may see no benefit or even slowdowns (−19%). The "jagged frontier" identified in the BCG experiment means the boundary between effective and ineffective use is irregular and hard to predict.
  2. Macro GDP estimates ranging from 0.53% to 7% reflect different assumptions, not different errors. The share of automatable tasks, per-task cost savings, adoption speed, and treatment of new tasks all diverge between conservative and optimistic frameworks. These numbers should be read as scenarios.
  3. Task-level efficiency gains do not automatically become organizational productivity gains. Sorting tasks, prioritizing deployment among less experienced workers, building quality monitoring into workflows, and running small-scale experiments are the practical conditions for extracting value from AI.
Closing

The empirical evidence on AI and productivity has moved beyond a simple "yes or no." Effects exist, but they are uneven. Structured tasks show double-digit improvements; unstructured tasks show slowdowns. The tenfold gap in macro estimates stems from how this unevenness is aggregated. For any organization, what matters is not whether global GDP will rise by some percentage, but which of its own tasks AI accelerates and which it does not. That knowledge comes not from published estimates but from experiments run on its own operations.

Sources & references
  1. Brynjolfsson, E., Li, D. & Raymond, L. R. Generative AI at Work. The Quarterly Journal of Economics, 2025. https://www.nber.org/papers/w31161
  2. Dell'Acqua, F. et al. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Organization Science, 2025. https://pubsonline.informs.org/doi/10.1287/orsc.2025.21838
  3. Cui, Z., Demirer, M., Jaffe, S., Musolff, L., Peng, S. & Salz, T. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers. 2024. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4945566
  4. METR. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. 2025. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
  5. Peng, S., Kalliamvakou, E., Cihon, P. & Demirer, M. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. 2023. https://arxiv.org/abs/2302.06590
  6. Acemoglu, D. The Simple Macroeconomics of AI. Economic Policy, 40(121), 13–58, 2025. https://academic.oup.com/economicpolicy/article-abstract/40/121/13/7728473
  7. Goldman Sachs. The Potentially Large Effects of Artificial Intelligence on Economic Growth. 2023. https://www.goldmansachs.com/insights/articles/generative-ai-could-raise-global-gdp-by-7-percent
  8. McKinsey Global Institute. The Economic Potential of Generative AI: The Next Productivity Frontier. 2023. https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/the-economic-potential-of-generative-ai-the-next-productivity-frontier