[research] · · 3 min read
Harvard study finds AI coding agents increase code volume but not software output
A large-scale analysis of 700 firms shows that while AI agents boost lines of code by 30%, human review bottlenecks absorb the efficiency gains, leaving overall software delivery unchanged.
By ByteBulletin Editor · Editor
The Efficiency Paradox
According to a new study published by Ars Technica, the widespread adoption of AI coding agents has created a paradox in software engineering: teams are generating significantly more code, but not necessarily more usable software. Researchers Fiona Chen and James Stratton from Harvard University analyzed aggregated analytics data from Jellyfish, covering 300 million individual work events across more than 700,000 employees at over 700 software development firms. The data spans from 2021 through March 2026, providing a comprehensive look at how AI tools have integrated into real-world engineering workflows. The central finding is stark: while AI agents drive a 30 percent increase in total lines of code generated and a 23 percent rise in pull requests, there is "little evidence that firms increase software output or reduce employment" as a result.
The Review Bottleneck
The study identifies the root cause of this stagnation in the code review process. As AI agents autonomously write and submit code, the volume of pull requests increases, but the quality or readiness of that code does not scale at the same rate. Consequently, the average time between a pull request being submitted and merged into the codebase balloons by 49 percent after the introduction of AI agents. This slowdown is driven by a significant increase in friction: the share of pull requests requiring changes nearly doubles, and the number of comments per pull request increases by 35 percent. The researchers describe this as a downstream constraint where efficiency gained in the coding phase is "absorbed by downstream constraints in the production process."
Employment and Workforce Impact
Contrary to fears of mass displacement, the study found no significant employment changes attributable to AI. By cross-referencing Jellyfish data with LinkedIn records, the researchers determined that total active workers at these firms did not decrease. Instead, the workforce adapted to the new workflow. There was a 14 percent increase in the share of workers performing code reviews, indicating that human oversight has become a more critical component of the development pipeline. The data suggests that rather than replacing engineers, AI tools have shifted the labor distribution toward verification and quality assurance.
The Limits of AI-Assisted Review
One might assume that AI could solve the review bottleneck by automating the inspection of AI-generated code. However, the data shows this is currently not the case. By March 2026, 80 percent of measured firms used some form of AI code review, yet AI agents were responsible for only 23.3 percent of all review comments and just 10.8 percent of all pull request resolutions. This indicates that humans remain responsible for the vast majority of the review workload. The complexity of modern codebases and the need for contextual understanding of business logic appear to exceed the current capabilities of automated review tools, keeping human engineers in the loop for critical decision-making.
Implications for Developers
For development teams, this study suggests that simply deploying AI coding agents without adjusting their review processes may yield diminishing returns. The "coding time versus review time" trade-off is currently neutral, meaning the time saved on writing code is offset by the time spent reviewing it. Teams should consider investing in better AI-assisted review tools or refining their prompt engineering to ensure AI-generated code is more likely to pass initial checks. Additionally, the finding that 95 percent of firms have implemented AI agents by early 2026 suggests that the learning curve is still active. Organizations that treat AI as a partner in the entire lifecycle, rather than just a code generator, are better positioned to unlock true efficiency gains.
What to Watch
- Improvements in AI Review Accuracy: As AI models improve, the percentage of review comments handled by AI may rise, potentially alleviating the human bottleneck.
- Changes in Code Complexity: If AI agents begin producing more complex or integrated features, the review burden may shift in different ways.
- Employment Trends: Longitudinal data may reveal subtle shifts in job roles, such as a greater emphasis on "AI orchestration" or "code auditing" positions.
- Tooling Evolution: The development of specialized AI tools designed specifically for reviewing AI-generated code could change the current dynamic.
Get the signal, not the noise.
One short email when it matters. No recaps of recaps.
SOURCES
SHARE
RELATED

[launches] ·
OpenAI launches Dots, a paid agent platform for business tasks

[models] ·
Anthropic releases Claude Opus 5.5 with 40% lower costs

[launches] ·
Anthropic merges Claude chat and Cowork into one interface

[models] ·
GPT-5.6 Luna vs GPT-6 Astra: Code Review Cost and Accuracy

[launches] ·
Cline v4.1.18 and CLI 3.0.62 ship Desktop app and major SDK updates

[tooling] ·
