Do AI coding agents improve software development productivity?
AI coding agents seem to help productivity, but the size of the gain varies a lot and studies measure it differently.
Covers: What controlled studies, field experiments and industry reports find about the effect of AI coding agents on developer productivity, code quality and review time, including evidence that gains in code generation can be absorbed by human review. Does not cover general-purpose chatbots or non-coding AI tools.
Also answers: How do AI coding agents affect software development productivity? · Do AI coding agents make developers more productive? · How much do AI coding assistants speed up software development? · Are AI coding tools worth it for developers?
- One page for this question6 other ways of asking lead here
- 6 independent sourcesEvery claim links to what supports it
- Joins the mapLinked as related pages appear
- Clean discussionScreened before anything appears
Before you read, make a guess
Fill in the blank: ?% of the task time saved by developers using GitHub Copilot
Drag the slider to fill in the blank
The short answer
Evidence-backed AI-prepared starting mapEvidence on AI coding agents and productivity points in a positive direction but with widely varying magnitudes and contested measurement. In a controlled experiment, developers with GitHub Copilot completed an HTTP server task 55.8% faster than controls (Peng et al., 2023). In field experiments with 1,974 developers at Microsoft and Accenture, those given Copilot completed 12.92%–21.83% more pull requests per week at Microsoft and 7.51%–8.69% at Accenture, though the authors call these estimates imprecise and only statistically significant under specifications that weight periods of differing Copilot uptake (Cui et al., 2024). An IBM internal deployment of watsonx Code Assistant found net productivity increases were common but not experienced by all users (Weisz et al., 2024). A systematic review argues the literature lacks comparability because productivity is operationalized through heterogeneous markers, settings and reporting styles (Wittig et al., 2026). Reporting on a study describes coding efficiency gains being "absorbed" by a human review "bottleneck" (Ars Technica, 2026).12345
- Evidence 17
Did this answer your question?
Be the first to voteIn brief
Gains are not universal: an IBM deployment found net increases were common but not experienced by all users.3
Evidence-backedProductivity is measured inconsistently across studies, which is a major reason results are hard to compare.4
Evidence-backedMore generated code does not automatically mean more shipped software: review can absorb the gains.5
Evidence-backedLong-term factors such as technical expertise and ownership of work are increasingly proposed as part of how productivity should be judged.6
Evidence-backed
At a glance
The picture in numbers
Live · updated just now
55.8%
56 in every 100
- Microsoft12.92–21.83%
- Accenture7.51–8.69%
2,989 developer responses
The evidence behind it
6 sources- Reviews of many studies1
- Other studies and data4
- Background1
When it was published
Newest from 2026
| Source | Kind | Year |
|---|---|---|
| AI coding agents generate more code, but not more software | Background | 2026 |
| Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants | Other studies and data | 2026 |
| Fragmented Markers, Mixed Results: A Systematic Review of AI Coding Assistants and Developer Productivity | Reviews of many studies | 2026 |
| Examining the Use and Impact of an AI Code Assistant on Developer Productivity and Experience in the Enterprise | Other studies and data | 2024 |
| The Impact of AI on Developer Productivity: Evidence from GitHub Copilot | Other studies and data | 2023 |
| The Productivity Effects of Generative AI: Evidence from a Field Experiment with GitHub Copilot | Other studies and data | 2024 |
The community around it
No one has added to this page yet. Firsthand experience, a newer study or a different reading of the numbers would show up here, credited to you.
What it means for you
Which fits you?
Pick the situation closest to yours. Each answer says what it rests on.
If you are evaluating a vendor's headline productivity claim
check whether it comes from a controlled task or a field setting, since the reported magnitudes differ sharply (55.8% on one controlled task versus 7.51%–21.83% more pull requests per week in field experiments).12
Evidence-backedIf your team's review capacity is already a constraint
expect code-generation speedups to be partly or fully absorbed by review rather than showing up as more shipped software.5
Evidence-backedIf you are choosing productivity metrics for an AI rollout
plan for multiple measures, because survey and interview evidence shows usefulness judgments conflict and productivity spans short-term and long-term dimensions including expertise and ownership.6
Evidence-backedIf you are rolling out an assistant across a large organization
do not assume uniform benefit, since an internal deployment found net increases were common but not experienced by all users.3
Evidence-backedIf you are comparing studies to decide what to expect
treat cross-study comparisons cautiously, because productivity is operationalized through heterogeneous markers, settings and reporting styles.4
Evidence-backedThe full story · 3 chapters
01
What controlled studies and field experiments find
AI summary:Controlled and field experiments show real but variable gains, from 55.8% faster on one task to smaller, imprecise pull-request increases.
Evidence-backed: The most-cited controlled result comes from a randomized experiment in which developers were asked to implement an HTTP server in JavaScript as quickly as possible. The group with access to GitHub Copilot finished 55.8% faster than the control group, and the authors report heterogeneous effects that they suggest may help people transition into software development careers (Peng et al., 2023).1
Evidence-backed: Field experiments at Microsoft and Accenture with 1,974 developers found smaller throughput gains: 12.92%–21.83% more pull requests per week at Microsoft and 7.51%–8.69% at Accenture, depending on specification. The authors describe these as suggestive and not very precise, noting low compliance in the Microsoft experiment and internal organizational changes at Accenture; the estimates reach statistical significance only when periods with differing Copilot uptake across control and treatment are weighted more heavily (Cui et al., 2024).2
Evidence-backed: An internal deployment of watsonx Code Assistant at IBM, studied through surveys of two user cohorts (N=669) and unmoderated usability testing (N=15), found that such tools often provide net productivity increases but that these benefits may not be experienced by all users. The study also surfaced new questions about ownership of and responsibility for generated code (Weisz et al., 2024).3
02
Why the numbers disagree: productivity is measured inconsistently
AI summary:Studies measure productivity in inconsistent ways, making results hard to compare, and some argue for long-term factors like expertise and ownership.
Evidence-backed: A staged systematic literature review argues the central problem is not a lack of studies but a lack of comparability. Prior work operationalizes productivity through heterogeneous markers, settings and reporting styles, which makes cross-study conclusions difficult. The review first synthesizes studies that explicitly examine developer productivity, then extends to productivity-related outcomes using SPACE as a theory-informed lens to capture human-centered evidence that explicit productivity studies often omit, and analyzes how productivity is operationalized, which confounders are discussed, and how reported effects are distributed (Wittig et al., 2026).4
Evidence-backed: A mixed-method study at BNY Mellon, combining a survey with 2,989 developer responses and 11 in-depth interviews, found that a multifaceted approach is needed: survey results exposed conflicting perspectives on AI tool usefulness, while interviews elicited six distinct factors spanning short-term and long-term dimensions of productivity. Unlike earlier work, these factors emphasize long-term metrics such as technical expertise and ownership of work (Chen et al., 2026).6
03
Code generation gains and the human review bottleneck
AI summary:Reporting describes coding gains being absorbed by a human review bottleneck, so more generated code does not mean more shipped software.
Evidence-backed: Reporting on a study describes a pattern in which coding efficiency gains get "absorbed" by a human review "bottleneck," so that AI coding agents generate more code but not more software (Ars Technica, 2026). This is consistent with the broader measurement critique: if output is counted in generated code rather than shipped, reviewed and maintained software, apparent gains can overstate delivered productivity.5
Your turn
Have your say
Quick votes, open to everyone. See where you stand the moment you vote. Only totals are ever shown.
How do you feel about this?
No votes yetQuick questions from connected pages
Before you go
What to remember
Try to recall each hidden figure before you reveal it. Remembering, not rereading, is what makes it stick.
Controlled and field evidence points to real but variable productivity gains: faster on one controlled task, and 7.51%–21.83% more pull requests per week in field experiments that the authors call imprecise.
Gains are not universal: an IBM deployment found net increases were common but not experienced by all users.
Productivity is measured inconsistently across studies, which is a major reason results are hard to compare.
Your reading
0 of 3 chaptersThis answer keeps changing
When new evidence or a better source comes in, this page is updated (it's on version 2, last changed 53 minutes ago). Follow it to be told when that happens.
Ask this Sylo
Still wondering about something?
Answers come only from this page's reviewed material, with citations, and say plainly when the page doesn't cover it yet.
More on AI
Everything on AI ›How much energy does ChatGPT use?
How much energy does ChatGPT use per query, and how does that compare with other everyday activities?
Does AI help students learn or make them lazier?
Does using AI tools help students learn more effectively, or does it make them lazier and undermine their learning?
Is using ChatGPT for homework cheating?
How accurate is ChatGPT?
How accurate is ChatGPT, and what does research show about its error rates and reliability?
Why do AI chatbots make things up?
Is AI dangerous?
Is artificial intelligence dangerous, and what does the evidence say about its risks?
Behind this page
Who's adding to it, where it comes from, how it changed and what would make it better. Always open to everyone.
Discussion
Sources
Numbers match the citations in the article. A working link isn't proof that a page supports a claim; check the quoted passage and date.
- 1The Impact of AI on Developer Productivity: Evidence from GitHub CopilotarXiv (Cornell University) (Peng et al.)Published Feb 13, 2023Checked Oct 10, 2026
“Generative AI tools hold promise to increase human productivity. This paper presents results from a controlled experiment with GitHub Copilot, an AI pair programmer. Recruited software developers were asked to implement an HTTP server in JavaScript as quickly as possible. The treatment group, with access to the AI pair programmer, completed the task 55.8% faster than the control group. Observed heterogenous effects show promise for AI pair programmers to help people transition into software development careers.”
- 2The Productivity Effects of Generative AI: Evidence from a Field Experiment with GitHub CopilotResearch paper (Cui et al.)Published Mar 27, 2024Checked Oct 10, 2026
“We are providing a preview of a project that analyzes two field experiments with 1,974 software developers at Microsoft and Accenture to evaluate the productivity impact of Generative AI. As part of our study, a random subset of developers was given access to GitHub Copilot, an AI-based coding assistant that intelligently suggests ‘completions’ for code. Our preliminary results provide suggestive evidence that these developers became more productive, completing 12.92% to 21.83% more pull requests per week at Microsoft and 7.51% to 8.69% at Accenture (depending on specification). Due to low compliance in the Microsoft experiment and internal organizational changes at Accenture, our estimates are not very precise and only reach statistical significance if we weight more heavily periods in which Copilot uptake differs across control and treatment.”
- 3Examining the Use and Impact of an AI Code Assistant on Developer Productivity and Experience in the EnterprisearXiv (Cornell University) (Weisz et al.)Published Dec 9, 2024Checked Oct 10, 2026
“AI assistants are being created to help software engineers conduct a variety of coding-related tasks, such as writing, documenting, and testing code. We describe the use of the watsonx Code Assistant (WCA), an LLM-powered coding assistant deployed internally within IBM. Through surveys of two user cohorts (N=669) and unmoderated usability testing (N=15), we examined developers' experiences with WCA and its impact on their productivity. We learned about their motivations for using (or not using) WCA, we examined their expectations of its speed and quality, and we identified new considerations regarding ownership of and responsibility for generated code. Our case study characterizes the impact of an LLM-powered assistant on developers' perceptions of productivity and it shows that although such tools do often provide net productivity increases, these benefits may not always be experienced by all users.”
- 4Fragmented Markers, Mixed Results: A Systematic Review of AI Coding Assistants and Developer ProductivityIEEE/ACM International Conference on Automated Software Engineering (ASE) (Wittig et al.)Published Oct 7, 2026Checked Oct 10, 2026
“AI coding assistants are increasingly embedded in software engineering practice, yet the literature still offers no cumulative, easily interpretable picture of their effects on software developer productivity. The central problem is not a lack of studies, but a lack of comparability: prior work operationalizes productivity through heterogeneous markers, settings, and reporting styles, making cross-study conclusions difficult. To address this, we conduct a staged systematic literature review. In the first stage, we synthesize studies that explicitly examine developer productivity. In the second, we extend the scope to productivity-related developer outcomes using SPACE as a theory-informed lens, capturing human-centered evidence that explicit productivity studies often omit. Across both stages, we analyze how productivity is operationalized, which confounding factors are discussed, and how reported effects are distributed.”
- 5AI coding agents generate more code, but not more softwareArs TechnicaPublished Oct 9, 2026Checked Oct 10, 2026
“Study finds coding efficiency gains get "absorbed" by human review "bottleneck."”
- 6Beyond the Commit: Developer Perspectives on Productivity with AI Coding AssistantsarXiv (Cornell University) (Chen et al.)Published Feb 3, 2026Checked Oct 10, 2026
“In the age of AI coding assistants, it has become even more important for both academia and industry to understand how to measure their impact on developer productivity, and to reconsider whether earlier measures and frameworks still apply. This study analyzes the validity of different approaches to evaluating the productivity impacts of AI coding assistants by leveraging mixed-method research. At BNY Mellon, we conduct a survey with 2989 developer responses and 11 in-depth interviews. Our findings demonstrate that a multifaceted approach is needed to measure AI productivity impacts: survey results expose conflicting perspectives on AI tool usefulness, while interviews elicit six distinct factors that capture both short-term and long-term dimensions of productivity. In contrast to prior work, our factors highlight the importance of long-term metrics like technical expertise and ownership of work. We hope this work encourages future research to incorporate a broader range of human-centered factors, and supports industry in adopting more holistic approaches to evaluating developer productivity.”
How it changed
Published 1 time since Oct 10, 2026.
- Version 2Oct 10, 2026Live now
AI-prepared Starting Map from live research.
- First published version.
Help improve it
The brief is open about what's uncertain. These are the specific gaps that new material would fill.
“Code generation gains and the human review bottleneck” rests on one independent source
A second, independent source that confirms or challenges it would make this part more reliable.
Open questions
How large is the review bottleneck effect in real teams, and under what conditions does it fully absorb code-generation gains versus only partly offset them?
No answers yet
Do short-term throughput gains translate into long-term gains in technical expertise and ownership of work, or do they trade off against them?
No answers yet
Which developers experience net productivity increases and which do not, and what distinguishes them?
No answers yet
What set of markers would make productivity results comparable across studies and settings?
No answers yet
Around this topic
Sylos connect: narrower topics report up to broader ones, so what's learned in one place shows up where it matters.