Real productivity gains — smaller and slower to arrive than the pitch decks suggest
GitHub Copilot (multi-org studies) · 2023–2025
What they did
As AI coding assistants moved from autocomplete to agentic "fix this, then run the tests" workflows, independent researchers ran longitudinal studies rather than taking vendor benchmarks at face value — tracking real engineering teams before and after adoption.
What happened
Results diverge by study and metric. ZoomInfo (Jan 2025) reported a 33% suggestion-acceptance rate and a 40–50% productivity boost on developer-reported measures, growing with task difficulty. A separate longitudinal study at NAV IT, tracking 100→250 Copilot users from 2023–2025, found no statistically significant change in commit-based activity metrics, even though developers subjectively felt more productive — a real gap between perceived and measured gain. Microsoft's own research found it takes developers roughly 11 weeks to realize the tool's full productivity value, but most judge it in the first week, at only ~20% of that value.
The so-what
"AI coding assistants boost productivity" is true and also not a single number — it depends which metric you trust (commits vs. self-report), how long the team has used it, and how hard the tasks are. Treat any single-study claim as a data point, not the answer.