Part 3 of the AI-Governed Enterprise Development Series. What the evidence shows, what it doesn’t, and why methodology matters more than tooling

Every published case of dramatic AI timeline compression involves modifications of existing systems, not greenfield development. The data that most enterprise leaders need to make planning decisions does not yet exist.
To separate proven outcomes from industry hype, it helps to look at the strength of the underlying evidence.
This analysis categorizes findings into four evidence tiers:
Throughout this discussion, findings are evaluated according to these tiers. Where evidence is limited, emerging, or inconclusive, that will be stated explicitly.
Understanding the quality of evidence is critical because AI adoption decisions are increasingly influencing budgets, timelines, governance strategies, and enterprise operating models.
The Perception Gap (Tier 1): The METR randomized controlled trial found that experienced developers using AI tools on their own repositories took 19% longer to complete tasks. Before starting, they predicted AI would make them 24% faster. After finishing, they still estimated they had been 20% faster. In other words, their perception was directionally wrong by roughly 40 percentage points.
The Effort Shift (Tier 3): Anthropic's self reported engineering survey found that AI is involved in 59% of daily work, with engineers reporting productivity gains of around 50%. However, more than half could fully delegate only 0 to 20% of their tasks to AI. Human interactions per session decreased by 33%, while AI tool usage increased by 116%. The effort is not disappearing. It is shifting from writing code to reviewing, validating, and orchestrating AI generated outputs.
The Organizational Paradox (Tier 2): Data from Faros AI, covering more than 10,000 developers across 1,255 teams, showed tasks completed increased by 21% and pull request volume rose by 98%. However, review times increased by 91%, bugs per developer rose by 9%, and any correlation with company level delivery performance disappeared. Individual productivity gains did not translate into organizational delivery improvements.
The 9% increase in bugs reported by Faros AI may be an early indicator of growing quality debt. Industry analyses suggest rework time has increased materially in teams without structured review processes. For systems expected to be maintained for five to ten years, the long term cost of rework can outweigh the initial gains achieved through faster code generation.
Amazon Java Migration (Tier 3): Amazon's large scale Java migration is one of the strongest examples of AI driven timeline compression. Across tens of thousands of applications, the initiative reportedly saved 4,500 developer years and delivered approximately $260 million in annualized value. Individual application upgrades that previously required around 50 developer days were completed in a matter of hours. However, this was a pattern based transformation with a fully defined target state, making it an ideal candidate for AI acceleration.
Fujitsu Medical Software (Tier 3): Fujitsu reported completing one software change request in approximately four hours, compared to a vendor estimate of three person months. While impressive, this result came from a highly constrained maintenance environment with known requirements and existing system context. It demonstrates significant compression for a specific task rather than an entire software project.
McKinsey G SIB Bank (Tier 3): One of the first publicly documented greenfield examples comes from a Global Systemically Important Bank highlighted by McKinsey. Using an AI agent factory with roughly 100 agent instances and three human engineers, the organization reported 10x faster delivery, 50% lower costs, and 40 to 70% productivity gains. However, important details such as codebase size, team structure, methodology, and independent verification were not disclosed. The case is an encouraging signal, but not yet a benchmark for enterprise planning.
A common thread across most published success stories is that AI operated within known constraints. Existing codebases, established architectures, and clearly defined transformation goals reduced ambiguity and guided decision making.
Greenfield development is fundamentally different. The challenge shifts from implementing known solutions to defining requirements, designing architecture, stabilizing interfaces, and governing uncertainty. These are precisely the areas where AI currently delivers the least compression, making it difficult to apply modification based results directly to greenfield projects.
The industry is arriving at a similar conclusion from multiple directions. Gartner's work on Multiagent Systems, IBM's Agent Stack, Singapore IMDA's governance framework, EY's concept of an "agentic AI OS," and Deloitte's research all point to the same underlying principle: AI can improve individual productivity, but governance is what enables organizational outcomes.
The challenge is that adoption is moving faster than coordination. According to Belitsoft, enterprises now run an average of 12 AI agents, yet half operate in isolation and only 11% of intended use cases have reached production. This mirrors what Faros AI observed at the developer level. Without governance, local productivity gains rarely translate into enterprise-wide delivery improvements.
The challenge is not interpreting the data we have. It is recognizing the data we do not. Some of the most important questions around AI driven software delivery remain largely unanswered, especially in complex enterprise settings.
For CEOs, CFOs, CIOs/CTOs, CSOs, and General Counsel:
For technology leaders:
This brief is part of the AI-Governed Enterprise Development Series by Technossus. Full white papers available upon request.
This document was developed with the assistance of AI tools for drafting and editing.