AI has made software engineering faster in ways that are easy to see.
Requirements can be drafted in hours instead of days. Test cases can be generated almost instantly. Developers can produce and refactor code at a pace that would have been difficult to imagine a few years ago.
That makes AI engineering productivity look straightforward: give teams better tools, measure how much faster individual tasks become, and calculate the gain.
In practice, that equation breaks down quickly.
At Teravision, our own AI transformation started with an internal experiment.
We created three independent software delivery teams and challenged them to rebuild in eight weeks a product that had previously taken seven months.
Each team could use the AI tools it considered useful across product, development, QA, and infrastructure. Every stage found ways to accelerate.
But those gains did not compound across the software development life cycle. One team could produce requirements dramatically faster, while the next stage was unable to absorb the additional work. QA could generate tests faster, while development became the new bottleneck.
Individual productivity had increased. End-to-end productivity had not.
That distinction matters because engineering leaders are now under pressure to answer a much harder question than whether developers are using AI:
Is the organization actually delivering more value because of it?
The AI productivity trap: measuring speed where it is easiest to see
The easiest AI metrics to collect tend to be the least useful ones.
How many engineers use Copilot? How much code was generated by AI? How many prompts were sent? How quickly did a developer complete a specific task?
These signals can help understand adoption. They do not tell you whether an engineering organization is delivering better software faster.
This is increasingly important because perception and measurable productivity do not always move together.
In a 2025 randomized study, METR found that a group of experienced open-source developers working on repositories they knew well took 19% longer to complete assigned tasks when AI tools were allowed, even though the developers believed AI had made them faster.
The study is narrow and should not be generalized to every engineering environment, but the gap between perceived and measured improvement is instructive.
DORA's research reaches a related conclusion from a broader organizational perspective: successful AI adoption is a systems problem, not simply a tooling problem.
Local productivity improvements create business value when the surrounding delivery system can absorb and compound them.
We saw this directly with client teams
In one case, a QA engineer reduced the creation of a test matrix from roughly two weeks to two days. That looked like an extraordinary productivity gain.
Then tickets began piling up for developers who were not prepared to process feedback at the same speed.
The measurement was correct. QA really had accelerated. The conclusion was wrong. The system had not.
This is why an AI engineering productivity strategy has to look beyond how fast one person or one stage operates.
When AI changes the speed of one part of the SDLC, the constraint frequently moves somewhere else.
The goal is to find out what actually ships
From AI activity to engineering outcomes
One useful way to think about measurement is to separate four levels: activities, outputs, outcomes, and impacts.
- Activities tell you what teams are doing: prompting models, writing code, reviewing pull requests, testing, documenting.
- Outputs are what those activities produce: releases, features, test cases, pull requests, resolved defects.
- Outcomes describe what changed for the engineering organization: shorter lead times, better maintainability, fewer defects, faster releases or greater ability to respond to new requirements.
- Impacts connect those improvements to the business: revenue, customer value, profitability, market responsiveness or another result leadership actually cares about.
The mistake is stopping at the first two levels.
For engineering teams, that means AI productivity measurement should combine several dimensions rather than search for a single universal metric:
- Delivery: Lead time, cycle time, release frequency and overall team velocity
- Quality: Defect trends, bug count, test coverage and rework
- Sustainability: Maintainability, scalability and technical debt
- Business value: Whether faster engineering translates into faster learning, customer outcomes or measurable organizational impact
At Teravision, one AI transformation baseline included velocity, code quality, bug count, unit test coverage, maintainability, scalability, clarity and documentation, alongside end-to-end measures such as lead time and cycle time.
The important part was not the number of metrics. It was the baseline.
Without knowing how the team performed before changing its way of working, leaders have no reliable way to separate genuine improvement from enthusiasm around a new tool.
The same principle appears in Ricardo Arcia's recent conversations about the transformation.
Teravision historically used measures such as story points to understand how much work a team could deliver in a two-week sprint.
AI has made it necessary to broaden that view. If product, design, coding and testing all accelerate independently, the relevant result is still whether the same team can increase what it delivers end to end.
In one example Ricardo discusses, a team that previously delivered approximately 22 story points per sprint increased that capacity toward 30.
That tells us more about system productivity than measuring how many lines of code an AI assistant generated.
Productivity improves when the system can absorb the speed
Measurement also exposes another reality: AI does not enter a neutral engineering system.
It amplifies what is already there.
If code review is poorly structured, AI can create more code for an already weak review process to handle.
If requirements lack clarity, agents can produce more output based on insufficient context.
If teams operate with different standards, faster execution can increase inconsistency rather than remove it.
This is why engineering productivity and engineering maturity are becoming increasingly difficult to separate.
In another Teravision implementation, teams initially progressed at very different speeds.
Giving multiple teams the same mandate to experiment with AI did not create standardized improvement.
Some advanced rapidly while others struggled because each group was discovering its own practices while also trying to meet delivery commitments.
The response was to centralize the initial learning through Team Zero: a real delivery team given explicit permission to experiment, adapt the Cognitive Engineering Framework™ to the organization's reality, establish baselines, and turn what works into a Transfer Package that subsequent teams can follow.
The objective is not to make every team experiment from scratch. It is to make learning reusable.
This also explains why giving teams access to AI tools is a weak proxy for transformation.
In one client example discussed by Ricardo, engineers were given Copilot without a common framework or structured training.
A small group found meaningful gains, many experimented and stopped after failing to get useful results, and another group barely adopted the tool at all.
The technology was available to everyone. The capability was not.
The role of the engineer is changing with it. Across Teravision's experiments, one of the recurring shifts has been from executor toward orchestrator and validator: people provide context, make architectural and product decisions, evaluate probabilistic outputs, and remain accountable for what goes into production.
As Ricardo described it in another recent conversation, engineers increasingly need to “think and orchestrate and be the strategist.”
Build a measurement loop, not a productivity dashboard
A dashboard can tell you what happened. Transformation requires a mechanism for deciding what to do next.
At Teravision, we structure that learning through 90-Day Loops: establish a baseline, introduce a deliberate change, execute it in real delivery conditions, measure the result, and evolve the approach for the next cycle.
Ninety days gives teams enough time to move beyond the initial productivity dip that comes with changing established ways of working, while remaining short enough to make assumptions testable.
This is especially important because AI engineering practices are changing faster than most organizations can standardize them.
The framework Teravision uses has already gone through more than twenty iterations as models, tools and real-world implementation experience have evolved.
The point is not to find the perfect AI productivity metric.
It is to create enough visibility to answer a series of better questions: Are we shipping faster? Are we maintaining quality? Did the bottleneck move? Is the team spending more capacity on new value or on correcting AI-generated work?
Are improvements surviving beyond individual power users? And can engineering leadership connect those gains to something the business recognizes as valuable?
AI engineering productivity becomes meaningful when those questions can be answered with evidence.
Because faster code is useful. Faster delivery is better. And measurable business value is what makes AI transformation worth scaling.
For engineering organizations trying to move from isolated AI experiments to measurable end-to-end improvement, Teravision's AI Transformation Accelerator Program is designed to build that foundation: from baselines and training to a structured adoption framework, Team Zero and repeatable measurement cycles.
