Est.

Measuring and Reporting Agent Productivity Impact

Companies are spending billions on AI agents while barely measuring whether they actually work.

Staff Writer · · 8 min read
Cover illustration for “Measuring and Reporting Agent Productivity Impact”
Enterprise Agent Governance · September 24, 2026 · 8 min read · 1,847 words

Enterprises are spending faster than they're learning. That's the whole story in one line: adoption of AI agents has raced past the evidence for what these things actually do to output, and the gap between what companies believe and what gets measured is where the real money gets made or lost. Nearly a third of enterprise software is expected to carry agentic AI by 2028, and more than half of employees are already using agents today. The market backing all this starts at a substantial base in 2025 and is projected to grow more than sixfold by 2030. Yet only 23% of organizations say agents have delivered significant ROI, even though 80% report some kind of measurable return. Ninety-seven percent of executives say they deployed agents in the past year. Eighty-eight percent plan to increase AI budgets in the next twelve months anyway. Spending is outrunning proof, and that's exactly the problem this piece is here to unpack.

What the METR randomized controlled trial revealed about perceived versus actual agent productivity

Start with the punchline, because it's the kind of result that deserves to lead: developers using AI coding tools took 19% longer to finish real tasks, and they walked away convinced they'd been faster. Not a little convinced, either. They forecast a 24% speedup going in. That's a striking gap between what people believed and what a stopwatch actually recorded, the kind of divergence METR's research has documented systematically.

This wasn't a study of people who'd just discovered what an agent was. METR's 2025 trial used 16 experienced open-source developers working across 246 real tasks in mature codebases, people with roughly five years of experience on average. So the usual explanation, that this is a novice-user problem and things sort themselves out with practice, doesn't hold up here. These were seasoned engineers working in projects they knew, and the AI still slowed them down while making them feel sped up.

METR's report covering 349 technical workers follows that, and the picture gets stranger, not clearer. Self-reported "value" from AI use landed somewhere between 1.4x and 2x. Self-reported "speed," on the same population, came in around 3x. Those numbers should roughly agree if people are measuring the same thing. They don't, because they're not the same thing: speed is about how fast a task got done, value is about whether the output was worth having, and most survey instruments are built to elicit the second even when they ask about the first. Conflating them is how a company ends up celebrating a number that doesn't mean what everyone assumes it means.

What large-scale enterprise deployments show when productivity is tracked end-to-end

Zoomed out from the lab and into a real rollout, the story gets messier in a useful way. Microsoft tracked tens of thousands of engineers during an early-2026 rollout of Copilot CLI and Claude Code, and adopters merged roughly 24% more pull requests than they otherwise would have, a real number reflecting a real gain. That's a real number, and it's a real gain. But the researchers behind it were careful to state directly that a merged pull request is a proxy for output, not proof that the output was any good.

Retention tracked more closely with how active someone already was as a coder than with anything demographic. That matters more than it sounds like it should, because it means the people showing up first in the data are self-selected, already-engaged engineers. Roll the same tool out to everyone, including the engineers who rarely touch the codebase, and the average result will look worse than the early numbers suggest. Early adopter data is a flattering mirror, not a forecast.

Some internal engineering reports have attempted a more honest ledger, combining extended telemetry across large developer populations and many teams. Whether throughput metrics like epics completed, task volume, and PR merge rate actually move in the right direction is the question most internal dashboards stop short of answering honestly. The ones that do check are what separates a dashboard from an audit.

Self-reported hours-saved metrics are a useful signal but a dangerous primary measure

Asking around in Q1 2026 gets a suspiciously tidy answer: a specific hours-saved figure per week, give or take, across most of the major surveys. It's specific enough to sound like a fact carved in stone. It isn't; it just sounds that way.

METR's controlled data says people overestimate AI's effect on their own time by 40 percentage points, on average. So when survey after survey converges on a tidy hours-saved figure, the honest read is that a lot of people are equally confident and equally wrong, in the same direction, for the same reason." It's "a lot of people are equally confident and equally wrong, in the same direction, for the same reason." Confidence isn't evidence. Enthusiasm isn't a stopwatch.

None of that means self-report is worthless, and METR itself doesn't argue that. Surveys capture something real: how people feel about the tool, whether they'll keep using it, whether morale is up. That's useful information for a rollout plan. It's a terrible stand-in for a productivity number, and treating it as one is the actual failure, not the surveys existing.

The sector breakdown makes the limitation impossible to ignore. Productivity gains tend to appear most clearly in tasks with fast feedback loops and low stakes per error. They appear weakest wherever a human still has to review everything before it counts, because that review step eats most of the time the AI supposedly saved. One tool, wildly different results, depending entirely on how much oversight the workflow requires. A single hours-saved figure flattens all of that into a number that describes nothing in particular.

The frameworks organizations are using in 2026 to move from proxy metrics to output-linked measurement

Every serious framework published in 2026 lands on the same conclusion from a different direction: stop trusting one metric to carry the whole story. Leading evaluation frameworks make this point directly: no single number can tell you whether an agent is actually working.

The most ambitious attempt at fixing this is the Intelligence Impact Quotient, or IIQ, published on arXiv in May 2026 by Chandan Rajah, Neha Sengupta, Federico Castanedo, and colleagues at Inception/G42. IIQ builds a single 0 to 1000 index out of several ingredients at once: a novelty-weighted, time-decayed measure of accumulated "token stock," usage frequency, a grace-period gate so recent activity counts more than stale activity, a factor for organizational leverage, task complexity, and how much autonomy the agent is actually given. It's a mouthful, but the underlying logic is simple. Seat counts and total tokens processed don't tell you whether an employee is poking at a chatbot once a week or a workflow is quietly running a real process every day. IIQ is an attempt to separate the tourists from the residents.

The same instinct is visible in a four-tier enterprise KPI structure that's becoming close to standard. Resolution metrics come first: resolution rate, deflection rate, reopen rate, first-contact resolution. Quality metrics come next: hallucination rate, conversation quality scoring. Then operational metrics: automation rate, cost, escalation rate. Then business impact: the change in customer satisfaction, repeat contact rate, total ROI, time to value. Four tiers, not one headline number, because collapsing all of that into a single figure is exactly how you end up with a plausible-sounding six-to-seven-hours-saved claim that nobody can actually defend under questioning.

What business outcome evidence looks like and where the documented enterprise cases stand

Stripped of the surveys and the proxies, the honest inventory of documented, workflow-level business outcomes is thinner than the adoption numbers would suggest. Broad executive surveys have consistently found that only a minority report tangible value from generative AI at scale. McKinsey's State of AI finds real EBIT impact among the companies that have adopted it, but only a small fraction of them have actually redesigned a workflow around the technology rather than bolting it onto the old one. Other large-scale CEO research is blunter still, finding that only a minority of AI initiatives delivered the ROI they were expected to.

The thread running through all three is not which platform a company picked. It's whether the company was willing to tear apart and rebuild at least one high-volume workflow from the ground up, instead of dropping an agent into the existing process and hoping. That single decision, redesign versus bolt-on, predicts documented ROI better than anything else in the data.

Where the evidence is strongest, it's strongest because it's measured at the level of the actual workflow, not the tool. Some widely cited vendor case studies report high resolution rates with low handoff to human agents, a resolution number, not a speed number. Other enterprise deployments point to significant HR hours saved alongside cost avoidance, workflow outputs rather than self-reported impressions. The strongest internal deployments report meaningful cuts in operating costs alongside high containment, with cost and containment measured together rather than one flattering number standing in for the whole picture.

Among the clearest evidence in the broader body of research comes from customer support randomized controlled trials. Such studies have found meaningful average productivity gains measured against genuine control groups, at the task level, not self-reported. The detail that matters most in such studies: gains often concentrate among particular worker subgroups. The average number alone would have hidden that. The distribution is the finding.

Building and presenting an agent productivity report that holds up to executive and CFO scrutiny

Lead with the measurement design before a single result appears on the page. A CFO's first question is never "what's the number," it's "how did you get it," and a report that can't answer that cleanly is dead on arrival. If survey data appears anywhere in the report, disclose METR's 40 percentage point overestimation finding right next to it. Hiding that context isn't neutral, it's a choice, and it's the kind that gets found out later.

Keep four categories of evidence visibly separate rather than folding them into one productivity headline: self-reported perception, behavioral and telemetry data, controlled experimental results, and documented business outcomes. They're answering four different questions. A report that blends "employees feel 3x faster" with "PR merge rate rose 24%" with "resolution rate hit 84%" into a single number isn't simplifying, it's laundering four different claims into one that none of them individually support.

Structure the report itself around the four-tier framework: resolution and deflection first, quality metrics including hallucination and error rate second, operational metrics including actual cost third, business impact fourth. All four tiers, every time, not just the ones that make the quarter look good.

And report the costs sitting next to the gains, not buried three appendices later. If pull request volume is up, say whether the reviewed-versus-unreviewed ratio moved too. A dashboard that shows throughput climbing while quietly leaving out that review coverage dropped is a dashboard that's going to get contradicted by an incident report eventually, and that's a far worse conversation to have with a CFO than the one where the numbers were honest from the start.

More in Enterprise Agent Governance