Stop Measuring AI Adoption, Start Measuring AI Outcomes

There is a slide appearing in board packs across large enterprises at the moment. It shows how many employees have an AI licence, weekly active users, perhaps token consumption and a reassuring trend line heading upwards.

It looks like evidence of progress.

The problem is that it tells you very little about whether AI is actually improving the business.

McKinsey’s 2026 State of AI survey makes the disconnect particularly clear. Eighty per cent of respondents say AI has improved their individual productivity. Yet only 37 per cent say it has contributed anything to their organisation’s EBIT, broadly unchanged from the previous year. More strikingly, just six per cent qualify as McKinsey’s AI “high performers”, organisations attributing at least five per cent of EBIT to AI and describing its impact as significant.

Meanwhile, adoption itself is becoming almost universal. Around 88 per cent of organisations now use AI in at least one business function.[1]

We have therefore become very good at achieving the thing we chose to measure. The evidence that this is translating into business performance is much less convincing.

That distinction matters.

When the measure becomes the target

Goodhart’s law is probably overquoted, but AI adoption is becoming a very good demonstration of it: when a measure becomes a target, it stops being a useful measure.

Adoption made sense as an early indicator. Then organisations started putting targets against it.

Amazon is one example. The company reportedly pushed for high levels of weekly AI use among developers and used internal leaderboards to monitor adoption. Employees described finding ways to increase their usage figures, while Amazon subsequently moved away from the leaderboard approach.[2]

Meta has wrestled with a similar problem. Its internal emphasis on heavy AI usage, sometimes described as “tokenmaxxing”, has subsequently been toned down in favour of assessing the impact employees produce rather than token consumption itself.[3]

There have been similar tensions elsewhere. Organisations understandably want employees to experiment with expensive new tools, particularly after making substantial investments in licences and infrastructure. But once usage becomes something employees believe they are being judged against, the behaviour changes.

That isn’t particularly surprising.

People respond to the incentives they are given. If an organisation tells employees that AI consumption demonstrates maturity, employees will manufacture consumption.

That doesn’t mean they are manufacturing value.

I wrote about the economics behind this in The Great Token Retrenchment. The broader point is the same: consumption is an input. Treating it as an outcome encourages organisations to optimise the wrong side of the equation.

Every programme built around adoption or token consumption risks optimising the numerator of a ratio whose denominator has never really been defined.

Worse, once adoption becomes a target, the number itself becomes less trustworthy. Some proportion of the activity is now happening because the activity is being measured.

The dashboard still goes up. Its meaning goes down.

Adoption wasn’t a bad metric

There is an important distinction here. Measuring adoption was not a mistake.

During the first phase of generative AI, diffusion really was the constraint. Organisations had bought technology that many employees didn’t understand, didn’t trust or simply didn’t use.

In 2023 and 2024, knowing whether people had crossed that threshold was useful. If nobody touches the technology, nothing downstream of it can happen.

Adoption also has several practical advantages. It is cheap to measure, available almost immediately, comparable across teams and easy to explain to a board. Outcome measures are generally none of those things.

The mistake was allowing adoption to remain the headline measure after adoption stopped being the problem.

It is a prerequisite for value, not evidence of value.

Attendance is necessary for education too. Nobody would seriously suggest measuring the quality of a school solely by counting how many pupils walked through the door.

At 88 per cent organisational adoption, saying “we use AI” is increasingly meaningless as a strategic position.[1]

Customers aren’t going to pay you more because your employees have Copilot licences.

Where did the productivity go?

There is a tempting explanation for the gap between 80 per cent reporting individual productivity improvements and 37 per cent reporting EBIT impact: perhaps the value exists but businesses simply can’t measure it.

There is certainly some truth in that.

Enterprise attribution is difficult. Benefits can be distributed across hundreds of small tasks, while financial measures such as EBIT are affected by dozens of things that have nothing to do with AI. A survey asking executives whether AI has affected EBIT is hardly a perfect measurement instrument either.

But measurement can’t explain the whole gap.

Some of the productivity gain appears to disappear elsewhere in the process.

Research from BetterUp Labs and Stanford’s Social Media Lab described the phenomenon as “workslop”: AI-generated work that looks complete but lacks enough substance to move the task forward.

Around 40 per cent of US desk workers surveyed said they had received it. Resolving each instance took around two hours on average. BetterUp modelled the resulting productivity cost at more than $9 million annually for an organisation of 10,000 workers.[4]

Workday’s research points in the same direction. It found that nearly 40 per cent of the time employees said AI had saved them was subsequently lost reviewing, correcting or rewriting AI-generated output.[5]

That exposes a weakness in individual productivity measures.

Someone can genuinely become faster while the overall process becomes slower.

If I produce something in thirty minutes rather than two hours, my productivity has improved. But if a colleague then spends ninety minutes checking, correcting and reconstructing it, the organisation may have gained very little.

The cost has simply moved.

And that matters because many of the productivity claims being made about AI are measured at the point where the work is produced, rather than at the point where the work is ultimately accepted, used or converted into an outcome.

An adoption dashboard can’t see that. In fact, the person generating the additional downstream work probably looks excellent on it.

The BetterUp research points to another cost too. Forty-two per cent of recipients of workslop said it made them regard the sender as less trustworthy.[4]

There isn’t an AI adoption dashboard on which that looks like anything other than healthy engagement.

Measure the workflow instead

The obvious response is to say organisations should “measure outcomes”.

That’s correct, but not particularly helpful. If it were easy, everybody would already be doing it.

The reason organisations measure adoption isn’t necessarily that they don’t understand the difference. Adoption is readily available. Outcomes are expensive to measure.

Doing this properly requires a different approach.

Start with the workflow

The most important change is moving the unit of measurement away from the employee and towards the workflow.

Per-user metrics, whether usage counts, licence utilisation or spending caps, are administratively convenient but rarely tell you whether a business process has improved.

Instead ask what happened to the process.

  • Did cycle time fall?
  • Did throughput increase?
  • Did the error rate change?
  • Were there fewer hand-offs or escalations?
  • Did conversion or retention improve?
  • Was risk reduced?
  • And, importantly, what happened to the capacity supposedly released?

That last question is often missing.

Saving somebody five hours a week does not automatically create five hours of economic value. If the capacity is simply absorbed into the working day, the individual may have had a better experience, which has value in itself, but it shouldn’t automatically be booked as a financial return.

It is also useful to keep different types of value separate.

Productivity might be measured through cycle time, throughput or capacity.

Quality might mean defects, rework, escalation rates or customer satisfaction.

Revenue might mean conversion, retention, average order value or the ability to serve markets that were previously uneconomic.

Risk might mean incidents, compliance exceptions or audit findings.

These outcomes behave differently and materialise over different timescales. Compressing all of them into a single AI ROI percentage can conceal more than it reveals.

Establish the baseline first

This sounds obvious, yet it is frequently missed.

If you don’t know how the process performed before introducing AI, demonstrating improvement afterwards becomes extremely difficult.

A baseline constructed retrospectively is vulnerable to selection and confirmation bias and will struggle under serious financial scrutiny.

If a workflow is going to be changed, measure it before the technology touches it.[6]

That means understanding the current cycle time, cost, error rate, human effort and whatever business outcome the intervention is supposed to affect.

Otherwise the inevitable conversation six months later becomes: “It feels faster.”

That may be true. It isn’t a business case.

Ask what would have happened anyway

A metric improving after an AI deployment doesn’t mean AI caused the improvement.

Most large organisations are changing multiple things simultaneously. Teams are restructured. Suppliers change. Systems are replaced. Prices move. Processes are redesigned. People become more experienced.

If conversion rises by five per cent after an AI deployment, how much of that five per cent belongs to AI?

Where practical, use control or holdout groups.

Where that isn’t possible, at least document the attribution assumption and acknowledge the uncertainty around it.

This doesn’t require turning every enterprise AI deployment into an academic experiment. It requires being explicit about what you know and what you don’t.

I’d rather see a credible range than a suspiciously precise ROI number.

Include the actual cost

This is where the economics become more useful.

Instead of concentrating on tokens, licences or API calls, measure the cost per successful outcome.

Imagine one model resolves 90 per cent of customer requests correctly on the first attempt. Another costs half as much to run but resolves only 40 per cent, with the remainder requiring human intervention.

Measured by token cost, the second model looks efficient.

Measured by cost per successfully resolved case, it may be considerably more expensive.

The cheap AI model hasn’t saved money. It has moved the cost from the technology budget into operations.

The same principle applies beyond customer service.

Cost per accepted software change is more useful than tokens per developer.

Cost per qualified sales opportunity is more useful than prompts per salesperson.

Cost per correctly processed claim is more useful than AI interactions per claims handler.

Once the denominator is a successful business outcome, arguments about models, tokens and licences start to look quite different.

Why aren’t organisations doing this already?

This is where the problem becomes more interesting.

Larridin’s State of Enterprise AI 2026 asked senior leaders what prevented their organisations from measuring AI effectively.

The leading problems weren’t technical.

Some 30.5 per cent cited unclear responsibility for measurement and another 27.7 per cent pointed to fragmented ownership across teams. Almost a quarter, 24.4 per cent, reported no correlation between usage and outcomes.

Inadequate data infrastructure came fourth, at 15 per cent.[7]

That suggests this isn’t primarily an analytics problem.

It is an operating-model problem.

A workflow might span three business functions. Technology belongs to another team. The budget sits somewhere else. Finance measures the eventual outcome. Nobody owns the complete chain.

And the only party with an obvious interest in proving a positive result may be the team that sponsored the investment in the first place.

Licence utilisation, by comparison, has a system of record, an owner and a number that can be produced every Monday morning.

So that’s what gets measured.

The McKinsey research provides an interesting counterpoint.

Its six per cent of high performers are much more likely to have fundamentally redesigned workflows around AI. They are also substantially more likely to have defined processes for measuring the impact of their AI initiatives.[1]

I don’t think those findings are unrelated.

It is difficult to redesign a workflow properly without understanding how it currently performs. Equally, there is little incentive to invest in measuring a workflow if nobody has the authority to change it.

Measurement and workflow ownership need to meet somewhere.

That is why this is ultimately as much an organisational question as a technology one.

What I’d change

First, I would stop presenting adoption as a headline board metric.

That doesn’t necessarily mean removing it. Adoption can still tell you useful things, particularly during rollout. But if it appears in the board pack, put the business outcome it is intended to influence next to it.

If AI usage rises 30 per cent while cycle time, quality, revenue and cost remain unchanged, that divergence should be visible immediately, not discovered eighteen months later when somebody asks where the return went.

I’d also give adoption metrics a sunset date.

Once usage has reached a reasonable level, their strategic usefulness falls rapidly.

Second, don’t respond to excessive AI consumption by simply moving to the opposite extreme.

A blanket restriction on tokens or model usage isn’t evidence of financial discipline any more than encouraging maximum consumption was evidence of AI maturity.

Both approaches are managing inputs.

There will almost certainly be a relatively small group of people producing disproportionate value from these tools. A crude cap risks constraining them while doing nothing about low-value usage elsewhere.

The objective isn’t maximum consumption or minimum consumption.

It is efficient conversion of consumption into outcomes.

Finally, don’t try to measure everything.

Pick a handful of important workflows and instrument them properly.

You don’t need an enterprise-wide AI value framework covering every possible interaction before you can start. Five to eight meaningful measures tied to deployed use cases will probably tell leadership far more than a comprehensive dashboard nobody really understands.

Establish the baseline.

Define what success means.

Decide how attribution will work.

Include the full cost.

Then review the measures at a cadence appropriate to them. Operational measures might move weekly, business outcomes monthly and strategic value quarterly.

This also changes the architecture conversation.

If you understand the economics of the workflow before selecting the technology, decisions about models, agents, human review and automation become much easier to defend.

I’ve made versions of that argument previously in Are AI Development Gains a Myth for Enterprise Software? and, on where scarce human attention should be spent, in The AI Differentiator Dividend.

The recurring theme is that the scarce resource isn’t necessarily compute or tokens. It is knowing where applying AI actually changes the economics of the business.

None of this is an argument against AI.

The individual productivity evidence is increasingly difficult to dismiss. If 80 per cent of respondents believe they are working more effectively with these tools, that matters.[1]

And the six per cent of high-performing organisations are perhaps the most important part of the McKinsey data. They demonstrate that meaningful financial value is achievable.

What they don’t demonstrate is that the value appears automatically once enough employees start using AI.

That is the assumption enterprises now need to leave behind.

The board question has changed.

It isn’t: How many of our people are using AI?

It isn’t even: How much AI are they using?

It is:

What is measurably different about the way this business operates because of AI, and how would we know if the answer were nothing?

Sources

[1] McKinsey & Company, The state of AI in 2026: On the road to ROI, August 2026. Survey of 1,719 respondents across 97 countries.

[2] Financial Times, Employers pushed staff to use AI more. That has backfired, August 2026. Reporting on enterprise AI adoption targets, incentives and Amazon’s experience with usage leaderboards.

[3] WIRED, Meta Pushes Its New AI Agent on Employees — but Eases Off on Tokenmaxxing, September 2026.

[4] BetterUp Labs & Stanford Social Media Lab, Workslop: The Hidden Cost of AI-Generated Busywork, research into the organisational cost of low-quality AI-generated work. See also Niederhoffer, Kellerman, Lee, Liebscher, Rapuano and Hancock, AI-Generated ‘Workslop’ Is Destroying Productivity, Harvard Business Review.

[5] Workday, The AI Productivity Paradox, January 2026. Global research examining AI time savings, rework and realised productivity.

[6] Shopify Enterprise, AI ROI: How to Calculate Returns in 2026, 2026. Guidance on establishing baselines and measuring AI at workflow level.

[7] Larridin, State of Enterprise AI 2026, February 2026. Survey of 365 senior leaders in organisations with more than 1,000 employees, including measurement ownership and governance.

Analysis and interpretation are my own. Figures are drawn from the cited research and reporting available at the time of publication.

Feel free to share on

Recent Posts

Related blog posts

aerial photography of building surrounded with body of water and trees during daytime

What Is My Moat, and Will It Protect Me From Disintermediation?

brown framed sunglasses on map

Why Is AI Trusted for Planning, but Not for Booking?

Prompt Injection Is Not SQL Injection, and the Difference Is Strategic