From Hours Saved to Value Captured
Artificial intelligence has become remarkably easy to make look productive.
A company can count users. Prompts. Copilot licenses. Automated tasks. Generated documents. Use cases. Pilots. Minutes saved. Estimated hours returned to employees.
Almost all of these numbers can move in the right direction while the economics of the business remain essentially unchanged.
That is the measurement problem now emerging behind enterprise AI.
The question is no longer whether AI can make individual activities faster. There is enough evidence to show that, in the right settings, it can. The more difficult question is whether companies can convert those local improvements into economically meaningful changes in throughput, cost, revenue, quality, risk, or strategic capability.
Those are not the same thing.
A large 2026 research project based on surveys of nearly 6,000 senior executives across the United States, United Kingdom, Germany, and Australia illustrates the gap. Sixty-nine percent of firms reported some current use of AI. Yet roughly nine in ten executives reported no impact from AI on either employment or labor productivity over the preceding three years. The same executives expected substantially larger effects in the years ahead. This is survey evidence, not a causal estimate, but the contrast is striking: adoption is already widespread; measurable enterprise effects are not.
That does not prove that AI has failed. It shows that AI adoption and AI value realization are different variables.
And that is where many AI business cases go wrong.
They measure the intervention rather than the result.
The Productivity Illusion
There is good empirical evidence that generative AI can improve performance on specific tasks.
In a preregistered experiment published in Science, Shakked Noy and Whitney Zhang assigned 453 college-educated professionals realistic writing tasks. Participants with access to ChatGPT completed the work about 40 percent faster, while independent quality ratings increased by 18 percent. That is a substantial result. But it is a result about a defined class of writing tasks, not evidence that the organizations employing those people would become 40 percent more productive.
A much larger workplace study by Erik Brynjolfsson, Danielle Li, and Lindsey Raymond, published in The Quarterly Journal of Economics in 2025, followed more than 5,000 customer-support agents at a Fortune 500 software company. AI assistance increased issues successfully resolved per hour by around 15 percent on average. Less experienced workers benefited substantially more, while the most experienced workers saw much smaller gains and, in some cases, slight quality deterioration. Importantly, the researchers themselves caution against generalizing a study of one tool, occupation, and company into broader labor-market conclusions.
The same pattern appears in knowledge work. A field experiment involving 758 Boston Consulting Group consultants found that, on tasks inside what the researchers called the AI capability frontier, participants using GPT-4 completed 12.2 percent more tasks and worked 25.1 percent faster, with improved quality. But on a deliberately selected task outside that frontier, AI users were 19 percent less likely to arrive at the correct answer. The paper, originally circulated in 2023, was published in Organization Science in 2026.
Software development provides an even clearer warning against generalization. A controlled GitHub Copilot experiment found developers completed one defined JavaScript programming task 55.8 percent faster with AI assistance. But a 2025 randomized study by METR, involving 16 experienced open-source developers working on 246 issues in repositories they knew well, found the opposite: allowing early-2025 AI tools increased completion time by 19 percent. METR later reported that its follow-up work using newer tools suffered from serious selection effects, making the magnitude of any newer speed-up difficult to estimate reliably.
The correct conclusion is not that AI is productive or unproductive.
It is that productivity is conditional.
It depends on the task, the worker, the model, the workflow, the quality threshold, the surrounding process and, ultimately, what the organization does with the improvement.
That last part is where management accounting begins.
⚖️ A Minute Saved Is Not a Euro Earned
Consider one of the most common AI ROI calculations.
Ten thousand employees use an AI assistant. Each reports saving thirty minutes per week. Finance multiplies those hours by an average loaded labor rate and presents the result as millions of euros of annual value.
The arithmetic may be correct.
The economics may not be.
If those employees remain employed at the same salaries, work the same contracted hours, produce the same business output and simply absorb the released time into the rest of their working week, almost none of that theoretical value appears in the P&L.
There is nothing wrong with measuring the time saving. It is useful information. The problem starts when a time-equivalent benefit is presented as a cash-equivalent benefit.
Salaried labor is largely fixed in the short term. Saving ten percent of a person’s time does not reduce ten percent of that person’s salary. The financial consequence appears only if the organization can capture the released capacity.
That might mean processing more cases with the same team. It might mean avoiding a planned hire. It might reduce overtime or external contractor spending. It might allow employees to perform higher-value activities. It might increase sales capacity. It might improve service levels enough to affect retention. Or several fragmented savings might eventually allow work to be consolidated into a smaller operating structure.
But until something like that happens, the organization has released time, not necessarily created economic value.
This suggests a distinction that should become standard in AI management reporting:
Gross time saved is the estimated amount of human effort no longer required for an activity.
Economically captured capacity is the portion of that released effort that the organization successfully converts into additional output, avoided cost, improved quality, increased revenue, reduced risk, or another observable business outcome.
The conversion rate between the two may be high.
It may also be close to zero.
And management, not the AI model, determines much of that difference.
Faster Work Does Not Automatically Mean a More Productive Company
The mistake becomes easier to see if AI productivity is measured at several levels.
Task → Workflow → End-to-End Process → Business Capability → P&L / Customer / Risk Outcome
At the task level, AI may reduce the time required to prepare a document from four hours to two.
At the workflow level, the relevant question becomes whether that document now moves through the surrounding activities faster.
At the process level, we ask whether the end-to-end cycle time, throughput, error rate or cost has changed.
At the capability level, we ask whether the organization can now handle more customers, settle more claims, launch products more quickly, manage greater complexity, improve forecasting or operate with a structurally different cost base.
Only then do we arrive at the outcomes that matter economically: revenue, margin, working capital, customer retention, loss avoidance, service performance or risk.
The farther management moves toward the right-hand side of that chain, the more meaningful the measurement generally becomes.
This does not make task metrics useless.
Task metrics are excellent diagnostic measures. They tell us whether the technology is doing something useful.
They simply should not automatically be called enterprise value.
⚙️ The Bottleneck Determines the Economics
Operations management has understood this problem long before AI existed.
Improving one activity does not necessarily improve the system.
Cycle time contains both processing time and waiting time. Bottlenecks, queues, capacity constraints and variability determine how local improvements propagate through an end-to-end process. MIT operations-management material, for example, explicitly treats bottleneck analysis as a way of understanding capacity and waiting rather than merely looking at the speed of individual activities.
Imagine that preparing a contract takes four hours and AI reduces this to two.
That sounds like a 50 percent productivity improvement.
But suppose the contract then waits three days for legal approval.
The customer experiences almost no difference.
Now imagine that AI increases the number of contracts prepared each day. The legal queue becomes longer. The local improvement may actually increase work in progress without improving completion time.
The same pattern occurs everywhere.
Marketing can generate more leads than Sales can process.
AI can produce more software code than testing and release processes can absorb.
Automated analysis can create more fraud alerts than investigators can review.
Customer-service AI can generate responses faster while mandatory human verification becomes the new bottleneck.
An agent can prepare decisions instantly while management committees still meet once per week.
This is value leakage: technically successful AI whose local benefit disappears elsewhere in the system.
The economically relevant question is therefore not simply:
What can we automate?
It is:
Where is the constraint that limits the outcome we care about, and does AI materially change that constraint?
The easiest task to automate is not necessarily the most valuable constraint to remove.
📊 Measure the Process, Not the Prompt
Once the unit of analysis changes, the KPI system changes with it.
For customer service, prompts per employee tell management very little about economic performance. Cost per resolved case, resolution rate, first-contact resolution, customer satisfaction, queue time and cases handled per paid hour are far more informative.
For claims handling, the relevant measures might be cost per claim, cycle time, loss-adjustment expense, leakage, rework and first-time-right rates.
For software development, generated lines of code are almost meaningless by themselves. Lead time to production, deployable features, escaped defects, rework, incident frequency and engineering cost per delivered capability move closer to the actual production system.
For sales, the relevant denominator might be qualified opportunities, conversion, sales-cycle duration, gross profit or revenue generated per seller.
For contracting, it may be cost per executed contract, turnaround time, exception rate or contracts processed per legal FTE.
There is no universal enterprise AI KPI.
That is precisely the point.
Enterprise AI needs a measurement system, not a magic KPI.
Unit economics are particularly powerful because the denominator forces the business case to touch operational reality.
If AI spending rises from €100,000 to €500,000 but cost per successfully completed transaction falls materially, the system may be becoming economically better.
If model costs fall by 70 percent while cost per business outcome remains unchanged because human review and exception handling have increased, the cheaper model has not solved the economic problem.
Tokens are a technology unit.
The business needs a business unit.
AI Has More Than One Economic Mechanism
Another reason productivity alone is insufficient is that labor efficiency is only one way AI can create value.
The first mechanism is efficiency: performing substantially the same output with fewer resources, less effort or lower external expenditure.
The second is capacity: producing more with essentially the same resource base. A customer-service operation that can absorb twenty percent volume growth without twenty percent more headcount has created genuine economic value even if nobody is laid off.
The third is effectiveness. AI may make the organization better rather than merely faster. Fewer errors, better prioritization, improved detection, better decision preparation, more consistent execution, reduced rework or improved first-time-right rates may have larger financial consequences than administrative time savings.
The fourth is growth and innovation. AI may allow more opportunities to be evaluated, shorten time to market, enable a previously uneconomic service, improve asset utilization or create a new product entirely.
There is also a fifth category that deserves caution: strategic option value.
A reusable AI platform, enterprise knowledge layer or agent infrastructure may not initially have a clean standalone P&L return. It can create an option to deploy future capabilities more cheaply and quickly.
That is legitimate.
But “strategic” cannot become the word companies use when they have stopped measuring.
A strategic investment should still have an explicit hypothesis: what capability is being created, which future use cases depend on it, what alternatives exist, what milestones demonstrate that the capability is becoming real, and what evidence would cause management to stop investing.
Exploration does not require fictional precision.
It still requires discipline.
Cost Reduction Is Not Cost Avoidance
AI also forces finance to distinguish mechanisms that are routinely mixed together.
If automation allows a team of 100 to become a team of 90, that is potentially a cost reduction, subject to transition costs and timing.
If the team remains at 100 but can absorb growth that would otherwise have required ten additional hires, that is cost avoidance.
If the team remains at 100 and produces more, that is capacity improvement.
If the team remains at 100, produces the same amount, but uses released time for work that improves quality or revenue, the value lies somewhere else.
These outcomes are economically different even if all four began with exactly the same task-level time saving.
This is why the debate about whether AI “reduces headcount” is often badly framed.
AI can eliminate activities without eliminating positions. It can alter skill mixes. It can reduce overtime. It can reduce external-service dependency. It can allow growth without proportional staffing growth. It can shift people from execution toward exception handling, customer interaction or decision work.
The financial effect depends on which of those mechanisms actually occurs.
Avoided hiring can be extremely valuable.
But avoided hiring is not the same accounting event as reducing today’s payroll.
Management reporting should stop pretending otherwise.
💶 Gross Benefits Are Not Net Benefits
The second half of AI economics is cost.
AI business cases still have a tendency to compare a gross productivity estimate with the price of an AI license.
That is rarely the correct denominator.
A useful management lens is:
Net AI Value = Attributable Economic Upside – Full Production Cost – Risk-Adjusted Downside
This is not meant as a universal accounting formula. Different use cases require different valuation methods. It is a reminder of what needs to be inside the business case.
Full production cost includes much more than the model.
Depending on the application, it includes model or API consumption, compute, platform licensing, data preparation, integration, engineering, evaluation, observability, security, infrastructure, human review, vendor management, training, organizational change and ongoing maintenance.
AI also changes the cost structure of software.
Compared with conventional per-seat SaaS, many model services introduce a meaningful usage-dependent component. Input, output and cached tokens can have different prices. Context size matters. Agentic workflows can call models repeatedly. Retrieval, search and external tools can add additional consumption. Model routing, caching and batching can materially change unit cost.
The FinOps Foundation now treats AI as a distinct cost-management domain precisely because token-based consumption, heterogeneous pricing, forecasting volatility and cost allocation create new management challenges. Its current guidance explicitly recommends moving beyond token metrics toward workload-specific unit economics such as cost per call or cost per completed business outcome.
This does not mean model costs will necessarily dominate AI economics. In many enterprise applications, they may be small compared with labor, engineering or integration.
But their variable nature means they should be observable.
An AI system whose execution cost is invisible to its product owner does not yet have mature unit economics.
🧪 The ROI of a Demo Is Not the ROI of an Operating Capability
There is another structural problem with AI business cases.
Proofs of concept are economically unusual environments.
A pilot may use a carefully selected dataset. It may operate on a few hundred transactions. Experts may supervise every output. Integration may be manual. Security controls may be temporary. Exceptions may simply be excluded. Innovation teams may absorb the work. Nobody has yet built 24/7 monitoring, fallback logic, auditability, production support, evaluation pipelines or change processes.
Then the pilot works.
And the pilot’s economics are projected onto production.
Long before generative AI, Google researchers described a similar pattern in machine learning. Their 2015 paper on hidden technical debt argued that machine-learning systems could be developed and deployed relatively quickly while accumulating substantial ongoing system-level maintenance cost through data dependencies, entanglement, configuration, feedback loops and monitoring requirements. Later Google work on production readiness made testing and monitoring explicit parts of operating reliable ML systems.
Generative and agentic systems do not remove this distinction. They add new variables.
A demo can ignore what happens when the model is wrong.
A production process cannot.
If the system needs a human to verify every output, that review time belongs in the economic model. If false positives create downstream work, that work belongs in the model. If an agent loop can unexpectedly make twenty calls rather than five, that variability belongs in the model. If confidential output requires additional controls, those controls belong in the model.
NIST’s Generative AI risk profile identifies issues including confabulation, information integrity, privacy, harmful bias and problematic human-AI configurations. Not every risk can or should be converted into a convenient euro amount. But relevant failure modes should be part of the investment decision rather than treated as somebody else’s governance problem.
The correct business case is therefore not:
What does success cost?
It is:
What does reliable production at the required quality and risk level cost?
Those numbers can be very different.
The AI Business Case Starts Before the AI
Good measurement begins with a baseline.
Before AI changes a process, management should know how that process performs today.
How many cases enter? How many leave? How much human effort is required? What is the end-to-end cycle time? Where does work wait? What is the error rate? What is reworked? What does the process cost? What service level is achieved? How many people are required at current volume? What happens when volume grows? Where is the binding constraint?
Without this baseline, AI is compared with memory and intuition.
That creates very weak business cases.
The measurement design should also match the size of the claim.
If the claim is that AI makes document drafting faster, a controlled before-and-after comparison may be adequate.
If the claim is that AI increases sales conversion, the organization should be much more careful about attribution. Sales performance changes for many reasons simultaneously.
Where practical, randomized trials, A/B designs, phased rollouts, matched comparison groups or difference-in-differences approaches can provide better evidence. The large customer-support study discussed earlier, for example, exploited a staggered deployment and used difference-in-differences techniques rather than simply comparing performance before and after AI appeared.
Not every corporate process needs an academic experiment.
But every serious investment needs a credible answer to:
Compared with what?
and
How do we know AI caused the change?
Quality Is an Economic Variable
One of the biggest weaknesses in productivity discussions is the assumption that quality is secondary.
It is often the opposite.
A process that becomes 30 percent faster while rework increases 20 percent may have become economically worse.
A fraud-detection system that identifies more cases but floods investigators with low-quality alerts may reduce effective investigative capacity.
A developer assistant that generates code faster but creates additional testing, security or maintenance work may move effort rather than remove it.
Conversely, an AI system that saves almost no labor may still be valuable if it materially reduces costly errors, improves consistency or prevents losses.
This is particularly important in high-consequence processes, where the economic value of one avoided failure may exceed thousands of small time savings elsewhere.
The BCG consultant experiment is useful precisely because it measured more than speed. Performance improved substantially on tasks within AI’s capability frontier while correctness deteriorated on a task outside it. Faster execution, therefore, cannot be interpreted independently of what is being executed.
Quality does not have to be monetized artificially.
Management can track defect rates, complaints, rework, first-time-right performance, escalation, losses or other observable consequences and only translate them into money where a defensible relationship exists.
False precision is not financial discipline.
It is merely precise-looking uncertainty.
From AI Use Cases to AI Portfolio Economics
Individual business cases are only the beginning.
A large enterprise will not eventually run five AI use cases. It may run hundreds or thousands of AI-enabled workflows, embedded across products and operations.
That requires portfolio discipline.
Some initiatives will have immediate economic impact but little reuse. Others may create reusable capabilities that make ten downstream initiatives cheaper. Some will have high theoretical value but poor technical feasibility. Some will be technically impressive but economically irrelevant. Some should continue because they are generating valuable learning. Others should be stopped because the learning has already answered the question.
This is where organizations need to distinguish exploration economics from scaling economics.
An experiment can legitimately be funded to reduce uncertainty.
Its return may initially be knowledge.
But once a capability moves into scaled production, “we are learning” cannot remain its indefinite business case.
A sensible portfolio process therefore changes the evidence requirement over time.
Early-stage initiatives need a credible problem, hypothesis and learning objective.
Scaling initiatives need evidence that the technology works in the real process.
Production initiatives need unit economics, attributable outcomes, operating costs, risk metrics and an accountable owner.
The threshold should rise as the investment rises.
That approach avoids both extremes: demanding a fully quantified NPV before anybody is allowed to experiment, and funding permanent experimentation because financial measurement is considered somehow inappropriate for AI.
Finance Must Enter the Conversation
This does not mean the CFO should own AI.
It means AI has matured far enough that Finance has to become part of value realization.
The business owner should own the business outcome and baseline.
Technology should understand platform, integration, infrastructure and operating cost.
Data and AI teams should understand model performance, evaluation, data dependencies and technical limitations.
Risk, Legal and Compliance should structure the relevant exposure and controls.
Finance should challenge benefit assumptions, prevent theoretical labor savings from masquerading as cash savings, distinguish cost reduction from cost avoidance, and help determine whether expected benefits actually arrive.
Management has another role that cannot be delegated to any of them.
It must decide what happens to released capacity.
If an AI project theoretically returns 40,000 hours to an organization and nobody has decided how that capacity will be used, the business case is incomplete.
The economic value of AI does not emerge when time is saved.
It emerges when the organization deliberately captures what the technology has released.
That is a management decision.
The Productivity Paradox Is Older Than AI
There is a useful historical warning here.
In 1987, economist Robert Solow famously observed:
“You can see the computer age everywhere but in the productivity statistics.”
The sentence became shorthand for the information-technology productivity paradox: enormous excitement and visible technological adoption without a correspondingly obvious productivity effect.
Later research produced a more nuanced explanation.
Erik Brynjolfsson, Daniel Rock and Chad Syverson argued that general-purpose technologies such as AI require complementary investment: new processes, business models, organizational structures, skills and other forms of intangible capital. Benefits can therefore arrive later than the technology itself. Their subsequent work described a “Productivity J-Curve,” in which those complementary investments can initially suppress measured productivity before their benefits are realized.
This is not proof that today’s AI investment will inevitably pay off later.
Bad investments do not become good investments simply by waiting.
The historical lesson is more useful than that:
Technology deployment and economic productivity are separated by an organizational conversion process.
A 2026 study of roughly 750 corporate executives reached a strikingly similar conclusion in today’s environment: reported perceived productivity gains from AI were larger than measured productivity gains, producing what the researchers themselves describe as a productivity paradox.
The gap between technological capability and economic performance is therefore precisely where management should focus.
🧭 A Practical AI Value Test
For every significant AI investment, management should be able to answer ten questions:
- What business outcome are we actually trying to change? Not what the model will do, but what the business will do differently.
- What is the baseline? Current cost, throughput, effort, cycle time, quality, service level, conversion, staffing and relevant risk.
- What is the right economic unit? A case, claim, transaction, customer, contract, feature, order, qualified lead or another business denominator.
- Where is the actual constraint? Which part of the system limits the outcome, and is the AI intervention changing that constraint or merely optimizing a non-bottleneck activity?
- What does AI change causally? Task time, throughput, quality, decision accuracy, conversion, loss rate, capacity or something else.
- How will released capacity be captured? Additional output, avoided hiring, reduced overtime, lower external spend, different work, higher service levels or a redesigned organization.
- What are the full production economics? Model consumption, technology, integration, data, evaluation, monitoring, review, security, governance, support and organizational change.
- What new failure modes and downstream costs are introduced? Including errors, rework, verification, customer effects and operational risk.
- How will attributable impact be measured? Through a baseline and, where justified, controlled or comparative measurement rather than post-hoc storytelling.
- What decision threshold has been agreed in advance? What evidence causes the organization to scale, redesign, pause or terminate the initiative?
If these questions cannot be answered, the problem is not necessarily that the AI initiative is bad.
It means management does not yet understand its economics.
Productivity Is Not the Wrong Idea. It Is Usually Measured at the Wrong Level.
The title of this article is deliberately provocative.
Productivity is not irrelevant.
The mistake is measuring it too locally.
A person completing a task faster is productivity.
A team creating additional usable capacity is productivity.
A process increasing throughput without increasing cost is productivity.
A company reducing cost-to-serve while maintaining quality is productivity.
A business making better use of its assets, people and capital is productivity.
But these are not interchangeable measures.
AI adoption measures activity.
Task productivity measures local performance.
AI economics measures consequences.
The management challenge for the next stage of enterprise AI is therefore not to collect increasingly spectacular statistics about prompts, users and theoretical hours saved.
It is to trace the chain from technological intervention to economically captured outcome.
That requires baselines.
It requires unit economics.
It requires understanding constraints.
It requires full production cost.
It requires quality and risk.
It requires attribution.
And it requires managers to decide what happens after AI has made something faster.
The question is no longer whether AI can save time. In specific settings, the empirical evidence says clearly that it can.
The more important question is whether the organization can convert that capability into measurable economic performance.
A minute saved is an operational metric.
What the organization does with that minute is the business case.




Join the discussion