At a glance
Source: McKinsey, The State of AI: Global Survey 2026 (n = 1,719 executives, 97 countries, fielded 4 May - 8 June 2026).
The same picture keeps coming back in conversations with companies. The licenses are paid for, people praise the tools, and the leadership team cannot point to where any of it shows up in the results. The question "where is the ROI on AI" usually arrives nine to twelve months in, and usually there is no good answer to it.
The answer nobody wants to hear is that the question came too late. Return on AI is not created when the tool is bought. It is created at the moment somebody decides what will happen to the recovered time, and measures the baseline so the claim can later be proved. A company that did neither does not have a technology problem. It has a missing design.
Chapter oneThree numbers that describe the whole phenomenon
The cleanest picture comes from McKinsey's global survey: 1,719 executives across 97 countries, fielded between 4 May and 8 June 2026 and published on 25 August 2026.1
Eighty percent of users report higher individual productivity. Thirty-seven percent of organizations attribute any EBIT impact to AI, and that number did not move year over year despite growing budgets. Six percent qualify as AI high performers: they attribute at least 5% of EBIT to AI and describe the impact as significant. That share is flat too.
These numbers do not say AI fails to work. They say it works at the level of the individual and stops before the level of the firm. Between the two sits a gap, and the gap is the work.
The louder figure - 95% of generative AI deployments with no measurable P&L impact, from the MIT NANDA report of August 2025 - makes the same point more forcefully on weaker evidence: 52 interviews, 153 survey responses, more than 300 initiative reviews, no peer review, and a definition of success narrowed to a return within roughly six months.2 Treat it as a directional signal, not as proof. The McKinsey numbers are enough to carry the argument.
Chapter twoWhat happens to the recovered hour
Assume an employee saves one hour a day thanks to AI. That assumption has support. In a study of 5,179 customer support agents, access to a generative assistant raised productivity by 14% on average, with 34% for the least experienced staff and close to nothing for the best.3 In a field experiment with 758 BCG consultants, tasks inside the model's capability frontier were completed 12.2% more often and 25.1% faster.4
The question is what happens to that hour next.
Capturing the value is a leadership decision, not a side effect of deployment.
The hour can be captured: converted into higher volume, shorter response times, lower unit cost, better quality, or people moved to higher-value work. It can also dissipate: absorbed as the task expands to fill the available time, spent by colleagues verifying and fixing output, converted into work nobody asked for, or simply invisible because nobody measured the starting point.
By default, the second happens.
This does not mean raising pressure or cutting headcount. Gartner estimates that around 80% of organizations report workforce reductions, and that those reductions do not translate into returns.5 Cutting cost without changing the process buys one-off budget relief, not durable advantage. A capture decision sounds like "we will handle 30% more cases with the same team" or "we will cut response time from 26 hours to 8". It only has to be made, written down and measured - that is the entire difference.
Chapter threeWhy "it feels faster" is not a measurement
METR ran a randomized controlled trial with 16 experienced open-source developers on 246 real tasks in their own repositories, where they averaged five years of prior experience.
Before the study, participants forecast a 24% speed-up. After doing the work, they judged themselves 20% faster. The measurement showed they were 19% slower.7 The distance between perception and fact was 39 percentage points - among competent, motivated people working on their own code.
This is one domain and a small sample. It should not be generalized to all office work, and the claim here is not that AI slows people down. The conclusion is methodological, and it is hard: an employee's self-report is not productivity data. If the only source of knowledge about AI's effect is a user survey, the company is not measuring the effect. It is measuring the mood.
The BCG consultant experiment showed the other side of the same coin. On a task deliberately placed beyond the model's reach, consultants using AI were 19 percentage points less likely to arrive at the correct answer - while producing answers that were more coherent and more persuasive.4 Confidence rising faster than accuracy is the most expensive effect a deployment can produce, because the cost surfaces at the customer.
Chapter fourThree archetypes of value capture
Companies differ less in which models they use than in which of three levels they stopped at.
Level one is the most common scenario: the company grants access to tools, sometimes runs a training session, and stops there. What is missing is usage measurement, knowledge of the real saving, a link between the tool and a specific process, and a mechanism for translating savings into KPIs. The cost is certain; the benefit is asserted.
Level two begins when AI stops being an employee's tool and becomes part of a process: the company measures handling time, volume, unit cost and quality before and after, and defines a new performance standard. This is where value shows up in metrics rather than in opinions.
Level three means changing the operating model: redesigning processes, changing roles, dividing work differently between people and agents. It is not mandatory and rarely a sensible starting point. But it is the level that correlates with financial results: among companies reporting material EBIT impact from AI, close to three-quarters have redesigned their workflows, against one-quarter of the rest.1
The order matters, and it runs against instinct. Not "we will deploy AI and then see what can be improved", but "we will decide what the process should look like, and only then put AI into it".
Chapter fiveThe J-curve: why year one looks like failure
Brynjolfsson, Rock and Syverson described the pattern in which a new general purpose technology first depresses measured productivity and only later raises it as the productivity J-curve.
The mechanism is simple. Spending on technology is visible at once: the license invoice, the integration cost. Complementary spending - redesigning processes, new roles, cleaning up data, quality control, training - is incurred early, is several times larger, and is not booked as an asset. For the earlier wave of computerization, the authors estimated that the organizational capital accompanying an IT investment can be up to ten times the investment itself.8
The practical consequence: in the first period the measured effect is often negative, and that is normal rather than proof of failure. The second consequence is less comfortable. A company building organizational capital and a company that bought licenses and is waiting look identical at this point. Both show cost without effect. The only thing separating them is whether the complementary investment is being made at all.
Which is why the six-month checkpoint question is not "how much have we saved". It is: what exactly did we change in the workflow, and do we have a baseline to compute against.
Chapter sixThe programme: one process, four moves
There is no need to rebuild the company. There is a need to take one process all the way through, measurement included, and only then scale what held up.
Find the people who have already automated their own work
Most organizations have them. They do it quietly and do not call it an AI project. Ask what they automated, which problem it solved, how much time it saves, and whether it transfers to a larger group. It is the cheapest source of good first ideas, because it comes from people who know the process from the inside.
Pick one process, not a portfolio
The candidate has to be repetitive, high in volume, staffed with predictable manual work, measurable in output, and important enough for an improvement to register at the business level. In practice: customer service, document processing, invoices, forms, offer generation, moving data between systems.
Measure the current state before deploying anything
This is the one step that cannot be made up later. Without a baseline, every subsequent result is an opinion. The eight questions in the table below are enough to start and usually take two weeks.
Put AI inside the process, not beside it
If an employee copies data from a system into a tool, waits, verifies, and moves the result back by hand, the saving eats itself. The human role belongs in the design, not bolted on afterwards: in customer service AI takes the repetitive queries and a person takes escalations; in documents AI extracts and a person approves; in content AI drafts and an expert owns quality and the decision.
| Dimension | The question you need a number for before deployment |
|---|---|
| Time | How many minutes does one case take, from intake to close? |
| Involvement | How many people touch one case, and at which stage? |
| Unit cost | What does handling one case cost at fully loaded labour cost? |
| Volume | How many cases per month, with what seasonality? |
| Quality | How many errors, reworks, complaints and escalations per hundred cases? |
| Delay | Where does a case wait, and how long does it sit in a queue rather than in work? |
| Variance | What separates the fastest from the slowest decile of cases? |
| What we do with the time | What exactly will the recovered hours be converted into - volume, response time, quality, or not hiring as scale grows? |
Chapter sevenArithmetic you can put in front of a CFO
The starting model does not need to be complicated:
| Line | Basis | Value |
|---|---|---|
| Volume | Documents per month | 4,000 |
| Time before | Minutes per document, manual handling | 11.0 |
| Time after | Minutes per document, AI extraction plus human verification | 4.5 |
| Time recovered | Hours per year | 5,200 |
| Rate | Fully loaded hourly cost | PLN 50 |
| Gross benefit | 5,200 h × PLN 50 | PLN 260,000 |
| Annual cost | Licenses, integration, maintenance | PLN 80,000 |
| Net result | Per year | PLN 180,000 |
| Unit cost | Per document, before and after (tooling included) | PLN 9.17 → 5.42 |
And this is where the calculation stops, because PLN 180,000 is so far a number on a slide. It turns into money only once the company decides what the 5,200 recovered hours are actually spent on. That is the eighth row of Table 1, and it is where most AI projects end without a result: not because the saving failed to materialize, but because nobody claimed it.
| Option | What the 5,200 hours become | The effect that shows up in results |
|---|---|---|
| Volume | The same team handles more documents | Throughput rises from 4,000 to roughly 9,800 documents a month with no increase in headcount |
| Cost | No new hires even as scale grows | About 3 FTE fewer than in the no-AI scenario, that is PLN 260,000 of cost not incurred per year |
| Time and quality | Shorter turnaround, part of the hours goes to control | Faster handling and fewer errors reaching the customer |
Chapter eightWho should act now
Not every company needs to rebuild its operating model today. Pressure rises fastest where a large share of the work is processing information at a computer: documents, customer queries, offers, data, content, repetitive actions in business systems. The smaller the physical component, the sooner AI moves cost, time and volume.
Pressure is also rising from the other direction - from costs and from disappointment. Gartner forecasts that more than 40% of agentic AI projects will be cancelled by the end of 2027 because of escalating costs, unclear business value and inadequate risk controls.9 At the same time, one company in five already limits its use of AI because of operating costs.1 The window in which "we are experimenting" justifies the spend is closing.
Companies that learn to measure and capture this value earlier will run faster and cheaper. Those that stop at buying licenses will eventually conclude that AI does not work. The problem will not be the technology. It will be the absence of a process, a metric, and a decision about what is supposed to change.
ConclusionStart with the process and the metric, not the license
Return on AI is not a function of model quality. It is a function of three decisions taken before deployment: which process, what baseline, and what happens to the recovered time. Companies that took those decisions are today in the six percent. Companies that did not have tools, costs and satisfied users - and a P&L in which none of it appears.
AI is not magic. It creates value when an organization knows where to apply it, how to measure the effect, and how to turn an individual's productivity into the performance of the whole.
Notes and sources
- McKinsey & Company, "The State of AI: Global Survey 2026", published 25 August 2026; fielded 4 May - 8 June 2026, n = 1,719 executives across 97 countries, weighted by each country's share of global GDP. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
- MIT NANDA, "The GenAI Divide: State of AI in Business 2025", August 2025. Base: 300+ initiative reviews, 52 interviews, 153 survey responses. The report was not peer reviewed; criticism centres on the narrow definition of success (measurable return within roughly six months) and the small interview base. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
- Erik Brynjolfsson, Danielle Li and Lindsey Raymond, "Generative AI at Work", The Quarterly Journal of Economics 140(2), 2025, pp. 889-942. Base: 5,179 customer support agents. https://academic.oup.com/qje/article/140/2/889/7990658
- Fabrizio Dell'Acqua et al., "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality", Organization Science, 2025 (earlier HBS Working Paper 24-013). Base: 758 BCG consultants. https://pubsonline.informs.org/doi/10.1287/orsc.2025.21838
- Gartner, "Autonomous Business and AI Layoffs May Create Budget Room, but Do Not Deliver Returns", press release, 5 May 2026. https://www.gartner.com/en/newsroom/press-releases/2026-05-05-gartner-says-autonomous-business-and-artificial-intelligence-layoffs-may-create-budget-room-but-do-not-deliver-returns
- Kate Niederhoffer, Gabriella Rosen Kellerman et al., "AI-Generated 'Workslop' Is Destroying Productivity", Harvard Business Review, September 2025. Study by BetterUp Labs and the Stanford Social Media Lab, n = 1,150 US full-time employees. The cost estimate rests on self-reported salaries and self-reported time. https://hbr.org/2025/09/ai-generated-workslop-is-destroying-productivity
- METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity", 10 July 2025. Base: 16 developers, 246 tasks, using Cursor Pro with Claude 3.5/3.7 Sonnet. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
- Erik Brynjolfsson, Daniel Rock and Chad Syverson, "The Productivity J-Curve: How Intangibles Complement General Purpose Technologies", American Economic Journal: Macroeconomics 13(1), 2021, pp. 333-372. https://www.aeaweb.org/articles?id=10.1257/mac.20180386
- Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027", press release, 25 June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027