Two numbers, both published in 2026, both nominally answering the same question about whether corporate AI spending is paying off. EY’s pulse survey reports 98% of AI-investing leaders seeing positive ROI. Bain’s April 2026 survey of 951 companies found just 4% achieving cost savings above 30%, while 90% of the underperformers plan to increase budgets anyway, a pattern Bloomberg dryly headlined a “circular bet.” Somewhere in between, McKinsey’s State of AI finds only 37% of organizations attributing any earnings impact to AI, which is to say that 63% report none.
These describe one economy in one year. The tempting conclusion is that someone is lying. The difference is in the instrument: surveys measure belief, stated by self-selected leaders about differently worded questions, while controlled experiments measure work, and the experiments agree with each other almost boringly well. Sorting the 2026 evidence into those two piles is most of the analysis.
Line up the year’s headline consultancy statistics, every one of them a self-reported survey of its own client-adjacent population, and the spread tells you more than any single bar:
To be fair to the firms, the fine print often converges even when the headlines don’t. McKinsey’s own detail shows only ~6% of companies are “high performers” with more than 5% of earnings attributable to AI, and BCG’s long-running “10-20-70” framing — 10% of the value in algorithms, 20% in data and technology, 70% in people and process — says in consultant-speak what I argued in the level-set essay about the 80/20 of AI value: the technology was never the constraint. KPMG’s Q2 2026 pulse found organizations with full AI cost visibility were five times more likely to report established ROI, which is either an insight about discipline or a tautology about who is capable of measuring anything. Probably both.
The year’s most viral number deserves the same scrutiny from the other direction. MIT-affiliated researchers claimed that “95% of enterprise generative AI pilots show no measurable P&L impact”. The report was self-labeled preliminary, not peer-reviewed, built on ~52 interviews and 153 survey responses, and published by a group whose own agentic-infrastructure agenda benefits from the conclusion. It earns the same scrutiny this series gave the “80% unstructured data” statistic. It may well be directionally right; it rhymes with McKinsey’s 63%. It is still a weak measurement that went viral because it was quotable, and the bar for numbers we repeat should not depend on whether we like the direction they point.
The second pile holds randomized and quasi-experimental studies, where somebody counted output rather than asking about it. Their consistency, across completely different settings, is striking:
Three findings repeat across nearly all of them. AI reliably speeds up bounded, well-defined tasks. The gains concentrate among novices: the landmark customer-support study in the Quarterly Journal of Economics found 34% gains for the least experienced agents and roughly zero for the most expert, the opposite skill-bias of every previous IT wave, with real implications for how teams develop talent. And the BCG/Harvard “jagged frontier” experiment supplies the caution: the same consultants who gained 40% on AI-suitable tasks did worse when they trusted AI on tasks just beyond its frontier. Readers of my decision-intelligence essay will recognize automation bias wearing a lanyard.
So tasks get 15–55% faster while most firms report no earnings impact, and Acemoglu’s macro arithmetic projects a total factor productivity gain of well under 1% over a decade. Those are not contradictions. They are three altitudes.
Task gains become firm profits only when workflows are redesigned around them, and that is the intangible, unglamorous investment that the productivity J-curve literature says always lags the technology and always makes early returns look worse than they are. Two other data points describe the middle of a J-curve rather than the end of a story: McKinsey’s detail line, roughly 90% of use cases still stuck in pilot, and Gartner’s first Hype Cycle for Agentic AI, which places agents at the peak of inflated expectations alongside its now-familiar 40%-canceled-by-2027 forecast. The market is making the same wager at monstrous scale: hyperscaler AI capital spending in the hundreds of billions against cloud AI revenues growing fast from a much smaller base. Sequoia’s “$600 billion question” is two years on, still open, and still widening before it narrows.
Discount every self-reported survey to directional sentiment, and three things survive. The experimental record is real: double-digit task-level gains, biggest for novices, bounded by a jagged frontier that punishes blind trust. The organizational record is consistent: value follows workflow redesign and measurement discipline, which is the 70% in BCG’s 10-20-70, the cost-visibility effect in KPMG’s data, and the subtraction playbook I’ve argued for all year. The macro record counsels patience without complacency, because general-purpose technologies pay off on the timescale of reorganization rather than procurement.
Which leaves the awkward part. On the evidence of 2026, most companies are not measuring at any altitude: per Goldman Sachs’ analysis of earnings calls, about 54% of the S&P 500 now discuss AI productivity, roughly 11% quantify a gain on any specific use case, and only about 2% quantify its effect on earnings. The institutions that stay ahead on profitability will be the ones measuring at the right altitude, not the ones with the most optimistic survey. On current numbers, that contest has barely started.
References