In 1954 the psychologist Paul Meehl published Clinical versus Statistical Prediction, a short book reviewing the evidence on whether expert judgment or simple actuarial formulas made better predictions. The formulas won. Not exotic formulas — often little more than weighted checklists. The finding was replicated for decades and formalized in Grove and colleagues’ 2000 meta-analysis: mechanical prediction was on average about 10% more accurate than clinical judgment, substantially outperformed the clinicians in a third to half of studies, and lost substantially in as few as 6–16%. Every modern argument about “trusting the model” is a rerun of an argument psychology settled while Eisenhower was president.
The result predates the current AI wave by half a century, and the industry with the most to gain from it has only recently gotten around to naming the problem. Gartner published its first Magic Quadrant for Decision Intelligence Platforms in January 2026, and predicts that by 2027 half of business decisions will be AI-augmented or AI-automated. The name it has settled on is decision intelligence: the discipline of engineering decisions — not dashboards, not models, but the full path from information to action — using whatever combination of data, analytics, AI, and human judgment the decision actually warrants. Cassie Kozyrkov, who served as Google’s first Chief Decision Scientist, compresses it to “turning information into better action at any scale”; Lorien Pratt, who co-originated the framing and wrote its first book-length treatment (Link, 2019), draws it as a chain linking data to actions to outcomes.
That is the market answer. The scientific answer is older, stranger, and far more useful — and since my day job sits exactly at this seam, what follows is the evidence I think every builder of decision systems should be forced to read.
If mechanical prediction has been better for seventy years, why isn’t everything decided by formula? Because two failure modes stand in the way, and they pull in opposite directions. Managing both is most of what decision intelligence amounts to.
Algorithm aversion: in Dietvorst, Simmons, and Massey’s 2015 experiments, people who watched an algorithm err lost confidence in it faster than they lost confidence in an erring human — and abandoned the algorithm even when it demonstrably outperformed. The 2018 follow-up in Management Science found the antidote almost embarrassingly cheap: let people modify the algorithm’s output, even slightly and within strict limits, and usage jumps — along with performance. A little control is enough to keep people engaged, not full command over the algorithm.
Automation bias: the opposite disease, documented in aviation before it reached AI. Parasuraman and Riley’s classic taxonomy named the axis — use, misuse, disuse — and Mosier, Skitka and colleagues’ cockpit studies gave the two signature errors their names: omission (missing what the automation fails to flag) and commission (following a wrong automated directive past visible contraindications). The clinical literature finds the same pattern in decision-support systems: better on average, plus a new class of errors from overreliance.
Between those failure modes sits the assumption underneath most enterprise AI: humans plus AI must be better than either alone. It was finally put to a preregistered test. Vaccaro, Almaatouq, and Malone’s 2024 meta-analysis in Nature Human Behaviour pooled 106 studies and 370 effect sizes, and the headline finding deserves to be framed on the wall of every AI steering committee: on average, human–AI combinations performed worse than the best of the human or the AI alone.
Read the second bar carefully, because it explains the industry’s confusion: combinations beat the human alone handily (g = +0.64), which is what every vendor demo shows you. They lost to the best performer — often the AI alone — which is what the deployment eventually discovers. The moderator analysis carries the practical lesson: when the human was the stronger performer, adding the AI helped; when the AI was stronger, adding the human dragged it down. And the one task family where combinations reliably won was open-ended creation work, not selection among options.
None of this puts humans out of the loop — I have argued elsewhere in this series that in high-stakes domains the human checkpoint is mandatory. Complementarity just has to be built on purpose, because it does not happen by default. The evidence points to a short list of mechanisms that work:
Divide by comparative advantage, not by ceremony. The meta-analysis is blunt: a human reviewing every AI output adds error to the cases the AI had right. Route by confidence and stakes — the machine decides where it is demonstrably strong, the human owns the exceptions, the low-confidence cases, and everything with a signature. Bayesian modeling work in PNAS shows hybrids can beat both parties even at unequal accuracy — but only within bounds set by how correlated their errors are, and eliciting confidence, not just answers, is what makes the combination work.
Give the human a dial, not a veto. Dietvorst’s antidote — bounded modification of the model’s output — converts aversion into engagement at almost no accuracy cost.
Make accountability explicit. In the cockpit studies, pilots who felt personally accountable for outcomes double-checked the automation more and committed fewer errors. Diffuse accountability is the substrate automation bias grows in.
Train the judgment itself. The Good Judgment Project showed forecasting accuracy responds to teachable technique — debiasing training, reference-class thinking, teaming, and tracking top performers — which is worth remembering in an era that treats human judgment as a fixed quantity to be routed around. (A 2025 reanalysis questions how much of those gains survive tighter controls — in this field, even the evidence about judgment needs judgment.)
Gartner’s new Magic Quadrant will pull the term “decision intelligence” through every vendor deck in 2026, and the research history insists on a distinction those decks will blur. A decision intelligence platform is software. Decision intelligence itself is the discipline of taking Meehl seriously — formulas where formulas win, humans where humans win, and honest measurement of which is which — while designing around the two documented failure modes of the humans involved. Organizations that buy the platform without the discipline will automate their decisions and keep their biases. The ones that get it right will look less like “AI making decisions” and more like a well-run cockpit: machine flying, human accountable, both calibrated, and every landing graded.
References