← Gautam Parab

What Is Decision Intelligence? The Science of Letting Machines Recommend and Humans Decide

In 1954 the psychologist Paul Meehl published Clinical versus Statistical Prediction, a short book reviewing the evidence on whether expert judgment or simple actuarial formulas made better predictions. The formulas won. Not exotic formulas — often little more than weighted checklists. The finding was replicated for decades and formalized in Grove and colleagues’ 2000 meta-analysis: mechanical prediction was on average about 10% more accurate than clinical judgment, substantially outperformed the clinicians in a third to half of studies, and lost substantially in as few as 6–16%. Every modern argument about “trusting the model” is a rerun of an argument psychology settled while Eisenhower was president.

The result predates the current AI wave by half a century, and the industry with the most to gain from it has only recently gotten around to naming the problem. Gartner published its first Magic Quadrant for Decision Intelligence Platforms in January 2026, and predicts that by 2027 half of business decisions will be AI-augmented or AI-automated. The name it has settled on is decision intelligence: the discipline of engineering decisions — not dashboards, not models, but the full path from information to action — using whatever combination of data, analytics, AI, and human judgment the decision actually warrants. Cassie Kozyrkov, who served as Google’s first Chief Decision Scientist, compresses it to “turning information into better action at any scale”; Lorien Pratt, who co-originated the framing and wrote its first book-length treatment (Link, 2019), draws it as a chain linking data to actions to outcomes.

That is the market answer. The scientific answer is older, stranger, and far more useful — and since my day job sits exactly at this seam, what follows is the evidence I think every builder of decision systems should be forced to read.

Where decision intelligence lives Five boxes in a chain: data, then analytics answering what happened, then prediction answering what will happen, then decision answering what to do, then outcome. A bracket under the decision box marks where decision intelligence lives — and where most analytics programs stop short. WHERE DECISION INTELLIGENCE LIVES data analytics what happened prediction what will happen decision what to do outcome decision intelligence lives here — and most analytics programs stop one box short
The chain every organization runs, whether it manages it or not. Instrumenting the first three boxes while leaving the fourth to vibes is the standard failure mode.

The seventy-year gap

If mechanical prediction has been better for seventy years, why isn’t everything decided by formula? Because two failure modes stand in the way, and they pull in opposite directions. Managing both is most of what decision intelligence amounts to.

Algorithm aversion: in Dietvorst, Simmons, and Massey’s 2015 experiments, people who watched an algorithm err lost confidence in it faster than they lost confidence in an erring human — and abandoned the algorithm even when it demonstrably outperformed. The 2018 follow-up in Management Science found the antidote almost embarrassingly cheap: let people modify the algorithm’s output, even slightly and within strict limits, and usage jumps — along with performance. A little control is enough to keep people engaged, not full command over the algorithm.

Automation bias: the opposite disease, documented in aviation before it reached AI. Parasuraman and Riley’s classic taxonomy named the axis — use, misuse, disuse — and Mosier, Skitka and colleagues’ cockpit studies gave the two signature errors their names: omission (missing what the automation fails to flag) and commission (following a wrong automated directive past visible contraindications). The clinical literature finds the same pattern in decision-support systems: better on average, plus a new class of errors from overreliance.

The reliance quadrant A two-by-two grid. Columns: the AI is right, the AI is wrong. Rows: the human relies, the human overrides. Relying when the AI is right is appropriate reliance. Relying when it is wrong is automation bias, a commission error. Overriding when the AI is right is algorithm aversion, or disuse. Overriding when it is wrong is appropriate skepticism. THE RELIANCE QUADRANT the AI is right the AI is wrong the human relies the human overrides appropriate reliance the quadrant you design for automation bias commission errors: following it off a cliff algorithm aversion disuse: firing the better forecaster appropriate skepticism the quadrant that justifies the human a system is well designed to the degree that behavior concentrates in the two outlined cells
Both failure modes are failures of calibration, in opposite directions. Neither is fixed by exhortation; both respond to design.

Putting the assumption to the test

Between those failure modes sits the assumption underneath most enterprise AI: humans plus AI must be better than either alone. It was finally put to a preregistered test. Vaccaro, Almaatouq, and Malone’s 2024 meta-analysis in Nature Human Behaviour pooled 106 studies and 370 effect sizes, and the headline finding deserves to be framed on the wall of every AI steering committee: on average, human–AI combinations performed worse than the best of the human or the AI alone.

What 370 effect sizes say A diverging bar chart of Hedges' g effect sizes from the Vaccaro, Almaatouq and Malone 2024 meta-analysis. Human-AI combinations versus the best of either alone: minus 0.23, with a confidence interval from minus 0.39 to minus 0.07. Combinations versus the human alone: plus 0.64. Decision tasks: minus 0.27. Creation tasks: plus 0.19. HUMAN + AI, MEASURED · HEDGES' g, VACCARO ET AL. 2024 vs best of either alone vs the human alone decision tasks creation tasks g = -0.23 (95% CI -0.39 to -0.07) g = +0.64 g = -0.27 g = +0.19 −0.23 +0.64 −0.27 +0.19 ← combination worse combination better → 106 studies, 370 effect sizes, preregistered; whisker = 95% CI on the headline estimate
Adding a human to a strong AI usually helped the human and hurt the system. The exception — open-ended creation tasks — is where the combination genuinely earned its keep.

Read the second bar carefully, because it explains the industry’s confusion: combinations beat the human alone handily (g = +0.64), which is what every vendor demo shows you. They lost to the best performer — often the AI alone — which is what the deployment eventually discovers. The moderator analysis carries the practical lesson: when the human was the stronger performer, adding the AI helped; when the AI was stronger, adding the human dragged it down. And the one task family where combinations reliably won was open-ended creation work, not selection among options.

Engineering the complement

None of this puts humans out of the loop — I have argued elsewhere in this series that in high-stakes domains the human checkpoint is mandatory. Complementarity just has to be built on purpose, because it does not happen by default. The evidence points to a short list of mechanisms that work:

Divide by comparative advantage, not by ceremony. The meta-analysis is blunt: a human reviewing every AI output adds error to the cases the AI had right. Route by confidence and stakes — the machine decides where it is demonstrably strong, the human owns the exceptions, the low-confidence cases, and everything with a signature. Bayesian modeling work in PNAS shows hybrids can beat both parties even at unequal accuracy — but only within bounds set by how correlated their errors are, and eliciting confidence, not just answers, is what makes the combination work.

Give the human a dial, not a veto. Dietvorst’s antidote — bounded modification of the model’s output — converts aversion into engagement at almost no accuracy cost.

Make accountability explicit. In the cockpit studies, pilots who felt personally accountable for outcomes double-checked the automation more and committed fewer errors. Diffuse accountability is the substrate automation bias grows in.

Train the judgment itself. The Good Judgment Project showed forecasting accuracy responds to teachable technique — debiasing training, reference-class thinking, teaming, and tracking top performers — which is worth remembering in an era that treats human judgment as a fixed quantity to be routed around. (A 2025 reanalysis questions how much of those gains survive tighter controls — in this field, even the evidence about judgment needs judgment.)

The discipline, not the platform

Gartner’s new Magic Quadrant will pull the term “decision intelligence” through every vendor deck in 2026, and the research history insists on a distinction those decks will blur. A decision intelligence platform is software. Decision intelligence itself is the discipline of taking Meehl seriously — formulas where formulas win, humans where humans win, and honest measurement of which is which — while designing around the two documented failure modes of the humans involved. Organizations that buy the platform without the discipline will automate their decisions and keep their biases. The ones that get it right will look less like “AI making decisions” and more like a well-run cockpit: machine flying, human accountable, both calibrated, and every landing graded.

References

  1. Vaccaro, Almaatouq & Malone (2024). DOI: 10.1038/s41562-024-02024-1. Effect sizes cited in this piece.
  2. Grove et al. (2000). PubMed. Clinical-vs-mechanical prediction statistics cited in this piece.