← Gautam Parab

The Same Robotaxi Data Says 4x Worse and 79% Better. The Gap Is the Denominator.

In 2024, two analyses looked at the same fleet, drawing on the same federal crash reports, and reached opposite conclusions. One found Waymo’s vehicles involved in property-damage crashes at roughly four times the rate of human drivers. The other found large, statistically significant reductions. Neither made an arithmetic mistake. They compared the robotaxi’s numerator to different human denominators, and in rare-event safety statistics the denominator is most of the answer.

That gap explains why you can read two credible-sounding accounts of robotaxi safety in the same week and come away with opposite impressions. The next round of numbers will have the same problem, so it helps to know where to look.

The flip

Waymo’s own published comparison of its first 7.14 million rider-only miles reports 2.1 police-reported crashes per million miles against a matched human benchmark of 4.68, and 0.6 any-injury-reported crashes against 2.80 (Kusano et al., Traffic Injury Prevention, 2024). The competing framing takes Waymo’s San Francisco crashes reported under NHTSA’s Standing General Order at any property damage or injury, 16.5 per million miles, and sets them beside a national police-reported benchmark of 4.10. Same vehicles. One comparison says the machine is four times worse.

The same Waymo fleet compared against three different human benchmarks Dumbbell chart of incidents per million miles. Comparing Waymo's any-property-damage-or-injury crashes reported under the Standing General Order (16.5 per million miles) against an unmatched national police-reported benchmark (4.10) makes the vehicle look about four times worse. Comparing police-reported crashes against a matched benchmark gives 2.1 for Waymo versus 4.68 for humans, about 55 percent lower. Comparing any-injury-reported crashes gives 0.6 versus 2.80, about 79 percent lower. INCIDENTS PER MILLION MILES Β· WAYMO RIDER-ONLY, 7.14M MI Β· KUSANO ET AL. 2024 Waymo human benchmark Any property damage or injury Police-reported Any injury reported unmatched benchmark → 4× worse matched benchmark → 55% lower matched benchmark → 79% lower 16.5 incidents per million miles β€” Waymo SGO any property damage or injury, San Francisco 2.1 incidents per million miles β€” Waymo police-reported 0.6 incidents per million miles β€” Waymo any-injury-reported 4.10 incidents per million miles β€” national police-reported benchmark, unmatched 4.68 incidents per million miles β€” matched human benchmark 2.80 incidents per million miles β€” matched human benchmark 16.5 4.10 2.1 4.68 0.6 2.80 Top row pairs a far lower reporting threshold against a police-reported one; the lower two match thresholds.
One fleet, three framings. Only the mismatched top row makes the vehicle look worse than a human driver.

The top row of that chart is comparing a smoke detector to a fire department. The Standing General Order’s triggers are broad: property damage, a tow-away, an airbag deployment, a vulnerable road user struck, a hospital transport. The operator applies them with sensors that miss nothing (Third Amended SGO 2021-01, effective June 2025). Roughly half the in-transport collisions Waymo reported in the 7.14-million-mile study involved a change in velocity under one mile per hour. Human crash databases, by contrast, have floors: $2,000 of damage in Arizona, $1,000 in California, tow-away in Pennsylvania. Below those floors most damage never enters a database at all. The machine reports the parking-lot scuff. The human does not.

The human baseline is a construction, not a measurement

There is also no clean number for how often human drivers crash. Police reports miss roughly 60% of property-damage-only crashes and about 32% of non-fatal injury crashes, with the undercount shrinking as severity rises (Scanlon et al., 2024). Naturalistic driving studies, which watch drivers with cameras rather than waiting for a police form, put property-damage underreporting higher still.

Then there is a unit mismatch that almost nobody notices, and it is my favorite artifact in this whole literature. Crash statistics come in two flavors: crash-level (a two-car collision counts once) and vehicle-level (it counts twice). Robotaxi data are vehicle-level by construction, since the operator reports its own vehicle. National human data are often quoted crash-level. In 2022 the United States recorded about 5.9 million police-reported crashes involving roughly 10.5 million crashed vehicles across 3.2 trillion vehicle-miles. That is 1.86 crashes per million miles, or 3.29 crashed vehicles per million miles: a 78% difference produced entirely by which noun you count. Published comparisons have used the smaller number as the human benchmark, which flatters human driving and understates the machine, and the error survives peer review because both figures are correct arithmetic on the same data.

Geography does similar work. San Francisco’s surface-street injury crash rate for passenger vehicles ran about 330% above the national average in 2022, at 5.82 per million miles against 1.76 nationally; Maricopa County, at 1.80, sits essentially on the national figure (Scanlon et al., 2024). Robotaxis operate almost entirely on dense urban surface streets, at speeds generally at or below 50 mph, excluding freeways and severe weather. Compare that fleet to a national all-roads, all-vehicles average that includes motorcycles, heavy trucks and interstate miles, and you have not measured anything about autonomy.

The mileage nobody has

Even with a perfectly matched benchmark, there is a hard statistical floor. Kalra and Paddock’s 2016 RAND analysis is still the canonical treatment (Driving to Safety): against a human rate of 1.09 fatalities per 100 million miles, demonstrating with 95% confidence that a system is no worse than human requires 275 million failure-free miles. Estimating the fatality rate to within 20% requires about 8.8 billion. Showing the system is 20% better than human, with 80% power, needs around 11.3 billion. A 5% improvement needs roughly 215 billion. Their conclusion was that developers cannot drive their way to safety, and a decade of additional data has not overturned it.

Miles required for fatality-rate claims versus miles actually driven Logarithmic dot plot. Demonstrating a fatality rate no worse than human, failure-free, requires 275 million miles. Estimating the fatality rate to within 20 percent requires 8.8 billion miles. Showing the system is 20 percent better than human requires 11.3 billion miles, and 5 percent better requires 215 billion miles. Waymo's cumulative rider-only mileage of 220.6 million miles as of March 2026 falls short of even the first threshold. MILES REQUIRED, LOG SCALE Β· KALRA & PADDOCK, RAND 2016 100M 1B 10B 100B Fatality rate no worse than human Fatality rate estimated to ±20% Prove 20% better than human Prove 5% better than human 275 million failure-free miles β€” RAND, 2016 8.8 billion miles β€” RAND, 2016 11.3 billion miles β€” RAND, 2016 215 billion miles β€” RAND, 2016 275M 8.8B 11.3B 215B Dashed line: 220.6M rider-only miles, Waymo's cumulative total as of March 2026.
Even the least demanding fatality claim sits beyond the miles anyone has driven; the rest sit orders of magnitude beyond.

Waymo reports 220.6 million rider-only miles as of March 2026 (Waymo safety impact), the largest such figure anyone has, and still short of the failure-free threshold on the least demanding fatality question. This is why serious analyses stop at injury-level outcomes, where events are common enough to count. At 56.7 million miles, Waymo’s peer-reviewed comparison found a 79% reduction in any-injury-reported crashes (95% CI 71–85%) and 81% fewer airbag deployments (CI 69–90%) (Kusano et al., 2025). A Swiss Re study of insurance claims across 25.3 million miles found 88% fewer property damage claims and 92% fewer bodily injury claims against a baseline of over 500,000 claims and 200 billion miles of exposure (Waymo, December 2024). Those are company-published or company-co-authored figures and should be read as such, but the methodology is in the open literature and the injury-level results survive the scrutiny that the property-damage ones do not.

Where the averages hide the structure

A matched case-control study of California crashes found automated vehicles over-represented in only two conditions: dawn and dusk, at 5.25 times the human odds, and turning maneuvers, at 1.98. They were substantially under-represented in rain, rear-end, broadside and run-off-road crashes (Abdel-Aty & Ding, Nature Communications, 2024).

Automated-vehicle crash involvement odds by scenario, relative to human drivers Logarithmic dot plot of odds ratios from a matched case-control study. Automated vehicles show odds ratios above parity at dawn or dusk (5.25) and while turning (1.98), and below parity for rear-end crashes (0.46), rain (0.34), broadside crashes (0.17) and run-off-road crashes (0.02). A vertical line marks parity at 1.0. ODDS RATIO VS. HUMAN-DRIVEN, LOG SCALE Β· ABDEL-ATY & DING 2024 0.01 0.1 1.0 10 Dawn / dusk Turning maneuver Rear-end Rain Broadside Run-off-road Odds ratio 5.25 β€” dawn/dusk Odds ratio 1.98 β€” turning maneuver Odds ratio 0.46 β€” rear-end Odds ratio 0.34 β€” rain Odds ratio 0.17 β€” broadside Odds ratio 0.02 β€” run-off-road 5.25× 1.98× 0.46× 0.34× 0.17× 0.02× Solid line: parity with human drivers. Right of it, automated vehicles are over-represented.
A single "safer than humans" ratio averages away a fivefold disadvantage at dusk and a fiftyfold advantage off-road.

Automated vehicles do get rear-ended more often than the national average, which is a favorite talking point. But naturalistic data show human struck-from-behind rates roughly double in urban and business-industrial environments, exactly where these fleets operate. Control for environment and most of the gap closes; what remains concentrates in stopped vehicles, at 8.9 per million miles against 1.8 for humans in comparable settings (Goodall, Accident Analysis & Prevention, 2021). The residual gap is about where and how long these vehicles choose to stop rather than how well they drive, which makes it a design decision and a fixable one.

The baseline everyone quotes was withdrawn

The statistic underwriting the whole enterprise, that 94% of crashes are caused by human error, comes from NHTSA’s 2015 report on the National Motor Vehicle Crash Causation Survey, built on 2005–2007 data (DOT HS 812 115). The report assigns a β€œcritical reason” to the driver in 94% of crashes, plus or minus 2.2 points. It also says, in the same document, that the critical reason is the last failure in a causal chain and β€œis not intended to be interpreted as the cause of the crash nor as the assignment of the fault to the driver, vehicle, or environment.”

In January 2022 NTSB chair Jennifer Homendy told the Associated Press that the statistic was being used wrongly, and that it β€œain’t 94 percent.” NHTSA removed the figure from its materials within days, as reported at the time. The number is still everywhere in industry decks. The agency that produced it stopped publishing it four years ago.

What would actually settle it

The older proxy metrics are being retired for good reason. California’s disengagement reports, which count the times a safety driver takes over, were never comparable across companies, since each defines its own trigger and test strategy, a critique the literature made as early as 2018. The DMV now says plainly that the reports are β€œnot designed for comparative analysis across companies,” and in February 2026 proposed replacing the requirement with metrics it says would give β€œa more precise picture of safety-relevant events arising during AV operation” (CA DMV, February 2026); permit holders logged over 9 million test miles in that reporting year. Whether the rulemaking has been finalized, I could not confirm.

The Standing General Order is a better instrument but a partial one. It collects crashes and no exposure at all, no miles, so no rate can be computed from it without merging in mileage counted on some other basis. It records self-reported narratives with no fault determination and few police reports, and its thirty-second automation window excludes crashes where the system disengaged earlier. It is a census of numerators (Cummings & Bauchwitz, IEEE Transactions on Intelligent Vehicles, 2025). It is also the instrument that caught Cruise: NHTSA issued a $1.5 million consent order in 2024 for failing to fully report the October 2023 pedestrian dragging, and the Justice Department secured a $500,000 criminal fine after Cruise admitted submitting a false report. Mandatory reporting has value independent of whether you can compute a rate from it.

So when the next headline arrives, ask narrow questions. What outcome level: property damage, injury, or fatality? What benchmark, matched on geography, road type and vehicle type, or a national average? Crash-level or vehicle-level counting? How many actual events underlie the percentage, and does the confidence interval clear parity? At fatality level the honest answer today is that nobody has the miles, and anyone claiming otherwise is extrapolating.

There is a last complication that no amount of methodology fixes. Even a perfectly demonstrated β€œas safe as the average driver” would not satisfy most people, because about 81% of drivers rate themselves above the median and want autonomous vehicles safer than that inflated self-assessment, typically somewhere in the 95th to 99th percentile of human performance (Nees, Journal of Safety Research, 2019). The statistical bar and the social one are different bars, and the second is higher.

This is the same shape of problem I keep running into elsewhere: in how progress toward AGI gets measured, and in the gap between two scores on the same benchmark. The metric is real, the measurement is real, and the number still does not mean what the headline says it means. With robotaxis the stakes are simply more physical.


References

  1. Kalra & Paddock. (2016). Driving to Safety. RAND.
  2. Kusano et al. (2024). Traffic Injury Prevention.
  3. Kusano et al. (2025). Traffic Injury Prevention.
  4. Scanlon et al. (2024). Traffic Injury Prevention.
  5. Abdel-Aty & Ding. (2024). Nature Communications.
  6. Goodall. (2021). Accident Analysis & Prevention.
  7. Nees. (2019). Journal of Safety Research.
  8. FavarΓ² et al. (2018). Accident Analysis & Prevention.
  9. Cummings & Bauchwitz. (2025). IEEE Transactions on Intelligent Vehicles.
  10. NHTSA. (2015). DOT HS 812 115.
  11. NHTSA. Third Amended Standing General Order 2021-01, effective June 2025.
  12. Waymo. Safety impact hub.
  13. California DMV. (2026, February). Autonomous vehicle permit holder mileage report.