In 2024, two analyses looked at the same fleet, drawing on the same federal crash reports, and reached opposite conclusions. One found Waymoβs vehicles involved in property-damage crashes at roughly four times the rate of human drivers. The other found large, statistically significant reductions. Neither made an arithmetic mistake. They compared the robotaxiβs numerator to different human denominators, and in rare-event safety statistics the denominator is most of the answer.
That gap explains why you can read two credible-sounding accounts of robotaxi safety in the same week and come away with opposite impressions. The next round of numbers will have the same problem, so it helps to know where to look.
Waymoβs own published comparison of its first 7.14 million rider-only miles reports 2.1 police-reported crashes per million miles against a matched human benchmark of 4.68, and 0.6 any-injury-reported crashes against 2.80 (Kusano et al., Traffic Injury Prevention, 2024). The competing framing takes Waymoβs San Francisco crashes reported under NHTSAβs Standing General Order at any property damage or injury, 16.5 per million miles, and sets them beside a national police-reported benchmark of 4.10. Same vehicles. One comparison says the machine is four times worse.
The top row of that chart is comparing a smoke detector to a fire department. The Standing General Orderβs triggers are broad: property damage, a tow-away, an airbag deployment, a vulnerable road user struck, a hospital transport. The operator applies them with sensors that miss nothing (Third Amended SGO 2021-01, effective June 2025). Roughly half the in-transport collisions Waymo reported in the 7.14-million-mile study involved a change in velocity under one mile per hour. Human crash databases, by contrast, have floors: $2,000 of damage in Arizona, $1,000 in California, tow-away in Pennsylvania. Below those floors most damage never enters a database at all. The machine reports the parking-lot scuff. The human does not.
There is also no clean number for how often human drivers crash. Police reports miss roughly 60% of property-damage-only crashes and about 32% of non-fatal injury crashes, with the undercount shrinking as severity rises (Scanlon et al., 2024). Naturalistic driving studies, which watch drivers with cameras rather than waiting for a police form, put property-damage underreporting higher still.
Then there is a unit mismatch that almost nobody notices, and it is my favorite artifact in this whole literature. Crash statistics come in two flavors: crash-level (a two-car collision counts once) and vehicle-level (it counts twice). Robotaxi data are vehicle-level by construction, since the operator reports its own vehicle. National human data are often quoted crash-level. In 2022 the United States recorded about 5.9 million police-reported crashes involving roughly 10.5 million crashed vehicles across 3.2 trillion vehicle-miles. That is 1.86 crashes per million miles, or 3.29 crashed vehicles per million miles: a 78% difference produced entirely by which noun you count. Published comparisons have used the smaller number as the human benchmark, which flatters human driving and understates the machine, and the error survives peer review because both figures are correct arithmetic on the same data.
Geography does similar work. San Franciscoβs surface-street injury crash rate for passenger vehicles ran about 330% above the national average in 2022, at 5.82 per million miles against 1.76 nationally; Maricopa County, at 1.80, sits essentially on the national figure (Scanlon et al., 2024). Robotaxis operate almost entirely on dense urban surface streets, at speeds generally at or below 50 mph, excluding freeways and severe weather. Compare that fleet to a national all-roads, all-vehicles average that includes motorcycles, heavy trucks and interstate miles, and you have not measured anything about autonomy.
Even with a perfectly matched benchmark, there is a hard statistical floor. Kalra and Paddockβs 2016 RAND analysis is still the canonical treatment (Driving to Safety): against a human rate of 1.09 fatalities per 100 million miles, demonstrating with 95% confidence that a system is no worse than human requires 275 million failure-free miles. Estimating the fatality rate to within 20% requires about 8.8 billion. Showing the system is 20% better than human, with 80% power, needs around 11.3 billion. A 5% improvement needs roughly 215 billion. Their conclusion was that developers cannot drive their way to safety, and a decade of additional data has not overturned it.
Waymo reports 220.6 million rider-only miles as of March 2026 (Waymo safety impact), the largest such figure anyone has, and still short of the failure-free threshold on the least demanding fatality question. This is why serious analyses stop at injury-level outcomes, where events are common enough to count. At 56.7 million miles, Waymoβs peer-reviewed comparison found a 79% reduction in any-injury-reported crashes (95% CI 71β85%) and 81% fewer airbag deployments (CI 69β90%) (Kusano et al., 2025). A Swiss Re study of insurance claims across 25.3 million miles found 88% fewer property damage claims and 92% fewer bodily injury claims against a baseline of over 500,000 claims and 200 billion miles of exposure (Waymo, December 2024). Those are company-published or company-co-authored figures and should be read as such, but the methodology is in the open literature and the injury-level results survive the scrutiny that the property-damage ones do not.
A matched case-control study of California crashes found automated vehicles over-represented in only two conditions: dawn and dusk, at 5.25 times the human odds, and turning maneuvers, at 1.98. They were substantially under-represented in rain, rear-end, broadside and run-off-road crashes (Abdel-Aty & Ding, Nature Communications, 2024).
Automated vehicles do get rear-ended more often than the national average, which is a favorite talking point. But naturalistic data show human struck-from-behind rates roughly double in urban and business-industrial environments, exactly where these fleets operate. Control for environment and most of the gap closes; what remains concentrates in stopped vehicles, at 8.9 per million miles against 1.8 for humans in comparable settings (Goodall, Accident Analysis & Prevention, 2021). The residual gap is about where and how long these vehicles choose to stop rather than how well they drive, which makes it a design decision and a fixable one.
The statistic underwriting the whole enterprise, that 94% of crashes are caused by human error, comes from NHTSAβs 2015 report on the National Motor Vehicle Crash Causation Survey, built on 2005β2007 data (DOT HS 812 115). The report assigns a βcritical reasonβ to the driver in 94% of crashes, plus or minus 2.2 points. It also says, in the same document, that the critical reason is the last failure in a causal chain and βis not intended to be interpreted as the cause of the crash nor as the assignment of the fault to the driver, vehicle, or environment.β
In January 2022 NTSB chair Jennifer Homendy told the Associated Press that the statistic was being used wrongly, and that it βainβt 94 percent.β NHTSA removed the figure from its materials within days, as reported at the time. The number is still everywhere in industry decks. The agency that produced it stopped publishing it four years ago.
The older proxy metrics are being retired for good reason. Californiaβs disengagement reports, which count the times a safety driver takes over, were never comparable across companies, since each defines its own trigger and test strategy, a critique the literature made as early as 2018. The DMV now says plainly that the reports are βnot designed for comparative analysis across companies,β and in February 2026 proposed replacing the requirement with metrics it says would give βa more precise picture of safety-relevant events arising during AV operationβ (CA DMV, February 2026); permit holders logged over 9 million test miles in that reporting year. Whether the rulemaking has been finalized, I could not confirm.
The Standing General Order is a better instrument but a partial one. It collects crashes and no exposure at all, no miles, so no rate can be computed from it without merging in mileage counted on some other basis. It records self-reported narratives with no fault determination and few police reports, and its thirty-second automation window excludes crashes where the system disengaged earlier. It is a census of numerators (Cummings & Bauchwitz, IEEE Transactions on Intelligent Vehicles, 2025). It is also the instrument that caught Cruise: NHTSA issued a $1.5 million consent order in 2024 for failing to fully report the October 2023 pedestrian dragging, and the Justice Department secured a $500,000 criminal fine after Cruise admitted submitting a false report. Mandatory reporting has value independent of whether you can compute a rate from it.
So when the next headline arrives, ask narrow questions. What outcome level: property damage, injury, or fatality? What benchmark, matched on geography, road type and vehicle type, or a national average? Crash-level or vehicle-level counting? How many actual events underlie the percentage, and does the confidence interval clear parity? At fatality level the honest answer today is that nobody has the miles, and anyone claiming otherwise is extrapolating.
There is a last complication that no amount of methodology fixes. Even a perfectly demonstrated βas safe as the average driverβ would not satisfy most people, because about 81% of drivers rate themselves above the median and want autonomous vehicles safer than that inflated self-assessment, typically somewhere in the 95th to 99th percentile of human performance (Nees, Journal of Safety Research, 2019). The statistical bar and the social one are different bars, and the second is higher.
This is the same shape of problem I keep running into elsewhere: in how progress toward AGI gets measured, and in the gap between two scores on the same benchmark. The metric is real, the measurement is real, and the number still does not mean what the headline says it means. With robotaxis the stakes are simply more physical.
References