In 1908 Felix Exner made a weather forecast that worked. Assuming geostrophic balance and constant thermal forcing, he derived a simple advection equation for the pressure pattern and got a realistic result for a four-hour period over the central United States on 3 January 1895. Fourteen years later Lewis Fry Richardson published the forecast everyone remembers instead: physically complete, hand-computed from the primitive equations, and wrong by a six-hour pressure change of 145 hPa. Richardson knew about Exnerβs work. He gave it one sentence, on page 43 β Exner βhas published a prognostic method based on the source of air supplyβ β and Peter Lynch, in his history of the episode, supplies the dry gloss: βIt would appear from this that Richardson was not particularly impressed by it!β
The crude method that worked got a line. The physically complete method that failed got a book, and the century.
The machine-learning weather models have now run that asymmetry in reverse. They really do beat the physics. GraphCast beat the ECMWF high-resolution forecast on 90% of 1,380 verification targets; GenCast beat ECMWFβs 50-member ensemble on 97.4% of 1,320 targets, and on 99.8% of them beyond 36 hours. ECMWF now runs two data-driven systems operationally. These are not vendor press releases β they are Science and Nature papers, and the incumbent being beaten is the same organization that went on to deploy the challenger.
So this time the crude method is winning, and it is being handed the century. The scoring rule those wins are measured with pays a premium for exactly the failure mode these models have; the models consume an atmospheric state produced by the system they are being compared against; and on ECMWFβs own scorecard the answer to βwhich one is betterβ changes depending on whether you check against the modelβs own analysis or against actual thermometers. This is the same problem I keep running into with benchmarks generally: the number is real, and it is still measuring something narrower than the claim it gets used to support.
Every one of these models is trained to minimize a squared-error-type loss, and most are scored with one. That combination has a known consequence that predates machine learning entirely. Minimizing mean-squared error against an uncertain future returns the conditional mean of the outcome distribution, not a sample from it β and the conditional mean of a chaotic system is smooth. The forecast that scores best is the forecast that stops committing to detail.
The verification literature has a name for the flip side. Elizabeth Ebert, in Neighborhood Verification (Weather and Forecasting, 2009), sets out the double-penalty problem: point-by-point scores βseverely penalize finescale differences that are not present in coarser-resolution forecasts,β so a sharp forecast that puts a rainband 40 km off gets charged twice β a miss where the rain was, a false alarm where it wasnβt β while a smeared forecast that commits to nothing gets charged once, softly. ECMWF states the same thing in its own science blog. A model trained on that objective learns the lesson.
Massimo Bonavita, at ECMWF, measured how far the lesson goes. In On some limitations of current machine learning weather prediction models (Geophysical Research Letters, 2024) he takes Pangu-Weather apart spectrally and finds that βthe effective resolution of Pangu-Weather forecasts is closer to 500-700 km than to the nominal 0.25 degβ β and that it degrades with lead time, most sharply in the first 24 hours. The IFS spectrum, by contrast, stays close to the ERA5 analysis out to about wavenumber 200. Panguβs diverges from wavenumber 60β80.
Bonavitaβs conclusion is worth quoting exactly, because it is stronger than the usual hedging: these forecasts βdo not have the fidelity and physical consistency of physics-based models and their advantage in accuracy on traditional deterministic metrics of forecast skill can be at least partly attributed to these peculiarities.β The blur is not a tolerated side effect of winning. It is part of how the winning happens.
Smoothing is harmless until the thing you need is an extreme. Two independent case studies land in the same place.
For Typhoon Doksuri, Bonavita compares a t+132h forecast: Pangu-Weather predicts a minimum central pressure of 986 hPa, a shallow low. The IFS predicts 957. IBTrACS records 944. For Storm CiarΓ‘n, Charlton-Perez and colleagues (npj Climate and Atmospheric Science, 2024) find all four ML models tested predicting maximum 10-metre winds of 25β26 m/s at 48 hours, against 34 m/s in the IFS analysis and 36 m/s in the IFS high-resolution forecast β and they specifically rule out the easy explanation, noting the weakness is not simply inherited from ERA5βs 31 km training resolution, because NWP models at comparable resolution do not show it.
The mechanism is not mysterious. Xu and colleagues (npj Climate and Atmospheric Science, 2025) find Pangu-Weather giving Doksuri a maximum surface wind of 23.9 m/s against 48.3 m/s from synthetic aperture radar, an intensity RMSE of 29.3 m/s β and then show that feeding those same Pangu large-scale fields into a 2 km WRF regional model with vortex initialization cuts the intensity RMSE to 5.1 m/s. The large-scale fields were fine. The model simply cannot put a real vortex on a grid that coarse, and the loss function never asked it to.
The detail I found most clarifying comes from ECMWFβs own operational release note for AIFS ENS (51 members, approximately 30 km, now running alongside the 9 km IFS ensemble). ECMWF reports improvements reaching up to 25% for upper-air variables. It also reports, in the same paragraph, that βfor early lead times, AIFS ENS forecasts can appear less skilful than IFS ENS forecasts when verified against IFS analyses; however, this degradation of AIFS ENS compared to IFS ENS is not visible when SYNOP and radiosonde observations are used for verification.β
The ranking flips with the yardstick. And it flips both ways: against station observations, ECMWF reports IFS ENS still more skilful than AIFS ENS for 10-metre wind speed. This is the ERA5 circularity problem made concrete β these models are trained on a reanalysis, scored against analyses from the same lineage, and the reanalysis is itself the output of the assimilation system belonging to the incumbent. Ben-BouallΓ¨gue and colleagues (BAMS, 2024) did the honest version of this test, verifying Pangu against both operational analyses and synoptic observations, and found comparable skill under both β so the circularity is not fatal. It is just never free, and it is rarely stated.
One more ECMWF self-report worth keeping: AIFS ENS βis currently overdispersive for a range of upper-air variables.β The broader ML literature mostly worries about underdispersion. The direction of the calibration error is not even settled, which is a reasonable sign that this is early.
The efficiency claims are the least disputed part of the story and the most misread. GraphCast produces a 10-day forecast in βless than a minute on a single Google TPU v4 machineβ; GenCast takes about eight minutes on a TPU v5 for a 15-day ensemble member set; ECMWF claims AIFS generates forecasts over ten times faster while using roughly a thousandth of the energy. I have no reason to doubt any of it.
But an ML forecast starts from an analysis β a physically consistent estimate of the current atmosphere, produced by assimilating millions of observations. The ML models do not do this. They are handed it, by the NWP centre they are being benchmarked against. Peter Bauer, in What if? Numerical weather prediction at the crossroads, sizes both halves. At 9 km, each ENS member uses about 50 compute nodes, so βcompleting the forecasts in about one hour requires therefore at least 2,550 nodes.β And the assimilation side is not a fraction of that β it is the same size: βas the EDA runs its outer-loop trajectories at the same resolution as ENS, the total node allocation is about the same as for ENS.β
So the ledger is roughly two equal halves, and the machine-learning model replaces one of them. Bauer adds the context that makes this bite: the two operational clusters comprise 3,840 nodes, and the assimilation, forecast and boundary-condition suites together βalready claim 80-85% of the available capability.β The β1,000Γ less energyβ claim is true of the part being replaced and silent about the part being inherited. That is a real and large saving. It is not the end of the supercomputer, and the people at ECMWF have never said it was β Bauerβs whole paper is an argument for using ML to relieve a machine that is running out of room.
The operational posture reflects this. NOAAβs National Hurricane Center describes integrating AI systems βas guidance when preparing operational forecasts, alongside all of the other critical tools in our toolbox,β while noting that verification over many storms shows βthe official forecast is the most skillful and consistent, surpassing any individual model forecast.β Nobody is issuing warnings off a neural network alone.
Which brings me back to 1922. Richardsonβs instantaneous pressure tendency was about 0.7 Pa/s, which is not absurd at all β Loehrer and Johnson later measured β1.3 Pa/s in a real mesoscale convective system. The catastrophe came from multiplying a noisy instantaneous value, contaminated by unfiltered gravity waves, by a six-hour timestep. His physics was fine. What he had was a number that meant something much narrower than the use he put it to.
That is the failure mode I would watch for now. Three things are true at once, and the press coverage tends to carry only the first. The data-driven models beat the physics on published scores. Those scores reward the smoothing that makes them dangerous precisely where forecasts matter most β the deep low, the intensifying cyclone. And they are one half of a pipeline whose other half is still a supercomputer solving equations, which is also the half that decides what βthe current state of the atmosphereβ even means.
That is not a story about AI replacing simulation. It is closer to what I argued about digital twins: the learned model earns its place by being fast enough to run a thousand times, on top of a physical model that stays responsible for being right. I would not want to forecast without these models. But if you are reading a forecast scorecard, ask what the verification target was, and whether anyone checked against a thermometer.
References