← Gautam Parab

The AI Weather Models Won. Look Closely at What They Won At.

In 1908 Felix Exner made a weather forecast that worked. Assuming geostrophic balance and constant thermal forcing, he derived a simple advection equation for the pressure pattern and got a realistic result for a four-hour period over the central United States on 3 January 1895. Fourteen years later Lewis Fry Richardson published the forecast everyone remembers instead: physically complete, hand-computed from the primitive equations, and wrong by a six-hour pressure change of 145 hPa. Richardson knew about Exner’s work. He gave it one sentence, on page 43 β€” Exner β€œhas published a prognostic method based on the source of air supply” β€” and Peter Lynch, in his history of the episode, supplies the dry gloss: β€œIt would appear from this that Richardson was not particularly impressed by it!”

The crude method that worked got a line. The physically complete method that failed got a book, and the century.

The machine-learning weather models have now run that asymmetry in reverse. They really do beat the physics. GraphCast beat the ECMWF high-resolution forecast on 90% of 1,380 verification targets; GenCast beat ECMWF’s 50-member ensemble on 97.4% of 1,320 targets, and on 99.8% of them beyond 36 hours. ECMWF now runs two data-driven systems operationally. These are not vendor press releases β€” they are Science and Nature papers, and the incumbent being beaten is the same organization that went on to deploy the challenger.

Published headline claims: share of verification targets on which the ML model beat the operational physics-based system Three published results on a shared zero to one hundred percent scale. GraphCast beat the ECMWF deterministic high-resolution forecast on 90 percent of 1380 verification targets, at 0.25 degree resolution out to 10 days. GenCast beat the ECMWF 50-member ensemble on 97.4 percent of 1320 targets across all lead times, scored by CRPS with 50 members on both sides. Restricted to lead times beyond 36 hours, GenCast beat the same ensemble on 99.8 percent of targets. A footnote records that a verification target is one variable at one pressure level at one lead time, so the count is dominated by upper-air fields on a smooth analysis grid. HEADLINE CLAIMS Β· AS PUBLISHED Β· SCIENCE 2023, NATURE 2024 GraphCast vs. IFS HRES GenCast vs. ENS, all leads GenCast vs. ENS, beyond 36 h 1,380 targets, to 10 days 1,320 targets, CRPS same 1,320 targets 90% of 1,380 targets β€” Lam et al., Science 2023 97.4% of 1,320 targets β€” Price et al., Nature 2024 99.8% of targets beyond 36 h β€” Price et al., Nature 2024 90% 97.4% 99.8% A "target" is one variable, one level, one lead time. The count is dominated by upper-air fields on a smooth grid.
The wins are large, published, and peer-reviewed. What a target is does most of the work.

So this time the crude method is winning, and it is being handed the century. The scoring rule those wins are measured with pays a premium for exactly the failure mode these models have; the models consume an atmospheric state produced by the system they are being compared against; and on ECMWF’s own scorecard the answer to β€œwhich one is better” changes depending on whether you check against the model’s own analysis or against actual thermometers. This is the same problem I keep running into with benchmarks generally: the number is real, and it is still measuring something narrower than the claim it gets used to support.

The score pays a premium for blur

Every one of these models is trained to minimize a squared-error-type loss, and most are scored with one. That combination has a known consequence that predates machine learning entirely. Minimizing mean-squared error against an uncertain future returns the conditional mean of the outcome distribution, not a sample from it β€” and the conditional mean of a chaotic system is smooth. The forecast that scores best is the forecast that stops committing to detail.

The verification literature has a name for the flip side. Elizabeth Ebert, in Neighborhood Verification (Weather and Forecasting, 2009), sets out the double-penalty problem: point-by-point scores β€œseverely penalize finescale differences that are not present in coarser-resolution forecasts,” so a sharp forecast that puts a rainband 40 km off gets charged twice β€” a miss where the rain was, a false alarm where it wasn’t β€” while a smeared forecast that commits to nothing gets charged once, softly. ECMWF states the same thing in its own science blog. A model trained on that objective learns the lesson.

Massimo Bonavita, at ECMWF, measured how far the lesson goes. In On some limitations of current machine learning weather prediction models (Geophysical Research Letters, 2024) he takes Pangu-Weather apart spectrally and finds that β€œthe effective resolution of Pangu-Weather forecasts is closer to 500-700 km than to the nominal 0.25 deg” β€” and that it degrades with lead time, most sharply in the first 24 hours. The IFS spectrum, by contrast, stays close to the ERA5 analysis out to about wavenumber 200. Pangu’s diverges from wavenumber 60–80.

Nominal grid spacing versus the scale at which forecast detail is actually real Three lengths on a shared zero to 800 kilometre scale. The nominal grid spacing shared by these models, 0.25 degrees, is about 31 kilometres. The ECMWF IFS deterministic forecast holds a realistic spectrum down to roughly 200 kilometres. Pangu-Weather's effective resolution is reported as 500 to 700 kilometres, and the bar is drawn at the conservative lower end of 500. A footnote records that the Pangu bar is drawn at the lower end of the reported range, and that its effective resolution degrades further with lead time. EFFECTIVE RESOLUTION Β· PANGU-WEATHER Β· BONAVITA, GRL 2024 Nominal grid, 0.25Β° IFS: spectrum stays faithful to Pangu: real detail stops near about 31 km β€” nominal 0.25 degree grid about 200 km β€” IFS deterministic, Bonavita 2024 500 to 700 km effective resolution β€” Pangu-Weather, Bonavita 2024 β‰ˆ31 km β‰ˆ200 km 500–700 km Pangu bar drawn at the conservative lower end of the reported range; its effective resolution degrades further with lead time.
The model is published at 31 km. It stops saying anything real somewhere around 500.

Bonavita’s conclusion is worth quoting exactly, because it is stronger than the usual hedging: these forecasts β€œdo not have the fidelity and physical consistency of physics-based models and their advantage in accuracy on traditional deterministic metrics of forecast skill can be at least partly attributed to these peculiarities.” The blur is not a tolerated side effect of winning. It is part of how the winning happens.

Where the blur becomes a different storm

Smoothing is harmless until the thing you need is an extreme. Two independent case studies land in the same place.

For Typhoon Doksuri, Bonavita compares a t+132h forecast: Pangu-Weather predicts a minimum central pressure of 986 hPa, a shallow low. The IFS predicts 957. IBTrACS records 944. For Storm CiarΓ‘n, Charlton-Perez and colleagues (npj Climate and Atmospheric Science, 2024) find all four ML models tested predicting maximum 10-metre winds of 25–26 m/s at 48 hours, against 34 m/s in the IFS analysis and 36 m/s in the IFS high-resolution forecast β€” and they specifically rule out the easy explanation, noting the weakness is not simply inherited from ERA5’s 31 km training resolution, because NWP models at comparable resolution do not show it.

Typhoon Doksuri minimum central pressure at a 132-hour lead time: how deep each system thought the storm was Three values shown as depth below a 1010 hectopascal reference, so a longer bar means a stronger storm. The observed best-track minimum central pressure from IBTrACS was 944 hectopascals, a depth of 66. The ECMWF IFS forecast at 132 hours was 957 hectopascals, a depth of 53. The Pangu-Weather forecast at the same lead time was 986 hectopascals, a depth of only 24 β€” about a third of the observed depth. A footnote notes that bars show depth below 1010 hectopascals and that a separate study found Pangu-Weather predicting a 23.9 metre per second maximum surface wind for the same storm against 48.3 observed by synthetic aperture radar. TYPHOON DOKSURI Β· T+132 H Β· BONAVITA 2024, IBTRACS v04r00 Observed (best track) ECMWF IFS forecast Pangu-Weather forecast 944 hPa β€” IBTrACS v04r00 best track 957 hPa β€” ECMWF IFS, t+132h 986 hPa β€” Pangu-Weather, t+132h 944 957 986 hPa Bars show depth below 1010 hPa β€” longer is a stronger storm. Single deterministic case, illustrative of a documented bias.
A 42 hPa error on a category-4 typhoon is not a rounding difference in a scorecard. It is a different storm.

The mechanism is not mysterious. Xu and colleagues (npj Climate and Atmospheric Science, 2025) find Pangu-Weather giving Doksuri a maximum surface wind of 23.9 m/s against 48.3 m/s from synthetic aperture radar, an intensity RMSE of 29.3 m/s β€” and then show that feeding those same Pangu large-scale fields into a 2 km WRF regional model with vortex initialization cuts the intensity RMSE to 5.1 m/s. The large-scale fields were fine. The model simply cannot put a real vortex on a grid that coarse, and the loss function never asked it to.

Change the yardstick, change the winner

The detail I found most clarifying comes from ECMWF’s own operational release note for AIFS ENS (51 members, approximately 30 km, now running alongside the 9 km IFS ensemble). ECMWF reports improvements reaching up to 25% for upper-air variables. It also reports, in the same paragraph, that β€œfor early lead times, AIFS ENS forecasts can appear less skilful than IFS ENS forecasts when verified against IFS analyses; however, this degradation of AIFS ENS compared to IFS ENS is not visible when SYNOP and radiosonde observations are used for verification.”

The ranking flips with the yardstick. And it flips both ways: against station observations, ECMWF reports IFS ENS still more skilful than AIFS ENS for 10-metre wind speed. This is the ERA5 circularity problem made concrete — these models are trained on a reanalysis, scored against analyses from the same lineage, and the reanalysis is itself the output of the assimilation system belonging to the incumbent. Ben-Bouallègue and colleagues (BAMS, 2024) did the honest version of this test, verifying Pangu against both operational analyses and synoptic observations, and found comparable skill under both — so the circularity is not fatal. It is just never free, and it is rarely stated.

One more ECMWF self-report worth keeping: AIFS ENS β€œis currently overdispersive for a range of upper-air variables.” The broader ML literature mostly worries about underdispersion. The direction of the calibration error is not even settled, which is a reasonable sign that this is early.

The half nobody replaced

The efficiency claims are the least disputed part of the story and the most misread. GraphCast produces a 10-day forecast in β€œless than a minute on a single Google TPU v4 machine”; GenCast takes about eight minutes on a TPU v5 for a 15-day ensemble member set; ECMWF claims AIFS generates forecasts over ten times faster while using roughly a thousandth of the energy. I have no reason to doubt any of it.

But an ML forecast starts from an analysis β€” a physically consistent estimate of the current atmosphere, produced by assimilating millions of observations. The ML models do not do this. They are handed it, by the NWP centre they are being benchmarked against. Peter Bauer, in What if? Numerical weather prediction at the crossroads, sizes both halves. At 9 km, each ENS member uses about 50 compute nodes, so β€œcompleting the forecasts in about one hour requires therefore at least 2,550 nodes.” And the assimilation side is not a fraction of that β€” it is the same size: β€œas the EDA runs its outer-loop trajectories at the same resolution as ENS, the total node allocation is about the same as for ENS.”

ECMWF HPC node allocation: data assimilation and ensemble forecast are about the same size Three quantities on a shared scale to 4000 nodes. The 51-member ensemble forecast requires at least 2550 compute nodes to finish in about an hour, at 9 kilometre resolution and about 50 nodes per member. The ensemble of data assimilations runs its outer-loop trajectories at the same resolution, and its total node allocation is about the same as the ensemble forecast, so roughly 2550 nodes. For scale, the two ECMWF operational clusters comprise 3840 nodes in total. A footnote records that machine-learning models replace the forecast bar and consume the output of the assimilation bar, and that the assimilation, forecast and boundary-condition suites together already claim 80 to 85 percent of available capability. ECMWF HPC ALLOCATION Β· BAUER, ARXIV:2407.03787 EDA (data assimilation) ENS (ensemble forecast) Both clusters, total capacity about the same allocation as ENS, so roughly 2,550 nodes β€” Bauer at least 2,550 nodes to complete in about one hour β€” Bauer 3,840 nodes across the two operational clusters β€” Bauer, Table 1 β‰ˆ2,550 β‰₯2,550 nodes 3,840 ML replaces the ENS bar and consumes the EDA bar's output. With the boundary-condition suite, the three claim 80–85%.
The two halves are the same size. The thousand-fold saving applies to one of them.

So the ledger is roughly two equal halves, and the machine-learning model replaces one of them. Bauer adds the context that makes this bite: the two operational clusters comprise 3,840 nodes, and the assimilation, forecast and boundary-condition suites together β€œalready claim 80-85% of the available capability.” The β€œ1,000Γ— less energy” claim is true of the part being replaced and silent about the part being inherited. That is a real and large saving. It is not the end of the supercomputer, and the people at ECMWF have never said it was β€” Bauer’s whole paper is an argument for using ML to relieve a machine that is running out of room.

The operational posture reflects this. NOAA’s National Hurricane Center describes integrating AI systems β€œas guidance when preparing operational forecasts, alongside all of the other critical tools in our toolbox,” while noting that verification over many storms shows β€œthe official forecast is the most skillful and consistent, surpassing any individual model forecast.” Nobody is issuing warnings off a neural network alone.

Richardson’s actual mistake

Which brings me back to 1922. Richardson’s instantaneous pressure tendency was about 0.7 Pa/s, which is not absurd at all β€” Loehrer and Johnson later measured βˆ’1.3 Pa/s in a real mesoscale convective system. The catastrophe came from multiplying a noisy instantaneous value, contaminated by unfiltered gravity waves, by a six-hour timestep. His physics was fine. What he had was a number that meant something much narrower than the use he put it to.

That is the failure mode I would watch for now. Three things are true at once, and the press coverage tends to carry only the first. The data-driven models beat the physics on published scores. Those scores reward the smoothing that makes them dangerous precisely where forecasts matter most β€” the deep low, the intensifying cyclone. And they are one half of a pipeline whose other half is still a supercomputer solving equations, which is also the half that decides what β€œthe current state of the atmosphere” even means.

That is not a story about AI replacing simulation. It is closer to what I argued about digital twins: the learned model earns its place by being fast enough to run a thousand times, on top of a physical model that stays responsible for being right. I would not want to forecast without these models. But if you are reading a forecast scorecard, ask what the verification target was, and whether anyone checked against a thermometer.


References

  1. Lam et al. (2023). Science. arXiv:2212.12794.
  2. Bi et al. (2023). Nature. arXiv:2211.02556.
  3. Pathak et al. arXiv:2202.11214.
  4. Price et al. (2024). Nature. arXiv:2312.15796.
  5. Bonavita, M. (2024). On some limitations of current machine learning weather prediction models. Geophysical Research Letters.
  6. Charlton-Perez et al. (2024). npj Climate and Atmospheric Science.
  7. Xu et al. (2025). npj Climate and Atmospheric Science.
  8. Ben-Bouallègue et al. (2024). BAMS. arXiv:2307.10128.
  9. Bauer, P. What if? Numerical weather prediction at the crossroads. arXiv:2407.03787.
  10. Ebert, E. (2009). Neighborhood Verification. Weather and Forecasting.
  11. Bracco et al. (2024). Nature Reviews Physics.
  12. Eyring et al. (2024). Nature Climate Change.
  13. Emanuel. (2023). Nature Climate Change.
  14. ECMWF Newsletter 183.
  15. ECMWF. Newsletter 185: AIFS ENS becomes operational.
  16. NOAA/NHC. AI Hurricane Forecasting. weather.gov.
  17. Lynch, P. History of the episode, from Weather Prediction by Numerical Process, ch. 1.