NHC verification finds Google DeepMind’s GDMI leading individual hurricane guidance in 2025
During its first season in the National Hurricane Center’s live forecasting workflow, Google DeepMind’s GDMI ensemble mean slightly outperformed the official Atlantic track forecast at 12–72 hours and beat the official forecast and all tested consensus aids in eastern North Pacific track forecasts at 48–120 hours. Its Atlantic intensity skill was comparable to the official forecast.
The Reality Check
Context
The 2025 season was the first in which NHC incorporated AI-based model output into real-time operations. Its post-season verification compared available guidance under homogeneous samples, so every model was evaluated only on cases where the required guidance was available.
In the Atlantic, GDMI had lower mean track errors than the official NHC forecast at 12–72 hours, while its intensity skill was described as comparable to the official forecast. In the eastern North Pacific, GDMI performed better than the other individual models and beat both the official forecast and every tested consensus aid from 48–120 hours.
These results varied by basin, lead time, and metric. NHC’s official forecast remained essential, and the final Melissa verification found that the human intensity forecast outperformed the models at nearly every lead time. The season therefore demonstrates operationally useful AI guidance, not replacement of forecasters or universally superior performance.
THE TAKEAWAY
GDMI crossed a meaningful operational threshold: an AI model became useful inside a high-stakes real-time hurricane workflow for both track and intensity, and achieved leading individual-model performance under formal NHC verification. The result remains conditional on basin, forecast range, availability, and the limited evidence of one season.
Continue the Thread
AI in Operational Weather ForecastingTracks the use and verified performance of AI systems inside real-world weather-forecasting workflows.
Sources
Forecast Verification Report: 2025 Hurricane Season
National Hurricane Center
Used for: Final post-season verification methodology; Atlantic and eastern North Pacific model comparisons; GDMI’s track and intensity performance; homogeneous sample boundaries; model availability limitations; official 2025 season findings. The report was published on 30 March 2026.
2025 NHC Verification Report Preview
National Hurricane Center
Used for: Confirmation that 2025 was NHC’s first season incorporating AI models in real time; preliminary rapid-intensification evaluation; Melissa forecast verification; statement that GDMI was useful but experimental systems were not always available on time.
Category 5 Melissa makes landfall in Jamaica
National Hurricane Center
Used for: Official observed landfall outcome: Category 5 intensity, estimated 185 mph sustained winds and 892-millibar central pressure.
Hurricane Melissa Forecast Discussion 13
National Hurricane Center
Used for: Real-time evidence that every DeepMind ensemble member forecast Category 4 intensity or higher; documentation that NHC blended GDMI into its track and intensity forecasts while Melissa was still weak.
Hurricane Melissa Forecast Discussion 18
National Hurricane Center
Used for: Real-time evidence that approximately four-fifths of the DeepMind ensemble projected Category 5 intensity and that NHC regarded GDMI as its best-performing intensity guidance to that point in the season.
Hurricane Melissa Forecast Discussion 22
National Hurricane Center
Used for: Confirmation that 48 of 50 DeepMind ensemble members projected Category 5 intensity and that GDMI remained part of NHC’s operational track blend.
An Operations-Based Evaluation of Tropical Cyclone Track and Intensity Forecasts from Artificial Intelligence Weather Prediction Models
DeMaria et al. / AMS
Used for: Previous-frontier baseline showing that GraphCast, Pangu-Weather and FourCastNet produced competitive tracks but severe intensity underprediction, substantial low bias and no improvement to the intensity consensus.
How we’re supporting better tropical cyclone prediction with AI
Google DeepMind
Used for: Weather Lab launch date; 50-scenario and 15-day output design; training-data description; retrospective 2023–2024 testing; commencement of live NHC evaluation. Its performance claims are treated as developer evidence, not as independent verification.
How WeatherNext helped NHC better predict Hurricane Melissa
Google DeepMind
Used for: DeepMind’s merged-basin interpretation of the 2025 results; its account of WeatherNext’s Melissa probabilities; the developer’s description of its continuing NHC collaboration.
Skillful joint probabilistic weather forecasting from marginals
Alet et al. / arXiv
Used for: Technical description of the Forecast Generative Network underlying Weather Lab’s probabilistic forecasts. It supports architectural context but is not used to establish the season-level operational result.
This AI weather model helped predict 2025’s huge hurricanes
The Weather Network
Used for: Independent journalistic explanation of the completed NHC verification report, including the exact Atlantic and eastern North Pacific comparisons. It is contextual coverage rather than the primary basis for confirmation.
Last checked Methodology 2.0.0