Skip to main content
← Back to Latest

NHC verification finds Google DeepMind’s GDMI leading individual hurricane guidance in 2025

During the first season in which the National Hurricane Center incorporated AI-based models into real-time operations, Google DeepMind’s GDMI ensemble mean slightly outperformed the official Atlantic track forecast at 12–72 hours, delivered comparable Atlantic intensity skill, and beat the official forecast and all consensus track aids in the eastern North Pacific at 48–120 hours.

Topic Trail: AI in operational weather forecasting

Claim: ConfirmedEvidence: Very StrongReview: StableState: Main Entry
Methodology: v1.0.0

Frontier Delta

Previous Frontier

Operational-style testing of GraphCast, Pangu-Weather, FourCastNet and related AI weather models found that their tropical-cyclone tracks could compete with strong operational models, but their intensity forecasts were worse than even simple climatology-and-persistence guidance. The models systematically weakened storms and did not improve the operational intensity consensus.

New Claim / Result

In its first season supplying real-time guidance to NHC, Google DeepMind’s GDMI became one of the leading individual hurricane models. It slightly outperformed NHC’s Atlantic track forecast at 12–72 hours, achieved comparable Atlantic intensity skill, and outperformed the official forecast and all consensus track aids in the eastern North Pacific at 48–120 hours.

Delta

AI hurricane forecasting moved from strong track predictions paired with largely unusable raw intensity forecasts to one AI ensemble providing operationally competitive guidance for both track and intensity during a live hurricane season.

Details

What Happened?

On 12 June 2025, Google DeepMind and Google Research launched Weather Lab and began sharing live predictions from an experimental tropical-cyclone model with NHC forecasters. The system generated 50 possible cyclone scenarios extending as far as 15 days and attempted to predict formation, track, intensity, size and structure within one probabilistic model. DeepMind’s launch evidence was retrospective and developer-controlled, so the 2025 season was the first important prospective operational test.

NHC later described 2025 as the first season in which it incorporated AI-based models into real-time operations. Its preliminary report called GDMI “very useful,” while warning that several AI systems remained under development and were not always delivered reliably enough for routine use. The completed seasonal verification subsequently showed that GDMI was not merely an interesting experimental output: it had become competitive with, and under defined conditions better than, established operational guidance.

What Does the Evidence Show?

In homogeneous Atlantic comparisons, meaning forecasts were compared only where every included model was available for the same cases, GDMI slightly outperformed the NHC official track forecast from 12 through 72 hours. Its Atlantic intensity forecasts had skill comparable to the official NHC forecast and placed it among the strongest individual intensity models at several lead times.

Its strongest relative result came in the eastern North Pacific. From 48 through 120 hours, GDMI performed considerably better than the other individual track models and beat both the official NHC forecast and every tested consensus aid. That is especially notable because consensus aids combine several strong models and are normally more difficult for one individual system to outperform consistently.

DeepMind later summarized the merged Atlantic and North Pacific results by saying WeatherNext was the season’s top-performing individual model for both track and intensity. That summary is useful, but it remains a Developer / Vendor Claim and compresses substantial variation between basins, forecast ranges and metrics. The more granular NHC findings above should therefore define the public VyDex claim.

Hurricane Melissa provides the clearest operational example. On 24 October, while Melissa was still a relatively weak tropical storm, every DeepMind ensemble member projected Category 4 strength or higher, and NHC said its track and intensity forecasts were being blended with GDMI guidance. By the following day, roughly four-fifths of the 50 members projected Category 5 strength, and an NHC discussion described GDMI as the best-performing intensity guidance to that point in the season. On 26 October, 48 of the 50 members still reached Category 5.

Melissa ultimately struck Jamaica on 28 October with estimated sustained winds of 185 mph and a central pressure of 892 millibars. NHC had issued almost three days of advance notice of a Category 5 Jamaican landfall—the first time it had forecast Category 5 intensity while a storm was still only Category 1.

What Context Changes Interpretation?

This result does not mean GDMI beat NHC or the best consensus systems everywhere. Performance depended on basin, forecast horizon and metric. The strongest verified superiority was in eastern North Pacific track forecasting at 48–120 hours. Atlantic track advantages were smaller and concentrated at shorter ranges, while Atlantic intensity performance was described as comparable to the official forecast rather than universally superior.

NHC’s official forecast remained the stronger overall decision product in several of the season’s hardest situations. Across Atlantic rapid-intensification cases, NHC achieved a higher probability of detection and critical success index than every real-time model. For Melissa specifically, its official intensity forecast outperformed every model at nearly every lead time. GDMI contributed valuable evidence to the human forecast; it did not independently produce the official warning or replace the observations, aircraft reconnaissance, physical models, consensus systems and forecaster judgment behind it.

GDMI was also unavailable during the early portion of the season. The homogeneous comparisons excluded the first two Atlantic storms and the first five eastern North Pacific storms, so this was not a complete all-storm evaluation. NHC separately noted that some experimental AI systems were not consistently available in time for routine operational use.

One season is important prospective evidence, but it cannot establish performance across every possible hurricane environment. The Atlantic had fewer forecast cases than normal, while an unusual concentration of Category 5 storms and rapid-intensification episodes made intensity prediction particularly difficult. Future seasons could expose weaknesses that the 2025 sample did not.

The system’s ability to generate scenarios extending 15 days should also not be confused with verified, high-accuracy hurricane intensity guidance across that entire horizon. The central operational comparisons in the NHC report extend through five days. The 15-day horizon describes what the model can generate, not a demonstrated 15-day intensity-forecasting threshold.

What Should the Reader Take From It?

The supported conclusion is that AI hurricane forecasting crossed a meaningful operational threshold in 2025. Earlier global AI weather models could provide useful track guidance but generally smoothed away the strongest winds and failed badly at intensity. GDMI became a system that forecasters could consult in real time for both dimensions and that subsequently achieved leading individual-model performance under formal NHC verification.

The result should not be interpreted as hurricane forecasting being solved or human forecasters becoming unnecessary. It shows something narrower and more consequential: an AI model overcame a central documented limitation of the previous model generation well enough to become useful inside one of the world’s highest-stakes operational forecasting workflows.

Significance

Confirmed Significance

This was the first formally verified season in which NHC incorporated AI-based model output into real-time hurricane operations. The system’s strongest contribution was not merely faster global weather prediction or improved track forecasting—areas where earlier AI models had already shown substantial capability—but competitive intensity guidance, historically the major weakness of global AI weather models.

The milestone is therefore the movement from retrospective model-development claims to prospective, real-time and institutionally verified performance. GDMI’s forecasts influenced operational discussions during major storms, survived comparison against final post-season storm records, and in some conditions outperformed both the official forecast and multi-model consensus guidance.

Potential Significance If Confirmed

If comparable performance is reproduced across additional seasons, basins, and operational centres, AI guidance could become a durable part of high-stakes tropical-cyclone forecasting for both track and intensity, giving human forecasters stronger probabilistic evidence earlier in a storm’s development. The 2025 result does not establish worldwide generalisation, consistently longer warning lead times, or reduced casualties.

Caveats

  • Performance was conditional. GDMI’s relative position changed across basins, forecast horizons and metrics. “Leading individual model” must not be rewritten as “best forecast system everywhere.”
  • The evaluation covered only part of the season. GDMI was unavailable for the first two Atlantic storms and first five eastern North Pacific storms.
  • Human forecasts remained stronger in crucial cases. NHC performed better across the rapid-intensification subset and outperformed all individual models for Melissa’s intensity at almost every lead time.
  • This is one season. Formal operational verification makes the result much stronger than a retrospective demonstration, but it does not yet establish durable performance across different multi-year storm distributions.
  • The system remained experimental. NHC reported availability and timeliness problems, and DeepMind described Weather Lab as a research tool rather than an official warning service.
  • The 15-day output horizon is easy to overread. The strongest verified comparisons concern forecasts through 120 hours, not consistently accurate hurricane intensity prediction 15 days ahead.
  • Melissa is supporting evidence, not the whole case. Its forecast demonstrates practical relevance, but the entry’s status and evidence strength rest on the complete season verification rather than one successful storm.

Sources

Forecast Verification Report: 2025 Hurricane Season

National Hurricane Center

Used For

Final post-season verification methodology; Atlantic and eastern North Pacific model comparisons; GDMI’s track and intensity performance; homogeneous sample boundaries; model availability limitations; official 2025 season findings. The report was published on 30 March 2026.

Open source: Forecast Verification Report: 2025 Hurricane Season →

2025 NHC Verification Report Preview

National Hurricane Center

Used For

Confirmation that 2025 was NHC’s first season incorporating AI models in real time; preliminary rapid-intensification evaluation; Melissa forecast verification; statement that GDMI was useful but experimental systems were not always available on time.

Open source: 2025 NHC Verification Report Preview →

An Operations-Based Evaluation of Tropical Cyclone Track and Intensity Forecasts from Artificial Intelligence Weather Prediction Models

DeMaria et al. / AMS

Used For

Previous-frontier baseline showing that GraphCast, Pangu-Weather and FourCastNet produced competitive tracks but severe intensity underprediction, substantial low bias and no improvement to the intensity consensus.

Open source: An Operations-Based Evaluation of Tropical Cyclone Track and Intensity Forecasts from Artificial Intelligence Weather Prediction Models →

Methodology Used

This entry uses Methodology v1.0.0.

View Methodology v1.0.0 →