Back to feed

The AI That Saw a Category Five a Week Out: Inside DeepMind's Weather Revolution

Google DeepMind's WeatherNext model called Hurricane Melissa a category five a week before landfall; the podcast walks through the physics, limits, and energy-to-agriculture promise of AI weather forecasting from GraphCast to GenCast.

Imported to Nodesdaily: (UTC+03:00)
Watch on YouTube — O_EWbnkjXdk
Reading options

Device speech is unavailable in this browser.

Concept lens

Choose a technical term in this view to read its general definition, teaching example and use in the article.

No terms from our glossary were found in this view. The glossary does not cover every term yet.

When a hurricane moves toward land, time is everything: a single extra day of warning clears the roads, empties the hospitals, opens the shelters. We can compute the exact second of a sunrise a thousand years out, yet knowing next week's weather with certainty remains one of the most stubborn problems in physics. A tiny wobble in today's air pressure can grow into a major storm a week later: the famous butterfly effect in person.

Google DeepMind has been wrestling with this problem for years: a journey that began in 2020 with models reading satellite images for short-term rainfall now reaches WeatherNext, which predicts the whole planet hour by hour. The podcast guest is Peter Battaglia, senior director of research at DeepMind. Host Hannah Fry opens with a concrete case: Hurricane Melissa.

Melissa was born in October 2025 as a tiny Atlantic disturbance and stayed unusually small for a while; systems like this can draw energy from warm seas and turn into the most dangerous storms. About a week before landfall, the DeepMind model began insisting this system would climb to category five. At that point there was not even a tropical storm, most likely just a tropical-depression-level feature. The meteorology community on X noticed the model's persistence first.

Other simulations were watching the same patch of ocean; some saw intensification, but none matched the same confidence, track, and intensity together. On Saturday the National Hurricane Center issued a category-five call while the system was not even category one: category five by Monday, landfall by Tuesday. That bought settlements roughly three days of preparation time.

The toll was heavy: Melissa came ashore over western Jamaica on October 28 with 185-mile-per-hour winds, and 45 confirmed deaths were recorded. The video cites a higher early figure, while independent reports documented 45 deaths in Jamaica plus missing people. According to the Center, it was their first category-five forecast issued from such a low starting intensity. Forecasters said watching the model's confidence scores, reaching toward 80 percent, had raised their own confidence.

Fry asks the key question here: does the AI model hold a structural edge in long-range hurricane prediction, or does it get different answers simply by doing something different? Battaglia's answer is plain: the physics that operated in the past will operate in the future. With enough evidence from history and models able to read the fine structure between inputs and futures, forecasts naturally get sharper. No magic is doing the work, only data discipline and modeling quality.

The case did not stay a one-off success story; the team turned cyclone prediction into a dedicated product. Four years of global models forecasting the surface and atmosphere out to 10-15 days matured first, then a focused team spent about two years on tropical cyclones alone. Partnerships with academics and operational centers were built, with months of trusted-tester trials. Experimental evaluations showed roughly a day of accuracy gained: like moving a two-day forecast's skill out to three days. Live forecasts appear on the Weather Lab site alongside temperature and wind variables; anyone can touch the same model through the WeatherNext feed or the weather app on their phone.

Why is forecasting becoming more urgent? The answer sits in physics: weather is energy. The sun carries electromagnetic energy, wind carries mechanical energy, temperature carries kinetic energy; a warmer planet means more energy, and more energy means more intense weather. The hottest-years list being filled by the last 10-15 years reads together with wildfires and floods.

So why is next Wednesday's rain still a mystery despite all this scientific muscle? Because small things cause big effects, and we cannot observe every butterfly. Limits in observation convert into limits in prediction. But this is not a hopeless picture: important processes leave small traces, statistical breadcrumbs. Battaglia's pond analogy sticks: the small ring from a pebble foretells the big wave reaching the shore, and machine learning is in the business of reading those rings.

Why would an AI lab enter the weather business? Battaglia gives two reasons. First, mission: weather prediction is science's oldest forecasting problem and touches everyone's life daily; abstract it is not, universal it is. Second, personal and technical: the team had spent years on machine learning for fluid simulation, and the atmosphere is one giant fluid problem. On top of that, the European Centre for Medium-Range Weather Forecasts held ready-made records spanning decades: a perfect fit for machine learning.

Battaglia also explains the classical method: numerical weather prediction is a supercomputer loop that takes the current estimate of the weather, computes a few hours ahead as an approximation to the fluid-motion equations, and feeds its output back in. The butterfly effect plus cross-scale coupling make exact solutions impossible, so coarse, blurry-scale approximations are used instead. Even so, this is a decades-long triumph of science and engineering: the horizon moved from a day or two out to two weeks.

What AI does is conceptually familiar: watch how past weather unfolded in time, learn the statistical patterns, extrapolate forward. Battaglia compares it to fitting a line to numbers, except the line is enormously complicated and the data enormous. The journey came in three phases: first machine learning patched numerical models, then came regional image models, and in the third phase came models simulating the entire planet. GraphCast pioneered this third phase: take the whole planet's weather and forecast out to ten days.

GraphCast's mechanics are elegant: the model takes the full state of global weather and applies the same learned local operation everywhere, reflecting that physics is the same everywhere. Then it fuses this into a planet-scale representation and predicts back down to the local level, say London. The west side of a hundred-kilometre hurricane says a lot about the east side's future, and seeing the storm in one representation lets the model use that. Classical prediction stays local while AI also reads the large scale; that is largely where the gap comes from.

Then forecasts turned probabilistic. A deterministic model gives one answer per input; a probabilistic model produces hundreds of scenarios showing the distribution of likely futures. For extreme events this is vital: even a 5-10 percent scenario deserves preparation. The hurricane cone and spaghetti plots on television are this idea's visual language. GenCast, arriving around 2023 after GraphCast, targeted scenario diversity rather than the average, and that is where the bar rose for extreme-event prediction.

Two techniques generate this diversity. First, diffusion models familiar from video generation: descending from a very noisy image toward a weather image, each different start lands on a slightly different scenario. Second, the team's own invention, functional generative networks: slightly perturbing the network's weights while trying hundreds of input scenarios. In Battaglia's phrasing, each run asks "what if the butterfly were over here"; hundreds of runs reveal the most likely future. In the Melissa case, overlapping lines read as the model's confidence, scattered lines as its humility.

WeatherNext 3 rebuilds the value chain instead. Raw satellite data used to become a current-weather estimate, an operational model ran, then application-specific models followed; Battaglia calls this the weather tree: roots are data, the trunk is the model, branches are applications, leaves are the phone screen. The new model goes root-to-leaf in one architecture: raw satellite images in, station measurements like airport readings directly out. Forecasts refresh hourly instead of every six hours, at higher resolution. Announced in September 2026, the model consumes real-time satellite data and feeds search and map products too.

The energy payoff is concrete: load forecasting, electricity-demand prediction, leans heavily on temperature, with heaters straining grids in the cold and air conditioners in the heat. The supply side watches wind speed and sunshine. The model's new wind and solar variables foresee supply, while high-resolution temperature and humidity foresee demand. Grid planning gets fed from both sides.

On the farm front, growers are the oldest forecaster class: planting and harvest calls depend on temperature and rain. Horizons like weekly rainfall or monthly mean temperature can even change seed choice. On the architecture side, not every novelty weighed the same: adding solar and wind variables was relatively easy, while satellite input and station output demanded deep structural change. The performance ledger stays honest: GraphCast and GenCast were the first big wins in deterministic and probabilistic forecasting; track and intensity met in a single cyclone model, but basin and regional differences matter to decision makers, so no one-line "we are the best" claim is made.

The error regime separates the two camps: in classical methods a small error can grow and blow up, since equations never say "return to normal weather." AI models regress to the mean when they err, behaving like seasonality, the oldest forecasting method of all. Here lies the paradox: how does a mean-reverting model catch storms stronger than anything seen? The mosaic theory answers: the model never saw that storm, but it saw the pieces in other corners of the planet and fused local statistics into a novel whole. Still, few-shot regimes like the yearly monsoon remain an open learning-efficiency problem. Impact forecasting waits on the horizon: predicting not the wind but the damage the wind does to power lines. Physics is not abandoned, only the way of solving its equations; what sets truth is not interpretation but evaluation discipline.

Visualization: nodesdaily AI

AI commentary

"What struck me most in this episode was not the forecast itself but the humility behind it: a model that runs hundreds of scenarios and also says what it cannot know. I cross-checked the figures against independent sources; this is written as a critical reading, not a single-case tribute."

AI assessment

Before joining the episode's enthusiasm, I look at the test reported by New Scientist: AI forecasts are fast and accurate in average weather but can miss rare, unprecedented events. The mosaic theory sounds powerful, yet as the climate shifts, the "unprecedented" stops being an exception and becomes the norm. Building tomorrow's whole from yesterday's pieces will structurally strain under a climate that moves outside history.

Battaglia himself concedes the methodology limits: little work exists on extreme heat, extreme cold, and fine-grained events like tornadoes. The July floods, where classical systems came out most accurate, and the 2026 German study finding physics-based models more reliable for extremes, back that concession. A trackable cyclone like Melissa and a flash flood developing in hours are not the same problem; on the second front AI is still a spectator.

On verifiability I stay cautious too: the story is told by the research director of the institution that built the model. The "the Center gained confidence watching our model" claim is a one-sided quote whose independent counterpart should be sought in the NHC's official assessment reports. The numbers also want checking: the death toll and wind speed cited in the video run higher than the 45 deaths and 185-mile-per-hour values in independent reports. Generalizing from a single case, especially while budget cuts at US weather agencies shake research continuity, deserves a careful read.

My takeaway is clear: for someone planning a grid, calling a planting date, or timing travel, these models are a usable gain today; hourly, high-resolution forecasts improve daily life. But anyone deciding under extreme events — a disaster manager, an insurer, an infrastructure operator — should not lean on a single model. Ensemble agreement, physics-based models, and the full probability distribution must be read together; AI is a strong instrument in the orchestra, not its conductor.

Sources

9 links; no other published story cites them. Stories sharing a link do not confirm each other; a source's origin is not inferred from how often it is cited.

weathernext · deepmind · weather forecasting · hurricane melissa · graphcast · gencast

Follow the topic

Before this story

A short reading order from earlier stories linked to this event by an editor.

Evidence and sources

Review permitted source passages, versions and origins.

KAYNAKLARLA OKU

Bu haberi açalım.

Hesap kontrol ediliyor…