A few years ago, Anima Anandkumar, a professor at the California Institute of Technology (Caltech), began developing the first open-source AI-based weather model. When she spoke with experts in the field, she encountered skepticism: weather is chaotic, physical simulation is difficult, the relevant technology had been developing for decades, and supercomputers were required; there simply was not enough data. Despite these reservations, Anima continued her work and completed the project. In less than a year, her team developed FourCastNet, a forecasting model whose performance was comparable to the best existing physics-based simulations. Thanks to Anima and the work that followed, anyone can now use a consumer-grade GPU to accurately predict weather over relatively short time horizons.[^1]
Over the roughly fifteen science episodes we have published on Latent.Space, we have discussed atoms, molecules, materials, biology, and mathematics. Anima is a pioneer in the study of continuous physical systems. Weather, fusion, and the flow of fluids or heat are all vast and extraordinarily difficult scientific domains to model: they are enormous, chaotic, and fundamentally multiscale. This is an area that the AI community has largely overlooked, but we expect it to develop rapidly. In the next few episodes, we plan to explore large-scale physical systems in greater depth.
Anima’s work illustrates why this area of AI does not lend itself to the scaling-centric approach that has permeated other fields. The data simply does not exist: open datasets in many of these domains contain only hundreds of thousands—or even hundreds—of examples, nowhere near the scale required by Transformers that “ingest” vast quantities of tokens. Worse still, the resolution demanded by physics drives context lengths into the hundreds of billions, so you cannot solve the problem simply by adding more tokens. But this is not a hard limit; it is merely a slower path forward. Progress here comes from incorporating structure and inductive biases. To the language-model researchers shaped by the “bitter lesson”: sorry.
“If there are only a few hundred grid points in each dimension—and that is already the starting point for industrial-scale problems … we are talking about context lengths in the hundreds of billions, or even a trillion. So don’t expect to build a Transformer for tasks at this scale; even the combined computing power of the entire world would not be enough.”
The Mathematics Behind It
To tackle these systems, Anima pioneered a technique known as Neural Operators, one of the most elegant developments in AI theory over the past decade.[^2] Neural operators can combine data with the laws of physics, enabling inputs and outputs at multiple scales. Rather than modeling a grid, we model a function that can evolve across multiple scales. This allowed Anima and her team to build priors based on physical intuition.
