The Neil Ashton Podcast
Prof. Anima Anandkumar — The future of AI and science
Watch on YouTube
Prof. Anima Anandkumar — The future of AI and science
YouTube video
Watch this episode
YouTube is contacted only after you choose to play the video, keeping this page fast and private by default.
Listen to the audio
Episode overview
Professor Anima Anandkumar is one of the world’s leading scientists in AI and machine learning, with more than 30,000 citations, an h-index of 80 and landmark papers including FourCastNet, which received worldwide coverage for demonstrating how AI can accelerate weather prediction. She is the Bren Professor at Caltech, leads the AI+Science Lab and previously served as Senior Director of AI Research at NVIDIA. In this episode I speak to her about her background in academia and industry, her journey into machine learning, and the importance of AI for science.
We discuss the integration of AI and scientific research, the potential of AI in weather modeling, and the challenges of applying AI to other areas of science. Prof Anandkumar shares examples of successful AI applications in science and explains the concept of AI + science. We also touch on the skepticism surrounding machine learning in physics and the need for data-driven approaches.
The conversation explores the potential of AI in the field of science and engineering, specifically in the context of physics-based simulations. Prof. Anandkumar discusses the concept of neural operators, highlights the advantages of neural operators, such as their ability to handle multiple domains and resolutions, and their potential to revolutionize traditional simulation methods.
She emphasizes the importance of combining AI with scientific knowledge and traditional numerical solvers, supported by collaboration between machine-learning specialists and domain experts. Finally, she offers advice for PhD students and highlights the value of smaller workshops and conferences for following emerging ideas.
Chapters
- 00:00 Introduction and Overview
- 04:29 Professor Anima Anandkumar's Career Journey
- 09:14 Moving to the US for PhD and Transitioning to Industry
- 13:00 Academia vs Industry: Personal Choices and Opportunities
- 17:49 Defining AI for Science and Its Importance
- 22:05 AI's Promise in Enhancing Scientific Discovery
- 28:18 The Success of AI-Based Wea
References and links
Transcript
This transcript was created from the corrected YouTube captions, with names and technical terminology reviewed. Download the corrected SRT file.
Hi and welcome to the Neil Ashton podcast. In each episode, we explained some of the fascinating ways that science and engineering are changing the world around us. We talk to leading engineers from elite level sports like cycling and Formula One to some of the world's top academics to understand how fluid dynamics, machine learning and supercomputing are bringing in a new era of discovery. We also hear some of their life stories, their career advice and lessons they've learned on the way that I hope will be helpful to you too. So sit back and enjoy this episode. Hi and welcome back to the Neil Ashton podcast. Today's episode is a continuation of the theme of the past few. This I guess mini
series focusing on AI for science. This idea that you can use machine learning and artificial intelligence to accelerate in hand, improve scientific discovery in fields such as fluid dynamics, computational fluid dynamics. So, you know, car design plane design that has been common throughout this podcast, but of course, also things like weather forecasting or drug discovery. And today's guest is most definitely one of those people who is seen as a leader and a pioneer and someone at the forefront of these methods. That person is Professor Anima Anandkumar, who is a Bren Professor at Caltech. But interestingly, and I think what has arguably made her such an important um person in this field
is that she has also been very close to, to industry. So she was first um at Amazon Web Services as a principal scientist focusing on, on AI and this was actually before, I suppose, the mad craze now, um, with AI and GenAI. Um and and recently was also senior director at NVIDIA for AI research. So she's spent not only time in academia but crucially at very senior roles at these tech companies that are at the forefront of machine learning. So that's why it's so interesting to get her perspective in this field because she has not purely been doing research in academics. And she's also been there in these companies and seen
um you know, its application to, to more real life problems. And um we, we discuss a number of things in, in this podcast as with any podcast, there's so many things that we didn't get to cover. So I'll be putting lots of links uh in into the comments on YouTube for you to, to read papers and look at other things that, that her group has done. But we talk quite a lot about one of the achievements and one of the things that she's arguably best known for which is uh neural operators, this this type of machine learning method that is uh very well suited to solving scientific problems like fluid dynamics. Uh And is um probably best known in a way for the work that was done for
FourCastNet. And if you look at the links below, you can read the paper that was one of the first papers or first approaches to really revolutionize um weather prediction. And so we talk a little bit about where things are at. You know, what, where does she see the current state? She focuses heavily on AI and science, not AI for science because she really sees them as complementary. And we dive into the usual questions of, you know, how close are we to being able to do in real life? How important is physics, some of these ideas of foundational models, some of the topics that I've been asking the other guests. But I,
I almost want to repeat the same question to the different guests because then you hear the different opinions and hopefully you as a listener can get a more rounded view of, you know, where the community feels it's at. Uh I'd say uh Anima is more bullish um on this. I think she has more confidence that we are closer to making some key breakthroughs. Um And, and we talk through that and we talk through some of the actual Pacific applications that, that have been made. Um interestingly, we also talk a little bit about her um career and her story so far. And I think it's very interesting her move through academia um to, to, to industry and,
yeah, really, um, really interesting. She's a fantastic speaker. She's actually just done a TED talk. So I'd highly encourage you to, to watch that because as we were discussing um more off air, I think it's hard to explain some of these things in either purely verbal, if you're just listening on the podcast or even just in a, in an interview fashion, you really want some slides and pictures to explain some of these topics. So if you are interested in AI for science or AI and Science, uh I would highly recommend you to watch her TED Talk cos I think it gives you a more visual uh impression of some of these topics.
So please sit back uh and en enjoy this episode. You're a full professor at Caltech, probably one of the most well known institutions in the world for, you know, science and engineering and many people, I'm sure have a dream of, of getting to that when they're, you know, when they're younger, when they're studying. Did you have academia in your mind from a, from a young age when you were in school? What was your, what did you want to be when you were growing up? Yeah. Yeah. No, I, you know, certainly Caltech was my dream too. I grew up being extremely inspired by Richard Feynman, you know, reading his books, but,
you know, his books on physics as well as his life stories. And, yeah, to me, you know, I wanted to be innovative. Right. And I kind of like, fully didn't think what that path would be because my background is kind of quite worried. My, both my parents are engineers and in fact, in India, it's quite rare for my mom, you know, in her generation as a woman to be an engineer. And that was just great because, you know, at home, it never felt like as a woman, you aren't good at something just because you're a woman. So it, it gave me a lot of uh just really good role model for someone who's a woman who is also a great engineer.
And both my parents um started a factories, manufacturing small components for automotive industry. And in the early nineties, they brought some of the first computerised machinery to my hometown in India. And so they really kind of like were forward looking and thinking. OK, how do we make this more efficient? How do we bring programming into manufacturing? Right. And uh so that was also a different way to get introduced to computers for me because it was always something more physical, like it did kind of like produce these components, it kind of manufacture them and machine them. And so that was also an aspect that was very interdisciplinary in what they did.
Um So that was great. Whereas, you know, my grandfather was a math teacher and I was always very excited about just solving puzzles and I would do that as fun. Although for a lot of people, math is like this homework they have to deal with. Um And I really liked him as a teacher and, uh you know, so I've had like role models both in industry and academia who, who were just great and I wasn't kind of set on one path, but it kind of led me to where kind of, I felt I would have the most access to doing minor work. Um You know, when I was in the process of finishing my PhD and going to the job market in 2008, you know,
those of us who know it was just such a difficult time, right? It was the first time to be on the job market. And at that point AI and machine learning was not a job description. So industry wasn't even an option. And of course, with all of this financial crisis, it was even less of an option. And, you know, I got a faculty position at the University of California at Irvine. And I was like, OK, you know, I can continue to do this work in machine learning, which right now many people think as science fiction, but that's what I want to do. I want to make that work. And here we are today. So I'm really glad how that that took me.
Uh And once deep learning started really taking off and there were all this industrial expat, I was also really lucky to be placed in many of the top industrial roles. You know, I uh went to Amazon Web Services, helped start the cloud AI group build some of the first AI products on the Cloud, then went to NVIDIA. So really kind of also following the journey of AI expansion in industry and first being in academia to build those foundations has been just a dream come true. Yeah, totally. And yeah, I really would love to dive into some of those uh those details. But what was it like moving to the US to do your PhD? That must have been
exciting daunting at the same time. Had you been to the US before or was that actually sort of a, a very big move for you to do that? Um You know, it always felt like I wanted to be in the top places to do A PhD, right? And do research and I had applied at various kind of schools. And Cornell is where I really connected with my advisor, you know, Lang Tong and I picked that uh but it was just uh culturally us is something that, you know, I've been close to, I've been to Europe several times. Before. Um And so, and also it's such an international place when you go to these universities. So, you know, that aspect wasn't an issue at all.
It really, he was like, ok, I can now learn from all the cultures around the world, meet people from all over the world. The only kind of issue was, oh, it's damn cold. Like, you know, coming from a tropical place it took in upstate New York was a, was a big, let's say surprise. But that also was, you know, I was like, what can we do about this? And I decided to take all kinds of winter sports. I took up ice climbing, which is really challenging, not to say I became an expert. But I was like, what can I do to really get out of my element and embrace this cold weather? Is that one of the reasons you then went to the
sunniest part of the US or was it completely separate? Yeah, I mean, like, as I mentioned in the height of the financial crisis, there weren't too many choices. So I was glad I got this faculty role and, you know, after that was Caltech and uh you know, there's a lot in California that's, you know, really great. So, yeah, happy to be here with them. Yeah. No, no, no, definitely. And did you, um at what point I, I've spoken to a few people about this so I'm kind of interested of this choice between academia and industry? Was it a conscious decision? Did you have a desire to um not only be a pure academic because you wanted to go
into tech or did you sort of fall into it by accident? Um You know, like I mentioned earlier, I wasn't like hard set ever to be in one versus the other and I don't see it as even an exclusive choice. Uh And, you know, like as I mentioned, uh straight out of PhD, industry wasn't even an option. So I had to go to academia, really build up the methods to show their work, right? That's when industry decides to start investing and when that happens, and there were roles where I could really contribute, like, you know, have this outsized impact of being able to start something entirely new, a new AI division or AI research,
you know, that was like to me, something that I was like, OK, I can really make a step change here and to the whole community and be able to publish open source. And that really was a great motivator as well for me to take that step while also keeping a foot in academia to really bridge the gap between industry and academia. And would you recommend that to people, do you think that's actually quite a wise and good thing to jump between academia and industry? You know, if, if you, if someone's listening to this now doing a PhD, would you, I know you said you kind of had to go into academia because 2008, 2009 was a
unique time. But would you actually recommend it in a way? It's really about personal choices? Right. And, and also in industry research as well, we're seeing more changes than before. It's not completely free for those guys. Complete freedom. Uh you know, at one point it used to be at least to some extent and we are seeing more of the closed models and uh kind of more, you know, targeted goals and for some people that's really what they want to do, right? And there are others who want to be working on areas that are like still new, still unexplored, making interdisciplinary connections. And perhaps for those cases, industry is not the best,
especially if you don't get assigned to a team that allows you to do that. So it's really one of like personal choices. So what is it that you value the most? And, and certainly in industry, there are more resources, of course, depends on the place. But in big tech, you do have a lot more GPUs. Uh but there is more of a focus, you know, aspect of, you can only work on a certain set of topics and areas that, that was actually, uh maybe before we move on to, to, to another topic. Uh where do you see that? Because I, I've read that that in a, in a strange way, academia in most other fields has been very much the developer of the fundamental science,
the fundamental methods. And then industry takes it once the technology readiness level is higher. When it comes to this current age of machine learning. And AI, there seems to be almost not the, not the opposite at all, but there is a huge amount of research and development going on in tech companies. And I'm, is there a brain drain essentially going on that because they have the resources to hire in the talent that in a way we're losing some people who would have naturally gone into academia. Um You know, for me, I don't like the term brain drying no matter between countries or between organizations, right? Because
there's always more brains to be brought in. So kind of like let me start with that. Uh you know, and, and sure, you know, like industrial roles speak to a broad set of students because there is resources and very well defined problems and you know, let's go work at it in academia. There's fewer resources, but that also means necessity is the mother of all inventions, right? Like how do we do more with less? So can we improve our training methods? Can we look at scenarios where there isn't as much data? How do we do learning in that? And it's always academia that works on problems that industry is typically not looking at because it's considered as too impractical
machine learning, all of machine learning. Was that at some point? Right. So I'm just saying that it's ideally in academia, it would be great to have more resources. But you know, that shouldn't be the only aspect when it comes to making the choices. So when did you first get into machine learning? And was that a conscious choice? When was it in your head that you really started to double down in that area compared to other areas of science and engineering? Yeah, I mean, during my undergrad too, I worked on signal processing which is essentially machine learning, right? Like either you work with the radars or image processing and
you know, there was not deep learning at that, although we knew neural networks as a concept, it wasn't considered practical, but the concepts were there those fundamentals were there, my undergraduate thesis was looking at Iris recognition biometrics, right? So it was really just these aspects that OK, there are all these important problems to be solved, maybe they're not practical today. But how do we start building the foundations of the algorithms? How do we frame learning as a problem? And what are the requirements in terms of data in terms of the right algorithms and constraints? Right. So we were doing a lot more of the theoretical analysis because
you couldn't like kind of run algorithms at scale at that point. Uh But that also meant we were thinking about it systematically and all of that led to the further developments. Hm. And what about, um, the movement into, I guess AI for science? So maybe pivoting a little bit more to, I guess what has been, um, I would argue you've been one of the leading voices in this field that's really pushed it along and given it more of a global um awareness because it a bit to your point of um companies work in some areas where there's more particular focus on others. And I guess the science one is arguably not immediately the stuff you see on the TV, the, you know,
the GenAI, the chatbots, et cetera. Um But I think voices like yours have helped to bring it out. Maybe we can, there is such a big topic. So maybe let's start by how would you define AI for science and why do you see it as being such an important area for us to, to work on? Yeah, I mean, to me, you know, being at Caltech, that was one of the first things I thought about when I came here, right? Like I came here thinking about Richard Feynman and all of the scientific developments that have happened here. And the natural question is what can AI do for enhancing and accelerating those scientific developments? And,
you know, back then in 2017, when I came here, there was still a lot of skepticism of even about AI or other areas, forget its usefulness to other areas. They were like, oh, is a, I, I even a real thing that was still early days and now of course, no one asked that question. Uh But I don't even think of it as AI for science instead. I think of it AI plus science. It's the deep integration of AI and what we call scientific research in all kinds of aspects. Um Because if you think about how scientific research is done since the time of Newton, you know, what do we call the scientific method? It's the aspect of coming up with ideas, right?
So this is where experts are better because either they get a eureka aha moment or they're just thinking and thinking and like saying, OK, I've kind of rejected all these hypothesis. So this one works. So, you know, there is this deep intuition and domain expertise that people develop over time and then that's not right. It's not just the great ideas that propel science, but the hard work spent in the labs to actually go test it out, validate it. And those experiments many times just disprove many of the theories that you have to go back and come up with new explanations, new ideas to test further. And this is great,
but it's extremely slow because the bottleneck is not ideas, you can keep generally thinking of all kinds of like, you know, our minds are creative. People have been coming up with all kinds of theories about this planet, but only a few of them are correct because you have to go measure and carefully validate those. And so to me, AI can help in all of these aspects, right? AI could come up with new ideas. You're seeing language models do that today. Uh You can come up with new proposals of materials, drugs, It can, you know, even design aircraft wings. If you ask DALL·E or stable diffusion, it can generate something that looks like an aircraft wing.
But you still have to go physically tested in the lab to validate that it's correct. And the problem of hallucination means many of those ideas that AI generates today is not that useful. You have to go through a lot of testing and discard a lot of the proposals made by deep learning before you get a success story. And that's why, you know, my kind of premises AI plus science when I say is not just using AI to come up with hypothesis or ideas, but AI that is deeply integrated with understanding the scientific models and processes that themselves, by which I mean, ideally one day we would reduce or completely remove the
physical testing that is needed because AI can internally simulate and really understand all the processes that are involved in that reasoning that yeah, I really would love to dive a little bit deeper into this. So could you maybe give some examples of where AI has already shown great promise in this area and there may be uh some areas where it's more challenging. Yeah, I mean, to me, like there's so many great success stories of AI and science coming together. Um by the way, also, I recently gave a TED talk that has come out so I encourage you to go check that out. Yeah. Yeah. No, I'll put it in a link. It's very, very good.
Yeah, I'll make sure if people are watching on YouTube, go to the comment section, the links and I'll put it there. OK. OK. Fantastic. Uh So to me like, you know, one of the really impressive success stories is the AI-based weather model, FourCastNet. That was the first AI-based model that we started working on more than three years ago and we released it first and other teams followed up and we now have a whole family of AI-based weather models, right? And what is impressive is, you know, there was a lot of skepticism as we were working on it because a lot of domain scientists felt oh there's decades of work that has gone into building these weather models through numerical methods.
And what those methods do is from ground up, try to simulate the physics, right? Like looking at fluid dynamics through NA Stokes equations, looking at heat transfer and all of these processes that you're simulating and using that to forecast the weather in the next time step. And there was a skepticism, how can AI learn all of this complexity? It's highly multi physics, high dimensional uh can AI really be able to do a good job in this? And to our surprise, you know, our very first attempt got us to being tens of thousands of times faster than weather numerical weather models. But not only that, it even ended up doing better on many aspects like extreme weather prediction.
Uh In fact, the recent Hurricane Beryl, our forecast model had like a better kind of uncertainty band around the where the true landfall happen compared to the traditional weather models. So we're not even having a question of trade off that AI is so much faster but worse. That's not the case AI is both faster and better, right? And why is this happening? So it's better because it's able to learn from all that historical data and able to adapt based on that. Whereas numerical models tend to be a bit more rigid and you can't easily adapt it with respect to the data. Um The other aspect why it's so much faster is instead
of like doing bottom up simulation of all of the processes, it's really learning to take, let's say bigger steps, right? You don't need the fine grid or the fine resolution that Nayer Stokes in a fluid simulation with the traditional methods would take, you can afford to take bigger jumps because AI learns nonlinear transformations. So it's able to learn what is the shortest path to get us to the correct answer rather than being forced to take these very small steps on a fine grid that numerical methods need to do because they're each time solving it from scratch and they have a fixed set of steps that is already presigned and not learned from data.
And maybe for people who are not so familiar, um what uh could you maybe describe a little bit the FourCastNet and the neural operators that are, that are behind it, the at a high level, the theory behind these these approaches. Yeah, absolutely. So you know, when it comes to training weather models on the data, right? There's historical weather data, we can ask, oh there's the current weather forecast, the pre the future weather, right? So you can do this uh forecasting model uh through training. But the question is of course, what is the right kind of model architecture to be used in these kinds of processes? So if you look at fluid dynamics, for instance, how the hurricane moves,
you can't just eyeball it and precisely predict where it's gonna move, right? If you just stare at a hurricane, we are not good at predicting where it's going to go. So this is a superhuman capability. So it's not just like, you know, we can do intuitive physics of very simple kind, but this is much more complex. And that's because it requires fine scale features, meaning you need to zoom into the details of the very what we call fine resolution or fine s and see how they move and how that relates to the macroscopic behavior. And if you use standard machine learning models like transformers, they have to be working on fixed size patches because the tokens are fixed there.
And that may be OK in some applications. But if you really want to get to the fine details of this fluid flow, you should be having the capability of working across resolutions. Meaning one model that learns information across multiple resolutions. And that's what neural operators that we designed are able to do because what they're learning is mapping between functions, meaning it represents data as continuous functions and maps them to answers that are also modeled as continuous functions. Meaning we are not just thinking of the globe as grid points but as continuous processes that happen everywhere along the globe, right,
not just at grid points. And by learning such a model, we can really be able to now capture these fine scale processes accurately. So one of them and it's fair to say that this really has been a seminal piece of work that is, as you say, seems like it's uh kicked off a a very positive and needed. Um I don't want to say race, but, you know, there's, there's a lot of people now trying to sort of outdo each other, which is sometimes good for, for science when that sort of thing happens. Um But one of the discussions I think people have, um and I'm sure you meet people like this all the time who are maybe slightly more skeptical of,
of machine learning is, is the physics angle. If I'm not mistaken that in that approach, you're not explicitly solving the PDEs, you're not explicitly encoding the boundary conditions. You, you're using a more data driven approach to do it. Is that correct? No, not entirely because you have the flexibility of doing it as a physics informed approach where you can add the physics losses along with data, right. So if you only did it on the physics losses, it's too difficult on optimization problems. So you just cannot solve it. So this is one of being practical that you can do both a mix of data and physics. But the aspect is no matter how you train the model,
you can always in each instance test whether it satisfies your loss of physics, right, or your your PDEs themselves. So you can always like test it for validity in physics, which is a great thing, especially for partial differential equations. We know what the ground through satisfied. So it's always easy to verify if the answer that's obtained is correct or not. Um So I don't get the skepticism because in cases where you can easily check if machine learning is correct or not, should be the ideal case to apply it because you can always say that, oh, this is not accurate enough for my application. In which case I can use that as a precondition
and initialize my software and further solve it even if you know, there are all kinds of ways to still use this model. But in areas like weather modeling, what you've seen is that it does better than what decades of numerical weather models have shown. And that's because it's able to learn and adapt from data. And what numerical models are not able to do is about the modeling error, right. So you can assume this is the model, but that's not how the planet is. There's always deviations from that model. So you should also account for modeling errors which is difficult with numerical methods compared to data driven approaches because you with data,
you can really bring down those errors and fit to the data. Well, I I think um what I was meaning more is uh for the weather, it seems a very ideal use case because there has been this collection of data for decades and at least currently openly available data. Um And so that's what I mean in in the FourCastNet, you don't have to use a physics informed type approach because you have a lot of data to learn it on. Um what a and I guess the geometry is always fixed because it's, it's always the earth. Um where if we now move to, let's say aircraft design or, or the bigger fluid dynamics,
uh where do you see, where do you see where we're at at the moment and where we can go, how do we, do we have, do we have enough data basically and therefore should it motivate other approaches? Yeah. So it's always a challenge of like, you know, how do you get enough data? Right? And we can use the current numerical solvers to do it. But the question is of course, what is the computational cost for it? And is there a way to reduce that? And this is where a number of techniques we've been developing has been helpful where we look at like progressive training. So you have a curriculum of training from simpler physics to more complex physics
and that simpler to complex can be in terms of, you know, thinking about lower no numbers and slow moving fluids and then fine tuning on like faster moving fluids. So that way you kind of reduce the requirements of data when it comes to the more complex physics that the numerical solvers need to generate. Um We also have like a recent work that will be releasing soon where we'll show that this approach is better than you know, trying to do closure modeling and other kinds of like traditional approaches, people do where they still keep a core skills, co solver and only use machine learning to refine that. Um Whereas this progressive approach just kind of
completely gets rid of any, let's say traditional solver and learns from data in a way that it takes it all the way to even getting to the complex. And along the way, you can always also do hybrid modeling where you add physics losses along with data. And of course, there is an art to it because you don't want to make the optimization landscape too difficult, right? So you need to kind of be much more thoughtful of the algorithms to make this work. And it's not as straightforward as text models where a lot of data is available. And there isn't a notion of curriculum, you just kind of take in all of that data and you just uh train the model.
Uh whereas in this case, it's more nuanced, but at the same time, there's an opportunity here to make it work other than just the brute force standard approach of like ingesting all of the data because it's just too expensive to do that. Hm. So where do you um I think what you're alluding to is this idea some would call and I'd be interested to know what you think about this more like a foundational models which you start to give it enough data from maybe other tasks that it is able to predict something quite, it's quite generalisable. Do you think that how, how, how broad can it truly be? How ambitious do you think we could, we could be in this area?
I mean, to me, the sky is the limit and just as we speak, my students are presenting at ICML. Uh you know, I don't know if anyone lives there right now. I know I in Vienna. But uh anyway, so they're presenting it right as we speak or around this time where we've created a, let's say, a GPT-2-sized model that is able to learn on multiple families of different partial differential equations. And also showing that one model can broadly learn across these domains rather than narrow models trained only on those narrow data sets. So you have this cross domain learning that gets that benefits from having that approach of learning multiple phenomena at the same time.
And do you think um one of the uh statements I've often heard um from, from various people is this idea of scale that um actually for a lot of the large language models, there's been a focus on any architecture that could scale for huge amounts of data, maybe at the expense of accuracy. I I would say I'd be interested to know you think, do you think the method we have now can truly scale because I've seen limited examples in academia or industry so far where it's at the scale of, you know, hundreds of millions of points of grids or, or thousands of different cases, is it just a matter of time to do it?
It's just a matter of time and resources. If you see our model, you know, there are two kind of aspects, right? You can think of like the weather model, which, like, the largest one is, let's say, a little less than a GPT-2-sized model is working very well for the narrow domain. And now we have a broader foundation model that is getting towards universal understanding of physics. But that's only GPT-2-sized, meaning it can't possibly have the ability to solve very complex three dimensional tasks, right. So it's kind of like able to do flow dynamics in a certain regime, but it's a limited regime because the model has only limited capacity.
But with this, what we are able to demonstrate is it can be scaled further and neural operators as a foundation can just work across resolutions in the same model. So you can have like different PDEs be queried in different grids, different resolutions and just one model can handle all of that, different geometries, all of that. So where does this bring to the future of traditional simulation methods? Traditional, you know, find it different, find an element in commercial or open source codes. So to me like it's there'll be a time where it won't be one versus the other. So like I mentioned AI plus science. It really be the ideas from these
numerical simulations would be deeply integrated into AI, right? Not just the solvers themselves in a black box, which I don't think it is a good idea for a variety of reasons. It's not a lot of researchers attempt to do it. I would just think that in future that's not a winning path for, you know, whole set of reasons that I'm happy to get into. Uh but it's really gonna be the aspects of how do we take the best of all the algorithmic ideas that have been developed for decades in numerical solvers and integrate it with AI algorithms together. And in fact, that was also our inspiration when we came up with neural operators,
right? This idea that numerical solvers are able to solve um on different grids, you can query any point in the domain. It need not be on a fixed grid or a fixed resolution. But neural operator before we invented neural operators, the neural networks that were standard didn't have that ability, right? They were always on a fixed resolution. So how do we bridge that gap? And how do we think about, like, numerical solvers like pseudospectral methods and then make them learn from data. So instead of a rigid like aspect of going between the frequency domain and the standard domain, we also have learnable parameters and nonlinear transformations in between.
And that came resulted in the Fourier neural operator. So we are continuing to take all kinds of inspirations from traditional methods to strengthen these methods. And to me, I think that is a better approach because by definition, that would be taking the best of everything. And so it wouldn't be just purely A I numerical methods deeply integrated into those ideas. So do you think that to um to achieve this goal? And I'm certainly interested to read that paper, maybe by the time this comes out, I can put a link to it, so maybe we can have a chat afterwards what you're allowed to share at the moment? Um Is it therefore a need for us
to somehow bring together the community to generate this training data? Because I guess a lot of this assumes that there is a broad enough set of problems and data to create these models or do you think this is going to be different than your large language models? Because people are not going to want to essentially release uh their data? Well, you know, in my view, so many of the solvers are open source, right? I mean, sure there are closed source solvers and there are some details of which one is that are capable of for a set of domain. But there's a lot that we can generate with open source solvers. And that's what we've done now recently and we've released many PD benchmarks,
other labs have done that and in our recent ICML paper, it's really the largest collection, right? We took all of the available PD benchmarks and trained a model on it. And to me, I think it's a continuing journey, I'm aware of different groups are gonna further working and are gonna release bigger data sets. So this is an exciting time. OK. That's good. And, and how about other domains? So we, you know, weather is obviously one. Um you were you were talking about, you know, some of the fluid dynamics, but are there any other areas that particularly excite you in terms of the potential of AI plus science? Yeah, I mean to me like if you look at partial differential equation,
there's so many different areas, right? So to give you some examples where there have been already success stories, uh nuclear fusion is one of them, we worked with UK Atomic Energy Agency and created these AI-based simulations of the tokamak— you know, how plasma evolution occurs in these nuclear fusion reactors. And how can we predict disruptions? Meaning when the plasma may escape confinement and could potentially damage the reactor, which you don't want it to happen and to be able to do that these AI models are should be very fast. In fact, they are faster than real time, right? That's why we can take corrective action.
On the other hand, if you try to run traditional magnetohydrodynamic simulations, which is what describes this plasma evolution that would be extremely slow. Uh In fact, our methods are a million times faster than what traditional simulations can do. And so that shows what we can kind of like now think about using these models from physics in applications that you earlier couldn't even conceive of. Right, like plasma is one control of drone is another, you wouldn't dream of putting a CFD solver on a drone because, first of all, it's kind of going to take a lot of time and energy to run anything and you know, and it's just going to be too slow and by the time the drone would have crashed.
But now we are in a place where we have used machine learning techniques that can help make the drone flights be better and safer. And so in all of these areas, speed is important, right? The cost of uh coming up with these predictions is important. The other aspect is going back to not just simulations but the aspect of design and discovery itself. So if you recall, I was mentioning the scientific method of like going back and forth between ideas and lab testing and the lab testing is the critical component. So it's important to come up with ideas that are already physically valid. And hopefully, you know, there's an internal simulation that certifies the model
to be the design to be correct. Otherwise you will spend a lot of time with going back and forth with lab testing. And we did that recently with the medical catheter where, uh we designed a better medical catheter than was previously available. Uh, so I'm not sure many of the listeners here know about the problems with medical catheter. Uh, it's a tube that takes fluids out of the human body. Very simple, but it's one of the, the most common cases of healthcare related infections, more than half a million cases just here in the US annually. And so, you know, this is a problem that is in fact described by physics pretty well.
Meaning bacteria tend to swim upstream near the wall of the pipe, right, where the slow fluid, the fluid outflow is slower and hence swim into the body and infect the human. And now we our collaborators in fluid dynamics had a simple idea. They were like, let's create like ridges, these triangular kind of shapes inside the wall. And with that, you can create vortices. So the bacteria doesn't have this ease of just swimming upstream and going into the body. But of course, the question is what is the optimal design for those shapes that best reduces bacterial contamination. And with our neural operator based model, because it's an AI model,
it's differentiable, meaning we can directly have gradients for improving our design and come up with an optimal design. So our AI model already proposed an optimal design and then we had to go to the lab and 3D print it just once and it resulted in 100 fold reduction in bacterial contamination. So, you know, think of like a future where AI imagines all kinds of new back designs, but it's not just hallucinations, it's not just a creative concept, it's physically grounded and valid and then we can actually bring it to the real world and show the impact in the real world. That's the future. That's very exciting to me.
Yeah, that, that, that is super exciting. And, and I think um it's probably the one that in, in some ways industry is most, not most interested, but I think the generative design I think is one, you know, most, whether it's an aircraft, a wind turbine, a fridge or whatever, you're trying to come up with a better design normally and simulation is just a tool to go that way. Um And I think it's there have been, you know, adjoint methods and, and other sort of methods that, that try to drive, but I think it's still not still not realized that dream of sort of a computer doing it for you. You know, there's always a very heavy human in the loop, et cetera.
And I think what you're alluding to is that potentially AI may help that come closer to reality because of the speed of the simulation. But one thing a previous guest mentioned to me, and it really got my sort of my thinking is in my head, I've split transformer large language models and you know, uh neural operators or PINNs or, or graph neural nets as if they're sort of two separate things. But how do you see because humans do ultimately like to verbalize something or the sort of prompts that we ask larger language models? How do you see the potential of integrating these together? Yeah. So first of all, I want to kind of make a clarification
that transformers are not separate from neural operators, right? Neural operators are the super class where you can now have transformers that work at all resolutions. That would be a neural operator. And in fact, if you one simple way to think of an example of that is you know, instead of the attention mechanism working on fixed tokens, you can work it in the four year space and like you know, go back to the standard domain with inverse Fourier transform. So if you do that now you can extend it to multiple resolutions, which essentially is a Fourier neural operator. But it's also a transformer if you parameters your attention with the four
Fourier transforms. So there's a lot of interesting math connections that mean that you know, they are not as separate and different as people may think it to be and neural operators are not as exotic as people may think it to be, right. So there's a lot of interesting nice connections. Another example is also like graph based methods also being operators when you have the flexibility to add new nodes to the graph and be able to adapt uh with new nodes, what are the new edges? And you can do that with like spatially defined graphs very easily. Uh So there's all these standard methods that you think of as quite separate that are really just special cases of neural operators.
So it's really that broader concept that we are describing. And uh now the question is can text be an interesting modality to me, like you can add all these modalities on top of a physics based model where there is physical understanding, right? I mean, text is a way for humans to interact. If you want to design an aircraft wing or a drone, you can kind of have a chat if you think that is the best interface. I mean of course, it depends on some certain people preferring that versus the other kinds of interfaces, right? So language can be one modality, there can be other modalities like looking at uh right, observational data from video feeds and so on and
all kinds of other aspects of multimodal model. But to me the foundation of all this is physical understanding. So you can build this and on top of it comes the other modalities, I guess what I'm I just want to see if you think this is purely fiction or is actually something that could happen, which is at the moment, you go to a, you know, various large language models. Um and you say create me a picture of a plane flying OV over Mars and it will create you something depending on how well you do the prompt, you know, pretty nice, pretty accurate. Now even more, create me a a movie of, of a plane going through. What about the reality of saying,
create me a design of an aircraft that can efficiently and then it goes off, creates a design, uses the model to go and simulate it. There's 1000 of those simulations come back and gives you the design that you go in 3d print. How much is that sort of science fiction? And how much could that actually happen? I mean, in fact, it's already happening, right? That's what we did with the medical catheter in a short smaller scale, I think to me, text is not the important aspect. Uh text is almost a distraction because if you can nicely describe and in fact, for, you know, design considerations, you want more clearly describe your design cost functions,
right? That you can go optimize with our neural operator foundation model as a way to get to those design goals. And sure you can try to specify it through text. But to me, like with text, you must still kind of end up making mistakes of translating it into the right design course. So if you could directly do it, I would just start with that because that is a much more cleanly defined problem. Yeah. Yeah. No, you, you, you're probably right. I think so. I think it's more the, um, the conceptual design, you know, the, the, the sort of market who just wants to, uh, very quickly go through designs, which I guess has been the dream of real time simulation, hasn't it,
this idea that a person in design studio? But, yeah, you're probably right. In reality, an engineer wouldn't just type it because it probably wouldn't be precise enough. Um You would, you would want to interface with, you know, some tool. Um So how what needs to be done then for this to happen at an industrial scale. So more practically now, where do you see the role of academia of industry of government funding? What's, what's, how is it now? And where would you like to see it go to really make this a reality? I mean, we certainly need more resources is the short answer, right? There's a lot of kind of attention paid to
language models and also now recently robotics, but you know, robotics is still kind of like uh data hungry and their data generation is even a harder problem than I would say with physics based simulations because you have the simulator, you can generate as much data as you want in our case. Uh But of course, the main constraint is the getting the resources to do that at scale. And you know, in various kind of aspects, we continue to scale these models and other teams are doing that. I think the question is how to get us to that next level where we can train it on much larger clusters and really get this to a place where
we can show those benefits of scale of being able to do cross domain models that can understand multiple physical phenomena as well as what we call multi physics stability to couple different partial differential equations together. Mm Yeah, that always seems like the holy grail that even with physics based simulations, it's still never really fully done. People do tend to be siloed, don't they into a fluids or structure or climate? Even though technically they are kind of connected at an engineering um level. Is that really your vision that there is—it's easier in a way to integrate it through an AI model rather than
these more traditional solvers. Yes, precisely. And, and to me that's the only fact to making also these AI models realizable and practical, right? Because you need to have the shared data from all of these different use cases to help learn a good representation. Otherwise the data needs would be enormous for narrow surrogates. So he here's a maybe a bit of a different question for you. But um you, you started doing your research and your academic career and you were arguably in the early days of ML where you were, you were moving into an incredibly exciting field that, you know, you've made massive contributions to,
if you were now giving advice to a PhD student or maybe an early postdoc, what, what area of, let's say, machine learning would you focus on bearing in mind? You know, the, the time it takes to, to sort of progress, are there certain areas that you see as being very promising that would be good for somebody to get into? Um Yeah, thanks Neil, I guess to me like, you know, there's the question I get a lot because also people are like, how do I distinguish myself? There's a million other people working on language models and how do I not have enough resources compared to people in the industry? Right? And it's not a level playing field. And to me, I think the main aspect is
doing a PhD that means you should be working on problems that most others are not paying their attention. And that could be anything, right? I mean, in my case in my lab, that would be AI plus science. And you know, not just uh you know, we're doing neural operators, but we're also looking at the fundamentals of learning itself. Like what are the optimization challenges and all the way to like problems in biology? You know how to train models for protein design for um being able to also generate new genome sequences of viruses and bacteria being able to like do drug discovery more efficiently. So there's all these different areas, right?
And there's application areas and then there's the fundamental problems within them and how do you formulate and make impact on those areas? And there's still such a green field. And to me, Caltech is an ideal place where because of its small size, it's highly interdisciplinary. And if you look at almost any of our pre papers in these different domains, they have people from those domains, right? So it's a very interdisciplinary collaboration and that means we are also tackling the right problem. We aren't just doing machine learning blindly putting a hammer, but looking at metrics that are relevant to the field.
And that's also what I think makes it difficult for many people because in machine learning L A there's one objective function, you optimize it, you're golden, but that's not how science works. Even if the weather model wasn't just like root mean square runner or just some metric, one metric, right? It's like, oh how does it do one extreme weather events of all kinds of things? Right? It's like all these different aspects of like, oh what happens if you keep running the model for a longer time? Is it stable? So you have to look at more than just one objective function. And that to me is a rich set of even foundational
questions to be answered. How do you formulate those kind of multi objective machine learning. How do you kind of prioritize different objectives because the optimization landscape becomes too difficult. And of course, in the, even in those application areas, there's a lot of deep thinking and working closely with the domain scientists to understand what matters for them. Um And so my advice would be like this is not something that's easily done in industry. Sure, in certain areas, there are industrial groups working on it. But you know, there's always some such broad areas of science and engineering that you know, academia has an upper hand.
And so to identify that uh uh but you know, similar argument could be also made for social sciences, for economics, right? Interdisciplinary problems are where I think you will be able to make unique contributions and have a Yeah, that's really uh that's, that's really good advice. And another because I like to make things quite practical. Where would, where would you recommend people go to stay abreast? Is it the standard? You know, ICLR, ICML, NeurIPS—are there other places that you're seeing emerging as where people should go to hear some of these newer ideas at? Yeah, very practical level conferences, symposiums.
Where would you recommend people go to, you know, to me, social media is not something I would recommend. Once in a while, it's still you may come up with something useful but days of course, a whole other conversation, it's too distracting for a variety of reasons. And so that leaves us with various of these academic conferences, right? So I, I think they are just still a great place to go. But the main events are quite crowded. We have, like, now NeurIPS with tens of thousands of people. So I would highly advise going to workshops, going to meetups, like smaller groups. And you know, it also not be very popular events worldwide.
It may be something local, you know, in New York City, different meetups, different uh just small workshops. So focus on attending the smaller events because that's when you can have real conversations and just, you know, get to know people who are working on related problems and what in their experience didn't work. I think that's the aspect that is very hard for people to get in an open forum like social media and you really have to go talk to people. And so I advise my students to just try and attend a lot of events locally here in California because they learn a lot from that. And what about um one of the things that sometimes come up in some of these workshops
is this idea of bringing the communities together, the domain specialist with the sort of ML specialist, do you think there is a need for the ML community to almost venture out from their comfort zone of NeurIPS and ICML and actually go to engineering conference or, or go to a, a Biology conference. I mean, I'm sure that is being done. But I, I was amazed when I went to, uh, ICLR so it was a huge number of people that then when you go to the big aerospace conferences or engineering, you see very little of, of those people yet they're both kind of working on the same problem. Do you think there are ways that we can improve this situation?
Yeah. No, I'm smiling because this was kind of my goal when I came here to Caltech, right? In 2017, we started the AI and Science initiative here. And that was one of the first across any campus to have that kind of an initiative to bring together people from many areas and really debate how AI can help them. You know, they may think, oh AI is not already, I don't need to worry about it and then there may be, you know, but we need to also kind of develop methods where it's currently not working. Like currently just doing a hammer approach, they didn't get a machine learning to work, doesn't mean it's not an interesting problem, right?
To me, those precisely are the interesting problems. So we created that initiative, we had regular workshops, we continue to have them. In fact, now we are doing it jointly with the University of Chicago. Thanks to a generous gift by the Pritzker Foundation. So, you know, the aspect I think is maybe the traditional venues are not so great for it. It would be my reaction to that because they're just too big, right? If you're going as a machine learning person, especially a student to this enormous event and you can't really make sense of like what aspects of it is relevant to you as a machine learning person. So you really have to go to talk to people who are invested in making that happen.
You know, as you said, there's a lot of skepticism in the field I have kind of, you know, face, let's say different kinds of personalities, you to pick people who are willing to disrupt their own methods. And here at Celtic, I'm lucky to kind of have collaborators in that realm who really are, you know, joint inventors of neural operators because they were willing to take this leap and say all they worked on numerical methods. Let me see what AI can do. And that's I think harder to combine a big conference. I think they should just talk to their colleagues in other areas, just grab a coffee, grab a drink, keep it casual, just go and learn and survey a lot of problems.
You know, I was lucky to do that because when I came into Caltech in 2017, I started this AI and science initiative. But I also had this generous gift from AWS to, you know, thanks to AWS have cloud credits and distribute that across campus. So I invited proposals and just learned everywhere. What are people doing, you know, how are they going to use the cloud, not just machine learning, right? All kinds of computational methods on the cloud. And you know that jump started a number of projects and collaborations. It also led people to thinking about using the cloud and machine learning methods more because the resources were available.
Um So I think we need to think of more creative initiatives like that at a smaller scale, you know, and then, you know, the question of course, how to go from there and do much bigger events. But at least at that point, that was a great way to jump start this initiative. Yeah, it definitely seems um the only way to truly convince a community is to work with them. And I, and I guess, you know, be, be inside of it rather than a than an outsider. I think people are always scared of change. And I think um it will be interesting to see and that, that's kind of what I was saying, you know, what, how close are we for things being ready? I, I guess the um
this is to some extent, which is why I was asking about the solvers in things like fluid. And I think weather is a bit different because I think weather codes are typically developed by the weather centers and, and so they have, they're in sort of control of their own destiny to a certain extent. Um Whereas I think it's, it's really interesting in the fluid dynamics or computational fluid dynamics community, because there is such a reliance on large commercial codes that in some ways there's a, there's a need for them to get on board to expose that if you know what I mean, like academia can do what it does, but you kind of need
those big powerhouse companies to sort of embrace it for it then for all these smaller companies to, to take advantage of it. Do you know what I mean? That there's, there's a, there's only a certain amount of academia you can do when you need these sort of software houses to, to get involved. Have, have you seen that or are you more optimistic? Now? I'm certainly more optimistic. Of course, I, I come from a different realm. Right. So academia has done wonders. So I wouldn't say that it's not possible. I mean, we continue to open source a lot of what we do and I'm also doing to, I think sometimes disruption happens from outside then with them.
Yeah. Yeah. Yeah. You know, maybe one way or the other but in the end progress will happen, which is great. Yeah. Yeah. Well, II, I would love to talk for hours and hours but, uh, I know you're a very busy person so I'll, uh, I'll make sure not to take too much of your time, but I wanted to really thank you for, for talking about this today. I I'm gonna share a lot of links because I think um I appreciate you stayed high level, you know, II I asked you to do that, so I appreciate it. But um I know there's a lot of content and papers that people could jump in to really dive into the work that your group's been doing.
So if people are watching this and you have a look, I'm gonna put a ton of links down there that you can dive into some of these papers and of course, watch the, the TED talk that I, we were discussing. Um, before we started that the only problem with doing a podcast and you talk about complex sciences, it's kind of hard to explain things without sort of illustrations and things. So I think your, your TED talk actually does a very good job of doing that. Thank you, Neil. No, this is so much fun and kind of took in all kinds of interesting directions. So I'm really glad that you're doing this and getting attention to this area.
You know, it's still a bit more niche, right? Compared to language models. But hopefully we can all together, bring more attention to this. Exactly. Thank you. Yes.