The Neil Ashton Podcast

Foundational AI Models for Fluids

Season 2, episode 11 00:22:33

Foundational AI Models for Fluids — The Neil Ashton Podcast

Spotify player

Foundational AI Models for Fluids

Spotify is contacted only after you load the player and may then set cookies.

Open in Spotify

Episode overview

In this episode of the Neil Ashton podcast, the discussion revolves around foundational models in fluid dynamics, particularly in the context of computational fluid dynamics (CFD). Neil shares insights from a recent panel discussion and explores the potential of AI in predicting fluid behavior. He discusses the evolution of AI in CFD, the challenges of data availability, and the differing adoption rates between industries.

The episode concludes with predictions about the future of foundational models and their impact on the engineering landscape.

Chapters

  1. 00:00 Introduction to the Podcast and Topic
  2. 01:09 Foundational Models in Fluid Dynamics
  3. 10:09 The Evolution of AI in CFD
  4. 19:52 Future Predictions and Industry Dynamics

Transcript

This transcript was generated by Spotify and may contain errors. Download the original SRT file.

0:00 Hi, and welcome to the Neil Ashton Podcast. In each episode, we explained some of the fascinating ways that science and engineering are changing the world around us. We talked to leading engineers from elite level sports like cycling in Formula One to some of the world's top academics to understand how fluid dynamics, machine learning, supercomputing are bringing in a new era discovery. We also hear some of their life stories, their career advice, and lessons they've learned on the way that I hope will be helpful to you too. So sit back and enjoy this episode. Hi, and welcome back to the Neil Ashton Podcast. So today I wanted to talk about

0:43 the topic of foundational models for fluids or for computational fluid dynamics. It's a topic that I brought up and asked quite a few of the people that I interviewed recently. But I, I wanted to give some personal thoughts on this, partially because I had the pleasure a couple of weeks ago of chairing a panel discussion at the second data-driven fluid dynamics conference, which was on a cough tac euromech thing that was done at Imperial Carter London. I was involved in the first one in Paris the previous year. And it's a great initiative that has tried to bring the fluid community and I guess the AI community together, particularly with a sort of European flavour,

1:28 I guess to it. But in this panel session, it was really interesting because we had representatives from, you know, Harvard, but also Caltech and, and some from industry and, and really representing the experimental and the computational side. And one of the things that I did was to try and ask people about their opinion of foundational models. But more interestingly or as interesting was I tried to ask the audience some of the questions and there was a couple of 100 people in the audience. And so we use this Menti meter app that I have to confess I haven't really used before, but it's great because you can essentially do real time polling of people.

2:18 So I asked the question, do you believe that we will have foundational models? And in fact, I'm just looking on the on the screen now because I have the results up. Well, I should tell you, the first thing I did ask was what's your favorite British food? Bit of a tongue in chic really. And people for fish and chips won, followed secondly by the vomit emoji, which I was a bit offended by. And then a full English breakfast. So there we go. But more seriously, on the actual topic of foundational models, I basically said, will you know, can we get foundational morphin fluids? How generalisable could it be? And what I asked was, could it?

3:01 Could this foundational model provide instantaneous pressure, velocity for any boundary conditional geometry? So as in you can say to it here's a geometry with a certain boundary condition, can it predict it and give you the instantaneous pressure and velocity? A bit like ACFD solver. Now you give it a geometry, you give it some boundary conditions, and it would go and solve it. Second option I gave people is exactly the same, but time averaged. So not instantaneous, but time averaged, and I'll get into the nuance of that in a little bit. Then the third option was yes, but only for a specific use case, as in maybe like a road,

3:40 car or a plane. And finally, option was essentially no physics. Sorry fluids are just too complex for an AI model to learn. And interestingly the out of the scores it was 38 and this is out of 100 and something. So the the the most popular answer was only for a specific use case. 38 people close 2nd. 37 was provide time average to any bound addition and geometry followed by instantaneous pressure. And finally 13 had fluids that are just too complex. So out of that audience, which is biased because they are researchers looking at fluids and AI, the sense was it really is only possible for a specific use case. And therefore the terminology of foundational model is perhaps

4:31 pushing it a little bit or needs redefining. OK. So that's why essentially I wanted to to talk about it and I wanted to share I guess some opinions of my own, some stuff I've seen along the way. I'm in a lucky position, I guess. I'm working in the tech sector that I get to engage with a lot of customers and academics from different groups. So I'd like to and, and go to conferences. And I'd like to say I'd therefore have a, a reasonable good view of what's going on in the community. And I would actually say that things are looking very interesting as in only a couple of years ago there was really no work into this idea of a foundational model.

5:17 And another way of thinking of foundational model is simply a pre train model. So what most people were looking at was simply can we come up with some sort of algorithm, some architecture that could allow somebody to go and train an AI model based on prior simulation data. So we're not talking about accelerating an existing code, we're talking about replacing the code with a surrogate model. And that was certainly the main focus probably five years ago or four years ago. And there was some Seminole work done with things like mesh, graphnet and others which took a graph neural net approach where you could essentially take the

5:51 nodes of the mesh as the input. And that was something or particles I guess because it was all particle examples of this. And this rapidly became one of the success stories in the weather and climate, but also was showing potential for CFD. And this led to a number of start-ups that that formed over the time, the sort of Navastos, Nora concepts and others who took some of these work or pioneered things themselves and, and released some of those things that people could go and, and try for themselves. And so you saw in the past, you know, several years, people really wanting to test out those methods and with some success for sure.

6:36 But obviously that requires you to have your own training data, your own ability to train a model and to have, I guess, both the software and the knowledge to do it. And I guess if you can compare that against what you would typically see in the LLM space where maturity of us are not training our own LLMS, but we are simply doing a prompt and asking a question, doing an inference, we're not training it. The logical question was, well, could it be possible to train a model and then just provide it to people as a pre trained model? And it's a really interesting discussion because it has pretty large it, it could have a potentially very large impact on

7:28 the CFT community. Now, there are some of you out there who still have very negative opinions of AI and sort of C as a fad, and it will come and go. But I would say those voices are slowly starting to be overtaken by the realization that actually these methods have a lot of potential. Only, like I said, five years ago with this was basically not happening. Nobody was even talking about it. I remember I got involved maybe three years ago in this topic, so still quite late to the game I guess. But even three years ago I remember most people had no idea what this was all about. They just thought of this as like reduced order modelling ROMs, real, no sense of what was

8:15 going on, didn't know what architecture would work. There was only maybe one or two start-ups in the space, wasn't very prominent. Fast forward to now and I would say that it is certainly rapidly expanding and I would say companies can be categorised as either trying to lead in this way, being very interested and the number who were just set against it is now sort of getting lower. But I would say that this is still a little bit split between the industries. And like with there's a bit of a parallel with high fidelity methods, I see the automotive sector as being much faster to adopt these new technologies or investigate these new technologies.

9:03 Automotive companies today use sort of hybrid around scale resolving methods on a day-to-day basis, whereas the aerospace sector is still sort of proving out that they think they can be used in in production. And some of that is because Reynolds numbers and complexities. Another is just a mindset, I believe. And interestingly, I see the same thing with the AII, see that the automotive companies are far more motivated to try and find AI approaches that can be faster potentially and are embracing it. Whereas I see aerospace is more in the mindset of I don't think this can work for me. And I guess part of me, the reason doing this podcast episode is really just to put

9:43 across that, and this is a genuine, this is not influenced by the company I work for. I've I've these all sort of independent opinions, personal thoughts, but I really do think that these are looking incredibly interesting for the obvious reason that, you know, now if you do an inference on a, on a model, you're going to get an answer in seconds. But the accuracy is debatable whether people have truly proved out that these these AI methods are accurate enough. They give a very good impression of, you know, approximately what the flow field is or approximately what the lift or drag is. But I would say that we are still not at the case where they're widely accepted or there

10:26 are very clear accuracy. I would say it's a bit like CFD in the early days where you could argue there was clear points where it worked well, but there were many examples where it didn't work. And hence why CFD took decades, you know, to to progress and be trusted enough to be a dominant design tool. Sai is still in that sense where it's probably not trusted as a key design tool yet. But I would say that the number of companies really investing in this is significantly growing. The number of start-ups who are emerging in this space is growing month on month. And the interest from the big is VS is also growing month on

11:08 month. But the challenge is the data. So do you have available data? And this has always been the challenge for the idea of a foundational model, as probably people know, you know, the reason that these LMS and others can do so well is that they have a very large data source to, to to, to get to. So typically scraping off the Internet, but nowadays maybe doing commercial deals with a certain repository of data. You know, if you want to do something on journal papers, you might do a deal with Elsevier or John Wiley. If you want to do something on video, you might do a deal with a video, you know, distributor or or a newspaper. If you want to do on text or

11:54 pictures, you might do with Adobe. You know, there's a lot of commercial arrangements going on now. And I think that's what's super interesting on the CFD side because on one side you could say the data is just not available. If you want to do combustion modelling or multi phase or hypersonic something. This is not publicly available data. So how do you you know, how do you solve that? And this is where I think there's an interesting technical but also commercial angle to this, which is, is it worth it for a company to independently run lots and lots of simulations to the fact that they can pre train a model and then supply the model and, and and charge it

12:41 at a certain rate. And we have seen some early signs of this happening. You know, there's a start up luminary cloud. Some people know that I spoke to to Juan Alonso, the founder, and they recently released a sort of foundational model for Rd. cars where they'd, you know, run a whole bunch of simulations themselves in collaboration with an end with an end customer of theirs and and then released that and pre trained the model. Now that it's probably too early to see, you know, how ultimately successful that that that will be, But it's it's they've sort of fired the starting gun, so to speak, on an actual company releasing a sort of pre trained

13:23 model. And I think it's really interesting to observe how useful that is because on one side, it removes a massive barrier for a company to to collect all the day, to find all the day to have the expertise to use an AI tool. It's much easier just to do inference on a pre trade model. So I think that's the first thing I predict that we're going to see many more of those come out. Many more companies will will, will do that both I think from a start up and a nice fee space. But how, how much data can a company generate and how can they incentivize it? Well, one interesting thing, when I spoke with Priff, who's the CTO of ANSYS, he said in one

14:10 of the talks that I gave that, you know, maybe the ISV needs to incentivize people to allow the ISV to have the data. So I think that's a really interesting proposition that, you know, what if you would take a box that was to say, well, you know, as you run your CFD simulation, I give permission for the ISV to use that data and perhaps in return they get, you know, some commercial incentive to do it. And I think that's a really interesting proposition. And I'm, I'm, I'm keen to see if some of the Isvs go down that route. And this actually links why I've often spoken about the cloud being an important Ave. and the

14:50 cloud being an important Ave. links to this because of the SAS bit. If you give somebody a on Prem binary, even if they tick some box to say we're happy for you to look at our data, what's the mechanism to share that file? Well, it's pretty difficult because it's running on their local network. How are they going to, you know, send that over to you? It's not practical where with a SAS solution, it's much easier because it's by definition running in, let's say someone's cloud. And so if you say I want to share some of my data with them, it's actually much easier to do it. So I predict that part of the motivation and I think we'll see

15:30 an acceleration of the sassification is to make this sort of AI and data collection easier both. If you're an existing AI start up, you know, you could say, hey, use my AI tool and I'll give you a discount if you share your training data with me and over time will collect more data. And also it links into one of the things I've said repeatedly on this podcast about fast CFD solvers or CAE solvers. Because now one of the big things is if you've got to generate 5000 CFD cases, if your CFD solver is twice as fast as another CFD solver, that's a big amount of money to save. Or for the same budget, you could run twice as many cases. So if we assume that people to

16:20 build these foundational models will have to run hundreds of thousands of cases, then actually there will be a huge focus on on enabling the code to be as efficient as possible. Now, one of the other things that I haven't brought up, but I think is another interesting one is everything I've spoken about now assumes that the CFD code is independent from the AI code. But as anybody knows anything about, for example, in situ visualisation, we'll know that there's a strong move now towards this idea of doing it, you know, online rather than saving everything to disk and then doing the visualisation. Can you do it whilst it's running?

17:01 And this is obviously something that's not just me, you know, saying this for the first time, it's known that I think there'll be a big increase in the Ori is certainly in the academic world of trying to build your CFD code in the same framework as your AI code. Some groups are doing this in Pytorch, you know, how do I write a code in there? Or they're using some sort of Python code that you can automatically, you know, differentiate and move between to pass gradients along and to the neural networks. But I predict this will be a will be important because if you are trying to build some pre trained model, some sort of

17:35 foundational model, you don't really want to be doing your CFD independently to your AI training. You want them to be tightly coupled and potentially adding more points as you need them and not having to be, you know, bound by IO issues, having to keep the data. Which is particularly true if you have the dream, and I think this is a much longer term dream of being able to do instantaneous. So time dependent, most training now is done on time average solutions, even if it's a time accurate simulation, because just the amount of data, I mean having created with colleagues, you know the driver ML data set, it's already 30 terabytes with one time step essentially as in

18:20 the time averaging, if you were going to try and do it for all of them, 200,000 iterations, well, you can do the maths, it's unbelievably large. But if at every time we were running those simulations, we were passing that data to a model to train, then the cost we we don't need to save the time steps out. We're just passing it in memory. So we actually could have done it because we were running the simulations time after time dependent anyway. So I think this is where there's a lot of interest in these. I feel this is where new technologies are converging with AI as the sort of motivating factor. So foundational models. I, I really would encourage any

19:03 academic who's listening to this, I think this is the topic to propose as a, you know, a big university project or European project or government project or, you know, anything that is big. Because if this could be made and there's so many debates on should be open source, should be closed source. It could transform the way that we are doing CFD today. If it is possible, it would change the commercial landscape, it would change the technical landscape. And if you imagine that it's possible for a certain class of applications, you've got to imagine the transfer learning, the fine tuning. You know, at some point how different is a car than a plane

19:46 or a city if you are starting with, if the model is able to learn some of these interactions. I haven't really spoken about the idea of including physics into it, but I I really do feel that just as five years ago AI was an interesting thing for people to look at, I think the whole concept of foundational models is becoming interesting. I believe at the beginning it will be for Pacific use cases as the audience voted, but I perceive that there will be a bit of an arms race between the highest fees and startups to build this, and it'll be interesting to see whether companies see this as their secret sauce. Whether you know, an aerospace manufacturer or an automotive

20:31 manufacturer says, well, I don't need the ISV, I'm going to do it myself. And this will be an interesting balance between the software suppliers, the companies, because on the other hand, the automotive company can say, well, we're not a software, we're a car designer. Why are we going to write our own codes? And that's certainly been the case even in the aerospace. There's a move towards commercial codes. Well, you know, it used to be the case, everybody would write their own codes and that was their IP. And then they realised their IP is making cars or planes, not writing software. So it, I still suspect that most

21:04 of this work, these foundations will still make their way into commercial sort of big software companies IS VS. But given that they don't have the data, it's the sort of engineering companies who have the data and they need to somehow get that data or produce that data. There'll be an interesting dynamic, I think in collaboration. So these are just my personal thoughts. I could be completely wrong, but I am quite bullish on the idea of some of these foundational models. And, and certainly, you know, in my day job, this is something I'm actively pursuing. I'm I'm academically interested. I'm interested in, you know,

21:44 NVIDIA doing its bit to help things along the way. And yeah, I'll be really interested to see what happens in a few years time. So maybe I I'll try and set a reminder in two years to record another one and see how much of this. I was right on and maybe it was it'll all not happen and I was completely wrong. But I'm, I'm going to make a bet that in a couple of years we will have progressed quite a bit further than than than we have done to date. So with that, thanks for thanks for listening. I would really enjoy your comments. Let me know what you think if I'm completely wrong, if you have a different viewpoint on

22:20 it, please put your comments in the YouTube or, or send me a message. And yeah, thanks for listening and hope you enjoyed it.