The Neil Ashton Podcast

Fluid Intelligence with Johannes Brandstetter and Siddhartha Mishra

Season 3, episode 9 01:24:54

Fluid Intelligence with Johannes Brandstetter and Siddhartha Mishra — The Neil Ashton Podcast

Watch on YouTube

Fluid Intelligence with Johannes Brandstetter and Siddhartha Mishra

Open on YouTube

YouTube video

Watch this episode

YouTube is contacted only after you choose to play the video, keeping this page fast and private by default.

Listen to the audio

Spotify player

Fluid Intelligence with Johannes Brandstetter and Siddhartha Mishra

Spotify is contacted only after you load the player and may then set cookies.

Open in Spotify

Episode overview

In this conversation, Neil Ashton and Prof. Siddhartha Mishra, and Prof. Johannes Brandstetter discuss their recent paper on AI foundation models in computational fluid dynamics (CFD).

They explore the backgrounds of the speakers, the journey to writing the paper, the role of AI in CFD, and the challenges of scaling laws and data generation. The discussion also covers model training costs, open questions, and future directions for research in this field.

References and links

Transcript

This transcript was created from the corrected YouTube captions, with names and technical terminology reviewed. Download the corrected SRT file.

0:00 Hi and welcome to the Neil Ashton Podcast. In each episode, we explain some of the fascinating ways that science and engineering are changing the world around us. We talk to leading engineers from elite level sports like cycling and Formula One, to some of the world's top academics to understand how fluid dynamics, machine learning, and supercomputing are bringing in a new era of discovery. We also hear some of their life stories, their career advice, and lessons they've learned on the way. That I hope will be helpful to you too. So sit back and enjoy this episode. Welcome back to the Neil Ashton Podcast. So this is, um, I guess a special episode, um, that I did because I've just published a paper with Sid and Johannes.

0:52 They'll give their full intros later. So I want to keep it a bit short, which is called Fluid Intelligence: A Forward Look on AI Foundation Models in Computational Fluid Dynamics. Bit of a long title. I had to look exactly what we call the title now, so I didn't say it wrong. But it's, it's something I'm quite proud of because it's been a real collaborative work to try and give a vision of how you could build a foundational model for CFD. And just to set the context, what that means and the way I described it some people is like a ChatGPT for fluids, you know, that, that dream of a model that could predict any possible

1:31 scenario of any plane, of any car or of any data center or of any pipe flow or anything. And, and not that our paper has the full answer to it, but we tried to look at the scaling laws. I mean, like what's the, what's the bottleneck to doing this from a compute, from a data point of view. Um, it is still a scientific piece of work. It's a piece of research. It has hypotheses, but it's something we deeply thought about, researched, discussed with people in the field, you know, to get their feedback. And something that I think the three of us do stand behind as a combination of CFD, ML and applied maths. And the paper's quite long.

2:14 It's nearly 40 pages. And so the point of this episode was to have the three of us talk through it really, so that hopefully if you're reading it, you understand it um a bit more and you understand where we came from and some of the decisions we made, what we put in, what we didn't put in. And so, yeah, I hope, I hope you enjoy this. This discussion does get into the weeds, but it's a topic that anyone who's listened to this podcast for a while has been a key theme, right? All the questions I've been asking guests have been about foundational models. It's been about can AI, what can AI do for CFD? What can it do for things like Formula One, the cycling, for aircraft design, for space.

2:52 So it's a combination of some of that work, you know, over the past couple of years, thinking in my head and Johannes and Sid had similar sort of thought processes and we all came together to. to write this paper. I'll put a link in the notes so you can get this paper on arXiv. It's a pre-print or if you just Google something like Fluid Intelligence and if you put in Neil Ashton and I mean, that'll bring it up, but you could put any of other co-authors. So I hope you enjoy this paper. I hope you found it useful. It's been a fun piece of work and I hope this helps to explain some of the logic behind it. So yeah, sit back and listen to this special episode.

3:34 Thanks, Sid, Johannes, for, for joining. This is, I guess a special episode because we put out a paper recently that I think has, you know, resonated well with people. And we said, why don't we actually talk about it? Talk about the journey, how we got to this point, some of the technical details, um, so that people, yeah, understand a little bit more, um, than, just reading the paper alone. But maybe before we begin, why don't we just set the background and for the two of you just to give a brief overview, you know, who you are, where you work, what you're doing and, and yeah, your background, I guess, and then we can get into the paper.

4:15 So, Sid, did you want to go first? Yes, I can go first. I'm Sid Mishra. I'm a professor for computational applied mathematics at ETH Zurich in Switzerland. And I run the computational applied mathematics laboratory here. And what I do is my basic background is in physics and math. I did my undergraduate degrees in those subjects. Then I shifted more or less to applied and computational math. Then after my PhD, a lot of the research was on numerical methods, numerical analysis, scientific computing. applied it to different areas, astrophysics, fluid dynamics, um material science, different topics. And for the last five, six years, I've been working on using AI for solving physics problems and with some success.

5:04 First with physics-informed neural networks, and now for the last several years with neural operators, diffusion models, foundation models, and so on. So it's been a lot of fun. Yeah, nice. Johannes? Hi, I'm Johannes Brandstetter. I am a professor at JKU in Linz and I'm also a co-founder and chief scientist of Emmi AI. I've been in the field of surrogates for uh engineering, if you want to call it like that, for approximately three years. I'm thinking about how to build these surrogates that they really fit to industrial scale problems, industrial size problems. That is where I met uh you Neil, think, approximately two years ago, and also where the paths of me and and Sid were overlapping.

5:53 And it's a tremendous pleasure for me to have a paper with both of you because it was very high up on the bucket list to have a paper with each of you. But to have it combined is like really true honor for me. Yeah, I think, think we met because you gave it a finger and like an eye clear workshop. I think I was organizing, but I remember where we really spoke was at a, um, what was the restaurant was it pizza or Italian or something in, um, in Silicon Valley, um, where we, where we met up. And that was where I think we first were like properly getting into this topic. And that was actually. Whilst I was still at AWS and I think I was just.

6:37 Was it like Christmas time or something or January time? I can't remember what it was. So actually when I look back, we'd sort of started having some of these discussions almost like a year ago, um, on this. And maybe what I thought we could do is each of us maybe can set the scene in what prerequisite knowledge we brought, right? What, before we really started to interact and, I guess, you know, um, blend this knowledge together, which is, think ultimately what the paper was. Maybe we can set the scene of like what our thoughts were before. So maybe if I can start it a little bit, I am not a machine learning specialist.

7:17 That's the first thing I'll put my hand up and admit. am not anywhere like the two of you really are like super deep in the knowledge, but I was always intrigued on the, on how far AI could go. And I sort of was thinking more from a very applied point of view. You know, you, you start to build these surrogates, you have like 500 cases. It has a, you know, a certain accuracy. I like all of these ideas of scaling a little bit. So like, if you take a, a solver, you know, it works well for a simple problem, but will that same method work for a really big problem? And we know if we do like a DNS simulation, it works really well for a small problem.

8:01 But the scaling means you can never really do it with a full aircraft, regardless of how good it is. And so with the AI, I was always like tempted of, well, is this just a matter of scale? You know, if you just throw more stuff at it, is that the solution? Just more compute or is there something that ultimately stops it? And I guess, um, the second bit was when I did go to ICLR or NeurIPS, I was sort of amazed by the amazing talent. But at the same time, the sense that they were two different worlds, that you were speaking to these people who didn't have that practical sense of like using CFD for engineering, you know, like actually designing something, but we're looking more theoretical.

8:47 And then you have the CFD side that were really good at that stuff, but just had no clue about all this amazing work that was going on. And I think that's why I was, I'm always excited when I speak to people like you two who have much more of the So applied math ML side and you can educate me. I feel like I'm absorbing off you. Um, but how about you, Johannes? Where did you come into this? What was your, before we really started to go in this sort of three way thing, what was your thoughts? Yeah, approximately, think, no, not approximately, pretty much exactly three years ago, ChatGPT came out and at the very same time I was at Microsoft research and there I pushed

9:30 really, really hard for these foundation models. First, it was the ClimaX model and then it was this Aurora model. And the gist of it was basically throwing as much data as possible into the mix and trying to see what comes out, which worked very well. um and which was a big motivation for me to go into this engineering. However, there's a lot of uh difficulties, uh differences. First of all, um Aurora worked so well because it was all based on vision transformers, which everyone at that point understood how they work and how to scale and how to operate. And there is not a lot of control of the weather data because in the end, you download some weather data from different uh offices at different simulations and they kind of all

10:13 model their similar physics. So they have different discretization schemes, but they all try to have the same turbulence modeling schemes baked in. However, when you go to engineering, you encounter two problems on this axis. First, there is no scalable architecture available for doing a 200 million surface mesh, volume mesh, Formula One car simulation. So you have to build the scalable architectures and there is a big, big difference on the data regime because data can have so many aspects how you discretize what turbulence model you use, what machine you use, what numerical scheme you use and so on and so forth which makes this very very fragile and

10:58 very hard to understand and on both of these fronts I tried to do my research and I tried to move forward and that's where from my side the discussion started. How about you, Sid? mean, obviously you've been working on this for a while and there was like the Poseidon paper and you've had these thoughts of foundational models, I guess. Yeah, so I come from it from sort of similar viewpoint as Johannes. That's why we discuss in general very well. So my perspective in the beginning was more I'm a mathematician, so I wanted to essentially understand if possible, rigorously prove how much of data is necessary for and what is the model size that is necessary for an ML model or an AI model for that matter to learn certain tasks, to excel at certain tasks, to generalize well.

11:51 So this was And I was not looking at vision and text like what most people look at, but I was looking at scientific data sets, physics, could be chemistry, and certainly engineering, because this is what I've worked on for long time. And this was sort of the foundational question. And uh we already, some years ago, had some theoretical results, which essentially said that, OK, things scale if you have enough data. But as Johannes just expressed, you never have enough data. mean, this is the... There's a challenge in some sense, also the opportunity in this field. And then this is where foundation models make sense.

12:27 The idea was that, you have a large corpus of pre-trained data, and then maybe you can fine tune it. And I have been working on this topic from different perspectives, just interested in the fact that can AI systems generalize to unseen circumstances, to unseen physics in particular, because that's my interest. And then of course, I discussed a little bit, Johannes approximately a year back, I think he was in Zurich. And uh then to his credit, he's the one who somehow brought the two of us together because Neil, I know your work very well from the past, but I never had the opportunity to sort of meet you in person or interact with you before.

13:07 So I think it is Johannes who sort of deserves the credit. Because you know, In a way, our knowledge base intersects pretty well. I know a little bit of CFD. You are the real CFD expert. Johannes is a real ML expert. I also know a little bit of ML. So we sort of complement each other very well. But I think the questions that sort of motivate us, drive us are very similar, I would say, to some extent. So I thought maybe, um, we're useful for people listening to this paper, Fluid Intelligence: A Forward Look on AI Foundation Models in Computational Fluid Dynamics, which by the way, I think it took us a while to figure out what title we should have. We had quite a few different titles.

13:49 you know, maybe what we should do is kind of go through it, section by section. Obviously we're not going to be able to go in full depth, but just describing maybe why, you know, we had some, some of the things, and. And hopefully then by the end of it, people will get more like how we came to this conclusion. Um, I think it's fair to say, isn't it that we. Our, way the papers turned out was not exactly how we went into it. It wasn't really the idea to do it this way. but I think it's sort of organically, we kept having meetings and then we would be like, yes, for sure. This is how, you know, yes, we've done it. Okay.

14:32 And then you're like, actually the data showing something different. And, um, yeah, so it definitely has been a bit of a journey, which is what all good research should be. Right. It was sort of live research over WhatsApp, essentially. but yeah, maybe just, uh, I guess the first section, the whole CFD process, I think was born out a little bit. Um, or at least partially when, you know, Johannes and I were talking originally, and I think there was this little bit of sad, and I don't want to put words in your mouth hands, but I guess you don't come from a traditional CFD background. So some of the stuff, you know, you, you kind of knew, but you wasn't as obvious like the industrial side to it.

15:24 And I think that's when we realized that it might be useful for people to have almost that little bit of a, yeah, a go-to reference. that highlighted this idea that it isn't just uh a single PDE that you just somehow need to model. And so it is difficult and probably some people who are reading this and they look at the geometry, the physics modeling, the meshing, people I'm sure will find holes or will find bits that couldn't be covered. You would sort of need to write a textbook and people have wrote textbooks, right, on this. But I think what we were trying to get across is just how broad the input space is. And I think that's where you, Johannes and Sid like this sort of distributional way of looking at this, like trying to look at it in a way that would set it up to have some link

16:12 to the LLMs, right? That was the way you wanted it, was it right, Johannes? Yeah, fully 100%. I think that for machine learning people, the most important equation of the first half of the paper is equation nine. So what equation nine is, we are basically saying you can deconstruct the CFD process into an input vector, which uh well, which has the turbulence model, the geometry, the meshing and all these parts in it. And this input vector you can use to have a look at the input output relations. which makes it very clear that if you want to have a foundation model, it surely needs to capture all these variations in the input vector.

16:55 It also makes it clear that if you fix a few of these conditions, for example, if you use same inflow condition or same boundary condition, that the space of variation gets just much smaller. And the first section, uh which was mostly written by Neil, is actually really, um thought to explain for machine learning people how to come to this input vector, which makes a lot of sense because suddenly you don't think of complex CFD simulation anymore. You think of input output relations and that makes it also easier to compare existing data sets and to understand which data sets can be actually mixed and which they're mixing is

17:36 resulting to very disjoint distributions. And the distribution point of view is Sid's way of thinking, I got from him. Yeah, so maybe just to add to the mix, actually engineers like this thinking, right? A systemic thinking. So you can imagine that there is a system and to a certain system we feed some inputs and we have some outputs at the end of the system's work, right? So if you think of machine learning or learning in particular, it's sort of task specific in that sense that a learning system has an input or a set of inputs, input vectors, and a set of outputs, the output vector, output functions, fields, whatever you call it.

18:15 And equation nine is sort of providing that, right? Where the distributional perspective comes in is that you cannot sort of have everything under the sun, right? Means you have to sample from a distribution. You can make this distribution as broad as possible. But once you sample from an input distribution, then in some sense, your output distribution is highly conditioned on it, right? Because you might add some noise at the time of measurement of your system. So it's a very natural sort of mathematical perhaps or algorithmic thinking about machine learning that you think of sampling from a distribution, take the samples, feed the inputs

18:55 into system, observe its output, and then all the AI system does is try to learn how the system sort of propagates or evolves. And this was sort of the thinking that I always keep and I also teach it to my students in class that think of everything in terms of systems. What's your input? What's your output? What's your input distribution? What's your output distribution? And then there's a lot of clarity because without this clarity, then we don't know what you're talking about, right? It becomes vague. And this is what we, I want that mathematical precision. And this is, think useful to have so that we know what our target is.

19:32 And probably, maybe I'll use this point to jump forward a little bit, just cause it seems an opportune time around the appendix B and the whole notion. Cause as soon as, just to say appendix B is this idea of FLOPS per cell per step. And I felt that was important because one of the objectives coming into this was to some way quantify the, you know, if you're generating a data set. Or you're trying to calculate the costs. How do you do it? And it is obviously massively dependent on those CFD inputs, you know, that equation nine. But one of the things that it is as well is the code itself. And this is a, I still don't think we have it perfect.

20:16 Um, but it was an attempt to say, right, if you are running a RANS solver, you are going to pick certain inputs to match. Like they're not, they're never done in isolation. If you pick RANS, you're probably going to pick an unstructured grid because you're probably going to be doing complex geometries that you need to run fast. And because it's a steady state, you're probably going to go implicit and because it's unstructured. So there's these sort of secrets. Yes, you could technically pick a different combination, but then they're more extreme. So the idea was to say, well, if you do that, you're more likely to do that.

20:55 So let's clump that as one category. which was how we had that sort of implicit RANS unstructured. It's sort of, if you look at many of the ISV codes out today, they have sort of centered around that choice. But then if you're going to be doing like a half a billion cell LES, whilst you could do it within an implicit unstructured code, it's not the optimum way of doing it. And it would have massively misrepresented it. So we thought, let's pick. You know, and there are examples of companies out there who do have, you know, an explicit, because now you are trying to time resolve. So using implicit doesn't make as much sense.

21:38 You're probably going to do Cartesian because in this we picked wall-modelled LES. So you don't need to resolve the boundary layer. So it's fine to use Cartesian and it's a GPU solver. we, we, know, and again, they're categorical choice in a way they're picking which inputs and clumping them. And so there's a million other combinations. But the idea was to say, and maybe we'll get onto that in a minute. A time step and a cell is not equivalent because in the explicit, you're doing hundreds of thousands of time steps and yet in the steady state, you may be just doing hundreds or low thousands, but that flops per step per cell, which is like how much it costs to do it is the important multiplier in this.

22:19 And, um, it's one that I think there's lots of holes in it. And you could calculate it in different ways, but I still think it at least gives you a picture. And I'm, I'm hoping that people listen to it. I'd love to see people test that theory, you know, like how close are we to it? And it'd be good to get feedback, you know, if there's certain things people disagree with or, know, which is, guess, part of the point of putting a pre-print out, isn't it? It's to say, here's something give us feedback, but we got to that bit. CFD and I was pushing for this more and saying, I still think we need to explain AI, you know, to the AI people reading this, they're like, I know that.

23:03 So that fundamentals bit, how did you sort of think about framing the whole, I think the, um, maybe the sentence that depicts it the most is CFD is not a language. I can't remember which of you wrote that, but I that was quite. I did, but... Maybe let me add before we jump to that. Let me add one thing for the ML community. So important to understand and this is also when I started to realize this, I don't know, one, two years ago in engineering and CFD, it's not that you make choices on fidelity and that basically results in everything what Neil just said. It also problem specific. So if you try to simulate the plane, which is in cruise condition, certain CFD choices are okay.

23:52 So you don't need higher resolution or better turbulence model because certain turbulence model are capturing what you actually want to have captured and then you can do all the results in your parameter vector is sort of limited. it's the, it really depends on the problem. It's very different than what we usually know from machine learning that you have, well, my high quality data and my low quality data, there is quality always comes with the problem set up. And this is very important to convey to the machine learning community that it really depends on the problem what you're using. And that is, yeah, that's already the going to the ML part of things.

24:37 um I think it was a quite eye-opener discussing these LLMs. uh what token means for LLMs? I in LLM, everything is so... straightforward in a way you have your documents and your documents you have your tokens you have a bunch of documents you have many tokens this is how you construct your learning tasks and then we tried to map that to CFD which was much harder because you have this data sets where you have suddenly this huge geometries with half a billion meshes or whatever so is this like one data point is this uh should you count the tokens but then the Additionally, can subsample many different input combinations from this data point.

25:25 So how is this all coming together? And at that point, we were really saying, okay, let's really write down what people do in language and then make the connection why CFD is not language. And this is, I think, where Sid should comment. Yeah, so why is CFD not language? Well, for starters, I think uh there are so many obvious differences, right? So in language modeling, you have this entire notion of tokenization, which is a very sort of clear paradigm, right? So what you have is you have these words or pieces of words, and then you convert them into vectors through the process of tokenization. And in fact, you convert them into entries of a codebook through what is called

26:08 quantized tokenization. So your tokens are essentially living in some very large codebook and they are sort of entries on that register, right? And then all that we do in language model is given a distribution on tokens, you sort of sample from the distribution, conditional distribution of the next token, right? So it's very sort of uh mathematically clear cut in some sense. In a way, it's very non-mathematical because language, but thanks to tokenization, we have been able to push that into a very sort of occurred learning objective. This is different in CFD or in physics in general, right, because for us the learning task, you everything is in equation nine in some sense, that is our master equation and so

26:51 on. And what does it say that when you sample, let's say that we fixed the categorical variables, this is always the case you have RANS or LES or the structured unstructured these sort of categorical choices. Once you fix the categorical choice, uh Then you condition that distribution, input distribution upon this category. And then when you sample from this, you are essentially sampling functions, the shape, for instance, right, of your object, or the flow conditions, which could be vectors or even single parameters, boundary conditions and so on and so forth. And your output could be a solution field, could be a sort of quantity of interest, drag, lift, and so on.

27:31 So the learning objective is very different. Just looking at equation number nine. So we cannot Maybe someday we will be able to quantize everything so that everything is just written in terms of distributions of tokens and we can do some next token prediction. But this is unclear to me at least and I have worked quite a bit on this whether we have the accuracy because you know it's not enough that the word is close enough to what because as humans we have a lot of slack in how we understand right, but if the drag is 20 % off we are finished. So we have to be So to me, think it's very important to remember that the sort of input-output setting, the learning task is very different and that leads to a very different kind of formulation,

28:17 which is what there are many similarities, right? Means, after all, means there are more similarities and differences, but the differences here are very, very important. And I think the biggest difference is that our distributions, our inputs, our outputs are very different. And just to keep that spirit is, Because as Johannes argued means, what does it mean? See, if you take a huge corpus in language, you have a lot of tokens. But if you look at the flow field, a lot of it is going to be void, right? So it's not very useful information. So you need that sort of global coupling, which is uh probably different from language.

28:56 Maybe one day when you have very good tokenization for physics, this might change. But at the moment, we are not yet there. I don't know if we ever get there, we are not yet there. And then I guess the next section was, again, it's impossible for us to do a complete job of it because our paper's dedicated to just this. But really that review of, I mean, should say, I mean, we put a, we had a much longer version of this at one point, but then we cut it down, which was the different use cases of AI. And I think we were conscious that we are focusing on the, I guess, surrogate use case. But it is fair to say that there are other use cases of AI for CFD around like, you know, post-processing vision stuff, you know, trying to automatically find patterns or

29:51 initializing solutions, or that I think the most famous one in the CFD world that I'm not sure if both of you have tracked much, but in some ways I feel has hurt AI, which was the turbulence modelling. I think there was this idea that, and I think, like Chris Rumsey and Spalart and people from NASA. I had some of these turbulence-modelling workshops and, Karthik Duraisamy was one of the first for this. think Turbulence Modeling in the Age of Data or I think it was that title, something like that. And, um, really great piece of work and it sort of made sense that a turbulence model being so empirical in some ways tuned, hand tuned to coefficients that ML could do a better job.

30:31 And, um, you know, for years people tried it and I think even today it never really generalized and. And when I speak to people who are maybe not deep ML people, that sits in their mind as like almost AI thought that it could build the ultimate turbulence model. And so I think that's why, at least personally, and I think you both agree, it's, steered it towards a surrogate modeling side because it felt like a well-defined problem that has got probably the most to gain. know, like the, idea, like we say, if you build a surrogate, The inference is in what seconds or less than second. Nowadays it can give you full volume, full surface prediction.

31:16 and making an improved turbulence model is great, but you're still going to be bound by the same time constraints that a CFD simulation gives you. Um, so I think it's just, we, it's probably, we want to make that clear. This is not a, of all AI for CFD, this is in some way a subset. Um, of the surrogate modeling, which is, guess, the foundation model question. Could I say something because since you raised this, maybe I get this off my chest. I've always wondered or often wondered why don't turbulence closure models with AI work so well. And there is probably an obvious reason for that. And let me sort of state what I believe could be a reason, right?

32:01 So the reason is that this mapping that takes data to a turbulence closure model is a very, badly behaved operator if it's a map, very, very badly behaved map. There are many, choices that you can make, many knobs that you can tune in the turbulence model that can give you the same flow outcome, right? So I think it's extremely ill-behaved operator. And in a sense, is not making use of, because eventually people are learning like three, four parameters, right? Or five, six parameters in these models. Machine learning is about big things. If you use neural networks to learn a one-dimensional function, it will always do very poorly compared to its competition.

32:46 You can just take a polynomial and get a better fit. So machine learning only works in very large dimensions, large scale. And this is why I think the bad behavior of the underlying map and the fact that you're trying to solve a low or an artificially low, I think the problem is really high dimensional, but you're trying to solve it in this low dimensional setting. meant that the chances of it generalizing beyond the specific flow regime are very, very low. perhaps that's one of the reasons. Whereas in a surrogate model, on the other hand, it's a really high dimensional problem. Your input vector could be millions of dimensions.

33:23 Your output vector could be a billion dimensions, Neil. In your DrivAerML data set, for instance, if you learn the volume field, this is over a billion points in principle. And this is where I think AI or ML models can excel compared to turbulence modeling. So our choice was perhaps based more on pragmatism, but I think there is a deeper underlying reason behind that. Yeah, I am. And there's two more things which we did not consider. So we were always in all our learning tasks, we were always in the large data limit. So basically, the model shouldn't learn any spurious correlation, but it should really generalize because it has enough data, which is usually in this surrogate tasks where you

34:07 have 100 samples, 50 samples, and then you should learn to generalize across geometry. That's quite hard for me. from a machine learning point of view. it's really in the large data limit. It is fully um Transformer-based in a sense that we train with a simple MSE loss and no physics information is added, just really standard training. And it's in distribution. So no out of distribution questions are asked. This is the setup. And in this setup, one can formalize things, otherwise it gets very, very tricky. And we got a lot of questions in how How do you think out of distribution works and so on and so forth. think this is beyond the scope and this is also where experiments are needed.

34:51 Yeah. And I think, again, the point of the paper was really, this is on the, every meeting I have with every commercial company or research company is just trying to get a sense of if is, is a foundational model possible. And that's where you start to get into the numbers game. Um, and that's where things did seem quite ill defined. You know, people may read this and say, oh yeah, that's obvious. Or I knew that, but I don't think it was even obvious to us before. And we kept sort of changing things around and we weren't fully aware of the, yeah, until you put pen to paper and start to do some of the maths, it's not obvious.

35:38 And I think, um, you know, I think it's fair to say that Sid, this was the bit where I really learned from you. And I think is the strength of the papers that Johannes and I, when we were talking on the side, we were like, wanted to it, to bring some of the maths, you know, to actually bring some of the rigor, because you can do, you know, some back of the envelope calculations, but, you need to try to formalize it more. So how did you go, you know, if we go into like the section on the actual scaling laws, so section five, so what, was your sort of process? How did you come up? with these scales. Yeah. Okay, this is a very, very good question.

36:20 Actually, this is a disclaimer, we know this, but I think our listeners and viewers should also know this, that till the night before the final submission, we were still working out the final numbers and so on. So it was a very iterative process. Nothing was set in stone. And I myself, I was surprised. And the whole point of research is to be surprised, right? No, I have, as I said, I have worked on a mathematical formulation of some of these questions in different uh domains before. So for me to understand, well, we should also say that these are hypothesis, right? We don't have end to end proofs. have not yet built such a foundation model, but we believe that these are very reasonable hypothesis because they hold for small scale models, they hold for language models perhaps

37:10 and so on. So the basic formulation we had very quickly, if you remember the first equations, how it scales with model size and how it scales with data size, this hypothesis we had pretty quickly. I think what was fun was to sort of try out different implications of these ideas, of these exponents and so on. The fact that you get power laws has been known for a while. m You can show something with statistical learning theory, but in the LLMs you have the Kaplan et al., you know, the famous scaling paper, have the Chinchilla scaling laws. So there is quite a bit of work also in the domain that we work on. It's much less, you know, um this is surprising.

37:56 Johannes and I are among the very few people who even discuss this question. If you look at some of the foundational texts for scientific machine learning, this question of scaling is not even there. If you look at some of original papers, there's no scaling. You have a table where you have at certain resolution or at a certain scale of data, these are our errors and we compare with 10 different models. Without any understanding that if you change the data or you change the model size, the picture can be very, very different. So I have been investigating this both theoretically and empirically. to do it at this scale with a proper foundation model to sort of have different scenarios, right?

38:39 In the paper we have a low fidelity, probably not the best way to express it. But I think the key was to be able to formulate everything in terms of a common quantity. So in the beginning we were looking at the model error. If you remember, we had an epsilon, which is the model error. We formulated all the complexity estimates in terms of the model error so that the theory can give us some predictions. But at the last moment, I shifted it to the number of samples so that everyone understands. No one understands what's model error, but everyone understands how much of data that you have so that you have a common dictionary to compare different things.

39:16 And we had these three scenarios. We had this, let's say, RANS, steady state RANS. We have an LES, but we only look at the time average. And finally, we did this full transient uh LES, and we compared the three. And we came up with. I would say some surprising conclusions. uh Perhaps we are going to talk more about it. But what really surprised me later on, it was obvious, but in the beginning, what is not obvious is I maintain that data generation was going to be the dominant cost. I think all three of us believe that. And we wanted our theory to show that, and our numbers to show that. But it's only by this iterative process of discussing that we realized that, in the small data regime and small model regime,

40:03 You have this actually holds true, but there is this crossover point, there is this critical limit, and that's why the paper is really, really interesting. Perhaps we should talk more about that. Maybe let me add here one thing because Sid you mentioned that people don't look at scaling laws so they try everything at low scale or with one resolution. And then there's actually the other extreme where people say, okay, we just burn millions and millions of dollars and we get lots of different data, we bunch them together and what we get out is a model which solves it all. And that was my entry point into this whole area.

40:39 That's why I pushed so hard to write everything as a composite vector. because that makes it very clear that it just doesn't work if you throw data together. um And therefore, we had set up this RANS versus LES type of things because we all agreed that you probably cannot mix RANS and LES simulation so easily and that the choice you make on this has a tremendous impact in the data generation and in everything you do. and all comes down to modeling error you take into account. basically so before doing that, you already take a modeling error into account, which has the consequence of all your future steps and all your scaling loss and everything you

41:26 obtain. And this is something which is not clear to the machine learning community, which is very, very different than how machine learning people think. So that's why we have these three different setups of this modeling paradigms. Yeah. And I mean, we can come to that a bit later as well. Some of the open questions and I think definitely the mixing of fidelities was something we kept going back and forth on, but maybe we should leave that a little bit towards the end. Cause I think that's definitely the open question. One thing that I did want to highlight is, that transient one and maybe to explain a little bit more because, and again, it's true what Sid and you said as well, Johannes,

42:06 about our assumption going in. I think I've even said it on this podcast in a few opinion things where I said, I, it felt to me that data generation was the biggest bottleneck and the idea that you could use transient data felt like intractable. Um, but I changed my tune a little bit when I speak into the two of you, because I think one, we started to realize that you run a RANS simulation. And you in some ways use all the information that the simulation gives you because you can't take an early mid simulation checkpoint because it doesn't have any physical sense. The solution is converging to a steady state with the transient.

42:57 On the other hand, we spend, you know, five times more, 10 times more, depending on the solver to generate this data. But we only use the time average checkpoint. All this data with, you know, was sticking out. don't use. And one, some of the early papers, I guess, uh, on like MeshGraphNets, you know, had this idea that you would train on the time steps, you know, cause it was like, know, a cylinder flow or something. And that was always intractable. Cause you said, I can't store all that data. I can't use all that data. And so I think it's to be fair today, all the state of the art work in an industrial way has always just been a steady state.

43:38 or some time average. Um, and I thought, okay, you'd need petabytes of storage, unbelievable amount. And, but I think one of the findings, uh there's nuance is the idea that if we do run a transient simulation, can we take, let's say, you know, a 50th or, you know, or some of the intermediate data, which can be extra samples. And I think that was something you both mentioned, but Sid in particular, I remember you mentioned that from some of your earlier work and that does seem to be key finding. We're not sure how much you could use, but more than one. Yeah, well, I can talk about it a little bit.

44:26 Again, I think a distributional perspective gives you some insight here, right? Because what we're learning when we are learning a time-dependent process is that we are learning the evolution of the time-dependent operator, how things vary in time. And in principle, we are learning, given a snapshot of the current state, what would be a future state given the lead time. This can be done directly. This can be done autoregressively. That's a detail, right? But essentially what you do is that you sample the data distribution in different ways. Of course, if you just have very, very tiny time steps, if you learn every single time step, this is useless because essentially you're learning the identity then, right?

45:08 But if you learn large time steps, which is what surrogates allow you to do, you're really sampling the distribution at different ways. And I think that makes a lot of intuitive sense. my previous work, we had demonstrated, think others have also demonstrated, you just imagine that you're learning the steady state from some inputs and just uh using this training where you learn the transient information too, adds value. You can reduce the number of samples, you can increase the accuracy and so on. I think Johannes likes the perspective of data augmentation. This is one valid way to think about it. I would just think that you are just getting more data from sampling more data from the underlying distribution.

45:51 And because of that, then you have this behavior. Of course, what happens is you cannot, means you lose information if you sample too much. But if you sample too little, down sample too much, if you down sample too little, then it's too much in terms of training and storage and so on. So the golden mean, which we don't know, could be problem dependent, is to sample a few of the things, 50th or whatever we said 50th, based on some ballpark. And then you do have some gain. And this is not 50 times better, but it's perhaps square root of 50, which is seven times better. This is what we put the number five, which is based on some previous work.

46:31 But I think it makes a lot of sense. ah intuitively as well as mathematically and I think it's a very important, well one shouldn't throw away the transient information. I'm pretty confident that it will help you. uh For me, it was very eye-opening because we always had this discussion. um In LLM, you basically assume you see each token once. every token is seen once. And what does this mean in CFD? That means every data point should be seen once if you go for large scale limits. But if you have 20,000 simulations, which is already a pretty big data set, that's kind of not possible. But then this traversing in time, what Sid mentioned that you need this time dimension is actually also what happens in space.

47:21 So just sub-sampling one uh data point is not giving you the whole information. So therefore you can see the data points multiple times and that's the same of traversing in space in time is the traversing in space and that still corresponds to this LLM paradigm. in the infinite data. this was quite enlightening for me when I understood this by the explanation of Sid over the temporal domain. I think the other one that maybe is a key thing is the, which comes later when we showed the numbers is the down sampling. And I think this matters from the training, but also the storage side that, you know, when I was generating the DrivAerML or Ahmed or whatever data sets, you know, we, we provided

48:07 the full data, the full volume, the full surface. And that was partially because even at the time we were just unsure of like, it felt like we were down sampling just because of memory limits. And that was probably more because of the graph neural net approaches that we were trying at the time. Um, but it, it seems like the, and I think you have this, Johannes, in a couple of your recent papers where there's like a saturation point of like how many tokens in this case do you need? Where if you go beyond it or lower, if you go beyond it, you don't gain anything. And so I think, you know, what, what was that? come up with a number now.

48:43 What was it like on a hundred million cell case with three million surface, you know, you're only needing hundreds of thousands of tokens rather than the tokens equals the number of cells. And that is a massive, um, difference, right? It fundamentally alters the cost of training versus if you had to take everyone. How much of that was, how important do you think that is that, that down sampling essentially? I see it again more from the perspective of a person who does not come from a CFD uh PhD. um CFD needs this fine resolution because otherwise simulations just don't converge.

49:32 um We artificially have this fine mesh because we have to enforce physics onto the computer and we as humans cannot do that with a mesh which a machine learning model would understand. So we need this. fine resolution for the CFD to converge. I mean, this is a bit bluntly put, but in a way. So the machine learning model does not need this fine resolution to understand the problem. in fact, the machine learning model is not solving the CFD. It's not about convergence. It's not about uh implicit time stepping or whatnot. It's just about getting the information. And information is not the same as a CFD convergence criteria.

50:10 um I have not understood this at the first, but I think it boils down to this that you can have the full information m in a down sampled case. And therefore I would assume that if you do a RANS or an LES simulation, it would need approximately the same number of tokens. However, you're modeling different physics and for the CFD, you need much more, finer resolution for the LES. So this is a very big difference between uh simulation and machine learning. Sid, is that how you say more or less. But from a purely mathematical perspective, I just uh formalize a little bit of what Johannes exactly said, right? So the reason why we have extreme grids in space as well as very tiny time steps with explicit methods is because of

51:01 accuracy and stability requirements. That's why we need extremely fine grids so that we can resolve small vortices, we can resolve and in time you have the CFL conditions. So if you have a very small spacing in space, mesh size in space, you have consequently a small time step. Now, once you have generated the simulation and I urge everyone to do this experiment, they can themselves down sample and then they can down sample, they can put it aside. And then they can up sample using some simple interpolant, right? It you don't even need a fancy interpolant. And you'll see that the errors that you make in this process are tiny compared to your modeling error.

51:40 Your modeling error is, of course, if you just down sample at 10 points, when you have 10 million points, you make a huge error. But as long as it is reasonable, and that's what we sort of uh argued in our exact numbers, I think we said that we down sampled to like 8 million, and then we do something more. And so we had some concrete numbers. The interpolation errors are so tiny that no information is lost. And this is the big difference because in a CFD setting, you sort of simulate it at that resolution, you get nothing, right? But once you have simulated, you can always down sample both in space as well as in time.

52:17 Sid, you should probably explain what you tell your PhD students what they should do with a data set because this is something I found very interesting and which I learned from you. My students, my students in my class, I always say when you get a sort of scientific data set, what you should really do is to first see what is the essential spatial scale that you need and what is the essential time scale that you need, right? And the best way you can do this is that you take your data, you keep on downsampling it till you can arrive at 1 % error. 1 % is a number that I've plucked out here. It could be 0.5 or 0.2 % error.

52:57 And that is the information. You don't need more than that. That's an extremis. It's the same with time. You don't need every single time step. How much can you down sample? And then you have an idea of the spatial and temporal scales of your problem. I also tell them to look at the sort of variation of the data. So compute things like mean and the variance. Say what is sort of the noise to signal ratio, what is standard deviation over mean so that They understand what the data has and it is very useful, but uh as always, students don't always do it. They want to take the data and they want to run a model and yeah.

53:34 m So if we go, you know, into maybe the numbers a little bit, and if we, think what was, it's true what Sid said and, um, it's true, things were changing, but it was almost a uh good thing that we were doing that because we were trying, there's always a temptation of, um, like group think where you all sort of don't want to challenge each other. You've sort of got to a conclusion and nobody wants to challenge it. Where I think we were all open to challenging each other's perceptions. And I know I, as you described, definitely came into this thinking the dataset was the biggest and that seemed to be true in the early numbers we had.

54:25 but then, so if we go through, so we have this low fidelity, high fidelity transient, but maybe the more interesting one is the table for. which is this large, extra large, extra, extra large, and then the graphs in figure three. And, um, yeah, what we were basically trying to show. So maybe for the data generation, I can go over some of the numbers and maybe Sid, you and Johannes can talk a little bit on the model training. The, they all, they are by definition. We did try to give some estimate based on the sort of error floor, I guess, but, but it is still tricky to estimate the number of samples. And that's why we did try a couple of approaches.

55:03 One was, I guess, the more statistical, more mathematical way of doing it. And then the second way, which is Appendix C was, I guess, the way that I think more CFD people have looked at it, which is just simply this sense of, well, how can my model predict a plane if it's only trained on cars? You know, doesn't matter how many cars you give it. It's never going to be able to predict a plane, right? I mean, that's just common sense. Um, so. And one of the challenging theories of CFD and it's an interesting intellectual exercise. CFD is very broad, very broad. you know, the possible geometries and flows and physics that you can do with CFD is quite extreme.

55:47 And so if you look at like appendix C, I'm sure I missed out some, but it is amazing when you start to go to thing and you think, okay, it's cars, right. But then it's also lorries. and trains and planes and space planes and rockets and engines, data centers, buildings, combustion. And then what's really interesting is some of that chemical side. There's a huge industry all doing these like multi-phase flows, multi-physics, cement mixing, ice cream making. there's just so many and each of those. The input, you know, that we described, which is essentially boundary conditions, initial conditions. physics modeling, we, we purposely summarized it in the sort of turbulence model, because turbulence model is one of the biggest impacts.

56:35 But once you get into these chemical and process, the, the additional source terms you need to model some of the physical processes, just get it big. So, you know, with those numbers, you can easily reach into the millions of samples. Um, now what I think is not entirely clear to me is just how much cross learning there is between, you know, there is obviously a, you could argue a low speed plane shares quite a lot with a car in some ways, but an ice cream is quite different. So there's obviously some, you know, learning, but anyway, that's how we got into the millions. we said 200,000, a million and 2 million. we purposely gave the Python code to this.

57:23 So you could change these numbers yourself. Um, because they will change if you went to now 20 million, it would change the numbers. Um, we went for the 500 million cells. This is based on some work that I've been doing, um, for, for an upcoming data set, which was on an aircraft and the 500 million is sort of at the, it's a, it's a number that encompasses quite a lot of cases in terms of, you know, the Reynolds numbers, the sort of structures you get, know, 500 million is quite a good number to represent anything in the, let's say, LES of buildings or planes of cars, or it, you could obviously change the number, but it is representative of that sort of case.

58:17 Steps is tricky because it depends on the mesh size, the total convective time you'd want to go for. But we have to pick some numbers to base it in. So that was a 200,000. The flops per cell per step is quite aggressive. If you do the maths on your own code, if anyone's listening to this, you'll probably find that that represents the ultimate today of efficiency of a code. You're probably fine if you run, I don't know, OpenFOAM, you'll definitely not be at that flops per cell per step. It'll be quite a bit higher. But if you go through the maths, it's not unreasonable, but it is quite aggressive. Um, so I say that because the number of hours you may need would be more if your code wasn't that well optimized.

59:02 then in terms of just to finish off some of the logic of this, the, we got to the hours. Now this is tricky and does get into a bit of the HPC side, but we essentially took the flops and then worked out if you were using FP 32, which is reasonable for like modern day CFD. There are some cases that would need FP64, but quite a lot use FP32. So we took the flops that was available on the system, but we realized that because many of these codes are memory bandwidth bound, the actual flops they can use on the system is still relatively low. So let's say 15%. em And that's ultimately then once you take a cost,

59:51 And I took these costs. you go across all different cloud vendors, all the cloud computing vendors give their costs publicly. So I think $8 was like probably the best price with a reserved instance. You know, if you sign up and you buy them upfront for three years. So, but my logic was if you're spending a hundred million dollars, you're probably going to do that sort of deal. You're not just going to pay like an off the shelf price. You are going to commit. And then the storage was the same. was like a one of cheapest. cloud pricing for storage. And so that's how you got to these 10, 50, a hundred million. And just to finish off on the data generation side, what we did do is we put into Figure 3 in, think the sub figures around D, E and F that if you did change the amount of

1:00:46 flops or the size of the mesh or the cell size, how that would change. And so those numbers could go to 200 million to 500 million. Um, but they sort of set a ballpark of where we're at. then for the model training, mean, one of the things that we were working on till the last moment was like the, the FLOPS question, right? That was a tricky one. Yeah, you want me to take it. But before that, maybe this is some people have commented on this to me privately is that our assumption is that they, so some of the practitioners, they claim that our numbers are too aggressive because most of the codes at the moment, or many of the codes don't have the sort of GPU capacity yet.

1:01:37 Right. And if you're doing this on a CPU, then this numbers change dramatically or they will change, right? So perhaps, maybe Neil, you are the expert in this. What's your thinking about this? We have made a big assumption, a big bet here that everything can be generated on a GPU at FP32 with a very aggressive flops per cell, which we don't vary much in figure three. So what happens if you are stuck with a CPU code? Well, yeah, interesting question. And, know, obviously I have to, as you do in papers, put like, uh, what's your conflicting thing. work for NVIDIA, a GPU company, right? Obviously. But, you know, people hopefully will know that I'm not in any, this is not a marketing or sales exercise.

1:02:23 I think it's pretty defensible to say that if you look across all of the supercomputing centers, all of the people developing codes, any new code. If you're going to start a code from today, are you going to write it on the CPU? No. I mean, that's, that's just the case. Um, so, but you are absolutely right, which is sort of the comment on OpenFOAM. If you were to run with like a CPU code, I'm pretty sure those numbers would be 10 times higher. Um, which would basically make it. You couldn't do it. That's sort of my logic, which is a little bit why the, It is irrelevant for the, for the GPU topic here. and, yeah, but you're right, probably in the revised version, we probably should give a bit of a caveat of it.

1:03:13 I just take it as an obvious thing, but maybe I'm a bit like surrounded by this data for me as a no-brainer that you do it. Um, but we could definitely add, but it would make the numbers quite scary. This is the point that I wanted to make. maybe I can walk and Johannes can just add. I can walk us through the choices we made regarding the model training and what are the different things that come in. So before that, regarding the number of samples, so it does turn out that it was a fortuitous coincidence that this 2 million is roughly, 2.5 million is what Neil arrived at by just sampling this entire data set of, you know, chemical reactors and rockets and cars and everything in between, right?

1:03:58 But independently, because we have some or we have some clue about what is this. So there are two key numbers of exponents here, beta in table four, if someone has to look at the paper, and alpha. So beta is a coefficient by which things scale with respect to data. And more or less central limit theorem that we learn law of large numbers that we learn at an undergraduate level tells us that this is no better than 0.5. There's usually a logarithmic correction. And so far, we looked at the literature, including my own work. we put the number 0.43 because this was sort of a representative of several test cases. So it turns out that by using some nominal errors, if you want to arrive at

1:04:44 below 1 % error in the field, not in the integral quantities, then you would need 2.7 million samples, which is very close to the number that you also arrived at, Neil. So this was not totally unscientific that these two things sort of coincide. So remember that we have this, the scale at which, or the exponent at which, power law at which, model size contributes, sorry, the data set size contributes to the error. So that's an important thing. The other thing is, of course, the scaling exponent with respect to model size. Now, is a fact that you need to grow the models a lot more to be able to ingest the data that you provide to them.

1:05:27 So because there is a slower decay with respect to model size than with respect to data set size. So that's why you see some of the big numbers when it comes to model size in the paper. So these were the two key numbers that we fit in. We also added something like the number of copies that you want to see in transient training. What is the compression ratio that you are going to see? We put some reasonable numbers there. And the important thing was the flops. This is a similar question. This is a training flops per step. And this was finally we put, I think, very, very aggressive numbers based on some use case that we had in the lab on some GH200s.

1:06:06 We have assumed GB200s. And we assumed a very, very sort of aggressive scaling. In general, if someone is interested in figure three, I think we have provided where we, no, we didn't provide this, probably in the next revision we can provide that. But this is also a caveat. We were very aggressive. So these numbers could grow. And also the values of alpha and beta are not set in stone. Again, if you look at figures three, A, B, We sort of vary alpha, keeping beta fixed, vary beta, keeping alpha fixed. And then you can see how these different numbers change. And the important thing is whatever you do asymptotically, at some point for large enough amounts of data, the model size has to be very big.

1:06:53 So training the model will be very expensive, and it will overtake or dominate the cost of data generation. And this is sort of the... We have a theoretical demonstration of this and we also see this empirically. So this was sort of the logic. Maybe, Johannes, you wanted to add something? mean, this was perfect. It's just one thing to underline. um All these coefficients can change. Change is not even the right word. They need to be um determined scientifically, experimentally. But what we are quite certain is the slopes of these two curves. And that is, we were quite um surprised that data generation is a different slope than the model training.

1:07:39 And the model training at some point, the slope overtakes the data-generation slope. And this has big impact in many, many ways. And this is one of the big findings, I would say. Yes, so just to paraphrase, data generation scales linearly with data set size. Model training scales superlinearly. How superlinear depends on this alpha and beta, but it will always go superlinearly because the model is always going to be slower ah in scaling than the data. And then eventually you'll have a crossover point. Yeah. That's kind of the real teaser, I guess, because you're right. Nobody or, you know, publicly anyway, is doing the sort of huge data set generation.

1:08:25 So it's kind of hard for us to, to put it out, but as soon as that does happen, it'd be very interesting to compare because as you say, if you take all you're doing is delaying it. So let's say, imagine the data generations on a CPU and all those numbers are 10 times higher. It just means from our prediction that the crossover point will come later, but it will come at some point. And so it's almost then the question is, well, from a usefulness point of view, how big does the model need to be to give a low enough error and a broad enough generalization? And depending on that, will determine where the biggest cost on investment is, but also, um, maybe this is a good segue into some of the open questions.

1:09:20 One of the topics that I went into writing this paper thinking would be a bigger one, but I think in the end, the scaling laws more dominated the discussion. And I think we agree that we just didn't have enough information to write on it was some of the online training. You know, that sense that the data-generation cost is so big. that you want and the cost of storing the data is so high that you need some way of doing it online. And I still sort of get that feeling, but the time scales difference. We assume that the model training is super fast and generating the data takes ages, but at some crossover point, when the model training takes forever or very long time in the day,

1:10:04 you you sort of wonder where these are going to come in. and it is interesting that the, I should also caveat. The storage numbers assume very heavy compression. Um, so we should put a sort of caveat out there that the storage may still be quite a large cost. also the storage I'm picking here is, you know, quite a cheap object, based storage system. This is not a like a Lustre file system that you're dumping on. That would be like seven times more expensive. So. Um, I don't know what, what did, what did you go, what did you feel as the key opening, uh, the open questions, you know, maybe comment on the online training.

1:10:54 Um, how much do you see that as being important, overblown? I just add one thing to the storage. um So there is this very interesting phenomenon that we're training surrogates and surrogates always make errors, right? So usually in compression, if you do um text or images or videos, you basically want to go for lossless compression. But if your surrogate is, your model is making an error anyways, that means your compression can also have an error, this error just needs to be smaller than the error the surrogate is making, which opens a huge field of research. I shouldn't probably say that loud, but I think that there's a lot of potential in that.

1:11:40 And the second point, which is super intriguing to me is the data mixing. I don't think what we did in Aurora, that bunching data together and just uh doing the mix, is the way forward. I think that having a clever formulation of how you can leverage different fidelities in that sense, or different that you stay as cheap as possible in the data generation side, while being as performant as possible with the highest fidelity on the modeling side is one of the key open research points to discuss, which is super, super hard because that means you have to operate on scale. bring these communities together. You cannot do that in a machine learning lab without CFD expertise.

1:12:29 You probably cannot do that. But that's where the fun begins, I would say. Yeah, and I don't know whether I'm allowed to say this aloud, but we discussed this quite a bit. We even had different versions of the write-up where we had different models on how we can sort of combine data and so on. We different strategies. Maybe that should be a separate paper altogether at some point. But I agree with you, Johannes, that this is a super interesting question and also this question of cross-learning that you raised, Neil, right? So how much of information can be In my day job as an academic, I'm very, very interested in that.

1:13:06 How much of physics can be learned from other physical effects? Can you learn diffusion from fluid flows? With Poseidon, we showed that this can be done because somehow it is hidden inside. So these are the sort of scientific questions which will be hopefully answered in the year or two, But uh we cannot wait for things to... So the best way to generalize is to... make your distribution so large that everything is more or less an in-distribution problem. We have seen that with language modeling, right? So I think putting these numbers out, telling people that, hey, look, this can be done provided that there are these, and people can make their own choices.

1:13:45 Maybe I don't scale my model at all. And I simply say that my model has 50 billion parameters come what may. And then at some level, as you generate more and more and more data, All that it does is that the model doesn't improve because it's dominated by your modeling error, right? So the model error, not the modeling error. So these are choices that people can make and they will make pragmatic choices, but at least it gives a ballpark of what people should aim for. So I think that's what is very interesting outcome of this project. And I think, at least from my perspective, that that still feels like the unanswered question.

1:14:24 We were hoping, or at least I was hoping that we would have a little bit more of a definitive answer on this topic, which is, you know, if I train the model with the RANS with a certain level of error, and then I also give it some LES with a similar error, how can the ML model know that one was RANS? And one was LES and we touched this a little bit in the paper, but this, we have, I feel set the question right with the equation nine, but we haven't still fully answered how you would use and mix all those different inputs. Um, that for me feels we, tried it and I think we felt it was just, we didn't have enough data to fully answer that question, but that feels like you're right.

1:15:15 Cause the point being is that. Um, you may not have the luxury of generating all this data from scratch. It may be pre-existing data where it has had a certain modeling error or a certain boundary conditions or a certain fidelity. And you need a way of the model knowing that rather than saying, everything will be an LES and everything will be, that still feels like, uh Yeah, a key thing which is not clear, right? My take is, so there's two opinions on that. take is, and I think I agree here with Sid, that you had to get rid of categorical distributions because categorical distributions is the arch enemy of generalization.

1:16:05 On the other hand, people say, okay, it's basically a multi-task learning because the model learns two different tasks and the weights share across these two tasks. um I cannot comment because it needs experiments, needs large scale experiments but I'm just convinced that categorical formalization is just very very hard because you never get then out of this distribution problem. Yeah. But the reason we did some of this scaling laws or example was even if you could make the data generation half the cost, would still at some point have the training to be the biggest, right?

1:16:53 So we're, there's no free lunch in that sense. You know, there's, if you do it all with RANS, yes, compared to the estimate we had, maybe it's half the price of the data generation, but Um, if you still believe that you need millions of samples, then you're just shifting it more to the model training. Um, but then how do you fix the, ultimate issue that the RANS error, you know, unless you come up, which is where the full circle is interesting in my, and we didn't talk about too much in the paper. There is still a reason to come up with the ultimate RANS model. If you could use AI to go with the ultimate RANS model.

1:17:33 You would alternate the data generation costs quite a bit lower, but, the reality of that happening, as you said, is a much harder problem, ironically. Um, but there is, there is some motivation for it. Um, maybe one of the other final topics that we should cover, maybe we should come back, um, and have another discussion on this when we've had more feedback from the community, um, is the inductive bias one. I think we were quite deliberate in saying. We'd already stretched ourselves with making hypotheses and making assumptions that I think we didn't feel this was the right time to have too much of dedicated opinions of

1:18:14 including physics. We wanted this to be a little bit more of a, like a data driven, like how much data do you need? What's the scaling law? It's still an open question, right? Um, but it feels one way you would need quite a lot more paper space to, get into this. I don't know what you both think. I answer because I would burn if I answer here. Okay, to put it as politely as possible, we don't know, right? This is a reality. I think the belief that one way to put physics is just to know what the governing equations are and stick them into the loss function. We know that there are some difficulties in training this.

1:18:58 The training is ill-conditioned. You have to precondition it somehow. There are some ways out there, but this is far from what we are already able to do with the data-driven approach. Far from it, right? With the data-driven approach, we are able to predict the weather now, which is a physics-based problem, and it's completely done in a data-driven approach. So I don't think the question of how to add physics has been answered yet, even in an academic setting, let alone in a large-scale setting, right? So you could argue that maybe we can put in conservation laws, symmetries. I think there are some comments about that also on the post.

1:19:39 Yes, that could be. We could use that as data augmentation and so on. But I think it would still be a low ball. It's not going to change things dramatically because maybe the physics has lot of hidden, explicit symmetries, but the boundary conditions break it, right? And then, where do you go? And so I... And sticking the sort of physics equation based laws may so far has not been able to be shown to scale at this limits, but maybe it's possible. I think the question of how to add physics into these models is a very interesting one. In my opinion, it has not been answered yet. And I think one of the points of this paper, which I was keen and I think maybe sets it apart from some other recent papers talking about foundational models is it is unashamedly

1:20:34 a more industrial focused, like, you know, if you want to build a model that could predict a car or a plane or a data center using the architectures available today. would you do that? And so that is why it doesn't focus on stuff that could be around. It is essentially saying there are architectures available today. How much compute and data would you need to throw at it to build the model? That is essentially the like summary, isn't it? Um, and, and what is undeniable is there will be future architectures and future ways of doing things that could incorporate it. But I think what we're setting out is if you've got enough.

1:21:18 you know, compute budget and data budget. feels like this could be done, right? I mean, that's the sort of takeaway I'm getting from this paper is that if you have a efficient enough data generation code and a scalable architecture, the only limitation is compute. And, uh, I think because of that, I would make a prediction that we will see people try and build these models because those numbers are not so high that they are beyond, especially in the age of AI, you know, stuff. I hopefully that formalizes a little bit that at least from our numbers, this, and I say this, maybe this is the final point.

1:22:10 It'd be good to get your final. There was an early version of the paper where I wrote the bottom of the abstract, something like. And we conclude that this is an intractable problem that you cannot actually build it. And, and I had to completely change it, you know, to basically the other way around because I believe now it is a tractable problem. We don't know how accurate it will be. Right. So I, so I think it can be done. What's your final thoughts? Yeah, that's a beautiful ending and not much to add. Just saying um it can be done, but not for all CFD. So not for all the cases you listed in appendix C together, but for certain specified um set of cases either together or separately.

1:23:02 It depends on what you're looking at. Industry doesn't need a model which works for all CFD. Industry needs a model which works for the type of problems they're looking at. This does not only apply for CFD, it applies for all sorts of other problems, semiconductors, crash testing, blah, blah, blah. Yeah, being a sort of the more the scientist here, maybe I can say that, yeah, the scientist dream is to have this large scale model. I think we are close. Maybe we, because our main assumption in terms of data generation, I think we have laid the case very well. In terms of model architecture, we have made the hypothesis very clear that your model has to scale in a certain manner with respect to data and with respect to model size, right?

1:23:51 There are architectures out there which can do that. And they can certainly do it in sort of specific domains, as Johannes rightly said. Whether they can do it cross-domain or not, this is essentially my research. I'm occupied with that question all the time. I think uh hopefully the answer is yes, but uh certainly in a sort of restricted setting that Johannes said, I think our numbers clearly show that this can certainly be done. But let's be optimistic and say that at least in some cross-domain settings, can still do it. Soon, soon enough, rather than waiting for five years. It's been a pleasure. Um, we'll have to, do this again.

1:24:30 So thanks guys. Thanks. It was an absolute pleasure not only writing this paper but also having the podcast with you. Yeah. Awesome. Hi, thanks.