The Neil Ashton Podcast

Prof. Michael Mahoney — Perspectives on AI for Science

Season 2, episode 7 01:16:44

Prof. Michael Mahoney — Perspectives on AI for Science — The Neil Ashton Podcast

Watch on YouTube

Prof. Michael Mahoney — Perspectives on AI for Science

Open on YouTube

YouTube video

Watch this episode

YouTube is contacted only after you choose to play the video, keeping this page fast and private by default.

Listen to the audio

Spotify player

Prof. Michael Mahoney — Perspectives on AI for Science

Spotify is contacted only after you load the player and may then set cookies.

Open in Spotify

Episode overview

In this episode of the Neil Ashton podcast, Professor Michael Mahoney discusses the intersection of machine learning, mathematics, and computer science. The conversation covers topics such as randomized linear algebra, foundational models for science, and the debate between physics-informed and data-driven approaches. Prof.

Mahoney shares insights on the relevance of his research, the potential of using randomness in algorithms, and the evolving landscape of machine learning in scientific disciplines. He also discusses the evolution and practical applications of randomized linear algebra in machine learning, emphasizing the importance of randomness and data availability. He explores the tension between traditional scientific methods and modern machine learning approaches, highlighting the need for collaboration across disciplines.

Prof Mahoney also addresses the challenges of data licensing and the commercial viability of machine learning solutions, offering insights for aspiring researchers in the field.

Chapters

  1. 00:00 Introduction to the Podcast and Guest
  2. 05:51 Understanding Randomized Linear Algebra
  3. 19:09 Foundational Models for Science
  4. 32:29 Physics-Informed vs Data-Driven Approaches
  5. 38:36 The Practical Application of Randomized Linear Algebra
  6. 39:32 Creative Destruction in Linear Algebra and Machine Learning
  7. 40:32 The Role of Randomness in Scientific Machine Learning
  8. 41:56 Identifying Commonalities Across Scientific Domains
  9. 42:52 The Horizontal vs. Vertical Application of Machine Learning
  10. 44:19 The Challenge of Common Architectures in Science
  11. 46:31 Data Availability and Licensing Issues
  12. 50:04 The Future of Foundation Models in Science
  13. 54:21 The Commercial Viability of Machine Learning Solutions
  14. 58:05 Emerging Opportunities in Scientific Machine Learning
  15. 01:00:24 Navigating Academia and Industry in Machine Learning
  16. 01:11:15 Advice for Aspiring Scientific Machine Learning Researchers

References and links

Transcript

This transcript was created from the corrected YouTube captions, with names and technical terminology reviewed. Download the corrected SRT file.

0:00 Hi, and welcome to the Neil Ashton Podcast. In each episode, we explained some of the fascinating ways that science and engineering are changing the world around us. We talked to leading engineers from elite level sports like cycling and Formula One to some of the world's top academics to understand how fluid dynamics, machine learning and supercomputing are bringing in a new era of discovery. We also hear some of their life stories, their career advice, the lessons they've learned on the way that I hope will be helpful to you too. So sit back and enjoy this episode. Hi, and welcome back to the Neil Ashton Podcast. So on today's episode, we have

0:45 Professor Michael Mahoney, who is one of the world's leading experts on machine learning, mathematics, and computer science. He's also somebody who I've got to know over the past year and has been a great help in understanding this space. He's among other things, an Amazon scholar. So we've managed to work together a little bit. And hopefully, as you can tell from this episode, he's a really nice guy and has a great sense of humor and such an intelligent person. Where do I start to talk about what he's done? Well, he's a professor at UC Berkeley in the Department of Statistics, but he's also at Lawrence Berkeley National Lab.

1:28 And as I mentioned, he's also an Amazon scholar and, and has a few more hats as well. He, if you look at his Google Scholar, which for a lot of academics is a way of, you know, getting a look at what they've done, he has some pretty impressive statistics. And one that sort of comes to mind, not only does he have more than 36,000 citations, which, which is a lot, his h-index is 84, which is very high. But what's more impressive is, and again, this is probably going into the details, but it makes a difference. If you go on Google Scholar and you look at the number of citations, his is exponentially growing. So not only does he have more than, you know, 30,000

2:14 citations, h-index of 84, which is already in the very high, sort of top 1% or probably even higher, it is increasing like this, which basically shows that his research is becoming ever more relevant every single year. It's not plateauing or going down. That's actually very impressive and that shows why he is so highly regarded because his work is really cutting edge. There are many, I think it's fair to say that because he has that mathematics background and that computer science background, he's put his hand to many different things. But one thing maybe relevant for the listeners or viewers of this podcast is a paper he did with

2:54 some colleagues on characterizing possible failure modes in physics-informed neural networks. And this is, you hear many people talk about PINNs: physics- informed neural networks. And so they did a great paper a couple of years ago looking at some places where it may not do so well. We talk actually about that in the podcast. We go through some of the themes I've discussed with other leading ML experts. And we talk about, for example, his opinion on, you know, do you really need to include the physics in this AI for science regime or is it good enough just to use data? We discussed this at length. And we also get into the topic

3:29 of foundational models for science, something that he's actually really being a leading voice on. He's given numerous keynotes, important seminars, and published papers in this area. So we have a good debate about that. But we actually start off the conversation talking about something he's also very well known for, which is randomized numerical linear algebra, which is quite a complex topic, to be, to be completely honest. And I don't think we got to the bottom of it in this short talk, but it really shows that some of his work from the pure math side or pure math side is coming through and being relevant as we increasingly look towards lower precision methods and ways of

4:07 doing mathematical tricks to try and improve the speed of many linear solvers, which are common, of course in many, many fields of science, including CFD. We also talk about his general opinions or advice for people looking to move between academia and industry. Just general career advice, something he's well placed given that he has an incredible academic record, but he's also very close to industry and sort of industrial applications as well. An hour or an hour and a bit is not really enough to go through all of this, but I, I certainly learnt a lot and I really enjoyed talking to him like I do every time I speak to him. And I hope you'll, you'll find

4:50 the same thing. I will add some links if you're watching this on YouTube because again, he's published so many papers and done so many great talks that I really want you after watching this, listening to this to go to his website and and look through all of those. And obviously if you are watching this on YouTube, you may prefer to listen to it. Some people don't realise that these podcasts are in video form and YouTube, but they're also on Spotify and Apple for audio and vice versa. If you normally listen to this and you didn't realise there's a full video version, you can go to YouTube. So yeah, I hope you enjoy this

5:26 conversation with Professor Michael Mahoney. One topic that you mentioned to me when we were chatting in the past that I was fascinated by, but if I'm being completely honest, I didn't fully understand. So it's good for my purpose and for everybody else was around the randomized linear algebra. What exactly? For people who don't know what, what is it and why is this becoming even mentioned on a Netflix show? So that yeah. So that's, this was a sign of success when it was finally mentioned on Netflix—no, the Lincoln Lawyer a couple of years ago, someone sent it to me and it was mentioned in in one of the courtroom scenes.

6:12 So randomized linear algebra, randomized numerical linear algebra is basically an area that uses randomness as an algorithmic resource to solve linear algebra problems. A lot, a lot of linear algebra problems are under the hood, whether you're doing machine learning or scientific computing, they manifest themselves in different ways. So the exact questions and numerical issues and so on are, are different in those areas. And that's the source of maybe tension in the area, but also synergy and, and, and, and, you know, part of the reason a lot of people are interested in it because it, it holds the potential to solve a lot of

6:48 problems people look at, But it's, it's solving core linear algebra problems and, and linear algebra ones appear in a lot of places. If you're solving partial differential equations, it's, you know, linear operators or iterated forms of linear operators appear. If you're solving machine learning, you may be interested in support vector machines or ensembling methods or these days deep neural networks. And so matrix multiplications are at the core of a lot of that stuff. So very core linear algebra problems. I think historically the way people thought about the relationship between linear algebra and and randomness or

7:21 noise was that there's randomness and noise in the world, meaning the data you measure, think of least squares and it's your job to clean it up. And then you call a linear algebra problem least squares or low rank approximation. And you get more or less a deterministic answer and you get more or less the exact answer. I mean, and I say exact in scare quotes because there's numerical issues and you can't represent square root of two on a computer, but you know, in so far as machine precision is exact and you get an exact deterministic answer. And for a lot of things that's just overkill. And so you can use randomness as an algorithmic resource, meaning

7:58 inside the algorithm to speed up computation. And this may be most notable historically, like in Monte Carlo or Markov Chain Monte Carlo, where you run simulations of of fluid dynamics. I mean, the Metropolis algorithm was developed in that context, but this is for core numerical linear algebra problems. And so most people, if they're running computations, sit on top of BLAS and LAPACK and, and, and related software. And if you're calling Python, you're calling something else, you're calling something that's calling something that's calling them typically. And so these are core libraries. And so the question is, can you,

8:36 you know, use good theory from randomness and measure concentration and high dimensional probability to improve those algorithms or improve, you know, variance of those algorithms? The short answer is that you can. So how? But so how does it work in practice then? So if I'm solving a PDE and I'm normally using some sort of linear algebra library and it takes me this time or this amount of flops or this compute, how is it that the randomness reduces that? Yeah, I mean, when you're solving a PDE, you're typically solving it for a particular application and you're calling certain core linear algebraic primitives in in one way or another.

9:23 So for example, if you're solving it with a splitting method or predictor corrector method, or you're solving a finite element or finite volume or you're solving a second order optimisation, too, as a piece of that PDE solver. These have least squares, you know, linear solves things like this under the hood. So typically you're improving those. You could, you could improve the modeling setup for the PDE also. But but you know, you could, you could try and solve the, you know, the improve the core primitive. So take least squares as an example, a low-rank approximation. A common motif in in a lot of these algorithms is you have the

10:00 data and you want to solve the least squares of the low-rank problem. And you could solve it exactly machine precision or whatever. But you might want to on the other hand, get what they call a sketch of the of the data, which is roughly a small number of data points. It could be a small number of actual data points, or it could be what they call a random projection. And if you're familiar with like if you're familiar with signal processing and electrical engineering and, and physics, think of this as like a randomized version of a, of a Fourier transform. So it takes some signal that that might be localized in space

10:34 and, and spreads it out everywhere. But you don't need this to be a physical space. This is just a linear algebraic problem. And so you have, you know, you have columns and rows, so you could select the important columns or rows, or you could do this random projection, which is essentially a, a basically a random type rotation that spreads the information out and then sample uniformly in that rotated space. And so you take this sketch and you can do one of a couple things that so some communities, the more theoretically inclined communities want to take that sketch and solve the subproblem exactly. And there's a range of

11:07 theoretical work that says the solution to the exact, the exact solution to that subproblem computed any which way. A traditional solver or something else is epsilon-close to the exact solution to the original problem if you set things up right. Now, in a lot of cases that is sort of coarse because you can't. It's hard to get machine precision just by drawing a sample and solving the sub problem unless you take a huge number of samples, just because Monte Carlo methods tend to converge slowly as a function of the error parameter. So you could take that, you know, low quality, low precision, but not trivially bad solution and ask yourself what

11:44 is a preconditioner? And by preconditioner I just mean a preconditioner for a PDE solver or or least squares or a linear solver. And a preconditioner is basically a low quality solution that you refine. So you can actually take this preconditioner that the that this is a sketch. The theoretically inclined people just say good, done; I'm epsilon-good. And you can say, now I'm going to take that epsilon-good where epsilon is 0.1 and just use any of a range of traditional iterative solvers and drive that epsilon down to 10⁻⁸ or 10⁻¹⁶. And, and just like most preconditioners, if, if the cost to construct it is less than the

12:20 cost to, you know, the cost to construct it plus iterate is less than you know, the, the algorithm you're competing with, you win. And it's a good preconditioner. So there's a range of ways people solve it that way. So those are called sketch-and- solve and sketch-and- precondition, respectively. Increasingly for machine learning applications that want medium precision, also relevant for scientific computing when you're interested in low precision data representations, you know, going to half- or quarter precision, which is, is increasingly seen in hardware. There's a more subtle interplay where you can get the solution

12:54 and then you can iterate it and get a second sketch and toggle back and forth. And if you set the parameters differently, you can get intermediate solutions 10 to the 10⁻², 10⁻⁴ and 10⁻⁶ quality. So there's a range of ways you can use the sketches, but roughly the idea is you get most of the information in the sketch and and solve it or do something with it. And is there any, you know, orders of magnitude of the savings that you could get? You know, if I'm doing a Navier– Stokes solver or solving some neural net using these randomized, you know, numerical linear algebra versus BLAS or LINPACK, are we talking, you know, a 10% saving?

13:33 Is it an order of magnitude or is it still under research? And so it's not delivering the full yet. Yeah. I mean, I think the question of of how well you'd improve upon something depends strongly on what your baseline is. And if you feed this into a big PDE solver, there's a lot of moving parts and you're competing with very mature code and they're doing a little bit better or even a lot better on a solver may or may not matter for the downstream use case that that the scientist that is looking at the PDE solver is interested in. Or go to the other extreme. And you want to say, I want to compete with BLAS or LAPACK.

14:10 These are extremely optimized pieces of code and it's, you know, it's hard to beat them. So one of the big successes in the area was with Blendenpik. And then soon after that was LSRN and Blendenpik, which said we wanted to ask whether these randomized sketches, you know, not in theory, not Big O notation, not whatever can they beat LAPACK Because this is not boutique code you have or I have or a solver that you don't share with the community. So well, this is something that's been stress-tested for decades. And the answer is yes. I mean, basically on any tall dense matrix, you can beat LAPACK with these techniques. You got to set parameters right

14:47 and be a little bit careful. But, but the short answer is yes, how much you improve it by depends on the aspect ratio and depends on properties of of the matrix. But think of it as ballpark 2 to 10. Now this was back in 2010 and so a lot's changed since then, in particular in the hardware landscape with respect to GPUs and hardware heterogeneity and the sunsetting of Moore's law and Dennard scaling and these sorts of things. And so it is the case that you know, you can be a factor of 10 better. Similarly in in in low rank type approximations. I think the the the way scientists versus machine learners use low rank

15:24 approximations is very different. Scientists tend to say low rank means you know 99% of the Frobenius norm, meaning 99.9% of the Frobenius norm and 99.99 I mean basically the whole matrix. When machine learners say low rank, they mean sorry, sorry, the spectral norm. When machine learners say low rank, they mean, you know, 50% of the Frobenius norm, in which case, you know, you lose a lot of information, but you might iterate and and do something better. So think of sort of scaling laws and the neural scaling context with, with neural networks. And so the way you'd use those low rank algorithms and feed them into other solvers would be very different.

16:03 And then that would lead to very different levels of improvement. And so I think that's largely open. And then with the memory wall that you're increasingly seeing in, in general, but most egregiously in the machine learning applications, one of the big wins and it's just starting to be explored for the randomized techniques is basically improving the memory properties. And so this could be just reordering algorithms in the sense of communication-avoiding linear algebra could be much broader because the randomness sort of decouples things, makes it easier to parallelise certain sorts of computations. And so you see much bigger than factors of 2 or 10 improvement

16:36 there. But then again the baseline's changing year to year as as as people explore different representations and and low precision and intermediate precision and stuff. Interesting, so how was it mentioned in Netflix then going back to the former? Well, the particular I, I think so, I'm not a movie star or a movie producer, but I, I have a sense that they, they want to, I, I mentioned this to someone and they said, you know what that means? They tried to take the most exotic thing that wouldn't make sense to anyone and, and, and, and cited it. So the context was I, I think it was early on and the, the, the main character, who was this,

17:12 the attorney needed to represent someone and the person had to get out of, of, of whatever they were charged with, which was not a, a, you know, relatively minor thing, but it was the 10th time they did it. And so they, the, the, the Lincoln lawyer said, why do you need to get out? And they said, I have a defense on Thursday. I need to get out. And they asked what's the, what's the topic? And they said randomization and numerical linear algebra. So that was the context of how it was mentioned there. I wonder how they found that out. I wouldn't know how they actually. Yeah, I don't know. They they, they, they didn't call me up.

17:45 So they must have had someone in movie producer land or something that did did a Google search on on something exotic and came across it. That's fine. Yeah. That reminds me of the Did you ever like Star Trek? Yeah, I used to see if there's a bifurcation which maybe would get us in into a separate discussion whether you like the original version or the the second version. But then there was this explosion of different, different versions of it. Yeah, I was more the next generation person, but I the reason I mention it is it was often whenever they wanted to explain how they were breaking the laws of physics, some stupidly complicated phrase was used.

18:26 And I'm sure it exists somewhere and it just reminds me of that, you know, explanation of how they're exceeding the warp drive. Oh, it's because the something something. Yeah. So I think it was a little bit like that. I mean, I, I think it's a sign of the times that what they're citing is not warp drive and quantum gravity, but randomized numerical linear algebra. This is a core statement about how areas are progressing and so on. Exactly. Exactly. Yeah, yeah. So maybe one of the topics would be good to chat with you about, and it's definitely a hot topic at the moment, is the the foundational models for science.

19:07 I've seen quite different viewpoints. There's one argument that is, well, large language models take all this data from the Internet. They train and they're remarkably good at doing what they do. What about if we could do it for scientific disciplines? Wouldn't we be able to just, you know, ask a prompt to go and simulate me a plane or, or, or do something? And then there's the the other avenue, which seems to be more the replacing simulation tools or augmenting simulation tools like GraphCast or FourCastNet or, you know, these sort of seminal pieces of work that have come out, but they were for very, they had to be trained, you know, for something.

19:56 Where do you see this? Where do you see this going? Are foundational models possible? Is it we just don't have enough data? This is a big topic. So yeah, I'm not sure how we want to start it. I mean, we're in the middle of this fray and I guess that's why we were asking about this. And I should start off with saying I don't know what a foundation model is. Actually. I know what different people say a foundation model is and, and it means very different things to to different communities. And so I think articulating that a little bit helps articulate some of the possible directions and some of the questions you're

20:38 asking. I mean, I think one way to think about a lot of these machine learning methods and it's highlighted particularly cleanly with these foundation models is machine learning methods. And in particular, these foundation models are, they're an infrastructure, right? The, the, the, the Stanford report said, call it foundation models, not foundational models, because it's a foundation in which you build things and as and as opposed to a half dozen other terms they could have used. Now, whether that's right or not, I mean, but that, that's what they said. And so since then people have said, well, it has to be this big and that big and, and

21:13 whatever. And, and then people want to get a foundation model, not for all of science, but for areas A, B and C and subareas D, E and F and get finer and finer granularity. So in a sense it's, it's a it's a term. And, and if you ask people, most people using it what it means, they sort of will acknowledge that, that they don't know. And it's not a standard definition. I think probably the best way to think about it is it's, it's, it's infrastructure, it's a foundation on which you build things. And so the computer's infrastructure, I mean, you can do different things with the computer than you can do with pencil and paper.

21:46 And so it wasn't obvious the way computer science, even before it existed, would evolve. And at one point, post World War 2, the US thought they'd need five computers and that would be it. And then it would solve all the nation's needs, right? So it evolved in ways different than expected. And it evolved in particular because it, it could complement what people did. I mean, you could ask different scientific questions than you could with pencil and paper and you could ask different engineering questions than you could. I mean, now you can simulate whole things and, and the very, very last step is building it, right. But in addition, you could do

22:17 lots of other things that were driven by industry. And, and so there's a bifurcation in the area whether you're doing more numerical things that are continuous or whether you're doing discrete things that, that, that are not continuous that were driven largely by business needs. So I think the, the right way to think about this is it's that sort of infrastructure and it's tied to the data more closely than than, you know, database systems because the foundation models have information cooked into them. And so it wasn't obvious some number of years ago that just with language you, you could do what you know, has caught the

22:47 popular attention in the last couple of years. So I think so to anyone who says that confidently, you, you can or you can't with science, I mean, probably you can't say that they're sort of reliably saying what what will hold for the future 'cause I think it's not at all. I think it's not at all clear for two reasons. One, there's certain technical things that really do matter. I mean that that are different about scientific data than, you know, language and, and vision data. And then there's just very different cultural, there's very different cultural things. I think every scientist thinks that their data is special and unique just like they think that

23:21 they're special and unique. If ML's taught you anything. These algorithms can predict what movie you'll watch and what stuff you'll buy better than you can. And so you're maybe slightly less unique, you know, than than you thought in some sense. And so I think there's a question, you know, what does foundation model mean and how can it be used so narrowly? You know, I have a big model I can train it in in domain A domain A could be weather and climate have gotten attention, but it could be fluid dynamics. It could be something else about stellar formation. It could be properties of materials and doing density functional theory and so on. And then there's a question

23:54 about whether you could maybe learn broad based models that cut across domains, which would of course be the more interesting thing. So there's something, Chronos said, I was involved with the AWS people that sort of says, I, I don't want to be the best at at the most extreme things. I don't want to be the 1% that's predicting the most extreme things I want to be. I want to do as well as as the majority of users for for time series analysis and oversimplify the story. I mean, time series is a complicated area clearly of interest in a lot of cases, but it's it's a little bit, you know, you got to be careful about off-by-one errors and a range of other technical things.

24:29 And what that says basically is take a language model, language model structure with a bunch of time series and just mean-centre it and variance-normalise it. So just do the simplest possible things you could and you got to do some data augmentation stuff and boom, you know, you do sort of comparably and, and better than a a wide range of public benchmarks. Now, the public benchmarks are not extraordinarily high because a lot of time series data is valuable and so companies tend not to release it. But but the fact that you can do that just by sort of variance- normalising and and mean- centring is is, you know, not obvious and sort of interesting.

25:05 It turns out sort of on a separate thread that you could say, could I do well compared to the 1%, you know, the most extreme things and, and the the answers that you can there too. You got to use different techniques than you do there. And so that leads to questions as to whether you could have a foundation model for sort of a broad swath of, of time series and forecasting analysis. And so, as you know, I said, I'm an Amazon Scholar working with the Scott team and we're looking at that there. And one of the interesting things I think about there, if you look at the details of the model, not the Chronos, but some other things, why would you

25:42 expect that articles scraped from Wikipedia, you know, public, publicly available language models? What would, why would that allow you to do better job predicting dog food demand? You know, I mean, it's not obvious they have anything to do with anything. You know, it turns out, I suspect the hypothesis is that the text, the linguistic structure that you're learning from the Wikipedia articles is strongly related to sequence to sequence modeling, which, you know, if you think of 1 dimension as a metric space, it's a very, very special metric space, a very specially structured metric space. And so you can learn not just recent information and not just

26:19 low frequency information over the past, but maybe in information at different scales. Because, you know, sometimes articles in text refer back 10 words or 100 words or 1000 words or 10,000 words. So you can learn these sort of heavy-tailed sort of structures and that gives you a better set of basis functions to learn dog food demand or whatever else. So in a sense, these, the language models give you better embeddings, you know, Fourier analysis and Laplace transforms and, and these sort of things are great for pencil and paper. They're all developed in the 1800s, right? These are data-driven embeddings that are good if you have

26:54 computers and and data not necessary for pencil and paper, but they're good for that. And so that leads you to the question about if, if I'm really careful about doing foundation model work in science, what do I need to worry about? It's not the details of this PDE or that PDE. It might be that I want to reproduce what happens in the natural language processing (NLP) models, which is roughly scale model and data and compute. So none of them saturate. That's very different than the usual strong scaling and weak scaling and high performance computing. I want to change the amount of data I'm putting in. I want to change the size of the

27:28 model. I want to change the size as it computes. So none of them saturate and, and the conjecture would be that if if any of them saturate, it's going to be harder to do transfer learning. So you really need the non saturation and you do that one way with NLP (natural language processing) and CV (computer vision). But in NLP and CV, there's much weaker control you have and that you need on the spatiotemporal geometry than PDEs, right? Is, is, you know, PDEs, if you're below a Mach transition or some other physical transition, things are sort of smooth. If you're above it, things are very messy. Dealing with that transition is hard and people spend their

28:04 whole careers dealing with that. So can you come up with data generation or tokenisation mechanisms that that respect the spatiotemporal properties? So I think if you're going to have a foundation model that applies across a broad range of sciences that's trained on weather and, and, and climate and, and simulations of different grid structures from the machine learning perspective, it's not so different than satellite data from satellites looking down, you know, it's a different grid and it's an image. And then you say, how does that couple with the spatiotemporal properties? So I think if, if you want to deliver on the promise in the same way as computer

28:38 science had to work out a range of numerical methods to really match and beat state-of-the-art, you're going to have to do a similar thing here. And so partly depends on these technical issues, but partly depends on cultural issues. But how much do you think? I mean you raised the 1% versus the 50%? I guess that is one argument as well, which is how close does it need to be to machine precision to be useful. You could argue that any of these large language models, the way that most people use them, there is still a correction you need to add. You don't typically ask it to write a document and literally take it word for word. You usually go in and go.

29:18 That's, that's remarkably close, but I'm still going to go and fix it. Or it writes you Python code. It's unusual, isn't it, that it's perfect. There's usually something you have to correct. So I guess with the science side, maybe the argument is how close does it need to be to be useful And the cost of getting the incremental increase in accuracy, you know, is it, is there a like a trade off where you need so much more? Data, I mean, again, I think the best analogy is look at the history of computer science and how that evolved, I think, I think framing the question to say how close does it need to be to be useful? You're already framing it in a

29:59 way that that makes certain one group comfortable, the numerical analysts and the PDE people that frame a question a certain way. That's not how people would have framed the question before. I mean that, you know, before you had, you know, represented continuous numbers on, on a computer discretely take a step back and say, I want to solve a certain problem. And, and the question is now that I have very different trade- off in terms of compute versus data and I have this new infrastructure that's language models. Can I ask a different question and, and, and push the science for it? I mean, so in, for example, in,

30:39 in chemistry historically, but also nuclear physics and, and in fluid mechanics, there's this notion of a semi-empirical theory, which is a theory sort of derived, you know, it's not just curve fitting, it's derived from an underlying more fundamental theory, maybe phenomenologically with parameters that are then fit empirically or semi-empirically. And So what you need there is not the theory to be right, but sort of right enough at the level of in the, in the chemistry is chemical accuracy, which is, is however many kilovolts or whatever depends on the reaction you're interested in so on. And so certain methods that

31:11 maybe were more principled were a bit too coarse for that. And other methods, when you combined it with other techniques achieved chemical accuracy. And so none of them were low enough in the stack that you they were better QR codes for QR from, from linear algebra. But they solved the downstream problem at that level of actually, probably what you'd see here is that right? So it's, it's probably not going to be so successful to say I want to get 10⁻¹⁶ or 10⁻² of my large language model, but but a foundation model for science need not be an LLM. It, it, it, you could be literally learning the embeddings from a range of

31:42 PDEs. You could, you could train to PDEs of different types. And you could say, I want to port if I, if I know the basics of hyperbolic and parabolic. And I know that in the transport equation, this can enter in a certain way. And you have nonlinearities and certain types of forcing functions, you know, model each of those and and solve each of the components separately. So I think the question is, what's the right level of abstraction to do that? And, and one of the modalities you could use is language, but of course you could use image or PDEs or simulations, any of a range of things. So I think in the context of, of scientific problems that the

32:13 foundation model need not be an LLM based. I mean, there's some people are pushing that, I think, but but it certainly need not be LLM based. We just query the model and say, you know, tell me what to look for, the top quark or whatever. Well that that's the bit maybe you brought up would be interesting to explore. There is, I would say, a relatively fierce or contested argument around physics-informed or physics in it versus data-driven. Where do you sit on that, on the argument? Do you, do you? Does the model need to implicitly have some of the boundary conditions, some awareness of continuity of energy? Or is it good enough just to

32:57 give it so much data that it essentially learns those laws by itself? Yeah, I mean, I, I think you could ask the same question about in natural language processing, computer vision did you just, I mean, there's a, there's a common story people tell which is just get more data and everything works. But that ignores the fact as you know, that a lot of companies and universities, a lot of people have put a lot of resources into NLP and CV to try and figure out how to make it work. And you know, convolutions convolve things, they spread stuff out. You know, if you have an image in 2D that may make sense, you might want to average over

33:38 nearby things. But if there's, if there's a sharp corner like, you know, the shirt against the background wall in this picture of me, you may smudge stuff out. And so there's a range of ways to try and sort of deal with that. That's core structure. The 2D structure that you're trying to average over is very different than sequence to sequence learning. That's also much more discrete in, in natural language processing. So it's not like computer vision or natural language processing, just learn stuff. You gave it very particular architectures that had very particular inductive biases and then you're careful about the

34:08 data and, and trying to get gobs of it and, and then then it worked. So I think if you, you know, you'd have to follow the same path here. You can't just take an existing architecture and press a button and hope it works. I mean, time series, forecasting space—temporal modelling, these sorts of things will have to be, you have to figure that out, how to do that. So I think the question, I mean, some people, you know, I think are, and, and maybe sometimes people that have a vested interest in, in moving forward or not. I mean, some people say, you know, just ignore everything, let the data do it all. Some people say, Oh, you'll

34:41 never match what I can do with my careful PDE solver. I mean, I think that's neither of those questions is the right one. I mean, the question is, given the fact that there's an infrastructure of code and experience numerically, but but you're also at a very different place. You've been sitting on top of certain essentially hardware and, and, and linear algebraic advances. I mean, for years, for decades. People working on PDEs call certain linear algebra libraries and have essentially had not to pay the technical debt associated with the fairly high level of of complexity with developing new linear algebra because just wait a year or two

35:15 and the machines are faster. And so if, if, if that's not the case, you're going to have to start thinking a little bit more carefully about the underlying linear algebraic computations. Maybe there's different trade-off points in the space and maybe bring something else to bear. You know, a model that has learned certain coarse-type functions of one dimension, say from the language model that I alluded to before that aren't Fourier modes or aren't Laplace modes or aren't something else like that that they're more familiar with, learn a different grid type discretization. So I think the real challenge in delivering on the promise of all

35:50 this stuff is figuring out what the right way to combine those two. So you can imagine rather than just writing down a physics- informed loss and pressing a button saying, well, the way people actually would solve this is to have some sort of splitting method and deal with the two types of physical things and the diffusion and advection in two slightly different ways. And machine-learn one and take the other as numerical. And then you couple them together, like maybe in a differentiable end-to-end. Yeah. The, the reason I ask this is because it does feel certainly at least in the CFD community, that there's a, there's a real sort of split within the

36:24 community. And in a sense that on one hand you have people saying, well, we have all these PDE solvers, you know, developed for decades. And you know what convince me how your machine learning or AI method is going to be as good and faster and cheaper versus on the other hand, you have quite sensational claims of, you know, 10,000 times faster, you know, sort of real-time. Well, of course, when you dig in, you know, you start asking the questions, well, how much data did you have? What was the time to create the data? And but there is a bigger question, I think more when you look to the future, which is, is

37:14 it just a matter of time? Is it a bit like in the 1980s, they could only simulate a plane to a reasonably low accuracy because the computers just weren't powerful enough. And as you said, you just almost keep the same effort almost. And you just literally add 20 years of compute on top and now you can simulate. Or is it a we'll never get there because we'll it's like an it's like impossible to have that much data? This is the bit I think a lot of VCs and start-ups are also asking that question if we pump 100 million into this or a billion, will we fix it? Or is it a trillion dollar problem and it's not worth doing? Yeah, Yeah.

37:59 I mean, OK. So I think there's a there's a bunch of things in the question there. Sorry. We'll we'll CFD people to BC and and you know that they have two very different utility. You. Know, I think, OK, so it when someone says to me do this, convince me of that I have these metrics that have been I've worked on for decades, I just say thank you, good to talk to you. I mean, because, because there's no way you're going to win that battle because they're very fine-tuned metrics. So they're very one little use case. Even if you do win that battle, they're not going to admit it because they'll, they'll tweak their method and, and you know,

38:32 be slightly better than you. And there's greener pastures few everywhere. And this is not a hypothetical statement. This is a very practical thing. You're asking about this in the context of PDEs. We saw this before in randomized linear algebra. I mean, so literally, I remember it was very clear that the techniques could potentially be useful, not just in theory, but in practice, and that the reason the numerical people were saying, oh, they won't work, You know, really. I mean, those are those number of methods that clearly wouldn't work. But but then there's the next generation of methods. And that's sort of when I

39:01 entered the the area and it was clear that they could potentially work and the objections people had before wouldn't work. And so this was right around the time of the Blendenpik paper that I mentioned that that basically said, can you beat LAPACK? And the short answer was yes. So now, for example, if you look at the SIAM Linear Algebra meeting, half the talks are on this topic, but there's been this process of creative destruction where this they're still asking the same questions, but putting randomness there. And so that's one group of community. I mean, a very different community is the is the people that say, I'm going to not do

39:28 that. I'm going to go and apply it to machine learning and 10 other problems. And, and, and so there's that tension I was alluding to before. So I think if you, if you try and, and satisfy an old school metric, it's, it's good to have a few examples of that as a proof of principle. And, and this is 15 years later, and I mentioned it twice here, right? So this was a clear proof of principle of the area. The area sort of accelerated after that. And there's been a lot of theory and empirical development. So I think there's been a few things that, if not that are starting to look almost like that in this general scientific ML area.

40:04 I I think. History is moving in that direction. So it's clear then that, you know, randomness would be important for linear algebra. Now it's even more so for all the historical and hardware trends you were talking about. I think likely the same thing's going to happen here, right? And so a question you could ask is, are the people who know the linear algebra, are the people who know the computational fluid dynamics? Are they at the table developing the methods? Are they just going to pooh- pooh it and say, well, they're not going to work because it's not going to satisfy my particular measure as opposed to here's a broad class of

40:34 techniques and it's hard to imagine that there's nothing in my area that will be improved by them. And so I think that's the that that latter question is the question to ask. And that latter is the question the questions VCs and and other people will ask. And I think the, an important maybe determinant of how this will evolve is, is how the different players interact with this in, in the sense that a lot of the machine learning, more broadly, scientific machine learning, not just for the foundation model, but more broadly machine learning really is, really is a horizontal. I mean, it's, it's designed to say, what can I do with lots of data and be relatively ignorant

41:15 about your application area? Because who knows why people click on ads or click on links on their social media account, right? I mean, you can tell the story, but really who knows why? And so a lot of machine learners implicitly or explicitly will say, I'm not interested in solving your scientific problem if it's just a one off problem. I mean, if you're a scientist using machine learning, you might be interested in that, but that's because you're interested in the domain. But if you're a machine learner, you might say, I mean, I see some commonalities between your fluid dynamics and this other solid-mechanics problem.

41:44 And I see some similarities there between that and something in in 3D in the atmosphere as opposed to, you know, something very different. And so identifying those commonalities I think is important. What I see in, in in some cases, and this is hampered by universities and it's hampered by industry, and it's hampered in a couple government labs in, in a couple different ways, is a lot of people want to solve a scientific problem. And so they're going to say, I'm only going to invest, invest literally or figuratively in machine learning in that one area. And that's just not how machine learning works. You don't see the, I mean, think

42:19 of it as horizontal business with horizontals and verticals. If machine learning is a horizontal, like high performance computing that will solve a wide range of problems, you got to apply it to a vertical to justify it. That's that's where that that's where you know, the rubber hits the road typically. And so you don't see the successes in machine learning, scientific machine learning at vertical companies, you know, companies working in genetics or or oil exploration or any of a range of other particular domain verticals. You see it in the horizontal companies. The companies have built out the technology infrastructure and they get a team of people that

42:53 know this is a vertical and applies it there. And then, you know, a team that knows a different vertical and applies it there. And so I think the question is, how does that play out? And that's going to be a non- trivial evolution in terms of. And your. That day. Your comment was interesting as well about the difference between sort of Transformers versus CNNs and you know, if you just throw the whole world's Internet at CNN, that isn't you needed to have the right architecture to be able to throw a lot of data at it. I guess the question is more for the science. Is there a common architecture across the sciences that then

43:33 would allow you to sort of build out that big which could be applied to verticals? Or is it the case that the needs for weather or genetics or chemistry are so different that you would need a different architecture for for each different science? And therefore the, the, the dream of a true foundational thing. Besides, it will never happen because you know. Yeah. So that's the big question. And I, I think there's a technical component to that and then non-technical component in terms of the areas, right? And so technically it's open-ended. I don't know, right? I mean, we're, we're in the middle of the fray and we're, we're voting with our feet and

44:14 placing a bet there that the, that the answer is or could be yes. To try and address that, you're going to have to deal with the issues we were talking about before. You know, one dimension is very, very special. two dimensions, the surface area to volume properties are very different. Three dimensions is very different. By the time you get four in scientific computing, that's sort of infinite, right? Machine learners, by the time you get to a million, you know, it still seems pretty small. So mapping the methods to one versus two versus three dimensions, open question, maybe there'll be no trade-off point where the surface area to volume and the

44:47 isoperimetry and other things like that wins. I suspect not. I suspect it could, but I mean, open question. And then non-technically, if you look at how computer science evolved, numerical analysis and scientific computing was core to computer science. Look at how numerical analysis and computer science evolved. Yeah, these areas gave birth to computer science and then computer science banished them because it was easier to build the structure, the theory around, around Turing machines and, and, and the Lambda calculus and, and discrete notions that talked about complexity independent of the data. And if you're talking about

45:17 complexity independent of the data, you don't need preconditioners because it's independent of the data. And so there's a cultural aspect that that says, you know, does the vertical or the horizontal culture win here? And so even if you stipulate that the answer is technically yes, I think there's an open question on how it'll look, it'll evolve and whether that'll win the day and solve the problems that you're asking. But I think it's not, it's not so much that I solved PDE type A and applied it to PDE type B. It would be certain coarse things. I mean, as you know, there's a class of things for which finite element methods work. There's a different class of

45:49 things for which finite volume methods work. And you know, you, you take a course, you learn a 101 method and then boom, they bifurcate after that. So what's the analogous taxonomy here? And it's probably neither of those. It's probably something that's more data-driven, but what's the right way to taxonomize that? And even if there's a technically a correct way, I mean, you know, computer science is not computational science. Computer science is a coherent area. Computational science is spread over many departments and many divisions. And so I think it's an open question about how this area evolves. Yeah, it is certainly.

46:21 It is certainly interesting. I think one of the sort of key topics to all of this though is the data availability and licensing and financial reward. So one thing that I've seen in I guess the early days is nobody knew what was going on. So I guess the companies were scraping half the Internet. There were a few CAPTCHAs or there was a few sort of, you know, firewalls to stop things. But in general, it was a sort of free-for-all it seemed, Whereas now people have, you know, cottoned on to that and some of the companies having to do deals with publishing companies or do deals with, you know, to try and legally get access or or permissively get access to

47:10 things. I saw the example on the weather side that, you know, some of those databases were released openly because they were released openly enabled, you know, these various teams to go and do things like, you know, FourCastNet and GraphCast. But those data may not be open in the future. So how much you know, and you work in academia, you see people want to have their own data. Some people want to release later. Do you see that that will ultimately be the blocker? I mean, how do you incentivize people to want to make their data available and should it even be done that way, you know, that the whole open source data versus closed?

47:58 You know, this is this, I think, is the elephant in the room in terms of this area because a lot of government grants require that you make the data publicly available in some sense of the word. Now, what does in some sense of the word mean for a lot of scientists? If the data is more than a year or two old, it's worthless because they're trying to push the cutting edge of the science and then they don't maintain the data for more than a couple years ago. And so no one maintains and it's hard to access. So there's a lot of public data that's not de facto public or accessible, you know, in a lot in those cases and maybe in, in, but, but certainly in, in

48:31 machine learning sort of applications more generally in, in Internet and social media and so on. Having older data is good, but you can't get more than a couple decades older, right? And so oftentimes the the most recent data is the most valuable and you just design a method around that. And so it's not obvious to me that the way the data are generated and, and collected, distributed now will be the one that is going to be developed going forward. I mean, if you get better telescopes, if you get better, you know, molecular sort of robotic systems to, to to make molecules, I can easily imagine a situation where data that exists now or prior to a few

49:10 years ago is basically obsolete and irrelevant. And so all of that data sits at whoever generated it. And so then the question is, is does a government lab generate it? And if so, do they make it easy de facto to access? Is it done at universities where, you know, you have a big infrastructure capital investment in in particular, essentially in particular science and verticals to generate it? Is it done at companies that are in different verticals or is there some sort of collaboration or interaction with, you know, horizontal infrastructure companies, hyperscalers and then they get their fingers in that and have access to it?

49:46 So I think these are all open questions in in terms of how this evolves. Yeah, it it seems to be very much, you know, traditionally and I guess if this is for me, the difference between the pretrained, quote- unquote foundation models versus just letting someone go and train them. Although I mean, I guess most people you and I would not go and train our own LLM. We're just going to go and use someone who's audit and we're just doing an inference on it. That's and I could see if I'm BMW and I have data from previous car designs or simulations or whatever, how they could train a model using their data. But I'm sort of wondering in my head, how could it ever become

50:34 the point where some person has access to not only BMW's data, but Fords and Audi's and Volvos and Boeings and Airbuses and you know, ExxonMobil and Shell and everyone else. Like how how could they collect the data? What would be the incentive for that to happen to or that's what I mean. Or is it the case that that's why there will never be a foundation and it will just be that individual people train their own models because the data is just so sensitive and difficult to collect in one? Yeah. I mean, this is, I think this is the big elephant in the room. I mean, it's not clear to me that one. OK. So I think if you were to try and generate a meaningful

51:17 scientific foundation model, you'd need data from a pretty broad range of use cases. And so you gave, you know, several that in particular are with a strong sort of financial backing. I'm always reminded of the story. Jim Gray was, I guess, the most prominent person in the Microsoft Research division years ago, such that, you know, he would have to have regular meetings with Bill Gates, the CEO at the time. And Bill Gates would say, why is the most prominent person in my research division applying database techniques to astronomy? This is when he started working with Alex Szalay and some others on, on databases for, for

51:56 astronomy and, and cosmology. Why don't you work on something more useful like, you know, any of our products? And, and Gray's answer was basically, I'm working on it because it's worthless, because there's no financial sort of incentive behind it. I don't have to talk to the lawyers and we can work out IP and all these things. It's public data and, and I can work on database techniques and figure out latency and bandwidth and all these trade-offs. So I don't know, but it would be interesting to see. Will you have an analogous thing here? So there's there, there is public, there's, there's various places where public data is available.

52:28 And I think if you can illustrate that you would have a proof-of-principle foundation- like model. Then the question is, will people come to it? Now clearly the incentives for scientists will be different than than for companies there. But I could imagine a situation where you build that out and then companies could take it and fine-tune to their particular data, right. It's not obvious to me that any of the companies that you mentioned have data that that's, that is that special in a scientific sense? I mean, clearly the particular data they have is particular to their particular situation and then the particular, you know, product they have.

53:04 But I, I can imagine that, you know, the PDEs from early solar system formation and, and hydrogen bomb explosions and hydrodynamics inside stars, you know, is, has some similarities and some differences with weather and climate and some similarities and some differences with crack propagation in the Earth. And that'll have some similarities and differences with, with various sorts of sensing modalities that you would, that have a geospatial component. And so can you stress-test that and, and get even even a proof of principle that, that if you would attack those, then, then the answer would be yes, that, that you could have a path to have a more foundation model.

53:41 If the answer is yes, then the question is you talk to various stakeholders. Some of them are government entities or government sponsors, some of them are corporate and try and figure that out. And I could, I could imagine that that basically it's too big a lift in the area stalls. So I think this is something to be determined. But I think having a nucleus of something is going to be important because right now I at least what I've seen, I mean, in most of companies like that, there's, there's, there may be some people a little bit intrigued, but there's not a sufficiently heavy lift that they go screaming at their CEO

54:11 to say we need to do this, probably because it's not extremely strong baselines that it can happen. And their competitors by talking to these public models and fine- tuning on their proprietary data will do better. I think at that point they'll be you can see a change. And I think this is, I mean, I guess this is affecting the current generative AI, which is how do you make money out of this? You know, like you can do it for the sake of science. That's nice. I developed this thing. But I'm always interested to see like the money that you put in to create the data to train the model, can you ultimately make a profitable thing out of it that

54:53 justifies its cost? Now, with standard, I guess, simulation software, there's an argument, well, if I simulate it, I don't have to build it. If I can predict the weather, I can help people figure out if it's there's going to be a hurricane. And I can say, you know, there's a, there's a clear thing. But because we can already simulate things, we can already predict the weather. We can always do it. For me, the machine learning bit is just the argument, well, it's faster or, or, or it's cheaper. That's where the amount of data do. Do you know what I mean? Is there a commercial? But I, I think, I think there's a faster question and I think

55:35 it's tempting to say, can I be faster? Because faster means there's a number you're comparing to. So am I bigger? Am I more on some axis? And back to the, the, the randomized linear algebra example, I mean, I, I was talking to various people and saying, don't worry about the randomness. You know, if, if you do this and that, that'll work. And the people would, you know, basically when you have a new idea, they'll beat you up. And so they beat us up and so on. And a few years later, one of those people in particular had some methods and showed it worked, you know, not the blend people, you know, something around the same time and their students came back to me and

56:08 didn't know who I was. was lecturing me and saying, oh, don't worry about the randomness, it'll all be fine. And so on. There's just a cultural shift. And so there was a cultural shift that really was not, is as faster as this is as better. And so I, I, I don't think you're going to see the win here just because something's faster and better. You might need that to point to, but the win will happen because you can solve other things that you couldn't solve before. So just as an example, say you're, you know, you want a better battery and you want to understand the way crack propagates better, you know, so that I mean, you want, you want, you want a longer lifetime, you

56:40 know, warranty. So you, you're working on developing batteries. One of the ways this fits in, not in a scientific, but in an engineering or commercial context is you want to write a warranty to say that that the batteries are going to last so and so long. And, and otherwise you'll, you know, you'll eat the cost. And so can you predict battery lifetimes. And maybe if you have a certain crack propagation heterogeneity structure, you know that it correlates better with how long the batteries will last. So now there's a fair variability and a particular run on the extreme tails. So that sounds sort of very different than I want to do oil

57:13 exploration and better fracking or something. I send sound waves down into the earth and I look at how they bounce back and I try and learn something about it. But I could imagine that the embeddings that you get from a core model that cuts across a couple different domains learn spatiotemporal properties in particular of complex systems that have sort of long-range, non-trivial, heavy-tailed correlations, which is what both of those examples have. And so say that you can come up with a model that captures those embeddings. Now, a company here may tweak it to their data, do some transfer learning in the battery space.

57:46 A company over here can do a better job adjusting it to their data, which is, you know, spatiotemporally very different in terms of, you know, better fracking. And, you know, for all I know, the same embeddings could be useful for, you know, you had a very different use case, right? And certainly in the United States, there's various places where insurance markets don't exist for floods or for earthquakes or for a range of things like that. I could imagine that you get these coarse embeddings. These are coarse functions. They're not sinusoids and Fourier modes. They're something that's data-driven embeddings. And you say on top of that, I

58:18 want to do transfer learning to the relatively small amount of data that I have. That's very high touch. You know, you have 10,000 stations across the country that measure something in chemicals, whatever I want to do transfer learning to that and I model, you know, model the system, the way information flows in the Mississippi Delta or something. And, and if I do that, I can do a better job, you know, predicting the tails of flood events. Why? Because as a model, I don't know how things propagate, but you know, subtle effects in terms of the density of the soil or, or the moisture of the soil coupled with, you know, 10,000 other things that the neural network

58:55 learns in a complicated way. What's hard to reason about scientifically means that now I can create that market. And that's, that's something that any of those companies I could imagine try and you know, buy out a startup company that does that. So I think you're going to see a win. Not just I'm running faster, something faster than you, but but in a space like that. Yeah, no, I, I, I think what you're getting at is you, what we've seen from current machine learning that it seems to do things that we don't fully understand, but are like you, like you gave the original example of the time series that an LLM that wasn't

59:27 explicitly trained on it actually ended up being quite useful. So I guess we shouldn't underestimate the potential of how some of these things may do something that we're, it's not just about replicating what an existing simulation tool could do, but maybe it could help to merge different disciplines together and see the links and the, and the, and the things between them. I did want to pivot to a different topic, which is when some of the other people I've spoken to, I'm always really fascinated by the curiosity or the uniqueness or the peculiarity of academia and, and the sort of different people's careers and, and, and progression through.

1:00:12 And I think you have quite an interesting one, because you are a very pure. Mathematician, academic, you know, you, you are very much in that, but you also have, you know, the Amazon Scholar position; you're at, you know, Lawrence Berkeley. Was this a a strategic thing in the sense that you wanted to work on real problems? Was this just you find more interesting, like is there, How has it benefited you? And would you advise it to others? You know? Yeah. Yeah. I mean, it's there's a lot of, again, a lot of questions. I mean, at least the narrow questions, universities versus not, yeah. Yeah, there you go. I'll start. That but, but, but yeah, you

1:01:03 tapped enough stuff into the question to make it forward- compatible with lots of answers as it sounds like. I mean, so my, my, you know, OK, so my PhD's in physics, I never actually studied computer science or statistics. And it was in computational statistical mechanics. And at the end of the dissertation I switched areas basically to theoretical computer science to work on originally Markov chain algorithms, some of the stuff I've been working on computationally, Markov chains and molecular dynamics. So it was a technical connection, but it was it was fairly loose. And then the randomization inside the Markov chains that

1:01:40 the people I was working with were also working on these very early versions of these randomized linear algebra algorithms. So I got involved with that and I'd done enough coding that it was clear certain ones are going to be useful and certain ones weren't. And so at that point I decided to switch areas partly because of what I've been working on. Well, two reasons there was a forcing function. This is I just mentioned this because if you have younger people listening, it's good to know that sometimes older people have had chequered past, let's say. So Long story short, I had career ending problems with the dissertation advisor and and the

1:02:11 supply-demand structure in the natural sciences and engineering is very different than in computer science. And so that was was motivation to, you know, jump off the Cliff. I also realized that that the particular work I've been doing in computational statistical mechanics, this is a great area. I loved it and and it's informed a lot of the stuff that I've done since then. But you know, it was a great area to be in 1970 because you know, forward-compatible with a huge expansion. There's going to be Nobel prizes given out. I mean, just lots of good stuff going on. Not so much in 2000. The area, it's sort of sad, but

1:02:45 it was clear that that the future is going to be data analysis and, and I joke that I never knew more about algorithms or data than the day I switched. And then I've just been becoming more ignorant since then. So I don't know quite what I meant by data analysis then, but but but clearly I was right. I mean, it was a fruitful area for the next couple decades. So I'd done work in theory of algorithms and, and with respect to the academic question, if you're desperately wed to an academic career, you shouldn't do this because the academic hiring is fairly siloed and, and it's, it's not a good idea to switch areas after your dissertation and, and so on.

1:03:28 If, if you, if there's a forcing function like I had, or you're interested in sort of problems more broadly, then you can think a little bit more broadly. So I've had a lot of students in postdocs and I sort of give them this advice and you try and carve out various projects that they're interested in. Some want an academic route, so they should work on slightly more conservative things in the general space. Some definitely want industry and some, you know, could go with either and, and the ones that could go with either, I've seen some go one way and some go the other. As a general rule, this is going to be a marketable space and

1:03:56 stuff I do. And so maybe it's a little bit less of an issue, but but, but that's something that they figure out. So I had done work in in theoretical computer science and it was clear that these algorithms would be useful more broadly because the randomness entered in a very different way than classical algorithms. And so then the question is, how does it evolve? And so it evolved. I was at Yale in the mathematics department as in one of the junior faculty positions. I spent time at Yahoo and Stanford and moved to Berkeley. I'm in the statistics department there and at the International Computer Science Institute and Lawrence Berkeley National Lab.

1:04:36 And as you mentioned, a couple years ago started with as an Amazon Scholar working in the supply chain optimization technology group, where, you know, we work on better forecasting, demand forecasting algorithms and supply chain decisions. So I think this is, this is something that's good for people to think about. And you know, there's at least, I guess 4 hats there. And each hat has pros and cons and pluses and minuses. And so figuring out how to navigate that space is something I try and encourage students and postdocs to think about sooner rather than later. How? How much should people expect to have to move?

1:05:14 You know, I always think this is one of the sad in a way, things about academia, at least from what I've observed, that the fact that you do your PhD and let's say you do a postdoc, how realistic is it that that person in the US or the UK should assume that they can stay at that same university and just move up a chain? Or how much should they expect that they're basically going to have to move somewhere else in the country to find the position? Yeah. I mean, the short answer is that when you do a postdoc, it's usually more targeted because you're working with a particular person, a particular group and and so you could be at the same place or not.

1:05:59 It's usually a good idea to go elsewhere to get experience with other people, the real filter's at the assistant professor level. And you got to move there because, you know, except in rare cases, you don't get a job at the same place. And that's not a statement about whether the place where you are at is good or bad or whether you're good or bad is just just run the numbers, right. If there's 10, 20, 30 or 40 places hiring and that you're looking at and and the chances are one and whatever that that you get an offer, it's just unlikely to happen at that same place even if you ultimately end up there. Sometimes people go and leave

1:06:29 and come back. So moving, for better or worse, is part of the part of the equation, yeah. But did you, some of the questions I get is the should they go into industry or, or just let's put this two ways, does academia value industry? So if you've been an assistant professor or you've been a postdoc and you then say, you know what, I'm going to go into industry, do they value then? And is the ability to come back or have you been out of the game too long? You haven't sort of, you know, got the publications or the teaching experience, You know, is that a good piece of advice or would you say no, no, no, If you really want to be in academia, you basically need to

1:07:16 stay in in academia. Yeah, I mean it, it's, it's more textured than the following. But I, I think the short answer is that that that you need to stay. The slightly longer answer is it depends on the area, like computer science versus statistics versus engineering are rather different. In some cases having a startup, in some cases having a postdoc in industry is good and you can go back to universities and that tends to correlate with computer science. I don't know as much in engineering that may be changing over time, but less so. And maybe statistics and applied math. I think in a sense, on the one hand, people value the

1:08:02 experience you might have, but not really. And by that I mean, you know, if, if you, if you check all, if if you check all my boxes. And in addition, you have this other stuff, good, but you got to check all my boxes and the boxes are the things you alluded to. And so there's sometimes there's postdocs in industry that are basically academic post docs, you write papers. So I'm not counting that because that's, that's effectively the same sort of thing. But if you're out more than a couple years and you have fewer papers, it just gets harder to publish. There are certainly cases where people do that, especially if they have some high profile ones

1:08:35 and, and can work their way back one way or the other. But it's, you know, you should know going in that you're, it's like a salmon swimming upstream. And this is going to be a hard one. So how does he, this is maybe people in the US are maybe a little bit more familiar, but particularly people who are not. So how does it work with the national labs? So you have a position at the national labs. How, how does that work between is this a, a quite common thing in a way for these things to happen? I I can only imagine how you juggle the e-mail inboxes and the sort of meetings, but I guess it's valuable. And interestingly, after an hour

1:09:14 today and I haven't been hit by e-mail. That's why I've turned it off. Except, e-mail is a little overwhelming. I mean, it's not so common. So, so I, I with the Berkeley hat, because I'm in the Statistics Department there, I'm at Lawrence Berkeley lab. It's literally just up the hill. And so I can, you can walk there. And so they were interested in scientific machine learning and, and I had done work and was well known in machine learning and, and algorithms and statistics, large scale statistics. And I had various projects over the years with people there and elsewhere on scientific ML problems. And so started about two years

1:09:53 ago. This is what we renamed, but the scientific machine learning group, basically MLA: Machine learning and analytics. And so it was partly facilitated by the fact that there there is a moderate amount of interaction between UC Berkeley campus and LBNL Lawrence Berkeley National Lab, Oak Ridge, Argonne. Some of the other labs have things like that, but it's not so, so, so common. Most people are there, meaning just there and 100% of their FTEs there and and there for long term but so it's not. I wouldn't say it's very common. OK, which makes your experience even more unique then to have done it. Yeah, what maybe is a sort of

1:10:39 finishing off comment or, or topic is if you have somebody now who was wanting to get into scientific machine learning, because this is I guess kind of what we've been largely talking about, what would you get them to focus on? They're doing a PhD or they're doing a postdoc. But is there any, doesn't have to be a very, you know, specific, specific, but what sort of area we'd say, you know, this is something you should look at. This is an area that is ripe for, you know, progress. Yeah, that's a good question. There's certain things, I mean, that are less ripe for progress

1:11:30 in scientific ML, even if they're important scientific problems. So what I, what I, what I would try and say is work. You should work on a problem that's of interest to both sides. Otherwise you're coming at it from one side or another. And this is true whether you're coming from CS or stats learning scientific problems or coming from one scientific area. You should work on trying to figure out how to frame what you're doing as something of interest to both sides. So I have projects that, you know, we, we, we construct projects, you know, a student or postdoc owns a piece of it. And, and a goal typically is we want a publication in the

1:12:06 particular scientific area, but also in an ML venue. And it needn't be the same one. Sometimes it is, it's a long version, a short version. Sometimes it's just two different things, but work on a method that's broad enough that machine learning people are interested in it and that by the definition of machine learning, people are interested. You convince 3 reviewers to say yes. And so it's an imperfect process, you know the review process, etcetera, but boom, you get it in a top machine learning venue. And similarly, on the scientific side, you know, a statement that it's of value to the scientist is you get three reviewers to say accept and then it's

1:12:36 accepted on the scientific side. And oftentimes the way we try and scope our projects is to say, you know, we don't on the machine learning side, we don't want to work on your problem if it's only of interest to you. If I can't apply it to some other area, someone else, some other domain, Because if that's the case, I probably need to know so much about your area that I become, you know, an expert in your particular area. And similarly, on the scientific side, you know, if you have a particular method that is only that's so heavily tailored to your domain that it's not going to be useful more broadly to machine learning people, that's a much harder sell.

1:13:07 Then you're not doing something that satisfies both sides. So something that I knew about from working on years ago was like sequence alignment in in genetics, right? Clearly an important problem, not something that's portable the particular algorithm, not something that's portable to lots of other scientific areas, but better machine learning methods for solving spatiotemporal forecasting problems with non-trivial boundary constraints. That's clearly of interest to a pretty wide range of applications. And also to solve it, you're going to have to introduce technical solutions that are probably, you know, go, can you come up with a differentiable

1:13:40 optimizer for an end-to-end differentiable system where you have hard constraints? I mean, that's clearly something of interest to machine learning and optimization. So I guess I'd try and focus on something there. And that is probably the way that you're going to be. And unless you know exactly what you want to be doing 30 years from now, that's probably the way to make yourself most forward-compatible with with however. That's a really, that's a really good point. I guess what you're saying is, yeah, it's true. Like later on in your career you've become established in a certain area, but most people during their PhD or the postdoc

1:14:12 at that time have no real idea, necessarily. So you're saying if you do something broad enough that can get you exposed to different groups and and have a foundation to your earlier comment, you can more easily go into verticals because you've built that foundation. Whereas I guess if you jump directly into a vertical right from the beginning for you to sort of reverse out that it's kind of harder, isn't it? So I I think that's what you're alluding to. It sets you up better. Yeah, yeah, because it's not at all. I mean, all the questions you're asking are fair ones. And it's not at all clear how the whole area will evolve.

1:14:51 And so depending on how it evolves, having experience in one versus another. And, and I think I mean, having worked in a bunch of areas, you know, it's, it's, but, but speaking substantially only English, but a but a little bit of other things I can imagine. And I see, you know, you learn a second language, it's hard. You learn a third language, the, the relative cost is a lot less. You learn a fourth. I mean, so there's a diminishing returns in terms of the amount of extra effort you need to do. So you work in one area, switch to the second. It's very hard switch to the third. You know, you can, you know, you can, you can do that. And after that, one thing I've

1:15:27 sort of been always struck by is it's amazing how little you, you need to know about a particular area, especially if you're working with good people who'll complement what you're doing. You can learn from them, right? And so it's amazing what you don't need to know in order to, to get interesting results in an area. And so, and so learning new things. By the time you know two or three things, it's easy to learn the fourth. Yeah, yeah, I know that that makes sense. Well, thank you so much. I it's the day after Labour Day, which means that probably there's a whole bunch of emails and stuff that's started to come through on essentially your

1:16:01 first day back. And I know we probably could have carried on for for another couple of hours. But yeah, thank you. Really, really appreciate it and look forward to catching up in person at some conference in the future. Sounds great, thanks for having me, this has been fun. All right. Cheers, Neil.