The Neil Ashton Podcast
Prof. Michael Mahoney — Perspectives on AI for Science
Episode overview
In this episode of the Neil Ashton podcast, Professor Michael Mahoney discusses the intersection of machine learning, mathematics, and computer science. The conversation covers topics such as randomized linear algebra, foundational models for science, and the debate between physics-informed and data-driven approaches. Prof.
Mahoney shares insights on the relevance of his research, the potential of using randomness in algorithms, and the evolving landscape of machine learning in scientific disciplines. He also discusses the evolution and practical applications of randomized linear algebra in machine learning, emphasizing the importance of randomness and data availability. He explores the tension between traditional scientific methods and modern machine learning approaches, highlighting the need for collaboration across disciplines.
Prof Mahoney also addresses the challenges of data licensing and the commercial viability of machine learning solutions, offering insights for aspiring researchers in the field. Prof. Mahoney website:
Chapters
- 00:00 Introduction to the Podcast and Guest
- 05:51 Understanding Randomized Linear Algebra
- 19:09 Foundational Models for Science
- 32:29 Physics-Informed vs Data-Driven Approaches
- 38:36 The Practical Application of Randomized Linear Algebra
- 39:32 Creative Destruction in Linear Algebra and Machine Learning
- 40:32 The Role of Randomness in Scientific Machine Learning
- 41:56 Identifying Commonalities Across Scientific Domains
- 42:52 The Horizontal vs. Vertical Application of Machine Learning
- 44:19 The Challenge of Common Architectures in Science
- 46:31 Data Availability and Licensing Issues
- 50:04 The Future of Foundation Models in Science
- 54:21 The Commercial Viability of Machine Learning Solutions
- 58:05 Emerging Opportunities in Scientific Machine Learning
- 01:00:24 Navigating Academia and Industry in Machine Learning
- 01:11:15 Advice for Aspiring Scientific Machine Learning Researchers
References and links
Transcript
This transcript was generated by Spotify and may contain errors. Download the original SRT file.
Hi, and welcome to the Neil Ashton Podcast. In each episode, we explained some of the fascinating ways that science and engineering are changing the world around us. We talked to leading engineers from elite level sports like cycling and Formula One to some of the world's top academics to understand how fluid dynamics, machine learning, supercomputing are bringing in a new era of discovery. We also hear some of their life stories, their career advice, the lessons they've learned on the way that I hope will be helpful to you too. So sit back and enjoy this episode. Hi, and welcome back to the Neil Ashton Podcast. So on today's episode, we have
Professor Michael Mahoney, who is one of the world's leading experts on machine learning, mathematics, and computer science. He's also somebody who I've got to know over the past year and has been a great help in understanding this space. He's among other things, an Amazon scholar. So we've managed to work together a little bit. And hopefully, as you can tell from this episode, he's a a really nice guy and has a great sense of humor and such an intelligent person. Where do I start to talk about what he's done? Well, he's a professor at UC Berkeley in the Department of Statistics, but he's also at Lawrence Berkeley National Lab.
And as I mentioned, he's also an Amazon scholar and, and has a few more hats as well. He, if you look at his Google Scholar, which for a lot of academics is a way of, you know, getting a look at what they've done, he has some pretty impressive statistics. And one that sort of comes to mind, not only does he have more than 36,000 citations, which, which is a lot, his h-index is 84, which is very high. But what's more impressive is, and again, this is probably going into the details, but it makes a difference. If you go on Google Scholar and you look at the number of citations, his is exponentially growing. So not only does he have more than, you know, 30,000
citations, H&X of 84, which is already in the very high, sort of top 1% or probably even higher, it is increasing like this, which basically shows that his research is becoming ever more relevant every single year. It's not plateauing or going down. That's actually very impressive and that shows why he is so highly regarded because his work is really cutting edge. There are many, I think it's fair to say that because he has that mathematics background and that computer science background, he's put his hand to many different things. But one thing maybe relevant for the listeners or viewers of this podcast is a paper he did with
some colleagues on characterizing possible failure modes in physics informed neural networks. And this is, you hear many people talk about pins, physics, informed neural networks. And so they did a great paper a couple of years ago looking at some places where it may not do so well. We talk actually about that in the podcast. We go through some of the themes I've discussed with other leading ML experts. And we talk about, for example, his opinion on, you know, do you really need to include the physics in this AI for science regime or is it good enough just to use data? We discussed this at length. And we also get into the topic
of foundational models for science, something that he's actually really being a leading voice on. He's given numerous keynotes, important seminars, and published papers in this area. So we have a good debate about that. But we actually start off the conversation talking about something he's also very well known for, which is randomized linear numerical algebra, which is quite a complex topic, to be, to be completely honest. And I don't think we got to the bottom of it in this short talk, but it really shows that some of his work from the pure math side or pure math side is coming through and being relevant as we increasingly look towards lower precision methods and ways of
doing mathematical tricks to try and improve the speed of many linear solvers, which are common, of course in many, many fields of science, including CFD. We also talk about his general opinions or advice for people looking to move between academia and industry. Just general career advice, something he's well placed given that he has an incredible academic record, but he's also very close to industry and sort of industrial applications as well. An hour or an hour and a bit is not really enough to go through all of this, but I, I certainly learnt a lot and I really enjoyed talking to him like I do every time I speak to him. And I hope you'll, you'll find
the same thing. I will add some links if you're watching this on YouTube because again, he's published so many papers and done so many great talks that I really want you after watching this, listening this to go to his website and and look through all of those. And obviously if you are watching this on YouTube, you may prefer to listen to it. Some people don't realise that these podcasts are in video form and YouTube, but they're also on Spotify and Apple for audio and vice versa. If you normally listen to this and you didn't realise there's a full video version, you can go to YouTube. So yeah, I hope you enjoy this
conversation with Professor Michael Mahoney. One topic that you mentioned to me when we were chatting in the past that I was fascinated by, but if I'm being completely honest, I didn't fully understand. So it's good for my purpose and for everybody else was around the randomized linear algebra. What exactly? For people who don't know what, what is it and why is this becoming even mentioned on a Netflix show? So that yeah. So that's, this was a sign of success when it was finally mentioned on a net no, the Lincoln Lawyer a couple of years ago, someone sent it to me and it was mentioned in in one of the courtroom scenes.
So randomized linear algebra, randomized numerical linear algebra is basically an area that uses randomness as an algorithmic resource to solve linear algebra problems. A lot, a lot of linear algebra problems are under the hood, whether you're doing machine learning or scientific computing, they manifest themselves in different ways. So the exact questions and numerical issues and so on are, are different in those areas. And that's the source of maybe tension in the area, but also synergy and, and, and, and, you know, part of the reason a lot of people are interested in it because it, it holds the potential to solve a lot of
problems people look at, But it's, it's solving core linear algebra problems and, and linear algebra ones appear in a lot of places. If you're solving partial differential equations, it's, you know, linear operators or iterated former linear operators appear. If you're solving machine learning, you may be interested in support vector machines or ensembling methods or these days deep neural networks. And so matrix multiplications are at the core of a lot of that stuff. So very core linear algebra problems. I think historically the way people thought about the relationship between linear algebra and and randomness or
noise was that there's randomness and noise in the world, meaning the data you measure, think of least squares and it's your job to clean it up. And then you call a linear algebra problem least squares or low rank approximation. And you get more or less a deterministic answer and you get more or less the exact answer. I mean, and I say exact and scare and scare quotes because there's numerical issues and you can't represent sqrt 2 on a computer, but you know, in so far as machine precision is exact and you get an exact deterministic answer. And for a lot of things that's just overkill. And so you can use randomness as an algorithmic resource, meaning
inside the algorithm to speed up computation. And this may be most notable historically, like in Monte Carlo or Markov Chain Monte Carlo, where you run simulations of of fluid dynamics. I mean, the Metropolis algorithm was developed in that context, but this is for core numerical linear algebra problems. And so most people, if they're running computations, sit on top of Glass and LA Pack and, and, and related software. And if you're calling Python, you're calling something else, you're calling something that's calling something that's calling them typically. And so these are core libraries. And so the question is, can you,
you know, use good theory from randomness and measure concentration and high dimensional probability to improve those algorithms or improve, you know, variance of those algorithms? The short answer is that you can. So how? But so how does it work in practice then? So if I'm sold in APDE and I'm normally using some sort of linear algebra library and it takes me this time or this amount of flops or this compute, how is it that the randomness reduces that? Yeah, I mean, when you're solving APDE, you're typically solving it for a particular application and you're calling certain core linear algebraic primitives in in one way or another.
So for example, if you're solving it with a splitting method or predictor corrector method, or you're solving a finite element or finite volume or you're solving a second order optimization to as a piece of that PDE solver. These have least squares, you know, linear solves things like this under the hood. So typically you're improving those. You could, you could improve the modeling setup for the PDE also. But but you know, you could, you could try and solve the, you know, the improve the core primitive. So take least squares as an example, a low rank approximation. A common motif in in a lot of these algorithms is you have the
data and you want to solve the least squares of the low rank problem. And you could solve it exactly machine precision or whatever. But you might want to on the other hand, get what they call a sketch of the of the data, which is roughly a small number of data points. It could be a small number of actual data points, or it could be what they call a random projection. And if you're familiar with like if you feel signal processing and electrical engineering and, and physics, think of this as like a randomized version of a, of a Fourier transform. So it takes some signal that that might be localized in space and, and spreads it out everywhere.
But you don't need this to be a physical space. This is just a linear algebraic problem. And so you have, you know, you have columns and rows, so you could select the important columns or rows, or you could do this random projection, which is essentially a, a basically a random type rotation that spreads the information out and then sample uniformly in that rotated space. And so you take this sketch and you can do one of a couple things that so some communities, the more theoretically inclined communities want to take that sketch and solve the sub problem exactly. And there's a range of theoretical work that says the solution to the exact, the exact
solution to that sub problem computed any which way. A traditional solver or something else is epsilon close to the exact solution to the original problem if you set things up right. Now, in a lot of cases that is sort of course because you can't. It's hard to get machine precision just by drawing a sample and solving the sub problem unless you take a huge number of samples, just because Monte Carlo methods tend to converge slowly as a function of the error parameter. So you could take that, you know, low quality, low precision, but not trivially bad solution and ask yourself what is a preconditioner? And by preconditioner I just
mean a preconditioner for APDE solver or or least squares or a linear solver. And a preconditioner is basically a low quality solution that you refine. So you can actually take this preconditioner that the that this is a sketch. The theoretically inclined people just say good, done M epsilon good. And you can say, now I'm going to take that epsilon good where epsilon is .1 and just use any of a range of traditional iterative solvers and drive that epsilon down to 10 to the -8 or 10 to the -16. And, and just like most preconditioners, if, if the cost to construct it is less than the cost to, you know, the cost to construct it plus iterate is
less than you know, the, the algorithm you're competing with, you win. And it's a good preconditioner. So there's a range of ways people solve it that way. So those are called sketch and solve and sketch and precondition, respectively. Increasingly for machine learning applications that want medium precision, also relevant for scientific computing when you're interested in low precision data representations, you know, going to half or quarter precision, which is, is increasingly seen in hardware. There's a more subtle interplay where you can get the solution and then you can iterate it and get a second sketch and toggle
back and forth. And if you set the parameters differently, you can get intermediate solutions 10 to the -210 to the -410 to the -6 quality. So there's a range of ways you can use the sketches, but roughly the idea is you get most of the information in the sketch and and solve it or do something with it. And is there any, you know, orders of magnitude of the savings that you could get? You know, if I'm doing a Nabia Stokes solver or solving some neural net using these randomized, you know, numerical linear algebra versus a blast rolling back, Are we talking, you know, a 10% saving? Is it an order of magnitude or is it still under research?
And so it's not delivering the full yet. Yeah. I mean, I think the question of of how well you'd improve upon something depends strongly on what your baseline is. And if you feed this into a big PD solver, there's a lot of moving parts and you're competing with very mature code and they're doing a little bit better or even a lot better on a solver may or may not matter for the downstream use case that that the scientist that is looking at the PD solver is interested in roll gold to the other extreme. And you want to say, I want to compete with blast or a late pack. These are extremely optimized pieces of code and it's, you
know, it's hard to beat them. So one of the big successes in the area was with blend and pick. And then soon after that was LSR and blend and pick, which said we wanted to ask whether these randomized sketches, you know, not in theory, not Big O notation, not whatever can they beat LA pack Because this is not boutique code you have or I have or a solver that you don't share with the community. So well, this is something that's been stress test for decades. And the answer is yes. I mean, basically on any tall dense matrix, you can beat LA pack with these techniques. You got to set parameters right and be a little bit careful. But, but the short answer is
yes, how much you improve it by depends on the aspect ratio and depends on properties of of the matrix. But think of it as ballpark 2 to 10. Now this was back in 2010 and so a lot's changed the since then, in particular in the hardware landscape with respect to GP us and hardware heterogeneity and the Sun sending of Moore's law and Dennard scaling and these sorts of things. And so it is the case that you know, you can be a factor of 10 better. Similarly in in in low rank type approximations. I think the the the way scientists versus machine learners use low rank approximations is very different. Scientists tend to say low rank means you know 99% of the
Frobenius norm, meaning 99.9% of the Frobenius norm and 99.99 I mean basically the whole matrix. When machine learners say low rank, they mean sorry, sorry, the spectral norm. When machine learners say low rank, they mean, you know, 50% of the Frobenius norm, in which case, you know, you lose a lot of information, but you might iterate and and do something better. So think of sort of scaling laws and the neural scaling context with, with neural networks. And so the way you'd use those low rank algorithms and feed this them into other solvers would be very different. And then that would lead to very different levels of improvement.
And so I think that's largely open. And then with the memory wall that you're increasingly seeing in, in general, but most egregiously in the machine learning applications, one of the big wins and it's just starting to be explored for the randomized techniques is basically improving the memory properties. And so this could be just reordering algorithms in the sense of communication, avoiding linear algebra could be much broader because the randomness sort of decouples things, makes it easier to paralyze certain sorts of computations. And so you see much bigger than factors of 2 or 10 improvement there. But then again the baselines
changing year to year as as as people explore different representations and and low precision and intermediate precision and stuff. Interesting, so how was it mentioned in Netflix then going back to the former? Well, the particular I, I think so, I'm not a movie star or a movie producer, but I, I have a sense that they, they want to, I, I mentioned this to someone and they said, you know what that means? They tried to take the most exotic thing that wouldn't make sense to anyone and, and, and, and cited it. So the context was I, I think it was early on and the, the, the main character, who was this, the attorney needed to represent someone and the person had to
get out of, of, of whatever they were charged with, which was not a, a, you know, relatively minor thing, but it was the 10th time they did it. And so they, the, the, the Lincoln lawyer said, why do you need to get out? And they said, I have a defense on Thursday. I need to get out. And they asked what's the, what's the topic? And they said randomization and numerical in your algebra. So that was the context of how it was mentioned there. I wonder how they found that out. I wouldn't know how they actually. Yeah, I don't know. They they, they, they didn't call me up. So they must have had someone in movie producer land or something
that did did a Google search on on something exotic and came across. It that's fine. Yeah. That's like reminds me of the Did you ever like Star Trek? Yeah, I used to see if there's a bifurcation which maybe would get us in into a separate discussion whether you like the original version or the the second version. But then there was this explosion of different, different versions of it. Yeah, I was more the next generation person, but I the reason I mention it is it was often whenever they wanted to explain how they were breaking the laws of physics, some stupidly complicated phrase was used. And I'm sure it exists somewhere and it just reminds me of that,
you know, explanation of how they're exceeding the warp drive. Oh, it's because the something something. Yeah. So I think it was a little bit like that. I mean, I, I think it's a sign of the times that what they're citing is not warp Dr. and quantum gravity, but randomization numerical. This is a core statement about how areas are progressing and so on. Exactly. Exactly. Yeah, yeah. So maybe one of the topics would be good to chat with you about, and it's definitely a hot topic at the moment, is the the foundational models for science. I've seen quite different viewpoints. There's one argument that is, well, large language models take
all this data from the Internet. They train and they're remarkably good at doing what they do. What about if we could do it for scientific disciplines? Wouldn't we be able to just, you know, ask her a prompt to go and simulate me a plane or, or, or do something? And then there's the the other Ave., which seems to be more the replacing simulation tools or augmenting simulation tools like graph cast or forecast net or, you know, these sort of Seminole pieces of work that have come out, but they were for very, they had to be trained, you know, for something. Where do you see this? Where do you see this going? Are foundational models possible?
Is it we just don't have enough data? This is a big topic. So yeah, I'm not sure how we want to start it. I mean, we're in the middle of this fray and I guess that's why we were asking about this. And I should start off with saying I don't know what a foundation model is. Actually. I know what different people say a foundation model is and, and it means very different things to to different communities. And so I think articulating that a little bit helps articulate some of the possible directions and some of the questions you're asking. I mean, I think one way to think about a lot of these machine learning methods and it's
highlighted particularly cleanly with these foundation models is machine learning methods. And in particular, these foundation models are, they're an infrastructure, right? The, the, the, the Stanford report said, call it foundation models, not foundational models, because it's a foundation in which you build things and as and as opposed to a half dozen other terms they could have used. Now, whether that's right or not, I mean, but that, that's what they said. And so since then people have said, well, it has to be this big and that big and, and whatever. And, and then people want to get a foundation model, not for all of science, but for area A&B&C
and sub area D&E&F and get finer and finer granularity. So in a sense it's, it's a it's a term. And, and if you ask people, most people using it what it means, they sort of will acknowledge that, that they don't know. And it's not a standard definition. I think probably the best way to think about it is it's, it's, it's infrastructure, it's a foundation what you build things. And so the computer's infrastructure, I mean, you can do different things with the computer than you can do with pencil and paper. And so it wasn't obvious the way computer science, even before it existed, would evolve. And at one point, post World War
2, the US thought they'd need 5 computers and that would be it. And then it would solve all the nation's needs, right? So it evolved in ways different than expected. And it involved in particular because it, it could complement what people did. I mean, you could ask different scientific questions than you could with pencil and paper and you could ask different engineering questions than you could. I mean, now you can simulate whole things and, and the very, very last step is building it, right. But in addition, you could do lots of other things that were driven by industry. And, and so there's a bifurcation in the area whether
you're doing more numerical things that are continuous or whether you're doing discrete things that, that, that are not continuous that were driven largely cited by business needs. So I think the, the right way to think about this is it's that sort of infrastructure and it's tied to the data more closely than than, you know, database systems because the foundation models have information cooked into them. And so it wasn't obvious some number of years ago that just with language you, you could do what you know, has caught the popular attention in the last couple of years. So I think so to anyone who says that confidently, you, you can
or you can't with science, I mean, probably you can't say that they're sort of reliably saying what what will hold for the future 'cause I think it's not at all. I think it's not all clear for two reasons. One, there's certain technical things that really do matter. I mean that that are different about scientific data than, you know, language and, and vision data. And then there's just very different cultural, there's very different cultural things. I think every scientist thinks that their data is special and unique just like they think that they're special and unique. M LS taught you anything. These algorithms can predict what movie you'll watch and what
stuff you'll buy better than you can. And so you're maybe slightly less unique, you know, than than you thought in some sense. And so I think there's a question, you know, what does foundation model mean and how can it be used so narrowly? You know, I have a big model I can train it in in domain A domain A could be weather and climate have gotten attention, but it could be fluid dynamics. It could be something else about stylist formation. It could be properties and materials and doing density functional theory and so on. And then there's a question about whether you could maybe learn broad based models that cut across domains, which would of course be the more
interesting thing. So there's something Kronos said, I was involved with the AWS people that sort of says, I, I don't want to be the best at at the most extreme things. I don't want to be the 1% that's predicting the most extreme things I want to be. I want to do as well as as the majority of users for for time series analysis and oversimplify the story. I mean, time series is a complicated area clearly of interest in a lot of cases, but it's it's a little bit, you know, you got to be careful about off by 1 errors and a range of other technical things. And what that says basically is take a language model, language
model structure with a bunch of time series and just mean center it and variance normalize it. So just do the simplest possible things you could and you got to do some data augmentation stuff and boom, you know, you do sort of comparably and, and better than a a wide range of public benchmarks. Now, the public benchmarks are not extraordinarily high because a lot of time series data is valuable and so companies tend not to release it. But but the fact that you can do that just by sort of variance normalizing and and mean centering is is, you know, not obvious and sort of interesting. It turns out sort of on a separate thread that you could
say, could I do well compared to the 1%, you know, the most extreme things and, and the the answers that you can there too. You got to use different techniques than you do there. And so that leads to questions as to whether you could have a foundation model for sort of a broad swaths of, of time series and forecasting analysis. And so, as you know, I said, I'm an Amazon scholar working with the Scott team and we're looking at that there. And one of the interesting things I think about there, if you look at the details of the model, not the chronos, but some other things, why would you expect that article scraped from Wikipedia, you know, public,
publicly available language models? What would, why would that allow you to do better job predicting dog food demand? You know, I mean, it's not obvious they have anything to do with anything. You know, it turns out, I suspect the hypothesis is that the text, the linguistic structure that you're learning from the Wikipedia articles is strongly related to sequence to sequence modeling, which, you know, if you think of 1 dimension as a metric space, it's a very, very special metric space, a very specially structured metric space. And so you can learn not just recent information and not just low frequency information over the past, but maybe in
information at different scales. Because, you know, sometimes articles in text refer back 10 words or 100 words or 1000 words or 10,000 words. So you can learn these sort of heavy tailed sort of structures and that gives you a better set of basis functions to learn dog food demand or whatever else. So in a sense, these, the language models give you better embeddings, you know, 4A analysis and Laplace transforms and, and these sort of things are great for pencil and paper. They're all developed in the 1800s, right? These are data-driven embeddings that are good if you have computers and and data not necessary for pencil and paper,
but they're good for that. And so that leads you to the question about if, if I'm really careful about doing foundation model work in science, what do I need to worry about? It's not the details of this PDE or that PDE. It might be that I want to reproduce what happens in the natural language processing the NLP models, which is roughly scale model and data and compute. So none of them saturate. That's very different than the usual strong scaling and weak scaling and high performance computing. I want to change the amount of data I'm putting in. I want to change the size of the model. I want to change the size as it
computes. So none of them saturate and, and the conjecture would be that if if any of them saturate, it's going to be harder to do transfer learning. So you really need the non saturation and you do that one way with NLP, natural language processing and CV computer vision. But an NLP and CV, there's much weaker control you have and that you need on the spatiotemporal geometry than PD ES, right? Is, is, you know, PD ES, if you're below a nice a Mach transition or some other physical transition, things are sort of smooth. If you're above it, things are very messy. Dealing with that transition is hard and people spend their
whole careers dealing with that. So can you come up with data generation or tokenization mechanisms that that respect the spatio temporal properties? So I think if you're going to have a foundation model that applies across a raw broad range of sciences that's trained on weather and, and, and climate and, and simulations of different grid structure from the machine learning perspective, it's not so different than satellite data from satellites looking down, you know, it's a different grid and it's an image. And then you say, how does that couple with the spatiotemporal properties? So I think if, if you want to deliver on the promise in the same way is, is the is computer
science had to work out a range of numerical methods to really match and beat state-of-the-art, you're going to have to do a similar thing here. And so partly depends on these technical issues, but partly depends on cultural issues. But how much do you think? I mean you raised the what percent versus the 50%? I guess that is one argument as well, which is how close does it need to be to machine accuracy to be useful. You could argue that any of these large language models, the way that most people use them, there is still a correction you need to add. You don't typically ask it to do something right, document and
literally take it word for word. You usually go in and go. That's, that's remarkably close, but I'm still going to go and fix it. Or it writes you a Python code. It's unusual, isn't it, that it's perfect. There's usually something you have to correct. So I guess with the science side, maybe the argument is how close does it need to be to be useful And the cost of getting the incremental increase in accuracy, you know, is it, is there a like a trade off where you need so much more? Data, I mean, again, I think the best analogy is look at the history of computer science and how that of all that, I think, I think framing the question to say how close does it need to be
to be useful? You're already framing it in a way that that makes certain one group comfortable, the numerical analysts and the PD people that frame a question a certain way. That's not how people would have framed the question before. I mean that, you know, before you had, you know, represented continuous numbers on, on a computer discreetly take a step back and say, I want to solve a certain problem. And, and the question is now that I have very different trade off in terms of compute versus data and I have this new infrastructure that's language models. Can I ask a different question and, and, and push the science
for it? I mean, so in, for example, in, in chemistry historically, but also nucleus physics and, and in fluid mechanics, there's this notion of a semi empirical theory, which is a theory sort of derived, you know, it's not just curve fitting, it's derived from an underlying more fundamental theory, maybe phenomenologically with parameters that are then fit empirically or semi empirically. And So what you need there is not the theory to be right, but sort of right enough at the level of in the, in the chemistry is chemical accuracy, which is, is however many kilovolts or whatever depends on the reaction you're interested
in so on. And so certain methods that maybe were more principled were a bit too course for that. And other methods, when you combined it with other techniques achieved chemical accuracy. And so none of them were low enough in the stack that you they were better QR codes for QR from, from linear algebra. But they solved the downstream problem at that level of actually, probably what you'd see here is that right? So it's, it's probably not going to be so successful to say I want to get 10 to the -16 or 10 to the -2 of my large language model, but but a foundation model for science and need not be an LLM. It, it, it, you could be
literally learning the embeddings from a range of PD ES. You could, you could train to PDS of different types. And you could say, I want to port if I, if I know the basics of hyperbolic and parabolic. And I know that in the transport equation, this can enter in a certain way. And you have nonlinearities and certain types of forcing functions, you know, model each of those and and solve each of the components separately. So I think the question is, what's the right level of abstraction to do that? And, and one of the modalities you could use is language, but of course you could use image or PD ES or simulations, any of a range of things.
So I think in the context of, of scientific problems that the foundation model need not be an LLM based. I mean, there's some people are pushing that, I think, but but it certainly need not be LLM based. We just query the model and say, you know, tell me what to look for, the top quark or whatever. Well that that's the bit maybe you brought up would be interesting to explore. There is, AI would say, a relatively fierce or contested argument around physics informed or physics in it versus data-driven. Where do you sit on that, on the argument? Do you, do you? Does the model need to implicitly have some of the boundary conditions, some awareness of continuity of
energy? Or is it good enough just to give it so much data that it essentially learns those laws by itself? Yeah, I mean, I, I think you could ask the same question about in natural language processing, computer vision does did you just, I mean, there's a, there's a common story people tell which is just get more data and everything works. But that ignores the fact as you know, that a lot of companies and universities, a lot of people have put a lot of resources into NLP and CV to try and figure out how to make it work. And you know, convolutions convolve things, they spread stuff out. You know, if you have an image
in 2D that may make sense, you might want to average over nearby things. But if there's, if there's a sharp corner like, you know, the shirt against the background wall in this picture in my of me, you may smudge stuff out. And so there's a range of ways to try and sort of deal with that. That's core structure. The 2D structure that you're trying to average over is very different than sequence to sequence learning. That's also much more discreet in, in natural language processing. So it's not like computer vision or natural language processing, just learn stuff. You gave it very particular architectures that had very
particular adaptive biases and then you're careful about the data and, and trying to get gobs of it and, and then then it worked. So I think if you, you know, you'd have to follow the same path here. You can't just take an existing architecture and press a button and hope it works. I mean, time series, forecasting space, you temporal modelling, these sorts of things will have to be, you have to figure that out, how to do that. So I think the question, I mean, some people, you know, I think are, and, and maybe sometimes people that have a vested interest in, in moving forward or not. I mean, some people say, you know, just ignore everything,
let the data do it all. Some people say, Oh, you'll never match what I can do with my careful PD solver. I mean, I think that's neither of those questions is the right one. I mean, the question is, given the fact that there's an infrastructure of code and experience numerically, but but you're also at a very different place. You've been sitting on top of certain essentially hardware and, and, and linear algebraic advances. I mean, for years, for decades. People working on PDS call certain linear algebra libraries and have essentially had not to pay the technical debt associated with the fairly high level of of complexity with
developing new linear algebra because just wait a year or two and the machines are faster. And so if, if, if that's not the case, you're going to have to start thinking a little bit more carefully about the underlying linear algebraic computations. Maybe there's different trade-offs points in the space and maybe bring something else to bear. You know, a model that has learned certain course type functions of one dimension, say from the language model that I alluded to before that aren't Fourier modes or aren't Laplace modes or aren't something else like that that they're more familiar with, learn a different grid type discretization.
So I think the real challenge in delivering on the promise of all this stuff is figuring out what the right way to combine those two. So you can imagine rather than just writing down a physics informed loss and pressing a button saying, well, the way people actually would solve this is to have some sort of splitting method and deal with the two types of physical things and the diffusion and invection in two slightly different ways. And machine one and one and take the other as numerical. And then you couple them together, like maybe in a differentiable end to end. Yeah. The, the reason I ask this is because it does feel certainly at least in the CFD community,
that there's a, there's a real sort of split within the community. And in a sense that on one hand you have people saying, well, we have all these PDE solvers, you know, developed for decades. And you know what convinced me how your machine learning or AI method is going to be as good and faster and cheaper versus on the other hand, you have quite sensational claims of, you know, 10,000 times faster, you know, sort of real time. Well, of course, when you dig in, you know, you start asking the questions, well, how much data did you have? What was the time to create the data? And but there is a bigger
question, I think more when you look to the future, which is, is it just a matter of time? Is it a bit like in the 1980s, they could only simulate a plane to a reasonably low accuracy because the computers just weren't power enough. And as you said, you just almost keep the same effort almost. And you just literally add 20 years of compute on top and now you can simulate. Or is it a we'll never get there because we'll it's like an it's like impossible to have that much data? This is the bit I think a lot of BCS and start-ups are also asking that question if we pump 100 million into this or a billion, will we fix it?
Or is it a trillion dollar problem and it's not worth doing? Yeah, Yeah. I mean, OK. So I think there's a there's a bunch of things in the question there. Sorry. We'll we'll CFD people to BC and and you know that they have two very different utility. You. Know, I think, OK, so it when someone says to me do this, convince me of that I have these metrics that have been I've worked on for decades, I just say thank you, good to talk to you. I mean, because, because there's no way you're going to win that battle because they're very fine-tuned metrics. So they're very one little use case. Even if you do win that battle, they're not going to admit it
because they'll, they'll tweak their method and, and you know, be slightly better than you. And there's greener pastures few everywhere. And this is not a hypothetical statement. This is a very practical thing. You're asking about this in the context of PDS. We saw this before in randomized linear algebra. I mean, so literally, I remember it was very clear that the techniques could potentially be useful, not just in theory, but in practice, and that the reason the numerical people were saying, oh, they won't work, You know, really. I mean, those are those number of methods that clearly wouldn't work. But but then there's the next
generation of methods. And that's sort of when I entered the the area and it was clear that they could potentially work and the objections people had before wouldn't work. And so this was right around the time of the Blunden pic paper that I mentioned that that basically said, can you beat LA pack? And the short answer was yes. So now, for example, if you look at the Simon linear Alger meeting, half the talks are on this topic, but there's been this process of creative destruction where this they're still asking the same questions, but putting randomness there. And so that's one group of community. I mean, a very different
community is the is the people that say, I'm going to not do that. I'm going to go and apply it to machine learning and 10 other problems. And, and, and so there's that tension I was alluding to before. So I think if you, if you try and, and satisfy an old school metric, it's, it's good to have a few examples of that as a proof of principle. And, and this is 15 years later, and I mentioned it twice here, right? So this was a clear proof of principle of the area. The area sort of accelerated after that. And there's been a lot of theory and empirical development. So I think there's been a few things that, if not that are starting to look almost like
that in this general scientific ML area. I I think. History is moving in that direction. So it's clear then that, you know, randomness would be important for linear algebra. Now it's even more so for all the historical and hardware trends you were talking about. I think likely the same thing's going to happen here, right? And so a question you could ask is, are the people who know the linear algebra, are the people who know the computational fluid dynamics? Are they at the table developing the methods? Are they just going to Pooh, Pooh it and say, well, they're not going to work because it's not going to satisfy my
particular measure as opposed to here's a broad class of techniques and it's hard to imagine that there's nothing in my area that will be improved by them. And so I think that's the that that latter question is the question to ask. And that latter is the question the questions VCs and and other people will ask. And I think the, an important maybe determinant of how this will evolve is, is how the different players interact with this in, in the sense that a lot of the machine learning, more broadly, scientific machine learning, not just for the foundation model, but more broadly machine learning really is, really is a horizontal. I mean, it's, it's designed to
say, what can I do with lots of data and be relatively ignorant about your application area? Because who knows why people click on ads or click on links on their social media account, right? I mean, you can tell the story, but really who knows why? And so a lot of machine learners implicitly or explicitly will say, I'm not interested in solving your scientific problem if it's just a one off problem. I mean, if you're a scientist using machine learning, you might be interested in that, but that's because you're interested in the domain. But if you're a machine learner, you might say, I mean, I see some commonalities between your
fluid dynamics and this other solid mechanics problem. And I see some similarities there between that and something in in 3DS in the atmosphere as opposed to, you know, something very different. And so identifying those commonalities I think is important. What I see in, in in some cases, and this is Hanford by universities and it's Hanford by industry, and it's Hanford in a couple government labs in, in a couple different ways, is a lot of people want to solve a scientific problem. And so they're going to say, I'm only going to invest, invest literally or figuratively in machine learning in that one area. And that's just not how machine
learning works. You don't see the, I mean, think of it as horizontal business with horizontals and verticals. If machine learning is a horizontal, like high performance computing that will solve a wide range of problems, you got to apply it to a vertical to justify it. That's that's where that that's where you know, the rubber hits the road typically. And so you don't see the successes in machine learning, scientific machine learning at vertical companies, you know, companies working in genetics or or oil exploration or any of a range of other particular domain verticals. You see it in the horizontal companies. The companies have built out the
technology infrastructure and they get a team of people that know this is a vertical and applies it there. And then, you know, a team that knows a different vertical and applies it there. And so I think the question is, how does that play out? And that's going to be a non trivial evolution in terms of. And your. That day. Your comment was interesting as well about the difference between sort of Transformers versus CNNS and you know, if you just throw the whole world's Internet at CNN, that isn't you needed to have the right architecture to be able to throw a lot of data at it. I guess the question is more for the science.
Is there a common architecture across the sciences that then would allow you to sort of build out that big which could be applied to verticals? Or is it the case that the needs for weather or genetics or chemistry are so different that you would need a different architecture for for each different size? And therefore the, the, the dream of a true foundational thing. Besides, it will never happen because you know. Yeah. So that's the big question. And I, I think there's a technical component to that and then non-technical component in terms of the areas, right? And so technically it's open-ended. I don't know, right? I mean, we're, we're in the
middle of the fray and we're, we're voting with our feet and placing a bet there that the, that the answer is or could be yes. To try and address that, you're going to have to deal with the issues we were talking about before. You know, one dimension is very, very special. 2 dimension, the surface area to volume properties are very different. 3 dimensions is very different. By the time you get 4 in scientific computing, that's sort of infinite, right? Machine learners, by the time until you get to a million, you know, it still seems pretty small. So mapping the methods to one versus 2 versus 3 dimensions, open question, maybe there'll be
no trade off point where the surface area to volume and the isoperimetry and other things like that wins. I suspect not. I suspect it could, but I mean, open question. And then non technically, if you look at how computer science evolved, numerical analysis and scientific computing was core to computer science. Look at how numerical analysis and computer science evolved. Yeah, these areas gave birth to computer science and then computer science banished them because it was easier to build the structure, the theory around, around Turing machines and, and, and the Lambda calculus and, and discreet notions that talked about complexity independent of the
data. And if you're talking about complexity independent of the data, you don't need preconditioners because it's independent of the data. And so there's a cultural aspect that that says, you know, does the vertical or the horizontal culture win here? And so even if you stipulate that the answer is technically yes, I think there's an open question on how it'll look, it'll evolve and whether that'll win the day and solve the problems that you're asking. But I think it's not, it's not so much that I solved PDE type A and applied it to PDE type B. It would be certain course things. I mean, as you know, there's a class of things for which finite
element methods work. There's a different class of things for which finite volume methods work. And you know, you, you take a course, you learn A1 O1 method and then boom, they bifurcate after that. So what's the analogous taxonomy here? And it's probably neither of those. It's probably something that's more data-driven, but what's the right way to taxonomize that? And even if there's a technically a correct way, I mean, you know, computer science is not computational science. Computer science is an area of the coherent. Computational science is spread over many departments and many divisions. And so I think it's an open question about how this area
evolves. Yeah, it is certainly. It is certainly interesting. I think one of the sort of key topics to all of this though is the data availability and licensing and financial reward. So one thing that I've seen in I guess the early days is nobody knew what was going on. So I guess the companies were scraping half the Internet. There were a few captures or there was a few sort of, you know, firewalls to stop thing. But in general, it was a sort of free for all it seemed, Whereas now people have, you know, cottoned on to that and some of the companies having to do deals with publishing companies or do deals with, you know, to try and
legally get access or or permissively get access to things. I saw the example on the weather side that, you know, some of those databases were released openly because they were released openly enabled, you know, these various teams to go and do things like, you know, forecast net and graph cast. But those data may not be open in the future. So how much you know, and you work in academia, you see people want to have their own data. Some people want to release later. Do you see that that will ultimately be the blocker? I mean, how do you incentivize people to want to make their data available and should it even be done that way, you know,
that the whole open source data versus closed? You know, this is this, I think, is the elephant in the room in terms of this area because a lot of government grants require that you make the data publicly available in some sense of the word. Now, what does in some sense of the word mean for a lot of scientists? If the data is more than a year or two old, it's worthless because they're trying to push the cutting edge of the science and then they don't maintain the data for more than a couple years ago. And so no one maintains and it's hard to access. So there's a lot of public data that's not de facto public or accessible, you know, in a lot
in those cases and maybe in, in, but, but certainly in, in machine learning sort of applications more generally in, in Internet and social media and so on. Having older data is good, but you can't get more than a couple decades older, right? And so often times the the most recent data is the most valuable and you just design A method around that. And so it's not obvious to me that the way the data are generated and, and collected, distributed now will be the one that is going developed going forward. I mean, if you get better telescopes, if you get better, you know, molecular sort of robotic systems to, to to make molecules, I can easily imagine
a situation where data that exists now or prior to a few years ago is basically obsolete and irrelevant. And so all of that data sits at whoever generated it. And so then the question is, is does a government lab generate it? And if so, do they make it easy de facto to access? Is it done at universities where, you know, you have a big infrastructure capital investment in in particular, essentially in particular science and verticals to generate it? Is it done at companies that are in different verticals or is there some sort of collaboration or interaction with, you know, horizontal infrastructure companies, hyper scalers and
then they get their fingers in that and have access to it? So I think these are all open questions in in terms of how this evolves. Yeah, it it seems to be very much, you know, traditionally and I guess if this is for me, the difference between the pre trained quote quote foundation models versus just letting someone go and train them. Although I mean, I guess most people you and I would not go and train our own LLM. We're just going to go and use someone who's audit and we're just doing an inference on it. That's and I could see if I'm BMW and I have data from previous car designs or simulations or whatever, how
they could train a model using their data. But I'm sort of wondering in my head, how could it ever become the point where some person has access to not only BM WS data, but Fords and Audi's and Volvos and Boeings and Airbuses and you know, Exxon Mobil and Shell and everyone else. Like how how could they collect the data? What would be the incentive for that to happen to or that's what I mean. Or is it the case that that's why there will never be a foundation and it will just be that individual people train their own models because the data is just so sensitive and difficult to collect in one? Yeah. I mean, this is, I think this is
the big elephant in the room. I mean, it's not clear to me that one. OK. So I think if you were to try and generate a meaningful scientific foundation model, you'd need data from a pretty broad range of use cases. And so you gave, you know, several that in particular are with is a strong sort of financial backing. I'm always reminded of the story. Gray was, I guess, the most prominent person in the Microsoft's research division years ago, such that, you know, he would have to have regular meetings with Bill Gates, the CEO at the time. And Bill Gates would say, why is the most prominent person in my research division applying databases techniques to
astronomy? This is when he started working with Alex Silly and some others on, on databases for, for astronomy and, and cosmology. Why don't you work on something more useful like, you know, any of our products? And, and crazy answer was basically, I'm working on it because it's worthless, because there's no financial sort of incentive behind it. I don't have to talk to the lawyers and we can work at IP and all these things. It's public data and, and I can work on database techniques and figure out latency and bandwidth and all these trade-offs. So I don't know, but it would be interesting to see. Will you have an analogous thing here?
So there's there, there is public, there's, there's various places where public data is available. And I think if you can illustrate that you would have a proof of principle foundation like model. Then the question is, will people come to it? Now clearly the incentives for scientists will be different than than for companies there. But I could imagine a situation where you build that out and then companies could take it and fine tune to their particular data, right. It's not obvious to me that any of the companies that you mentioned have data that that's, that is that special in a scientific sense? I mean, clearly the particular
data they have is particular to their particular situation and then the particular, you know, product they have. But I, I can imagine that, you know, the Pdes from early solar system formation and, and hydrogen bomb explosions and hydrodynamics inside stars, you know, is, has some similarities and some differences with weather and climate and some similarities and some differences with crack propagation in the Earth. And that'll have some similarities and differences with, with various sorts of sensing modalities that you would, that have a geospatial component. And so can you stress test that and, and get even even a proof of principle that, that if you
would attack those on, then the answer would be yes, that, that you could have a path to have a more foundation model. If the answer is yes, then the question is you talk to various stakeholders. Some of them are government entities or government sponsors, some of them are corporate and try and figure that out. And I could, I could imagine that that basically it's too big a left in the area stalls. So I think this is something to be determined. But I think having a nucleus of something is going to be important because right now I at least what I've seen, I mean, in most of companies like that, there's, there's, there may be
some people a little bit intrigued, but there's not a sufficiently heavy lift that they go screaming at their CEO to say we need to do this, probably because it's not extremely strong baselines that it can happen. And their competitors by talking to these public models and fine tuning on their proprietary data will do better. I think at that point they'll be you can see a change. And I think this is, I mean, I guess this is affecting the current Gen. AI and AI, which is how do you make money out of this? You know, like you can do it for the sake of science. That's nice. I developed this thing. But I'm always interested to see
like the money that you put in to create the data to train the model, can you ultimately make a profitable thing out of it that justifies its cost? Now, with standard, I guess, simulation software, there's an argument, well, if I simulate it, I don't have to build it. If I can predict the weather, I can help people figure out if it's there's going to be a hurricane. And I can say, you know, there's a, there's a clear thing. But because we can already simulate things, we can already predict the weather. We can always do it. For me, the machine learning bit is just the argument, well, it's faster or, or, or it's cheaper.
That's where the amount of data do. Do you know what I mean? Is there a commercial? But I, I think, I think there's a faster question and I think it's tempting to say, can I be faster? Because faster means there's a number you're comparing to. So am I bigger? Am I more on some axis? And back to the, the, the randomized linear algebra example, I mean, I, I was talking to various people and saying, don't worry about the randomness. You know, if, if you do this and that, that'll work. And the people would, you know, basically when you have a new idea, they'll beat you up. And so they beat us up and so on. And a few years later, one of
those people in particular had some methods and showed it worked, you know, not the blend people, you know, something around the same time and their students came back to me and didn't know who I was. I was lecturing me and saying, oh, don't worry about the randomness, it'll all be fine. And so on. There's just a cultural shift. And so there was a cultural shift that really was not, is as faster as this is as better. And so I, I, I don't think you're going to see the win here just because something's faster and better. You might need that to point to, but the win will happen because you can solve other things that
you couldn't solve before. So just as an example, say you're, you know, you want a better battery and you want to understand the way crack propagates better, you know, so that I mean, you want, you want, you want a longer lifetime, you know, warranty. So you, you're working on developing batteries. One of the ways this fits in, not in a scientific, but in an engineering or commercial context is you want to write a warranty to say that that the batteries are going to last so and so long. And, and otherwise you'll, you know, you'll eat the cost. And so can you predict battery lifetimes. And maybe if you have a certain crack propagation heterogeneity
structure, you know that it correlates better with how long the batteries will last. So now there's a fair variability and a particular run on the extreme tails. So that sounds sort of very different than I want to do oil exploration and better fracking or something. I send sound waves down into the earth and I look at how they bounce back and I try and learn something about it. But I could imagine that the embeddings that you get from a core model that cuts across a couple different domains learn spatiotemporal properties in particular of complex systems that have sort of long ranged, non trivial, heavy tailed correlations, which is what both of those examples have.
And so say that you can come up with a model that captures those embeddings. Now, a company here may tweak it to their data, do some transfer learning in the battery space. A company over here can do a better job adjusting it to their data, which is, you know, spatiotemporally very different in terms of, you know, better fracking. And, you know, for all I know, the same embeddings could be useful for, you know, you had a very different use case, right? And certainly in the United States, there's various places where insurance markets don't exist for floods or for earthquakes or for a range of things like that. I could imagine that you get these course embeddings.
These are course functions. They're not sinusoids and 4:00 AM modes. They're something that's data-driven embeddings. And you say on top of that, I want to do transfer learning to the relatively small amount of data that I have. That's very high touch. You know, you have 10,000 stations across the country that measure something in chemicals, whatever I want to do transfer learning to that and I model, you know, model the system, the way information flows in the Mississippi Delta or something. And, and if I do that, I can do a better job, you know, predicting the tails of flood events. Why? Because as a model, I don't know
how things propagate, but you know, subtle effects in terms of the density of the soil or, or the moisture of the soil coupled with, you know, 10,000 other things that the neural network learns in a complicated way. What's hard to reason about scientifically means that now I can create that market. And that's, that's something that any of those companies I could imagine try and you know, buy out a start up company that does that. So I think you're going to see a win. Not just I'm running faster, something faster than you, but but in a space like that. Yeah, no, I, I, I think what you're getting at is you, what we've seen from current machine that it seems to do things that
we don't fully understand, but are like you, like you gave the original example of the time series that an LM that wasn't explicitly trained on it actually ended up being quite useful. So I guess we shouldn't underestimate the potential of how some of these things may do something that we're, it's not just about replicating what an existing simulation tool could do, but maybe it could help to merge different disciplines together and see the links and the, and the, and the things between them. I did want to pivot to a different topic, which is when some of the other people I've spoken to, I'm always really fascinated by the curiosity or the uniqueness or the
peculiarity of academia and, and the sort of different people's careers and, and, and progression through. And I think you have quite an interesting one, because you are a very pure. Mathematician, academic, you know, you, you are very much in that, but you also have, you know, the Amazon scholar position you at the, you know, Lawrence Berkeley. Was this a a strategic thing in the sense that you wanted to work on real problems? Was this a just you find more interesting, like is there, How has it benefited you? And would you advise it to others? You know? Yeah. Yeah.
I mean, it's there's a lot of, again, a lot of questions. I mean, at least the narrow questions, universities versus not, yeah. Yeah, there you go. I'll start. That but, but, but yeah, you tapped enough stuff into the question to make it Ford compatible with lots of answers as it sounds like. I mean, so my, my, you know, OK, so my PF2's in physics, I never actually studied computer science or statistics. And it was in computational statistical mechanics. And at the end of the dissertation I switched areas basically to theoretical computer science to work on originally Markov chain algorithms, some of the stuff I've been working on computationally, Markov chains
and molecular dynamics. So it was a technical connection, but it was it was fairly loose. And then the randomization inside the Markov chains that the people I was working with were also working on these very early versions of these randomized linear algebra algorithms. So I got involved with that and I'd done enough coding that was clear certain ones are going to be useful and certain ones weren't. And so at that point I decided to switch areas partly because of what I've been working on. Well, two reasons there was a forcing function. This is I just mentioned this because if you have younger people listening, it's good to
know that sometimes older people have had textured past, let's say. So Long story short, I had career ending problems with the dissertation advisor and and the supply demand structure in the natural sciences and engineering is very different than in computer science. And so that was was motivation to, you know, jump off the Cliff. I also realized that that the particular work I've been doing in computational system mechanics, this is a great area. I loved it and and it's informed a lot of the stuff that I've done since then. But you know, it was a great area to be in a 1970 because you know, Ford compatible with a huge expansion.
There's going to be Nobel prizes given out. I mean, just lots of good stuff going on. Not so much in 2000. The area, it's sort of sad, but it was clear that that the future is going to be data analysis and, and I joke that I never knew more about algorithms or data than the day I switched. And then I've just been becoming more ignorant since then. So I don't know quite what I meant by data analysis then, but but but clearly I was right. I mean, it was a fruitful area for the next couple decades. So I done work in theory of algorithms and, and with respect to the academic question, if you're desperately wed to an academic career, you shouldn't do this because the academic
hiring is fairly siloed and, and it's, it's not a good idea to switch areas after your dissertation and, and so on. If, if you, if there's a forcing function like I had, or you're interested in sort of problems more broadly, then you can think a little bit more broadly. So I've had a lot of students in postdocs and I sort of give them this advice and you try and carve out various projects that they're interested in. Some want an academic route, so they should work on slightly more conservative things in the general space. Some definitely 1 industry and some, you know, could go with either and, and the ones that could go with either, I've seen
some go one way and some go the other. As a general rule, this is going to be a marketable space and stuff I do. And so maybe it's a little bit less of an issue, but but, but that's something that they figure out. So I had done work in in theoretical computer science and it was clear that these algorithms would be useful more broadly because the randomness entered in a very different way then classical algorithms. And so then the question is, how does it evolve? And so it evolved. I was at Yale in the mathematics department as in one of the junior faculty positions. I spent time at Yahoo and Stanford and moved to Berkeley.
I'm in the statistics department there and at the International Computer Science Institute of Lawrence Berkeley National Lab. And as you mentioned couple years ago started with as an Amazon scholar working in the supply chain optimization technology group, where, you know, we work on better forecasting, demand forecasting algorithms and supply chain decisions. So I think this is, this is something that's good for people to think about. And you know, there's at least, I guess 4 hats there. And each hat has pros and cons and pluses and minuses. And so figuring out how to navigate that space is something I try and encourage students and post docs to think about sooner
rather than later. How? How much should people expect to have to move? You know, I always think this is 1 of this sad in a way, things about academia, at least from what I've observed, that the fact that you do your PhD and let's say you do a postdoc, how realistic is it that that person in the US or the UK should assume that they can stay at that same university and just move up a chain? Or how much should they expect that they're basically going to have to move somewhere else in the country to find the position? Yeah. I mean, the short answer is that when you do a postdoc, it's usually more targeted because you're working with a particular person, a particular group and
and so you could be at the same place or not. It's usually a good idea to go elsewhere to get experience with other people, the real filters at the assistant professor level. And you got to move there because, you know, except in our cases, you don't get a job at the same place. And that's not a statement about whether the place where you are at is good or bad or whether you're good or bad is just just run the numbers, right. If there's 10:20, 30-40 places hiring and that you're looking at and and the chances are one and whatever that that you get an offer, it's just unlikely to happen at that same place even if ultimately end up there. Sometimes people go and leave
and come back. So moving for better for us as part of the part of the equation, yeah. But did you, some of the questions I get is the should they go into industry or, or just let's put this two ways, does academia value industry? So if you've been an assistant professor or you've been a postdoc and you then say, you know what, I'm going to go into industry, do they value then? And is the ability to come back or have you been out of the game too long? You haven't sort of, you know, got the publications or the teaching experience, You know, is that a good piece of advice or would you say no, no, no, If you really want to be an academia, you basically need to
stay in in academia. Yeah, I mean it, it's, it's more texture than the following. But I, I think the short answer is that that that you need to stay. The slightly longer answer is it depends on the area, like computer science versus statistics versus engineering are rather different. In some cases having a startup, in some cases having a postdoc in industry is good and you can go back to universities and that tends to correlate with computer science. I don't know as much in engineering that may be changing over time, but less so. And maybe statistics and applied math. I think in a sense, on the one hand, people value the
experience you might have, but not really. And by that I mean, you know, if, if you, if you check all, if if you check all my boxes. And in addition, you have this other stuff, good, but you got to check all my boxes and the boxes are the things you alluded to. And so there's sometimes there's post docs that at industry that are basically academic post docs, you write papers. So I'm not counting that because that's, that's the fact of the same sort of thing. But if you're out more than a couple years and you have fewer papers, it just gets harder to publish. There are certainly cases where people do that, especially if they have some high profile ones
and, and can work their way back one way or the other. But it's, you know, you should know going in that you're, it's like a salmon swimming upstream. And this is going to be a hard one. So how does he, this is maybe people in the US are maybe a little bit more familiar, but particularly people who are not. So how does it work with the national labs? So you have a position on the national labs. How, how does that work between is this a, a quite common thing in a way for these things to happen? I I can only imagine how you juggle the e-mail inboxes and the sort of meetings, but I guess it's valuable. And interesting off an hour
today and I haven't been hit by e-mail. That's what I've turned it off. That's the accept e-mail is a little overwhelming. I mean, it's not so common. So, so I, I with the Berkeley hack, because I'm in the statistic department there, I'm at Lawrence Berkeley lab. It's literally just up the hill. And so I can, you can walk there. And so they worked in scientific machine learning and, and I had done work and was well known in machine learning and, and algorithms and statistics, large scale statistics. And I had various projects over the years with people there and elsewhere on scientific ML problems. And so started about two years
ago. This is all we rename, but the scientific machine learning group, basically MLA machine learning and analytics. And and so it's it's it's part was partly facilitating for the fact that there there is a moderate amount of interaction between UC Berkeley campus and LBNL Lawrence Berkeley National Lab, Oak Ridge Argon. Some of the other labs have things like that, but it's not so, so, so common. Most people are there, meaning just there and 100% of their FT ES there and and there for long term but so it's not. I wouldn't say it's very common. OK, which makes your experience even more unique then to have done it. Yeah, what maybe is a sort of
finishing off comment or, or topic is if you have somebody now who was wanting to get into scientific machine learning, because this is I guess kind of what we've been largely talking about, what would you get them to focus on? They're doing a PhD or they're doing a postdoc. But is there any, doesn't have to be a very, you know, Pacific, Pacific, but what sort of area we'd say, you know, this is something you should look at. This is an area that is ripe for, you know, progress. Yeah, that's a good question. There's certain things, I mean, that are less ripe for progress
in scientific ML, even if they're important scientific problems. So what I, what I, what I would try and say is work. You should work on a problem that's of interest to both sides. Otherwise you're coming at it from one side or another. And this is true whether you're coming from CS or stats learning scientific problems are coming from one scientific area. You should work on trying to figure out how to frame what you're doing as something of interest to both sides. So I have projects that, you know, we, we, we construct projects, you know, student, a postdoc owns a piece of it. And, and a goal typically is we want a publication in the
particular scientific area, but also in an ML venue. And it need to be the same one. Sometimes it is, it's a long version, a short version. Sometimes it's just two different things, but work on a method that's broad enough that machine learning people are interested in it and that by the definition of machine learning, people are interested. You convince 3 reviewers to say yes. And so it's an imperfect process, you know the review process, etcetera, but boom, you get it in a top machine learning band room. And similarly, on the scientific side, you know, a, a statement that it's a value of the scientist is you get 3 reviewers to say accept and then it's
accepted on the scientific side. And oftentimes the way we try and scope our projects is to say, you know, we don't on the machine learning side, we don't want to work on your problem if it's only of interest to you. If I can't apply it to some other area, someone else, some other domain, Because if that's the case, I probably need to know so much about your area that I become, you know, an expert in your particular area. And similarly, on the scientific side, you know, if you have a particular method that is only that's so heavily tailored to your domain that it's not going to be useful more broadly to machine learning people, that's a much harder sell.
Then you're not doing something that satisfies both sides. So something that I knew about from working on years ago was like sequence alignment in in genetics, right? Clearly an important problem, not something that's portable the particular album, not something that's portable to lots of other scientific areas, but better machine learning methods for solving spatiotemporal forecasting problems with non trivial boundary constraints. That's clearly of interest to a pretty wide range of applications. And also to solve it, you're going to have to introduce technical solutions that are probably, you know, go, can you come up with a differential
optimizer for an end to end differentiable system where you have hard constraints? I mean, that's clearly something of interest to machine learning and optimization. So I guess I'd try and focus on something there. And that is probably the way that you're going to be. And unless you know exactly what you want to be doing 30 years from now, that's probably the way to make yourself most forward compatible with with however. That's a really, that's a really good point. I guess what you're saying is, yeah, it's true. Like later on in your career you've become established in a certain area, but most people during their PhD or the postdoc
at that time have no real necessary idea. So you're saying if you do something broad enough that can get you exposed to different groups and and have a foundation to your earlier comment, you can more easily go into verticals because you've built that foundation. Whereas I guess if you jump directly into a vertical right from the beginning for you to sort of reverse out that it's kind of harder, isn't it? So I I think that's what you're alluding to it, it sets you up better. Yeah, yeah, because it's not at all. I mean, all the questions you're asking are fair ones. And it's not at all clear how the whole area will evolve.
And so depending on how it evolves, having experience in one versus another. And, and I think I mean, having worked in a bunch of areas, you know, it's, it's, but, but speaking substantially only English, but a but a little bit of other things I can imagine. And I see, you know, you learn a second language, it's hard. You learn a third language, the, the relative cost is a lot less. You learn a fourth. I mean, so there's a diminishing returns in terms of the amount of extra effort you need to do. So you work in one area, switch to the second. It's very hard switch to the third. You know, you can, you know, you can, you can do that. And after that, one thing I've
sort of been always struck by is it's amazing how little you, you need to know about a particular area, especially if you're working with good people who'll complement what you're doing. You can learn from them, right? And so it's amazing what you don't need to know in order to, to get interesting results in an area. And so, and so being learning new things. By the time you know two or three things, it's easy to learn the 4th. Yeah, yeah, I know that that makes sense. Well, thank you so much. I it's the day after Labour Day, which means that probably there's a whole bunch of emails and stuff that's started to come through on essentially your
first day back. And I know we probably could have carried on for for another couple of hours. But yeah, thank you. Really, really appreciate it and look forward to catching up in person at some conference in the future. Sounds great, thanks for having me, this has been fun. All right. Cheers. None.