The Neil Ashton Podcast
Prof. Max Welling — Machine Learning Pioneer and AI for Science Visionary
Watch on YouTube
Prof. Max Welling — Machine Learning Pioneer and AI for Science Visionary
YouTube video
Watch this episode
YouTube is contacted only after you choose to play the video, keeping this page fast and private by default.
Listen to the audio
Episode overview
In this episode, Neil interviews Professor Max Welling, one of the foremost experts in Machine Learning about AI4Science: the use of machine learning and AI to solve challenges in various scientific disciplines. They discuss and debate between data-driven and physics-driven approaches, the potential for foundational models, the importance of open sourcing models and data, the challenges of data sharing in science, and the ethical considerations of releasing powerful models. The conversation covers the role of academia, industry, and startups in driving innovation, with a focus on the field of AI.
Professor Welling discusses the advantages and limitations of each sector and shares his experience in academia, big tech companies, and startups. The conversation then shifts to Professor Wellings new company; CuspAI, which focuses on material discovery for carbon capture using metal organic frameworks and machine learning. Prof.
Welling provides insights into the potential applications of this technology and the importance of addressing sustainability challenges. The conversation concludes with a discussion on career advice and the future of AI for science.
Chapters
- 00:00 Introduction to the Neil Ashton Podcast
- 00:39 Guest Introduction: Professor Max Welling
- 11:12 Data-Driven vs. Physics-Driven Approaches in Machine Learning for Science
- 17:00 Foundational models for science
- 23:08 Discussion around Open-Sourcing Models and Data
- 29:26 Ethical Considerations in Releasing Powerful Models for Public Use
- 33:14 Collaboration and Shared Resources in Addressing Global Challenges
- 34:07 The Role of Academia, Industry, and Startups
- 43:27 Material Discovery for Carbon Capture
- 52:02 Career Advice for Early-stage Researchers
- 01:01:07 The Future of AI for Science and Sustainability
References and links
Transcript
This transcript was created from the corrected YouTube captions, with names and technical terminology reviewed. Download the corrected SRT file.
Hi and welcome to the Neil Ashton podcast. In each episode, we explained some of the fascinating ways that science and engineering are changing the world around us. We talk to leading engineers from elite level sports like cycling and Formula One to some of the world's top academics to understand how fluid dynamics, machine learning and supercomputing are bringing in a new era of discovery. We also hear some of their life stories, their career advice and lessons they've learned on the way that I hope will be helpful to you too. So sit back and enjoy this episode. Welcome back to the Neil Ashton podcast. Today's guest is Professor Max Welling,
one of the foremost experts in machine learning. And I was speaking to him in this episode about AI for science, the use of machine learning and AI to solve some of the biggest challenges in science. Now, if you're in the machine learning world, you'll already know who Professor Welling is. But just in case you're, you're not, I'll just briefly give his background of why he is such an expert. Well, as we actually discussed in the podcast cos I I always like to ask people about their academic careers and how they got to, you know, where they got to, uh, he started off, you know, the, the standard academic track, I guess,
doing a master's and a PhD. Uh, and then as he said, he actually did many postdocs, um, moving between different universities. I, I guess, most notably he was with, um, Geoffrey Hinton who is seen as the sort of godfather of, um, of AI at U of T. Uh But then he went over to, to the US uh to UC Irvine and he started being a professor there. Um And then he ended up actually going um and being, being, yeah, an assistant professor there and then coming back to the University of Amsterdam to become a professor and he's actually stayed there uh pretty much ever since. But one of the interesting thing is that he has taken routes into industry and,
and start ups. So he was a VP at Qualcomm. Uh And actually very recently, he was also a VP and distinguished scientist at Microsoft. He recently left um that position and to start a new start up called CuspAI, which is working on machine learning for material discovery. In, in the topic of carbon capture, we actually go into what his startup's about what they're doing, why he's interested in this area. Um I probably could spend a long time going through all his achievements. Um But he, he's one of those people who is, you know, at the forefront of, of machine learning. So he's been the people in charge of all the major conference like
NeurIPS, ICML and ICLR conferences. Um He has been leading, if people haven't seen that, I would highly recommend that you do look some of these workshops on AI for science, uh, particularly at NeurIPS the past few years. So these bring together all the world leading experts on, you know, uh machine learning and its and its application to science and bring them all together in workshops that's has been really good. Um And what I like about him is he has that unique experience that we, that we touch about being not just a pure academic or being a pure industrialist, but someone who has jumped between. And I really do believe that that gives him uh somewhat a unique insight
into not just the science but the practicalities of how do you go about it and that's exactly what we try and discuss a little bit in this. We purposely keep it reasonably high level. So we're not going into the super super details but rather looking at the the big topics. So for example, we start off by discussing data driven versus physics driven approach. It as in do you need to include the physics, physical equations into the loss functions into machine learning? Or can you just let the data do it? We talk about uh foundational models as in, could you create a foundational model, the science that you can just use or, or just, or just fine tune.
And then we pivot a little bit and what I think anyway, at least have some interesting discussions around. Should you open source the models and the data? We have a good discussion about the merits of that um ethical considerations, et cetera. Uh And then we get into the different mode of operation, what the advantages of academic industry and start ups in terms of reaching these benefits. Um We talk about his start up about CuspAI we talk about what they're doing, why he's doing this, his excitement and motivation to go back to a start up. We briefly took on some of the funding, you know, what's it like getting funding from VCs in Europe versus the US?
Uh And then we, we sort of finish a little bit discussing career paths. He goes into more details about his personal journey through academia, through industry, his advice for people now. Uh and finishing really on what his uh thoughts are on the future of machine learning and how this could help scientific discovery. So a pretty wide ranging uh just topic like anything we probably could have spoke for many more hours. But I, I certainly learned a lot through this conversation and I, I hope you enjoy it too. So uh sit back and please listen to this episode with Professor Max Welling. do we start it now? We have? All right.
Uh Yeah. AI for science, I think, um, is the application of the tools that are being developed in machine learning and AI um to basically help the discovery process and the analysis process in the sciences. Um The interesting thing is that there's also signs for AI in a way which is the, you know, using principles from the sciences and the mathematics such as um you know, symmetries, um diffusion processes like thermodynamics to build better machine learning models. So this thing actually go goes two ways, the causal direction is going two ways. Um Now, what we see is that in the sciences um in in a very broad
spectrum of the science is going all the way from the very small uh you know, PICO secons uh femto meters, which is basically uh high energy physics where a lot of machine learning is u used to do data analysis all the way up. Uh you know, let's say going through molecules and liquids and fluids and then all the way to sort of earth sciences and maybe going all the way up to uh astronomy um which is the at the scale of the universe uh machine learning mo um models and methods are being used to, yeah, model that particular domain, but also analyze the data better. So it it it has been having a huge and very broad impact in the science as I would say,
and how much overlap is there, would you say between the work and the methods that have been used to develop, let's say some of the large language models and those which are in the sciences, is it a matter of applying the same tools and technologies? And it's just essentially a different focus area or do you see that fundamentally there are differences and new approaches that are need to be developed. There's a surprising amount of overlap I would say um which I find interesting. So the fir the first thing to say is that even within this incredibly broad range of um scientific disciplines, of course, the tools that are being used,
the mathematical tools that are being used are often very similar too. So you start from you, you know, it's either a partial differential equation or stochastic differential equation or ordinary differential equation. That's because physics is causal, you know, you try to predict the future from the past. Um It's typically continuous but not necessarily, but many, you know, processes are modeled continuously like fluids a me and it's local, which means that you, you typically don't have something at a very far distant place, instantaneously impacting things here. And so these three aspects make the tool set typically quite similar,
even even across all those scales. Um So now to the question, um some of these tools which are being developed for image generation um perhaps more than for large language models. So they're actually different tools, although they're all being, you know, put in the same bucket typically as generative AI but actually quite different. Um But especially the methods that have been used for video and and image generation, which you can actually somewhat think of as a physical object with pixels being microscopic degrees of freedom which have some dynamics. A me. So the tools and diffusion models for instance, which is the method that generates all
these beautiful images and videos these days, those can be almost literally used to generate molecules um and fluids. So um maybe not surprising because of fluid, you can think of it as a sort of a set of pixels or it's typically modeled as a set of pixels, you know, it evolves over time as well. Um For molecules you could argue, you know, those are, you know, it's more like a graph uh where the nodes are at certain locations. Um So maybe they're slightly more surprising, but it is almost literally the same tools. In fact, we developed uh what's called equi equivariant methods which incorporate the symmetries of the world,
which is that a molecule or an image. So it may be best for an image if you look at sort of a histopathology slide, which is a slide which has all all these kind of stain cells in them and some cells might be uh have a tumor in them and and other ones don't. And your task is to find the cells which are malignant. Um So, you know, if you turn the slide upside down, you know, you don't even know. Right. It's just the same set of cells, but this rotate a little bit and, but a neural network doesn't know that usually. And so we've built models which were called equi equivariant models, which build in sort of symmetries.
So for instance, you know, those methods, these symmetries which we've developed for fission very naturally also transfer to but back to physics. So they came from physics in some sense and we turned them into tools for, for, for images. And now, now, now we apply them again to molecules, right? And so a molecule upside down is the same molecule, a molecule displaced from here to here is still the same molecule with the same properties typically. So that's why these tools are very useful, you know, across the board. But one of the um topics that seemed to be very much an active debate in at least in the machine learning for
fluid dynamics, which is maybe what I'm most familiar with. But I, I'd be keen to get your opinion on this and more broadly to other scientific disciplines is this sort of physics driven versus data driven. I see on one side, the logical argument there are physics laws surely, you know, the solutions should bound by them. But I I saw you gave a talk recently and you were alluding to the idea of there's a sort of cut off point where if you don't have much data, maybe it makes sense to use them. But if you have a lot of data, you may even be constraining the problem and, and, and solving it. What, what are your current thoughts on this debate around
physics or data driven approaches? Yeah, very similar. So I, I'm still, I agree with my older self. So that's good, at least some level of consistency. Um But uh so, for instance, uh you know, when I was still working at Microsoft with a wonderful team there, um uh we worked on models for the atmosphere. Um And this was recently published as a paper called aurora. Um So there, there is a lot of data. So it's like a uh petabytes of data for because, you know, we predict the weather every day, every hour. Um and we store all that data and then of course, we know what the actual weather was, at least we have measurements.
And so it's a very rich um and outdated set of data in some sense. Um in that domain, you can basically let go of all the laws of physics. You can just say, let's think of this as a big machine learning problem. We just try to predict the state of the atmosphere, you know, a day ahead, you know, 10 days ahead, whatever, right? Um And you can just use your transformers and everything that you use for image and videos as well. Um And that worked really well. In fact, it's, it's basically as good as a me numerical soul verse. I should be a bit careful saying that. Um But it's as good um with if, if the, if the prediction stays within, you know, the,
the the sort of the input that it was trained on, if it goes far beyond that, you have to be much more careful, of course. But the big advantage is just it's way, way faster. So it can be up to 10,000 like four orders of magnitudes faster. Now, there's other domains where this is much harder where you don't have this huge amount of data in fact. And in that case, um you should really build in um the what we call inductive biases that are coming from the laws of physics. And, and I've seen some interesting examples where, for instance, we know the continuity equations should hold typically, if you write down sort of the Navier–Stokes equations,
many of them have to shape of some kind of continuity equation. So that, that's the, the partial differential equation that describes basically a velocity field over which the fluid moves um in three dimensions. And then there's a source term basically which you know, applies extra forces to that. And um and so um sort of if you do it with the transformers, you know, it has to learn all that. But um it makes a lot of sense to say, well, we know the world is three dimensional and, and, and it has this sort of, and we know how fluids move around. We just may not know precisely what the forces are that, that are acting on them.
And so we learn the forces we, but we constrain it to be a fluid that moves around uh according to the laws of physics. And there you know, if you have little data, then that's a very good prior. Now, the flip side is there's always a flip side. So if you, you know, we don't know the precise laws of physics at that scale, even for the weather, right, we have written down or we, the demetri ologists have written down and I don't know order 1020 equations which describe all the complicated interactions of the, you know, the atmosphere with, you know, the surface with, you know, the radiation with chemicals and all these things.
So everything is written down in these equations, but these are approximations. They are not actually, you know, we we might know the laws of physics at the very microscopic level, you know how one atom sort of interacts with another atom using quantum mechanics. And all that, we, we know it very precise, but it's completely hopeless to simulate the atmosphere at that level. So we have to Coors grain is what we call it. We have to go to a much more um sort of a less precise level of description. And at that level, things become approximate. And so if you impose more and more laws of physics, which is, you know, you've approximated
at some point, uh you might have so much data that, you know, you, you sort of build in a ceiling to this modeling, you say I want you to fulfill these constraints. But, but if these constraints are approximations, um at some point, that means that you cannot get out of that box. And if you have so much data that the model can actually learn something that's more precise than that box than those constraints you're giving it. Um Then you're constraining the model over constraining the model and that's, that's hurtful all of these, at least from what I can see most of us, when we use a large language model, we never need to train the model.
We are just doing an inference essentially on a pre trained model, a foundational model. And most of us never need to, to train it ourselves. Whereas if we take a fluid dynamics or, or a weather problem, do you think that we will ever be able to get to a truly foundational model which, which is trained on such massive data that it, it could work for a plane or a car or a building or the weather could be done in, you know, different parts of the world or anything or, or is, is science a problem that's just much more challenging than um than what we've done for the large language models to date. Yeah, it was a very interesting question, I think on the one.
So to me, this is kind of a, I think it does make a lot of sense to train a very large foundation model. And this means plenty of evidence now that if you collect data from diff from very different domains, different regimes um training such a pre training, such a very large foundation model um is a huge benefit, especially if you can then fine tune it to the particular problem that you're interested in. So with a little bit of data and a little bit of fine tuning, you can get a very good model. So I think that's a paradigm that has now been established. And I think that's, you know, I think that that that will just work for the foreseeable future.
Um But there's also this thing that um which are called the 8020 rule, which is that it's 80% of getting it right is quite easy, you can do it with 20% maybe of the effort, but then the last 20% is incredibly hard. Um And you may be you may you need a lot more effort, right? And so I think you can see this very clearly for self driving cars, right? So in self driving cars, you know, people get very enthusiastic for the 1st 80% because, hey, we can drive a car without hands on a, on a street. How wonderful. And then you realize that to get it to drive inside Amsterdam, um, you know, that's not gonna work with that first,
you know, with this first, you know, model you have. And so you now need an enormous amount of training, but actually probably just rules based methods as well. You know, to make this thing drive safely inside a complicated city like Amsterdam, maybe San Francisco works. I don't know, the streets are all rectangular and everything, but in Amsterdam it's all big chaos. Yeah, people run red lights, bikes everywhere, you know, it's like massively complicated. Um And so it will take 80% of the time maybe to get this long till of issues correct. And this is where humans really are very good somehow, this has to do with this generalization.
Somehow we do this very naturally and very good. Now, in large language models, I'm sort of seeing something somewhat similar, right? Which is we can get amazingly good at, you know, and, you know, I didn't predict this to be really honest, you know, to, to build these, these um these chat bots, but it is, it is uh perhaps 80% of the way. Um Now the last 20% is you should get it such that, you know, you can't trick it into saying all sorts of rubbish, like, you know, put glue in your pizza or eat rocks, right? These are the kind of things that are big tech companies are now struggling with because, you know, if you put this out there,
people will try to, you know, to game it. Um, and that's very negative publicity and all that. So, so this last 20% to make, to make it really understand common sense. Um, you know, it's, it's arguably whether this even can be done by just looking at text, maybe this needs to be some system that grows up among humans, you know, with the body and sort of interact with humans and understand social interactions and, you know, ethical rules and all these kinds of things. Maybe it's just, you can just learn that from text and that last 20% might be very, very hard. And so perhaps in, you know, in physics, it's the same, right?
So, you know, you can get it, you can get it right? 80% and that hopefully is useful. Um But then how do you protect yourself against that 20%? You know, an unseen initial condition with incredibly important consequences like OK, it might just, we might just get a heat wave, 50 °C, you know, do we, you know, do we warn the public um or do we not because we, we don't know precisely what's going on, right? And that I think is incredibly important to get Right. And we haven't really get our heads around how to do that. And I think uncertainty quantification and reliable, uncertainty quantification is going to be key for that.
Very interesting. I haven't really thought about it that way. It's true. The 8020 and now I think about it, how I use these large language models is I very rarely take the exact output. It gives me, I normally will iterate on it, but it enables me to get somewhere much faster, I suppose maybe from the physics side, we're expecting it to get perfect where really maybe our expectation should be more similar that it gets us someone closer and maybe then we feed that solution into a traditional solver precisely to do the next bit. Um And its initialization like a smart initialization or something. Yeah. Yeah. But all of this,
I still wonder about the data challenge because it seems because of the internet and maybe things have been clamped down, you know, so many images and so many text are publicly available that you can get. Whereas with science, it's very rare to have so much of that data in the public. To me, most companies have that behind firewalls. Even people publish papers, but they very rarely publish the entire data that created it. So how can we get around that issue of of the data? Yeah, that's the absolutely key. So that's the, that's the key for progress I think. So, so you can see that the domains were, you know, researchers have managed to do a lot more data sharing
or whether it's just more data available have progressed fast. Um certainly, you know, with machine learning tools. Um But for many other domains, uh companies are hoarding that data or even universities are hoarding their data because it's expensive to generate and they want to, you know, milk that cow for a couple of years in terms of papers before they give it to others. Um And so that is really um slowing down science and progress. Um There is, there is great initiatives where, you know, for instance, in, in bio where people aren't sharing a lot of data, there is a lot of omics, data available, et cetera. So there's, there's also good examples,
there's some materials project for instance, and other materials projects um around the world where data about materials are being collected. Um But um I think in the end, we will have to invent the technology that makes data sharing much easier. And, and I, you know, would envision some kind of marketplace um where people have their data behind a firewall. Um But um if I want to train a model, I just go to your, your data source and I'm going to say um I'm gonna pay you whatever amount of money to access your data source. Now, um note that I will do it in a differentially privacy private way which is which is to guarantee that
I'm gonna update my model parameters, but I'm gonna do it in such a way that I can never reconstruct data that were, that were, that were used inside your data source to, to build my model now, that can be guaranteed. But that technology also needs to be improved. So that means that, you know, you know, your agent and my agent, they negotiate a bit about the price and then I enter, you know, do my model updating and then I move to the next uh you know, data source, right? So now we can train really excellent models. Um but you know, there is kind of trading going on about how much that is worth. In fact, you could do it for your own data. So I could say
I've collected all my data over my lifetime, including my medical data, whatever I'm, I'm quite happy to help that, that hospital, let's say for free. Um But when Google knocks on my door, I want money because I want a fair share, right? And so basically this means or actually for, for open AI kind of, you know, ChatGPT models, right? Just like they now sort of there's lawsuits now, right? So basically, you know, it's unclear, you know, what these models like, you know, let's so I think there's a recent lawsuit against some of these companies that generate songs from, you know, from different rights from songs that are being created by artists.
And so these artists feel, you know, I've created this content and as being, you know, and somebody else runs away with it and it's, and you know, where am I in this process? This is a fundamental problem that we have to solve now. And it would be much better if an artist, you know, or a person who collects data in any way or generates content or data gets paid in a, in a reasonable way or can choose to make their data available to somebody who wants that data to train their models on. This is a fundamental problem across the whole machine learning. And AI think that space that needs to be solved. That's fascinating. I
it seems very interesting and maybe complimentary but in some ways also at odds in that, on one hand, there is a movement of open source open data, open source. So the idea being that everybody should make their data sets available. But and I believe that some of the um weather forecasting companies who released those data sets are now debating whether the next data should be closed and not open because it's true that many companies in our published models have have, you know, great success, but the people who generate the data have no financial reward for that. Um And I feel that for the sciences, particularly most applications of science get into industrial use
where there is a commercial side to it. So just saying open source your data, well, they're probably not gonna open up their data. So that's a very interesting point about it. A way of a way of people sharing their data but also being able to make money and privacy because I can't see a way of, let's say fluid dynamics. Why would Boeing release all the aircraft data for train? Why would Airbus, why would anybody else do it? It's a competitive advantage, it's a competitive advantage, which is why I'm always wondering and then universities typically do more fundamental problems. So if you train the data on just university data,
it's only gonna work the simple test cases because universities are incentivized to publish pure research, typically not full aircrafts. Um Yeah, I think the Blockchain could play an interesting role here somehow, right? So you, you could imagine that you could sort of ii I believe there should be some kind of marketplace for data. So first of all, we have to realize that data is perhaps the most important, you know, uh what do you say, resource? You know, it's, it's kind of the the the the new oil on which these machines, you know, and compute, I guess, right? So it's, it's computing and data that these machine learning methods need
and we just need to put a real value on it. We have to say, you know, it's, it is valuable if you share your data or give your data just pushing people to put it open source is I don't think it's gonna happen because everybody will always, you know, do what's good for them and we just need to make the incentive structures correct. I think that's the thing. You don't force people to do things to make the incentive structures correct. So people who are open sourcing their data probably, you know, they work at universities or they work at companies and then the open sourcing works for them somehow anyway, because even companies who are open sourcing their stuff probably do it
because in the long run, it will benefit them as well. What is your general opinion of that open source though? Versus close source? It seems to be a bit of a debate in the community of it. Does it help science to open source? But then how do companies fund themselves if everything's open source and how easily it can be copied? Do you have a, a thought on the, the optimum approach that you've seen different companies over the years take? Um I mean, are you open to talking about open sourcing data or open sourcing their models models more? But yes, I guess also, well, we maybe we just discussed the data, but how about the models themselves?
I guess the discussion is about whether it's dangerous to release a model, right? Um Because the, the model can be used for at the zeal purposes and is being used for a zeal purposes. You have, you know, you have to be careful that if you enable bad agents um to do horrible things. You know, you, you just have to really think about that. So, in other words, um if I would have the recipe to build in a lab, um you know, a terrible disease that can be, you know, unleashed on the world. I'm not going to argue for open sourcing that knowledge, you know, clearly, I just don't want that to be open out there. Um And so you could argue similar things will apply at some point
to machine learning models or actually maybe already, I mean, at some, you know, these models can be used for fishing attacks and all these kinds of things, right? So, so, or, or crime where you sort of um uh basically uh imitate somebody's voice and then you, you ask them to, you know, pretend to be your son and then, you know, ask to, to, um give some money. So you have to be extremely careful about, you know, when you open source something and maybe there just needs to be some kind of ethical board that decides whether a particular very powerful technology can be open source, yes or no. Now, within that, you know, once, you know, if you're within those bounds of safety,
um it's up to everyone to open source. It's, you know, it's a com you cannot tell a company to open source or not. Right. It's like it's, I think it's, it's not, not a very meaningful question to say companies should open source their stuff. It doesn't because of course they won't if it doesn't, you know, improve their bottom line. Um So it's up to them to open source, whatever they need to open or want to open source. But maybe a more specific question then is there a can methods move faster that have been open source because essentially you're outsourcing the development to a community effort or is it the case that
that sounds nice, but you can actually move faster by having an internal team with their own developers? Yeah. Yeah. So um so I think thinking very hard about the incentive structure. So if, if it's true that if you're a whole number of small play, you know, a small player cannot compete with a big player, right? So it does make a lot of sense to put a whole bunch of small players in a group, you know, uh share resources and compete with the big players. Let's say you all built one huge large language model uh within your group of smaller players and then everybody benefits from that large language model and now you can compete with the big players. So that make that's a clear incentive to
share in that group. Um Now there's also other reasons why people might want to share, which is like we're all stuck with a problem together, which is, uh, you know, climate change. Right. So for some strange reason we think that, um, we can pollute this earth, um, and, uh, you know, push carbon into an atmosphere which is a shared resource, um, but not pay the bill basically for doing that. Right. Um, we, we, we, we somehow think that's ok to do and, but I think there is a sort of ethical awareness within a lot of companies, big companies, big tech companies definitely who are saying, well, we should, we should all act together to do something about that, right?
Um So most of these companies have climate uh sort of um promises or what do you call them? Um sustainability, sustains and things like this. Um So to be carbon neutral or negative by 2030 all these kinds of things. So, so they, they're trying, although they're tied up in an AI arms race and so it's very hard to actually make those goals, but let's say that they're, you know, they're trying to do this. And so there it makes some sense to pool resources, I think and say, OK, let's be good actors in this world with good citizens and actually, you know, you know, release some of the data to the public so others can,
you know, help, for instance, a, you know, improve the technology to capture the carbon out of the atmosphere. Hm The um I wanted to maybe pivot a little bit um to, to, to talk about your new company, but maybe as a way of getting there, one of the things that I admire about what you've done and I'm always interested to hear the different logic is the role of academia, industry, big industry, like big tech companies and start ups. What, what can a university achieve? What's the advantage of a start up? Essentially? What, what, what, what drove you to create this new company? Is there a specific thing? I, I can just move faster by doing it as a start up versus
I'm interested to see how you've seen over your career, the sort of benefits of these three different places. Yeah, I find it very interesting because there really three modes of operation. I've tried all of them now so I can sort of sampled them a little bit. Um Now universities have a very special role of educating the next generation of talent that's really important and, you know, they use taxpayer money for that. So, you know, that's incredibly important, but they also do um so much more exploratory research, right? So they, they can, you know, uh if you like to think about the inside of a black hole, you know, fine, do it in a,
in a university setting. There's no company that will, well, there's very few companies that will fund that in some sense, basically. No, I think at this point. Um So, um but the impact is very diffuse. Right. So you, you, you write a paper, um, and that paper gets picked up maybe by different groups and they build on it and that's a very indirect, uh but important way to make impact and, you know, it's also, you know, of course, you, you, uh, it's also a different style of doing research and, and, and being engaged in intellectual activity. Um now the big companies and the start ups, you know, they can um command much bigger resources to really go after a specific target, right? So, um
no, but they are commercial, right? So they are tied to commercial goals. And so you have to really align the commercial goal with the thing that you want to achieve. Um Now some people like, you know, if you, if you like to achieve AGI, you know, great. So you can work for one of these big companies. Um that's a perfect alignment because that's what these big companies want. Do. They think they can make big money on that. So, you know, that's where you can sort of do that, that research and, and you can use the actual resources like massive amounts of GPUs um to achieve that. But big companies are also in a way slow and political
beasts, right? There's like, uh there's a lot of people that need to, uh you know, agree on something and, you know, a lot of legal involvement if you wanna change something small. And so it's, it's, it's, it's very uh fisc this whole thing talking about fluid. You know, it's like it's very fisc thing. So, um and a and a start, the great advantage of a start up is that um you're your own boss. So if you have a vision, you can execute on your own vision and you don't have to constantly calibrate with other people higher up in your organization, you just need a great set of co-founders and a great team to execute who are aligned with that vision. And then you can really go.
Um So to my surprise, um perhaps is that um the amount of resources you can get in a start up these days in the field of AI is no less than what you would get in a big tech company. So, in fact, um you know, if you can tap into VC capital, um I've had until now, incredibly positive experiences with this, which means that uh first of all, you know, these are very smart people, you know, of course, they, you know, they want, of course get return on investment. So there's no idealism here mostly. Um but they're very smart people, they help you find a market, they help you build a healthy business. Um And uh and there's lots of money available as well, right,
for that particular purpose that you want to pursue. Um And so, and so I think the agility and also the positive attitude that you find in start up land I really like. So I, I guess I've now converged on that particular model and I'll, I'll, I'll stick with that now for a while and, and just before we get into the Pacifics um of, of your company, is there any differences or challenges between, let's say, in your experience, us, Bay area, Europe, you know, other advantages, disadvantages in terms of finding the staff, finding the, the the money. Well, we actually do use uh American and somewhat and English VCs,
mostly there's also a few continental Europe, Dutch investors, but mostly it's uh we tapping into uh you know, um UK and US based VC firms, although our lead investor is actually is a UK based investor A me. So um so I don't think you need to make that distinction in some sense. The nice thing is that these big investors, they're turning their side to Europe. And the reason is that there's just incredible talent pool in Europe. Many of many of these talent pool likes to stay in Europe because Europe is just great, just, just fun place to be and the salaries are a lot lower than in Bay Area.
Um And so with relatively little investment, well, less investment than what you would have to invest in Bay Area, you can get a excellent team. Um And so it only makes a lot of sense that also these investors turn their eyes on, on Europe. So I think we will see a lot of growth and activity in these, in these sort of uh start up and scale ups in, in Europe in the coming years. I've often um I tease some of my American colleagues that why are they all based in Seattle? It seems like the most crazy time zone and position in the world to have such a, if you were going to invent a company to have a global reach, surely you wouldn't put it
there. I still don't understand how this, but yeah, something to do with ecosystems, right? So, um so, and, and that's really the thing. So uh if you build a particular nucleus or particular geographic area where, where it's lots of things are happening, right? In Silicon Valley, you have the big universities like Berkeley and Stanford, um which are providing a lot of talent and there is a lot of entrepreneurial spirit in those places and there's a whole, it think of it as a supply line sort of issue where there's a lot of, you know, other things which work for them. So there's a very small, you know, everybody knows each other in the VC.
So there's very, there's a lot of VCs which can fund these things and, you know, so I think it has to do with the fact that once you build a really good ecosystem locally, it then becomes very attractive to be there to do your start up because it's very easy to get good ideas from, you know, people help each other to grow these start ups. It's very collaborative atmosphere, it's easy. You know, if I need a new investment, I talk to a few people and they point me to other people and then you, you know, you before you know it, you're in front of new investors. So I think that kind of uh you know, lubricant in some sense is really good in, in these ecosystems.
So what we need to do in Europe is to build these ecosystems. And I think in Cambridge and in London, these ecosystems exist so they are very good as well. Um Amsterdam, Berlin, you know, some other places, Paris for sure. Um maybe Stockholm. So some of these places in Europe also start to develop these ecosystems, but we need to get better at that in Europe. Yeah, I I do definitely agree that there seems to be a lot of um a lot of very talented researchers and engineers and scientists in Europe that maybe feel a pressure to perhaps to move to, to the US where probably a lot of them would probably rather stay in Europe if they could do.
So. There is that I agree there seems to be a desire and if you could tap into that, you've got a big uh pool as and the other benefits of time zones and access to, to different markets. But yeah, I'd, I'd love to learn more a little bit um about your company. And, and maybe also, I think listeners to this podcast may be a little bit more familiar with uh some of the topics we discussed around fluid dynamics. So it would also be good perhaps for you to align how some, maybe the material discovery and material science has some overlaps and maybe uh differences to tonics. So the start up is about material discovery
um for carbon capture. Um And so we focus on uh uh um a sort of a family of molecules called uh metal–organic frameworks. Um These consist of uh sort of, you can think of them as a sort of a lattice structure with, on the nodes, you typically put some kind of uh you know, you know, small little molecule piece with a metal in it. Um And then they get connected by organic linkers is what they called. Um And so there's a huge design space, like many trillions of possible things that you can sort of even, you know, imagine uh with very different properties. And these um these frameworks uh if you blow, let's say air through them,
they tend to bind the molecules in the air uh to the to the framework. So for instance, carbon dioxide binds but hydrogen uh water, um sorry, I say water also binds nitrogen, binds. So all of these molecules bind to them and you want to find the uh metal organic framework that preferentially binds carbon dioxide um relative to water and nitrogen and all these other molecules. And then you want to do this at a particular temperature and pressure. Um And then you change the temperature and pressure. Um let's say you, you have less pressure and higher temperature that would sort of release then the carbon again and then you can sort
of sequester it and sort of put it away or reuse it for something. Um So, so that's the basic idea. And now we are going to use machine learning to accelerate the process a lot. All right. So we're gonna, we're gonna use the same generative models that are being used to generate images to generate new materials. Um And then we're gonna test those materials in a chemistry pipeline, but those, those are also um infused with machine learning methods to speed up this, this testing framework of how good this particular proposed material is. And then we have a search agent that tries to, you know, orchestrate this, this, this,
you know, searching through the space of molecules to find the best molecule with certain properties as quickly as possible. Um And you can, so how does this connect to fluids? Um Well, you a gas is not quite a fluid. Um but you can, you know, you could imagine actually a fluid in a, in a mo that's actually what happens. So you can sort of uh you can also put fluids in MOFs. Um But let's imagine a gas where you can sort of describe a gas but somewhat similar equations as a fluid. Um And, you know, there's sort of a, a, you know, an equation that describes how this gas would blow through this particular um framework and then, you know, and how it binds the molecules to the framework.
So that, that is not very unlike how you would uh simulate a fluid maybe um in, in some, you know, in some space with certain boundaries or something. And, and how do people do it now, you know, because there are companies trying to do these kind of, so are they relying on more traditional physics based simulations to do these discoveries? Yeah. So um there is already people who are um having these kind of pipelines. So uh they are mostly chemists. So it's, it's a chemistry pipeline or workflow um where um you know, there's a database of um you know, known um metal–organic frameworks and you push them through a chem chemistry
simulation pipeline that sort of simulates all the molecules as they wiggle around in this particular uh environment. And then you can sort of measure how many of them would would bind at certain temperatures and pressures. Um And that's, and that turns into a number that says how good this thing is. And then, you know, you have some, some sim simple way of searching through the space. Now, you can see that that can be massively improved with machine learning techniques. That's why we think it's, it's such an exciting field because they are not generating new materials with diffusion models. Um They don't use very sophisticated search algorithms either.
Um in order to um you know, to do, only do the simulations for the ones that are important. Um And you can even accelerate the simulation models by, you know, predicting the properties directly using machine learning models. Um or by what's called training force fields, these are methods that compute the forces on atoms as they wiggle. Typically, you have to do this with quantum mechanics, it's very expensive. Um but you could do it with a machine learned force field which is again way, way faster, many orders of magnitude faster than a normal uh force field. Uh if you use quantum mechanical calculations. Um but almost as accurate as that quantum mechanical calculation.
OK. Interesting. And what made you target this area? Is it more as the climate change sustainability? And is that a sort of a personal thing of yours that you've wanted to do? Cos I know you've had a focus in many areas of science. You know, I'm just wondering what made you focus on this area in particular. Yeah, it's definitely uh a personal uh sort of preference for me and my co-founder Chad Edwards a me. But also we think there is a big market for this because I think, you know, companies do want to be a good citizen and, you know, compensate their carbon emissions. And at some point, it might actually be enforced by government, which is what I hope,
you know, we need to tax pollution. Um And then having um a solution available to actually capture that carbon is, is there going to be a big market, I think as well? So it's, it's it's an alignment again between something you really want to do and work on with a commercial angle because otherwise you cannot build that company, that a healthy company that would actually make that impact. So that's the way I view that. Um But having said that, you know, there is many other directions in which this platform could be pointed even for metal–organic frameworks, um you can store hydrogen, you can catalyze, you can um um you can sort of deliver drugs, uh you can detect toxins, uh you can
clean water and that's only metal–organic frameworks, right? But then we can also go to other markets at some point um that are completely different materials. And how does it feel? I think you alluded to um what many people who work for large companies feel, which is the benefits of working for large companies, the, you know, access to a great talent pool, but maybe the slight slowing down and sort of meetings to decide other meetings. Have you felt a certain sort of energy of being uh back in a, in a start up again. Definitely, I've definitely felt a lot of energy being back in a start up. It is, uh I think it is no comparison
in terms of, uh you know, the, I don't know the way in which, you know, you can motivate people to all work on this problem. And um and it's, it's, it's somewhat small group at this point. Of course, when these companies grow much bigger, there's also, of course, you know, how do you maintain that spirit is not so easy, but, you know, as a start up, that's very special because you're small, um, you want to change the world, you have that vision to do that and you have, you know, you can do it quick and fast with a lot of money actually under your wings. So it's a big responsibility on the one hand, on the other hand,
it's also an incredible opportunity. And, yeah, and, and really a little bit of a different feel than working in a big company. Yeah, I wanted to pivot a little bit to get your advice really. And some of your, uh uh, yeah, maybe your advice, I guess to, to people who are early on in their careers, I think you've had a very interesting trajectory and, and, and kind of movement if you, what would you recommend? So if you're, you're going back, you're doing your undergraduate degree. Let's say you've already decided to do a PhD. Would you recommend go do a postdoc, try and go the academic route? Would you suggest to go to a start up?
How do you navigate this for people who wanted to sort of follow in your footsteps? What was, is there anything that you would do again or do differently? So, I should say I did many, many postdocs, right. So, um I did, uh I, I graduated in 98 and then I did a couple of years at Caltech and then I did and another three years with Geoffrey Hinton and do two different places and I just loved it. So, so the first thing I want to say, it depends a lot on who you are. Um And also, um so that really suited me at that stage in my life. It was incredibly interested in fundamental science and um I didn't get paid a lot at all.
Um But I didn't care at all, you know, it's just, it was a beautiful life, you know, uh having very little money to spend and just being, you know, being a scientist. Um So these days, but, but, but I didn't have that choice, right? So that was actually different and interesting. So these days it's much harder for young people because they see, you know, this other opportunity, clear clearly in front of them, right? And they say I could, I do that, you know, and earn like 10 times as much as I can earn there. Is it worth, is it worth for me to just go to this do this academic route? And you know, honestly academics
working in academia is also not all positive, right? At some point, you, you know, you're asked to juggle many balls, like you have to write grants, you have to teach, you have to basically, you know, run your own business. You have to be good people, a good manager. You know, you have to be good at research. You have to be good at everything basically. And then you, you don't have job security until you get tenure. It's really in that sense, a pretty lousy deal if you think about it. So you have, you have to really, you know, want, you want, you, you should be an educator. So the good thing about academia is it's great to work with young people,
right? It's fantastic to help people grow and I still sometimes get emails from people. I was in your master class then and then, and it's because of you that I took this route and I now have this fantastic job, right? That is, you know, unbelievably rewarding to get that. Um And that's what you have in academia, you can help people grow, you can turn them into great researchers and you can see them have stellar careers, which is very rewarding. I think so. So, so, you know, is that if that's you, you want to work on the inside of a black hole, you know what's going on on the inside of a black hole. And you like to work with young people
and you can put up with some annoying nuisances in, in academia. Then academia is your, is your thing, I honestly believe and you should, and especially, um, you should not worry about money. It's easy for me to say now. Um, but I didn't do it then I didn't worry about it. But it was the best thing I decided because I could really focus on, you know, building a, um, a good foundation for science learning. A lot about a lot about different topics and, and, and having a sort of a research strategy for myself. So I think having a postdoc on the Ariel Rings is, is really fantastic. And I don't think if you're good postdoc
and you publish well, you know, it, it's not a disadvantage either. I would say at, at all, but you can also do something else. You can also first work for a start up or work first work for, you know, for company and then go back to academia. You should, you should feel free to, you know, you're not closing any doors. I feel, you know, going back and forth between East East may maybe make sure you publish it now and then something exciting if you want to go back to academia. But do you think there should be more that the complaint I hear, uh, or the issue is that post doc, the funding model, the contract model is such that, as you said, for a certain time period,
you can sort of deal with it and I know from my own experiences but there comes a point where you, you're sort of, you want to buy a house or you want, you want to get married or something? There comes a point. Do you think universities should do more? Oh. Is there any way that you can do more to sort of keep those researchers who are great? But give them more security or, or does that break just the model of how universities work? Yeah. I don't know. It's a very hard one. Yeah. So, in Europe, you know, especially in the Netherlands tenure isn't, you know, it's very likely you'll get tenure, right. There's even something that was only recently introduced in some sense, but I've,
I've never seen people not get tenure. So it's, it's not at all like MIT, or Harvard where I know half of the people get kicked out or something. But then if you're kicked out of Harvard, you still have a great career afterwards, right? Because you work at Harvard in the first place. So, yeah, I've never really worried about tenure and all these things. I went through the 10 year process in the U.S. I just—maybe it's just blissful ignorance. I just never worried about it. I just, but whatever happened, it happened and just take the next step. Um, I don't know whether we can give postdoc permanent positions because that's not the definition of a postdoc. So,
what we have done at the University of Amsterdam a little bit is to allow people more like flexible contracts. Right. So, for instance, you could say half of my time I'm opposed or half of my time, you know, that particular construction doesn't exist, but it's a half time. You work in a start up, half your time, you work at the university or something like that. I think that could be very helpful because half of the time you can make a lot of money, you know, working in a start up or a big tech company or whatever or run your own business. And then the other half of the time you teach or, you know, do research at the university
that's already incredibly helpful and it doesn't cost university anything. So I was going to say that model, is it by chance or is it something? So if you look at an engineering company, look at Boeing, look at Airbus, look at Rolls Royce, look at Shell, look at almost any and I don't think there are very many people there unless I'm wrong who are, let's say their head of their science, but also a professor somewhere else. They seem to be purely at that company. Whereas I've noticed in tech companies and I guess your previous role of Microsoft was an example of that. They almost always is a, you know, extremely distinguished scientists like yourself who also has
a link to university because it enables them, by definition, they, that's how they still become distinguished because, you know, they, they're doing their own research. Do you think this is a model that should be taken more broadly or is it just something unique to the tech world with AI? If that's the case, I've just noticed this is a, a trend that, yeah, we do have some professorships in the Netherlands that are one day a week at the university um that are subsidized by a company. So I think it, it does exist a little bit more broadly, of course, in the tech world, it gives you also access to the talent pool, which is, which is a very scarce resource
in AI which may not be the same thing in other companies. So that, that's one way in which it's already helpful to have some leg in the university. Um Yeah, and the other thing is that this field is developing so fast, you just need to be at the front, you know, you know, frontier of all of that to know what's going on. And the best way to do that is to have, you know, to have students and you know, let them explain to you what the new trend is, right? So it's, it's really important that you're, you don't get, you know, sort of a, how do I say, um, fix it into your own, um, sort of ideology. Um But you keep listening to
young researchers which pick, pick up the new trends and want to work on new things and, and you, you go with that flow in some sense. So I think there's many advantages of having a foot in academia also for the company where you work for. Uh In fact, yeah, well, maybe um coming, coming to a final question, if in term, I think I heard you say in um one of the recent um talks that you gave that you had underestimated where machine learning could get to. So with that in mind, if we were to fast forward five years, where do you think we will be at and particularly from the AI for science side, do you think there will be a breakthrough,
huge change almost like, uh, OpenAI did with ChatGPT, which sort of sent this inflection point? Where do you see as being uh in the next few years? That's a tough one. Predicting the future is always extremely tough. Um I do see a very steady improvement and uh this field will grow definitely a lot over the next five years. So I see, but that's also because, you know, students need to get interested in it, you know, new start ups need to start, you know, big tech companies needs to get interested in this, which is what's happening now. Um And so that field will steadily grow, um whether there will be sort of a watershed moment.
Um Now, the reason that happened in, you know, in the other field is because there's a lot of data available there. Um So I think in some sense, the weather forecasting is a good example of a somewhat of a watershed moment. Um where, you know, these models are un you know, much better than people would have predicted. Um Now I'm betting myself on materials clearly. Um And I think my prediction would be that there's going to be enormous pool or pressure if you wish from society to work on, you know, problems and sustainability because we are running into a wall in this whole climate change problem, we are still increasing our
carbon output. Um you know, um since over the last 12% more, over the last five years, something like this. Um And there's a lot of countries which still have to go through sort of uh industrial revolution almost. So I I'm not seeing that politics will solve this. Um And the problem with this is that it's a shared resource and um it, it's, you know, there might be a dynamics where by the time it's in your face, it's too late because you're past tipping points and all these kinds of things and humans have the annoying, proper, you know, um as a property of only responding to things when it's right in their face, right?
You know, when the danger is right there in front of you and the lion is there to eat you, you know, that's when you run away. And so it's very hard to predict and say, oh, this is gonna be tricky, right? And so some people do that and some governments even do that. But two is that a global scale is incredibly hard. Now, what we will see is the impact of climate change, increasing and increasing. And that would mean that um this awareness also it's more in your face. So it become, you know, the awareness starts to increase as well. And that means that people are really going to look for solutions in the tech business as well.
So, you know, and there will be more and more investments and more and more people will go into that to try to do fusion or carbon capture or, you know, uh clean energy, all these kinds of things, right? And that will accelerate and that's a good thing. Um So in that sense, I think we will find amazing new technologies and you know, where indeed you you could actually find completely weird new materials that you know, weren't imaginable some time ago. But now with these new tools, we can actually create them. Um So I think we'll see amazing things where there's going to be an absolutely watershed moment that's really hard to predict.
Um But um yeah, I, I do think that with the new technology that we're building in AI um and maybe quantum computing as well uh in, you know, 5 to 10 years, this field will accelerate and, and, and, and produce very exciting results. Yeah. Do, do you feel almost that um one person told me we've globally tech companies and start ups have hired so many people to work on large language models. Do we almost need to wait for that to plateau for all those staff to then be given another direction to work on and then they can double down more on the scientific problems where for now, I guess most companies are trying to chase the current goal, which is,
you know, the the large language model type thing. Do, do you feel there's also a little bit of a resourcing and pivotal moment that we need to shift to focus on the science or can they be be done in parallel? Um I think they have to be done in parallel because I have no illusions that um you know, AGI will be pursued hotly by all these big companies because it's just, I don't know, you can make too much money if you find that, you know, this is too uh um you know, you can, how do you say too lucrative to, to not do it. Um But I think there is also a lot of talent in the market and I have found that a lot of young people in particular,
they want to do something that is meaningful with their life. And it's that subset of people that can align their goals in life, you know, contributing meaningfully to society with their talent and make a good, you know, salary on the side. And so I think there will be an increasing group of people who want to pursue these kinds of goals rather than work on advertisement placement or something like that. Well, yeah, I, I agree. I, I hope so. I think uh it's amazing what machine learning has been able to achieve. But I agree that if we could um push all that knowledge and that amazing thing to things that will help the world and society and the average individual,
then I think that will be seen as a very positive outcome from all this rather than, as you said, things that are maybe still useful, but maybe less an impact on society and on climate change. So, thank you so much. I know you're a very busy man. So I appreciate you taking the time to speak. I certainly learned a lot and I found it very interesting. So, yeah, thank you so much. My pleasure, Neil. It was, it was really fun talking to you
that