1
00:00:00,240 --> 00:00:02,360
Hi, and welcome to the Neil
Ashton Podcast.

2
00:00:03,000 --> 00:00:05,880
In each episode, we explained
some of the fascinating ways

3
00:00:05,880 --> 00:00:09,080
that science and engineering are
changing the world around us.

4
00:00:09,760 --> 00:00:12,720
We talked to leading engineers
from elite level sports like

5
00:00:12,800 --> 00:00:16,840
cycling and Formula One to some
of the world's top academics to

6
00:00:16,840 --> 00:00:20,960
understand how fluid dynamics,
machine learning, supercomputing

7
00:00:21,360 --> 00:00:22,960
are bringing in a new era of
discovery.

8
00:00:23,920 --> 00:00:27,040
We also hear some of their life
stories, their career advice,

9
00:00:27,640 --> 00:00:30,040
the lessons they've learned on
the way that I hope will be

10
00:00:30,040 --> 00:00:33,800
helpful to you too.
So sit back and enjoy this

11
00:00:33,800 --> 00:00:40,840
episode.
Hi and welcome back to the Neil

12
00:00:40,840 --> 00:00:43,800
Ashton podcast.
So today's guest is Professor

13
00:00:43,800 --> 00:00:47,000
Neil Torre.
He's an associate professor at

14
00:00:47,000 --> 00:00:51,000
the Technical University of
Munich TUM and has been a

15
00:00:51,000 --> 00:00:53,920
pioneer in physics based deep
learning.

16
00:00:54,240 --> 00:01:00,160
We had a really interesting
discussion today on first of

17
00:01:00,160 --> 00:01:03,200
all, some of his background,
Very interesting that he

18
00:01:03,600 --> 00:01:08,720
actually won an, an Oscar
working in the visual effects

19
00:01:08,720 --> 00:01:14,560
industry for, for movies after
his postdoc and PhD before he

20
00:01:14,560 --> 00:01:20,160
then turned to academia full
time, which is kind of

21
00:01:20,360 --> 00:01:23,880
incredible actually, and is
interesting because I've noticed

22
00:01:24,640 --> 00:01:28,080
the lessons learned from the
visual effects industry in terms

23
00:01:28,080 --> 00:01:31,720
of photorealistic
representations of, of, of the

24
00:01:31,720 --> 00:01:36,920
world is actually something we,
we then talked right the end of

25
00:01:36,920 --> 00:01:39,160
the, the podcast.
So make sure you listen to the

26
00:01:39,160 --> 00:01:43,920
end about the link between
foundation models and surrogates

27
00:01:43,920 --> 00:01:46,400
to world models.
These world models that are

28
00:01:46,800 --> 00:01:51,080
being created essentially as a
synthetic environment to train

29
00:01:51,160 --> 00:01:54,360
autonomous vehicles and robots.
And we actually talked a lot

30
00:01:54,360 --> 00:01:57,240
about whether there needs to be
more physics in these world

31
00:01:57,240 --> 00:01:59,560
models, which I actually thought
was a very interesting

32
00:01:59,560 --> 00:02:01,360
discussion.
And we only got to it at the

33
00:02:01,360 --> 00:02:05,320
end.
But you know, he he's been one

34
00:02:05,320 --> 00:02:11,440
of these people who was doing
deep learning in the mid twenty

35
00:02:11,440 --> 00:02:15,920
10's and early twenty 20s before
ChatGPT and the sort of big

36
00:02:15,920 --> 00:02:18,240
rise.
And it was interesting to him

37
00:02:18,240 --> 00:02:22,600
talk about how his work was
looked in those times compared

38
00:02:22,600 --> 00:02:25,440
to now and some of the barriers,
you know, and, and skepticism

39
00:02:25,440 --> 00:02:28,840
maybe people had.
He's also been someone who's

40
00:02:29,200 --> 00:02:34,200
really pioneered open source and
pushing the boundaries of

41
00:02:34,200 --> 00:02:36,920
differentiable physics and some
of the codes that him and his

42
00:02:36,920 --> 00:02:41,920
team has developed and now have
moved into both generating data

43
00:02:41,920 --> 00:02:45,800
sets like the Super wing data
set, but also addressing the

44
00:02:45,800 --> 00:02:50,920
challenge of foundation models
also from a sort of PDE point of

45
00:02:50,920 --> 00:02:53,400
view in the latest Tadpole
paper.

46
00:02:53,880 --> 00:02:57,840
So we, we, we kind of go through
some of those topics around

47
00:02:58,560 --> 00:03:02,440
those different strategies of
sort of PDE based methods, you

48
00:03:02,440 --> 00:03:06,480
know, pre training on those
versus doing more of the

49
00:03:06,480 --> 00:03:08,760
traditional, I guess, building
out large data sets.

50
00:03:09,120 --> 00:03:12,240
We talk about the role of
startups and industry and

51
00:03:12,240 --> 00:03:14,200
academia.
Interesting to get his

52
00:03:14,200 --> 00:03:18,000
perspectives on that and how all
three can work together to solve

53
00:03:18,000 --> 00:03:20,560
some of these problems.
And we also talk about the open

54
00:03:20,560 --> 00:03:23,680
source, why he's so pro open
source and how we think that can

55
00:03:23,680 --> 00:03:26,640
help progress the, the, the
field along.

56
00:03:27,360 --> 00:03:33,040
He's, he's actually contributed
together with his team, many

57
00:03:33,120 --> 00:03:36,840
important papers covering not
just on fluids, but also on the,

58
00:03:37,000 --> 00:03:39,040
on the weather side, interesting
like weather bench.

59
00:03:39,320 --> 00:03:42,840
So I, I put a whole link list of
papers in the chat notes that

60
00:03:42,840 --> 00:03:45,520
you can in the show notes that
you can have a look at.

61
00:03:45,920 --> 00:03:47,760
But I, I really enjoyed this
conversation.

62
00:03:47,760 --> 00:03:50,880
He's someone who is very modest
in, in what he's done.

63
00:03:51,040 --> 00:03:53,400
But if you actually read his
papers and look through, you can

64
00:03:53,400 --> 00:03:57,120
see he's being quite influential
in steering these topics and

65
00:03:57,120 --> 00:04:01,600
always seems to be one step
ahead of where most of the field

66
00:04:01,600 --> 00:04:04,240
is AT.
And so if you look at his 2025

67
00:04:04,240 --> 00:04:08,600
and 2026 papers, you'll know
what I mean by that.

68
00:04:08,600 --> 00:04:12,920
So hope you enjoy this episode
with Professor Neil Torry.

69
00:04:13,440 --> 00:04:16,480
Well, yeah, thanks for thanks
for joining today.

70
00:04:16,480 --> 00:04:21,079
I, I've been very keen to speak
to you because as I've been

71
00:04:21,079 --> 00:04:24,040
getting up to speed myself over
the past years on machine

72
00:04:24,040 --> 00:04:26,000
learning and differentiable
side.

73
00:04:26,000 --> 00:04:30,280
Your name is a, a constant, you
know, top of the list in terms

74
00:04:30,280 --> 00:04:34,560
of influential papers, in terms
of doing things what seemed to

75
00:04:34,560 --> 00:04:37,920
be earlier than most others have
done, you know, setting the

76
00:04:37,920 --> 00:04:41,880
scene, which suggests to me that
you have quite a good outlook

77
00:04:42,040 --> 00:04:46,160
and, and, and thought on it.
So, yeah, thanks very much for

78
00:04:46,160 --> 00:04:50,880
for for joining.
And yeah, maybe as a starting

79
00:04:50,880 --> 00:04:55,520
question, you took a diversion.
Well, not diversion, but you

80
00:04:55,520 --> 00:04:59,400
spent some time doing visual
effects after your PhD and

81
00:04:59,400 --> 00:05:02,160
postdoc maybe.
What did you do your PhD?

82
00:05:02,160 --> 00:05:04,920
What was the postdoc?
And then I'm I'm intrigued on on

83
00:05:04,920 --> 00:05:09,320
the movie stuff.
Thanks for the invite.

84
00:05:09,480 --> 00:05:11,880
First of all, Neil.
Yeah, I'd be happy to tell about

85
00:05:11,880 --> 00:05:15,520
this.
My background, right or over my

86
00:05:15,520 --> 00:05:18,040
career, I have switched fields a
couple of times.

87
00:05:18,040 --> 00:05:23,400
My PhD was in a in a
computational numerics

88
00:05:23,600 --> 00:05:28,480
computational physics lab,
basically targeting multi grid

89
00:05:28,520 --> 00:05:32,240
methods and and fluids.
And then I switched to computer

90
00:05:32,240 --> 00:05:34,360
graphics for a while.
I guess that that actually was

91
00:05:34,360 --> 00:05:37,080
one of my driving my motivations
from the start.

92
00:05:37,280 --> 00:05:39,520
I guess I'm a visual person.
I'd like to see how things

93
00:05:39,520 --> 00:05:41,720
evolve.
And I think that's that's still

94
00:05:41,720 --> 00:05:44,240
actually one of the main things
that's fascinate me about

95
00:05:44,240 --> 00:05:47,840
fluids, these swirling motions
and the things that actually do

96
00:05:47,840 --> 00:05:53,000
come out once it's working.
So we also worked on computer

97
00:05:53,000 --> 00:05:55,800
graphics, BASIC for computer
based animation for quite a

98
00:05:55,800 --> 00:05:58,000
while.
My post I get ETH and then

99
00:05:58,000 --> 00:06:01,200
afterwards also did actually
work in industry for a while

100
00:06:01,200 --> 00:06:04,840
because I was curious how things
would actually work out in

101
00:06:04,840 --> 00:06:06,760
industry in a practical
environment.

102
00:06:06,760 --> 00:06:08,560
And partially also because I
couldn't really get a good

103
00:06:09,560 --> 00:06:14,840
faculty position at the time.
So right, not not completely

104
00:06:14,840 --> 00:06:19,160
voluntary in a while in a way,
but right, it was good to see in

105
00:06:19,160 --> 00:06:21,400
the end.
I noticed for me, research is

106
00:06:21,400 --> 00:06:24,320
the more I think the better
field.

107
00:06:24,760 --> 00:06:28,080
But yes, that's that's why we
did my poster basically.

108
00:06:28,080 --> 00:06:32,760
And then I think that if you 10
years ago, we started looking at

109
00:06:33,120 --> 00:06:36,920
machine learning.
So previously we did in this ETH

110
00:06:36,920 --> 00:06:39,880
group what was called
data-driven approaches in

111
00:06:39,880 --> 00:06:42,520
computer graphics, essentially
also trying to work with data

112
00:06:42,520 --> 00:06:47,160
and right, somewhat aligned, but
we didn't have the tools back

113
00:06:47,160 --> 00:06:48,880
then.
So it was, it didn't really

114
00:06:48,880 --> 00:06:51,360
work, to be honest, but we
tried.

115
00:06:51,360 --> 00:06:54,840
And then when Alphago came
along, I thought, oh, this has

116
00:06:54,840 --> 00:06:57,040
to be something that works.
I didn't understand it at the

117
00:06:57,040 --> 00:06:59,440
time, but I thought this,
there's got to be something

118
00:06:59,440 --> 00:07:02,160
there.
And then, yeah, since 10 years,

119
00:07:02,160 --> 00:07:06,240
I think we've been working on,
on basically finding these

120
00:07:06,240 --> 00:07:08,360
techniques to the
reconciliations and food

121
00:07:08,360 --> 00:07:11,760
mechanics specifically.
And I think in a way, it's a

122
00:07:12,240 --> 00:07:14,280
good time at the world.
It's finally starting to work.

123
00:07:14,400 --> 00:07:19,720
But right, it took a while.
So you then joined TUM and

124
00:07:19,720 --> 00:07:24,520
you've been there ever since.
That was your your period.

125
00:07:24,520 --> 00:07:30,560
And so where did you?
What was it like?

126
00:07:30,760 --> 00:07:33,320
I'm trying to put it into the
picture because it's very hard

127
00:07:33,320 --> 00:07:38,480
to the people who were working
on these things before now, you

128
00:07:38,480 --> 00:07:41,920
know, before chat GB Teemo, what
was the field like?

129
00:07:41,920 --> 00:07:44,120
What was the community like?
What was the reception like when

130
00:07:44,120 --> 00:07:47,680
you were bringing up ideas of
sort of data-driven and machine

131
00:07:47,680 --> 00:07:49,480
learning?
Was there a lot of skepticism

132
00:07:49,480 --> 00:07:52,000
and push back?
Yes, actually there was.

133
00:07:52,480 --> 00:07:56,800
So I mean, initially, especially
in the early phases, all these

134
00:07:56,800 --> 00:08:00,280
AIO back then, it was called
deep learning techniques were

135
00:08:01,520 --> 00:08:03,280
focusing on, on learning
descriptors.

136
00:08:03,280 --> 00:08:05,360
I remember our very first papers
and graphics, they, they

137
00:08:05,360 --> 00:08:07,160
basically couldn't synthesize
anything.

138
00:08:07,160 --> 00:08:09,560
You learned some reduced
representation.

139
00:08:09,800 --> 00:08:12,840
They told you about what happens
in high dimensional data, but

140
00:08:12,840 --> 00:08:14,960
right, it, it couldn't really
generate anything.

141
00:08:14,960 --> 00:08:18,080
Also, if you, if you think about
alpha gold, the early success

142
00:08:18,080 --> 00:08:21,400
stories, it was basically 8 by 8
with, with three states.

143
00:08:21,400 --> 00:08:25,720
That's a tiny domain and tiny
state space.

144
00:08:26,360 --> 00:08:28,240
And back then that was
challenging.

145
00:08:28,680 --> 00:08:30,240
So in a way, it really couldn't
do much.

146
00:08:30,760 --> 00:08:34,400
And I, I remember quite well
talking to colleagues and

147
00:08:34,400 --> 00:08:36,080
telling them, look, we're
interested in this deep learning

148
00:08:36,080 --> 00:08:38,679
stuff.
We, we're doing simulations.

149
00:08:38,679 --> 00:08:41,280
And then some of my colleagues
literally laugh me in the face

150
00:08:41,280 --> 00:08:43,559
like, you're doing learning for
Pdes.

151
00:08:43,559 --> 00:08:45,760
What, what could be more
ridiculous?

152
00:08:46,520 --> 00:08:49,280
If you have a PDE and you can
solve it, why, why, why on earth

153
00:08:49,280 --> 00:08:54,800
would you do anything right
learning based And I think so

154
00:08:54,800 --> 00:08:57,480
there was a lot of scepticism
for years.

155
00:08:57,480 --> 00:09:00,600
I always had to defend in the
very first slides of my talks

156
00:09:01,160 --> 00:09:04,280
that this makes sense at all, at
least makes sense to consider

157
00:09:04,280 --> 00:09:07,880
this combination.
And luckily I interestingly, I

158
00:09:07,880 --> 00:09:11,400
think for the for the white
community, I think the picture

159
00:09:11,400 --> 00:09:15,280
really changed with chat GPTI
think in computer science and

160
00:09:15,280 --> 00:09:18,640
and computational fields.
People have been aware earlier,

161
00:09:18,640 --> 00:09:23,920
but since chat ChatGPT it's
really on everybody's plate and

162
00:09:23,920 --> 00:09:26,640
in every view but before.
But you but you were

163
00:09:26,640 --> 00:09:29,760
interestingly was looking
through some of the papers and

164
00:09:29,760 --> 00:09:33,400
you you had slightly earlier on
probably some of the first ones

165
00:09:33,400 --> 00:09:36,480
are sort of like in the AI AA,
the deep learning of like some

166
00:09:36,680 --> 00:09:43,720
random airfoils and so quite a
few in the 20 twenties before

167
00:09:44,120 --> 00:09:47,760
ChatGPT.
So what was it like sort of pre

168
00:09:47,760 --> 00:09:50,600
and post?
Was it was the issue partly the

169
00:09:50,600 --> 00:09:54,840
compute side?
Was it the the architectures?

170
00:09:54,840 --> 00:09:56,680
What?
What was holding back?

171
00:09:56,760 --> 00:09:59,280
Exactly the the architectures
were a fundamental issue.

172
00:09:59,680 --> 00:10:03,600
I think the whole infrastructure
right in the beginning, CNNS

173
00:10:03,600 --> 00:10:06,280
were were a big breakthrough,
yes, for your MLPS.

174
00:10:06,280 --> 00:10:10,640
So CNNS definitely for feel like
data worked quite nicely, but

175
00:10:10,640 --> 00:10:13,720
also right the, the, the shift
from basically computer graphics

176
00:10:13,720 --> 00:10:17,920
to things like aerodynamics.
The a paper you mentioned, we're

177
00:10:17,920 --> 00:10:20,880
a bit motivated by the
observation that it's, it's

178
00:10:20,880 --> 00:10:23,760
really tough to synthesize stuff
in 3D plus time.

179
00:10:23,760 --> 00:10:27,240
You essentially have 40 fields.
Even if you simplify it and and

180
00:10:27,240 --> 00:10:31,520
treat regular geometry as
regular grids, this is really

181
00:10:31,520 --> 00:10:34,680
difficult to do.
It's still challenging these

182
00:10:34,680 --> 00:10:36,320
days actually with, with neural
networks.

183
00:10:36,560 --> 00:10:39,880
So back then it was right bridge
out of the question to do this.

184
00:10:41,680 --> 00:10:43,600
And in graphics, you actually
have extremely good

185
00:10:43,600 --> 00:10:47,200
approximations that aim for the
visual aspect of, of things.

186
00:10:47,840 --> 00:10:51,280
So it, it was extremely tough to
compete with that.

187
00:10:51,280 --> 00:10:54,320
And it was interestingly in the
computer graphics field also, if

188
00:10:54,320 --> 00:10:57,920
you cannot directly apply it on
a scale and, and with a fidelity

189
00:10:57,920 --> 00:11:02,200
that's fit for movies, then your
papers typically get rejected.

190
00:11:02,760 --> 00:11:04,840
It's really, it's computer
graphics makes sense, but it

191
00:11:04,840 --> 00:11:09,080
needs to look good in the end.
So write this 2D proof of

192
00:11:09,080 --> 00:11:12,040
concept simulations.
Even if they were fairly

193
00:11:12,040 --> 00:11:13,400
accurate, you you couldn't do
much.

194
00:11:14,200 --> 00:11:16,760
It had had to be right movie
quality basically.

195
00:11:16,760 --> 00:11:19,000
And and that just was extremely
difficult.

196
00:11:19,040 --> 00:11:21,880
Only also for niche applications
like super resolution.

197
00:11:21,880 --> 00:11:24,080
This was was one that worked
very nicely.

198
00:11:24,080 --> 00:11:28,560
But they can look at repeating
structures and very local fields

199
00:11:29,480 --> 00:11:33,760
for synthesis on on larger
scales, it was basically not

200
00:11:33,760 --> 00:11:35,680
possible.
And that's why we also started

201
00:11:35,680 --> 00:11:39,560
actually doing these proof
concept papers like the the

202
00:11:39,560 --> 00:11:45,960
Iran's predictions by AIA back
then also these were tiny

203
00:11:45,960 --> 00:11:49,920
resolutions, right?
32128 square basically.

204
00:11:49,920 --> 00:11:54,560
So very small domains.
And also I remember back then

205
00:11:54,560 --> 00:11:58,040
actually it's I'm, I'm glad it's
got into a a, but there were

206
00:11:58,040 --> 00:11:59,760
also critical questions even
back then.

207
00:12:00,120 --> 00:12:01,760
Why would you even do this at
all?

208
00:12:01,760 --> 00:12:04,960
This is crude approximations.
Does this make sense?

209
00:12:07,840 --> 00:12:11,600
But yeah, basically we noticed
in in aerodynamics and in

210
00:12:11,600 --> 00:12:14,600
engineering applications,
there's much more demand for

211
00:12:14,760 --> 00:12:17,800
getting things right and really
going forward converged and

212
00:12:17,840 --> 00:12:21,200
solutions that are accurate and
can be evaluated.

213
00:12:21,960 --> 00:12:25,720
And there are many problems
there that that are like really

214
00:12:25,720 --> 00:12:29,440
important basically for for a
variety of industries.

215
00:12:29,440 --> 00:12:32,960
So we basically switched from
from graphics trying to apply

216
00:12:32,960 --> 00:12:34,680
these for engineering
applications.

217
00:12:37,440 --> 00:12:40,880
And where did the, you know, Phi
flow when there's a

218
00:12:40,880 --> 00:12:45,040
differentiable physics, where
did that start to come into the

219
00:12:45,040 --> 00:12:46,520
grid?
Because that seems to be now

220
00:12:46,520 --> 00:12:49,840
quite a hot topic.
But you were publishing in 2020

221
00:12:49,840 --> 00:12:52,760
on these, so where did that
first come from?

222
00:12:52,760 --> 00:12:53,840
I.
Don't even remember, it seemed

223
00:12:53,840 --> 00:12:57,520
like a very obvious thing.
We were working on computational

224
00:12:57,560 --> 00:13:02,080
and numerical methods, the the
networks couldn't really

225
00:13:02,080 --> 00:13:05,640
generate a whole simulation.
So why not start with the

226
00:13:05,640 --> 00:13:08,920
simulator, combining it with the
network and then if you want to

227
00:13:08,920 --> 00:13:12,480
get a gradient through it
actually need differentiability

228
00:13:12,480 --> 00:13:17,240
on on the solver side and right
the the available ones couldn't

229
00:13:17,240 --> 00:13:19,520
do this.
So we just started experimenting

230
00:13:19,520 --> 00:13:24,440
with on solvers since 5:00 flow
became a good testing ground for

231
00:13:24,440 --> 00:13:28,920
very basic methods across
different, different AP is back

232
00:13:28,920 --> 00:13:33,680
then, right there was 10s of
flow and, and Pytorch and JAX

233
00:13:33,680 --> 00:13:37,560
and the whole bandwidth of of
different methods.

234
00:13:38,720 --> 00:13:42,120
So I think it was good for
experimentation and flexibility.

235
00:13:42,640 --> 00:13:46,680
And ultimately we notice here
it's a certain trade off.

236
00:13:46,680 --> 00:13:49,600
Flexibility and generality don't
go too well together.

237
00:13:50,800 --> 00:13:53,480
And so now we're specialising a
bit more.

238
00:13:53,480 --> 00:13:56,280
But five floor was one of these
early testing grounds.

239
00:13:56,280 --> 00:14:00,400
But for us, it really seemed
like in a way obvious, obvious

240
00:14:00,400 --> 00:14:02,000
way to go.
And then later on, we notice,

241
00:14:02,000 --> 00:14:04,440
yeah, there's actually quite
some work to do and also

242
00:14:04,440 --> 00:14:06,400
interesting questions.
How do you get gradients through

243
00:14:07,120 --> 00:14:10,880
long chains of solver
operations, which is still a

244
00:14:10,880 --> 00:14:16,320
topic?
So maybe shifting a little bit

245
00:14:16,320 --> 00:14:19,640
towards the what's going on
right now.

246
00:14:19,880 --> 00:14:25,200
I mean, there's 22 fundamental
questions that I see you've also

247
00:14:25,200 --> 00:14:27,680
been trying to tackle and be
interested to get your thoughts

248
00:14:27,680 --> 00:14:29,560
on it.
And that is around, I guess the

249
00:14:29,560 --> 00:14:32,440
foundation model topic.
The idea of I guess it's

250
00:14:32,440 --> 00:14:34,840
logical, isn't it?
You look at ChatGPT and you see

251
00:14:34,840 --> 00:14:39,120
it as essentially a data-driven
problem that you know, applying

252
00:14:39,120 --> 00:14:42,240
up data, have a decent
architecture and you solve it.

253
00:14:44,120 --> 00:14:50,760
But you well, two questions.
One was on the idea of the

254
00:14:50,760 --> 00:14:53,760
ground truth and can you ever
surpass the ground truth?

255
00:14:53,800 --> 00:14:57,400
And you know, one of your recent
papers looked at can the model

256
00:14:57,400 --> 00:15:00,520
do better than essentially the
quality of the training data?

257
00:15:01,080 --> 00:15:05,800
And the second one is more just
a general question on where you

258
00:15:05,800 --> 00:15:08,320
see the likelihood of foundation
models going.

259
00:15:08,320 --> 00:15:11,000
How, how realistic is it?
But on the first one, maybe you

260
00:15:11,000 --> 00:15:13,560
could explain that paper because
I that really intrigued me

261
00:15:13,560 --> 00:15:15,920
actually, right?
There was actually a great

262
00:15:16,120 --> 00:15:18,920
observation by Phoenix, one of
one of the PhD students from my

263
00:15:18,920 --> 00:15:22,280
group, that he noticed in some
situations.

264
00:15:22,280 --> 00:15:25,360
He also looked at basic
combinations of solvers and

265
00:15:25,360 --> 00:15:29,560
networks and that's the networks
really seem to give extremely

266
00:15:29,560 --> 00:15:33,240
good situations in certain
situations extremely good

267
00:15:33,960 --> 00:15:37,680
results that seem to be even
better than the the training

268
00:15:37,680 --> 00:15:39,640
data we essentially used for
these cases.

269
00:15:40,200 --> 00:15:46,080
And looking into it a bit more,
it's unfortunately not a topic

270
00:15:46,080 --> 00:15:49,920
that I think can be generally
applied and when analysing it in

271
00:15:49,920 --> 00:15:53,160
terms of numerical errors and,
and how these pieces fit

272
00:15:53,160 --> 00:15:55,720
together.
It's basically a combination of

273
00:15:55,720 --> 00:15:59,520
the priors imposed by neural
networks in terms of smoothness,

274
00:15:59,520 --> 00:16:04,080
in terms of right, being able to
over sharp more reproduce and be

275
00:16:04,080 --> 00:16:07,760
trained for producing certain
parts of the solutions that

276
00:16:07,760 --> 00:16:13,280
benefits the the outcome.
So in a way, if you can leverage

277
00:16:13,280 --> 00:16:16,560
or if you understand your
solutions enough to leverage

278
00:16:16,680 --> 00:16:21,440
these particular behaviours of
neural networks, then basically

279
00:16:21,440 --> 00:16:25,200
the outputs can be better than
what you're trained with.

280
00:16:25,880 --> 00:16:29,600
But it's really this interplay
of, of what neural networks do

281
00:16:29,680 --> 00:16:32,640
as essentially also numerical
methods to compute a solution

282
00:16:33,160 --> 00:16:36,600
and what, what's in your data
and what you like to get out.

283
00:16:37,120 --> 00:16:40,760
So it also it raises some
interesting questions on how to

284
00:16:40,760 --> 00:16:44,040
evaluate it because if
especially in the numerical

285
00:16:44,040 --> 00:16:46,760
field, we're trusting the ground
truth, there are also typically

286
00:16:46,760 --> 00:16:51,800
approximation errors in there.
And it it can actually happen

287
00:16:51,800 --> 00:16:55,920
that solutions output by the
network might be even better if

288
00:16:55,920 --> 00:17:00,200
you consider other solves or or
other ways to approximate the

289
00:17:00,200 --> 00:17:05,119
same problems.
But I think it's we also for

290
00:17:05,119 --> 00:17:08,960
this paper thought about it.
It is very difficult to really

291
00:17:08,960 --> 00:17:12,040
do this in a target to way.
So we had some new cases where

292
00:17:12,040 --> 00:17:14,560
it worked for some, some basic
infection fusion and burgers.

293
00:17:14,760 --> 00:17:17,560
Burgers equation was a good
example actually because the the

294
00:17:17,560 --> 00:17:22,960
shock representation and
smoothness bogus is non bogus

295
00:17:22,960 --> 00:17:25,640
equation.
It's non trivial but seem to be

296
00:17:25,640 --> 00:17:30,200
a very good showcase for this.
Directly applying it to other

297
00:17:30,200 --> 00:17:34,000
scenarios is is challenging so
we haven't unfortunately we

298
00:17:34,000 --> 00:17:37,200
haven't gotten around to finding
a good way to to leverage this

299
00:17:37,200 --> 00:17:39,480
for it.
Reword never Stokes problems or

300
00:17:39,480 --> 00:17:46,920
so but but not to say.
But what about foundation mods

301
00:17:46,920 --> 00:17:48,480
in general?
Like, well, how was your group

302
00:17:48,480 --> 00:17:49,360
tackling this?
What?

303
00:17:49,360 --> 00:17:53,560
What's your strategy in a sense,
you know where, where, where do

304
00:17:53,560 --> 00:17:57,640
you see this going?
Recently we switched pretty much

305
00:17:57,640 --> 00:18:00,160
from these differential solvers
to foundation models.

306
00:18:00,760 --> 00:18:03,840
I was quite skeptical for a TI
for for quite a while.

307
00:18:04,120 --> 00:18:07,640
I think it's also there's also
just quite some skepticism

308
00:18:07,640 --> 00:18:09,840
talking to other people in the
field with foundation models.

309
00:18:10,280 --> 00:18:11,840
I think partly it's because of
the name, right?

310
00:18:11,840 --> 00:18:16,440
Foundation model sounds super
general and I'm sure you and and

311
00:18:16,840 --> 00:18:20,560
most people know these fields of
Pdes are extremely diverse.

312
00:18:21,320 --> 00:18:23,600
The thought of having one model
that just is everything

313
00:18:23,600 --> 00:18:31,960
out-of-the-box is yeah, but is
is a bit unbelievable and I

314
00:18:31,960 --> 00:18:34,680
think also still out of quite a
bit out of reach.

315
00:18:35,440 --> 00:18:39,160
Nonetheless, if you just right
foundation was also by now I

316
00:18:39,160 --> 00:18:41,200
think it's actually not so well
defined.

317
00:18:41,320 --> 00:18:44,440
People use it in in all various
forms and where I think it makes

318
00:18:44,440 --> 00:18:48,280
a lot of sense and in a way it's
also the direction we are

319
00:18:48,280 --> 00:18:52,280
typically tech typically
tackling is to see it as a model

320
00:18:52,280 --> 00:18:55,520
that's pre trained on on
something and then applicable to

321
00:18:55,520 --> 00:19:01,600
other tasks, other problems.
So there's just some disparity

322
00:19:01,600 --> 00:19:05,160
between training data and the
actual application we saw.

323
00:19:05,160 --> 00:19:07,440
We want models to generalise.
Also, generalisation is not a

324
00:19:07,440 --> 00:19:11,760
well defined topic, right?
It's in distribution is clear,

325
00:19:11,760 --> 00:19:14,240
but how far out of distribution
you need to go to have proper

326
00:19:14,240 --> 00:19:16,560
generalization.
It's completely open.

327
00:19:16,960 --> 00:19:19,360
But in the past that's been
really tricky.

328
00:19:19,600 --> 00:19:21,720
Anyway, it's a classical
transfer learning problem.

329
00:19:22,080 --> 00:19:24,600
Now, I think with these
foundation model tools borrowing

330
00:19:24,600 --> 00:19:27,680
a lot from large vendors models,
we can actually do this if we

331
00:19:27,760 --> 00:19:31,520
pre train on on one part and
then apply it to something else.

332
00:19:31,760 --> 00:19:38,200
And how else how how different
is is up for very specific to

333
00:19:38,200 --> 00:19:40,920
problems and and application
domains.

334
00:19:41,320 --> 00:19:43,680
But that seems to work.
And that's specifically what

335
00:19:43,680 --> 00:19:48,160
what I'm quite excited about is
that we notice you can actually

336
00:19:48,160 --> 00:19:51,320
pre train on very cheaply
generated data, offer us

337
00:19:52,200 --> 00:19:54,960
training foundation models in in
3D plus time.

338
00:19:55,080 --> 00:19:59,760
The amount of data needed seems
very scary and also somewhat

339
00:19:59,760 --> 00:20:03,160
impractical given given the
current cost of of doing these

340
00:20:03,160 --> 00:20:07,440
training runs.
But we noticed that in a way you

341
00:20:07,440 --> 00:20:09,960
don't really need these huge
amounts of data.

342
00:20:09,960 --> 00:20:11,840
And I think that that really
changes the picture, that

343
00:20:11,840 --> 00:20:14,400
changes how we can approach
foundation model training.

344
00:20:15,240 --> 00:20:19,120
And I think in a way it makes it
much more attractive as as a

345
00:20:19,120 --> 00:20:21,600
starting point.
Could you maybe go into a little

346
00:20:21,600 --> 00:20:25,120
bit more detail when you when
you talk about not needing as as

347
00:20:25,120 --> 00:20:27,960
much data or all the OR high
quality data?

348
00:20:28,680 --> 00:20:31,840
Right.
So you also thought about this

349
00:20:31,840 --> 00:20:35,000
quite a bit, right, as you
recent paper discussing the the

350
00:20:35,000 --> 00:20:39,680
scaling of of of data and and
computer requirements for fluids

351
00:20:40,240 --> 00:20:43,160
is is a good starting point to
think about it.

352
00:20:43,760 --> 00:20:48,280
What we noticed in this tadpole
work, we we called it typically

353
00:20:48,520 --> 00:20:54,000
was basically that if you have a
good solver.

354
00:20:54,000 --> 00:20:59,440
So it is actually also based in
a way on on previous work on on

355
00:20:59,440 --> 00:21:03,320
this eight bench benchmark with
very synthetic data.

356
00:21:03,320 --> 00:21:07,200
So if you and if you take the
right solver, these spectral

357
00:21:07,200 --> 00:21:12,080
solvers, then you can extremely
efficiently compute super good

358
00:21:12,280 --> 00:21:15,960
solutions for right, a class of
basic PD, ES affection fusion,

359
00:21:15,960 --> 00:21:20,880
these ETDRK solvers in a way,
they're really perfect.

360
00:21:20,880 --> 00:21:25,120
So also doing the learning on
this on this level does not

361
00:21:25,120 --> 00:21:27,280
really make sense because the
solvers are so good, the errors

362
00:21:27,280 --> 00:21:31,880
are so low, you can prove that
they are basically as accurate

363
00:21:31,880 --> 00:21:35,120
as it gets for for these basic
solutions.

364
00:21:35,120 --> 00:21:40,040
So we can basically generate
this data extremely quickly and

365
00:21:40,040 --> 00:21:43,680
very broadly if you bury the
parameters and we notice that

366
00:21:43,680 --> 00:21:47,440
right, you can then pre train a
model with this just generating

367
00:21:47,440 --> 00:21:53,320
a lot of different solutions for
for varying parameter ranges for

368
00:21:53,320 --> 00:21:57,480
these different PD ES from
affection diffusion to transport

369
00:21:58,240 --> 00:22:02,320
to some some higher order terms
with chaos promoters, Yushinski

370
00:22:02,320 --> 00:22:05,120
and some some polynomial
chemical reaction like

371
00:22:05,120 --> 00:22:09,360
constraints.
And basically pre training this

372
00:22:09,360 --> 00:22:15,480
was on this very broad class of
very basic synthetic solutions.

373
00:22:15,960 --> 00:22:19,720
Nonetheless, it's actually not
not a very easy task and

374
00:22:19,720 --> 00:22:23,920
benefits or seems to benefit in
in quite a bit of downstream

375
00:22:23,920 --> 00:22:26,400
tasks.
And the nice thing then is

376
00:22:26,400 --> 00:22:28,440
right, these, these solutions to
these synthetic prones are

377
00:22:28,440 --> 00:22:30,760
really not interested in or
interesting in a way you can

378
00:22:30,760 --> 00:22:33,880
compute them almost as quickly
with the solver from from

379
00:22:33,880 --> 00:22:36,160
scratch basically.
So you can train with it, you

380
00:22:36,160 --> 00:22:38,840
can run it through a model to
get a gradient basically, but

381
00:22:38,840 --> 00:22:41,200
then you can throw it away.
So it's basically just running

382
00:22:41,200 --> 00:22:44,360
the solver alongside you
typically have written and all

383
00:22:44,360 --> 00:22:46,360
the four GPU.
So one GPU does the data

384
00:22:46,360 --> 00:22:49,520
generation, the other 3 train
and just generates data on the

385
00:22:49,520 --> 00:22:52,920
fly, throws it away right after
it produces more data.

386
00:22:53,120 --> 00:22:56,680
But the nice thing is it's
produced on the fly and it's

387
00:22:56,680 --> 00:22:58,960
really just a continuous stream
of data to train with.

388
00:22:59,880 --> 00:23:03,800
You don't, you don't have
limitations on the bandwidth

389
00:23:03,800 --> 00:23:08,560
side on the hardware disk space
to manage 10s of terabytes of or

390
00:23:08,560 --> 00:23:12,640
hundreds of terabytes to upload
to HPC centres for for anyone

391
00:23:12,640 --> 00:23:15,200
who's tried this, it's, it's
very painful actually to do this

392
00:23:15,400 --> 00:23:20,080
and it's not trivial.
So right, having this online

393
00:23:20,080 --> 00:23:24,360
data generation with the
training is is really quite neat

394
00:23:24,360 --> 00:23:27,000
and worked quite well in our
experiments.

395
00:23:27,520 --> 00:23:31,480
But where?
Do you, how do you see that

396
00:23:31,920 --> 00:23:33,720
progressing?
I mean, I think this has often

397
00:23:33,720 --> 00:23:39,880
been my challenges and you've
been on both sides of it.

398
00:23:39,880 --> 00:23:44,960
You know, one side of it is,
let's take realistic in well, in

399
00:23:45,280 --> 00:23:47,840
if we assume that the model
should work for industrial scale

400
00:23:47,840 --> 00:23:53,160
problems, then you know, the the
idea is you generate data like

401
00:23:53,160 --> 00:23:56,280
you did with the Super wing or,
or now this highlight aero ML or

402
00:23:56,280 --> 00:23:59,080
any or company, you know,
generating data.

403
00:23:59,080 --> 00:24:03,840
And then you you train on that.
And you know that is one

404
00:24:03,840 --> 00:24:05,280
argument.
And the other one is you're

405
00:24:05,280 --> 00:24:08,320
starting from the far left side
of basic PD ES.

406
00:24:09,040 --> 00:24:13,720
But where do you see how can
they go in the middle?

407
00:24:13,720 --> 00:24:17,400
The basic PDS are just can you
go to Navier Stokes?

408
00:24:17,400 --> 00:24:19,400
How does that scale to more
complexity?

409
00:24:19,920 --> 00:24:22,160
Yeah.
So what we basically notice is

410
00:24:22,160 --> 00:24:25,240
that even if we pre trained also
one of the one of the tricks

411
00:24:25,240 --> 00:24:29,160
basically that I didn't mention
just now was also that we don't

412
00:24:29,160 --> 00:24:31,880
train for a time evolution.
It's really more of a of a

413
00:24:31,880 --> 00:24:36,320
latent space representation that
is learned for for single

414
00:24:36,320 --> 00:24:39,640
spatial fields.
So there's no time in the pre

415
00:24:39,640 --> 00:24:45,480
training and then the the basic
architecture is basically so so

416
00:24:45,480 --> 00:24:47,960
general that you can just put
time on top of that, right.

417
00:24:47,960 --> 00:24:50,360
So we learn in this latent
space, one step to the next

418
00:24:51,040 --> 00:24:55,760
seems to work nicely.
And also time in a way is also

419
00:24:55,760 --> 00:24:58,880
specific application already I
would argue, right.

420
00:24:58,880 --> 00:25:03,080
There are steady state problems
that that essentially go for an

421
00:25:03,080 --> 00:25:05,960
equilibrium already or an
average solution where time does

422
00:25:05,960 --> 00:25:09,040
not really the the evolution of
time doesn't really play a role.

423
00:25:09,040 --> 00:25:10,880
It's more a mapping of some
initial conditions in the

424
00:25:10,880 --> 00:25:13,040
geometry or so to the steady
state.

425
00:25:13,920 --> 00:25:17,560
You might need time and other
solutions, the occasions where

426
00:25:17,560 --> 00:25:19,320
you really want to resolve what
happens over time.

427
00:25:19,320 --> 00:25:25,520
But in a way taking that out of
pre training already makes sense

428
00:25:25,520 --> 00:25:30,320
in in retrospect I think.
And then with the synthetic PD

429
00:25:30,320 --> 00:25:34,560
ES, probably you also might not
get the out-of-the-box best

430
00:25:34,960 --> 00:25:39,440
accuracy for right high accuracy
aerodynamic predictions.

431
00:25:40,000 --> 00:25:42,240
But I think it, it seems to be a
very good starting point also

432
00:25:42,240 --> 00:25:45,560
for Navia Stokes.
So what I see as a practical

433
00:25:46,160 --> 00:25:49,160
kind of approach in in the
future would be to pre train

434
00:25:49,160 --> 00:25:51,800
some very generic models as a
starting point.

435
00:25:51,800 --> 00:25:55,520
And then step by step was first
further pre training stages move

436
00:25:55,520 --> 00:25:59,880
towards more application
specific models, aerodynamics,

437
00:25:59,880 --> 00:26:03,520
maybe other one, maybe aero
acoustics, other one, maybe

438
00:26:03,600 --> 00:26:08,240
structural structural problems
or branch off different

439
00:26:08,240 --> 00:26:10,560
directions.
But pre training doesn't need to

440
00:26:10,560 --> 00:26:13,560
be restricted to these canonical
PD ES, but could go in certain

441
00:26:13,560 --> 00:26:16,720
stages and then at a company you
might for a certain use case

442
00:26:16,720 --> 00:26:21,840
then have 10 final models to
fine tune the the last one for a

443
00:26:21,840 --> 00:26:25,560
certain application.
I just think just this this

444
00:26:25,560 --> 00:26:29,440
outlook is basically not
starting from scratch but pre

445
00:26:29,440 --> 00:26:32,600
training and then downstream
needing less data.

446
00:26:33,600 --> 00:26:39,160
I think it's potentially a good.
So do you so because of that, do

447
00:26:39,160 --> 00:26:43,840
you feel that there can be, I
mean how general can these

448
00:26:43,840 --> 00:26:46,760
models go do you think?
I mean if you if you put a

449
00:26:46,760 --> 00:26:49,120
looking glass.
Admittedly, that's an open

450
00:26:49,400 --> 00:26:54,080
question and in in a way also we
were surprised that the bonus

451
00:26:54,080 --> 00:26:59,000
can learn something useful from
these admittedly relatively

452
00:26:59,000 --> 00:27:02,920
useless synthetic solutions.
But it's a really interesting

453
00:27:02,920 --> 00:27:05,360
question to analyse what they
actually learned.

454
00:27:05,360 --> 00:27:07,200
So there are we had we had some
hypothesis.

455
00:27:07,200 --> 00:27:11,160
Basically, do they really just
learn or less random abstract

456
00:27:11,160 --> 00:27:15,280
fields or physical structures
like the key modes?

457
00:27:15,280 --> 00:27:18,040
Or do they learn differential
operators like Also, these PD ES

458
00:27:18,040 --> 00:27:22,440
are built from certain
derivatives and it's difficult

459
00:27:22,440 --> 00:27:27,080
to disentangle.
But for example, the initial

460
00:27:27,080 --> 00:27:28,960
conditions are relatively easy
to test.

461
00:27:28,960 --> 00:27:32,320
So the PD ES are initialized.
If you only train with these

462
00:27:32,320 --> 00:27:35,240
random fields, the models
actually also learn something,

463
00:27:35,520 --> 00:27:40,120
but it clearly does less good on
the Stokes than if you pre train

464
00:27:40,120 --> 00:27:42,760
with the PD ES.
So basically a transformation of

465
00:27:42,760 --> 00:27:46,400
initial conditions into actual
states of a of APDE seems to be

466
00:27:46,400 --> 00:27:48,600
beneficial.
I think this is just the first

467
00:27:48,600 --> 00:27:50,480
step.
It's it's actually not right

468
00:27:51,960 --> 00:27:54,560
right now difficult to
disentangle what they learned.

469
00:27:54,560 --> 00:27:57,320
And intuitively, I think the
network network currently don't

470
00:27:57,320 --> 00:28:01,120
has any reason to clearly
disentangled modes or

471
00:28:01,120 --> 00:28:04,280
differential operators or so
probably just mixes the most

472
00:28:04,720 --> 00:28:06,960
prominent structures that that
come up.

473
00:28:07,440 --> 00:28:12,280
But in a way also most of the
problems we're dealing with in

474
00:28:12,280 --> 00:28:14,840
in engineering and in real world
applications are typically built

475
00:28:14,840 --> 00:28:18,120
from certain like building
blocks to differential equations

476
00:28:18,520 --> 00:28:20,800
in a way have a have a very
limited vocabulary.

477
00:28:21,720 --> 00:28:24,960
So my guess is that the
structures that come from these

478
00:28:24,960 --> 00:28:29,680
actually in a way can be pre
trended and reused.

479
00:28:32,200 --> 00:28:38,600
Yeah.
And I guess the the other, I

480
00:28:38,600 --> 00:28:43,320
mean, so maybe to put it in a
nutshell, in terms of your

481
00:28:43,320 --> 00:28:48,080
research and your thinking, do
you think it is more better to

482
00:28:48,080 --> 00:28:53,800
go down this sort of PDE route
with an integrated solver in the

483
00:28:53,800 --> 00:28:57,960
loop rather than coming from the
other direction, which is just

484
00:28:57,960 --> 00:29:02,320
generating lots and lots of, you
know, if if we're trying to

485
00:29:02,320 --> 00:29:07,200
ultimately get to predict a
wing, practically speaking, do

486
00:29:07,200 --> 00:29:10,920
you think it's better just to
simulate a million wings in all

487
00:29:10,920 --> 00:29:15,720
different conditions or to sort
of build up from the ground from

488
00:29:15,720 --> 00:29:19,320
a sort of PDE based that may be
pre trained?

489
00:29:19,320 --> 00:29:23,520
And then you go to more complex,
more complex, because they still

490
00:29:23,520 --> 00:29:26,560
do feel like different
directions to go, right?

491
00:29:26,920 --> 00:29:33,080
The PD ES definitely aim for
generality in a way I think also

492
00:29:33,280 --> 00:29:36,040
based on previous work, if your
data set is large enough, if you

493
00:29:36,040 --> 00:29:38,680
have enough data to really train
the large scale model

494
00:29:39,040 --> 00:29:42,840
specifically for the task you're
interested in, I don't think you

495
00:29:42,840 --> 00:29:48,320
actually need this this much
more generic starting points.

496
00:29:50,920 --> 00:29:55,240
So if you can do this, I think
probably it's not not necessary.

497
00:29:56,440 --> 00:29:59,000
It's really more motivated by
the practical constraints.

498
00:29:59,000 --> 00:30:02,920
I think at the moment that's
especially in in 3D for large

499
00:30:02,920 --> 00:30:07,880
scale Pdes, it's actually
difficult to gather these really

500
00:30:07,880 --> 00:30:12,640
large collective, Yeah.
I, I'm, I'm, I'm kind of

501
00:30:12,640 --> 00:30:16,240
intrigued on that point just
because, you know, as that fluid

502
00:30:16,240 --> 00:30:19,360
intelligence paper tried to
point out, there is a certain

503
00:30:21,440 --> 00:30:27,920
economic challenge, which is
feels to be different than

504
00:30:27,920 --> 00:30:34,000
ChatGPT where, you know, chat
GBTS cost does not really

505
00:30:34,000 --> 00:30:38,000
include the, the data, the token
cost per SE.

506
00:30:38,000 --> 00:30:41,400
And, and therefore it's all on
the trading side.

507
00:30:43,040 --> 00:30:47,600
And I, I still, I still
fundamentally wonder, is it

508
00:30:47,640 --> 00:30:50,920
massively inefficient?
Because someone said to me,

509
00:30:50,920 --> 00:30:54,640
well, if you run cases, maybe
I'll pose this question to you.

510
00:30:54,640 --> 00:31:00,520
So someone said, OK, so if you
need to generate 2,000,000 cases

511
00:31:01,720 --> 00:31:07,000
in order to fully scope out what
is a reasonable definition of

512
00:31:07,000 --> 00:31:10,720
problems within fluids, So all
possible planes and cars and

513
00:31:10,720 --> 00:31:15,480
data centers and all the rest of
it, and then you use all that to

514
00:31:15,480 --> 00:31:18,000
train a model.
Haven't you already got the

515
00:31:18,000 --> 00:31:20,560
solution from CFD for most of
these problems anyway?

516
00:31:20,560 --> 00:31:22,400
Because you probably run it
like, what?

517
00:31:22,400 --> 00:31:27,880
Where is it just a, you know,
you're almost so massively in

518
00:31:27,880 --> 00:31:30,800
distribution that you almost
could just map that solution to

519
00:31:30,800 --> 00:31:31,960
a new.
Yeah.

520
00:31:33,360 --> 00:31:36,920
But you're right that that that
might happen.

521
00:31:36,920 --> 00:31:43,240
So 2 million seems seems a
little bit scarce for for actual

522
00:31:43,240 --> 00:31:47,320
geometry aside.
But if if you could pretend

523
00:31:47,360 --> 00:31:50,040
enough of them, I guess you
could hope that any new one you

524
00:31:50,040 --> 00:31:53,880
run would actually just hit
close to one of the existing

525
00:31:53,880 --> 00:31:56,560
samples basically.
Then I think a practical problem

526
00:31:56,560 --> 00:31:59,320
would still be that essentially
you need to look up this data

527
00:31:59,320 --> 00:32:02,000
somehow, right?
So these two all right, millions

528
00:32:02,000 --> 00:32:04,960
of simulations would probably be
I don't know how how many

529
00:32:04,960 --> 00:32:09,080
petabytes or gigantic lots of
storage.

530
00:32:09,080 --> 00:32:12,400
So it could be beneficial to pre
train just to have a reduced

531
00:32:12,400 --> 00:32:15,480
compressed representation to be
able to efficiently look up

532
00:32:15,800 --> 00:32:19,480
these data points and maybe also
give give some smoothness in

533
00:32:19,480 --> 00:32:21,880
between, right.
That's even if your exact

534
00:32:21,880 --> 00:32:25,760
geometry is not in there, you
get in in between solutions.

535
00:32:25,760 --> 00:32:30,240
So it, it could actually still
be beneficial from a from a

536
00:32:30,240 --> 00:32:36,000
practical standpoint.
But yeah, given how difficult it

537
00:32:36,000 --> 00:32:41,520
is and, and the amount of data
involved with these online

538
00:32:41,520 --> 00:32:45,720
generated data sets, it seems
that you can actually get away

539
00:32:45,720 --> 00:32:49,800
with much less data in a way.
The the hope would be that you

540
00:32:49,800 --> 00:32:55,200
do this as a first step and then
right maybe you can get away

541
00:32:55,200 --> 00:33:00,120
with 1,000,000 off for CFD
simulations or significantly

542
00:33:00,120 --> 00:33:02,520
less at least to to get to the
same level of accuracy.

543
00:33:02,520 --> 00:33:06,440
Practically in your maybe one of
the topics that also interests

544
00:33:08,080 --> 00:33:11,320
people and be interested to see
how from a coding point of view

545
00:33:11,320 --> 00:33:16,080
you've seen this progress.
So you raise the point of

546
00:33:16,280 --> 00:33:19,280
training the data, sorry,
generating the data, throwing it

547
00:33:19,280 --> 00:33:21,720
away because you're ultimately
using it.

548
00:33:21,720 --> 00:33:25,000
So maybe you could talk a little
bit more what you saw at the

549
00:33:25,000 --> 00:33:27,600
challenges from an
implementation from a coding

550
00:33:27,600 --> 00:33:30,840
point of view, because that
that's I guess also part of the

551
00:33:30,840 --> 00:33:33,240
problem if everybody's
generating these hundreds of

552
00:33:33,240 --> 00:33:37,600
thousands of cases offline and
then trying to bring into a

553
00:33:37,600 --> 00:33:40,400
model that also seems quite
inefficient, so.

554
00:33:42,160 --> 00:33:46,520
The infrastructure work is quite
substantial and I think it's

555
00:33:46,520 --> 00:33:51,440
getting better also largely
thanks to all the methodology is

556
00:33:51,440 --> 00:33:56,640
converging on this LLM
transformer style processing,

557
00:33:58,560 --> 00:34:01,920
which probably is is one of the
reasons why now finally we're

558
00:34:01,920 --> 00:34:05,800
we're in a stage where it starts
to work because also we have the

559
00:34:06,600 --> 00:34:09,719
the tools and and outlooks how
to stay up things up to these

560
00:34:09,719 --> 00:34:15,560
larger resolutions.
It's nonetheless still quite

561
00:34:15,639 --> 00:34:19,199
tricky in practice.
So also in my group took quite a

562
00:34:19,199 --> 00:34:22,560
while to to figure out how to
put these pieces together.

563
00:34:22,560 --> 00:34:25,880
And there are numerous caveats
and, and things that can go

564
00:34:25,880 --> 00:34:30,080
wrong.
So actually, admittedly, one of

565
00:34:30,080 --> 00:34:33,080
the things we're still fighting
with is the non linear scaling.

566
00:34:33,440 --> 00:34:36,280
So even with Transformers, which
which are demonstrated to scale

567
00:34:36,280 --> 00:34:40,159
up to billions of parameters, if
you take one model and you just

568
00:34:40,159 --> 00:34:43,520
try to increase the size of of
layers and the overall capacity,

569
00:34:44,199 --> 00:34:46,840
it's highly non linear.
So it's not guaranteed to to

570
00:34:46,840 --> 00:34:51,280
work if you rerun this, just
longer with an increased size.

571
00:34:51,600 --> 00:34:55,960
Most likely at least need to
adjust the learning rate, the

572
00:34:55,960 --> 00:34:59,400
additional hyper parameters like
smoothing of of network states

573
00:34:59,880 --> 00:35:02,800
over time.
How to actually adjust how to

574
00:35:02,800 --> 00:35:05,800
scale up the size of the network
in terms of this bedding space

575
00:35:05,800 --> 00:35:08,440
that the Transformers have or
the the weights for the

576
00:35:08,440 --> 00:35:12,560
additional components that
unfortunately or it seems it's

577
00:35:12,560 --> 00:35:17,640
necessary to revisit for, for
every new case again and also

578
00:35:17,640 --> 00:35:20,080
for the for the PD case,
unfortunately.

579
00:35:21,400 --> 00:35:26,920
So it's still not not right, not
not really trivial to.

580
00:35:27,880 --> 00:35:30,240
Do this.
And then right, also just

581
00:35:31,360 --> 00:35:33,800
ability, we fought quite a bit
just with with the data

582
00:35:33,800 --> 00:35:37,520
management.
So we rely on these official

583
00:35:37,520 --> 00:35:41,080
compute infrastructures here
from the very end.

584
00:35:41,080 --> 00:35:43,720
And in Germany we now have
supercomputers with a fair

585
00:35:43,720 --> 00:35:46,520
number of GPUs.
But your someone need to get the

586
00:35:46,520 --> 00:35:49,200
data over there and then you can
just store it on on temporary

587
00:35:49,200 --> 00:35:51,480
drives.
And if you don't pay attention

588
00:35:51,480 --> 00:35:55,240
then suddenly your trainer data
has gone if you don't train

589
00:35:55,240 --> 00:35:58,760
often enough and issues issues
like this basically.

590
00:35:58,800 --> 00:36:02,240
But how did you solve that with
some of the just wanted to

591
00:36:02,240 --> 00:36:06,640
double click on the you know one
GPU to generate the data, 3 to

592
00:36:06,640 --> 00:36:07,440
trade.
How?

593
00:36:07,600 --> 00:36:11,760
How were you getting around the
passing the data between?

594
00:36:13,560 --> 00:36:16,240
Yeah, I'm just trying to learn a
little bit more how you approach

595
00:36:16,240 --> 00:36:19,040
that that topic of the online
training.

596
00:36:19,080 --> 00:36:21,720
With the with the online
training, so right, the classic

597
00:36:21,720 --> 00:36:24,240
approach would be right there.
You, you have your huge data

598
00:36:24,240 --> 00:36:28,480
sets on disk and eventually if
you don't want to overfit to

599
00:36:28,480 --> 00:36:30,960
what fits into memory, you
somehow need to get it from the

600
00:36:30,960 --> 00:36:35,920
disk, get it to the to the GPU's
with this large data set that

601
00:36:35,920 --> 00:36:40,680
that becomes a bottleneck, at
least for right, actually for

602
00:36:40,680 --> 00:36:44,120
even if you've hundreds of
millions of parameters on HPC

603
00:36:44,120 --> 00:36:47,520
systems, I think we notice
especially and the the hardware,

604
00:36:48,040 --> 00:36:51,760
the interconnects are good, but
not fast enough to really get

605
00:36:51,760 --> 00:36:54,000
the data in quickly enough for
training.

606
00:36:54,280 --> 00:36:55,600
So the standard set up would be
right.

607
00:36:55,600 --> 00:36:58,040
You have your DP us they, they
need to load the data from this,

608
00:36:58,440 --> 00:37:00,480
then shuffle it through the
network to get a gradient and

609
00:37:00,480 --> 00:37:02,920
they need to load the next
sample basically.

610
00:37:03,400 --> 00:37:07,680
And for this online training, we
are already basically forced to

611
00:37:07,880 --> 00:37:10,440
deal with a separate process
that does the generation.

612
00:37:10,880 --> 00:37:12,600
So it's natural to put a buffer
in between.

613
00:37:12,800 --> 00:37:16,080
So you basically have a, a
buffer that is filled up by the

614
00:37:16,080 --> 00:37:20,840
simulator and then the training
sets just pull the data from

615
00:37:20,840 --> 00:37:24,800
there or you take a random
sample that's available, then

616
00:37:24,800 --> 00:37:27,000
train with it.
And we try to replace it as

617
00:37:27,000 --> 00:37:29,960
quickly as possible.
So in practice this this doesn't

618
00:37:29,960 --> 00:37:34,640
guarantee every sample is used
only once, but you can measure

619
00:37:34,840 --> 00:37:37,640
right the the throughput of your
simulator versus the training

620
00:37:37,640 --> 00:37:41,000
and it's typically at least
below 2, so somewhere between

621
00:37:41,400 --> 00:37:44,160
11.5.
The South descendants get reused

622
00:37:45,960 --> 00:37:48,480
before they get replaced.
So did you buffering?

623
00:37:49,120 --> 00:37:50,640
That's interesting.
I mean, I need to look a little

624
00:37:50,640 --> 00:37:54,600
bit more at the the paper and
the code, but that was, that was

625
00:37:54,600 --> 00:37:59,560
the bit that I guess I was
assuming maybe for your PDE

626
00:37:59,560 --> 00:38:03,080
problem is easier.
But if we're doing what is,

627
00:38:03,480 --> 00:38:06,720
let's say if we want to have
something really accurate and

628
00:38:06,720 --> 00:38:11,080
we're doing an LES simulation,
the actual LES simulation might

629
00:38:11,080 --> 00:38:15,640
need to run on 64 GPUs for like
8 hours, but the training is

630
00:38:15,640 --> 00:38:19,840
obviously way faster than that.
So do you have any thoughts on

631
00:38:19,840 --> 00:38:21,480
that?
How you would balance from a

632
00:38:21,480 --> 00:38:24,640
time from a loading point of
view?

633
00:38:25,560 --> 00:38:29,240
So my guess is that that's why
this online training so far

634
00:38:29,240 --> 00:38:33,600
hasn't really taken off before
because like you mentioned for

635
00:38:33,600 --> 00:38:36,480
all classic simulations, exactly
you have the the warm up time

636
00:38:36,480 --> 00:38:40,080
until you get some equilibrium
that's physically valid and that

637
00:38:40,080 --> 00:38:44,200
that could be used and it might
take for for realistic reward

638
00:38:44,200 --> 00:38:48,200
case might take hours until
you're in that regime and need

639
00:38:48,360 --> 00:38:52,680
fair number of of CPUs at least
or DP US if you have a modern

640
00:38:52,680 --> 00:38:54,840
solver.
So that's, that's really

641
00:38:54,840 --> 00:38:58,160
unattractive because the, the
speed of generating the data is

642
00:38:58,160 --> 00:39:02,320
way below what you need for
training, which is why I think

643
00:39:02,320 --> 00:39:05,520
it's actually so interesting to
use these canonical PDS with

644
00:39:05,520 --> 00:39:10,280
these Spectra solvers because
it, it's really, the solver is

645
00:39:10,280 --> 00:39:13,520
basically as fast as, as a
network or typically we run

646
00:39:13,720 --> 00:39:18,080
actually the, the training we
run on little regions like 64 ^3

647
00:39:18,640 --> 00:39:21,560
and the solver typically produce
large, produces larger ones like

648
00:39:21,560 --> 00:39:26,080
256 or so.
And even those are pretty close

649
00:39:26,080 --> 00:39:29,280
to the, to the training speed.
And then you can basically cut

650
00:39:29,280 --> 00:39:31,600
out different pieces for data
augmentation.

651
00:39:33,080 --> 00:39:37,680
But it, it matches quite nicely.
So I think without such a solver

652
00:39:38,080 --> 00:39:41,520
doing a large scale pre training
is is infeasible.

653
00:39:42,080 --> 00:39:44,840
I think there are some
approaches doing this, but you

654
00:39:45,080 --> 00:39:48,920
if you want to match the
generation capacity with the

655
00:39:48,920 --> 00:39:52,000
training capacity you would need
for a supercomputer to generate

656
00:39:52,000 --> 00:39:56,160
data on the fly just to feed a
decent sized model.

657
00:39:56,800 --> 00:40:00,400
Yeah, that though then my other
thought was actually whilst it

658
00:40:00,400 --> 00:40:05,520
takes 8 hours on 64 GPUs, let's
say, to do it, that is the

659
00:40:05,520 --> 00:40:08,720
entire simulation.
But actually to generate one

660
00:40:08,720 --> 00:40:12,960
time step is probably only a
second or two seconds.

661
00:40:13,840 --> 00:40:19,640
So part of my interest, and
maybe some groups already doing

662
00:40:19,640 --> 00:40:27,000
this is training per time step.
Now, the variation from one time

663
00:40:27,000 --> 00:40:28,480
step to another is pretty
minimal.

664
00:40:28,560 --> 00:40:32,800
So you could argue that you're
not really learning that much

665
00:40:33,440 --> 00:40:37,120
between, but you are.
That's the only way I could see

666
00:40:37,120 --> 00:40:42,520
that the time scales being
similar, You know, which I guess

667
00:40:42,520 --> 00:40:45,840
ultimately addresses one of the
challenges that I think many

668
00:40:45,840 --> 00:40:50,880
people have, which is almost all
of the standard data sets and

669
00:40:50,880 --> 00:40:54,560
standard machine learning
approaches do take some time

670
00:40:54,560 --> 00:40:58,920
averaged solution.
You know, which if you're trying

671
00:40:58,920 --> 00:41:03,520
to get to real problems and real
accuracy, turbulence is

672
00:41:03,720 --> 00:41:08,680
transient, you know, and giving
just a steady state answer or

673
00:41:08,680 --> 00:41:14,400
time average answer is, you
know, not, not really

674
00:41:14,400 --> 00:41:18,520
representing the ultimate goal
of like weather forecasting.

675
00:41:18,520 --> 00:41:23,000
I guess you would have the
temporal evolution of it.

676
00:41:23,000 --> 00:41:26,600
So I don't know how you how we
can get over that.

677
00:41:26,680 --> 00:41:28,640
Issue.
I think it's a great order to to

678
00:41:28,640 --> 00:41:32,200
really target time.
I think there are some probably

679
00:41:32,680 --> 00:41:36,680
what I see there are quite some
hurdles for for practitioners

680
00:41:36,680 --> 00:41:39,480
because we have all these
pipelines set up for for

681
00:41:39,480 --> 00:41:42,480
averaged quantities.
All right, potentially I think

682
00:41:42,480 --> 00:41:44,560
this would be would be neat to
have in place.

683
00:41:45,560 --> 00:41:48,360
Regarding your first point
though, with the kind of

684
00:41:49,600 --> 00:41:53,680
exploiting the, the fast solves
over time for data generation.

685
00:41:54,160 --> 00:41:56,360
Unfortunately he was tadpole
with his online training.

686
00:41:58,480 --> 00:42:03,240
We have some, some, some data
that is pointing to to problems

687
00:42:03,240 --> 00:42:04,520
there.
So what we noticed even with

688
00:42:04,520 --> 00:42:09,240
this online generation if we
don't replace the data fast

689
00:42:09,240 --> 00:42:12,440
enough.
So if this reuse is too high of

690
00:42:12,440 --> 00:42:16,480
our our pre trained buffer
basically the models do start to

691
00:42:16,480 --> 00:42:19,320
overfit.
So typically for foundation

692
00:42:19,320 --> 00:42:22,040
models you're dealing with
fairly large models.

693
00:42:22,080 --> 00:42:24,800
So ours are not even extreme,
but on the order of maybe 100

694
00:42:24,800 --> 00:42:28,680
million parameters.
And if the samples are reused

695
00:42:28,680 --> 00:42:32,160
too often, the bars already we
saw some signs of of performance

696
00:42:32,160 --> 00:42:35,600
deteriorating if the generator
didn't catch up.

697
00:42:36,400 --> 00:42:41,160
And basically this was just out
of pure luck because of of

698
00:42:41,160 --> 00:42:43,400
scheduling on these high
performance systems of right.

699
00:42:43,400 --> 00:42:47,760
If there's some one of the GPUs
is is for some reason slower and

700
00:42:47,760 --> 00:42:50,960
suddenly on when training run
you reuse the data more often.

701
00:42:51,360 --> 00:42:56,040
We saw a drop in performance and
even there the the correlation

702
00:42:56,040 --> 00:42:57,880
between the samples was not
overly strong.

703
00:42:57,880 --> 00:43:00,840
So it's still generative fair
amount of data.

704
00:43:01,640 --> 00:43:04,320
But I would be worried if you
have these strongly correlated

705
00:43:04,320 --> 00:43:07,560
samples over time.
Even if you're able to swap them

706
00:43:07,560 --> 00:43:12,040
out and reload reload different
configurations for a solver, I

707
00:43:12,040 --> 00:43:16,720
think this would be would be
difficult to ensure that the

708
00:43:16,720 --> 00:43:20,040
variability in the data is large
enough to right not overfit a

709
00:43:20,240 --> 00:43:24,040
large model.
Yeah, that's a very good point

710
00:43:24,040 --> 00:43:26,720
actually, that that's, that's
the gang.

711
00:43:27,000 --> 00:43:31,360
Yeah.
The drawback that we, the

712
00:43:31,360 --> 00:43:35,240
temporal evolution of the PDE
needs to be small in for

713
00:43:35,240 --> 00:43:39,560
numerical reasons, you know, to,
to avoid blowing up, you know,

714
00:43:39,560 --> 00:43:41,840
for because of the CFL
constraints, etcetera.

715
00:43:41,840 --> 00:43:46,440
So you you're kind of forced to
slowly March in time, whereas I

716
00:43:46,440 --> 00:43:48,920
guess you're right, if you feed
the machine learning model that

717
00:43:49,240 --> 00:43:53,040
it's just going to keep seeing
essentially the same solution,

718
00:43:53,320 --> 00:43:54,520
very similar.
Ones that.

719
00:43:55,120 --> 00:44:00,960
And and yeah, that's, but that's
where I still see a bit of a

720
00:44:00,960 --> 00:44:02,960
fundamental.
I just don't see.

721
00:44:03,120 --> 00:44:05,040
That's why I see almost as
that's why I was interested in

722
00:44:05,040 --> 00:44:08,080
your differentiable physic, your
your code writing essentially

723
00:44:08,080 --> 00:44:14,440
because I feel some of the
breakthroughs in this probably

724
00:44:14,440 --> 00:44:20,160
will come through a clever,
clever use of the training being

725
00:44:20,160 --> 00:44:23,520
linked to the data generation in
a way that.

726
00:44:25,080 --> 00:44:28,240
I think right now our focus is
on foundation was I'm also

727
00:44:28,240 --> 00:44:30,800
confident that at some point
it's going to make sense to

728
00:44:30,800 --> 00:44:37,080
bring the solvers back in the
our our tests back back then a

729
00:44:37,120 --> 00:44:40,280
few years ago basically did did
pretty short.

730
00:44:40,280 --> 00:44:42,960
If you have a solver, there's
anything decent, the learning

731
00:44:42,960 --> 00:44:47,040
task is just simpler.
So once once we have kind of

732
00:44:48,160 --> 00:44:51,960
reached the decent reasonable
capacity for this pre training,

733
00:44:52,840 --> 00:44:56,960
I think it's might make a lot of
sense to put a solver back into

734
00:44:56,960 --> 00:44:59,120
the loop to just make it more
accurate in the end.

735
00:45:02,840 --> 00:45:05,240
Also, they are it's it is
challenging at the moment at the

736
00:45:05,240 --> 00:45:09,760
at the scope given the current
infrastructures and and large

737
00:45:09,880 --> 00:45:13,000
scale models.
It's typically challenging to

738
00:45:13,000 --> 00:45:17,440
just train a single model and in
a decent amount of time and then

739
00:45:17,440 --> 00:45:20,520
trying to get a solver into the
picture and maybe going over

740
00:45:20,520 --> 00:45:23,840
multiple time steps.
Things like this are are tricky,

741
00:45:25,600 --> 00:45:28,480
but it's also going to change
next year's.

742
00:45:29,240 --> 00:45:31,280
Yeah, yeah.
No, I I would agree.

743
00:45:31,280 --> 00:45:38,640
And maybe the the other topic is
around what were some of your

744
00:45:38,640 --> 00:45:43,280
lessons learned from doing, you
know, weather bench and AP bench

745
00:45:43,280 --> 00:45:46,400
and, and some of these more
benchmarking efforts because I

746
00:45:46,400 --> 00:45:51,520
guess that seems to have helped
quite a bit on the weather and

747
00:45:51,520 --> 00:45:53,240
climate side.
So what?

748
00:45:53,240 --> 00:45:56,840
What was some of the genesis
behind those efforts?

749
00:45:58,040 --> 00:46:03,120
So the weather bench effort,
yeah, in retrospect, it's, it's

750
00:46:03,120 --> 00:46:05,040
great that it's, it's doing so
well.

751
00:46:06,640 --> 00:46:11,640
I, I think it's especially, I
guess I should think that so,

752
00:46:11,760 --> 00:46:12,960
right.
What I'm trying to say is

753
00:46:12,960 --> 00:46:16,920
basically the, these benchmarks
have a big impact in the field.

754
00:46:16,920 --> 00:46:19,440
I think by now that is widely
accepted across all the fields.

755
00:46:20,160 --> 00:46:24,040
Back then it was probably more
clear in the vision area where

756
00:46:25,080 --> 00:46:27,520
fair number of of benchmarks
have been around for quite a

757
00:46:27,520 --> 00:46:29,480
while.
But in other fields, like

758
00:46:29,480 --> 00:46:34,640
whether there was very little,
basically having established

759
00:46:34,640 --> 00:46:39,360
benchmarks and libel evaluations
is really important for for any

760
00:46:39,400 --> 00:46:43,680
discipline in the field.
I think that's a nice pointer

761
00:46:43,680 --> 00:46:46,160
and right, it's civic.
It's it's quite some work

762
00:46:46,560 --> 00:46:48,680
collecting the data, right and
thinking about what should be in

763
00:46:48,680 --> 00:46:51,840
there and how to evaluate it.
So the way those efforts are

764
00:46:52,320 --> 00:46:57,440
really important for any
subfield within any data-driven

765
00:46:57,840 --> 00:46:59,920
discipline.
And I think by now it's great

766
00:46:59,920 --> 00:47:05,040
also in this scientific Yeah, I
feel that this is noticed that

767
00:47:05,040 --> 00:47:07,720
by now more and more benchmarks
are coming out.

768
00:47:08,480 --> 00:47:10,200
I think that's.
Really important you you

769
00:47:10,200 --> 00:47:15,560
recently generated to your group
the Super wing data set.

770
00:47:16,000 --> 00:47:17,520
I mean, where do you see this
going?

771
00:47:17,520 --> 00:47:21,720
And it's a little bit of a
philosophical debate around data

772
00:47:21,720 --> 00:47:27,640
because on one hand you could
argue that we the progress in

773
00:47:27,640 --> 00:47:31,120
this field does seem limited by
data to a certain extent.

774
00:47:31,720 --> 00:47:38,640
And, you know, does it mean that
there should be some coordinated

775
00:47:38,640 --> 00:47:46,920
effort to generate data?
And if so, who should be doing

776
00:47:46,920 --> 00:47:48,720
that?
And what should be the license

777
00:47:48,760 --> 00:47:52,520
attached to it, given that there
is clearly also a commercial

778
00:47:53,400 --> 00:47:55,400
benefit to be had?
That's a good question.

779
00:47:55,400 --> 00:47:58,440
I mean coming from from a
university, I think it's great

780
00:47:58,440 --> 00:48:00,640
if it's all public and as open
as possible.

781
00:48:00,680 --> 00:48:04,520
But given the commercial
interest and also by now the the

782
00:48:04,520 --> 00:48:07,360
key outlook that this will be
useful, that would of course be

783
00:48:07,360 --> 00:48:12,000
great to have support from from
commercial partners in this.

784
00:48:12,320 --> 00:48:15,720
I think it's getting better.
Also more, more companies see

785
00:48:15,960 --> 00:48:20,480
the need for AI and then also
the need for or benchmarks and

786
00:48:20,480 --> 00:48:23,880
data sets on that front.
So I think it's improving, but

787
00:48:23,880 --> 00:48:27,400
definitely an issue where things
could be done and could improve

788
00:48:27,400 --> 00:48:30,880
a lot.
And unfortunately there's so

789
00:48:30,880 --> 00:48:33,840
much to quickly wrap this up.
But I think there's so much

790
00:48:34,200 --> 00:48:37,800
proprietary data with IP rights
and so on that are probably on

791
00:48:37,800 --> 00:48:40,840
some servers and companies that
they that cannot be used.

792
00:48:41,760 --> 00:48:44,280
So it probably needs an effort
to generate data that's free and

793
00:48:44,280 --> 00:48:48,200
and usable.
Yeah, that I keep jumping

794
00:48:48,200 --> 00:48:53,040
between that because on one hand
you could say that it is to the

795
00:48:53,080 --> 00:48:56,200
a bit like maybe open science
experiments, you know, with,

796
00:48:56,280 --> 00:49:00,760
with astronomy or, or, or things
where there's a sort of public

797
00:49:00,760 --> 00:49:03,880
good and a lot of it is funded
through taxpayers.

798
00:49:03,880 --> 00:49:06,760
Ultimately, you know, through
sort of science funding and the

799
00:49:06,760 --> 00:49:13,880
data's made available.
Part of it feels that that would

800
00:49:13,880 --> 00:49:16,480
help all companies, you know, to
be doing it.

801
00:49:16,640 --> 00:49:20,400
But at the same time, is it
really the responsibility of

802
00:49:21,840 --> 00:49:24,160
governments and science to do
this?

803
00:49:24,960 --> 00:49:31,320
It it's it's kind of.
Yeah, that's good to argue about

804
00:49:31,320 --> 00:49:32,960
the companies.
So pay for it if they benefit

805
00:49:32,960 --> 00:49:36,040
from it, that's.
Yeah, that's it.

806
00:49:37,160 --> 00:49:39,280
I just look at, I don't know
about you, but I look at all the

807
00:49:39,280 --> 00:49:43,640
supercomputers in Europe, just
as an example, you know, and you

808
00:49:43,640 --> 00:49:47,720
think how much data could be
generated given that these

809
00:49:47,720 --> 00:49:51,720
systems now are quite large
because they're being scaled up

810
00:49:51,720 --> 00:49:56,960
for the task of also, you know,
large language model training,

811
00:49:56,960 --> 00:49:58,240
etcetera.
You know, if you've got a

812
00:49:58,240 --> 00:50:02,600
cluster with 10,000 GPUs, you
think how much data could you

813
00:50:02,600 --> 00:50:07,720
generate, you know, from fluid
simulations to a 10,000 GPUs,

814
00:50:08,720 --> 00:50:12,760
even just, you know, for a few
weeks or, or you know, or a

815
00:50:12,760 --> 00:50:20,920
month, you know, that does feel
like it could be a, a huge way

816
00:50:20,920 --> 00:50:24,800
of doing it.
And if, if it's only a

817
00:50:24,800 --> 00:50:28,160
commercial company that does it,
they probably have no incentive

818
00:50:28,160 --> 00:50:34,400
to release that data.
And it then becomes hard to

819
00:50:34,400 --> 00:50:37,840
progress the field if only one
company has that data where it

820
00:50:37,840 --> 00:50:42,840
feels like with large language
models, the data has actually

821
00:50:42,840 --> 00:50:44,360
had quite an open movement,
right?

822
00:50:44,360 --> 00:50:47,240
There's quite a lot of open data
obviously around this.

823
00:50:47,720 --> 00:50:50,520
And so, yes, you still could say
you can only make models if

824
00:50:50,520 --> 00:50:54,680
you've got lots of compute, but
there was already a movement now

825
00:50:54,680 --> 00:51:00,560
with open source LMS, you know,
going on, whereas I feel now how

826
00:51:00,560 --> 00:51:04,880
can there be a movement of open
source surrogate models if the

827
00:51:05,560 --> 00:51:10,040
data is not there in an open
way, You know, it it it's.

828
00:51:11,680 --> 00:51:16,160
It's a good point.
It would be a pity if that gets

829
00:51:16,680 --> 00:51:19,280
hidden away due to commercial
interests.

830
00:51:19,320 --> 00:51:22,720
So it's almost the.
Progression of science, I guess,

831
00:51:22,720 --> 00:51:25,400
and this is, I don't know if
you've had this and certainly

832
00:51:25,400 --> 00:51:29,720
now my myself being a commercial
company or the last two

833
00:51:29,720 --> 00:51:33,520
companies that you know that I'm
at, there is always this debate

834
00:51:33,600 --> 00:51:39,040
of open source versus
commercial, and I still find

835
00:51:39,040 --> 00:51:41,640
that a tricky one.
I wondered how you've seen it.

836
00:51:41,640 --> 00:51:44,600
You know what, what at what
point does it help everybody to

837
00:51:44,600 --> 00:51:49,680
be open source?
Yeah, OK, right.

838
00:51:49,720 --> 00:51:52,560
I have AI have a good.
Argument for open source from

839
00:51:52,560 --> 00:51:55,720
from my career at least.
We used to work in computer

840
00:51:55,720 --> 00:51:59,560
graphics basically with movie
studios and I remember 1 of was

841
00:51:59,600 --> 00:52:03,160
my postdoc, one of the first
papers at once at Arthur

842
00:52:03,160 --> 00:52:06,840
conference seeker there.
We basically put up with an open

843
00:52:06,840 --> 00:52:11,840
source code and company
DreamWorks at the time actually

844
00:52:11,840 --> 00:52:14,240
picked this up pretty quickly
and directly used it for for one

845
00:52:14,240 --> 00:52:16,280
of the shots.
And I said, oh, can't we get, or

846
00:52:16,280 --> 00:52:18,160
maybe they can put us in the
credit somewhere.

847
00:52:18,600 --> 00:52:23,200
They send us a poster in the end
and also quite a few of my

848
00:52:23,200 --> 00:52:26,880
friends and kind of laugh or
look all all you got for for

849
00:52:26,880 --> 00:52:28,960
this work.
You put up your code, you got a

850
00:52:28,960 --> 00:52:31,680
poster for it.
Why did you sell it or

851
00:52:32,320 --> 00:52:35,160
something?
A couple of years later, these,

852
00:52:35,200 --> 00:52:38,680
all these movies where it was
used were one of the key things

853
00:52:38,680 --> 00:52:40,840
to apply for one of these test
tech Oscars.

854
00:52:41,200 --> 00:52:42,840
There's also application
procedure.

855
00:52:42,840 --> 00:52:45,520
But in the end, we could show,
look, it's been used in all

856
00:52:45,520 --> 00:52:49,200
these movies.
And this this Oscar actually did

857
00:52:49,200 --> 00:52:52,400
help a lot also for applying for
faculty positions and so on.

858
00:52:52,400 --> 00:52:56,160
So this is largely due to to
open source availability.

859
00:52:56,160 --> 00:52:58,600
It just might impact in the
field.

860
00:52:58,800 --> 00:53:02,880
So I've had very good experience
with this and I'm actually very

861
00:53:02,880 --> 00:53:06,720
happy that now that, yeah, I
feel this is so obvious. 15

862
00:53:06,720 --> 00:53:10,480
years ago it was very, it was
actually the exception that

863
00:53:10,480 --> 00:53:12,840
papers came to code and it's
great.

864
00:53:12,840 --> 00:53:16,040
It's changed so much.
Yeah, it, it does seem an

865
00:53:16,040 --> 00:53:19,800
interesting 1 though, because
it's, given the large amount of

866
00:53:19,800 --> 00:53:23,640
money that's required to develop
models and develop training

867
00:53:23,640 --> 00:53:30,240
data, you know, there has to be
some route to that money coming

868
00:53:30,240 --> 00:53:32,440
back.
Which is why for a commercial

869
00:53:32,440 --> 00:53:38,120
company, if they spend $100
million on compute and they, you

870
00:53:38,120 --> 00:53:39,840
know, they have to think, well,
how am I going to make a

871
00:53:40,040 --> 00:53:43,000
business out of this?
And so it's, it's kind of an

872
00:53:43,000 --> 00:53:49,040
interesting one that is the
business that they, that the

873
00:53:49,040 --> 00:53:53,000
value to the company is not the
model, but how you sort of tweak

874
00:53:53,000 --> 00:53:56,560
the model, so to speak.
And therefore, just because

875
00:53:56,560 --> 00:54:00,640
there's a foundation model out
there which is open source, the

876
00:54:00,640 --> 00:54:03,880
real value is that every single
company needs to customize it,

877
00:54:03,880 --> 00:54:06,880
which I think is the value at
the moment for open source LMS

878
00:54:07,320 --> 00:54:10,720
that ultimately companies still
make money out of it because

879
00:54:11,880 --> 00:54:17,560
everybody realizes that it isn't
fully ready in its base form.

880
00:54:17,560 --> 00:54:20,560
And that actually you take it
and you tweak it and you

881
00:54:20,560 --> 00:54:25,160
optimize it and you know, you so
that so there's money still to

882
00:54:25,160 --> 00:54:27,240
be made.
And that's what I kind of wonder

883
00:54:27,240 --> 00:54:31,280
on the fluids surrogate
modelling side or even beyond

884
00:54:31,280 --> 00:54:36,720
that, it would it still be in
the interest of some company to

885
00:54:36,720 --> 00:54:40,000
make it because ultimately the
money will be made customizing

886
00:54:40,000 --> 00:54:46,160
it and therefore being open
helps the whole community to

887
00:54:46,240 --> 00:54:50,320
develop it faster and make it
more competitive against close

888
00:54:50,320 --> 00:54:52,200
models.
Right, that's definitely.

889
00:54:52,200 --> 00:54:56,440
But I hope that, that we can go
towards some foundation models

890
00:54:56,440 --> 00:55:00,320
or some generalizing models in
the field that can then be

891
00:55:00,520 --> 00:55:05,480
fine-tuned and adaptive to
adapted to different, different

892
00:55:05,480 --> 00:55:09,480
applications quite easily.
I think that's, that's been

893
00:55:09,480 --> 00:55:15,440
super useful in the LLM field.
And I, I do see potential there

894
00:55:15,560 --> 00:55:18,680
on the, on the PDE front and
through its front.

895
00:55:20,640 --> 00:55:23,600
So, and I think if, if at some
point it's clear that you can

896
00:55:23,600 --> 00:55:26,840
get a really good result, right,
if if you actually download this

897
00:55:26,840 --> 00:55:30,280
model and then right fine tune
it on your couple of wing data

898
00:55:30,280 --> 00:55:33,640
set cases or so for certain
regime you're interested in,

899
00:55:34,880 --> 00:55:39,120
then.
But there could be a clear

900
00:55:39,440 --> 00:55:43,800
commercial also benefit for a
company and the use case of

901
00:55:43,800 --> 00:55:46,320
supporting these for supporting
these open models.

902
00:55:46,920 --> 00:55:50,840
Yeah.
So that could imagine in an

903
00:55:50,840 --> 00:55:53,960
infrastructure environment
working quite well also in this

904
00:55:53,960 --> 00:55:58,120
area, but that definitely needs
a coordinated and a large scale

905
00:55:58,120 --> 00:56:02,400
effort to build it up right now.
Yeah, that that's seems that

906
00:56:02,640 --> 00:56:05,840
there's lots of separate
efforts, I guess going on, but

907
00:56:05,880 --> 00:56:10,680
but not uncoordinated.
I I mean, maybe moving to some

908
00:56:10,680 --> 00:56:14,960
of the final questions for you
is where well, look, I, I guess

909
00:56:14,960 --> 00:56:17,840
a couple of things before we get
to the future looking one.

910
00:56:17,840 --> 00:56:24,080
I'm just interested, where's
your side on the like startups

911
00:56:24,080 --> 00:56:26,200
or academia industry?
Because I've noticed there's a

912
00:56:26,200 --> 00:56:33,600
lot of, in recent years, there's
even more interest in startups

913
00:56:34,080 --> 00:56:36,120
and, and, and the value of
startups.

914
00:56:36,120 --> 00:56:40,720
But then that also in some ways
can conflict with the open

915
00:56:40,720 --> 00:56:43,040
source academic, you know,
mindset.

916
00:56:43,520 --> 00:56:45,240
Where, where have you seen that
in the world?

917
00:56:45,240 --> 00:56:48,320
Have you been tempted with
startups of, of industry?

918
00:56:48,320 --> 00:56:53,080
Where do you see that role of
academia, start-ups and and

919
00:56:53,080 --> 00:56:54,880
industry?
Yeah, it's a good question.

920
00:56:54,880 --> 00:56:57,040
Exactly.
Especially in last one or two

921
00:56:57,040 --> 00:57:00,200
years we've seen a very nice
rise I think in terms of funding

922
00:57:00,200 --> 00:57:04,920
and and also just founding of
all kinds of spin off companies

923
00:57:05,040 --> 00:57:08,520
that are pivots of quite
existing companies towards this

924
00:57:08,520 --> 00:57:11,640
physically eye direction.
So I think also in a way

925
00:57:11,640 --> 00:57:15,360
confirmation that now it's
really starting to work, right,

926
00:57:15,720 --> 00:57:17,480
It's ready for practical
applications.

927
00:57:18,520 --> 00:57:21,720
I've definitely toyed with the
idea so far or also on my side,

928
00:57:21,720 --> 00:57:23,920
there are no immediate plans for
this.

929
00:57:24,240 --> 00:57:26,480
In a way.
I, I had my experience with

930
00:57:26,480 --> 00:57:30,760
industry I for, for visual
attacks and before I'm quite

931
00:57:30,760 --> 00:57:34,920
happy with the Open University
research side.

932
00:57:34,920 --> 00:57:38,160
So I've, I've always been a big
fan of open source and just

933
00:57:38,160 --> 00:57:41,520
being able to put out things and
for, for research.

934
00:57:41,520 --> 00:57:45,040
I mean, it's also effectively a
market with this openness and

935
00:57:45,040 --> 00:57:47,600
impact, right?
You, if you can generate this

936
00:57:47,600 --> 00:57:50,160
later on, you can say, look,
give me more money for research

937
00:57:50,160 --> 00:57:54,280
to do more of, of this.
So it's not the commercial

938
00:57:54,280 --> 00:57:58,800
market, but also their own
market in terms of research

939
00:57:58,800 --> 00:58:01,000
money.
And but I like the openness of

940
00:58:01,000 --> 00:58:04,840
that.
So for now, I think for me

941
00:58:04,840 --> 00:58:07,640
that's personally just because I
like this openness the the

942
00:58:07,640 --> 00:58:10,640
better direction.
But we also, by now we're

943
00:58:10,640 --> 00:58:14,880
working with all kinds of
companies due to the commercial

944
00:58:14,880 --> 00:58:16,320
interests.
Yeah.

945
00:58:16,480 --> 00:58:18,880
But we were basically trying to
do this on the more Open

946
00:58:18,880 --> 00:58:21,520
University side and then work
with individual companies for

947
00:58:21,520 --> 00:58:25,720
maybe adopting or or adapting
these these techniques.

948
00:58:26,840 --> 00:58:29,560
Yeah, I was going to say because
the decking jury if all top

949
00:58:29,560 --> 00:58:32,840
academics create start-ups, then
there will be no academics left

950
00:58:32,840 --> 00:58:36,160
to do that.
The sort of so that that I'm

951
00:58:36,160 --> 00:58:38,800
sure that has been discussed a
little bit in the broader LM

952
00:58:38,800 --> 00:58:41,600
space.
If the only companies doing the

953
00:58:41,600 --> 00:58:45,200
latest stuff because of the
scale are the big tech

954
00:58:45,200 --> 00:58:51,640
companies, then there is a risk
to open science, I guess because

955
00:58:51,640 --> 00:58:56,520
everyone's tempted to go to a
commercial company that maybe

956
00:58:56,520 --> 00:59:00,000
aren't, as you said,
incentivized to be open.

957
00:59:00,080 --> 00:59:03,040
I guess as academia, you are
incentivized to be open because

958
00:59:03,880 --> 00:59:07,440
the reward structure is based
around it publishing grant

959
00:59:07,440 --> 00:59:09,680
money.
Like if everything's closed, I

960
00:59:09,680 --> 00:59:12,480
guess it may help the university
a little bit because of grant

961
00:59:12,480 --> 00:59:16,520
capture or something, but it's
not really the the core aim, is

962
00:59:16,520 --> 00:59:19,920
it?
It's difficult actually for for

963
00:59:19,920 --> 00:59:22,840
universities.
I think in the past quite a few

964
00:59:22,840 --> 00:59:25,920
labs had their in house solvers
and then basically earned money

965
00:59:25,920 --> 00:59:30,680
and and funding.
With direct we basically support

966
00:59:30,680 --> 00:59:34,240
or consulting contracts with
companies.

967
00:59:35,320 --> 00:59:39,560
But right now I think especially
in the fast moving AI field, it

968
00:59:39,560 --> 00:59:43,160
was, it was closed solutions,
you would see very little impact

969
00:59:45,920 --> 00:59:50,680
at the at the moment, Stephanie.
Works nicely no I, I, I think

970
00:59:50,680 --> 00:59:54,440
that's why I'm always
championing the academic side

971
00:59:54,440 --> 00:59:58,440
because I think people look at
research coming out of tech

972
00:59:58,440 --> 01:00:02,960
companies or research coming out
of the but forget that actually

973
01:00:02,960 --> 01:00:07,560
most of the things start in the
university most most fundamental

974
01:00:07,560 --> 01:00:13,560
ideas and.
They need to be kept being

975
01:00:13,720 --> 01:00:16,720
promoted and that's where I feel
the data's the issue, because

976
01:00:16,720 --> 01:00:21,480
the they if you don't have open
data, then you can't actually

977
01:00:21,720 --> 01:00:25,160
universities can't progress and
can't publish and can't advance.

978
01:00:25,160 --> 01:00:27,760
The state-of-the-art.
If everything's done closed

979
01:00:27,760 --> 01:00:30,640
door, it actually hurts
progress, doesn't it?

980
01:00:31,360 --> 01:00:33,240
Yeah.
So LLMS will be interesting to

981
01:00:33,240 --> 01:00:34,960
some extent, right.
That's there have been

982
01:00:34,960 --> 01:00:39,480
initiatives of training open
LLMS, but some of them I guess

983
01:00:39,480 --> 01:00:42,560
did fairly well.
But the top models are not open.

984
01:00:43,520 --> 01:00:47,520
And right now I think there are
very few people who really can

985
01:00:47,520 --> 01:00:53,440
or try to train their own LLMS.
That's really become a very

986
01:00:53,440 --> 01:00:54,840
tough field, at least to operate
in.

987
01:00:55,640 --> 01:00:59,960
Yes, yes.
So where do you see if you have

988
01:00:59,960 --> 01:01:03,800
your looking glass or your, you
know, crystal ball?

989
01:01:04,560 --> 01:01:08,400
If we were to talk again in five
years time, where where do you

990
01:01:08,400 --> 01:01:10,640
think the conversation will be
going?

991
01:01:10,640 --> 01:01:14,160
What what do you think will be
the breakthrough moments?

992
01:01:14,680 --> 01:01:17,040
Do you feel like we're going to
be just incrementally over the

993
01:01:17,040 --> 01:01:21,800
next five years or do you
perceive some big breakthroughs

994
01:01:21,800 --> 01:01:23,400
or leap?
Just hope that we're going to

995
01:01:23,400 --> 01:01:25,280
see this breakthrough in terms
of adoption.

996
01:01:25,480 --> 01:01:29,960
So now also these all these
initiatives on commercial side I

997
01:01:29,960 --> 01:01:32,400
think point towards the field
really starting to work.

998
01:01:32,400 --> 01:01:36,840
So I hope that five years we we
really see widespread adoption

999
01:01:36,840 --> 01:01:42,920
for for actual applications in a
way sounds a bit boring, right,

1000
01:01:42,920 --> 01:01:44,320
just being it used in in
practice.

1001
01:01:44,320 --> 01:01:46,760
But I think right also working
on this for almost 10 years.

1002
01:01:46,760 --> 01:01:50,120
I think it's about time that we
that we do see that it's got

1003
01:01:50,400 --> 01:01:53,840
it's really useful for for real
world world things.

1004
01:01:54,920 --> 01:01:58,080
So I just hope that now we are
at a stage where this really can

1005
01:01:58,080 --> 01:02:02,600
be can be pulled off, but in the
field is really progressing

1006
01:02:02,720 --> 01:02:05,920
extremely quickly.
And so one of the, one of the

1007
01:02:05,920 --> 01:02:08,760
outlooks that I've been
intrigued about are these world

1008
01:02:08,760 --> 01:02:11,360
models, right?
Like vision by now, they're

1009
01:02:11,440 --> 01:02:14,720
already going beyond foundation
models or specific language and,

1010
01:02:14,720 --> 01:02:18,480
and just visual models towards
ones that just capture the whole

1011
01:02:18,480 --> 01:02:20,440
world.
And right, also vision now

1012
01:02:20,440 --> 01:02:22,600
notices actually physics is an
important part.

1013
01:02:22,600 --> 01:02:25,360
You only see so much.
A lot of the complexity comes

1014
01:02:25,360 --> 01:02:30,560
from things you don't directly
see like air moving or so I

1015
01:02:30,560 --> 01:02:34,320
think that's a really
interesting outlook to to really

1016
01:02:34,320 --> 01:02:37,240
simulate basically the the whole
world around us on a larger

1017
01:02:37,240 --> 01:02:38,880
scale with the physics
accurately.

1018
01:02:39,120 --> 01:02:43,520
I think that could also be super
useful for engineering and and

1019
01:02:43,720 --> 01:02:46,160
real applications in a bunch of
areas.

1020
01:02:46,600 --> 01:02:51,000
But getting that right seems
like a whole different level on

1021
01:02:51,000 --> 01:02:54,280
top of what we currently dealing
with just only doing the physics

1022
01:02:54,280 --> 01:02:55,360
right with the foundation
models.

1023
01:02:55,360 --> 01:02:56,840
But I think it's a really
interesting outlook.

1024
01:02:58,120 --> 01:02:59,520
Yeah, it's actually a good point
there.

1025
01:02:59,520 --> 01:03:00,840
Yeah, I forgot to ask you about
that.

1026
01:03:00,840 --> 01:03:05,480
That I still don't fully get.
I mean, I understand the world

1027
01:03:05,480 --> 01:03:12,240
models are obviously one very
big use case is robotics and

1028
01:03:12,240 --> 01:03:15,680
driving drop, you know,
autonomous vehicles and sort of

1029
01:03:15,680 --> 01:03:19,360
giving that synthetic data to
help them to operate in the

1030
01:03:19,360 --> 01:03:26,200
world, The bit that I guess I'm
less sure on.

1031
01:03:26,200 --> 01:03:29,800
And I wonder whether
economically and scientifically

1032
01:03:29,800 --> 01:03:35,120
is the physics side, you know,
does it make a difference how

1033
01:03:35,120 --> 01:03:39,400
good the physics is like, as
long as it looks right?

1034
01:03:39,560 --> 01:03:41,360
And this, This is why I'm
interested, because your movie

1035
01:03:41,360 --> 01:03:46,520
background, like does it, does
it matter to the robots and to

1036
01:03:46,520 --> 01:03:52,240
the autonomous vehicles?
Actually, I do think that at

1037
01:03:52,240 --> 01:03:56,280
least for robots, the physics
play a big role because

1038
01:03:56,680 --> 01:04:00,640
ultimately the, the action that
you're generating is, is the

1039
01:04:00,640 --> 01:04:03,160
control right of your, of the
actuators of the actual things

1040
01:04:03,160 --> 01:04:07,920
that the robot should do.
And you need to estimate surface

1041
01:04:07,920 --> 01:04:11,800
properties, the weight of an
object, weight distribution

1042
01:04:11,800 --> 01:04:16,680
stability and all these things
to at least in a, in a variable

1043
01:04:16,680 --> 01:04:19,440
environment to interact with
the, with the world.

1044
01:04:19,440 --> 01:04:24,720
So think for robots by now rigid
body type or in some somewhat

1045
01:04:24,720 --> 01:04:29,840
deformable and physics are
playing quite a role.

1046
01:04:29,960 --> 01:04:35,120
So I would I would guess that it
does make a difference for

1047
01:04:35,120 --> 01:04:37,840
autonomous driving.
We could argue at the end you

1048
01:04:37,880 --> 01:04:40,360
only need to take right?
Is it a dangerous situation or

1049
01:04:40,360 --> 01:04:42,720
would something have gone wrong?
You can probably do a lot

1050
01:04:44,760 --> 01:04:47,640
without having to go through the
whole crash or actually

1051
01:04:47,640 --> 01:04:50,800
simulating how the car would
tumble or crash into a building

1052
01:04:50,800 --> 01:04:52,440
or so.
That might not be so important,

1053
01:04:52,440 --> 01:04:54,840
right, As long as you can detect
this would have gone wrong.

1054
01:04:56,920 --> 01:04:59,680
But for robots actually having
to really interact with objects,

1055
01:04:59,680 --> 01:05:01,200
I can imagine that plays a
larger role.

1056
01:05:01,200 --> 01:05:06,760
And then again looking towards
fluids, I, I think the challenge

1057
01:05:06,760 --> 01:05:10,920
also in, in the current
development pipelines is this

1058
01:05:10,920 --> 01:05:14,480
interaction with the real world.
I mean, I think in the end, the

1059
01:05:14,480 --> 01:05:16,720
current validation and
certification still goes through

1060
01:05:16,720 --> 01:05:19,800
a lot of testing status and then
you have all kinds of real world

1061
01:05:20,320 --> 01:05:24,000
flights and and experience to
make sure that the systems

1062
01:05:24,000 --> 01:05:25,840
really do what they should in
the real world.

1063
01:05:26,600 --> 01:05:32,800
But if you could shift some of
that into an earlier stage, I

1064
01:05:32,800 --> 01:05:35,840
think that that could also be
definitely interesting.

1065
01:05:35,840 --> 01:05:39,800
If right, you could have
different environment

1066
01:05:39,800 --> 01:05:42,760
conditions, different flow
conditions, and you could have

1067
01:05:42,760 --> 01:05:46,280
an have an object interact
dynamically in in such an

1068
01:05:46,280 --> 01:05:48,880
environment.
Yeah, I guess maybe maybe you're

1069
01:05:48,880 --> 01:05:52,080
right to maybe the examples I've
looked at, you know, if you look

1070
01:05:52,080 --> 01:05:58,960
at a robot walking around a
factory floor or it has no the

1071
01:05:59,000 --> 01:06:03,000
the drag, the lift, the thermal
properties are very second or

1072
01:06:03,000 --> 01:06:05,680
third order effects.
They don't really influence how

1073
01:06:05,680 --> 01:06:08,160
it's performing.
But I guess you're right.

1074
01:06:08,160 --> 01:06:13,440
If it was like a drone in the
sky or a boat in the water, then

1075
01:06:13,440 --> 01:06:17,640
the physics actually do make
quite a bit of a difference in

1076
01:06:18,720 --> 01:06:20,960
and then I guess it really is
depending what's the use of

1077
01:06:20,960 --> 01:06:26,600
these world moles, if they
really are like a true, true

1078
01:06:26,600 --> 01:06:33,280
representation of the world for
that drone or or helicopter or,

1079
01:06:34,000 --> 01:06:36,760
you know, I guess the moment the
autonomous vehicle, you're

1080
01:06:36,760 --> 01:06:40,480
right, is it's more sensory.
You know, like I see there's a

1081
01:06:40,480 --> 01:06:43,560
person walking.
It's not to design the car,

1082
01:06:43,560 --> 01:06:46,680
right?
It's not to say I see now that

1083
01:06:46,680 --> 01:06:49,800
in this real world, by changing
the shape, the drag is lower

1084
01:06:50,000 --> 01:06:53,120
because it's actually I guess
maybe that's where you're

1085
01:06:53,120 --> 01:06:57,760
thinking, right, that the world
model could be used as a true

1086
01:06:57,760 --> 01:07:01,920
synthetic environment or cars, I
guess if you could.

1087
01:07:01,920 --> 01:07:05,320
Simulate a rainy and stormy
situation and then really get

1088
01:07:05,320 --> 01:07:09,000
feedback on how the rain would
splash around a car, or how gust

1089
01:07:09,000 --> 01:07:14,160
of wind would influence driving
stability in extreme conditions.

1090
01:07:14,160 --> 01:07:18,600
Or right, all kinds of of
landing and starting scenarios

1091
01:07:18,600 --> 01:07:20,360
for planes.
If you could really get feedback

1092
01:07:20,360 --> 01:07:24,200
on how all the forming plane
with straining conditions would

1093
01:07:24,640 --> 01:07:28,680
would behave.
I think that's that's still

1094
01:07:29,080 --> 01:07:34,480
beyond even regular simulation
capabilities that full structure

1095
01:07:34,480 --> 01:07:38,520
interactions was changing
conditions are are really

1096
01:07:38,520 --> 01:07:42,760
challenging.
You could get some feedback on

1097
01:07:42,760 --> 01:07:45,880
these and then ideally put an
optimization loop around them,

1098
01:07:45,880 --> 01:07:47,840
right?
You you would want to actually

1099
01:07:47,840 --> 01:07:53,000
optimize your, you know, your
wing profile or the shape of of

1100
01:07:53,000 --> 01:07:57,480
a vehicle tour behavior in the
whole longer sequence.

1101
01:07:57,480 --> 01:08:00,720
I feel like this is also the big
debate, well not big debate, but

1102
01:08:00,720 --> 01:08:05,160
a debate around the classic tool
calling thing.

1103
01:08:05,160 --> 01:08:08,120
You know, like isn't LLM good at
adding numbers?

1104
01:08:08,960 --> 01:08:11,400
No, so just go and call a
calculator.

1105
01:08:11,840 --> 01:08:15,000
That's like ultimately, I guess
if you call Chachi BT to add 2

1106
01:08:15,000 --> 01:08:18,439
numbers together, it's just
calling a tool of a calculator,

1107
01:08:18,439 --> 01:08:22,160
adding the 2 numbers together.
I kind of wonder it's sometimes

1108
01:08:22,160 --> 01:08:27,200
in this world of, you know,
world models or optimizing, you

1109
01:08:27,200 --> 01:08:30,960
know, is it, is it really that
it's truly integrated into a

1110
01:08:30,960 --> 01:08:33,000
world model?
Or is it more likely that there

1111
01:08:33,000 --> 01:08:37,160
is a sort of agent workflow
where it's calling a surrogate

1112
01:08:37,160 --> 01:08:41,439
model and maybe that's the where
they're two separate models, if

1113
01:08:41,439 --> 01:08:43,399
you know what I mean, rather
than just a world model.

1114
01:08:43,399 --> 01:08:47,720
It's it's more than a gentic
way, if you know what I mean.

1115
01:08:47,880 --> 01:08:50,279
Or not that that would already
be a big step, right, If if you

1116
01:08:50,359 --> 01:08:54,760
had any visual model that could
on demand call some physics sub

1117
01:08:54,760 --> 01:08:58,920
models surrogates to to get
feedback on how these different

1118
01:08:58,920 --> 01:09:03,479
pieces is recognised.
Should should interact of

1119
01:09:03,479 --> 01:09:05,319
original bodies.
Actually these days you would

1120
01:09:05,319 --> 01:09:06,800
just call it original body
simulator.

1121
01:09:07,160 --> 01:09:09,359
You would probably not need a
surrogate.

1122
01:09:10,680 --> 01:09:14,680
Yeah, that's.
That's kind of where I always

1123
01:09:14,680 --> 01:09:16,640
debate it and that and that's
the traditional one that you

1124
01:09:16,640 --> 01:09:21,000
said if the solvers get so fast,
like the spectral solver you

1125
01:09:21,000 --> 01:09:24,439
mentioned or now with, you know,
advancements in GPUs, some

1126
01:09:24,439 --> 01:09:27,000
people are saying I can do a
simulation in a minute.

1127
01:09:27,880 --> 01:09:31,279
And then you think, well, do I
need a surrogate?

1128
01:09:31,359 --> 01:09:34,720
If if I can, if the solver can
run so fast?

1129
01:09:35,680 --> 01:09:37,560
Right, I think for World War,
it's an interesting challenge

1130
01:09:37,560 --> 01:09:41,920
will be to to quickly switch
from from the proximate fidelity

1131
01:09:41,920 --> 01:09:45,319
kind of for for rigid bodies or
the former objects of winds or

1132
01:09:45,319 --> 01:09:49,080
different feedbacks needed in
the environment.

1133
01:09:51,279 --> 01:09:55,000
Sure, if I know I need perfect
rigid bodies for a certain case

1134
01:09:55,000 --> 01:09:59,000
that might be ideal, but but
right I need some feedback on on

1135
01:09:59,000 --> 01:10:03,960
wind or how right my plastic cup
should deform or should it break

1136
01:10:03,960 --> 01:10:08,360
predictions across the all these
different physical phenomena.

1137
01:10:09,320 --> 01:10:13,080
I could imagine that a flexible
surrogate that at least gives a

1138
01:10:13,080 --> 01:10:16,440
gives the first prediction of
estimates of of what what might

1139
01:10:16,440 --> 01:10:18,440
happen there could be
beneficial.

1140
01:10:18,440 --> 01:10:21,880
I think that would be tricky to
do with classical simulators.

1141
01:10:22,400 --> 01:10:26,680
Yes, yeah, yeah, yeah, that.
I think this is where it's the

1142
01:10:26,680 --> 01:10:29,280
classic what's the, what's the
use case of real time?

1143
01:10:29,280 --> 01:10:32,560
And sometimes if you go to an
engineering team and I've had

1144
01:10:32,560 --> 01:10:35,200
this where you tell them that a
surrogate model prediction in

1145
01:10:35,200 --> 01:10:37,960
one second, sometimes they'll
say, well, we don't need it to

1146
01:10:37,960 --> 01:10:40,480
do in one second.
You know, it's actually fine If

1147
01:10:40,480 --> 01:10:44,160
it takes 5 minutes or even half
an hour, it's OK.

1148
01:10:44,160 --> 01:10:47,480
It's not a bottleneck.
Whereas I guess if it's

1149
01:10:47,480 --> 01:10:50,840
integrated in a world model or
something, you need it to be

1150
01:10:51,240 --> 01:10:57,480
essentially real time or else
the whole value breaks of some

1151
01:10:57,480 --> 01:11:00,080
simulation.
You know, like testing a car

1152
01:11:00,080 --> 01:11:01,600
moving.
If you have to sort of pause the

1153
01:11:01,600 --> 01:11:06,080
simulator for 30 minutes, then
it's not that.

1154
01:11:06,440 --> 01:11:09,560
Feel right?
Yeah, I mean real time is is

1155
01:11:09,560 --> 01:11:13,280
extremely challenging.
Then you need I guess to to be a

1156
01:11:13,280 --> 01:11:16,360
realistic virtual environment
for a person.

1157
01:11:16,840 --> 01:11:20,560
Yes, you need milliseconds.
Yeah, very real.

1158
01:11:21,080 --> 01:11:26,440
Yeah.
So maybe a final question to you

1159
01:11:26,880 --> 01:11:30,960
looking more if, if there's a
student listening to this or

1160
01:11:31,320 --> 01:11:34,440
someone doing HD or early in
their career, given, given all

1161
01:11:34,440 --> 01:11:39,200
these changes, what would you
recommend someone was would

1162
01:11:39,200 --> 01:11:41,480
study as a PhD topic?
You know what was going to

1163
01:11:41,480 --> 01:11:46,600
future proof them in this world.
Maybe I'm biased here, but I

1164
01:11:46,600 --> 01:11:48,800
think these are really important
techniques.

1165
01:11:48,800 --> 01:11:52,320
I, I do see a lot of potential
for AI based techniques.

1166
01:11:52,360 --> 01:11:54,400
So I think it's important to
have an understanding to know

1167
01:11:54,400 --> 01:11:56,760
the classic, the classic physics
numerics.

1168
01:11:56,920 --> 01:12:00,920
But by now within these
numerical tools, I think AI is,

1169
01:12:02,000 --> 01:12:04,960
is a super important component.
So I could, could highly

1170
01:12:04,960 --> 01:12:09,920
recommend looking at at least
combinations of classic and AI

1171
01:12:09,920 --> 01:12:15,880
based methods.
And my, my guess is also that in

1172
01:12:15,880 --> 01:12:21,560
the future or the next couple of
of years, we were not going to

1173
01:12:21,560 --> 01:12:24,240
be completely replaced by I
don't think chap DPT is going to

1174
01:12:24,240 --> 01:12:26,840
take over fluid simulation too
soon.

1175
01:12:26,840 --> 01:12:29,080
So having experts that we
understand what's happening

1176
01:12:29,080 --> 01:12:31,960
there, I think it's it's going
to be needed for quite a while.

1177
01:12:31,960 --> 01:12:34,720
So I think it's still a good
field to to work on.

1178
01:12:35,520 --> 01:12:39,760
Yeah, I guess that that's the
the truth, I guess the, the

1179
01:12:39,760 --> 01:12:43,000
short, medium, long term, I
guess I see at least now the

1180
01:12:43,000 --> 01:12:45,800
short term, there's even more
need for specialists because

1181
01:12:46,120 --> 01:12:52,440
frankly, it's almost become that
simulation is more of a popular

1182
01:12:52,440 --> 01:12:53,960
topic now.
The fact that all these startups

1183
01:12:53,960 --> 01:12:56,440
are getting funded, the fact
that, you know, there's more

1184
01:12:56,440 --> 01:13:00,360
money in this space, there's
more need to find out who's an

1185
01:13:00,360 --> 01:13:03,760
expert in crash simulations,
who's an expert in acoustics,

1186
01:13:03,760 --> 01:13:08,160
who's an expert to help computer
scientists, you know, to develop

1187
01:13:08,280 --> 01:13:11,480
these models.
I guess it's only in the very

1188
01:13:11,480 --> 01:13:13,600
long term future that you could
imagine.

1189
01:13:13,600 --> 01:13:18,120
Maybe you don't need as many
specialists because so much of

1190
01:13:18,120 --> 01:13:20,520
that knowledge is now ingrained
in these models.

1191
01:13:21,080 --> 01:13:24,200
But there's a, in the short
term, it's almost even more

1192
01:13:24,360 --> 01:13:29,480
specialists to help.
But deep, deep specialists, I

1193
01:13:29,480 --> 01:13:34,360
guess you really understand it
that, that at least that's the

1194
01:13:34,360 --> 01:13:36,480
way I I'm seeing things at the
moment.

1195
01:13:36,480 --> 01:13:38,920
And I think that's usually a
great topic for a PhD, really

1196
01:13:38,920 --> 01:13:41,320
getting to the bottom of things.
So I just, it's a nice

1197
01:13:41,320 --> 01:13:45,520
opportunity to work on one topic
for a couple of years and yeah,

1198
01:13:45,840 --> 01:13:47,480
get as much understanding as
possible.

1199
01:13:48,320 --> 01:13:49,800
So on.
Right now I think PhD is in this

1200
01:13:49,800 --> 01:13:51,080
area.
Still a very good idea.

1201
01:13:51,680 --> 01:13:53,920
Yeah, yeah.
No, no, I agree.

1202
01:13:54,280 --> 01:13:56,960
Well, thank you so much for
taking the time to speak.

1203
01:13:57,080 --> 01:14:00,320
I mean, I, I find all these
topics fascinating and one of

1204
01:14:00,320 --> 01:14:04,840
the things I'll do is put a list
of the papers in the, in the

1205
01:14:04,840 --> 01:14:08,200
show notes because I think
actually there's a lot more to

1206
01:14:08,200 --> 01:14:10,800
gather from looking in the
details and, and looking at the

1207
01:14:10,800 --> 01:14:12,880
codes that your, your, your
group did.

1208
01:14:12,880 --> 01:14:15,240
So I know we didn't have time to
cover all in, you know, full

1209
01:14:15,240 --> 01:14:18,000
detail, but I'll, I'll put a
list and hopefully people can

1210
01:14:18,000 --> 01:14:20,240
find the time to read through
the papers.

1211
01:14:20,360 --> 01:14:22,440
That's a good idea.
Oh, great discussion.

1212
01:14:22,560 --> 01:14:25,440
Thanks for the invite.
Yeah, interesting.

1213
01:14:25,720 --> 01:14:27,680
We'll speak again in two years
time and we'll see if the

1214
01:14:27,680 --> 01:14:31,560
predictions are.
That's interesting.

1215
01:14:32,160 --> 01:14:33,720
Thank you.
Thanks.
